all the models — AI benchmark observatory
← Benchmarks

SEC-bench Pro

SEC-bench Pro is a self-evolving software-security benchmark that measures agent bug hunting on critical, high-complexity systems. It instantiates validated vulnerabilities across the V8 and SpiderMonkey JavaScript engines as reproducible vulnerability-discovery and proof-of-concept-generation tasks with oracle-based validation.

id sec-bench-pro · max 1 · 7 models reported

#ModelScore
1GPT-6 Astra
OpenAI
0.85
2GPT-5.6 Sol
OpenAI
0.71
SEC-bench Pro Leaderboard · all the models