← Benchmarks
SEC-bench Pro
SEC-bench Pro is a self-evolving software-security benchmark that measures agent bug hunting on critical, high-complexity systems. It instantiates validated vulnerabilities across the V8 and SpiderMonkey JavaScript engines as reproducible vulnerability-discovery and proof-of-concept-generation tasks with oracle-based validation.
id sec-bench-pro · max 1 · 7 models reported
| # | Model | Score |
|---|---|---|
| 1 | GPT-6 Astra OpenAI | 0.85 |
| 2 | GPT-5.6 Sol OpenAI | 0.71 |
