← Benchmarks
ExploitBench
ExploitBench is a cybersecurity benchmark that evaluates a model's ability to discover and exploit software vulnerabilities, reported as the fraction of challenges where the model captures the target (Cap%).
id exploitbench · max 1 · 8 models reported
| # | Model | Score |
|---|---|---|
| 1 | GPT-6 Astra OpenAI | 1.00 |
| 2 | Claude Fable 5 Anthropic | 0.78 |
| 3 | GPT-5.6 Sol OpenAI | 0.73 |
