all the models — AI benchmark observatory
← Benchmarks

ExploitBench

ExploitBench is a cybersecurity benchmark that evaluates a model's ability to discover and exploit software vulnerabilities, reported as the fraction of challenges where the model captures the target (Cap%).

id exploitbench · max 1 · 8 models reported

#ModelScore
1GPT-6 Astra
OpenAI
1.00
2Claude Fable 5
Anthropic
0.78
3GPT-5.6 Sol
OpenAI
0.73