← Benchmarks
Terminal-Bench 2.0
Terminal-Bench 2.0 is an updated benchmark for testing AI agents' tool use ability to operate a computer via terminal. It evaluates how well models can handle real-world, end-to-end tasks autonomously, including compiling code, training models, setting up servers, system administration, security tasks, data science workflows, and cybersecurity vulnerabilities.
id terminal-bench-2 · max 1 · 53 models reported
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.5 OpenAI | 0.83 |
| 2 | Claude Mythos Preview Anthropic | 0.82 |
| 3 | Claude Sonnet 5 Anthropic | 0.80 |
| 4 | GPT-5.3 Codex OpenAI | 0.77 |
| 5 | Gemini 3.5 Flash Google | 0.76 |
| 6 | GPT-5.4 OpenAI | 0.75 |
| 7 | Claude Opus 4.8 Anthropic | 0.75 |
| 8 | Qwen3.7-Plus Alibaba Cloud / Qwen Team | 0.70 |
