← Benchmarks
Terminal-Bench 4.0
Terminal-Bench 4.0 evaluates AI agents on real-world work in terminal and command-line containerized environments across 66 tasks, with emphasis on science-adjacent and frontier engineering problems. Relative to earlier Terminal-Bench releases, 4.0 increases timeouts and adaptively raises RAM/CPU on selected tasks to reduce harness and resource confounds.
id terminal-bench-4.0 · max 1 · 19 models reported
| # | Model | Score |
|---|
No scores for this benchmark yet.
