all the models — AI benchmark observatory
← Benchmarks

Terminal-Bench 4.0

Terminal-Bench 4.0 evaluates AI agents on real-world work in terminal and command-line containerized environments across 66 tasks, with emphasis on science-adjacent and frontier engineering problems. Relative to earlier Terminal-Bench releases, 4.0 increases timeouts and adaptively raises RAM/CPU on selected tasks to reduce harness and resource confounds.

id terminal-bench-4.0 · max 1 · 19 models reported

#ModelScore

No scores for this benchmark yet.

Terminal-Bench 4.0 Leaderboard · all the models