all the models — AI benchmark observatory
← Benchmarks

Terminal-Bench 3.0

Terminal-Bench 3.0 is a release of the Terminal-Bench benchmark that tests AI agents' ability to operate a computer via the terminal on real-world, end-to-end tasks.

id terminal-bench-3.0 · max 1 · 4 models reported

#ModelScore

No scores for this benchmark yet.