all the models — AI benchmark observatory
← Benchmarks

Terminal-Bench

Terminal-Bench is a benchmark for testing AI agents in real terminal environments. It evaluates how well agents can handle real-world, end-to-end tasks autonomously, including compiling code, training models, setting up servers, system administration, security tasks, data science workflows, and cybersecurity vulnerabilities. The benchmark consists of a dataset of ~100 hand-crafted, human-verified tasks and an execution harness that connects language models to a terminal sandbox.

id terminal-bench · max 1 · 25 models reported

#ModelScore

No scores for this benchmark yet.

Terminal-Bench Leaderboard · all the models