all the models — AI benchmark observatory
← Benchmarks

Terminal-Bench Hard

Terminal-Bench Hard is a harder terminal-agent benchmark variant evaluated with the Terminus-2 harness in Cohere's Command A+ and North Mini Code releases.

id terminal-bench-hard · max 1 · 2 models reported

#ModelScore

No scores for this benchmark yet.