all the models — AI benchmark observatory
← Benchmarks

Terminal-Bench-Science 0.1

Terminal-Bench-Science 0.1 is a Stanford-led community benchmark of 70 tasks drawn from scientific research workflows across the life, physical, earth, mathematical, and engineering sciences. Tasks are authored and reviewed by scientists and researchers; scores are typically reported as accuracy with relatively large standard error due to strongly bimodal per-task outcomes.

id terminal-bench-science-0.1 · max 1 · 3 models reported

#ModelScore

No scores for this benchmark yet.