all the models — AI benchmark observatory
← Benchmarks

Agents' Last Exam

Agents' Last Exam is a challenging benchmark for AI agents on hard, long-horizon tasks that test sustained reasoning, planning, and tool use, reported with and without tool access.

id agents-last-exam · max 1 · 21 models reported

#ModelScore

No scores for this benchmark yet.