← Benchmarks
Agents' Last Exam
Agents' Last Exam is a challenging benchmark for AI agents on hard, long-horizon tasks that test sustained reasoning, planning, and tool use, reported with and without tool access.
id agents-last-exam · max 1 · 21 models reported
| # | Model | Score |
|---|
No scores for this benchmark yet.
