all the models — AI benchmark observatory
← Benchmarks

HLE-Verified

HLE-Verified evaluates multidisciplinary expert reasoning on a verified subset of Humanity's Last Exam.

id hle-verified · max 1 · 3 models reported

#ModelScore

No scores for this benchmark yet.