all the models — AI benchmark observatory
← Benchmarks

HealthBench

An open-source benchmark for measuring performance and safety of large language models in healthcare, consisting of 5,000 multi-turn conversations evaluated by 262 physicians using 48,562 unique rubric criteria across health contexts and behavioral dimensions

id healthbench · max 1 · 11 models reported

#ModelScore

No scores for this benchmark yet.

HealthBench Leaderboard · all the models