← Benchmarks
HealthBench
An open-source benchmark for measuring performance and safety of large language models in healthcare, consisting of 5,000 multi-turn conversations evaluated by 262 physicians using 48,562 unique rubric criteria across health contexts and behavioral dimensions
id healthbench · max 1 · 11 models reported
| # | Model | Score |
|---|
No scores for this benchmark yet.
