← Benchmarks
Wild Bench
WildBench is an automated evaluation framework that benchmarks large language models using 1,024 challenging, real-world tasks selected from over one million human-chatbot conversation logs. It introduces two evaluation metrics (WB-Reward and WB-Score) that achieve high correlation with human preferences and uses task-specific checklists for systematic evaluation.
id wild-bench · max 1 · 8 models reported
| # | Model | Score |
|---|
No scores for this benchmark yet.
