all the models — AI benchmark observatory
← Benchmarks

Wild Bench

WildBench is an automated evaluation framework that benchmarks large language models using 1,024 challenging, real-world tasks selected from over one million human-chatbot conversation logs. It introduces two evaluation metrics (WB-Reward and WB-Score) that achieve high correlation with human preferences and uses task-specific checklists for systematic evaluation.

id wild-bench · max 1 · 8 models reported

#ModelScore

No scores for this benchmark yet.