all the models — AI benchmark observatory
← Benchmarks

LiveBench 20241125

LiveBench is a challenging, contamination-limited LLM benchmark that addresses test set contamination by releasing new questions monthly based on recently-released datasets, arXiv papers, news articles, and IMDb movie synopses. It comprises tasks across math, coding, reasoning, language, instruction following, and data analysis with verifiable, objective ground-truth answers.

id livebench-20241125 · max 1 · 15 models reported

#ModelScore

No scores for this benchmark yet.

LiveBench 20241125 Leaderboard · all the models