all the models — AI benchmark observatory
← Benchmarks

LMArena Text Leaderboard

LMArena Text Leaderboard is a blind human preference evaluation benchmark that ranks models based on pairwise comparisons in real-world conversations. The leaderboard uses Elo ratings computed from user preferences in head-to-head model battles, providing a comprehensive measure of overall model capability and style.

id lmarena-text · max 2000 · 2 models reported

#ModelScore
1Grok-4.1 Thinking
xAI
1483.00
2Grok-4.1
xAI
1465.00