← Benchmarks
LMArena Text Leaderboard
LMArena Text Leaderboard is a blind human preference evaluation benchmark that ranks models based on pairwise comparisons in real-world conversations. The leaderboard uses Elo ratings computed from user preferences in head-to-head model battles, providing a comprehensive measure of overall model capability and style.
id lmarena-text · max 2000 · 2 models reported
| # | Model | Score |
|---|---|---|
| 1 | Grok-4.1 Thinking xAI | 1483.00 |
| 2 | Grok-4.1 xAI | 1465.00 |
