← Benchmarks
HellaSwag
A challenging commonsense natural language inference dataset that uses Adversarial Filtering to create questions trivial for humans (>95% accuracy) but difficult for state-of-the-art models, requiring completion of sentence endings based on physical situations and everyday activities
id hellaswag · max 1 · 28 models reported
| # | Model | Score |
|---|---|---|
| 1 | Claude 3 Opus Anthropic | 0.95 |
| 2 | GPT-4 OpenAI | 0.95 |
| 3 | Gemini 1.5 Pro Google | 0.93 |
