all the models — AI benchmark observatory
← Benchmarks

HellaSwag

A challenging commonsense natural language inference dataset that uses Adversarial Filtering to create questions trivial for humans (>95% accuracy) but difficult for state-of-the-art models, requiring completion of sentence endings based on physical situations and everyday activities

id hellaswag · max 1 · 28 models reported

#ModelScore
1Claude 3 Opus
Anthropic
0.95
2GPT-4
OpenAI
0.95
3Gemini 1.5 Pro
Google
0.93