all the models — AI benchmark observatory
← Models

Claude 3.5 Sonnet

Anthropic · proprietary · claude-3-5-sonnet-20241022

Claude 3.5 Sonnet is a powerful AI model with industry-leading software engineering skills. It excels in coding, planning, and problem-solving, with significant improvements in agentic coding and tool use tasks. The model includes computer use capabilities in public beta, allowing it to interact with computer interfaces like a human user.

GSM8kDocVQAAI2DHumanEvalBIG-Bench HaMGSM

Benchmark scores

BenchmarkScore
GSM8k0.96
DocVQA0.95
AI2D0.95
HumanEval0.94
BIG-Bench Hard0.93
MGSM0.92
ChartQA0.91
MMLU0.90
MATH0.78
MMLU-Pro0.78
MMMU0.68
GPQA0.67
SWE-Bench Verified0.49

Pricing

  • No provider pricing.

AA metrics

No Artificial Analysis link yet.

Arena Elo

  • No Arena snapshot linked.
Claude 3.5 Sonnet Benchmarks · all the models