all the models — AI benchmark observatory
← Models

Claude Opus 4.1

Anthropic · proprietary · claude-opus-4-1-20250805

Claude Opus 4.1 is a hybrid reasoning model that pushes the frontier for coding and AI agents, featuring a 200K context window. It delivers superior performance and precision for real-world coding and agentic tasks, handling complex multi-step problems with rigor and attention to detail. With extended thinking capabilities, it offers instant responses or extended step-by-step thinking visible through user-friendly summaries. It advances state-of-the-art coding performance to 74.5% on SWE-bench Verified, excels at agentic search and research, and produces human-quality content with exceptional writing abilities. It supports 32K output tokens and adapts to specific coding styles while delivering exceptional quality for extensive generation and refactoring projects.

MMMLUGPQAAIME 2025SWE-Bench Ve

Benchmark scores

BenchmarkScore
MMMLU0.90
GPQA0.81
AIME 20250.78
SWE-Bench Verified0.74

Pricing

  • No provider pricing.

AA metrics

No Artificial Analysis link yet.

Arena Elo

  • No Arena snapshot linked.