all the models — AI benchmark observatory
← Models

Claude Opus 4.6

Anthropic · proprietary · claude-opus-4-6

Claude Opus 4.6 is Anthropic's most intelligent model, improving on its predecessor's coding skills with more careful planning, longer agentic task sustenance, more reliable operation in larger codebases, and better code review and debugging skills. First Opus-class model with 1M token context window (beta), 128K output tokens, and adaptive thinking. Features effort controls (low/medium/high/max) and context compaction for long-running tasks. State-of-the-art on Terminal-Bench 2.0, Humanity's Last Exam, GDPval-AA, and BrowseComp. Pricing: $5/$25 per million tokens (input/output).

Vending-BencAIME 2025Tau2 TelecomGraphwalks pDeepSearchQAGPQA

Benchmark scores

Pricing

  • Anthropic$5.00 / $25.00

Input / output per 1M tokens

AA metrics

No Artificial Analysis link yet.

Arena Elo

  • text1497
  • vision1293
  • code1537
Claude Opus 4.6 Benchmarks · all the models