Claude Mythos Preview
Anthropic · proprietary · claude-mythos-preview
Claude Mythos Preview is an unreleased general-purpose frontier model from Anthropic, a new tier above Opus (internal codename 'Capybara'). It identified thousands of zero-day vulnerabilities across every major operating system and web browser as part of Project Glasswing, a cross-industry cybersecurity initiative with 12 partners including AWS, Apple, Microsoft, and Google. State-of-the-art on SWE-bench Verified (93.9%), GPQA Diamond (94.6%), USAMO (97.6%), Terminal-Bench 2.0 (82.0%), CyberGym (83.1%), and Cybench (100% pass@1, saturated). Represents a 4.3x increase over the previous trendline for model performance. Deployed under ASL-3 Standard. Best-aligned Claude model to date per Anthropic's risk report, with the first-ever 24-hour internal alignment review before deployment. Not planned for general availability. Pricing for participants: $25/$125 per million tokens (input/output). 244-page system card.
Benchmark scores
| Benchmark | Score |
|---|---|
| CyBench | 1.00 |
| USAMO25 | 0.98 |
| GPQA | 0.95 |
| SWE-Bench Verified | 0.94 |
| CharXiv-R | 0.93 |
| MMMLU | 0.93 |
| FigQA | 0.89 |
| SWE-bench Multilingual | 0.87 |
| BrowseComp | 0.87 |
| CyberGym | 0.83 |
| Terminal-Bench 2.0 | 0.82 |
| OSWorld-Verified | 0.80 |
| SWE-Bench Pro | 0.78 |
| Humanity's Last Exam | 0.65 |
Pricing
- No provider pricing.
AA metrics
No Artificial Analysis link yet.
Arena Elo
- No Arena snapshot linked.
