← Models
Mercury 2
Inception · proprietary · mercury-2
Mercury 2 is the fastest reasoning LLM, built on diffusion-based language model (dLLM) architecture. Instead of generating text token-by-token, it refines multiple text blocks simultaneously, achieving over 1,000 tokens per second on Nvidia Blackwell GPUs — 5x faster than leading speed-optimized LLMs. Supports tool usage and JSON output with 128K context window.
Benchmark scores
| Benchmark | Score |
|---|---|
| AIME 2025 | 0.91 |
| GPQA | 0.74 |
| LiveCodeBench | 0.67 |
Pricing
- Inception$0.25 / $0.75
Input / output per 1M tokens
AA metrics
No Artificial Analysis link yet.
Arena Elo
- No Arena snapshot linked.
