all the models — AI benchmark observatory
← Models

Mercury 2

Inception · proprietary · mercury-2

Mercury 2 is the fastest reasoning LLM, built on diffusion-based language model (dLLM) architecture. Instead of generating text token-by-token, it refines multiple text blocks simultaneously, achieving over 1,000 tokens per second on Nvidia Blackwell GPUs — 5x faster than leading speed-optimized LLMs. Supports tool usage and JSON output with 128K context window.

AIME 2025GPQALiveCodeBenc

Benchmark scores

BenchmarkScore
AIME 20250.91
GPQA0.74
LiveCodeBench0.67

Pricing

  • Inception$0.25 / $0.75

Input / output per 1M tokens

AA metrics

No Artificial Analysis link yet.

Arena Elo

  • No Arena snapshot linked.