all the models — AI benchmark observatory
← Models

Inkling-Small

Thinking Machines Lab · open weight · inkling-small

Inkling-Small is Thinking Machines Lab's efficient open-weights MoE multimodal model (276B total / 12B active parameters) released under Apache 2.0. It accepts text, image, and audio inputs and generates text, with native reasoning, variable thinking effort, and a context window up to 1M tokens (Tinker exposes 64K and 256K configurations). Hugging Face weights: thinkingmachines/Inkling-Small and thinkingmachines/Inkling-Small-NVFP4. Vendor self-host VRAM: BF16 ≥ ~600 GB aggregated; NVFP4 ≥ ~180 GB aggregated. SWE-Bench Verified 80.2% vs Inkling 77.6% (same bash-only harness). Fine-tuning and playground chat are available via Tinker. Official Tinker serverless inference (256K, list): $0.30 / $1.20 per 1M input/output tokens ($0.06 cached).

GDPval-AAAA-BriefcaseAIME 2026GPQASWE-Bench VeMCP Atlas

Benchmark scores

Pricing

  • Thinking Machines Lab$0.30 / $1.20
  • DeepInfra$0.45 / $1.20

Input / output per 1M tokens

AA metrics

No Artificial Analysis link yet.

Arena Elo

  • No Arena snapshot linked.