Inkling-Small
Thinking Machines Lab · open weight · inkling-small
Inkling-Small is Thinking Machines Lab's efficient open-weights MoE multimodal model (276B total / 12B active parameters) released under Apache 2.0. It accepts text, image, and audio inputs and generates text, with native reasoning, variable thinking effort, and a context window up to 1M tokens (Tinker exposes 64K and 256K configurations). Hugging Face weights: thinkingmachines/Inkling-Small and thinkingmachines/Inkling-Small-NVFP4. Vendor self-host VRAM: BF16 ≥ ~600 GB aggregated; NVFP4 ≥ ~180 GB aggregated. SWE-Bench Verified 80.2% vs Inkling 77.6% (same bash-only harness). Fine-tuning and playground chat are available via Tinker. Official Tinker serverless inference (256K, list): $0.30 / $1.20 per 1M input/output tokens ($0.06 cached).
Benchmark scores
| Benchmark | Score |
|---|---|
| GDPval-AA | 1269.00 |
| AA-Briefcase | 917.00 |
| AIME 2026 | 0.95 |
| GPQA | 0.90 |
| SWE-Bench Verified | 0.80 |
| MCP Atlas | 0.80 |
| BrowseComp | 0.77 |
| MMMU-Pro | 0.74 |
| Humanity's Last Exam | 0.32 |
Pricing
- Thinking Machines Lab$0.30 / $1.20
- DeepInfra$0.45 / $1.20
Input / output per 1M tokens
AA metrics
No Artificial Analysis link yet.
Arena Elo
- No Arena snapshot linked.
