all the models — AI benchmark observatory
← Models

LongCat-Flash-Thinking-2601

Meituan · open weight · longcat-flash-thinking-2601

LongCat-Flash-Thinking-2601 is an upgraded version of LongCat-Flash-Thinking with 560B total parameters (MoE, ~27B activated). It achieves open-source SOTA performance on core evaluation benchmarks including Agentic Search, Agentic Tool Use, and Tool-Integrated Reasoning (TIR). Features Heavy Thinking mode that contributes +4-6 points on demanding agentic reasoning benchmarks. Mid-training with structured agentic trajectories improves pass@k by up to +12 points, and context management yields +17.5 improvement.

AIME 2025Tau2 TelecomLiveCodeBencGPQASWE-Bench VeHumanity's L

Benchmark scores

Pricing

  • No provider pricing.

AA metrics

No Artificial Analysis link yet.

Arena Elo

  • No Arena snapshot linked.