all the models — AI benchmark observatory
← Models

LongCat-Flash-Lite

Meituan · open weight · longcat-flash-lite

LongCat-Flash-Lite is a lightweight MoE model from Meituan with 68.5B total parameters and only 2.9B-4.5B activated per token. It explores N-gram embedding expansion as a new scaling direction, supporting 256K context length via YaRN. Optimized for agent tooling and programming tasks, achieving 500-700 tokens per second inference speed while maintaining strong performance on coding, math, and agentic benchmarks.

MATH-500MMLUMMLU-ProAIME 2024GPQAAIME 2025

Benchmark scores

BenchmarkScore
MATH-5000.97
MMLU0.86
MMLU-Pro0.78
AIME 20240.72
GPQA0.67
AIME 20250.63
SWE-Bench Verified0.54

Pricing

  • Meituan$0.10 / $0.40

Input / output per 1M tokens

AA metrics

No Artificial Analysis link yet.

Arena Elo

  • No Arena snapshot linked.