all the models — AI benchmark observatory
← Benchmarks

Kimi Code Bench v2

Kimi Code Bench v2 is Moonshot AI's in-house benchmark for evaluating coding agents on realistic software engineering tasks across 10+ mainstream programming languages and a production tech stack spanning backend services, infrastructure, performance engineering, systems programming, security, frontend development, and ML/data engineering.

id kimi-code-bench-v2 · max 1 · 2 models reported

#ModelScore
1Kimi K3
Moonshot AI · open
0.73