all the models — AI benchmark observatory
← Benchmarks

Kimi Claw 24/7 Bench

Kimi Claw 24/7 Bench is Moonshot AI's in-house benchmark for evaluating long-horizon agentic performance in persistent, multi-day coworking tasks. It spans 17 professional scenarios across 610 evaluation points, covering software engineering, ML research, recruiting, trading, and marketing tasks executed through the OpenClaw harness.

id kimi-claw-24-7-bench · max 1 · 1 models reported

#ModelScore

No scores for this benchmark yet.

Kimi Claw 24/7 Bench Leaderboard · all the models