all the models — AI benchmark observatory
← Benchmarks

CursorBench 4.0

CursorBench 4.0 evaluates coding agents on long-horizon coding tasks from real Cursor sessions. Version 4.0 adds longer-horizon tasks; scores are not comparable with CursorBench 3.2.

id cursorbench-4.0 · max 1 · 2 models reported

#ModelScore

No scores for this benchmark yet.