all the models — AI benchmark observatory
← Benchmarks

SWE-Dev

SWE-bench development split consisting of 225 software engineering problems drawn from real GitHub issues across 12 popular Python repositories. Language models are given a codebase along with a description of an issue to be resolved and must edit the codebase to address the issue, often requiring understanding and coordinating changes across multiple functions, classes, and files.

id swe-dev · max 1 · 0 models reported

#ModelScore

No scores for this benchmark yet.