all the models — AI benchmark observatory
← Benchmarks

NL2Repo

NL2Repo evaluates long-horizon coding capabilities including repository-level understanding, where models must generate or modify code across entire repositories from natural language specifications.

id nl2repo · max 1 · 23 models reported

#ModelScore

No scores for this benchmark yet.

NL2Repo Leaderboard · all the models