all the models — AI benchmark observatory
← Benchmarks

FrontierSWE V2

FrontierSWE V2 evaluates agents on open-ended technical projects and reports mean@5. Kept separate from FrontierSWE V1 dominance-score results, which are not comparable.

id frontierswe-v2 · max 1 · 2 models reported

#ModelScore

No scores for this benchmark yet.