← Benchmarks
FrontierSWE V2
FrontierSWE V2 evaluates agents on open-ended technical projects and reports mean@5. Kept separate from FrontierSWE V1 dominance-score results, which are not comparable.
id frontierswe-v2 · max 1 · 2 models reported
| # | Model | Score |
|---|
No scores for this benchmark yet.
