all the models — AI benchmark observatory
← Benchmarks

SWT-Bench

Software Test Benchmark evaluating LLM ability to write tests for software repositories

id swt-bench · max 1 · 1 models reported

#ModelScore

No scores for this benchmark yet.