all the models — AI benchmark observatory
← Benchmarks

CWE-Bench

CWE-Bench evaluates autonomous agents on finding and patching real-world software vulnerabilities.

id cwe-bench · max 1 · 1 models reported

#ModelScore

No scores for this benchmark yet.