← Benchmarks
CyberSecEval 4
CyberSecEval 4 is an evaluation suite covering cybersecurity-related capabilities and risks of large language models. The insecure-code-generation tracks measure whether a model produces vulnerable code: the Instruct track presents coding requests designed to elicit known insecure patterns, while the Autocomplete track prompts the model with code context leading up to a known insecure pattern, with vulnerabilities detected via static analysis.
id cyberseceval-4 · max 1 · 1 models reported
| # | Model | Score |
|---|
No scores for this benchmark yet.
