all the models — AI benchmark observatory
← Benchmarks

Kernel Bench L3

Kernel Bench L3 evaluates agentic GPU kernel optimization across 50 problems. Qwen reports two metrics for this benchmark: median per-problem speedup over the PyTorch eager reference and the fraction of problems faster than torch.compile.

id kernel-bench-l3 · max 1 · 1 models reported

#ModelScore
1Qwen3.7 Max
Alibaba Cloud / Qwen Team
0.96