all the models — AI benchmark observatory
← Benchmarks

LongCodeBench

LongCodeBench evaluates the code understanding and comprehension abilities of large language models at very long context windows, scaling up to 1M tokens. It tests whether models can reason about extensive codebases provided in a single prompt by answering multiple-choice questions about the code.

id longcodebench · max 1 · 2 models reported

#ModelScore
1Nova 2 Lite
Amazon
0.84
2Nova 2 Pro
Amazon
0.84
LongCodeBench Leaderboard · all the models