all the models — AI benchmark observatory
← Benchmarks

ZebraLogic

ZebraLogic is an evaluation framework for assessing large language models' logical reasoning capabilities through logic grid puzzles derived from constraint satisfaction problems (CSPs). The benchmark consists of 1,000 programmatically generated puzzles with controllable and quantifiable complexity, revealing a 'curse of complexity' where model accuracy declines significantly as problem complexity grows.

id zebralogic · max 1 · 9 models reported

#ModelScore
1Qwen3 VL 235B A22B Thinking
Alibaba Cloud / Qwen Team · open
0.97
2LongCat-Flash-Thinking
Meituan · open
0.95
3Qwen3-235B-A22B-Instruct-2507
Alibaba Cloud / Qwen Team · open
0.95
ZebraLogic Leaderboard · all the models