all the models — AI benchmark observatory
← Benchmarks

CritPT

CritPT is a challenging reasoning benchmark reported by Qwen for evaluating frontier mathematical and critical problem-solving capability.

id critpt · max 1 · 6 models reported

#ModelScore

No scores for this benchmark yet.

CritPT Leaderboard · all the models