all the models — AI benchmark observatory
← Benchmarks

FunctionalMATH

A functional variant of the MATH benchmark that tests language models' ability to generalize reasoning patterns across different problem instances, revealing the reasoning gap between static and functional performance.

id functionalmath · max 1 · 2 models reported

#ModelScore

No scores for this benchmark yet.

FunctionalMATH Leaderboard · all the models