all the models — AI benchmark observatory
← Benchmarks

HiddenMath

Google DeepMind's internal mathematical reasoning benchmark that introduces novel problems not encountered during model training to evaluate true mathematical reasoning capabilities rather than memorization

id hiddenmath · max 1 · 13 models reported

#ModelScore

No scores for this benchmark yet.