← Benchmarks
HiddenMath
Google DeepMind's internal mathematical reasoning benchmark that introduces novel problems not encountered during model training to evaluate true mathematical reasoning capabilities rather than memorization
id hiddenmath · max 1 · 13 models reported
| # | Model | Score |
|---|
No scores for this benchmark yet.
