all the models — AI benchmark observatory
← Benchmarks

PostTrainBench

PostTrainBench evaluates a model's ability to autonomously post-train base models. Given pretrain-only base models, the agent must complete the full pipeline of data synthesis, training, evaluation, and iteration within a time budget, scored across downstream benchmarks such as AIME2025, BFCL, GPQA Main, GSM8K, and HumanEval.

id posttrainbench · max 1 · 7 models reported

#ModelScore

No scores for this benchmark yet.

PostTrainBench Leaderboard · all the models