all the models — AI benchmark observatory
← Benchmarks

PostTrainBench Lite

PostTrainBench Lite measures whether an agent can design and execute a full post-training strategy (data, prompts, RL recipe, and eval loop) for a pretrained base model under a constrained time budget, scored as normalized mean reward over the improvement window.

id posttrainbench-lite · max 1 · 3 models reported

#ModelScore

No scores for this benchmark yet.

PostTrainBench Lite Leaderboard · all the models