all the models — AI benchmark observatory
← Benchmarks

MLS-Bench Lite

MLS-Bench Lite is the official 30-task subset of MLS-Bench for evaluating whether AI systems can invent generalizable and scalable machine learning methods across LLM pretraining and post-training, robotics, world models, computer vision, reinforcement learning, optimization, ML systems, and AI for Science.

id mls-bench-lite · max 1 · 3 models reported

#ModelScore

No scores for this benchmark yet.