AI Research Agents Push State of the Art on MLE-bench via Search Strategy Design
The paper formalizes AI research agents as search policies navigating a space of candidate ML solutions, applying operators to iteratively modify them. By systematically testing combinations of operator sets and search policies (Greedy, MCTS, Evolutionary) on MLE-bench—a benchmark where agents solve real Kaggle competitions—the authors show that the interplay between search strategy and operator design is critical for performance. Their best configuration raises the Kaggle medal success rate on MLE-bench lite from 39.6% to 47.7%, a new state of the art. This is the third installment (AIRA₃) in a research line the authors are visibly excited about on Twitter, following AIRA₁ (a NeurIPS spotlight) and AIRA₂ (previous SOTA on MLE-bench and AIRS Bench). Commentary from the authors frames this version as reaching elite human-level performance on a live, uncontaminated benchmark, positioning it as a step toward recursive self-improvement (RSI) in AI research automation, though such framing comes from the paper's own contributors rather than independent evaluation.
Discussion: 2 tweets from 2 authors · @EdanToledo, @RishiHazra95