TailRL Trains Policies to Chase Rare High-Reward Rollouts, Not Just the Mean
The paper argues that optimizing average reward in RL can obscure differences between policies that share the same mean but differ in their likelihood of producing rare, high-reward outcomes—a gap that matters as sampling scales at training and inference time. The authors propose Tail-Likelihood RL (TailRL), which maximizes the log-probability of exceeding a randomly chosen reward threshold, effectively turning reward into a family of binary success events across thresholds. This requires only a small change to the advantage function, and the gradient can be interpreted as a mixture of Best-of-k gradients, letting models benefit more from inference-time sampling without fixing k in advance. Experiments span object localization, maze navigation, GUI grounding, and code optimization, showing TailRL avoids suboptimal solutions by exploiting rare high-reward samples. On Twitter, the authors and collaborators framed TailRL as a drop-in modification to existing RL pipelines that aligns training more closely with Best-of-k style inference, generating interest for its simplicity and broad applicability across domains rather than sparking notable public pushback.
Discussion: 2 tweets from 2 authors · @rsalakhu, @di_zhang_fdu