OpenAI's T2MLR Adds Latent Recurrence to Transformers for Reasoning
The paper introduces T2MLR, a transformer variant that caches a middle-layer representation from the previous token and feeds it into an earlier layer at the current position, letting abstract intermediate computation persist across decoding steps with minimal inference overhead. The authors report consistent gains over parameter- and data-matched baselines on language pretraining and multi-hop reasoning, find that recurring only a small (~20%) middle-layer block often beats full-layer recurrence, and show the recurrent pathway can be retrofitted into an existing pretrained 1.7B model to improve math reasoning without retraining from scratch. On Twitter, researchers noted that the core idea—persisting latent state across token positions via layer recurrence—has several concurrent precedents outside Microsoft Research, including work from Sanjeev Arora's group and others, though OpenAI's version is distinguished by being pretrained from scratch with greater compute resources.
- T^2MLR: Transformer with Temporal Middle-Layer Recurrence
- T^2MLR: Transformer with Temporal Middle-Layer Recurrence
- T^2MLR: Transformer with Temporal Middle-Layer Recurrence
Discussion: 2 tweets from 2 authors · @prfsanjeevarora, @xidulu