Open Recipe Trains Qwen2.5-1.5B-Level LLM for Under $7K on Consumer GPUs
The authors present an open, fully-reproducible pretraining recipe for a 2B-parameter language model, Puro-2B, trained from scratch on up to 1.4 trillion tokens using FP8 precision on consumer RTX 5090 GPUs. Combining hardware selection, low-precision training, a 'hyperball' optimization approach, curriculum model averaging, and a tailored data recipe, their best model approaches Qwen2.5-1.5B performance at a training cost under $6.9K—versus over $1.5M for comparable models like Llama-3.2-3B. They also derive a cost-scaling law suggesting ~$4.4K suffices to match Qwen2-1.5B, and use their full pipeline access (not just weights) to study how pretraining data curricula affect downstream post-training performance. Data, code, and weights are released under Apache 2.0. Commentary highlighted the detailed technical report and its use of the MuonH optimizer, drawing comparisons to related 'Effective Learning Rate' (ELR) research, with some noting convergent insights across independent efforts on this topic.
Discussion: 2 tweets from 2 authors · @_reachsumit, @zhanpeng_zhou