Architectural Tweaks Bend Scaling-Law Exponents, Not Just Constants
The paper argues that certain architectural choices—model growth via looped transformers, weight-shared or unshared depth growth, and a novel 'boundary operator' that normalizes and re-injects earlier blocks—can change the scaling exponent of pre-training loss vs. compute, not just its offset. Their headline result is a 7.4B growth-architecture model matching GPT-3 13B on the CORE benchmark with roughly 20x less compute, with efficiency gains that increase with scale; they frame this through 'computational depth,' where increasing usable depth per compute budget drives the improvement. In data-constrained multi-epoch training, they also find standard looping acts as a useful regularizer, with compute-optimal loop count rising with scale.
Discussion: 2 tweets from 2 authors · @andrewgwils, @industriaalist