Extra-Merge: Tracing the Rank-1 Subspace for a "Free Lunch" in LLM Pre-training
Extra-Merge: Tracing the Rank-1 Subspace of Model Merging in Language Model Pre-Training
The paper introduces Extra-Merge, a training-free model merging strategy that enhances Large Language Models (LLMs) by extrapolating along a newly discovered "Rank-1 Subspace." By analyzing late-stage pre-training trajectories, the authors demonstrate that while raw optimization is chaotic, merged checkpoints collapse onto a stable, one-dimensional linear manifold, achieving SOTA-level performance gains on GPT-2, LLaMA, and Pythia models.
TL;DR
Researchers have discovered that the chaotic path of a Large Language Model (LLM) during pre-training hides a remarkably simple secret: when you average recent checkpoints, the optimization trajectory collapses into a nearly perfect straight line (a Rank-1 Subspace). By simply extending this line—a strategy called Extra-Merge—we can lower the model's loss and boost accuracy without spending a single cent more on GPU training.
The "Oscillation" Problem: Why Training is Noisy
Standard LLM training is a violent affair. Even in the late stages, the model doesn't drift smoothly toward the bottom of the loss mountain; instead, it "bounces" off the steep walls of the loss valley.
If you interpolate between two raw checkpoints ( and ), you often find a "convex basin" where the midpoint is better than the endpoints. This proves that the raw optimizer is zig-zagging. Previous methods like PMA (Pre-trained Model Averaging) tried to fix this by taking the average (the centroid) of these points, which typically lands in a flatter, better region. But they stopped there—treating the optimization history as a static cloud rather than a directed path.
The Breakthrough: The Rank-1 Subspace
By applying Principal Component Analysis (PCA) to merged checkpoints, the authors found a startling geometric shift. While raw checkpoints have their variance scattered across many dimensions, merged checkpoints concentrate over 94% of their variance into a single dimension.

In the figure above, note how merging "rectifies" the trajectory. The chaotic blue raw steps become a smooth, monotonic pink flow.
The "River-Valley" Intuition
Think of the loss landscape as a long, narrow river valley.
- The Mountains: High-curvature directions where the model bounces back and forth (Noise).
- The River: The flattest direction along the valley floor where real progress happens (Signal).
Averaging acts as a geometric low-pass filter. It cancels out the "mountain" oscillations, leaving only the "river" drift. This allows us to see exactly where the model wants to go.
Methodology: How Extra-Merge Works
Extra-Merge doesn't just average; it extrapolates.
- Direction Estimation: It looks at the last merged checkpoints and uses PCA to find the primary axis of descent ().
- Line Search: It moves the model parameters further along this direction.
- Adaptive Step: The distance moved is scaled by the "velocity" of the model's recent progress.
The illustration shows the raw optimizer bouncing off the "mountains" while Extra-Merge cruises straight down the "river" floor.
Experimental Proof: Better Performance for Free
The authors tested this on models ranging from GPT-2 (124M) to Pythia (12B).
- Validation Loss: Across GPT-2 and LLaMA, Extra-Merge consistently stayed below the raw baseline and the standard PMA average.
- Downstream Tasks: On tasks like ARC and PIQA, Pythia-12B saw a +0.59% accuracy boost simply by re-calculating the weights at the end of training.
- Optimizer Agnostic: It even works with Muon, a modern optimizer that uses orthogonal updates, proving the Rank-1 phenomenon is an intrinsic property of LLM training, not just a quirk of AdamW.
The charts show Extra-Merge (light blue) achieving significantly lower loss than both the Raw (dark blue) and PMA (pink) models.
Critical Analysis & Takeaways
The beauty of Extra-Merge lies in its simplicity. It requires no extra GPU memory for gradients and no extra data.
Limitations:
- The "straight-river" approximation holds best during the late stages of training. Early on, the landscape might be too curved for 1D extrapolation to be safe.
- It requires saving multiple checkpoints, which increases storage overhead.
Future Outlook: This research opens the door to Subspace-Aware Optimizers. Instead of discovering the river floor after training, we could potentially design optimizers that actively align themselves with this Rank-1 manifold from the start, potentially slashing pre-training costs by orders of magnitude.
Conclusion: Extra-Merge proves that we are often underestimating our models. The information for a better model is already present in the training history; we just need the right geometric lens to see it.
