Extra-Merge: Tracing the Rank-1 Subspace for a "Free Lunch" in LLM Pre-training

Extra-Merge: Tracing the Rank-1 Subspace of Model Merging in Language Model Pre-Training

2026-05-01
Wenjie Zhou, Bohan Wang, Hongtao Zhang, Chenxi Jia, Wei Chen, Xueqi Cheng
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces Extra-Merge, a training-free model merging strategy that enhances Large Language Models (LLMs) by extrapolating along a newly discovered "Rank-1 Subspace." By analyzing late-stage pre-training trajectories, the authors demonstrate that while raw optimization is chaotic, merged checkpoints collapse onto a stable, one-dimensional linear manifold, achieving SOTA-level performance gains on GPT-2, LLaMA, and Pythia models.

TL;DR

Researchers have discovered that the chaotic path of a Large Language Model (LLM) during pre-training hides a remarkably simple secret: when you average recent checkpoints, the optimization trajectory collapses into a nearly perfect straight line (a Rank-1 Subspace). By simply extending this line—a strategy called Extra-Merge—we can lower the model's loss and boost accuracy without spending a single cent more on GPU training.

The "Oscillation" Problem: Why Training is Noisy

Standard LLM training is a violent affair. Even in the late stages, the model doesn't drift smoothly toward the bottom of the loss mountain; instead, it "bounces" off the steep walls of the loss valley.

If you interpolate between two raw checkpoints ( and ), you often find a "convex basin" where the midpoint is better than the endpoints. This proves that the raw optimizer is zig-zagging. Previous methods like PMA (Pre-trained Model Averaging) tried to fix this by taking the average (the centroid) of these points, which typically lands in a flatter, better region. But they stopped there—treating the optimization history as a static cloud rather than a directed path.

The Breakthrough: The Rank-1 Subspace

By applying Principal Component Analysis (PCA) to merged checkpoints, the authors found a startling geometric shift. While raw checkpoints have their variance scattered across many dimensions, merged checkpoints concentrate over 94% of their variance into a single dimension.

Geometric shift from Convexity to Monotonicity

In the figure above, note how merging "rectifies" the trajectory. The chaotic blue raw steps become a smooth, monotonic pink flow.

The "River-Valley" Intuition

Think of the loss landscape as a long, narrow river valley.

  1. The Mountains: High-curvature directions where the model bounces back and forth (Noise).
  2. The River: The flattest direction along the valley floor where real progress happens (Signal).

Averaging acts as a geometric low-pass filter. It cancels out the "mountain" oscillations, leaving only the "river" drift. This allows us to see exactly where the model wants to go.

Methodology: How Extra-Merge Works

Extra-Merge doesn't just average; it extrapolates.

  1. Direction Estimation: It looks at the last merged checkpoints and uses PCA to find the primary axis of descent ().
  2. Line Search: It moves the model parameters further along this direction.
  3. Adaptive Step: The distance moved is scaled by the "velocity" of the model's recent progress.

Model Architecture/Concept The illustration shows the raw optimizer bouncing off the "mountains" while Extra-Merge cruises straight down the "river" floor.

Experimental Proof: Better Performance for Free

The authors tested this on models ranging from GPT-2 (124M) to Pythia (12B).

  • Validation Loss: Across GPT-2 and LLaMA, Extra-Merge consistently stayed below the raw baseline and the standard PMA average.
  • Downstream Tasks: On tasks like ARC and PIQA, Pythia-12B saw a +0.59% accuracy boost simply by re-calculating the weights at the end of training.
  • Optimizer Agnostic: It even works with Muon, a modern optimizer that uses orthogonal updates, proving the Rank-1 phenomenon is an intrinsic property of LLM training, not just a quirk of AdamW.

Performance Comparison The charts show Extra-Merge (light blue) achieving significantly lower loss than both the Raw (dark blue) and PMA (pink) models.

Critical Analysis & Takeaways

The beauty of Extra-Merge lies in its simplicity. It requires no extra GPU memory for gradients and no extra data.

Limitations:

  • The "straight-river" approximation holds best during the late stages of training. Early on, the landscape might be too curved for 1D extrapolation to be safe.
  • It requires saving multiple checkpoints, which increases storage overhead.

Future Outlook: This research opens the door to Subspace-Aware Optimizers. Instead of discovering the river floor after training, we could potentially design optimizers that actively align themselves with this Rank-1 manifold from the start, potentially slashing pre-training costs by orders of magnitude.

Conclusion: Extra-Merge proves that we are often underestimating our models. The information for a better model is already present in the training history; we just need the right geometric lens to see it.

Find Similar Papers

Try Our Examples

  • Which recent papers explore the "River-Valley" loss landscape or the "Edge of Stability" phenomenon in the context of scaling laws for Large Language Models?
  • Find the original paper introducing Stochastic Weight Averaging (SWA) and subsequent works that transitioned from interpolation to extrapolation in neural network optimization.
  • Are there any studies applying PCA-based subspace extrapolation to domain adaptation or fine-tuning of multi-modal models like CLIP or stable diffusion?
Contents
Extra-Merge: Tracing the Rank-1 Subspace for a "Free Lunch" in LLM Pre-training
1. TL;DR
2. The "Oscillation" Problem: Why Training is Noisy
3. The Breakthrough: The Rank-1 Subspace
3.1. The "River-Valley" Intuition
4. Methodology: How Extra-Merge Works
5. Experimental Proof: Better Performance for Free
6. Critical Analysis & Takeaways