PowLU: Taming the Quadratic Beast for Stable LLM Training

PowLU: An Activation Function for Stable Pre-Training of LLMs

2026-05-01
Peijie Jiang, Yuqi Feng, Cunyin Peng, Qian Zhao, Jia Liu, KunLong Chen, Zhiqiang Zhang, Jun Zhou
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces PowLU (Power Linear Unit), a novel activation function designed to stabilize Large Language Model (LLM) pre-training. By utilizing a rational power function, PowLU achieves adaptive non-linearity and effectively limits output ranges for large positive inputs, rivaling SwiGLU in performance while significantly enhancing numerical stability.

TL;DR

Researchers from the Ling Team at Ant Group have proposed PowLU (Power Linear Unit), a drop-in replacement for the popular SwiGLU activation function. While SwiGLU’s approximate quadratic growth () boosts expressivity, it often triggers catastrophic "loss spikes" in large-scale and low-precision (FP8) training. PowLU uses a rational power function to maintain non-linearity while enforcing a more stable, near-linear growth for large inputs, ensuring smooth pre-training of models up to 124B parameters.

The Problem: The Hidden Cost of SwiGLU's Expressivity

In the current LLM landscape, SwiGLU is the de facto standard. Its success stems from its ability to approximate a quadratic function for large positive inputs, providing the "non-linear muscle" needed for complex reasoning.

However, this behavior is a double-edged sword. As models grow deeper and training moves toward lower precision (FP8/FP4), this quadratic amplification:

  1. Exacerbates Outliers: Small anomalies in earlier layers are squared and magnified.
  2. Triggers Numerical Overflow: Values quickly exceed the narrow dynamic range of low-precision formats.
  3. Causes Training Collapse: Uncontrolled gradient and activation spikes manifest as the dreaded "loss spikes," forcing researchers to restart training runs.

SwiGLU vs PowLU Forward Pass Outliers Figure 1: SwiGLU (left) shows a wide band of high-magnitude outliers compared to the constrained distribution of PowLU (right).

Methodology: The Logic of Bounded Power

The core innovation of PowLU lies in its mathematical formulation for :

Why this specific form?

  • Adaptive Growth: For small , the exponent remains high, providing strong non-linearity. As , the exponent approaches 1, meaning the function transitions gracefully to linear growth.
  • Differentiability: The term is a clever inclusion. Without the , the derivative would approach infinity as , leading to new instabilities.
  • Smooth Suppression: Unlike SwiGLU-Clip, which simply "chops off" large values (hard truncation), PowLU smoothly suppresses growth, preserving more information in the latent space.

PowLU Architecture and Derivatives Figure 2: First-order derivatives highlighting the smooth, controlled gradients of PowLU compared to the explosive growth of SwiGLU.

Experiments: Performance Without the Spikes

The authors tested PowLU across several scales, including a heavy-weight 124B MoE model.

1. Stability at Scale

In FP8 training, standard SwiGLU encountered significant stability issues. While SwiGLU-Clip delayed spikes, PowLU maintained a perfectly smooth loss trajectory. Loss Curve Comparison Figure 3: Training loss curves. Notice the red line (PowLU) stays flat and stable while others spike.

2. Benchmark Excellence

Stability didn't come at the cost of "intelligence." At 124B parameters:

  • ARC-Challenge: SwiGLU (77.29) vs PowLU (83.05).
  • MATH: SwiGLU (42.22) vs PowLU (44.98).
  • HumanEval (Coding): SwiGLU (54.27) vs PowLU (55.49).

PowLU consistently matched or outperformed SwiGLU, proving that "near-linear" growth for large activations is sufficient for high-level reasoning.

Critical Insight: The Outlier Channel Solution

One of the most profound visualizations in the paper involves Outlier Channels. LLMs often suffer from a few "screaming" channels that carry disproportionately large magnitudes. PowLU significantly homogenizes these channel magnitudes, which is likely the primary reason it works so well for low-precision quantization.

Conclusion

PowLU is a mathematically grounded and empirically robust evolution of the activation function. It addresses the fundamental tension between expressive capacity and numerical stability. For any team currently pre-training LLMs in FP8 or seeking to scale beyond the 100B barrier without constant supervision of loss spikes, PowLU represents a significant step forward.

Takeaway: The era of "bigger is better" is evolving into "more stable is better." PowLU proves that by slightly curbing the wilder tendencies of SwiGLU, we can build models that are both smarter and more reliable.

Find Similar Papers

Try Our Examples

  • Search for recent papers or SOTA methods that propose alternatives to SwiGLU specifically to address numerical instability in FP8/FP4 low-precision training.
  • Which paper first identified the "quadratic amplification" issue of GLU variants in Transformer architectures, and how does PowLU's theoretical approach differ from SwiGLU-Clip?
  • Explore whether the PowLU activation function has been applied to other architectures like Vision Transformers (ViT) or State Space Models (SSM) to improve training stability.
Contents
PowLU: Taming the Quadratic Beast for Stable LLM Training
1. TL;DR
2. The Problem: The Hidden Cost of SwiGLU's Expressivity
3. Methodology: The Logic of Bounded Power
3.1. Why this specific form?
4. Experiments: Performance Without the Spikes
4.1. 1. Stability at Scale
4.2. 2. Benchmark Excellence
5. Critical Insight: The Outlier Channel Solution
6. Conclusion