Beyond Tokens: Why Predicting Your Own Latents is the Key to Data Efficiency

Learn from your own latents and not from tokens: A sample-complexity theory

2026-05-01
Daniel J. Korchinski, Alessandro Favero, Matthieu Wyart
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces a sample-complexity theory for Self-Supervised Learning (SSL) methods that predict their own latent representations (e.g., data2vec, JEPA). Using the Random Hierarchy Model (RHM) as a benchmark, the authors prove that latent prediction reduces sample complexity from exponential to constant relative to the depth of the data's hierarchical tree, achieving a scaling of .

TL;DR

Why can a child learn a language from a few million words while an LLM requires trillions? This paper argues the culprit is token-level prediction. By switching to a paradigm where models predict their own internal latent representations (as seen in data2vec or JEPA), we can reduce the data needed to learn complex hierarchical structures from exponential to constant.

The Problem: The "Token Bottleneck"

In current SOTA models, even if the goal is to understand high-level concepts (e.g., "justice" or "gravity"), the supervision signal comes from predicting surface-level tokens. In a hierarchical world—modeled here by the Random Hierarchy Model (RHM)—semantic concepts live at the root of a tree, while tokens are the leaves.

The authors show that for a tree of depth , token-level SSL (like GPT's next-token prediction) requires a sample complexity of . As the hierarchy gets deeper, the amount of data needed explodes. This is the Token Bottleneck: the high-level correlation is "diluted" through every layer of the tree before it reaches the tokens.

The Solution: Learn from Your Own Latents

The core insight is simple: Once you've learned a concept at level , use it as a target to learn level .

By predicting latents instead of tokens, the "distance" between the predictor and the target remains constant () regardless of the total depth of the tree. This collapses the complexity from down to a manageable .

1. Iterative Latent Clustering (ILC)

The authors first prove this mathematically using an algorithm that clusters "context vectors." If two fragments of data (tuples) share the same parent in the latent tree, they are "synonyms" and should have similar context vectors.

Comparison of Sample Complexity Figure 1: Comparison of tree-distance and sample complexity across different objectives. Note how latent supervision (right) maintains a fixed distance, bypassing the depth penalty.

2. Stacked Latent-Clustering (SLC)

To bring this into the neural domain, the authors designed the SLC Network. It consists of a stack of modules, each containing:

  • A Predictor: Guesses the latent representation of neighboring "cousin" patches.
  • A Clusterer: Maps these predictions into a discrete codebook.

Surprisingly, they found that this works even with local learning rules (stop-gradients between layers), which has profound implications for biological plausibility in neuroscience.

Evidence: Decoding data2vec

The paper's most impressive feat is providing a theoretical framework for data2vec. They prove that data2vec isn't just a clever engineering trick; it implicitly performs hierarchical clustering.

data2vec Scaling Collapse Figure 2: Performance of data2vec on RHM. When the x-axis is rescaled by , the curves for different parameters collapse, confirming the theoretical scaling law.

Experimental results show that:

  1. The encoder internally clusters synonyms at every level of the hierarchy.
  2. The sample complexity to reach adult-level competence on the RHM remains constant as tree depth increases.

Critical Insight: Is H-JEPA Redundant?

The authors drop a bombshell for architecture designers: Explicitly stacking JEPA modules (like in H-JEPA) might be redundant. Since a single latent-prediction loop (like data2vec) naturally climbs the hierarchy through an EMA teacher-student dynamic, the added complexity of multi-scale architectures may not provide the "scaling break" researchers hoped for—because the original algorithm was already doing it.

Conclusion & Future Outlook

This paper provides the first formal "Sample Complexity" proof for why latent-space prediction is superior to pixel or token reconstruction.

Takeaways for the Industry:

  • The Next Frontier: Moving away from "Next-Token Prediction" toward "Next-Latent Prediction" is likely necessary to solve the data-scarcity problem in specialized domains (e.g., scientific reasoning, rare languages).
  • Neuroscience Link: This reinforces the "Predictive Coding" theory of the human cortex, suggesting our brains are efficient because we predict abstract neural states, not raw sensory input.

Limitations: The RHM is a "clean" grammar. Natural language involves ambiguous, recursive, and context-dependent rules that might complicate the bound. However, as a baseline, this work sets a new theoretical gold standard for SSL efficiency.

Find Similar Papers

Try Our Examples

  • Find recent papers that compare the data efficiency of Joint-Embedding Predictive Architectures (JEPA) against standard Autoregressive Transducers in large-scale language modeling.
  • Which paper first introduced the Random Hierarchy Model (RHM), and how does the current work's analysis of self-supervised latent prediction differ from the original supervised learning complexity bounds?
  • Explore research applying data2vec or similar latent-prediction objectives to multi-modal tasks like video-audio alignment to see if hierarchical clustering emerges across different signal types.
Contents
Beyond Tokens: Why Predicting Your Own Latents is the Key to Data Efficiency
1. TL;DR
2. The Problem: The "Token Bottleneck"
3. The Solution: Learn from Your Own Latents
3.1. 1. Iterative Latent Clustering (ILC)
3.2. 2. Stacked Latent-Clustering (SLC)
4. Evidence: Decoding data2vec
5. Critical Insight: Is H-JEPA Redundant?
6. Conclusion & Future Outlook