From One to Millions: Scaling Personal Models on Trillion-Parameter Priors

On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters

2026-06-01
Mind Lab, :, Song Cao, Vic Cao, Kaijie Chen, Bunny Fan, Hera Feng, Huan Feng, Arthur Fu, Jun Gao, Hongquan Gu, Aaron Guan, Mutian Hong, Hailee Hou, Peixuan Hua, Charles Huang, Miles Jiang, Nora Jiang, Yuyi Jiang, Autumn Jin, Fancy Kong, Kyrie Lei, Alexy Li, Dawn Li, Ray Li, Theo Li, Wenhao Li, Jiayi Lin, Domini Liu, Heshan Liu, Kairus Liu, Logan Liu, Maeve Luo, Runism Lv, Pony Ma, Verity Niu, Anson Qiu, Vincent Wang, Maxwell Yao, Regis Ye, Wenlin Ye, Yanying Ye, Josh Ying, Danney Zeng, Salmon Zhan, Anya Zhang, Ruijia Zhang, Shiyang Zhang, Sueky Zhang, Ya Zhang, Wei Zhao, Ada Zhou, Sizer Zhou, Xinyue Zhu, Murphy Zhuang, Anya Zhang, Di Zhang, Ruijia Zhang, Shiyang Zhang, Sueky Zhang, Ya Zhang, Wei Zhao, Ada Zhou, Adrian Zhou, Yuhua Zhou, Xinyue Zhu, Murphy Zhuang
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a paradigm shift in PEFT, positioning small adapters as persistent local state for personal models rather than just cheap fine-tuning alternatives. Through the Mind Lab's "MinT" infrastructure, they demonstrate "Scale Up, Scale Down, and Scale Out" strategies, achieving SOTA results in trillion-scale MoE reinforcement learning and million-user simulation.

TL;DR

The frontier of AI is shifting from building "one model to rule them all" to "millions of models that know you." This paper from Mind Lab presents a comprehensive blueprint for Parameter-Efficient Fine-Tuning (PEFT) as the cornerstone of persistent personal identity. By scaling Up (1T+ MoE models), Down (rank-1 stability), and Out (addressing millions of users), they demonstrate how millions of unique "adapter-brains" can thrive on a single shared biological-scale prior.

Evolution of the Personal Model: Why Prompts Aren't Enough

A capable assistant is not automatically a personal one. While long contexts and RAG help, they are transient. To achieve true continuity, a model needs persistent state. The authors argue that LoRA (Low-Rank Adaptation) is not just a budget hack; it is the "genetic variation" (the 0.1% difference) that allows one shared "human biology" (the base model) to support billions of individuated lives.

The Three-Axis Scaling Framework

1. Scale Up: Bridging the Trillion-Parameter Gap

Personalization is high-leverage only if the base model is powerful. The authors successfully implemented LoRA RL on a 1.04 Trillion parameter MoE (Mixture-of-Experts).

  • The Problem: Training-Inference Mismatch (TIM). In MoE, if the training engine routes tokens differently than the inference engine, the gradients become meaningless.
  • The Solution: Router Replay (R3). By recording routing decisions during rollout and replaying them during training, they stabilized the KL divergence and ensured valid policy updates.

Kimi K2 MoE LoRA RL training curves

2. Scale Down: Stability at the Edge of Collapse

To support millions of users, adapters must be tiny (Rank 1). Usually, Rank 1 adapters are unstable and fail across different seeds.

  • OLoRA-tail: The authors found that standard initialization wastes the limited "direction" of a Rank-1 adapter. By focusing on the minor singular vectors (the "tail" of the representation) and removing aggressive scaling, they made Rank-1 adaptation reliable even at large batch sizes.
  • δ-mem: Beyond static weights, they introduced stateful adapters that act like a "working memory," writing residuals of the history into a compact associative state.

Rank-1 Stability Comparison

3. Scale Out: Diversity as Collective Intelligence

When you have 200 different adapters, you have more than just 200 users; you have a "Council of Experts."

  • Collective Performance: By aggregating the majority vote of 198 diverse LoRA variants, the team saw a massive jump in AIME24 accuracy (from 36% to 48%). This proves that distinct adapter trajectories learn complementary reasoning paths that simple "repeated sampling" from a single model cannot replicate.
  • User Simulation: In the OASIS social simulator, per-user LoRAs prevented the "personality collapse" seen in prompt-only agents, producing richer interaction topologies and more realistic "echo chamber" dynamics.

Majority Vote Scaling Law

Infrastructure: The MinT System

You cannot manage a million personal models with a folder full of files. Mind Lab's MinT infrastructure treats policies as identities with separate lifecycles:

  1. Resident Base Models: Keep the 1T+ model warm in memory.
  2. Tiered Residency: Orchestrate adapters between GPU slots (hot), CPU cache (warm), and shared storage (cold).
  3. Two-Phase Readiness: Pre-warming adapters before exposing them to users to prevent "TTFT stalls" that ruin user experience.

Critical Insight & Future Outlook

The core takeaway is that Diversity is a Resource. By making adaptation cheap (Scale Down) and infrastructure robust (MinT), we stop trying to build a single "perfect" model and instead foster a population of specialized ones.

Future Challenges:

  • Memory Capacity: LoRA has a "hard ceiling" for how many facts it can store (roughly tokens per parameter). We need better "Write Policies" (Context Learning) to decide what deserves to be carved into the weights.
  • Drift: Ensuring that millions of personal models don't eventually regress to the "average" prior over long interaction horizons.

This paper provides the technical foundation for a future where your AI isn't just a leased endpoint from a big tech company, but a persistent, evolving extension of your own digital identity.

Find Similar Papers

Try Our Examples

  • Search for recent papers that address training-inference mismatch (TIM) and expert routing stability in Mixture-of-Experts (MoE) reinforcement learning.
  • Which paper first proposed the "square-root scaling rule" (rsLoRA) for rank stabilization, and how has its application evolved in RL-based fine-tuning?
  • Find research studies exploring the use of low-rank adapters as a persistent memory substrate in multi-agent social simulations or user-modeling tasks.
Contents
From One to Millions: Scaling Personal Models on Trillion-Parameter Priors
1. TL;DR
2. Evolution of the Personal Model: Why Prompts Aren't Enough
3. The Three-Axis Scaling Framework
3.1. 1. Scale Up: Bridging the Trillion-Parameter Gap
3.2. 2. Scale Down: Stability at the Edge of Collapse
3.3. 3. Scale Out: Diversity as Collective Intelligence
4. Infrastructure: The MinT System
5. Critical Insight & Future Outlook