LatentMAS: Breaking the Text Bottleneck for Multi-Agent Intelligence

Latent Collaboration in Multi-Agent Systems

2025-01-01
Jiaru Zou, Xiyuan Yang, Ruizhong Qiu, Gaotang Li, Katherine Tieu, Pan Lu, Ke Shen, Hanghang Tong, Yejin Choi, Jingrui He, James Zou, Mengdi Wang, Ling Yang
Summary
Problem
Method
Results
Takeaways
Abstract

LatentMAS is an end-to-end, training-free framework for multi-agent systems (MAS) that enables Large Language Models (LLMs) to collaborate directly within the continuous latent space. By replacing text-based communication with latent thoughts and shared KV-cache working memory, it achieves state-of-the-art performance across 9 benchmarks including math, science, and coding.

TL;DR

Multi-agent systems (MAS) have long been constrained by the "textual bottleneck"—the slow, lossy process of agents talking to each other through natural language. LatentMAS changes the game by enabling agents to "think" and "collaborate" entirely in the continuous latent space. The result? A framework that is 4x faster, uses 80% fewer tokens, and is actually more accurate than traditional text-based systems—all without any retraining.

The "Lingua Franca" Problem in Agentic AI

In the current LLM landscape, agents communicate like humans: they write out thoughts, pass them as text, and the next agent re-encodes that text. This is fundamentally inefficient.

  1. Discretization Loss: Forcing a high-dimensional internal representation into discrete tokens loses semantic nuance.
  2. Computational Waste: Decoding and re-encoding steps consume significant FLOPs and time.
  3. Inflexible Reasoning: Natural language is linear, while latent thoughts can be O(dh/log|V|) more expressive.

LatentMAS asks a radical question: Can agents collaborate directly through their internal "brains" (hidden states) instead of their "mouths" (text outputs)?

Methodology: High-Fidelity Latent Collaboration

The technical core of LatentMAS rests on two pillars: Latent Thought Generation and Latent Working Memory.

1. Auto-regressive Latent Reasoning

Instead of generating tokens, each agent generates a sequence of last-layer hidden embeddings. To prevent the model from getting "confused" by these high-level embeddings when fed back as input, the authors use a clever Alignment Operator (). Using a pseudo-inverse mapping, realigns the output hidden states back into the input embedding space.

Model Architecture Figure: The LatentMAS Pipeline showing internal thought generation and cross-agent memory transfer.

2. Lossless Memory Transfer

In text-based MAS, a "Critic" agent reads a "Planner's" text. In LatentMAS, the Critic inherits the Planner's layer-wise KV-cache. This "Latent Working Memory" ensures that the successor agent's computation is conditioned on the exact internal state of the predecessor. Theorem 3.3 in the paper proves this is mathematically equivalent to re-processing the entire sequence, but without the re-computation overhead.

Empirical Superiority: Faster, Cheaper, Better

The results across 9 benchmarks (GSM8K, HumanEval+, GPQA, etc.) are striking.

  • Inference Speed: LatentMAS consistently delivers a 4.3x speedup because it skips the expensive token-by-token decoding process in intermediate steps.
  • Efficiency: Total token usage drops by over 70%. Intermediate agents generate zero text; only the final agent decodes the actual answer.
  • Accuracy: In complex tasks like AIME25 and MBPP+, LatentMAS sees up to 14.6% improvement.

Efficiency Results Figure: Performance comparison across accuracy, speed, and token usage.

Visualizing Semantic Meaning

One might worry that "latent thoughts" are just noise. The authors used t-SNE to compare LatentMAS embeddings with traditional text-generated embeddings.

Semantic Consistency Figure: Latent thoughts share the same semantic region as text but offer higher density and diversity.

The visualization confirms that latent thoughts stay within the "meaningful" region of the embedding space while actually offering richer representational diversity than constrained tokens.

Critical Insight: The End of Human-Centric Agents?

LatentMAS suggests a future where "System 2" reasoning happens in a dark, high-dimensional space where humans can't read every step. While this raises challenges for interpretability, the authors address this with a "Debug Mode" that can probe latent thoughts into text.

The Takeaway: If you want efficient multi-agent systems, stop making your agents talk to each other in English. Let them communicate in the language they were born with: Vectors.

Conclusion

LatentMAS provides a scalable, training-free paradigm that effectively separates reasoning from communication. By treating multi-agent collaboration as a distributed latent process rather than a dialogue, it clears the path for highly efficient, system-level intelligence.

Find Similar Papers

Try Our Examples

  • Find recent papers investigating training-free cross-model communication through KV-cache sharing or latent state alignment.
  • What are the seminal works proposing the Linear Representation Hypothesis in Transformer hidden states, and how has this theory evolved for multi-agent contexts?
  • Explore research that applies latent reasoning or continuous-space collaboration to multi-modal agents (e.g., Vision-Language Models) rather than text-only LLMs.
Contents
LatentMAS: Breaking the Text Bottleneck for Multi-Agent Intelligence
1. TL;DR
2. The "Lingua Franca" Problem in Agentic AI
3. Methodology: High-Fidelity Latent Collaboration
3.1. 1. Auto-regressive Latent Reasoning
3.2. 2. Lossless Memory Transfer
4. Empirical Superiority: Faster, Cheaper, Better
5. Visualizing Semantic Meaning
6. Critical Insight: The End of Human-Centric Agents?
7. Conclusion