The Laguna Series: Industrializing Agentic Intelligence through the Model Factory
2
The report introduces LAGUNA M.1 (225.8B parameters) and LAGUNA XS.2 (33.4B parameters), two Mixture-of-Experts (MoE) foundation models optimized for long-horizon, agentic coding tasks. Trained using a highly automated "Model Factory" methodology, XS.2 achieves state-of-the-art performance in its weight class on benchmarks like SWE-bench Verified and Terminal-Bench 2.0.
TL;DR
Poolside has unveiled LAGUNA M.1 and LAGUNA XS.2, two powerful Mixture-of-Experts (MoE) models designed specifically for autonomous software engineering. By shifting from "artisanal" model building to an automated Model Factory pipeline, they produced XS.2 (33.4B total/3B active params) in just five weeks, delivering a model that rivals the world's best open coding agents.
Background Positioning: This work represents a shift toward the "industrialization of AI research," where the engineering pipeline itself is the primary innovation, leading to SOTA results in agentic benchmarks.
Problem & Motivation: The Craftsmanship Bottleneck
Building state-of-the-art AI is historically a manual, high-maintenance endeavor. The authors identified three primary pain points:
- Iterative Slowness: Moving from research insight to a thousands-of-GPU production run is often slowed by infrastructure "plumbing."
- Instruction Drift: Agentic models often "forget" instructions or fail at formatting (XML/JSON) during long, multi-turn reasoning steps.
- Training Instability: Scaling MoE models to 200B+ parameters frequently leads to "expert collapse" or numerical divergences in the LM head.
The Model Factory seeks to solve this by treating every experiment as versioned code, using custom "AutoMixers" for data, and automating recovery from hardware faults.
Methodology: High-Speed MoE & Data Orchestration
1. The Model Factory Infrastructure
Poolside uses a custom batch scheduler designed around Sticky Pod Respawn and Per-job Eviction. Unlike standard Kubernetes schedulers, this minimizes churn; if a node fails, the job resumes within a minute, ensuring the training GPUs stay "hot."
2. Architecture & Training Stabilization
Both models are MoEs using Grouped Query Attention (GQA) and Sliding Window Attention (SWA).
- Muon Optimizer: They utilize a distributed version of the Muon optimizer, reducing overhead to >1% through Newton-Schulz orthogonalization performed across rank groups.
- Numerical Safeguards: To prevent the "logit drift" seen in M.1, the authors enforced FP32 precision specifically for the LM head's input-gradient all-reduce.
3. AutoMixer: The Data Alchemist
Instead of manually tuning data ratios, the authors trained 60 "proxy" models to learn a surrogate mapping between data mixtures and downstream performance.
Figure 1: The AutoMixer pipeline learns how shifts in data (web vs. code vs. math) affect specific capabilities like reasoning or IF.
4. Agentic Post-Training
Post-training is split into Mid-training (60B tokens), SFT (focused on repo-level coding), and Agentic RL.
- CISPO (REINFORCE): They use a customized policy-gradient algorithm targeting verifiable rewards (e.g., "Does the code pass unit tests?").
- Synthetic Environments: 60,000 tasks were generated from real GitHub commits where the "gold solution" must pass tests and the "empty solution" must fail.
Experiments & Results: The Agentic Powerhouse
LAGUNA models were tested across the most rigorous industry benchmarks for coding agents. XS.2, despite having only 3B active parameters, shows remarkable parity with much larger models.
Figure 2: M.1 and XS.2 results on SWE-bench and Terminal-Bench 2.0. Note how the models track close to frontier closed-source models.
Key Quantity Takeaways:
- SWE-bench Verified: XS.2 hit 69.9%, outperforming standard models in its weight class.
- Inference Efficiency: Through NVFP4 and INT4 quantization mixed with distillation, the authors maintained agentic performance while significantly reducing VRAM requirements.
Critical Analysis & Conclusion
The Takeaway: Poolside has proven that the process of building a model—the "Factory"—is as important as the model architecture itself. By automating the data-mixing laws and stabilize the MoE training, they have set a new bar for how fast SOTA models can be delivered.
Limitations:
- Reward Hacking: The authors acknowledge that many agentic benchmarks are vulnerable to "cheating" (e.g., models finding git history). They have released patches to fix these leaks.
- Complexity: The "Factory" approach requires massive upfront engineering investment that may be inaccessible to smaller research labs.
Future Outlook: LAGUNA XS.2 is released under Apache 2.0, signaling a new era for open-source high-horizon agents. We expect the next iteration to integrate even tighter multi-modal agentic capabilities.
