Blog

No more endless PDFs. Discover the core value of the latest top-tier research in one article.

Multi-Turn On-Policy Distillation with Prefix Replay
Do Modules Stay in Their Lane? Role Drift in Compound LLM Systems
LLMs can't jump
From Score Approximation to Distribution Approximation in Score-Based Diffusion Models
ACM: Agentic Context Management for Long Horizon Tasks
Not All LLM Reasoning is Visible in the Chain-of-Thought
LLMs Get Lost in Evolving User Intent
Matryoshka Agent: Unfolding Sub-Agents for Long-Horizon Machine Learning Engineering
Training Compute-Optimal Large Language Models
AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use
Can AI agents conduct open-ended AI research? Early evidence from two case studies
SkillSmith: Learning to Compose Parametric Skills and Textual Knowledge
Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents
Inducing language models to assert their own consciousness restores human beliefs and values
Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions
Memory Caching: RNNs with Growing Memory
ReToken: One Token to Improve Vision-Language Models for Visual Retrieval
Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers
Emergent Capabilities Arise Randomly from Learning Sparse Attention Patterns
Zero-Mem: Zero-Token Memory Operations for LLM Agents
Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B
TRAINING nGPT
The Case Against Generation for Retrieval: Discriminative Language Models as Effective Retrievers
MatrAIx: Simulating the World with 8.3 Billion Persona Agents
The Tragedy of the Cognitive Commons: How AI Could Disrupt the Regeneration of Professional Expertise
Can Induced Emotion Bias LLM Behaviors in Sequential Decision Making?
Stealing Reasoning Traces from Proprietary LLM APIs
Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements
Meta n : Recursive Self-Improvement through Emergent Depth