Blog

No more endless PDFs. Discover the core value of the latest top-tier research in one article.

The Art of Efficient Reasoning: Data, Reward, and Optimization
LeRobot: An Open-Source Library for End-to-End Robot Learning
Probing the Geometry of Diffusion Models with the String Method
GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL
Vision Transformers Need More Than Registers
Learning from Trials and Errors: Reflective Test-Time Planning for Embodied LLMs
WHOLE: World-Grounded Hand-Object Lifted from Egocentric Videos
Why Diffusion Language Models Struggle with Truly Parallel (Non-Autoregressive) Decoding?
Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the Wild
Architecting AgentOS: From Token-Level Context to Emergent System-Level Intelligence
SymTorch: A Framework for Symbolic Distillation of Deep Neural Networks
From Blind Spots to Gains: Diagnostic-Driven Iterative Training for Large Multimodal Models
tttLRM: Test-Time Training for Long Context and Autoregressive 3D Reconstruction
Transcoder Adapters for Reasoning-Model Diffing
Actor-Curator: Co-adaptive Curriculum Learning via Policy-Improvement Bandits for RL Post-Training
UCM: Unifying Camera Control and Memory with Time-aware Positional Encoding Warping for World Models
Unleashing the Potential of Diffusion Models for End-to-End Autonomous Driving
The Design Space of Tri-Modal Masked Diffusion Models
Causal Motion Diffusion Models for Autoregressive Motion Generation
Flow Matching is Adaptive to Manifold Structures
Behavior Learning (BL): Learning Hierarchical Optimization Structures from Data
AgentVista: Evaluating Multimodal Agents in Ultra-Challenging Realistic Visual Scenarios
BetterScene: 3D Scene Synthesis with Representation-Aligned Generative Model
Prompt-Level Distillation: A Non-Parametric Alternative to Model Fine-Tuning for Efficient Reasoning
AgentDropoutV2: Optimizing Information Flow in Multi-Agent Systems via Test-Time Rectify-or-Reject Pruning
Towards Better RL Training Data Utilization via Second-Order Rollout
UFO: Unifying Feed-Forward and Optimization-based Methods for Large Driving Scene Modeling
Towards Long-Form Spatio-Temporal Video Grounding
TrajTok: Learning Trajectory Tokens enables better Video Understanding
EmbodMocap: In-the-Wild 4D Human-Scene Reconstruction for Embodied Agents