Blog

No more endless PDFs. Discover the core value of the latest top-tier research in one article.

PivotRL: High Accuracy Agentic Post-Training at Low Compute Cost
Does This Gradient Spark Joy?
How Out-of-Equilibrium Phase Transitions can Seed Pattern Formation in Trained Diffusion Models
Towards Practical World Model-based Reinforcement Learning for Vision-Language-Action Models
MOSS-TTSD: Text to Spoken Dialogue Generation
P^2O: Joint Policy and Prompt Optimization
RefracGS: Novel View Synthesis Through Refractive Water Surfaces with 3D Gaussian Ray Tracing
CounterScene: Counterfactual Causal Reasoning in Generative World Models for Safety-Critical Closed-Loop Evaluation
EvoIdeator: Evolving Scientific Ideas through Checklist-Grounded Reinforcement Learning
Dreaming the Unseen: World Model-regularized Diffusion Policy for Out-of-Distribution Robustness
Relax Forcing: Relaxed KV-Memory for Consistent Long Video Generation
Text-Image Conditioned 3D Generation
Uncertainty in wind and solar projections depends on global and regional climate models
A transformer architecture alteration to incentivise externalised reasoning
Two Experts Are Better Than One Generalist: Decoupling Geometry and Appearance for Feed-Forward 3D Gaussian Splatting
OmniCodec: Low Frame Rate Universal Audio Codec with Semantic-Acoustic Disentanglement
Delightful Distributed Policy Gradient
Numerically stable equations for the orbital evolution of compact object binaries
Predicting States of Understanding in Explanatory Interactions Using Cognitive Load-Related Linguistic Cues
Agentic AI and the next intelligence explosion
CayleyPy-4: AI-Holography. Towards analogs of holographic string dualities for AI tasks
mSFT: Addressing Dataset Mixtures Overfiting Heterogeneously in Multi-task SFT
OpenEarth-Agent: From Tool Calling to Tool Creation for Open-Environment Earth Observation
Thinking Deeper, Not Longer: Depth-Recurrent Transformers for Compositional Generalization
Physion-Eval: Evaluating Physical Realism in Generated Video via Human Reasoning
Fiber-optic quantum interface with an array of more than 100 individually addressable atoms on an optical nanofiber
TiCo: Time-Controllable Training for Spoken Dialogue Models
UniGRPO: Unified Policy Optimization for Reasoning-Driven Visual Generation
CanViT: Toward Active-Vision Foundation Models
Effective Strategies for Asynchronous Software Engineering Agents