Blog

No more endless PDFs. Discover the core value of the latest top-tier research in one article.

ClawBench: Can AI Agents Complete Everyday Online Tasks?
In-Place Test-Time Training
Towards Real-world Human Behavior Simulation: Benchmarking Large Language Models on Long-horizon, Cross-scenario, Heterogeneous Behavior Traces
Scal3R: Scalable Test-Time Training for Large-Scale 3D Reconstruction
HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents
Your Agent Is Mine: Measuring Malicious Intermediary Attacks on the LLM Supply Chain
Learning is Forgetting: LLM Training As Lossy Compression
SIM1: Physics-Aligned Simulator as Zero-Shot Data Scaler in Deformable Worlds
What do Language Models Learn and When? The Implicit Curriculum Hypothesis
Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory
LPM 1.0: Video-based Character Performance Model
Synthetic Data for any Differentiable Target
Small Vision-Language Models are Smart Compressors for Long Video Understanding
DMax: Aggressive Parallel Decoding for dLLMs
ViVa: A Video-Generative Value Model for Robot Reinforcement Learning
ELT: Elastic Looped Transformers for Visual Generation
PhysInOne: Visual Physics Learning and Reasoning in One Suite
Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability
KnowU-Bench: Towards Interactive, Proactive, and Personalized Mobile Agent Evaluation
Social Reality Construction via Active Inference: Modeling the Dialectic of Conformity and Creativity
Stochastic Thermodynamics for Autoregressive Generative Models: A Non-Markovian Perspective
Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models
Envisioning the Future, One Step at a Time
VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images
Gym-Anything: Turn any Software into an Agent Environment
HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation
The Illusion of Stochasticity in LLMs
Exponential quantum advantage in processing massive classical data
VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis
Self-Improving 4D Perception via Self-Distillation