Blog

No more endless PDFs. Discover the core value of the latest top-tier research in one article.

Evidence of an Emergent "Self" in Continual Robot Learning
One View Is Enough! Monocular Training for In-the-Wild Novel View Generation
ViHOI: Human-Object Interaction Synthesis with Visual Priors
Compression is all you need: Modeling Mathematics
Strong-to-Weak Spontaneous Symmetry Breaking in a $(2+1)$D Transverse-Field Ising Model under Decoherence
Towards Training-Free Scene Text Editing
DA-VAE: Plug-in Latent Compression for Diffusion via Detail Alignment
VOLMO: Versatile and Open Large Models for Ophthalmology
Gaze-Regularized Vision-Language-Action Models for Robotic Manipulation
SOMA: Strategic Orchestration and Memory-Augmented System for Vision-Language-Action Model Robustness via In-Context Adaptation
Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs
TAG: Target-Agnostic Guidance for Stable Object-Centric Inference in Vision-Language-Action Models
ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling
Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale
Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
MACRO: Advancing Multi-Reference Image Generation with Structured Long-Context Data
CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use Agents
Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time Performance
Calibri: Enhancing Diffusion Transformers via Parameter-Efficient Calibration
PixelSmile: Toward Fine-Grained Facial Expression Editing
S2D2: Fast Decoding for Diffusion LLMs via Training-Free Self-Speculation
Many-body Josephson diode effect in superconducting quantum interferometers
Natural-Language Agent Harnesses
MuRF: Unlocking the Multi-Scale Potential of Vision Foundation Models
Chameleon: Episodic Memory for Long-Horizon Robotic Manipulation
PackForcing: Short Video Training Suffices for Long Video Sampling and Long Context Inference
Less Gaussians, Texture More: 4K Feed-Forward Textured Splatting
Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills
Vega: Learning to Drive with Natural Language Instructions
RealRestorer: Towards Generalizable Real-World Image Restoration with Large-Scale Image Editing Models