Blog

No more endless PDFs. Discover the core value of the latest top-tier research in one article.

LagerNVS: Latent Geometry for Fully Neural Real-time Novel View Synthesis
A Subgoal-driven Framework for Improving Long-Horizon LLM Agents
Beyond Single Tokens: Distilling Discrete Diffusion Models via Discrete MMD
How Well Does Generative Recommendation Generalize?
WorldAgents: Can Foundation Image Models be Agents for 3D World Models?
The $\mathbf{Y}$-Combinator for LLMs: Solving Long-Context Rot with $λ$-Calculus
Hyperagents
EgoForge: Goal-Directed Egocentric World Simulator
Kolmogorov-Arnold causal generative models
StreetForward: Perceiving Dynamic Street with Feedforward Causal Attention
VAMPO: Policy Optimization for Improving Visual Dynamics in Video Action Models
LumosX: Relate Any Identities with Their Attributes for Personalized Video Generation
DynFlowDrive: Flow-Based Dynamic World Modeling for Autonomous Driving
MME-CoF-Pro: Evaluating Reasoning Coherence in Video Generative Models with Text and Visual Hints
ReLi3D: Relightable Multi-view 3D Reconstruction with Disentangled Illumination
Plug-and-Steer: Decoupling Separation and Selection in Audio-Visual Target Speaker Extraction
What If Consensus Lies? Selective-Complementary Reinforcement Learning at Test Time
AGILE: A Comprehensive Workflow for Humanoid Loco-Manipulation Learning
From Masks to Pixels and Meaning: A New Taxonomy, Benchmark, and Metrics for VLM Image Tampering
MagicSeg: Open-World Segmentation Pretraining via Counterfactural Diffusion-Based Auto-Generation
Listen First, Then Answer: Timestamp-Grounded Speech Reasoning
Teaching an Agent to Sketch One Part at a Time
MedSPOT: A Workflow-Aware Sequential Grounding Benchmark for Clinical GUI
Casimir-Induced Quintessence in Dark Dimension
OrbitNVS: Harnessing Video Diffusion Priors for Novel View Synthesis
Serendipity by Design: Evaluating the Impact of Cross-domain Mappings on Human and LLM Creativity
Advances in the Worldline Approach to Quantum Field Theory: Strong Fields, Amplitudes and Gravity
CoVR-R:Reason-Aware Composed Video Retrieval
Four Fermi Theory in Four Dimensions is Renormalisable
How Auditory Knowledge in LLM Backbones Shapes Audio Language Models: A Holistic Evaluation