Blog

No more endless PDFs. Discover the core value of the latest top-tier research in one article.

Synthetic Mixed Training: Scaling Parametric Knowledge Acquisition Beyond RAG
Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models
Bound states of anyons: a geometric quantization approach
RefAlign: Representation Alignment for Reference-to-Video Generation
Wan-Weaver: Interleaved Multi-modal Generation via Decoupled Training
Self-Distillation for Multi-Token Prediction
Confidence-Based Mesh Extraction from 3D Gaussians
Probing Interacting Dark Sectors with upcoming Post-Reionization and Galaxy Surveys
BizGenEval: A Systematic Benchmark for Commercial Visual Content Generation
Towards Embodied AI with MuscleMimic: Unlocking full-body musculoskeletal motor learning at scale
From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents
ColBERT-Att: Late-Interaction Meets Attention for Enhanced Retrieval
SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
Sign Errors in "The Four Laws of Black Hole Mechanics"
Persistent Robot World Models: Stabilizing Multi-Step Rollouts via Reinforcement Learning
Understanding the Challenges in Iterative Generative Optimization with LLMs
A Deep Dive into Scaling RL for Code Generation with Synthetic Data and Curricula
What happens to wavepackets of fermions when scattered by the Maldacena-Ludwig wall?
Vision Hopfield Memory Networks
Notes on Diagrammatic Coaction for Cosmological Wavefunction Coefficients: A Two-Site Prelude
AnyHand: A Large-Scale Synthetic Dataset for RGB(-D) Hand Pose Estimation
BiFM: Bidirectional Flow Matching for Few-Step Image Editing and Generation
Beyond Matching to Tiles: Bridging Unaligned Aerial and Satellite Views for Vision-Only UAV Navigation
Drive My Way: Preference Alignment of Vision-Language-Action Model for Personalized Driving
LanteRn: Latent Visual Structured Reasoning
Slow-down of expanding bubbles in the early Universe
SlotVTG: Object-Centric Adapter for Generalizable Video Temporal Grounding
R-C2: Cycle-Consistent Reinforcement Learning Improves Multimodal Reasoning
Microergodicity implies orthogonality of Matérn fields on bounded domains in $\mathbb{R}^4$
DCARL: A Divide-and-Conquer Framework for Autoregressive Long-Trajectory Video Generation