Blog

No more endless PDFs. Discover the core value of the latest top-tier research in one article.

Truth as a Trajectory: What Internal Representations Reveal About Large Language Model Reasoning
FluxMem: Adaptive Hierarchical Memory for Streaming Video Understanding
Suffix-Constrained Greedy Search Algorithms for Causal Language Models
Speculative Speculative Decoding
Chain of World: World Model Thinking in Latent Motion
CuTe Layout Representation and Algebra
How Well Does Agent Development Reflect Real-World Work?
OneRanker: Unified Generation and Ranking with One Model in Industrial Advertising Recommendation
UniG2U-Bench: Do Unified Models Advance Multimodal Understanding?
ACE-Brain-0: Spatial Intelligence as a Shared Scaffold for Universal Embodiments
Kling-MotionControl Technical Report
HoMMI: Learning Whole-Body Mobile Manipulation from Human Demonstrations
ULTRA: Unified Multimodal Control for Autonomous Humanoid Whole-Body Loco-Manipulation
Next Embedding Prediction Makes World Models Stronger
NOVA: Sparse Control, Dense Synthesis for Pair-Free Video Editing
PRISM: Pushing the Frontier of Deep Think via Process Reward Model-Guided Inference
Track4World: Feedforward World-centric Dense 3D Tracking of All Pixels
LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory
DuoMo: Dual Motion Diffusion for World-Space Human Reconstruction
Expanding LLM Agent Boundaries with Strategy-Guided Exploration
Rhythm: Learning Interactive Whole-Body Control for Dual Humanoids
Diagnosing Retrieval vs. Utilization Bottlenecks in LLM Agent Memory
BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?
Think-as-You-See: Streaming Chain-of-Thought Reasoning for Large Vision-Language Models
Modular Memory is the Key to Continual Learning Agents
Surgical Post-Training: Cutting Errors, Keeping Knowledge
DREAM: Where Visual Understanding Meets Text-to-Image Generation
Improving Diffusion Planners by Self-Supervised Action Gating with Energies
DeepResearch-9K: A Challenging Benchmark Dataset of Deep-Research Agent
ShareVerse: Multi-Agent Consistent Video Generation for Shared World Modeling