Blog

No more endless PDFs. Discover the core value of the latest top-tier research in one article.

Anny-Fit: All-Age Human Mesh Recovery
Flow Sampling: Learning to Sample from Unnormalized Densities via Denoising Conditional Processes
Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior
LongSeeker: Elastic Context Orchestration for Long-Horizon Search Agents
Lightning Unified Video Editing via In-Context Sparse Attention
ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation
SymptomAI: Towards a Conversational AI Agent for Everyday Symptom Assessment
cotomi Act: Learning to Automate Work by Watching You
Information Theory and Statistical Learning
Fast, accurate, high-resolution simulation of large-scale Fermi-Hubbard models on a digital quantum processor
Potential Hessian Ascent III: Sampling the Sherrington--Kirkpatrick Model at Beta < 1/2
Dream-MPC: Gradient-Based Model Predictive Control with Latent Imagination
Reinforcement Learning for LLM-based Multi-Agent Systems through Orchestration Traces
Segmenting Human-LLM Co-authored Text via Change Point Detection
Information-Geometric Signatures of Nonconservative Driving
Nonexistence of Whirling-Knight Tours at Half Coil Count for $n \equiv 4, 6 \pmod 8$
DINO Soars: DINOv3 for Open-Vocabulary Semantic Segmentation of Remote Sensing Imagery
Structured 3D Latents Are Surprisingly Powerful: Unleashing Generalizable Style with 2D Diffusion
Steer Like the LLM: Activation Steering that Mimics Prompting
From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models
StateVLM: A State-Aware Vision-Language Model for Robotic Affordance Reasoning
Learning Time-Inhomogeneous Markov Dynamics in Financial Time Series via Neural Parameterization
KernelBench-X: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels
Entropic Riemannian Neural Optimal Transport
HDFlow: Hierarchical Diffusion-Flow Planning for Long-horizon Tasks
RAG over Thinking Traces Can Improve Reasoning Tasks
Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs
Open-Source Image Editing Models Are Zero-Shot Vision Learners
Consistent Scattering Amplitudes, Yang-Mills, the Higgs Mechanism and the EFTs Beyond
Rose-SQL: Role-State Evolution Guided Structured Reasoning for Multi-Turn Text-to-SQL