Blog

No more endless PDFs. Discover the core value of the latest top-tier research in one article.

Model Merging in the Era of Large Language Models: Methods, Applications, and Future Directions
Stepping VLMs onto the Court: Benchmarking Spatial Intelligence in Sports
Exclusive Self Attention
OpenClaw-RL: Train Any Agent Simply by Talking
DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot Control
Training Language Models via Neural Cellular Automata
Ranking Reasoning LLMs under Test-Time Scaling
Just-in-Time: Training-Free Spatial Acceleration for Diffusion Transformers
CLIPO: Contrastive Learning in Policy Optimization Generalizes RLVR
LLM2Vec-Gen: Generative Embeddings from Large Language Models
LiTo: Surface Light Field Tokenization
DynVLA: Learning World Dynamics for Action Reasoning in Autonomous Driving
CodePercept: Code-Grounded Visual STEM Perception for MLLMs
4DEquine: Disentangling Motion and Appearance for 4D Equine Reconstruction from Monocular Video
COMIC: Agentic Sketch Comedy Generation
In-Context Reinforcement Learning for Tool Use in Large Language Models
Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable Rewards
Flash-KMeans: Fast and Memory-Efficient Exact K-Means
World2Act: Latent Action Post-Training via Skill-Compositional World Models
The Coupling Within: Flow Matching via Distilled Normalizing Flows
FutureVLA: Joint Visuomotor Prediction for Vision-Language-Action Model
V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation
Cross-Hand Latent Representation for Vision-Language-Action Models
AR-VLA: True Autoregressive Action Expert for Vision-Language-Action Models
UniCom: Unified Multimodal Modeling via Compressed Continuous Semantic Representations
Geometric Autoencoder for Diffusion Models
$V_{0.5}$: Generalist Value Model as a Prior for Sparse RL Rollouts
SCALAR: Learning and Composing Skills through LLM Guided Symbolic Planning and Deep RL Grounding
Violating the All-or-Nothing Picture of Local Charges in Non-Hermitian Bosonic Chains
AlphaFlowTSE: One-Step Generative Target Speaker Extraction via Conditional AlphaFlow