Blog

No more endless PDFs. Discover the core value of the latest top-tier research in one article.

WorldMesh: Generating Navigable Multi-Room 3D Scenes via Mesh-Conditioned Image Diffusion
UniScale: Synergistic Entire Space Data and Model Scaling for Search Ranking
SpatialReward: Verifiable Spatial Reward Modeling for Fine-Grained Spatial Consistency in Text-to-Image Generation
3D-Mix for VLA: A Plug-and-Play Module for Integrating VGGT-based 3D Information into Vision-Language-Action Models
Sparser, Faster, Lighter Transformer Language Models
LensWalk: Agentic Video Understanding by Planning How You See in Videos
DreamerAD: Efficient Reinforcement Learning via Latent World Model for Autonomous Driving
MemCollab: Cross-Agent Memory Collaboration via Contrastive Trajectory Distillation
VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models
Local Bernstein theory, and lower bounds for Lebesgue constants
Model Predictive Control with Differentiable World Models for Offline Reinforcement Learning
Dual-Teacher Distillation with Subnetwork Rectification for Black-Box Domain Adaptation
Finite-Degree Quantum LDPC Codes Reaching the Gilbert-Varshamov Bound
VFIG: Vectorizing Complex Figures in SVG with Vision-Language Models
GenMask: Adapting DiT for Segmentation via Direct Mask
Composer 2 Technical Report
Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning
CAM3R: Camera-Agnostic Model for 3D Reconstruction
ScrollScape: Unlocking 32K Image Generation With Video Diffusion Priors
Toward Physically Consistent Driving Video World Models under Challenging Trajectories
PERMA: Benchmarking Personalized Memory Agents via Event-Driven Preference and Realistic Task Environments
The Specification Gap: Coordination Failure Under Partial Knowledge in Code Agents
ProcureGym: A Multi-Agent Markov Game Framework for Modeling National Volume-based Drug Procurement
Causal Evidence that Language Models use Confidence to Drive Behavior
Not All Tokens Are Created Equal: Query-Efficient Jailbreak Fuzzing for LLMs
Spectral Scalpel: Amplifying Adjacent Action Discrepancy via Frequency-Selective Filtering for Skeleton-Based Action Segmentation
ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers
HGGT: Robust and Flexible 3D Hand Mesh Reconstruction from Uncalibrated Images
UW-VOS: A Large-Scale Dataset for Underwater Video Object Segmentation
KARMA: Knowledge-Action Regularized Multimodal Alignment for Personalized Search at Taobao