Blog

告别冗长 PDF,一文读懂最新顶会与核心期刊的创新价值。

RIVER: A Real-Time Interaction Benchmark for Video LLMs
MEM: Multi-Scale Embodied Memory for Vision Language Action Models
Architecting Trust in Artificial Epistemic Agents
ManipulationNet: An Infrastructure for Benchmarking Real-World Robot Manipulation with Physical Skill Challenges and Embodied Multimodal Reasoning
Mozi: Governed Autonomy for Drug Discovery LLM Agents
Phys4D: Fine-Grained Physics-Consistent 4D Modeling from Video Diffusion
ArtHOI: Articulated Human-Object Interaction Synthesis by 4D Reconstruction from Video Priors
InfinityStory: Unlimited Video Generation with World Consistency and Character-Aware Shot Transitions
Pretrained Vision-Language-Action Models are Surprisingly Resistant to Forgetting in Continual Learning
Structural Action Transformer for 3D Dexterous Manipulation
Memex(RL): Scaling Long-Horizon LLM Agents via Indexed Experience Memory
EmbodiedSplat: Online Feed-Forward Semantic 3DGS for Open-Vocabulary 3D Scene Understanding
Gaussian Wardrobe: Compositional 3D Gaussian Avatars for Free-Form Virtual Try-On
Heterogeneous Agent Collaborative Reinforcement Learning
Beyond Pixel Histories: World Models with Persistent 3D State
Proact-VL: A Proactive VideoLLM for Real-Time AI Companions
EgoPoseFormer v2: Accurate Egocentric Human Motion Estimation for AR/VR
What Capable Agents Must Know: Selection Theorems for Robust Decision-Making under Uncertainty
T2S-Bench & Structure-of-Thought: Benchmarking and Prompting Comprehensive Text-to-Structure Reasoning
SWE-CI: Evaluating Agent Capabilities in Maintaining Codebases via Continuous Integration
MemSifter: Offloading LLM Memory Retrieval via Outcome-Driven Proxy Reasoning
X-Loco: Towards Generalist Humanoid Locomotion Control via Synergetic Policy Distillation
Transformer Neural-Network Quantum States for lattice models of spins and fermions: Application to the Ancilla Layer Model
AgentIR: Reasoning-Aware Retrival for Deep Research Agents
CubeComposer: Spatio-Temporal Autoregressive 4K 360° Video Generation from Perspective Video
MiM-DiT: MoE in MoE with Diffusion Transformers for All-in-One Image Restoration
Sleeping Beauty in One or Many Worlds: A Defense of the Halfer Position
CFG-Ctrl: Control-Based Classifier-Free Diffusion Guidance
NOVA3R: Non-pixel-aligned Visual Transformer for Amodal 3D Reconstruction
Generalized Bayes for Causal Inference