Blog

告别冗长 PDF,一文读懂最新顶会与核心期刊的创新价值。

FDIF: Formula-Driven supervised Learning with Implicit Functions for 3D Medical Image Segmentation
MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion Decoding
ABot-PhysWorld: Interactive World Foundation Model for Robotic Manipulation with Physics Alignment
SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning
VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs
TETO: Tracking Events with Teacher Observation for Motion Estimation and Frame Interpolation
Group Editing : Edit Multiple Images in One Go
3DCity-LLM: Empowering Multi-modality Large Language Models for 3D City-scale Perception and Understanding
DA-Flow: Degradation-Aware Optical Flow Estimation with Diffusion Models
Demystifying Reinforcement Learning for Long-Horizon Tool-Using Agents: A Comprehensive Recipe
Foveated Diffusion: Efficient Spatially Adaptive Image and Video Generation
CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detection
EVA: Efficient Reinforcement Learning for End-to-End Video Agent
OccAny: Generalized Unconstrained Urban 3D Occupancy
Graph Puzzles II.1: Counterexamples to Jain's Second Unit Vector Flows Conjecture
Off-Policy Value-Based Reinforcement Learning for Large Language Models
CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation
SIMART: Decomposing Monolithic Meshes into Sim-ready Articulated Assets via MLLM
Holistic Scaling Laws for Optimal Mixture-of-Experts Architecture Optimization
ViKey: Enhancing Temporal Understanding in Videos via Visual Prompting
Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos
Reasoning over Semantic IDs Enhances Generative Recommendation
AI Agents Can Already Autonomously Perform Experimental High Energy Physics
Phases of itinerant anyons in Laughlin's quantum Hall states on a lattice
Group3D: MLLM-Driven Semantic Grouping for Open-Vocabulary 3D Object Detection
The Fate of the Milky Way--Andromeda System: To Merge or Not?
A Conformal Bridge for the Light Transform of QCD Correlation Functions
SG-VLA: Learning Spatially-Grounded Vision-Language-Action Models for Mobile Manipulation
UniQueR: Unified Query-based Feedforward 3D Reconstruction
DualCoT-VLA: Visual-Linguistic Chain of Thought via Parallel Reasoning for Vision-Language-Action Models