Blog

告别冗长 PDF,一文读懂最新顶会与核心期刊的创新价值。

Chain of Event-Centric Causal Thought for Physically Plausible Video Generation
SpaceSense-Bench: A Large-Scale Multi-Modal Benchmark for Spacecraft Perception and Pose Estimation
RAE-NWM: Navigation World Model in Dense Visual Representation Space
Covenant-72B: Pre-Training a 72B LLM with Trustless Peers Over-the-Internet
PIRA-Bench: A Transition from Reactive GUI Agents to GUI-based Proactive Intent Recommendation Agents
Impermanent: A Live Benchmark for Temporal Generalization in Time Series Forecasting
DexHiL: A Human-in-the-Loop Framework for Vision-Language-Action Model Post-Training in Dexterous Manipulation
StreamReady: Learning What to Answer and When in Long Streaming Videos
From Semantics to Pixels: Coarse-to-Fine Masked Autoencoders for Hierarchical Visual Understanding
FrameDiT: Diffusion Transformer with Frame-Level Matrix Attention for Efficient Video Generation
Task Aware Modulation Using Representation Learning for Upsaling of Terrestrial Carbon Fluxes
EvoDriveVLA: Evolving Autonomous Driving Vision-Language-Action Model via Collaborative Perception-Planning Distillation
OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning
PureCC: Pure Learning for Text-to-Image Concept Customization
Revealing Behavioral Plasticity in Large Language Models: A Token-Conditional Perspective
ImprovedGS+: A High-Performance C++/CUDA Re-Implementation Strategy for 3D Gaussian Splatting
Instanton construction of the mapping cone Thom-Smale complex
Learning Next Action Predictors from Human-Computer Interaction
Skip to the Good Part: Representation Structure & Inference-Time Layer Skipping in Diffusion vs. Autoregressive LLMs
Training-free Motion Factorization for Compositional Video Generation
Towards Unified Multimodal Interleaved Generation via Group Relative Policy Optimization
MSSR: Memory-Aware Adaptive Replay for Continual LLM Fine-Tuning
Entanglement Measure Response to Modular Flow and Chiral Topological Phases
Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR
How Contrastive Decoding Enhances Large Audio Language Models?
3ViewSense: Spatial and Mental Perspective Reasoning from Orthographic Views in Vision-Language Models
AutoAgent: Evolving Cognition and Elastic Memory Orchestration for Adaptive Agents
3PoinTr: 3D Point Tracks for Robot Manipulation Pretraining from Casual Videos
Speeding Up the Learning of 3D Gaussians with Much Shorter Gaussian Lists
ICLR: In-Context Imitation Learning with Visual Reasoning