WisPaper
WisPaper
Search
Features
Resources
Pricing
Download
Workspace
Blog
No more endless PDFs. Discover the core value of the latest top-tier research in one article.
User Shared
Trends
3DrawAgent: Teaching LLM to Draw in 3D with Early Contrastive Experience
MolmoWeb: Open Visual Web Agent and Open Data for the Open Web
Learning Vision-Language-Action World Models for Autonomous Driving
A Decomposition Perspective to Long-context Reasoning for LLMs
OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering
ETCH-X: Robustify Expressive Body Fitting to Clothed Humans with Composable Datasets
EXAONE 4.5 Technical Report
PokeGym: A Visually-Driven Long-Horizon Benchmark for Vision-Language Models
Online3R: Online Learning for Consistent Sequential Reconstruction Based on Geometry Foundation Model
WorldMAP: Bootstrapping Vision-Language Navigation Trajectory Prediction with Generative World Models
Rays as Pixels: Learning A Joint Distribution of Videos and Camera Trajectories
Visually-grounded Humanoid Agents
MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding
INSPATIO-WORLD: A Real-Time 4D World Simulator via Spatiotemporal Autoregressive Modeling
Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics
SUPERNOVA: Eliciting General Reasoning in LLMs with Reinforcement Learning on Natural Instructions
Novel View Synthesis as Video Completion
DialogueSidon: Recovering Full-Duplex Dialogue Tracks from In-the-Wild Dialogue Audio
Nexus: Same Pretraining Loss, Better Downstream Generalization via Common Minima
LMGenDrive: Bridging Multimodal Understanding and Generative World Modeling for End-to-End Driving
When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models
Beyond Stochastic Exploration: What Makes Training Data Valuable for Agentic Search
Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest
CT-1: Vision-Language-Camera Models Transfer Spatial Reasoning Knowledge to Camera-Controllable Video Generation
Faithful GRPO: Improving Visual Spatial Reasoning in Multimodal Language Models via Constrained Policy Optimization
Claw-Eval: Toward Trustworthy Evaluation of Autonomous Agents
OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks
Scaling-Aware Data Selection for End-to-End Autonomous Driving Systems
Gravitational Memory from Hairy Binary Black Hole Mergers
From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models
←
1
...
45
46
47
...
69
→