WisPaper
WisPaper
Search
Features
Resources
Pricing
Download
Workspace
Blog
No more endless PDFs. Discover the core value of the latest top-tier research in one article.
User Shared
Trends
Mamoda2.5: Enhancing Unified Multimodal Model with DiT-MoE
RLDX-1 Technical Report
HeavySkill: Heavy Thinking as the Inner Skill in Agentic Harness
A Theory of Generalization in Deep Learning
T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning
OpenSeeker-v2: Pushing the Limits of Search Agents with Informative and High-Difficulty Trajectories
Representation Fréchet Loss for Visual Generation
World Model for Robot Learning: A Comprehensive Survey
ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration
On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length
Perceptual Flow Network for Visually Grounded Reasoning
Learning to Theorize the World from Observation
Audio-Visual Intelligence in Large Foundation Models
Mix3R: Mixing Feed-forward Reconstruction and Generative 3D Priors for Joint Multi-view Aligned 3D Reconstruction and Pose Estimation
A Benchmark for Interactive World Models with a Unified Action Generation Framework
Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning
Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe
Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs
Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL
UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors
Linearizing Vision Transformer with Test-Time Training
Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation
MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model
ProgramBench: Can Language Models Rebuild Programs From Scratch?
From Context to Skills: Can Language Models Learn from Context Skillfully?
RoboAlign-R1: Distilled Multimodal Reward Alignment for Robot Video World Models
Large Language Models are Universal Reasoners for Visual Generation
Model Merging: Foundations and Algorithms
LLM-Emu: Native Runtime Emulation of LLM Inference via Profile-Driven Sampling
Embody4D: A Generalist 4D World Model for Embodied AI
←
1
...
58
59
60
...
69
→