LS-Imagine: Ending the "Short-Sightedness" of Visual RL via Jumpy Imagination
Open-World Reinforcement Learning over Long Short-Term Imagination
LS-Imagine is a novel model-based reinforcement learning (MBRL) framework designed for high-dimensional open-world tasks like Minecraft. It introduces a "Long Short-Term World Model" that enables agents to perform jumpy state transitions in imagination, significantly outperforming SOTA baselines like DreamerV3 in exploration efficiency and success rates.
TL;DR
In high-dimensional open worlds like Minecraft, visual reinforcement learning (RL) agents often struggle with exploration because their internal "dreams" (imaginations) are only a few steps long. LS-Imagine solves this by teaching a world model to "jump" across time. By combining standard short-term transitions with long-term goal-directed jumps, the agent can see further and act more decisively, achieving up to a 30% boost in success rates over previous SOTA models like DreamerV3.
The Short-Horizon Trap in Open-World RL
Open-world decision-making is notoriously difficult. Agents must process raw pixels, navigate vast spaces, and handle sparse rewards. Current Model-Based Reinforcement Learning (MBRL) methods, such as DreamerV3, rely on learning a "World Model" to predict the future. However, these models are typically optimized on short snippets of experience (e.g., 15 time steps).
If your target—say, a tree to be harvested—is 100 steps away, a 15-step imagination horizon is effectively blind to the goal. This "short-sightedness" leads to aimless wandering and inefficient exploration.
The Core Insight: Jumpy State Transitions
The authors of LS-Imagine argue that an agent shouldn't have to simulate every single micro-step to understand the value of a distant goal. Their solution? Long Short-Term Imagination.
1. Affordance Maps via "Virtual Zoom"
To identify where to jump, the agent needs to know what is "interesting" in its field of view. The researchers developed a way to generate Affordance Maps without needing expert data:
- Virtual Exploration: They take a single image and "zoom in" on different regions, simulating the visual experience of moving toward that spot.
- Goal Alignment: Using MineCLIP (a pre-trained vision-language model), they score these "zoom-ins" based on how well they match a text instruction like "cut a tree."
- Multimodal U-Net: To make this fast enough for real-time training, they train a U-Net to predict these affordance maps instantly from raw pixels and text.
2. The Integrated Architecture
The world model is split into two specialized pathways. A Jumping Flag () acts as a traffic controller:
- Short-term Branch: Predicts the immediate next state (standard dynamics).
- Long-term Branch: Predicts a "post-jump" state—the latent representation of the agent having reached a distant area of interest.
Figure: The dual-branch world model architecture showing short-term transitions and long-term jumps.
Behavior Learning: Mixing the Dreams
When the agent "dreams" during policy training, it now uses a mix of both branches. If the Jumping Flag is triggered (when a target is spotted), the model performs a long-term jump.
Crucially, the model also predicts:
- Interval (): How many real environmental steps the jump covers.
- Cumulative Reward (): The total reward expected during those skipped steps.
This allows the Critic to calculate a modified -return that accurately values the distant future, bridging the gap between current actions and long-term payoffs.
Experimental Results on MineDojo
The researchers tested LS-Imagine against a suite of strong baselines, including VPT, STEVE-1, and DreamerV3, across various Minecraft tasks.
Performance Gains
As shown in the charts below, LS-Imagine (red line) consistently reaches higher success rates faster than any other model.
- Harvesting Log: Reached ~80% success vs DreamerV3's ~53%.
- Efficiency: In "Shear Sheep," LS-Imagine completed the task in roughly 633 steps, significantly faster than the 841 steps required by the baseline.
Figure: Comparison of success rates across different MineDojo tasks.
Visualizing the "Jump"
One of the most impressive parts of the paper is the qualitative visualization. We can actually see the model's decoded latent states. Before the jump, the tree is a small green blur in the distance; after the jump, the model successfully imagines being right in front of the trunk.
Figure: Visualizing the sequence of imagination, including the sudden jump to the target.
Critical Analysis & Takeaways
Strengths:
- Efficiency: It avoids the "chicken-and-egg" problem of exploration by creating its own training data for jumps via image manipulation (zooming).
- Flexibility: The system works with raw pixels and natural language, making it applicable to human-centric task descriptions.
Limitations:
- Domain Specificity: The "zoom-in" heuristic is tailored for 3D navigation and embodied agents. It might not translate directly to a 2D Atari game or a task like stock trading where "zooming" has no physical meaning.
- Computational Cost: Training the additional U-Net and dual-pathway world model increases VRAM requirements (requiring ~23GB).
Conclusion
LS-Imagine proves that for an AI to conquer an open world, it doesn't just need to react—it needs to envision the destination. By allowing the world model to "jump" across the timeline of its own imagination, we can train agents that are far more strategic and drastically more efficient than those limited to one-step lookaheads.
