LS-Imagine: Defeating Short-Sightedness in Open-World Reinforcement Learning
Open-World Reinforcement Learning over Long Short-Term Imagination
LS-Imagine is a novel Model-Based Reinforcement Learning (MBRL) framework designed for open-world visual control tasks like Minecraft. It achieves state-of-the-art success rates on MineDojo tasks by introducing "jumpy" state transitions and a long short-term world model that extends the agent's imagination horizon.
TL;DR
Visual Reinforcement Learning (RL) agents often struggle in open-world environments like Minecraft because their "imagination" is limited to the immediate future. LS-Imagine solves this by allowing agents to "jump" through time in their latent world models. By combining a Long Short-Term World Model with Affordance Maps, it reaches distant goals faster and more reliably, significantly outperforming baselines like DreamerV3 and VPT in the MineDojo benchmark.
The "Short-Sighted" Agent Problem
Modern Model-Based RL (MBRL) agents, such as DreamerV3, learn by simulating future trajectories in a latent space. While powerful, these simulations are usually restricted to short snippets (e.g., 15 steps). In an open-world setting where a tree might be hundreds of steps away, a 15-step lookahead is practically blind.
The core challenge is exploration efficiency. Without a way to bridge the gap between "here" and "there," agents wander aimlessly in the vast state space of raw pixels, failing to encounter the sparse rewards necessary for learning.
Methodology: High-Speed Jumps in Latent Space
The researchers introduced Long Short-Term Imagination (LS-Imagine), which fundamentally changes how an agent "dreams."
1. The Long Short-Term World Model
Unlike standard models that only predict from , LS-Imagine features two branches:
- Short-term branch: Predicts the immediate next state.
- Long-term branch: Simulates "jumpy" transitions to a pivotal future state where a goal object is closer.
2. Affordance Maps via "Virtual Exploration"
To know where to jump, the agent needs a spatial prior. The authors developed a method to generate Affordance Maps by:
- Scanning the current image with a sliding window.
- "Zooming in" on these windows to simulate the agent moving closer.
- Using a pretrained MineCLIP model to score the relevance of these "zoomed" clips to the text instruction (e.g., "cut a tree").
Figure 1: The general framework of LS-Imagine, enabling policy optimization in a joint space of short and long-term imaginations.
3. Jumpy State Transitions
When the affordance map indicates a high-value target is in view (measured by Kurtosis), a "jumping flag" is triggered. The model then predicts:
- The post-jump state (what it looks like near the target).
- The interval (how many real steps it takes to get there).
- The cumulative reward expected during that jump.
Results: Efficiency in Action
LS-Imagine was tested on five challenging MineDojo tasks. The results show a dramatic improvement in both success rates and efficiency.
Figure 4: Performance comparison. LS-Imagine (red) reaches high success rates much faster than DreamerV3 (blue).
- Sparse Target Advantage: In tasks like "Harvest log," LS-Imagine achieved an 80.6% success rate, whereas the previous SOTA, DreamerV3, hit only 53.3%.
- Reduced Path Length: The agent doesn't just succeed more often; it succeeds faster. By "imagining" the long-term value, the learned policy is more direct, reducing the steps per episode by roughly 30%.
Critical Insight: Series vs. Parallel Imagination
A fascinating finding in the paper is the comparison between Series and Parallel imagination pathways.
- Parallel: Jumping to a new sequence but keeping it independent.
- Series: Integrating the jump directly into the sequence used for value estimation.
The Series approach (used in LS-Imagine) allows the gradients and value estimates from the "post-jump" future to flow back and inform the "pre-jump" actions. This is why the agent learns to orient itself toward targets even before it makes the jump.
Conclusion & Perspectives
LS-Imagine proves that we don't necessarily need more data or bigger models to solve open-world RL; we need smarter imagination. By encoding temporal leaps and spatial affordances into the world model, the agent gains a "telephoto lens" on its future.
Limitations: The current zoom-in heuristic for affordance is tailored to 3D navigation. Future work will need to generalize this to tasks like "driving" or "manipulation" where a simple zoom doesn't represent reaching a goal.
As an editor's note: This work bridges the gap between high-level planning (like LLM-based agents) and low-level visual control, providing a purely RL-based solution to long-horizon problems.
