LOVON: Revolutionizing Legged Navigation via Open-Vocabulary Intelligence

LOVON: Legged Open-Vocabulary Object Navigator

2025-01-01
Daojie Peng, Jiahang Cao, Qiang Zhang, Jun Ma
Summary
Problem
Method
Results
Takeaways
Abstract

LOVON (Legged Open-Vocabulary Object Navigator) is a novel framework for long-horizon object navigation in legged robots, integrating Large Language Models (LLMs) for task planning with open-vocabulary vision models. It achieves SOTA performance on the Gym-Unreal benchmark and successfully deploys on multiple hardware platforms (Unitree Go2, B2, H1-2) with a significant 240x reduction in training time compared to prior work.

TL;DR

Navigating unstructured environments is a "holy grail" for legged robots. LOVON (Legged Open-Vocabulary Object Navigator) bridges the gap between high-level reasoning and low-level stability. By combining DeepSeek R1 for task planning with a specialized Language-to-Motion Model (L2MM) and a Laplacian filtering technique, it achieves robust, long-horizon navigation. Remarkably, it delivers SOTA performance with a fraction of the training cost (1.5 hours vs. 360 hours for previous leaders).

The "Jitter" Problem: Why Legged Navigation is Hard

Most navigation research assumes a stable camera mount (like a wheeled tray). Legged robots (quadrupeds and humanoids) inherently oscillate during movement. This creates motion blur and visual jittering, which causes open-vocabulary detectors (like YOLO or Grounding DINO) to lose confidence or fail entirely. Furthermore, translating a vague command like "find my bag, then find the chair" into continuous motor commands requires a sophisticated hierarchy that most systems lack.

Methodology: The Hierarchical Operating System

LOVON functions as an "operating system" for navigation, split into three layers:

  1. High-Level Planning (LLM): Uses DeepSeek R1 to take a long-horizon description and break it into a sequence of subtasks.
  2. Instruction Object Extractor (IOE): A transformer-based module that maps linguistic subtasks to specific visual object classes.
  3. The L2MM Core: An encoder-decoder architecture that takes processed visual data and mission states to output velocity vectors ().

LOVON Architecture

Visual Stability: Laplacian Variance Filtering

To fight motion blur, LOVON computes the variance of the Laplacian response for each frame. If a frame is too blurry (variance below a threshold ), it is discarded and replaced by the last clear frame. This simple yet effective physics-based intuition ensures the L2MM receives stable inputs even when the robot is running at 0.7 m/s.

Experimental Results: SOTA Performance at 1/240th the Cost

In the Gym-Unreal benchmark, LOVON achieved a perfect success rate (1.00) in nearly all scenes. While the previous leader, TrackVLA, required 360 hours of training on high-end GPUs, LOVON achieved comparable or better results in just 1.5 hours.

Performance Comparison Table

Real-World Robustness

The system was tested on Unitree Go2, B2, and H1-2 (humanoid). It handled:

  • Dynamic Targets: Following a person or a moving ball.
  • Disturbances: Re-calculating paths after being kicked or when the target was moved.
  • Complex Terrains: Spiral stairs and wild grass fields.

Threshold and Qualified Ratio

Deep Insights: Why It Works

The success of LOVON boils down to two factors:

  • State-Aware Logic: By explicitly defining "Searching" and "Running" states, the robot doesn't just "fail" when it loses an object; it enters a bi-directional rotation search pattern that dramatically increases recovery speed.
  • Separation of Concerns: By decoupling the heavy lifting of class detection from the high-frequency control of the L2MM, the system remains responsive (real-time execution on Jetson Orin) while maintaining GPT-level intelligence in planning.

Conclusion and Future Outlook

LOVON provides a blueprint for making legged robots truly autonomous in human environments. By addressing the physical realities of robot locomotion (blur) alongside the cognitive requirements of task planning (LLMs), it creates a robust framework for embodied AI. Future iterations will likely integrate even more powerful Vision-Language Models (VLMs) to handle even more nuanced environmental semantic understanding.

Takeaway: Real-world navigation isn't just about "seeing" the goal; it's about maintaining a stable "memory" of the goal despite the chaos of motion.

Find Similar Papers

Try Our Examples

  • Search for recent papers that address visual motion blur mitigation specifically for legged robot perception using deep learning or filter-based methods.
  • What are the primary differences in architecture between the Language-to-Motion Model (L2MM) used in LOVON and original Vision-Language-Action (VLA) models like RT-2?
  • Explore current research on using DeepSeek R1 or other reasoning-focused LLMs for real-time hierarchical task decomposition in embodied AI.
Contents
LOVON: Revolutionizing Legged Navigation via Open-Vocabulary Intelligence
1. TL;DR
2. The "Jitter" Problem: Why Legged Navigation is Hard
3. Methodology: The Hierarchical Operating System
3.1. Visual Stability: Laplacian Variance Filtering
4. Experimental Results: SOTA Performance at 1/240th the Cost
4.1. Real-World Robustness
5. Deep Insights: Why It Works
6. Conclusion and Future Outlook