LOVON: Revolutionizing Legged Navigation via Open-Vocabulary Intelligence
LOVON: Legged Open-Vocabulary Object Navigator
LOVON (Legged Open-Vocabulary Object Navigator) is a novel framework for long-horizon object navigation in legged robots, integrating Large Language Models (LLMs) for task planning with open-vocabulary vision models. It achieves SOTA performance on the Gym-Unreal benchmark and successfully deploys on multiple hardware platforms (Unitree Go2, B2, H1-2) with a significant 240x reduction in training time compared to prior work.
TL;DR
Navigating unstructured environments is a "holy grail" for legged robots. LOVON (Legged Open-Vocabulary Object Navigator) bridges the gap between high-level reasoning and low-level stability. By combining DeepSeek R1 for task planning with a specialized Language-to-Motion Model (L2MM) and a Laplacian filtering technique, it achieves robust, long-horizon navigation. Remarkably, it delivers SOTA performance with a fraction of the training cost (1.5 hours vs. 360 hours for previous leaders).
The "Jitter" Problem: Why Legged Navigation is Hard
Most navigation research assumes a stable camera mount (like a wheeled tray). Legged robots (quadrupeds and humanoids) inherently oscillate during movement. This creates motion blur and visual jittering, which causes open-vocabulary detectors (like YOLO or Grounding DINO) to lose confidence or fail entirely. Furthermore, translating a vague command like "find my bag, then find the chair" into continuous motor commands requires a sophisticated hierarchy that most systems lack.
Methodology: The Hierarchical Operating System
LOVON functions as an "operating system" for navigation, split into three layers:
- High-Level Planning (LLM): Uses DeepSeek R1 to take a long-horizon description and break it into a sequence of subtasks.
- Instruction Object Extractor (IOE): A transformer-based module that maps linguistic subtasks to specific visual object classes.
- The L2MM Core: An encoder-decoder architecture that takes processed visual data and mission states to output velocity vectors ().

Visual Stability: Laplacian Variance Filtering
To fight motion blur, LOVON computes the variance of the Laplacian response for each frame. If a frame is too blurry (variance below a threshold ), it is discarded and replaced by the last clear frame. This simple yet effective physics-based intuition ensures the L2MM receives stable inputs even when the robot is running at 0.7 m/s.
Experimental Results: SOTA Performance at 1/240th the Cost
In the Gym-Unreal benchmark, LOVON achieved a perfect success rate (1.00) in nearly all scenes. While the previous leader, TrackVLA, required 360 hours of training on high-end GPUs, LOVON achieved comparable or better results in just 1.5 hours.

Real-World Robustness
The system was tested on Unitree Go2, B2, and H1-2 (humanoid). It handled:
- Dynamic Targets: Following a person or a moving ball.
- Disturbances: Re-calculating paths after being kicked or when the target was moved.
- Complex Terrains: Spiral stairs and wild grass fields.

Deep Insights: Why It Works
The success of LOVON boils down to two factors:
- State-Aware Logic: By explicitly defining "Searching" and "Running" states, the robot doesn't just "fail" when it loses an object; it enters a bi-directional rotation search pattern that dramatically increases recovery speed.
- Separation of Concerns: By decoupling the heavy lifting of class detection from the high-frequency control of the L2MM, the system remains responsive (real-time execution on Jetson Orin) while maintaining GPT-level intelligence in planning.
Conclusion and Future Outlook
LOVON provides a blueprint for making legged robots truly autonomous in human environments. By addressing the physical realities of robot locomotion (blur) alongside the cognitive requirements of task planning (LLMs), it creates a robust framework for embodied AI. Future iterations will likely integrate even more powerful Vision-Language Models (VLMs) to handle even more nuanced environmental semantic understanding.
Takeaway: Real-world navigation isn't just about "seeing" the goal; it's about maintaining a stable "memory" of the goal despite the chaos of motion.
