Emotion-Motivated Robots: Solving the Sparse Reward Puzzle with Brain-Inspired Computing
Computational Modeling of Emotion-Motivated Decisions for Continuous Control of Mobile Robots
The paper proposes a brain-inspired emotion-motivated decision-making framework for mobile robots. It combines a computational model of the amygdala-hippocampus interaction to generate intrinsic emotional rewards (valence, novelty, and motivational relevance) with a hybrid Model-Based (MB) and Model-Free (MF) control architecture.
TL;DR
Researchers have developed a new framework that gives mobile robots a form of "artificial emotion" to help them learn in environments where feedback is nearly non-existent. By mimicking the interaction between the brain's amygdala (emotion center) and hippocampus (memory center), robots can generate their own internal rewards based on novelty and motivation, allowing them to solve complex navigation tasks that leave standard AI stumped.
The Problem: The "Desert" of Sparse Rewards
In ideal AI training, a robot gets a "treat" (reward) for every small correct move. In the real world, however, rewards are rare. Imagine a robot in a massive warehouse that only gets a signal when it finds one specific small box. Without frequent feedback, standard Model-Free (MF) algorithms like DDPG wander aimlessly, never stumbling upon the goal.
The authors argue that humans don't have this problem because our emotions reshape these sparse external signals into a dense stream of internal "subjective values."
Methodology: The Amygdala-Hippocampus Loop
The core innovation lies in a biologically plausible computational model that simulates the neural circuits of the human brain.
1. Generating Artificial Emotion
The system uses "shunting dynamics" to simulate four key amygdalar nuclei (LA, BA, CeM, and ITCs). It processes three specific psychological dimensions:
- Valence: Replicating the external reward (is this good or bad?).
- Novelty: Driven by episodic memory—if a state is "new," it generates a curiosity-like emotional spike.
- Motivational Relevance: Determining if the current position is near a previously discovered "goal" stored in memory.
Figure 4: The schematic shows how memory-based information (EHip) from the hippocampus interacts with amygdala nuclei to produce a final emotional response (xCeM).
2. Emotion-Motivated Decisions
The framework integrates two types of thinking:
- Model-Free (Habitual): A policy trained offline to act quickly.
- Model-Based (Goal-Directed): An online Model Predictive Control (MPC) that "imagines" future trajectories.
The "magic" happens when the online MB controller uses the intrinsic emotional rewards to evaluate its imagined paths, while being "tethered" to the MF policy via a KL-divergence constraint to prevent it from making wildly unrealistic choices.
Experimental Results: Breaking the Limits
The authors tested their agent in a "Very Sparse" navigation task where the robot only receives a reward upon reaching the goal—no "getting warmer" signals allowed.
- Baselines Fail: Standard DDPG couldn't find the goal once.
- Emotion Succeeds: The emotion-motivated agents (both MF and MB) successfully learned the task.
- Dynamic Adaptation: In tasks where the goal moved randomly, the Model-Based (MB) version was superior. It could evaluate "imaginary" paths using its emotional memory, finding the new goal location much faster than even advanced curiosity-based models like ICM.
Figure 7: The complete workflow showing how external sparse rewards are transformed into dense internal emotional responses for better policy training.
Critical Insight: Why Does This Work?
Traditional "curiosity" models often just look for surprise (prediction error). However, this paper's Motivational Relevance component acts as a "biological compass." Once a robot finds a goal once, its memory keeps it emotionally "excited" about areas near that goal, turning a one-time lucky find into a consistent learning signal.
Conclusion & Future Outlook
This work moves us closer to "self-aware" robots. By moving away from rigid, designer-defined reward functions and toward subjective, emotion-driven values, we enable robots to explore more like animals and less like calculators.
The primary downside is computational cost: "imagining" trajectories online is slow (taking ~0.6s to 1s per step). However, as edge computing power grows, this "emotional" architecture could become the standard for robots operating in unpredictable, real-world environments.
