SOLVER: Decoding the Relational Language of Visual Emotions
10931_SOLVER: Scene-Object Interrelated Visual Emotion R
The paper introduces SOLVER, a Scene-Object interreLated Visual Emotion Reasoning network for Visual Emotion Analysis (VEA). It utilizes an Emotion Graph with Graph Convolutional Networks (GCN) to model object-object interactions and a scene-based attention mechanism to integrate scene-object relationships, achieving SOTA performance across eight public datasets.
TL;DR
Human emotions are rarely triggered by a single object in isolation. Instead, our feelings are evoked by how objects interact with each other and their environment. SOLVER (Scene-Object interreLated Visual Emotion Reasoning network) is a novel architecture that mimics this cognitive process by using Graph Convolutional Networks (GCN) and Scene-based Attention to bridge the "affective gap" in image analysis.
The Problem: The Affective Gap
In the field of Visual Emotion Analysis (VEA), most deep learning models treat an image as a bag of pixels or a collection of isolated regions. They map these features directly to emotion labels like "Happy" or "Sad." However, psychology tells us that emotion is contextual.
For example, a "red rose" might evoke Contentment at a wedding but Sadness at a funeral. The object is the same, but the interaction with other objects and the scene changes the emotional output. Previous models struggled with this nuance, leading to a performance bottleneck known as the affective gap.
Methodology: Reasoning over Graphs and Scenes
SOLVER addresses this by treating an image as a structured system of interactions.
1. The Emotion Graph (Object-Object Interaction)
The model first detects objects using Faster R-CNN. It then builds an Emotion Graph where:
- Nodes: Represent semantic concepts (e.g., "dog," "mountain") using GloVe embeddings.
- Edges: Represent the emotional relationship (affinity) between these objects, calculated using visual features in an emotional embedding space.
By applying Graph Convolutional Networks (GCN), the model allows object features to "communicate" with their neighbors. A "rose" node learns about the "bride" node nearby, enhancing its feature representation with relational context.

2. Scene-Object Fusion (Scene-Object Interaction)
While objects provide local triggers, the scene sets the global "tone." SOLVER uses a Scene-Object Fusion Module. It doesn't just concatenate features; it uses the global scene feature as a "query" to weight the importance of different objects. This scene-based attention ensures that if the scene is a "stadium," the model pays more attention to "players" and "crowds" to infer excitement.

Experimental Performance
The researchers tested SOLVER against a battery of SOTA methods (including WSCNet and MldrNet) across 8 datasets.
- Large-scale datasets: Achieved a significant lead (e.g., 72.33% on FI, 86.20% on Flickr).
- Ablation Study: The results showed that adding the Emotion Graph + Fusion module improved performance significantly over using just a standard ResNet-50 backbone (which scored 67.53% on FI).

Deep Insight: Interpretable AI
One of the most impressive aspects of the SOLVER paper is its interpretability. By visualizing Weighted Frequencies, the authors show which concepts actually drive specific emotions.
- Awe is consistently linked to "mountains," "cliffs," and "horizons."
- Excitement correlates with "surfboards," "rafts," and "microphones."
The attention maps demonstrate that the model "looks" at the most emotionally relevant interactions—such as the bared teeth of a leopard when predicting "Anger"—rather than just the animal's body.

Conclusion & Limitations
SOLVER proves that relational reasoning is the key to mastering high-level cognitive tasks like emotion analysis. However, the authors honestly note a limitation: the model currently misses micro-expressions (facial cues) and body language, which are vital for human-centric datasets like LUCFER. Future iterations will likely integrate these "human-centric" features into the existing scene-object graph to create a truly holistic emotional AI.
