Deep Sentiment: Decoding the Multi-Layered Structure of Emotions via Tumblr
Multimodal Sentiment Analysis To Explore the Structure of Emotions
2018-07-19
Summary
Problem
Method
Results
Takeaways
Abstract
This paper introduces "Deep Sentiment," a multimodal neural network for sentiment analysis that combines visual features from Inception and textual embeddings via LSTM. Unlike standard binary polarity classification, the model predicts latent emotional states across 15 categories, achieving 72% test accuracy on a massive, novel Tumblr dataset.
## TL;DR
Researchers from Oxford and Imperial College London have developed **Deep Sentiment**, a multimodal deep learning framework that decodes human emotions by analyzing the synergy between images and text. By training on over 250,000 Tumblr posts, the model moves beyond simple "positive/negative" sentiment to predict 15 distinct emotional states, reaching **72% accuracy** and providing new insights into how emotions are structured in the digital age.
## Problem & Motivation: Beyond "Thumbs Up or Down"
Most sentiment analysis treats human emotion as a linear scale from sad to happy. However, human expression is far more complex—an image of a sunset could represent "Calm," "Optimism," or even "Sadness" depending on the accompanying caption.
The authors identified two major gaps in existing research:
1. **Lack of Nuance**: Traditional models ignore latent emotional states (e.g., "Annoyed" vs. "Disgusted").
2. **Modal Isolation**: Text-only or image-only models often miss the "Inter-modal" context required to resolve ambiguity.
By using Tumblr—a platform where users frequently tag their own "self-reported" emotions—the researchers found a goldmine of labeled data that represents how humans *actually* express feelings online.
## Methodology: The Fusion of Vision and Language
The "Deep Sentiment" architecture is a textbook example of effective **Late Fusion**.
### 1. Visual Branch (CNN)
The model uses a pre-trained **Inception** network. This branch captures the "low-dimensional manifold" of images—the colors, shapes, and arrangements that evoke specific moods before a single word is read.
### 2. Textual Branch (GloVe + LSTM)
To handle the variable length of captions, the authors used **GloVe** word embeddings followed by a **Long Short-Term Memory (LSTM)** network. This is critical because word order (e.g., "not happy") fundamentally flips the sentiment, a nuance lost in simple Bag-of-Words models.
### 3. Multimodal Fusion
The outputs from both branches are concatenated into a dense layer, which learns the correlation between the two.

*Figure 1: The Deep Sentiment architecture illustrating the parallel processing of visual and textual data.*
## Experiments & Results: Text is King, but Multi-modality is Ace
The evaluation showed a clear hierarchy of performance:
* **Image Model Only**: 36% Accuracy
* **Text Model Only**: 69% Accuracy
* **Deep Sentiment (Multimodal)**: 72% Accuracy
The text serves as the primary anchor for sentiment, but the image provides the "emotional edge" that pushes the accuracy higher.

*Figure 2: Performance gains showing the advantage of the combined Deep Sentiment approach over unimodal baselines.*
## Critical Analysis: Challenging Psychological Norms
One of the most fascinating aspects of this paper is its application to the **Circumplex Model of Emotion**. In psychology, it is often posited that two dimensions—**Valence** (positive/negative) and **Arousal** (high/low energy)—can explain all emotions.
The authors performed **Principal Component Analysis (PCA)** on their model's outputs and discovered:
* **PC1** mapped strongly to Valence.
* **PC2 and PC3** did not map neatly to Arousal alone.
This suggests that social media emotion is higher-dimensional than traditional laboratory models suggest. The model also identified modern vocabulary like "woke" as being statistically significant for emotions like "Scared" or "Amazed," proving that AI can track the evolving landscape of human language better than static, hand-coded dictionaries like LIWC.
## Deep Insights & Conclusion
**Deep Sentiment** proves that while "a picture is worth a thousand words," its emotional meaning is often locked behind the text that accompanies it.
**Limitations**: The authors acknowledge "performative bias"—people post how they *want* to be seen, not necessarily how they truly feel.
**Legacy**: This work provides a scalable tool for social scientists to analyze the "meme culture" and visual language of the 21st century, bridging the gap between computer science and affective psychology.
