Decoding the Heart of a Story: Emotional Book Classification via Blurbs
Emotional book classification from book blurbs
This paper presents a supervised learning framework for classifying books into emotional categories ("moods") using only their back-cover blurbs. By leveraging Italian lexical resources like MultiWordNet and OntoEmotion, the authors demonstrate that brief promotional texts can effectively predict reader-assigned emotional tags, achieving high accuracy with the J48 decision tree algorithm.
TL;DR
Can we predict if a book will make you cry, laugh, or think just by reading the short description on the back cover? Researchers Valentina Franzoni and Valentina Poggioni developed a system that uses lexical and semantic analysis of book blurbs to automatically assign "mood tags." By mapping Italian text to emotional ontologies like WordNet-Affect and Plutchik’s model, they proved that these short, promotional snippets are powerful enough to predict a reader's subjective emotional response.
The Problem: The "Missing Mood" in Book Discovery
Most book recommendation systems categorize works by genre (e.g., "Mystery" or "Bio") or author. However, readers often search for a specific "feeling"—something "uplifting" or "thought-provoking."
Current methods for capturing this emotional data face two major hurdles:
- Data Scarcity: On social networks like Zazie or Goodreads, many books lack user-generated tags.
- Computational Complexity: Processing the entire text of a 400-page novel for sentiment is inefficient and often illegal due to copyright issues.
The authors' Insight: The blurb is a condensed, emotionally-charged "hook" designed to attract readers. If the blurb reflects the core emotional "qualia" of the book, it can serve as a proxy for the whole work.
Methodology: From Words to Emotions
The system architecture follows a sophisticated pipeline to handle the complexities of the Italian language and emotional nuance.
1. Preprocessing and Lemmatization
Unlike English, Italian features complex adjective declensions. The authors utilized Morph-it! to reduce words to their base forms (lemmata), ensuring that "triste" (sad) and "tristi" (sad, plural) are treated as the same feature.
2. Emotional Extraction (The Core)
The researchers tested two distinct paths:
- WordNet-Affect Strategy: Mapping lemmata to a hierarchy of 296 nodes, eventually pruning them down to Ekman’s 8 basic emotions (Happiness, Anger, Disgust, Fear, Sadness, Surprise, Neutral, Ambiguous).
- OntoEmotion Strategy: Using an OWL-based ontology to map terms directly to Plutchik’s model, which includes complex emotions like "submission" and "optimism."
Figure 1: The multi-stage system architecture from text normalization to classification.
Experiments and Results
The study utilized a dataset from Zazie, focusing on 7 core emotional tags: angry, cry, lol, love, sad, smile, think.
Key Findings:
- Algorithm Performance: J48 (a decision tree algorithm) surprisingly outperformed SVM. This suggests that the emotional classification in this domain relies on specific hierarchical "breakpoints" in term frequency rather than a high-dimensional hyperplane separation.
- Accuracy Peaks: The system was remarkably accurate at identifying the "think" and "smile" categories, likely because these blurbs use very distinct, repeatable linguistic patterns.
- Class Imbalance: High-intensity emotions like "angry" had fewer samples but showed extremely high Precision (0.938) when the dataset was resampled to be more uniform.
Table 1: Detailed precision and recall for different emotional moods using the J48 classifier.
Critical Analysis & Future Outlook
This work demonstrates that affective semantics can bridge the gap between institutional book metadata and user-centered emotional tagging.
Limitations:
- Dataset Bias: The study is influenced by the "Marketing Bias" of blurbs—publishers often use hyperbolic emotional language, which might skew the actual sentiment of the interior text.
- Small Sample for "Angry/Cry": The minor classes need more data to be as robust as the "think" class.
The Future of "Emotion-Driven Search"
The authors suggest that future iterations could move beyond fixed ontologies by using Collective Knowledge (like Wikipedia) and Semantic Proximity Measures to dynamically link new words to emotions. This would transform digital libraries from simple databases into "intention-aware" systems that truly understand how a book will make you feel.
Conclusion
By treating the back-cover blurb as an emotional fingerprint, Franzoni and Poggioni have opened a new path for automated book discovery. It proves that even in the age of Big Data, the "vibe" of a book can be calculated one lemma at a time.
