Beyond Keywords: Explainable Emotion-Based Music Retrieval via Lyric Analysis
Emotion-Based Music Information Retrieval Using Lyrics
This paper presents a personalized music information retrieval system that recommends Japanese songs based on 9 emotions and 9 impressions extracted from lyrics. The core method uses NLP to identify adjectives and intransitive verbs in the "bridge" part of songs, maps them to representative words via a synonym dictionary, and employs AdaBoost with Decision Stumps to model individual user preferences.
TL;DR
Music discovery is shifting from "What is the title?" to "How does it feel?". This paper introduces a system that retrieves music by modeling 18 different emotional states (9 emotions, 9 impressions) from lyrics. By focusing on the "bridge" of a song and using specific linguistic markers (adjectives/verbs), the authors created a personalized recommendation engine that doesn't just suggest a song but explains why—using the very words that triggered the match.
Background: The Gap in Search
Most music libraries are indexed by metadata: Artist, Album, Year. But as listeners, our needs are often affective. We want music that makes us feel "Refreshed" or "Romantic." The authors identify two types of emotional search:
- Direct Emotion: Intentional state changes (e.g., "I want to feel Happy").
- Ambiguous Impression: Seeking a vibe (e.g., "Give me something Spectacular").
The challenge? Emotion is subjective. What feels "Joyful" to Subject A might feel "Sensational" to Subject B.
Methodology: The Linguistics of Feeling
The researchers developed a pipeline that prioritizes linguistic function over simple keyword matching.
1. The "Bridge" Hypothesis
In J-Pop, the bridge is often the emotional peak. Instead of analyzing the whole song, the system focuses its NLP resources here to maximize the "signal-to-noise" ratio of emotional content.
2. Feature Extraction: Verbs & Adjectives
While many systems use nouns (e.g., "summer," "sea"), this work argues that Intransitive Verbs (actions like "laughing") and Adjectives (states like "cheerful") are the true carriers of human emotion in text.
3. The Framework
The system maps extracted words to 4,626 "representative words" using a synonym dictionary to handle the diversity of the Japanese language.
Fig 1: The three-section framework for detecting, modeling, and retrieving music based on user feelings.
Experiments & Quantifiable Success
The authors tested the system on 9 subjects using 79 Japanese pop songs from the RWC music library.
Personalized vs. Baseline
The study proved that "one size does not fit all." Individual models (tailored to how specific users interpret lyrics) consistently outperformed the baseline model. For emotions like "Refreshed," accuracy reached as high as 98.7%.
The Power of Reduction (Information Gain)
One of the most striking results came from the "Reduction Model." By using Information Gain to select only the most relevant representative words, the researchers dramatically improved the baseline performance.
Table 6: The "Reduction Model" significantly improved accuracy across almost all categories (e.g., "Impressed" jumped from 43.0% to 79.7%).
Critical Insight: Explainable AI (XAI) in Music
Perhaps the most valuable contribution of this work is recommender transparency. When the system suggests a "Serenity" song, it presents the words "keep" and "think" as the justification. This builds user trust and helps the listener understand the system's logic, moving away from "black-box" recommendations.
Conclusion & Limitations
While the system shows high accuracy in many categories, it struggled with complex emotions like "Joyful." The authors noted that words like "laughing" are sometimes used in sad contexts (irony/contrast), which a simple dictionary mapping can't capture.
Future Outlook: The next step in this evolution is context-awareness—using LLMs to understand how a word is used, rather than just if it is present. However, this paper provides a robust foundation for building explainable, affective music discovery tools that respect the subjectivity of human emotion.
