Beyond Keywords: Explainable Emotion-Based Music Retrieval via Lyric Analysis

Emotion-Based Music Information Retrieval Using Lyrics

2015-01-01
Akihiro Ogino, Yuko Yamashita
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a personalized music information retrieval system that recommends Japanese songs based on 9 emotions and 9 impressions extracted from lyrics. The core method uses NLP to identify adjectives and intransitive verbs in the "bridge" part of songs, maps them to representative words via a synonym dictionary, and employs AdaBoost with Decision Stumps to model individual user preferences.

TL;DR

Music discovery is shifting from "What is the title?" to "How does it feel?". This paper introduces a system that retrieves music by modeling 18 different emotional states (9 emotions, 9 impressions) from lyrics. By focusing on the "bridge" of a song and using specific linguistic markers (adjectives/verbs), the authors created a personalized recommendation engine that doesn't just suggest a song but explains why—using the very words that triggered the match.

Background: The Gap in Search

Most music libraries are indexed by metadata: Artist, Album, Year. But as listeners, our needs are often affective. We want music that makes us feel "Refreshed" or "Romantic." The authors identify two types of emotional search:

  1. Direct Emotion: Intentional state changes (e.g., "I want to feel Happy").
  2. Ambiguous Impression: Seeking a vibe (e.g., "Give me something Spectacular").

The challenge? Emotion is subjective. What feels "Joyful" to Subject A might feel "Sensational" to Subject B.

Methodology: The Linguistics of Feeling

The researchers developed a pipeline that prioritizes linguistic function over simple keyword matching.

1. The "Bridge" Hypothesis

In J-Pop, the bridge is often the emotional peak. Instead of analyzing the whole song, the system focuses its NLP resources here to maximize the "signal-to-noise" ratio of emotional content.

2. Feature Extraction: Verbs & Adjectives

While many systems use nouns (e.g., "summer," "sea"), this work argues that Intransitive Verbs (actions like "laughing") and Adjectives (states like "cheerful") are the true carriers of human emotion in text.

3. The Framework

The system maps extracted words to 4,626 "representative words" using a synonym dictionary to handle the diversity of the Japanese language.

The Recommendation Framework Fig 1: The three-section framework for detecting, modeling, and retrieving music based on user feelings.

Experiments & Quantifiable Success

The authors tested the system on 9 subjects using 79 Japanese pop songs from the RWC music library.

Personalized vs. Baseline

The study proved that "one size does not fit all." Individual models (tailored to how specific users interpret lyrics) consistently outperformed the baseline model. For emotions like "Refreshed," accuracy reached as high as 98.7%.

The Power of Reduction (Information Gain)

One of the most striking results came from the "Reduction Model." By using Information Gain to select only the most relevant representative words, the researchers dramatically improved the baseline performance.

Performance Comparison Table 6: The "Reduction Model" significantly improved accuracy across almost all categories (e.g., "Impressed" jumped from 43.0% to 79.7%).

Critical Insight: Explainable AI (XAI) in Music

Perhaps the most valuable contribution of this work is recommender transparency. When the system suggests a "Serenity" song, it presents the words "keep" and "think" as the justification. This builds user trust and helps the listener understand the system's logic, moving away from "black-box" recommendations.

Conclusion & Limitations

While the system shows high accuracy in many categories, it struggled with complex emotions like "Joyful." The authors noted that words like "laughing" are sometimes used in sad contexts (irony/contrast), which a simple dictionary mapping can't capture.

Future Outlook: The next step in this evolution is context-awareness—using LLMs to understand how a word is used, rather than just if it is present. However, this paper provides a robust foundation for building explainable, affective music discovery tools that respect the subjectivity of human emotion.

Find Similar Papers

Try Our Examples

  • Search for recent studies that combine lyrical NLP with audio feature analysis (multimodal) to achieve SOTA music emotion recognition.
  • Which paper first established the 'bridge' of a song as the primary carrier of emotional climax in Japanese Pop (J-Pop) structure, and how has this influenced recent MIR research?
  • Investigate how transformer-based models (like BERT or RoBERTa) have replaced dictionary-based synonym mapping for Japanese sentiment analysis in the context of song lyrics.
Contents
Beyond Keywords: Explainable Emotion-Based Music Retrieval via Lyric Analysis
1. TL;DR
2. Background: The Gap in Search
3. Methodology: The Linguistics of Feeling
3.1. 1. The "Bridge" Hypothesis
3.2. 2. Feature Extraction: Verbs & Adjectives
3.3. 3. The Framework
4. Experiments & Quantifiable Success
4.1. Personalized vs. Baseline
4.2. The Power of Reduction (Information Gain)
5. Critical Insight: Explainable AI (XAI) in Music
6. Conclusion & Limitations