EmoBGM: Bridging the Affective Gap Between Music and Memories
EmoBGM: Estimating sound's emotion for creating slideshows with suitable BGM
EmoBGM is an automated system designed to recommend suitable background music (BGM) for photo slideshows by aligning the emotional content of both media types. The core contribution is a machine learning model that achieves 88% classification accuracy in estimating emotions from instrumental music using acoustic features.
Executive Summary
TL;DR: EmoBGM is a framework that automatically matches background music (BGM) to photo slideshows by analyzing the "emotional resonance" of both. By using a Random Forest model for audio feature extraction and facial recognition APIs for photos, the system achieves an 88% accuracy in music tagging and high user satisfaction in recommendation appropriateness.
Background Positioning: This work addresses the intersection of Affective Computing and Multi-modal Retrieval. While previous works focused on tagging music or images in isolation, EmoBGM focuses on the alignment of the two to enhance user experience in automated content creation.
The Motivation: Why Generic Music Fails
We’ve all seen those auto-generated "Year in Review" videos from smartphone apps. Often, a bittersweet memory is paired with an overly energetic pop track, or a high-octane sports clip is set to generic elevator music.
The authors argue that music is a powerful emotional amplifier. The mismatch occurs because current systems treat music as a generic "audio filler" rather than an emotional counterpart to the visual narrative. The challenge is specifically difficult for Instrumental BGM, where there are no lyrics to provide semantic clues, forcing the system to rely purely on raw acoustic signals.
Methodology: Decoding the Language of Sound and Sight
The process is divided into two distinct pipelines: Emotion Estimation (Music) and Emotion Extraction (Images).
1. The BGM Emotion Model
The researchers utilized a dataset of 50 movie scores, spanning multiple genres.
- Tagging: 250 participants used a refined version of Parrot’s emotion list (mapping 100+ tertiary emotions down to 14 core tags) to label the tracks.
- Feature Extraction: Using JAudio, they extracted low-level acoustic features, specifically focusing on MFCC-OSD (Mel-Frequency Cepstral Coefficients Overall Standard Deviation).
- Classification: Experimental results showed that Random Forest outperformed SVM and J48, especially when narrowed down to the 9 most impactful "BestFirst" features.

2. Image Sentiment and Matching
For the visual side, the system leverages the Microsoft Emotion API to analyze facial expressions. The matching logic is elegantly simple—it calculates a proximity score between the intensity percentage of the image's dominant emotion and the BGM's predicted emotion:
The goal is to find the BGM that matches the intensity of the visual sentiment, ensuring a more natural pairing.
Experimental Evidence: SOTA Validation
The model's classification performance is summarized in the table below, highlighting the robustness of the Random Forest approach:

User Evaluation Findings:
- Set 1 (Wedding): 4.1/5 rating. Success was attributed to clear, smiling faces providing strong "Happy" signals.
- Set 3 (Soccer Game): 3.0/5 rating. This drop revealed a critical insight: when facial expressions are obscured or the context (athletics) is specific, simple facial emotion analysis isn't enough.
Critical Analysis & Looking Ahead
Takeaway
EmoBGM proves that acoustic features alone can provide a high degree of emotional classification (88%) without needing metadata or lyrics. This is a significant win for processing independent or royalty-free music libraries.
The Context Gap (Limitations)
The study's results highlight the "Contextual Bottleneck." A "happy" wedding requires different instrumentation than a "happy" soccer victory. The authors correctly identify that future iterations must solve for Scene Understanding (e.g., using Google Cloud Vision or GPS data) to differentiate between social contexts.
Future Outlook
The shift toward multi-modal architectures (like CLIP or CLAP) suggests that the next step for EmoBGM would be an End-to-End Latent Space Alignment, where music and image features are projected into a shared emotional space, allowing for even more nuanced recommendations without explicit tagging.
Summary for the Reader: EmoBGM successfully moves us closer to "Director-level" automation in slideshow creation, proving that our devices can—and should—understand the vibe of our memories.
