You Are What You Listen To: Decoding Personal Traits from Music History

Inferring personal traits from music listening history

2012-11-02
Jen-Yu Liu, Yi-Hsuan Yang
Summary
Problem
Method
Results
Takeaways
Abstract

This paper explores the feasibility of automatically inferring personal traits, specifically age and gender, from music listening history. Using a dataset of 1K Last.fm users, the researchers utilize Support Vector Machines (SVM) with three feature categories: temporal patterns (context), song/artist metadata (content), and acoustic signal features, achieving a peak accuracy of 71.1% for age classification.

TL;DR

Can your Spotify or Last.fm history reveal your age and gender? This classic study demonstrates that machine learning models can accurately predict personal traits by analyzing not just what you listen to, but when you listen to it. By extracting temporal context and artist metadata, the researchers achieved up to 71% accuracy in identifying age groups, highlighting a new frontier for both personalization and digital privacy.

The "Digital Breadcrumbs" of Music

We are living in an era where our digital footprints—ranging from search queries to blog posts—are used to build sophisticated profiles of our identity. While text analysis (Natural Language Processing) has long been used to guess user demographics, music remained a relatively untapped signal.

The authors argue that music is an intensely personal choice influenced by emotion, social context, and culture. If a human can guess a student's personality just by looking at their iTunes playlist (as psychologist Sam Gosling famously suggested), a computer should be able to do the same at scale.

Methodology: Context vs. Content

The study breaks down music listening history into three distinct layers of data:

  1. Context (Temporal Patterns): Novel features such as "Working-hour ratio" or "Day-of-week entropy." These capture the rhythm of your life. Do you listen to music mostly during office hours or late at night on weekends?
  2. Content (Metadata): Histograms of specific artists and songs. This captures your taste and cultural alignment.
  3. Content (Acoustic Signals): High-level features like danceability, tempo, and timbre extracted from the audio signal itself.

Data distribution of Last.fm-1K

The researchers used a dataset of approximately 1,000 Last.fm users, formulating the task as binary classification: Adolescent (<24) vs. Adult (>=24) and Male vs. Female.

Key Findings: The Power of Timing

The results revealed a fascinating split in how different features contribute to different traits:

  • Age is about the Week: "Day-of-week" features were surprisingly effective for age detection. Interestingly, adults showed a higher "working-day ratio"—likely because they listen to music while at their desks during the work week, whereas adolescents' usage was more distributed.
  • Gender is about the Hour: Gender classification was better served by "hour-of-day" features. The study found that females tended to listen to more music between 17:00 and 22:00, while males were more active in the mornings (06:00–12:00).

Average hour-of-day histogram for male and female users

Performance Comparison

When it comes to raw predictive power, Artist Histograms (Metadata) were the clear winners. Knowing that you listen to Kanye West (popular among adolescents/males in this dataset) or Madonna (favored by the adult group) provides a stronger signal than the acoustic "danceability" of the tracks.

Performance Results Table

Audio signal features performed the worst, suggesting that while we might share a preference for "timbre," those preferences don't align as neatly with demographic boundaries as specific artists do.

Critical Insight & Privacy Warning

The study concludes with a sobering thought: music streaming data is essentially a mirror of our daily schedules and identities. While this allows for better recommendation engines (e.g., suggesting "Commute" playlists to workers), it also creates a surface for de-anonymization. Even if your name is removed, the unique "stamps" of your listening habits could potentially be used to identify you.

Conclusion

This work bridges the gap between sociology and computer science, proving that our musical preferences are not just aesthetic choices, but markers of our stage in life and our gender identity. As MIR moves forward, the challenge will be to leverage these insights for better user experience without compromising the inherent privacy of our listening sanctuaries.

Find Similar Papers

Try Our Examples

  • Which recent studies have used Deep Learning or Graph Neural Networks to improve the accuracy of demographic inference from streaming music metadata compared to traditional SVMs?
  • What is the theoretical origin of using temporal "routine patterns" or circadean rhythms for user profiling, and how has this been expanded beyond the Last.fm dataset?
  • How do modern privacy-preserving techniques like Differential Privacy or Federated Learning mitigate the risks of trait inference in music recommendation systems?
Contents
You Are What You Listen To: Decoding Personal Traits from Music History
1. TL;DR
2. The "Digital Breadcrumbs" of Music
3. Methodology: Context vs. Content
4. Key Findings: The Power of Timing
5. Performance Comparison
6. Critical Insight & Privacy Warning
7. Conclusion