You Are What You Listen To: Decoding Personal Traits from Music History
Inferring personal traits from music listening history
This paper explores the feasibility of automatically inferring personal traits, specifically age and gender, from music listening history. Using a dataset of 1K Last.fm users, the researchers utilize Support Vector Machines (SVM) with three feature categories: temporal patterns (context), song/artist metadata (content), and acoustic signal features, achieving a peak accuracy of 71.1% for age classification.
TL;DR
Can your Spotify or Last.fm history reveal your age and gender? This classic study demonstrates that machine learning models can accurately predict personal traits by analyzing not just what you listen to, but when you listen to it. By extracting temporal context and artist metadata, the researchers achieved up to 71% accuracy in identifying age groups, highlighting a new frontier for both personalization and digital privacy.
The "Digital Breadcrumbs" of Music
We are living in an era where our digital footprints—ranging from search queries to blog posts—are used to build sophisticated profiles of our identity. While text analysis (Natural Language Processing) has long been used to guess user demographics, music remained a relatively untapped signal.
The authors argue that music is an intensely personal choice influenced by emotion, social context, and culture. If a human can guess a student's personality just by looking at their iTunes playlist (as psychologist Sam Gosling famously suggested), a computer should be able to do the same at scale.
Methodology: Context vs. Content
The study breaks down music listening history into three distinct layers of data:
- Context (Temporal Patterns): Novel features such as "Working-hour ratio" or "Day-of-week entropy." These capture the rhythm of your life. Do you listen to music mostly during office hours or late at night on weekends?
- Content (Metadata): Histograms of specific artists and songs. This captures your taste and cultural alignment.
- Content (Acoustic Signals): High-level features like danceability, tempo, and timbre extracted from the audio signal itself.

The researchers used a dataset of approximately 1,000 Last.fm users, formulating the task as binary classification: Adolescent (<24) vs. Adult (>=24) and Male vs. Female.
Key Findings: The Power of Timing
The results revealed a fascinating split in how different features contribute to different traits:
- Age is about the Week: "Day-of-week" features were surprisingly effective for age detection. Interestingly, adults showed a higher "working-day ratio"—likely because they listen to music while at their desks during the work week, whereas adolescents' usage was more distributed.
- Gender is about the Hour: Gender classification was better served by "hour-of-day" features. The study found that females tended to listen to more music between 17:00 and 22:00, while males were more active in the mornings (06:00–12:00).

Performance Comparison
When it comes to raw predictive power, Artist Histograms (Metadata) were the clear winners. Knowing that you listen to Kanye West (popular among adolescents/males in this dataset) or Madonna (favored by the adult group) provides a stronger signal than the acoustic "danceability" of the tracks.

Audio signal features performed the worst, suggesting that while we might share a preference for "timbre," those preferences don't align as neatly with demographic boundaries as specific artists do.
Critical Insight & Privacy Warning
The study concludes with a sobering thought: music streaming data is essentially a mirror of our daily schedules and identities. While this allows for better recommendation engines (e.g., suggesting "Commute" playlists to workers), it also creates a surface for de-anonymization. Even if your name is removed, the unique "stamps" of your listening habits could potentially be used to identify you.
Conclusion
This work bridges the gap between sociology and computer science, proving that our musical preferences are not just aesthetic choices, but markers of our stage in life and our gender identity. As MIR moves forward, the challenge will be to leverage these insights for better user experience without compromising the inherent privacy of our listening sanctuaries.
