Decoding the Instagram Fashionista: A Data Mining Approach to Customer Behavior

Analysis of the behavior of customers in the social networks using data mining techniques

2016-08-01
Leidys del Carmen Contreras Chinchilla, Kevin Andrey Rosales Ferreira
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a data mining framework using the CRISP-DM methodology to analyze customer behavior on Instagram for a fashion company. By applying K-means clustering and FP-Growth association rules, the study successfully identifies high-engagement product categories like "Dressing Gowns" and "Long Cocktail Dresses" to inform B2C marketing strategies.

TL;DR

In the fast-paced world of fashion, understanding consumer "likes" is more than a vanity metric—it's a goldmine for strategy. This paper leverages the CRISP-DM methodology and K-means clustering to dissect Instagram data from a leading fashion brand, identifying which specific clothing styles (like cocktail dresses) drive the highest engagement and how association rules can predict future marketing success.

Background: Beyond the Feed

As companies shift toward Social CRM, the goal has moved from simple broadcasting to personalized engagement. The authors pose a powerful hypothesis: The predictions of future customer behavior are rooted in the past behaviors of others. By analyzing interactions on Instagram, businesses can move away from guesswork and toward a scientific understanding of user preferences.

Problem & Motivation: The Data-Information Gap

The digital era presents a paradox: companies are drowning in data but starving for insights. Many fashion brands post content based on intuition rather than empirical evidence. The specific challenge addressed here is characterizing which attributes (media type, tags, filters) actually trigger a customer to move from passive scrolling to active engagement (liking/commenting).

Methodology: The CRISP-DM Framework

The authors adopted the CRISP-DM (Cross Industry Standard Process for Data Mining) workflow, ensuring a structured transition from business understanding to deployment.

1. Data Acquisition & Preparation

Using the Instagram API and Python, the researchers gathered 1,435 records. Key attributes included creation date, media type, filter type, likes, comments, and tags.

2. The Modeling Core

The study implemented two primary descriptive techniques using RapidMiner:

  • Clustering (K-means): To segment the data into groups with similar interaction profiles.
  • Association Rules (FP-Growth): To find frequent relations between different attributes (e.g., how "media type" relates to "likes").

Table showing Cluster Centroids Above: Table 1 reveals the different clusters based on engagement metrics (Likes/Comments).

Experiments & Results: What Makes a Post Viral?

The analysis yielded specific, actionable segments. By evaluating the Sum of Squared Errors (SSE), the authors determined that K=5 provided the optimal number of clusters.

  • Cluster 0 (High Impact): Focused on "Dressing Gowns," showing high engagement (200-250 likes).
  • Cluster 4 (The Powerhouse): Focused on "Long Cocktail Dresses." This was identified as the company’s core strength, with 299 images generating peak interaction levels.
  • Temporal Insights: Association rules (via FP-Growth) indicated that posts from 2012-2013 had specific characteristics that led to historical engagement highs, suggesting a need to revisit those content styles.

Association Rules Analysis Above: The FP-Growth results showing the support and size of frequent itemsets.

Deep Insights & Conclusion

Takeaways for the Industry

The study proves that unsupervised learning is an efficient alternative to traditional market research. By identifying that "Long Cocktail Dresses" are the primary engagement driver, the brand can optimize its inventory and marketing spend toward these high-performing assets.

Limitations & Future Work

While effective, the study is limited to numerical and categorical metadata. A significant future expansion would be to incorporate Computer Vision (CNNs) to analyze the visual features of the clothing itself—such as color, texture, and pattern—to see how they influence the K-means clustering. Additionally, expanding the scope to Facebook and TikTok would allow for a cross-platform understanding of customer "social personas."

Final Summary

This paper serves as a practical roadmap for SMEs in the fashion industry to transition into data-driven powerhouses, proving that even simple data mining techniques can yield significant competitive advantages in the B2C landscape.

Find Similar Papers

Try Our Examples

  • Find recent studies that integrate Deep Learning-based image recognition with K-means clustering for fashion trend prediction on Instagram.
  • Which seminal paper established the CRISP-DM methodology, and how has its application evolved in the context of Big Data and Social Media Analytics?
  • Explore research papers and case studies that apply Association Rule Mining to multi-platform social media data (Instagram vs. TikTok) for consumer behavior comparison.
Contents
Decoding the Instagram Fashionista: A Data Mining Approach to Customer Behavior
1. TL;DR
2. Background: Beyond the Feed
3. Problem & Motivation: The Data-Information Gap
4. Methodology: The CRISP-DM Framework
4.1. 1. Data Acquisition & Preparation
4.2. 2. The Modeling Core
5. Experiments & Results: What Makes a Post Viral?
6. Deep Insights & Conclusion
6.1. Takeaways for the Industry
6.2. Limitations & Future Work
6.3. Final Summary