Overlap vs. Partition: Decoding the Geometry of Market Behavior

Overlap versus partition: Marketing classification and customer profiling in complex networks of products

2014-03-01
Diego Pennacchioli, Michele Coscia, Dino Pedreschi
Summary
Problem
Method
Results
Takeaways
Abstract

This paper investigates the practical utility of partition-based vs. overlapping community discovery in complex product co-purchase networks. Using data from a massive Italian retail chain, the authors demonstrate that the Infomap (partition) and Hierarchical Link Clustering (HLC) (overlap) algorithms serve distinct strategic marketing purposes rather than competing for "accuracy."

TL;DR

Is a bottle of Cola just a "Soft Drink," or is it a "Party Essential," "Mixer," and "Lunch Item" all at once? This paper argues that how we cluster products depends on the marketing goal. By comparing Partition-based and Overlapping community discovery on 80 million shopping sessions, the researchers found that partitions excel at refining product catalogs, while overlaps are the key to unlocking true customer profiling.

Background: The Dual Nature of Products

In the era of Big Data, retailers like the Italian chain Coop are no longer just selling items; they are managing complex networks of human needs. Traditionally, marketing clusters (communities) are treated as distinct silos. However, this paper challenges the "obsolete" notion that one method fits all, positioning community discovery as a flexible lens through either a structural or behavioral perspective.

The "Why": Why Choice of Algorithm Matters

The researchers identified a fundamental tension in network science:

  1. Structural Purity: Some products are functionally identical (e.g., different types of yogurt). They belong together in a rigid taxonomy.
  2. Contextual Fluidity: Some products are bought together across categories (e.g., sunscreen and charcoal for a summer trip).

Previous SOTA (State of the Art) focuses on which algorithm is "better." This paper asks a more practical question: What is each algorithm actually finding?

Methodology: Mapping the Supermarket Network

To test their hypothesis, the authors built a network of 14,949 products using two primary metrics:

  • Absolute Support: How many times two items were bought together.
  • Lift: How much more likely they are to be bought together compared to random chance.

The Two Pillars of Discovery:

  • Infomap (Partition): Uses the flow of information (random walks) to find the most efficient way to "compress" the network. It forces every product into exactly one "street" or cluster.
  • Hierarchical Link Clustering (HLC): Focuses on the edges rather than the nodes. Since a product (node) can have multiple relationships (edges), it can belong to multiple communities simultaneously.

Conceptual Data Model Figure 1: The star schema of the retail data warehouse used to build the network.

Experiments & Results: Distribution and Diversity

The team found startling differences in the output of the two methods:

1. Size Distribution

Infomap produced a power-law distribution—many small, tight clusters. HLC, however, favored larger, sprawling communities. This suggests that partitions find "micro-niches," while overlaps find "macro-behaviors."

Community Size Distribution Figure 2: Infomap (Red) generates smaller units; HLC (Green) captures more expansive structures.

2. The Entropy Gap

The researchers measured Information Entropy relative to the supermarket's own marketing hierarchy.

  • Infomap (Lower Entropy): Clusters were homogenous (e.g., a community consisting only of organic jams). It effectively "rediscovered" and refined the existing marketing categories.
  • HLC (Higher Entropy): Clusters were diverse. A single HLC community might include items from five different sections of the store, representing a specific "Customer Profile" (e.g., the "Healthy Breakfast" profile including fruit, yogurt, and specialized cereal).

Infomap Community Extract Figure 3: Infomap extracts (#37 and #80) show high-density, single-category cliques like biological jams or liquid yogurts.

Critical Insight & Conclusion

The "takeaway" is a framework for Business Intelligence:

  • Need to fix your catalog? Use a Partition approach like Infomap. It will highlight where your current categories are inconsistent with how people actually buy.
  • Need to build a Recommender System? Use an Overlapping approach like HLC. It captures the "cross-pollination" of products that define lifestyle-based purchasing.

Limitations: The study is limited to undirected networks. Future work could benefit from analyzing the temporal aspect—does a "Customer Profile" overlap change between winter and summer?

Final Thought: In complex networks, truth isn't found in a single partition, but in the tension between how we categorize the world (partition) and how we actually live in it (overlap).

Find Similar Papers

Try Our Examples

  • Find recent studies that compare different community discovery definitions in the context of retail recommender systems.
  • Which paper first proposed the Hierarchical Link Clustering (HLC) algorithm, and how does its use of link-based similarity differ from node-based overlapping methods like DEMON?
  • Explore how overlapping community discovery has been applied to multi-layered networks where customers and products are modeled as separate node types.
Contents
Overlap vs. Partition: Decoding the Geometry of Market Behavior
1. TL;DR
2. Background: The Dual Nature of Products
3. The "Why": Why Choice of Algorithm Matters
4. Methodology: Mapping the Supermarket Network
4.1. The Two Pillars of Discovery:
5. Experiments & Results: Distribution and Diversity
5.1. 1. Size Distribution
5.2. 2. The Entropy Gap
6. Critical Insight & Conclusion