Familiarity vs. Similarity: Why Your Friends Might Not Share Your Taste

A comparison study on familiarity-based and similarity-based social networks

2011-11-01
Jia Huang, Xiaohua Hu, Caimei Lu
Summary
Problem
Method
Results
Takeaways
Abstract

This study presents a comparative analysis between "familiarity-based" (contact lists) and "similarity-based" (co-reading patterns) social networks within the Douban platform. Utilizing Social Network Analysis (SNA), it demonstrates that while both exhibit small-world properties, there is a surprising lack of overlap between a user's explicit friends and their latent interest-based peers.

Executive Summary

In the world of collaborative filtering, we often assume that our social circles reflect our personal preferences. This paper, "A Comparison Study on Familiarity-Based and Similarity-Based Social Networks," puts that assumption to the test. By analyzing the Douban ecosystem, the researchers uncovered a startling structural gap: the people you follow (your "contacts") and the people who read like you (your "interest-sharers") are almost entirely different groups.

Background Positioning: This is a foundational comparative study that moves beyond "how to recommend" to "how these networks are fundamentally built." It challenges the traditional reliance on social proximity as a proxy for taste similarity.

Problem & Motivation: The Social Correlation Myth

Most recommender systems operate on the premise that users prefer items liked by "similar" others. But how do we define "similar"?

  1. Familiarity: People in your contact list or friend group.
  2. Similarity: People who consume the same content, even if they are strangers.

Previous literature was divided. Some argued friends vary significantly in taste, while others found strong correlations. The authors identified a need to look at the topology of these networks—how they are clustered and centralized—to see if they are truly mirrors of one another.

Methodology: Mapping the Social vs. the Semantic

The researchers constructed three distinct networks from a dataset of 1,080 users and over 73,000 books:

  • Contactship Network: A directed graph based on the "Add Contact" feature.
  • Co-occurrence Network: Based on the raw number of shared books (Threshold: 52).
  • Similarity Network: Based on the Cosine Similarity of reading vectors (Threshold: 0.169).

Architecture of the Analysis

The study utilized Regular Equivalence Blockmodeling to identify roles within the networks and employed Spearman’s correlation for centrality measures, acknowledging the non-normal distribution of social data.

Contactship Network Structure Figure 1: The Contactship network showing high centralization around seed users (node size indicates indegree).

Key Insights & Results

1. The Small-World Phenomenon

Despite their different origins, all three networks are "Small Worlds." They feature short path lengths and high clustering coefficients compared to random graphs. This confirms that information (or tastes) can theoretically flow quickly through either network.

2. Centrality Divergence

In the Similarity-based networks, top users across all centrality measures (Degree, Closeness, Betweenness) were consistent. However, in the Familiarity-based (contact) network, users with high "attractiveness" (Indegree) were rarely the ones with high "social activity" (Outdegree).

3. The Great Disconnect (Multiplexity)

The most critical finding was the lack of multiplexity. The cosine similarity between a user's friends and their interest-sharers was near zero (max 0.044).

Co-reading network visualization Figure 2: The Co-reading network colored by city area, showing no clear geographic clustering in reading habits.

4. Location Doesn't Matter

Through blockmodeling optimization, the authors proved that geographic location (Beijing, Shanghai, Overseas) has no correlation with a user's position in the interest graph. You are just as likely to share a "reading soul" with someone across the globe as with your neighbor.

Critical Analysis & Conclusion

Takeaways

  • Social is not Taste: Designers of recommender systems should be cautious about using "friendship" as a primary feature for item similarity.
  • Untapped Potential: There is a massive opportunity to recommend friends based on reading patterns, as these "interest-sharers" are currently invisible to each other in the social graph.

Limitations

The study is a snapshot in time. Contact lists are dynamic, and the co-reading thresholding (setting high bars like 52 shared books) might filter out subtler "weak tie" connections that are also valuable for discovery.

Future Outlook

The next frontier is Cross-Domain Similarity. Does your music taste correlate with your contact list even if your book taste doesn't? By applying this structural analysis to multi-modal data (audio, video, text), we can build more holistic models of human connection.

Find Similar Papers

Try Our Examples

  • Search for recent studies that measure the overlap between social graphs and interest graphs in modern platforms like Instagram or Twitter.
  • Which paper originally introduced Social Regularization in Recommender Systems, and how does this study's finding of low multiplexity impact that theory?
  • What are the performance differences between Jaccard similarity and Cosine similarity when constructing similarity-based networks for sparse user-item matrices?
Contents
Familiarity vs. Similarity: Why Your Friends Might Not Share Your Taste
1. Executive Summary
2. Problem & Motivation: The Social Correlation Myth
3. Methodology: Mapping the Social vs. the Semantic
3.1. Architecture of the Analysis
4. Key Insights & Results
4.1. 1. The Small-World Phenomenon
4.2. 2. Centrality Divergence
4.3. 3. The Great Disconnect (Multiplexity)
4.4. 4. Location Doesn't Matter
5. Critical Analysis & Conclusion
5.1. Takeaways
5.2. Limitations
5.3. Future Outlook