Familiarity vs. Similarity: Why Your Friends Might Not Share Your Taste
A comparison study on familiarity-based and similarity-based social networks
This study presents a comparative analysis between "familiarity-based" (contact lists) and "similarity-based" (co-reading patterns) social networks within the Douban platform. Utilizing Social Network Analysis (SNA), it demonstrates that while both exhibit small-world properties, there is a surprising lack of overlap between a user's explicit friends and their latent interest-based peers.
Executive Summary
In the world of collaborative filtering, we often assume that our social circles reflect our personal preferences. This paper, "A Comparison Study on Familiarity-Based and Similarity-Based Social Networks," puts that assumption to the test. By analyzing the Douban ecosystem, the researchers uncovered a startling structural gap: the people you follow (your "contacts") and the people who read like you (your "interest-sharers") are almost entirely different groups.
Background Positioning: This is a foundational comparative study that moves beyond "how to recommend" to "how these networks are fundamentally built." It challenges the traditional reliance on social proximity as a proxy for taste similarity.
Problem & Motivation: The Social Correlation Myth
Most recommender systems operate on the premise that users prefer items liked by "similar" others. But how do we define "similar"?
- Familiarity: People in your contact list or friend group.
- Similarity: People who consume the same content, even if they are strangers.
Previous literature was divided. Some argued friends vary significantly in taste, while others found strong correlations. The authors identified a need to look at the topology of these networks—how they are clustered and centralized—to see if they are truly mirrors of one another.
Methodology: Mapping the Social vs. the Semantic
The researchers constructed three distinct networks from a dataset of 1,080 users and over 73,000 books:
- Contactship Network: A directed graph based on the "Add Contact" feature.
- Co-occurrence Network: Based on the raw number of shared books (Threshold: 52).
- Similarity Network: Based on the Cosine Similarity of reading vectors (Threshold: 0.169).
Architecture of the Analysis
The study utilized Regular Equivalence Blockmodeling to identify roles within the networks and employed Spearman’s correlation for centrality measures, acknowledging the non-normal distribution of social data.
Figure 1: The Contactship network showing high centralization around seed users (node size indicates indegree).
Key Insights & Results
1. The Small-World Phenomenon
Despite their different origins, all three networks are "Small Worlds." They feature short path lengths and high clustering coefficients compared to random graphs. This confirms that information (or tastes) can theoretically flow quickly through either network.
2. Centrality Divergence
In the Similarity-based networks, top users across all centrality measures (Degree, Closeness, Betweenness) were consistent. However, in the Familiarity-based (contact) network, users with high "attractiveness" (Indegree) were rarely the ones with high "social activity" (Outdegree).
3. The Great Disconnect (Multiplexity)
The most critical finding was the lack of multiplexity. The cosine similarity between a user's friends and their interest-sharers was near zero (max 0.044).
Figure 2: The Co-reading network colored by city area, showing no clear geographic clustering in reading habits.
4. Location Doesn't Matter
Through blockmodeling optimization, the authors proved that geographic location (Beijing, Shanghai, Overseas) has no correlation with a user's position in the interest graph. You are just as likely to share a "reading soul" with someone across the globe as with your neighbor.
Critical Analysis & Conclusion
Takeaways
- Social is not Taste: Designers of recommender systems should be cautious about using "friendship" as a primary feature for item similarity.
- Untapped Potential: There is a massive opportunity to recommend friends based on reading patterns, as these "interest-sharers" are currently invisible to each other in the social graph.
Limitations
The study is a snapshot in time. Contact lists are dynamic, and the co-reading thresholding (setting high bars like 52 shared books) might filter out subtler "weak tie" connections that are also valuable for discovery.
Future Outlook
The next frontier is Cross-Domain Similarity. Does your music taste correlate with your contact list even if your book taste doesn't? By applying this structural analysis to multi-modal data (audio, video, text), we can build more holistic models of human connection.
