SNE: Bridging Topology and Homophily in Social Network Embedding

Attributed Social Network Embedding

2018-03-27
Lizi Liao, Xiangnan He, Hanwang Zhang, Tat-Seng Chua
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces SNE (Social Network Embedding), a deep learning framework designed to learn node representations by jointly preserving structural and attribute proximities. It achieves State-of-the-Art (SOTA) results on node classification and link prediction across four real-world social datasets.

TL;DR

Existing network embedding tools like node2vec and LINE often ignore the "who" behind the "nodes." SNE (Social Network Embedding) changes this by treating user attributes (age, major, text) as equal citizens to network links. By using a deep neural architecture to fuse identity and attributes early in the pipeline, SNE yields representations that are significantly more informative, specifically boosting node classification accuracy by 12.7%.

Background: Why Links Aren't Enough

In social science, the "homophily principle" states that "birds of a feather flock together." If two users in a university network share the same major and graduation year, they are likely to connect, even if a random walk hasn't reached them yet.

Standard embeddings rely on Structural Proximity (if A and B are connected, they are similar). However, they miss Attribute Proximity—the intrinsic similarity between actors. As shown in the Facebook dataset visualizations below, nodes grouped by attributes like "Class Year" form distinct, dense clusters (blocks) that structural-only models struggle to capture in sparse regions.

Attribute Homophily in Facebook Data (a) Re-ordering of users by 'Class Year' reveals clear block structures that represent attribute-driven connectivity.

Methodology: Deep Early Fusion

The core innovation of SNE is its Early Fusion architecture. Instead of training a structural model and an attribute model separately and "stitching" them together (late fusion), SNE combines them at the input.

1. Feature Encoding

SNE converts all attributes into a generic vector.

  • Discrete attributes (e.g., Gender) use one-hot encoding.
  • Continuous attributes (e.g., TF-IDF text features) are treated as real-valued entries.

2. The SNE Framework

The model projects the node ID into a dense vector (structure) and the attributes into vector (homophily). These are concatenated and fed through a high-capacity Multi-Layer Perceptron (MLP).

SNE Architecture Figure: The SNE framework. Input features undergo early fusion before passing through hidden layers to capture non-linear structure-attribute interactions.

The hidden layers follow a tower structure (halving the neurons per layer) which encourages the model to learn abstract, high-level features. The objective function uses negative sampling to efficiently optimize the likelihood of observed social ties.

Experiments: Does Complexity Pay Off?

The authors tested SNE against benchmarks like node2vec, LINE, and TriDNR on Friendship (Facebook) and Citation (DBLP, CiteSeer) networks.

Key Performance Gains

SNE consistently outperformed all baselines. In Link Prediction, SNE maintained high performance even when the network was extremely sparse, thanks to its ability to "fall back" on attribute similarity when link data was missing.

Link Prediction Results Comparison on Oklahoma and UNC datasets show SNE (top line) significantly leading in AUROC scores.

The Power of Depth

A critical finding (RQ3) was the impact of hidden layers. Moving from a linear model (no hidden layers) to a 2-layer deep model increased the AUROC from 0.927 to 0.954 on DBLP. This suggests that the relationship between who we are (attributes) and who we know (structure) is non-linear and requires the expressive power of deep learning to decode.

Critical Analysis & Conclusion

SNE serves as a vital bridge between traditional graph theory and modern deep learning. Its ability to subsume models like SVD++ and node2vec proves its theoretical robustness.

Limitations:

  • Computational Cost: More layers lead to higher training times (a 3x increase from 1 to 3 layers).
  • Scalability: While negative sampling helps, the fully-connected layers might encounter bottlenecks on billion-node graphs.

Looking Ahead: The framework opens the door for multi-modal embeddings—integrating images and video content from platforms like Instagram into the social graph representation. For researchers, SNE underscores that in social networks, the "node" is not just an ID; it is a rich, multifaceted entity.


Summary Takeaway: By fusing topology with attribute homophily via deep learning, SNE provides more robust node vectors that excel in link prediction and node classification, specifically addressing the link sparsity problem.

Find Similar Papers

Try Our Examples

  • Search for recent papers that extend Attributed Network Embedding (ANE) using Graph Convolutional Networks (GCNs) or Attention mechanisms.
  • Which paper first established the "homophily principle" as a theoretical basis for social network analysis, and how does SNE's mathematical objective align with it?
  • Explore how the SNE framework's early fusion strategy has been adapted for multi-modal social networks containing image or video attributes.
Contents
SNE: Bridging Topology and Homophily in Social Network Embedding
1. TL;DR
2. Background: Why Links Aren't Enough
3. Methodology: Deep Early Fusion
3.1. 1. Feature Encoding
3.2. 2. The SNE Framework
4. Experiments: Does Complexity Pay Off?
4.1. Key Performance Gains
4.2. The Power of Depth
5. Critical Analysis & Conclusion