HRSN: Redefining User Behavior Prediction via Meta-path-free Embedding and Self-Attention Denoising
User behavior prediction via heterogeneous information in social networks
This paper introduces the Heterogeneous Residual Self-Attention Shrinkage Network (HRSN) for user behavior prediction in social networks. It features a novel embedding method, User Heterogeneous Information Embedding (UHIE), and a Multi-head Self-attention Soft Thresholding (MSST) mechanism to achieve SOTA performance on datasets like IMDB and AMiner.
TL;DR
User behavior prediction is a cornerstone of social network analysis, yet traditional Graph Neural Networks (GNNs) often "blind" themselves by using only a single node attribute or relying on restrictive, human-defined meta-paths. This paper introduces the Heterogeneous Residual Self-Attention Shrinkage Network (HRSN). By combining a meta-path-free embedding method (UHIE) with a feature-level denoising gate (MSST), HRSN achieves massive performance gains—specifically outperforming traditional methods by over 20% on complex datasets like AMiner.
Problem & Motivation: The "Meta-path" Bottleneck
In a real-world social network, an "Author" isn't just a name; they are defined by their research fields, publication influence, social connections, and more.
- Attribute Neglect: Most existing frameworks (e.g., HAN, MAGNN) pick one primary attribute and discard the rest.
- The Meta-path Trap: To handle heterogeneous types (e.g., Author-writes-Paper), researchers usually define "meta-paths." However, defining these is labor-intensive and prone to missing subtle, complex relationships that the data might otherwise reveal.
The authors' insight is simple yet powerful: Aggregate everything first, but build a smart "shredder" (Soft Thresholding) to discard the noise later.
Methodology: The Core Innovations
1. User Heterogeneous Information Embedding (UHIE)
Instead of following a pre-defined path, UHIE looks at all neighbors and treats their various attributes as diverse signal sources.
- Intra-neighbor Aggregation: It uses an attention mechanism to fuse different attributes of a single neighbor type into a low-dimensional representation.
- Inter-neighbor Aggregation: It calculates the importance of different relationship types (e.g., a "Co-author" relationship vs. a "Citing" relationship) to update the target node's feature vector.
In Fig 1, we see the workflow: Node Content Transformation -> Intra-neighbor Aggregation -> Inter-neighbor Aggregation.
2. Multi-head Self-attention Soft Thresholding (MSST)
Aggregating "everything" introduces noise. Standard denoising uses a single threshold for all features, but HRSN argues that each feature deserves its own threshold. The MSST uses Multi-head Self-attention to look at the feature map and determine which dimensions are crucial and which are noise. It then applies a "Soft Thresholding" function:
- If a feature value is high, it’s kept (minus the threshold).
- If a feature value is within the "noise zone" (), it’s set to zero.
Fig 2 illustrates how multi-head attention generates the specific threshold matrix for the soft shrinkage operation.
Experiments & Results: Crushing the Baselines
The HRSN was tested against heavyweights like MAGNN, HAN, and GraphSAGE across three datasets: DBLP, IMDB, and AMiner.
Key Performance Wins:
- AMiner: HRSN achieved a Macro-F1 of 70.93% at 20% training ratio, while MAGNN (a leading meta-path method) struggled at 27.07%. This highlights that when meta-paths are hard to define, UHIE’s direct aggregation is far superior.
- IMDB: High-performance gains (approx. 9% improvement) show that the MSST denoising is highly effective for movie-related multi-attribute data.
Table 5 shows a consistent lead for HRSN, particularly as the complexity of the heterogeneous network increases.
Critical Analysis & Conclusion
Why it works:
The success of HRSN lies in its Inductive Bias. It assumes that the raw data in social networks is inherently "messy" but contains all the necessary signals. By replacing rigid human assumptions (meta-paths) with a flexible, attention-driven "aggregation-then-shrinkage" pipeline, the model learns the most discriminative features automatically.
Limitations:
- Computation: Multi-head self-attention on graph features adds a layer of complexity that might scale poorly on ultra-large graphs without further optimization (like sparse attention).
- Hyperparameter Sensitivity: The performance peaks at 4 MRSU units and 16 attention heads; beyond that, performance drops, suggesting a risk of overfitting.
Future Outlook:
The authors suggest applying these denoising techniques to Neuromorphic Computing. As we move towards spike-driven learning and energy-efficient AI, the ability to "shrink" unimportant information at the feature level will be crucial for reducing memory and power costs.
