MDM: Cracking Social Spammer Detection via Multi-Level Relation Dependencies

Leveraging multi-level dependency of relational sequences for social spammer detection

2020-11-05
Jun Yin, Qian Li, Shaowu Liu, Zhiang Wu, Guandong Xu
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces the Multi-level Dependency Model (MDM) for social spammer detection in multi-relational social networks. It leverages skip-gram, LSTM, and Residual Networks (ResNet) to extract features from heterogeneous relation sequences, achieving SOTA performance on the Tagged.com dataset.

TL;DR

Detection of social spammers has shifted from "what they say" (content) to "how they interact" (behavior). The Multi-level Dependency Model (MDM) represents a major leap in behavior-based detection by analyzing sequences of heterogeneous relations. By combining LSTM-based long-term context with ResNet-powered short-term dependency analysis (at both individual and union levels), MDM achieves a superior balance of precision and recall on massive real-world datasets like Tagged.com.

The Evolution of Spammer Detection

Spammers are evolving. While early filters looked for keywords, modern malicious actors are adept at hiding their intentions behind a cloak of anonymity. Content is also increasingly encrypted or gated for privacy.

Prior works faced two major hurdles:

  1. Graph-based methods often treat different relations (Add Friend, View Profile, Send Message) as separate, homogeneous graphs, ignoring how one action leads to another.
  2. Sequence-based methods (like k-grams) are either too computationally expensive for long sequences or fail to see the "big picture" of a user's intent.

MDM’s core insight is that while a spammer might mimic a single normal action (like adding a gift), they struggle to replicate the complex union-level sequences that normal users organically generate over time.

Methodology: The Three-Pillar Architecture

MDM processes raw quad-tuples (timestamp, usrc, udest, rm) through a sophisticated deep learning pipeline.

1. User-Relation Representation

Using a skip-gram approach within a Recurrent Neural Network, the model learns a -dimensional latent vector for each type of relation. This ensures that the semantic meaning of "Message" vs. "Report Abuse" is captured relative to the sequence context.

2. Long-term Dependency Modeling

A standard LSTM layer processes the entire history of a user's relations. This creates a hidden state that represents the "general intent" of the user across the whole day or week. Architecture Overview

3. Short-term Dependency Modeling (The Core Innovation)

The model focuses on the most recent relations (where was found optimal) using two specific perspectives:

  • Individual-level: Uses a Multi-order Attention Network with ResNetR to capture high-order nonlinear interactions of single recent actions.
  • Union-level: Uses ResNetE to model the collective influence of a combination of actions (e.g., how the specific combination of "Profile View + Friend Request" leads to a "Message").

Multi-order Attention Network

Experimental Validation

The model was tested on a massive Tagged.com dataset containing 85 million interactions.

SOTA Comparison

MDM was compared against traditional graph-based and sequence-based (k-gram) methods using Logistic Regression (LR) and XGBoost.

  • Performance: MDM reached an F-measure of 0.7750, a significant jump from the 0.7421 achieved by combining prior SOTA methods.
  • Precision vs. Recall: While k-gram methods often provide high recall at the cost of many false positives, MDM provides much higher Precision (0.7385), meaning fewer normal users are incorrectly flagged.

Results Comparison

Ablation Study: What makes it work?

The authors demonstrated that every layer counts. Adding the long-term context improved precision, but the inclusion of Individual-level and Union-level modeling provided the most significant performance boosts (Precision jumped from 0.59 to 0.73).

Critical Insight: The "Pet Game" Anomaly

In an interesting qualitative analysis, the authors noted that spammers on Tagged.com often obsessively repeat a specific "Pet Game" (Relation ID: 5) to gain rewards and appear on celebrity lists. MDM successfully flagged sequences like [5, 5, 5, 5, 5, 5] as highly suspicious, even when the spammer attempted to hide by occasionally adding a friend or gift.

Conclusion and Future Outlook

MDM proves that relational sequences are a goldmine for social network security. Its ability to extract high-order dependencies without relying on text content makes it a powerful tool for privacy-conscious platforms.

Future Directions: The next frontier is incorporating multi-modal data—combining these relational sequences with anonymized text embeddings (like BERT) to create an even more resilient detection framework.


Paper cited: "Leveraging multi-level dependency of relational sequences for social spammer detection", 2020.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Graph Neural Networks (GNNs) alongside sequence modeling for social spammer detection in multi-relational networks.
  • Which research first introduced the use of Residual Networks for modeling high-order interactions in sequential recommendation, and how did this paper adapt that framework for security tasks?
  • Explore how the Multi-level Dependency Model (MDM) framework could be extended to detect fraudulent activities in financial transaction sequences or multi-platform e-commerce logs.
Contents
MDM: Cracking Social Spammer Detection via Multi-Level Relation Dependencies
1. TL;DR
2. The Evolution of Spammer Detection
3. Methodology: The Three-Pillar Architecture
3.1. 1. User-Relation Representation
3.2. 2. Long-term Dependency Modeling
3.3. 3. Short-term Dependency Modeling (The Core Innovation)
4. Experimental Validation
4.1. SOTA Comparison
4.2. Ablation Study: What makes it work?
5. Critical Insight: The "Pet Game" Anomaly
6. Conclusion and Future Outlook