Unmasking the Anomalous: A Probabilistic Approach to Detecting Unexplained Behaviors in Social Networks

Detecting Unexplained Human Behaviors in Social Networks

2014-06-01
Flora Amato, Aniello De Santo, Vincenzo Moscato, Fabio Persia, Antonio Picariello
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a two-stage probabilistic framework for detecting anomalous human behaviors in Online Social Networks (OSNs). It utilizes Markov Chains to model "normal" user activities from clickstream data and a "Possible Worlds" reasoning engine to identify unexplained activity sequences (Totally or Partially Unexplained Behaviors) that deviate from established patterns.

TL;DR

As Online Social Networks (OSNs) move from simple connection hubs to critical infrastructure for personal and professional life, the risk of identity theft and malicious exploitation rises. This paper presents a sophisticated two-stage framework that learns "normal" clickstream patterns via Markov Chains and identifies anomalies using a Possible Worlds probabilistic model. It doesn't just look for bad behavior; it identifies unexplained behavior.

Context & Motivation: Beyond the Social Graph

Most OSN security research focuses on the Social Graph—who you are connected to. However, the authors argue that the real story is in the Clickstream—the sequence of actions a user takes.

The problem is that human behavior is stochastic and often overlapping. A user might be checking notifications while uploading a photo. Traditional "if-then" rules fail to capture this complexity. The authors' intuition is that we can define "unexplained" events as subsequences of data that existing behavior models simply cannot account for with high confidence.

Methodology: The Two-Stage Engine

1. Learning the "Normal"

First, the system needs to know what typical behavior looks like. The authors group low-level actions into high-level categories (e.g., "Status & Friends", "Shares", "Like") to reduce dimensionality. They then build First-order Markov Chains where:

  • Nodes are action categories.
  • Edges represent the probability of transitioning from one action to another within a specific time bound ().

2. The Possible Worlds Reasoning

When a new log arrives, the system attempts to match it against these Markov models. If multiple models match the same sequence of events, a Conflict occurs. To resolve this, the paper uses the concept of Possible Worlds:

  • A Possible World is a subset of behavior occurrences that do not conflict (i.e., they don't share the same log entry).
  • The system solves a system of Non-Linear Constraints (NLC) to assign a probability distribution to these worlds.

Model Architecture Figure: The modular system architecture from Data Acquisition to TUB/PUB Detection.

3. TUB vs. PUB

The paper introduces two critical metrics:

  • Totally Unexplained Behavior (TUB): A sequence where no known model explains the entries. This is the "maximal" sequence of weirdness.
  • Partially Unexplained Behavior (PUB): A sequence where at least one entry is unexplained. This is the "minimal" unit of anomaly.

Experiments and Results

The authors tested their prototype on real Facebook user sessions gathered over a 2-year period.

Performance & Scalability

The framework was tested on logs ranging from 6 to 24 hours. As shown in the figures below, the execution time for detecting TUBs and PUBs increases with the length of the log but remains within efficient bounds for near-real-time monitoring.

Experimental Results Figure: Execution time comparison varying log length (L) and probability threshold ().

Accuracy Through Ablation

To verify if the system actually catches unknown behavior, the authors performed a "leave-one-out" test. They removed a known behavior model (e.g., "Behavior 1") and checked if the system flagged those specific activities as "unexplained." The results confirmed high accuracy, particularly as the threshold () was adjusted, proving the models can effectively segregate novel patterns.

Critical Insight: The "Silence" Matters

The most profound takeaway is that this method treats unexplained activity as a first-class citizen. In cybersecurity, we often look for "signatures" of known attacks. This paper flips the script: by rigorously modeling the "known good," it defines a mathematical boundary for the "unknown," which is where attackers, bullies, and bots typically reside.

Future Outlook

While the Markov Chain approach is robust, the authors acknowledge the need to scale to even larger datasets. The integration of this "Possible Worlds" logic with modern Deep Learning backends could potentially allow OSN providers to automate the identification of new social trends AND new security threats simultaneously.


Keywords: Anomaly Detection, Markov Chains, Online Social Networks, Cyber Security, Clickstream Analysis.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Clickstream analysis and Markov Models for Sybil detection in decentralized social networks (DeSo).
  • Which original studies established the "Possible Worlds" probabilistic framework for video surveillance, and how does this paper adapt those constraints for discrete social log events?
  • Are there any modern implementations of this behavioral anomaly detection using Graph Neural Networks (GNNs) or Transformers instead of Markov Chains?
Contents
Unmasking the Anomalous: A Probabilistic Approach to Detecting Unexplained Behaviors in Social Networks
1. TL;DR
2. Context & Motivation: Beyond the Social Graph
3. Methodology: The Two-Stage Engine
3.1. 1. Learning the "Normal"
3.2. 2. The Possible Worlds Reasoning
3.3. 3. TUB vs. PUB
4. Experiments and Results
4.1. Performance & Scalability
4.2. Accuracy Through Ablation
5. Critical Insight: The "Silence" Matters
6. Future Outlook