Unmasking the Anomalous: A Probabilistic Approach to Detecting Unexplained Behaviors in Social Networks
Detecting Unexplained Human Behaviors in Social Networks
This paper introduces a two-stage probabilistic framework for detecting anomalous human behaviors in Online Social Networks (OSNs). It utilizes Markov Chains to model "normal" user activities from clickstream data and a "Possible Worlds" reasoning engine to identify unexplained activity sequences (Totally or Partially Unexplained Behaviors) that deviate from established patterns.
TL;DR
As Online Social Networks (OSNs) move from simple connection hubs to critical infrastructure for personal and professional life, the risk of identity theft and malicious exploitation rises. This paper presents a sophisticated two-stage framework that learns "normal" clickstream patterns via Markov Chains and identifies anomalies using a Possible Worlds probabilistic model. It doesn't just look for bad behavior; it identifies unexplained behavior.
Context & Motivation: Beyond the Social Graph
Most OSN security research focuses on the Social Graph—who you are connected to. However, the authors argue that the real story is in the Clickstream—the sequence of actions a user takes.
The problem is that human behavior is stochastic and often overlapping. A user might be checking notifications while uploading a photo. Traditional "if-then" rules fail to capture this complexity. The authors' intuition is that we can define "unexplained" events as subsequences of data that existing behavior models simply cannot account for with high confidence.
Methodology: The Two-Stage Engine
1. Learning the "Normal"
First, the system needs to know what typical behavior looks like. The authors group low-level actions into high-level categories (e.g., "Status & Friends", "Shares", "Like") to reduce dimensionality. They then build First-order Markov Chains where:
- Nodes are action categories.
- Edges represent the probability of transitioning from one action to another within a specific time bound ().
2. The Possible Worlds Reasoning
When a new log arrives, the system attempts to match it against these Markov models. If multiple models match the same sequence of events, a Conflict occurs. To resolve this, the paper uses the concept of Possible Worlds:
- A Possible World is a subset of behavior occurrences that do not conflict (i.e., they don't share the same log entry).
- The system solves a system of Non-Linear Constraints (NLC) to assign a probability distribution to these worlds.
Figure: The modular system architecture from Data Acquisition to TUB/PUB Detection.
3. TUB vs. PUB
The paper introduces two critical metrics:
- Totally Unexplained Behavior (TUB): A sequence where no known model explains the entries. This is the "maximal" sequence of weirdness.
- Partially Unexplained Behavior (PUB): A sequence where at least one entry is unexplained. This is the "minimal" unit of anomaly.
Experiments and Results
The authors tested their prototype on real Facebook user sessions gathered over a 2-year period.
Performance & Scalability
The framework was tested on logs ranging from 6 to 24 hours. As shown in the figures below, the execution time for detecting TUBs and PUBs increases with the length of the log but remains within efficient bounds for near-real-time monitoring.
Figure: Execution time comparison varying log length (L) and probability threshold ().
Accuracy Through Ablation
To verify if the system actually catches unknown behavior, the authors performed a "leave-one-out" test. They removed a known behavior model (e.g., "Behavior 1") and checked if the system flagged those specific activities as "unexplained." The results confirmed high accuracy, particularly as the threshold () was adjusted, proving the models can effectively segregate novel patterns.
Critical Insight: The "Silence" Matters
The most profound takeaway is that this method treats unexplained activity as a first-class citizen. In cybersecurity, we often look for "signatures" of known attacks. This paper flips the script: by rigorously modeling the "known good," it defines a mathematical boundary for the "unknown," which is where attackers, bullies, and bots typically reside.
Future Outlook
While the Markov Chain approach is robust, the authors acknowledge the need to scale to even larger datasets. The integration of this "Possible Worlds" logic with modern Deep Learning backends could potentially allow OSN providers to automate the identification of new social trends AND new security threats simultaneously.
Keywords: Anomaly Detection, Markov Chains, Online Social Networks, Cyber Security, Clickstream Analysis.
