Recognizing Human Behaviours in OSNs: A Proactive Defense via "Normality" Modeling
Recognizing human behaviours in online social networks
The paper proposes a novel "two-step" framework for identifying human behaviors in Online Social Networks (OSNs) to detect anomalies. It utilizes Markov Chains to model normal interaction patterns and a probabilistic unexplained activity detection framework based on "possible worlds" theory to flag deviations as potential cybersecurity threats.
TL;DR
Instead of playing a perpetual game of "cat and mouse" by trying to model every possible malware pattern, this research flips the script. By using Markov Chains to learn the signature of a "normal" user’s actions on Facebook, the authors developed a framework that flags anything it can't explain. This probabilistic approach excels at catching unknown threats (Zero-day attacks) and identifying new, legitimate platform features.
The "Bad Behavior" Bottleneck
In the world of Online Social Networks (OSNs), cybersecurity traditionally focuses on signature-based detection—identifying known "bad" patterns like spamming or phishing. However, this work identifies a critical flaw: attackers are adaptive. As soon as a defense is built against one pattern, they pivot to another.
The authors argue that it is mathematically and practically easier to define what a normal human does than to catalog every possible way a bot or a malicious actor might behave.
Methodology: The Architecture of Normality
The proposed framework operates in two distinct stages: Model Estimation and Unexplained Behaviour Detection.
1. Modeling the Baseline
The system processes raw OSN logs into high-level categories (e.g., "Status Update", "Messages", "Shares"). By grouping low-level actions, the authors reduce the state space, making the models more generalizable. These sequences are then used to train First-Order Markov Chains.
Figure: A simplified Markov transition graph representing a known, normal interaction sequence.
2. The Possible Worlds Logic
Detecting "unexplained" behavior isn't just about finding a sequence that doesn't exist in the training data. It’s about dealing with uncertainty. A single log entry might be part of several different potential behaviors.
The authors use a system of non-linear constraints (NLC) to calculate the probability of "possible worlds"—scenarios where certain sequences are explained and others are not. If a sequence persists as "unexplained" across the most probable scenarios and exceeds a threshold , it is flagged.
Figure: The system architecture, integrating Spark for distributed data processing and the TUB (Totally Unexplained Behaviours) algorithm.
Experimental Validation
The authors didn't just use synthetic data; they built a Facebook application to collect 2 years of real user interaction logs ( volunteers).
Key Performance Metrics:
- Accuracy: When a known behavior was intentionally hidden from the system (treated as an anomaly), the model identified it as unexplained with ~68-80% accuracy.
- Scalability: Using Apache Spark and HDFS, the system demonstrated it could handle the computational load of large logs, though execution time increases as the probability threshold decreases (due to the larger search space of "possible worlds").
Figure: Experimental results showing the framework's ability to detect malicious vs. non-malicious unexplained sequences.
Critical Insight: Malicious vs. Non-Malicious
One of the most interesting aspects of this research is the distinction between "Unexplained" and "Malicious."
- Non-Malicious Unexplained: A user starts using a new Facebook feature (like a "Superlike"). The system flags it because it hasn't seen it before. This signals the need to update the model.
- Malicious Unexplained: A spam bot follows a pattern of
Login -> Share Link (n times) -> Logout. This deviates significantly from the complex graph of a human user and is flagged for security intervention.
Conclusion & Future Outlook
This work provides a solid mathematical foundation for behavior-based security in social media. By focusing on the stochastic nature of human interaction, it moves away from rigid rules and toward a flexible, probabilistic detection system.
Future Work will likely involve clustering these models to identify specific user profiles—such as distinguishing between legitimate power users and sophisticated automated bots that attempt to mimic human pacing.
