CONcISE: Balancing Speed and Accuracy in Cyberbullying Detection via Sequential Analysis
Robust Detection of Cyberbullying in Social Media
The paper introduces CONcISE, a novel framework for cyberbullying detection that treats harassment as a sequential process rather than isolated events. It utilizes sequential hypothesis testing and online streaming feature selection to achieve high classification accuracy while significantly reducing computational overhead and detection latency.
TL;DR
Cyberbullying is not a single toxic comment; it is a repetitive, evolving process. Traditional AI models treat every message as an isolated data point, leading to slow response times and excessive resource consumption. This paper proposes CONcISE, a framework that treats detection as a sequential hypothesis testing problem. By optimizing when to "stop" looking at features, it detects bullying faster and more accurately than traditional SOTA models using only a fraction of the computational power.
Background: The Scalability-Timeliness Paradox
Modern social media platforms like Instagram generate millions of comments per hour. Current detection systems face a three-way struggle:
- Accuracy vs. False Positives: Spotting a single swear word is easy, but labeling it "bullying" without context creates false alarms.
- Latency: Waiting for a whole "session" to finish before analyzing it (offline detection) means the victim has already been harmed.
- Scalability: Evaluating hundreds of features for every single comment in a stream of millions is computationally prohibitive.
The author's insight is that we don't need to look at every feature to know if a comment is aggressive, and we don't need to see every comment to know if a user is being bullied.
Methodology: The Optimization of "Stopping"
CONcISE operates on a two-stage logic. Instead of a "one-size-fits-all" feature set, it dynamically decides how much information it needs for each comment.
1. Comment-Level Sequential Testing
The model uses a Sequential Probability Ratio Test (SPRT)-like logic. For each comment, it evaluates features (like profane unigrams) one by one. After each feature, it calculates the belief (posterior probability) of whether the comment is aggressive.
- The Decision Rule: It computes the cost of stopping now (potential misclassification) versus the cost of evaluating one more feature (computational cost).
- The Math: By applying Bellman’s principle of optimality, the system finds the "Stopping Time" () where the expected total cost is minimized.

2. Session-Level Alerting
To avoid "trigger-happy" alerts, CONcISE maintains a session-level counter. Only when the number of detected aggressive comments exceeds a predefined threshold () is a cyberbullying alert raised to moderators.
Experiments & Results: Doing More with Less
Tested on a massive Instagram dataset (over 9M comments), CONcISE was compared against traditional classifiers like Random Forests (RF) and Robust Discriminative Feature Selection (RDFS).
Key Performance Hits:
- Feature Efficiency: While traditional models used 10 to 115 features for every decision, CONcISE variants reached decisions using an average of only 2.97 to 3.92 features.
- Detection Speed: In session-level detection, it saved a significant number of comments compared to offline baselines, essentially "pruning" the conversation history to reach a conclusion.
- Throughput: The system can process between 30,000 to 45,000 comments per second, proving it is ready for the "staggering rates" of modern social media.

As shown in the table above, CONcISE-10 achieved a Recall of 0.794, significantly higher than the RDFS baseline (0.449), while using far fewer comments to reach that conclusion.
Critical Analysis & Professional Insight
The brilliance of the CONcISE framework lies in its Inductive Bias: it assumes that the most informative signals appear early or in clusters. By formulating feature selection as an Optimal Stopping problem, it mirrors how a human moderator might skim a thread—looking for just enough evidence to hit the "block" button.
Limitations:
- The current model relies heavily on a priori probabilities (expert-labeled seeds), which might not capture evolving slang or "dog whistles" effectively.
- The methodology currently assumes features are independent under each hypothesis, a common simplification in sequential testing that might overlook complex linguistic dependencies.
Future Outlook
The next frontier for this work is moving into Unsupervised Learning. If the system can learn the "cost of bullying" without thousands of human-labeled examples, it could be deployed across diverse languages and niche platforms (like Discord or Twitch) where linguistic norms change daily.
Takeaway for Architects
When building real-time moderation pipelines, stop asking "how accurate can this model be?" and start asking "how little data do I need to be accurate enough to act?"
