Beyond the Bot: Refining Sentiment Analysis through Social Media Signals

Evaluation of a Reusable Technique for Refining Social Media Query Criteria for Crowd-Sourced Sentiment for Decision Making

2019-07-01
Kimberley Hemmings-Jarrett, Julian Jarrett, M. Brian Blake
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces a reusable statistical technique to refine social media filtration criteria for sentiment analysis, focusing on the 2016 US Presidential Election data. It leverages Social Media Signals (SMS)—Origin, Originality, and Participation—to improve the accuracy of decision-making for various data consumers.

TL;DR

In the noisy ecosystem of social media, determining genuine public sentiment is a minefield of "data pollution." This paper presents a reusable statistical framework to evaluate how Origin (Human vs. Bot), Originality (Mention vs. Retweet), and Participation (User activity levels) influence decision-making. Testing against 2016 US Election data, the study surprisingly finds that retweets—not bots—are the primary drivers of sentiment distortion in aggregated data.

Background: The Three Types of Data Consumers

The authors categorize Twitter data users into three groups:

  1. Primary: Direct users who view or create content.
  2. Secondary: Researchers and pollsters who aggregate data into sentiment graphs.
  3. Tertiary: Decision-makers who consume those high-level summaries.

The "gap" exists when secondary users fail to filter signals properly, leading tertiary users to make decisions based on skewed or hyper-inflated data.

The "Pollution" Problem: Insights vs. Noise

Existing sentiment analysis often treats every tweet as an equal "vote" of opinion. However, the authors argue that the "humanity" of a tweet is often less impactful than its "repetitiveness." While much of the academic hype focuses on Bot Detection, this study investigates whether the noise is actually coming from the structural characteristics of the platform, such as the retweet mechanism.

Methodology: The Reusable Technique

The paper proposes a specialized pre-processing pipeline to categorize every tweet based on three "Social Media Signals" (SMS):

  • Origin: Using Botometer, accounts are scored 0-5. A score >3 is classified as a bot.
  • Originality: Distinguishing between "Mentions" (original thoughts) and "Retweets" (propagation).
  • Participation: Categorizing users as Passive (1 tweet), Active (2-7), or Junkies (>7).

Model Architecture

The workflow below demonstrates the iterative relationship between the system and the user to arrive at "High Confidence Data" (HCD).

Reusable technique for interpreting influence of SMS

Experimental Results: The Retweet Revelation

The researchers analyzed 6,375 tweets from a 2016 Presidential Debate. The findings were stark:

  • Distribution: Humans dominated bots 11:1, but retweets outnumbered original mentions 3:1.
  • Correlation: There was a near-perfect correlation (r=0.837) between participation levels and retweet counts, suggesting "Junkies" primarily retweet rather than create.
  • Regression Analysis: Using Binary Logistic Regression, the presence of a retweet increased the odds of predicting sentiment by a factor of over 4.0, significantly higher than the impact of an account being a bot.

Performance Comparison

The most dramatic evidence of the model's value appeared when comparing unfiltered vs. filtered sentiment.

Filter CriteriaSentiment Result
Unfiltered DatasetPrimarily Positive
HCD Filtered (Human + Passive + Mentions Only)Inverted to Negative

This "sentiment inversion" proves that without filtering for original, human-generated thoughts, analysts are likely seeing a mirage created by platform mechanics.

Human:bot distribution across different election-cycle datasets

Deep Insight: Why Why This Matters

The core takeaway is that automation is not the only enemy. The structural nature of social media—where a single "Junkie" user can amplify a specific sentiment through hundreds of retweets—is enough to flip the results of a high-stakes sentiment analysis.

Limitations & Future Work

The study was conducted post-hoc, and many accounts were "purged" by Twitter before they could be analyzed by Botometer. Future iterations of this reusable technique aim to integrate qualitative "Tertiary User" preferences to better understand how decision-makers interpret these aggregated visualizations.

Final Conclusion

To achieve "High Confidence Data," researchers must stop treating all social media signals as equal. By prioritizing the removal of retweets and highly active "junkie" accounts, we can bridge the gap between social media noise and actual human opinion.

Find Similar Papers

Try Our Examples

  • Find recent papers that compare the impact of social media bots versus retweet cascades on the spread of political misinformation.
  • Which study first defined the "Information Retrieval by Reformulation" principle, and how has it been adapted for modern social media APIs?
  • Explore how the "Social Media Signals" (origin, originality, participation) framework has been applied to sentiment analysis in the finance or public health sectors.
Contents
Beyond the Bot: Refining Sentiment Analysis through Social Media Signals
1. TL;DR
2. Background: The Three Types of Data Consumers
3. The "Pollution" Problem: Insights vs. Noise
4. Methodology: The Reusable Technique
4.1. Model Architecture
5. Experimental Results: The Retweet Revelation
5.1. Performance Comparison
6. Deep Insight: Why Why This Matters
6.1. Limitations & Future Work
7. Final Conclusion