Social Media Mining: How Business Models and Privacy Settings Gatekeep Big Data

Social Media Mining: Impact of the Business Model and Privacy Settings

2015-08-27
Carsten Ellwein, Benedikt Noller, Ben Noller
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a comparative case study of social media mining on Facebook and Twitter, focusing on how different business models and privacy configurations affect data accessibility. Using the 2015 Nepal earthquake as a benchmark, the study highlights the shift from open academic access to restricted, licensed data models.

Executive Summary

TL;DR: This paper investigates a fundamental but often ignored variable in Data Science: the platform itself. By comparing Facebook and Twitter during the 2015 Nepal earthquake crisis, the authors demonstrate that a platform’s business model (advertising vs. data licensing) and privacy settings are not just technical hurdles—they are filters that fundamentally bias the results of data mining.

Background: Positioned at the intersection of Information Systems and Data Mining, this work acts as a critical reality check for researchers, moving from "blindly mining data" to "critically assessing the source."

The Hidden Barriers: Why "Free" Data Isn't Free

The authors argue that social media platforms are no longer just communication hubs; they are profit-driven entities where users are the primary asset.

  • Twitter's Strategy: Relies heavily on data licensing (the "Firehose"). While it offers a free Streaming API, it is restricted to a 1% sample, potentially skewing results for niche topics.
  • Facebook's Strategy: A "walled garden" approach. Their pivot to Graph API v2.0+ effectively shut down automated mining for independent researchers to protect user privacy and proprietary data value.

The motivation for this study was to prove that these corporate decisions directly lead to Sampling Bias, making the "digital copy" of society provided by these platforms incomplete.

Methodology: A Tale of Two APIs

The researchers used the R statistical environment to pull data using specific search parameters related to the Nepal earthquake (e.g., #nepal, #earthquake).

The Technical Split

  1. Twitter (twitteR package): Automated, metadata-rich, but limited by rate-limiting windows.
  2. Facebook (Rfacebook & Manual): Automated access was blocked by API changes during the study. The authors were forced into manual collection, revealing that keyword searches on Facebook are heavily influenced by the researcher's own social graph—a massive blow to objectivity.

Table 1: Search Parameters for FB and Twitter

Experiments and Key Findings: Facebook vs. Twitter

The study compared the frequency and sentiment of words found in the collected datasets.

Quantifying the Impact

  • Character Limits: Twitter’s 140-character limit (at the time) resulted in lower word frequencies per post but higher metadata density.
  • The Keyword Trap: On Facebook, keyword searches returned results primarily from "connected pages," whereas hashtag searches were more "profile independent."
  • Volume Disparity: The word "Nepal" appeared 1,242 times in the Facebook set compared to only 336 times on Twitter, showcasing how different platform architectures allow for more or less detailed text mining.

Word Cloud for FB Hashtags Figure: The dominance of specific disaster-related stems on Facebook demonstrates high information density but highlights the lack of automated filtering tools compared to Twitter.

Critical Analysis & Conclusion

Takeaway

The core contribution of this paper is the verification that Twitter is the superior platform for academic social media mining due to its API's support for automation and metadata, despite its 1% sampling limit. Facebook, while richer in personal sentiment, has effectively "legally and technically" locked its gates to the research community.

Limitations

  • Temporal Scope: The study was conducted during a specific API transition period (2015), and the landscape has since become even more restrictive.
  • Manual Bias: The forced manual collection on Facebook introduces human error that mirrors the very bias the authors intended to study.

Future Outlook

As major platforms (including Twitter/X) move toward paid API models, the "impact of the business model" discussed here will become the dominant factor in the feasibility of future social media research. The era of "free" academic access to the global digital conversation is rapidly closing.

Find Similar Papers

Try Our Examples

  • Search for recent studies comparing the representativeness of Twitter's 1% Streaming API sample against the full Firehose data in 2024.
  • Which paper first established the conceptual framework for 'Algorithm Bias' or 'Platform Bias' in social media data mining, and how does this study extend that theory?
  • Explore research that applies cross-platform social media mining to humanitarian disaster relief tracking in the era of restricted LLM-based API access.
Contents
Social Media Mining: How Business Models and Privacy Settings Gatekeep Big Data
1. Executive Summary
2. The Hidden Barriers: Why "Free" Data Isn't Free
3. Methodology: A Tale of Two APIs
3.1. The Technical Split
4. Experiments and Key Findings: Facebook vs. Twitter
4.1. Quantifying the Impact
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations
5.3. Future Outlook