Social Impact: Unmasking Misattributed and Distorted Quotes in the Social Media Wild

Social Impact - Identifying Quotes of Literary Works in Social Networks

2015-01-01
Carlos Barata, Mónica Abreu, Pedro Torres, Jorge Teixeira, Tiago João Guerreiro, Francisco M. Couto
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces "Social Impact," a Service-Oriented Architecture (SOA) framework designed to identify and link literary quotes shared on social networks to their original works. Utilizing SocialBus for data collection and Apache Lucene for indexing, the system achieved high precision in mapping noisy social media text to authoritative corpora.

TL;DR

"Social Impact" is a robust architectural framework designed to solve the problem of "floating quotes"—literary snippets shared on Twitter and Facebook without proper attribution or with significant textual corruption. By treating social media messages as search queries against a canonical library (indexed via Apache Lucene), the researchers achieved up to 100% precision in linking noisy social posts to the original works of authors like Fernando Pessoa.

The Problem: The "Digital Whisper" Effect

Social networks are massive repositories of cultural heritage, yet information there is often degraded. Quotes are frequently:

  • Incomplete or Erroneous: Users add slang, synonyms, or make typos.
  • Misattributed: Quotes are often linked to the wrong author, creating a cycle of misinformation.
  • Context-Poor: Short messages and a lack of hashtags (only 4%-25% of tweets use them) make automated discovery a nightmare.

Current search methods often fail because they look for exact matches, ignoring the "noisy" nature of human sharing.

Methodology: High-Precision Retrieval

The authors proposed the Social Impact Platform, built on a Service-Oriented Architecture (SOA). The core innovation lies in its Quotes Detector module.

The Two-Fold Workflow

  1. Indexation of Truth: The system first ingests a clean "External Knowledge" base (e.g., the complete poems of Fernando Pessoa). These are pre-processed (stopwords removed) and indexed using Apache Lucene.
  2. Querying the Noise: Incoming messages from SocialBus (a specialized crawler) are treated as search queries. The system calculates a relevance score; if the score exceeds a specific threshold, the message is officially "linked" to the literary original.

Social Impact Architecture Figure 1: The three-layer architecture spanning External Knowledge, the Backend processing core, and the Application API.

Case Studies: Poetry and Music

The framework was validated through two distinct lenses:

  • O Mundo em Pessoa: Focused on the complex work of Portuguese poet Fernando Pessoa and his many heteronyms.
  • Lusica: Focused on Lusophone (Portuguese-speaking) music lyrics.

Quotes Detector Detail Figure 2: The inner workings of the Quotes Detector, showing the dual flow of canonical indexing and real-time social message matching.

Experimental Results & Insights

The system showed a remarkable ability to filter "noise" from "signal."

  • Precision vs. Recall: In both studies, precision was near perfect (98% to 100%). This means if the system says a tweet is a quote, it almost certainly is. However, recall was lower (53%-59%), indicating that highly distorted quotes or very short fragments still evade detection.
  • Performance: With an execution time of 0.01s to 0.02s per message, the system is capable of handling real-time social media streams at scale.
  • Data Reality Check: Only about 5%-8% of messages mentioning an author actually contain a quote; the rest are merely mentions or general discussion.

Critical Analysis & Future Outlook

The "Social Impact" framework proves that Information Retrieval (IR) techniques are highly effective for cultural data cleaning. However, the reliance on a manual Lucene score threshold is a limitation.

The Path Forward:

  • Machine Learning Integration: Future iterations could use ML to dynamically set thresholds, potentially catching those "low-score" matches that currently hurt recall.
  • Beyond Text: While currently focused on literature and music, the architecture is "abstract enough" to be applied to plagiarism detection or tracking the spread of political slogans and "fake news" quotes.

By bridging the gap between disorganized social streams and structured literary archives, Social Impact provides a vital tool for digital humanities and social media hygiene.

Find Similar Papers

Try Our Examples

  • Search for recent papers that use Machine Learning or Transformers to improve the recall of noisy quote detection in social media beyond traditional Lucene scoring.
  • Which early studies first defined the "SocialBus" architecture for focused crawling of regional social network data, and how does this paper extend its utility?
  • Examine how the "Social Impact" framework's indexing approach could be applied to automated plagiarism detection in academic or journalistic contexts.
Contents
Social Impact: Unmasking Misattributed and Distorted Quotes in the Social Media Wild
1. TL;DR
2. The Problem: The "Digital Whisper" Effect
3. Methodology: High-Precision Retrieval
3.1. The Two-Fold Workflow
4. Case Studies: Poetry and Music
5. Experimental Results & Insights
6. Critical Analysis & Future Outlook