Beyond Sentiment: Linking High-Quality Civic Feedback to Political Discourse

What's Public Feedback? Linking High Quality Feedback to Social Issues Using Social Media

2012-09-01
Swapna Gottipati, Jing Jiang
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a novel framework for linking high-quality public feedback from social media to specific issues discussed in political speeches. By combining supervised learning for quality assessment with a semi-supervised topic model (SLDA) for issue alignment, the authors achieve SOTA performance in ranking relevant, well-justified public opinions.

TL;DR

Government officials often struggle to process the "firehose" of social media comments following major policy speeches. This paper presents a framework that filters out "noise" to find high-quality, justified public feedback and automatically links it to specific issues (like Healthcare or Economy) using a specialized topic model called SLDA.

Background: The Web 2.0 Policy Dilemma

In the era of Web 1.0, governments spoke, and citizens listened. In Web 2.0, the surge of online forums and social networks has created a goldmine of public opinion. However, for a policy maker reading 550 comments on a National Day Rally speech, the challenge isn't finding opinions—it's finding substance. Most comments are either repetitive ("I agree!") or lack justification. This paper addresses the "Needle in the Haystack" problem: Identifying the feedback that actually provides reasoning and linking it to the correct part of the policy speech.

The Core Insight: Quality + Relevance

The researchers argue that a comment is only useful to a policy maker if it satisfies two conditions:

  1. High Quality: It provides justification through elaboration, comparison, or examples.
  2. Relevance: It clearly maps to a specific issue discussed in the speech.

1. The Quality Sieve

Instead of just looking at keywords, the authors use Discourse Features. They look for structural markers of reasoning (e.g., "because," "for example," "however"). By training a Logistic Regression model on features like verb phrase density and discourse relations, they can automatically flag comments that offer "well-justified" perspectives.

2. SLDA: Structuring the Conversation

Standard Topic Models (like LDA) often miss the mark because they don't know the "context" of the speech. The authors introduced Speech-aligned LDA (SLDA).

Model Architecture

As shown in the graphical model above, SLDA treats each pre-defined issue in the speech as a "fixed anchor" for topics, then forces the model to align user comments to these specific anchors. This ensures that the "relevance" score is grounded in the actual content of the political address.

Experimental Battleground: Obama vs. Lee Hsien Loong

The model was tested on two diverse datasets: US President Obama’s State of the Union Speech and the Singapore Prime Minister’s National Day Rally Speech.

Key Performance Metrics

The results demonstrated that the proposed SLDA model, especially when weighted with the Quality score (), significantly boosted the precision of retrieved feedback.

Experimental Results

  • The Precision Boost: For popular issues like "Immigration," the model achieved a staggering 100% Precision@10.
  • The Robustness: Unlike Bag-of-Words (BOW) models, SLDA was able to link comments containing informal language or slang to the correct formal policy topics by understanding the underlying distribution of words.

Deep Insight: Discovering "Feedback Words"

One of the most valuable outputs of this research is the extraction of "feedback words"—terms used by the public that weren't in the original speech but are highly correlated with specific issues.

  • Example: In Obama's speech regarding "Taxes," the model identified the word "food" as a high-probability feedback word. This revealed that the public's primary concern regarding tax policy was its impact on food security for the poor—a nuance that a simple keyword search might miss.

Critical Analysis & Conclusion

This paper serves as a bridge between traditional Political Science (surveys) and modern Data Science. By focusing on justification rather than just sentiment, it elevates social media mining from "counting likes" to "understanding why."

Limitations: The approach currently relies on manual segmentation of the target speech. Future iterations could benefit from automated speech segmentation and the use of Transformer-based embeddings (like BERT or GPT) to better handle the nuances of informal internet slang.

Future Outlook: As e-governance portals become the primary touchpoint for civic engagement, systems like SLDA will be essential for turning millions of scattered comments into a coherent "Public Report Card" that governments can actually act upon.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Large Language Models (LLMs) to perform zero-shot linking of social media comments to long-form policy documents or articles.
  • Which study first introduced the use of discourse relations (like those in the Penn Discourse Treebank) to measure the "justification quality" or persuasiveness of online text?
  • Explore how semi-supervised Latent Dirichlet Allocation (LDA) variants have been adapted for multi-modal public feedback analysis involving both text and images.
Contents
Beyond Sentiment: Linking High-Quality Civic Feedback to Political Discourse
1. TL;DR
2. Background: The Web 2.0 Policy Dilemma
3. The Core Insight: Quality + Relevance
3.1. 1. The Quality Sieve
3.2. 2. SLDA: Structuring the Conversation
4. Experimental Battleground: Obama vs. Lee Hsien Loong
4.1. Key Performance Metrics
5. Deep Insight: Discovering "Feedback Words"
6. Critical Analysis & Conclusion