ASTC & ASTCx: Unveiling Social Communities through the Lens of Topics and Sentiments

Community discovery using social links and author-based sentiment topics

2014-08-01
Baoguo Yang, Suresh Manandhar
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces ASTC and ASTCx, two probabilistic generative models for community discovery that integrate social links, author-recipient interactions, and sentiment-topic distributions. ASTCx specifically decouples sentiment and topic vocabularies, achieving superior interpretability in uncovering "Topic-Sentiment Unambiguous Communities" across the Enron and Twitter datasets.

TL;DR

In the era of social networking, a community isn't just a set of links; it's a shared conversation with distinct emotional undertones. This paper presents ASTC and ASTCx, two generative models that go beyond network topology to discover communities by blending social links with author-specific sentiment topics. By separating "what" people talk about from "how" they feel, the authors provide a clearer map of social group dynamics.

Background: Why Links Aren't Enough

For years, community detection was a graph theory problem—finding dense clusters of nodes. However, in a corporate email network like Enron or a volatile space like Twitter, a link only tells half the story. Two people might communicate frequently but hold diametrically opposite views on a topic, potentially belonging to different "sentiment communities." Prior works integrated content (topics), but they often ignored the sentiment bias that defines social cohesion or friction.

Methodology: The ASTCx Architecture

The authors propose two versions of their model. While ASTC mixes sentiment and topic words in a single distribution, ASTCx (the extended version) introduces a vital refinement: it treats sentiment words and topic words as separate entities.

1. The Generative Intuition

The model assumes that for every document:

  1. A Community is sampled.
  2. An Author and their Recipients are sampled based on that community.
  3. For every word, a Topic is chosen based on the author's preference within that community.
  4. A Sentiment label is then sampled, conditioned on that specific topic.

2. De-coupling Content and Emotion

By using a subjectivity lexicon (like MPQA) and WordNet, ASTCx separates adjectives and adverbs. This allows the model to learn that "iphone" and "nexus" are topic words, while "amazing" and "terrible" are sentiment markers, preventing the "topic noise" from diluting the "sentiment signal."

Model Architecture Figure 1: Plate notation for the ASTC model, showing the dependency between communities, authors, topics, and sentiments.

Experimental Insights

The researchers validated their models on the Enron email dataset and the Sanders-Twitter Sentiment Corpus.

Diverse Roles in Single Communities

A fascinating finding was the analysis of "active authors." In the Enron dataset, a single community might discuss multiple topics, but different authors within that community have different "dominant" topics. For example, a Vice President (Steven Kean) showed a broad distribution across all topics, reflecting his managerial oversight, while technical heads were more specialized.

The Clarity of ASTCx

The superiority of the ASTCx model is most visible in its output tables. Traditional models produce a "soup" of words. ASTCx produces structured "Sentiment-Topic" clusters:

  • Topic (Google): Positive words: "Sweet", "Nice"; Negative: "Sold", "Wrong".
  • Topic (California Energy): Positive words: "Successful", "Strategic"; Negative: "Crisis", "Refunds".

Results Comparison Table 1: Mixed topic-sentiment words from the basic ASTC model.

ASTCx Separated Results Table 2: Refined sentiment words for selected topics in the ASTCx model, showing much higher readability.

Professional Opinion: Bridging the Gap

The core value of this work lies in its Inductive Bias. By assuming that sentiment is topic-dependent (e.g., your sentiment towards "Microsoft" may differ from your sentiment towards "Apple"), the model mirrors human psychology more accurately than generic clustering.

Limitations

  • Data Sparsity: As the authors noted, the author-recipient relationship is often sparse. In a massive network, many users might only interact once, making the Dirichlet priors difficult to settle.
  • Lexicon Dependency: ASTCx relies on external tools like WordNet. Its performance is capped by the quality of these linguistic resources.

Conclusion

ASTCx provides a robust framework for managers and social analysts to not only see who is talking to whom, but to understand the core ideological or emotional consensus of those groups. Future directions, such as automatically determining the number of communities (M) and topics (K), will be essential for scaling this to the "Big Data" of modern social media.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Author-Recipient-Topic (ART) models specifically for sentiment-based user clustering in social networks.
  • Which paper first proposed the Joint Sentiment/Topic (JST) model, and how does the ASTC model adapt its generative process for community detection?
  • Explore research that applies ASTCx-like sentiment-topic modeling to multi-modal social data, such as images combined with text.
Contents
ASTC & ASTCx: Unveiling Social Communities through the Lens of Topics and Sentiments
1. TL;DR
2. Background: Why Links Aren't Enough
3. Methodology: The ASTCx Architecture
3.1. 1. The Generative Intuition
3.2. 2. De-coupling Content and Emotion
4. Experimental Insights
4.1. Diverse Roles in Single Communities
4.2. The Clarity of ASTCx
5. Professional Opinion: Bridging the Gap
5.1. Limitations
6. Conclusion