SCMGR: Rethinking Social Summarization through the Lens of Multi-Granularity Relations
SCMGR: Using Social Context and Multi-Granularity Relations for Unsupervised Social Summarization
The paper introduces SCMGR (Social Context and Multi-Granularity Relations), an unsupervised framework for social media summarization. It utilizes a GCN-based encoder to aggregate social context and a multi-granularity decoder to reconstruct semantic and social structures, achieving new SOTA results on English (TWEETSUM) and Chinese (Weibo) corpora.
TL;DR
Summarizing social media is a nightmare due to "informational poverty"—posts are too short to stand alone. SCMGR solves this by treating a collection of posts not just as a bag of words, but as a living social graph. By using Graph Convolutional Networks (GCN) to "borrow" meaning from a post's neighbors (social context) and decoding both semantic and social relations, it produces summaries that are more diverse and representative than text-only methods.
The "Short-Text" Bottleneck
In traditional NLP, we assume a document contains enough signal to summarize itself. On Twitter or Weibo, a post like "It’s finally here!" is semantically useless without knowing the user was replying to an iPhone launch announcement.
Existing SOTA methods usually fail because:
- Sparsity: Short texts result in high-dimensional, sparse vectors.
- Isolation: They treat posts as independent units, ignoring the "social contagion" and "social consistency" that link users and their opinions.
Methodology: The Core of SCMGR
The authors propose an encoder-decoder framework that doesn't just look at what was said, but who said it and who they talk to.
1. Constructing the Social Context
The model identifies relationships using two sociological meta-paths:
- P-U-P (Post-User-Post): Connects posts by the same author (consistency of opinion).
- P-U-U-P (Post-User-User-Post): Connects posts by friends (influence and shared interests).
These are fed into a Residual GCN. Unlike standard GCNs, the residual connection prevents "over-smoothing," ensuring the model doesn't lose the unique signal of the original post while aggregating context from 1-hop neighbors.

2. Multi-Granularity Decoding
The innovation lies in the Decoder. It forces the hidden representations to be powerful enough to reconstruct two different graphs:
- The Semantic Graph: A bipartite graph linking posts to specific words.
- The Social Graph: The actual network structure of user interactions.
By minimizing the reconstruction loss for both, the final post embeddings are "double-charged" with both meaning and social importance.
Performance & Insights
The results on TWEETSUM and Weibo datasets show SCMGR consistently beating baselines like LexRank and PacSum.

Ablation Deep-Dive:
- Social > Semantic? Interestingly, the study found that removing social relation guidance caused a bigger performance drop than removing semantic guidance. This suggests that in the chaotic world of social media, who is talking is often a stronger signal of "what is important" than the specific vocabulary used.
- Context is King: Removing the GCN-based context aggregation led to the most dramatic failure, confirming that social context is the primary "de-noiser" for short-form text.
Visualization: The "Global View"
The authors visualized the social context graph, marking the selected summary posts in red. The summaries aren't just clustered in one high-density area; they are evenly distributed across the network's key "bridge" positions (high betweenness centrality). This ensures the summary covers multiple viewpoints rather than just echoing the loudest user.

Conclusion
SCMGR proves that unsupervised social summarization is a structural problem as much as a linguistic one. By moving from text-centric to graph-centric representations, we can generate summaries that finally capture the diversity and nuances of social discourse without needing a single labeled training example.
