Inferring Lurkers’ Gender: Solving the "Silent User" Problem via Interest Tags

Inferring Lurkers’ Gender by Their Interest Tags

2016-01-01
Peisong Zhu, Tieyun Qian, Zhenni You, Xuhui Li
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a specialized framework for predicting the gender of "lurkers"—social media users who consume content without posting—using only their self-selected interest tags. By employing a "Conceptual Class" expansion method based on association rule mining, the authors successfully condense sparse tag spaces to improve classification accuracy on Sina Weibo data.

TL;DR

How do you identify the gender of a social media user who never posts a single word? Researchers from Wuhan University have addressed this "lurker" problem by shifting the focus from what users say (text) to who they want to be (interest tags). By grouping sparse, diverse tags into "Conceptual Classes" and expanding them via association mining, they achieved a significant boost in classification accuracy on Sina Weibo.

The "Lurker" Motivation: Beyond the Vocal Minority

Most social media research suffers from "participation bias"—we study the people who talk. However, in platforms like Sina Weibo, nearly 7.4% of the population are lurkers: users who browse but never post.

Current SOTA methods for gender detection rely on:

  • Lexical features: N-grams and vocabulary.
  • Stylistic features: Punctuation usage and slang.
  • Syntactic features: Sentence structure.

For a lurker, these features are non-existent. The only clue left is the Interest Tags in their profile. But tags are a nightmare for machine learning: they are sparse (usually <6 per person) and highly idiosyncratic (e.g., two fans of different F1 teams might not share a single tag).

Methodology: The Power of Conceptual Classes

The core insight of this paper is that while tags are diverse, the underlying concepts are gender-linked. The authors propose a "Conceptual Class" (CC) framework to condense the sparse tag space.

1. Building the Foundation

They defined 27 conceptual classes based on social and psycholinguistic traits. For instance:

  • Female-oriented: Beauty, Cosmetic, Family, Fashion.
  • Male-oriented: Cars, Technology, Sports(M), Politics.

2. Expanding the Vocabulary (ECC)

To handle the "Long Tail" of tags, the authors developed an Expanding Conceptual Class (ECC) algorithm. Using the Apriori algorithm on a massive unlabeled dataset, they looked for tags that frequently co-occur with the "seed" tags in their conceptual classes.

Concept Expansion Examples Figure 1: Examples of initial vs. expanded tags. Notice how specific product names (Estee Lauder) are correctly associated with the "Cosmetic" concept.

Experiments and Results

The researchers tested their framework on a dataset of 1,000 certified celebrities from Sina Weibo. They compared their method against typical baselines like Screenname N-grams and raw Tag N-grams.

Performance comparison:

Feature MethodAccuracy (%)
Screenname (Char 1-gram)63.67
Tag (Char 1-gram)68.33
Tag Vector (Expanded Conceptual Class)71.33

Minimum Support Table Table 1: The effect of minimum support () on expansion. Choosing an optimal (5%) ensures the conceptual classes remain relevant without becoming noisy.

Why it Works

Raw tag vectors (without CC) achieved only 65.33% accuracy. By introducing the Original Conceptual Class, accuracy jumped to 70.67%, and the Expanded version pushed it even further. This proves that "densifying" the feature space by grouping individual tags into semantic buckets is the key to handling sparse user data.

Critical Insight & Conclusion

This paper highlights a critical shift in User Profiling: when behavioral data (posts) is missing, identity data (tags) becomes the primary signal.

Key Takeaways:

  • Sparsity is the Enemy: Standard N-gram methods fail when users have only 2-3 features.
  • Association Mining as a Bridge: Leveraging unlabeled data to expand small "Seed" sets allows models to understand synonymous or related interests.
  • Future Path: While effective, the 27 classes were manually defined. Future work could benefit from unsupervised topic modeling (like LDA) or Word2Vec embeddings to discover these conceptual classes automatically.

For advertisers and developers, this study provides a blueprint for understanding the "silent majority" of their user base using nothing but a handful of account tags.

Find Similar Papers

Try Our Examples

  • Search for recent studies on "silent user" or "lurker" profiling in social media platforms beyond gender prediction.
  • Which paper first established the psycholinguistic basis for gender-specific interests in social media, and how do modern tag-based methods validate those theories?
  • Investigate how Knowledge Graph embeddings or Transfer Learning could be used to expand "Conceptual Classes" more effectively than traditional association rule mining.
Contents
Inferring Lurkers’ Gender: Solving the "Silent User" Problem via Interest Tags
1. TL;DR
2. The "Lurker" Motivation: Beyond the Vocal Minority
3. Methodology: The Power of Conceptual Classes
3.1. 1. Building the Foundation
3.2. 2. Expanding the Vocabulary (ECC)
4. Experiments and Results
4.1. Performance comparison:
4.2. Why it Works
5. Critical Insight & Conclusion