Inferring Gender from Sparse Interest Tags: Beyond Lexical Analysis
Inferring Users’ Gender from Interests: A Tag Embedding Approach
This paper introduces a conceptual class (CC) based framework for predicting user gender on social media by leveraging user interest tags. By combining generalization/specification operations with Word2Vec-based tag embeddings, the authors successfully map sparse user profile tags into a dense, semantically enriched feature space.
TL;DR
Gender prediction in social media is typically a text-mining task, but what if a user never posts? This paper presents a logic-driven embedding approach that uses Conceptual Classes to bridge the gap between sparse, idiosyncratic user tags (like "Red Devil") and broad demographic categories. By expanding tags into a dense vector space, the authors boosted classification accuracy from a mediocre 62.75% to a robust 82.25%.
Problem & Motivation
Most gender classification research focuses on how people write—analyzing punctuation, slang, or n-grams in their posts. However, many users are "lurkers" who consume content without creating it. For these users, the only available data is their interest tags.
The challenge? These tags are:
- Extremely Short: Users usually list fewer than 10 tags.
- Sparse: The vocabulary of tags is vast and unnormalized (e.g., "PES" vs. "Pro Evolution Soccer").
- Disconnected: Standard models treat tags as isolated tokens, failing to realize that "High-heeled shoes" and "LotionSPA" belong to the same gendered interest manifold.
Methodology: The Conceptual Class Framework
The researchers' core insight is that tags are not just strings; they are pointers to broader concepts. They propose a three-stage pipeline to concentrate the feature space.
1. Building Initial Conceptual Classes (ICC)
The authors identified the most "discriminating" tags—those that appear disproportionately in one gender. They then performed Generalization (finding a superclass) and Specification (finding a subclass/instance).
- Logic: If "Kobe" is the tag, the superclass is "Basketball" and another instance might be "Yao Ming".
2. Tag Embedding with Skip-Gram
To avoid purely manual work, they used Word2Vec (Skip-Gram) to learn tag representations. By setting a window size of 10 (the max tags per user), the model learns that tags appearing together in a profile are semantically related.
3. Expansion and Condensation
The initial classes were expanded iteratively. The authors tested three strategies to represent a Conceptual Class as a single vector:
- AVG: Average of tag vectors.
- DPT: Dot product of vectors.
- MST: Finding the most similar single tag.

Experiments & Results
The study utilized a dataset from Sina Weibo, comparing their method against content-based (FRC) and distributed tag (DRT) baselines.
SOTA Comparison
The results clearly show the superiority of the "Expanded Conceptual Class" (ECC):
- FRT (Frequency-based Tag): 62.75% (Baseline)
- FRC (Microblog Content): 77.25% (Semantic text)
- ECC (Our Method): 82.25% (Winner)

Why AVG Won
In the ablation study of expansion strategies, the AVG (Average) strategy consistently outperformed DPT and MST. This suggests that a conceptual class is best represented by the centroid of its members in the embedding space rather than a single representative point.
Critical Analysis & Conclusion
Takeaway
The paper proves that structured knowledge (Conceptual Classes) and unsupervised learning (Embeddings) are a powerful duo. By "pseudo-labeling" sparse tags with their higher-level concepts, the model effectively performs a form of data augmentation that makes the classifier's job much easier.
Limitations
Despite its success, the initial seed generation still requires manual intervention (WikiTaxonomy alignment). In the era of Modern LLMs, this step could likely be automated using few-shot prompting to generate the initial classes.
Future Outlook
This framework isn't limited to gender. It can easily be extended to age prediction, career classification, or personality trait inference. As social media platforms move toward privacy-preserving "minimal data" profiles, being able to infer much from very few tags is a vital capability for recommendation engines.
