LTPA: Building Precision User Profiles Through Social Influence and Knowledge Graphs
An Integrated Tag Recommendation Algorithm Towards Weibo User Profiling
The paper introduces LTPA (Local Tag Propagation Algorithm), an integrated tag recommendation framework designed for profiling users on Sina Weibo. It combines social homophily, tag co-occurrence, and a Chinese Knowledge Graph (CKG) to provide personalized tags, achieving superior precision and recall compared to traditional Collaborative Filtering and content-based methods.
TL;DR
Profiling users in microblogging systems like Sina Weibo is a double-edged sword: tags are vital for personalized marketing, yet users rarely provide them due to privacy or effort. This paper presents LTPA, a robust recommendation algorithm that treats user profiling as a social propagation task. By combining social "birds of a feather" logic with a Chinese Knowledge Graph (CKG), it solves the twin hurdles of data sparsity and semantic redundancy.
Problem & Motivation: Why Is User Tagging Different?
Most tag recommenders (like those for Flickr) are designed for objects. If a photo has no tags, we look at similar photos. But if a user has no tags, the problem is deeper.
- Data Sparsity: Nearly half of Weibo users are "blank slates."
- Multi-facet Characteristics: A person isn't just one thing; their tags must cover career, hobbies, and religion without being repetitive.
- The Synonym Trap: Recommending "travel," "voyage," and "trip" to a user who only has 10 tag slots is a waste of digital real estate.
The authors' core insight is that homophily—the tendency for similar people to connect—and influence (retweet behavior) can be used to "flow" tags from active users to silent ones.
Methodology: The Three-Step Pipeline
1. Recommendation by Homophily
Instead of just looking at what a user says, the model looks at who they follow. It uses a Local Tag Propagation mechanism. If User A follows Influencer B, tags from B propagate to A. The weights are determined by interaction intensity (e.g., retweet frequency).
2. Expansion by Co-occurrence
To ensure diversity, the model expands the initial set. If "Machine Learning" is a likely tag, it looks for tags that frequently co-occur in the global dataset, such as "Big Data" or "AI," using a TF-IDF variant to punish generic terms like "Internet."
3. Removing Semantic Redundancy via CKG
This is the most "academic" contribution. The authors built a Chinese Knowledge Graph (CKG) from online encyclopedias (like Baike).
- Mapping: Tags are mapped to CKG concepts.
- Scoring: Using Explicit Semantic Analysis (ESA), the system calculates the cosine similarity of the "concept vectors" of two tags.
- Pruning: If two tags are too semantically close, the lower-ranked one is discarded.
Figure: The entities and relationships within the LTPA framework showing propagation from users to tags via the CKG.
Experiments & Results
The authors compared LTPA against CF (Collaborative Filtering), TF-IDF, and TWEET (keyword extraction).
- Human Assessment: In tests on 500 Weibo users, LTPA consistently achieved the highest Precision and MAP (Mean Average Precision).
- Profiling Accuracy: The algorithm was surprisingly effective at guessing "hidden" traits. It achieved 99.21% accuracy for religion and 95.24% for education by analyzing the social circle's tag clouds.
- Ablation Insight: Without Step 3 (Deduplication), about 15% of the recommended tags were redundant synonyms, proving the necessity of the Knowledge Graph filter.
Figure: Comparison of LTPA against baselines, showing superior performance in Precision and MAP across different top-k recommendations.
Critical Analysis & Conclusion
The Takeaway: Personalization is not just about what you do, but who you are with. LTPA proves that social relationships are high-signal proxies for identity.
Limitations:
- The "Echo Chamber" Effect: By relying heavily on homophily, the system might fail to capture the unique, eccentric interests of a user that differ from their social circle.
- Computational Cost: While limited to a 2-hop radius (r=2), scaling this to a real-time system with 600 million users requires massive graph-processing infrastructure.
Future Outlook: Integrating this with Graph Neural Networks (GNNs) could further refine how "influence" is calculated, potentially moving away from manual TF-IDF weighting toward learned latent representations.
