Cognos: Mastering Expert Search via the Wisdom of the Twitter Crowd

Cognos: Crowdsourcing Search for Topic Experts in Microblogs

2016-01-08
Saptarshi Ghosh, Fabricio Benevenuto, Krishna Gummadi
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces Cognos, a specialized search system designed to identify topic experts in microblogging platforms like Twitter. By leveraging the "wisdom of crowds" through Twitter Lists and their metadata (names and descriptions), Cognos successfully maps users to specific areas of expertise and ranks them using a customized similarity metric.

TL;DR

Finding a real expert in the haystack of millions of social media users is notoriously difficult. While official tools often rely on what users say about themselves (bios), Cognos flips the script by looking at what the crowd says about them. By mining Twitter Lists, Cognos identifies and ranks experts with surgical precision, outperforming official search engines in identifying niche specialists.

The Problem: The "Bio" Trap and the Noise of Tweets

How do you currently find an expert on Twitter? You likely search for keywords in user bios or look for high follower counts. However, this approach has two fatal flaws:

  1. The Satire Problem: A comedian like Jimmy Fallon might put "Astrophysicist" in his bio as a joke. Traditional search engines will then erroneously rank him as a top scientist.
  2. The Sparse Profile Problem: Many true domain experts (e.g., specialized researchers) have minimalist bios or don't tweet the query keywords frequently, making them invisible to content-based algorithms.

Methodology: Crowdsourcing Expertise

The core insight of Cognos is that Twitter Lists are essentially crowdsourced directories. When a user adds an account to a list named "Machine Learning," they are providing a high-quality human annotation of that account's expertise.

The Cognos Pipeline

  1. Metadata Extraction: The system crawls list names (e.g., "TechGurus") and descriptions.
  2. Semantic Processing: Since list names are often short, Cognos uses CamelCase splitting (TechGurus → Tech, Gurus) and edit-distance matching to unify terms across languages (e.g., "Politics" and "Politica").
  3. Ranking Logic: Unlike standard TF-IDF, Cognos uses Cover Density Ranking. It calculates the affinity between the query and the user's list-derived topic vector, weighted by the logarithm of the total number of lists the user appears in.

Model Architecture and Data Flow

Experiments & Results: Cognos vs. The Giants

The authors deployed Cognos "in the wild" and gathered over 2,000 human relevance judgments.

1. Cognos vs. Twitter WTF (Who To Follow)

In a blind test where users didn't know which results came from which system:

  • Cognos won or tied in 52% of cases.
  • Cognos excelled at finding individuals and specialists, whereas Twitter WTF tended to over-prioritize large media organizations or brands.
  • Cognos maintained a Mean Average Precision (MAP) of 0.905, showing extreme consistency across diverse categories like Science, Hobbies, and Business.

2. Scalability and Efficiency

To keep the system updated without crawling all 500 million users, the authors used the HITS algorithm to identify "Hubs"—users who are excellent at curating lists. By only following the lists created by these top hubs, Cognos could discover 70% of new experts while making significantly fewer API calls.

Success Rate Comparison Figure: Distribution of expertise coverage using the hub-based update strategy.

Deep Insight: Why it Works

The "Wisdom of the Crowd" acts as a natural noise filter. While one person might mislabel an account, the aggregate mapping of thousands of users creating lists provides a robust, multifaceted view of a person's professional identity.

Limitations and Future Work

The system's primary vulnerability is List Spamming—the potential for bots to create fake lists to boost an account's ranking. Future iterations will need to incorporate "creativity reputation," weighing lists from established, trustworthy users more heavily than new or suspicious accounts.

Conclusion

Cognos demonstrates that the most valuable data on social networks isn't always what we post, but how we are categorized by others. By moving from "self-assertion" to "crowd-annotation," expert search becomes not just more accurate, but more human.

Find Similar Papers

Try Our Examples

  • Find recent papers that extend the use of crowdsourced folksonomies or list-style metadata for expert discovery in modern social media platforms like LinkedIn or Mastodon.
  • What are the current state-of-the-art methods for defending against "List spamming" or malicious annotation attacks in social network recommendation systems?
  • Which research works first formalized the HITS algorithm for identifying hubs and authorities, and how has this lineage evolved in the context of hypergraph social networks?
Contents
Cognos: Mastering Expert Search via the Wisdom of the Twitter Crowd
1. TL;DR
2. The Problem: The "Bio" Trap and the Noise of Tweets
3. Methodology: Crowdsourcing Expertise
3.1. The Cognos Pipeline
4. Experiments & Results: Cognos vs. The Giants
4.1. 1. Cognos vs. Twitter WTF (Who To Follow)
4.2. 2. Scalability and Efficiency
5. Deep Insight: Why it Works
5.1. Limitations and Future Work
6. Conclusion