Automatic Reputation Computation: A Social Network Approach to Knowledge Mining
Automatic Reputation Computation through Document Analysis: A Social Network Approach
The paper introduces two social network-based algorithms for automatically computing author reputations within the cyber-security domain by analyzing email exchanges. It constructs a directed reference network through entity extraction (IPs, URLs, Domains) and evaluates reputation using both direct referential activities (Sporas-based) and indirect trust propagation (TrustMail-based).
TL;DR
Quantifying "reputation" in a technical community is notoriously difficult. This paper proposes a methodology to move beyond manual votes by building a social network from email contents. By tracking how experts reference specific technical entities (IPs, URLs, Domains) and applying time-decayed algorithms, the authors can automatically identify authorities in the cyber-security space.
Motivation: Moving Beyond Simple Counts
In professional mailing lists, a high volume of posts doesn't necessarily equal high expertise. Previous systems struggled with two main issues:
- Context Blindness: They ignored what was being discussed.
- Temporal Decay: A breakthrough insight today is "old news" tomorrow.
The authors' insight was to treat the referencing of a technical entity (like a specific malware URL) as a "vote" of confidence or a follow-up on expertise. If Expert A mentions a zero-day exploit and Expert B later references that same entity, a directed edge of reputation is formed.
Methodology: The Reference Network
1. Network Construction
Using the UIMA framework, the system extracts three specific entities: IP addresses, URLs, and Domains. A directed graph is built where an edge exists if references an entity first introduced by .

2. The Weight of Time
To solve the "originality" problem, the authors introduced a Time Decaying Function. If two people mention the same IP independently and nearly simultaneously, they both deserve credit. However, if the second mention occurs much later, the weight of that edge decreases: This ensures that "copycat" behavior or late-coming comments don't inflate reputation as much as timely contributions.
3. Direct vs. Indirect Reputation
The authors implemented two distinct logic flows:
- Direct (Sporas-based): Updates reputation based on the authority of the person doing the referencing. If a high-reputation user references you, your reputation gains a larger boost.
- Indirect (TrustMail-based): Uses transitive properties. If trusts , and trusts , then can infer a level of trust for , even without a direct interaction.
Experiments and Structural Insights
The authors analyzed 2,415 emails, extracting 426 unique authors. The most fascinating results came from the Community Analysis:
- Power Law Distribution: Within every sub-community identified, reputation followed a power law. This implies that every technical niche has 1-2 "super-experts" and a long tail of participants.
- Negative Assortativity: Unlike typical social networks (where "popular" people hang out with "popular" people), this expert network showed that high-reputation nodes often connect to lower-reputation nodes (likely answering questions or being referenced by learners), rather than just forming an exclusive clique of authorities.
Critical Analysis & Conclusion
Takeaway: This work provides a template for "Actionable Intelligence" in HR and Knowledge Management. By analyzing the content of communication rather than just the frequency, companies can identify "hidden" experts who provide high-value technical pointers but might not be the most "vocal" in terms of message count.
Limitations: The study noted a discrepancy between human ratings and algorithmic results. This highlights a classic "Data Silos" problem: the human expert likely considered the authors' reputations outside the specific 2,415 emails analyzed, whereas the algorithm only knew what it saw in the text.
Future Outlook: With today's LLMs, we could extend this work by analyzing the sentiment and accuracy of the references, not just the occurrence of keywords, leading to even more robust reputation metrics in decentralized autonomous organizations (DAOs) and open-source communities.
