MEFDM: Enhancing Social Network Diffusion Modeling via Multi-Evidence Fusion
Multiple evidence fusion based information diffusion model for social network
The paper introduces the Multiple Evidence Fusion based Diffusion Model (MEFDM), a social network information propagation framework that utilizes Dempster-Shafer (D-S) evidence theory. By fusing neighbor influence and user activity, the model achieves a fitting goodness of 0.79 on the Enron email dataset.
TL;DR
Information diffusion in social networks is rarely the result of a single factor. This paper presents MEFDM, a model that moves beyond simple probability by using Dempster-Shafer (D-S) evidence theory to fuse user activity and interpersonal influence. Tested on the Enron email dataset, the model demonstrates high fidelity (0.79 fitting goodness) in simulating how information cascades through a professional network.
Problem & Motivation: Beyond Static Probabilities
Most classical models, such as the Independent Cascade (IC) or Linear Threshold (LT) models, suffer from a "parameter estimation" bottleneck. They assume fixed probabilities for information "jumping" from one person to another, often ignoring the context of who is sending the message and when the recipient is likely to be active.
The authors argue that a realistic model must account for the multi-faceted nature of human behavior. For instance, a colleague might have a high influence over you, but if you receive an email at 2:00 AM when your "activity level" is zero, the diffusion process hits a wall. Bridging these two dimensions—influence and activity—is the core focus of this research.
Methodology: The Core of Evidence Fusion
The MEFDM architecture is structured into four layers: Data Source, Evidence Generation, Evidence Reasoning, and Forwarding Decision.
1. The Evidence Components
The model quantifies two primary "mass functions":
- Influence Evidence (): Calculated based on the ratio of successfully forwarded messages between a specific sender-recipient pair versus the total messages forwarded by the recipient.
- Activity Evidence (): A temporal factor based on the time of day, reflecting the probability of a user being online and engaged.
2. D-S Theory Fusion
Instead of simple multiplication, the authors use D-S Evidence Reasoning to combine these factors into a single Belief Function (). This mathematical framework is particularly adept at handling uncertainty and conflicting evidence.
Figure 1: The MEFDM architecture showing the fusion of behavior data and historical interactions.
The belief function is defined as:
Experiments and Results
The model was validated using the Enron Email Dataset, which contains 250,000 messages.
The Temporal Impact
The study highlights how "work habits" dictate information flow. In the Enron network, activity peaks at 8:00 AM and troughs at 10:00 PM. Experiments showed that messages initiated at 8:00 AM reached significantly more nodes and spread faster than those initiated late at night.
Figure 2: Information spread varies drastically depending on the starting time of the message.
Simulating Real-World Cascades
The most impressive result is the comparison between real diffusion data and MEFDM simulation. The model tracks the growth of "infected" nodes (users who forwarded the email) with high precision.
Figure 3: MEFDM achieves a 0.79 fitting goodness score compared to actual Enron cascades.
Critical Analysis & Conclusion
Takeaway
MEFDM succeeds because it recognizes that diffusion is a stochastic process driven by multi-modal evidence. By combining temporal activity with relational influence, it provides a much more granular "forwarding probability" than baseline models.
Key Insights:
- Influence Matters: Nodes with even slightly higher average influence mass (e.g., node 9 in the study) act as catalysts for much broader cascades.
- D-S Theory Utility: The use of evidence theory allows for a "discrimination framework" that handles the binary decision (forward vs. non-forward) through a rigorous fusion of independent data streams.
Limitations & Future Work
While 0.79 is a strong fitting score, the authors acknowledge that user preference (the content of the message) was not included in this iteration. Future versions of MEFDM could potentially integrate Natural Language Processing (NLP) to extract "Interest Evidence," further refining the belief function.
