Decoding the Pulse of Social Networks: A Hawkes-Based Framework for Information Diffusion
A framework for information dissemination in social networks using Hawkes processes
This paper introduces a unified Hawkes-based framework for modeling information diffusion in social networks. By combining multivariate linear Hawkes processes with nonnegative tensor factorization (NTF) and topic models (LDA/Author-Topic), the method captures sub-second temporal dynamics and hidden influence structures between users and contents across multiple interconnected networks.
TL;DR
Researchers have developed a comprehensive framework that uses Hawkes processes to model how information spreads through social networks. By treating social interactions as "self-exciting" events—where one post triggers another—and applying low-rank tensor factorization, the model can uncover hidden communities (like the factions in Game of Thrones) and content influences without needing to know the social graph beforehand.
Context: Why Discrete Time is Not Enough
Most traditional models for information cascades treat time as a sequence of steps (). However, real-world social media is a continuous-time phenomenon. A tweet might be retweeted in seconds or days, and the "bursty" nature of these interactions suggests that events are not independent.
The authors argue that social influence is fundamentally self-exciting. Much like an earthquake triggers aftershocks, a message from an influential user increases the probability of subsequent broadcasts. To capture this, they turn to Hawkes Processes, a class of point processes perfect for modeling "contagious" discrete events in continuous time.
Methodology: The Architecture of Influence
The core innovation lies in how the authors decompose the intensity function , which represents the rate of new messages. Instead of estimating a massive, intractable matrix of every user's influence on everyone else, they factorize the problem:
- User-User Interaction (): They use a low-rank factorization to project users into a latent space of dimension . This acts as an implicit community detection mechanism.
- Topic-Topic Interaction (): They model how specific topics (e.g., Politics vs. Sports) excite one another.
- Fuzzy Topic Modeling: By coupling the Hawkes process with an Author-Topic (AT) model, the framework can handle messages that are mixtures of latent topics, using Variational Bayes or Gibbs Sampling to refine the topic distributions based on the timing of the posts.

Scalable Estimation via NTF
A major bottleneck in point processes is the likelihood estimation for millions of users. The authors solve this by discretizing the timeline into small bins and proving that maximizing the log-likelihood is equivalent to minimizing the Kullback–Leibler (KL) divergence in a Nonnegative Tensor Factorization (NTF) problem. This allows them to derive multiplicative updates—simple, parallelizable matrix operations that are much faster than standard convex optimization.
Experimental Proof: From Game of Thrones to MemeTracker
The most striking validation of the model involves the Game of Thrones dataset. By analyzing only the timestamps and speakers of lines in the pilot episode, the model reconstructed a "hidden influence graph."

As shown in the heatmap above, the model successfully clustered characters into their respective "Great Houses" (Stark and Lannister) without any prior knowledge of the plot. The "self-excitation" in the dialogue—who responds to whom and how quickly—revealed the underlying social structure.
Furthermore, in the MemeTracker dataset (tracking 5,000 websites), the framework identified the most influential news hubs and how different memes (topics) competed for dominance across the web.
In synthetic tests (above), the proposed method (Left) shows significantly cleaner recovery of the influence matrix compared to previous SOTA methods (Center).
Critical Insight & Future Outlook
The beauty of this framework is its unification. It bridges the gap between language modeling (what is being said) and point processes (when it is being said).
Limitations: The model assumes a linear excitation, which might not capture the "saturation" effect where a user becomes overwhelmed by too many notifications.
Future Impact: This work paves the way for real-time "influence monitoring" tools. By observing the temporal ripples of a message, companies or researchers can identify key opinion leaders and community shifts in near real-time, even when the explicit "follower" graph is hidden or manipulated.
