TSD-CRP: Unifying Time and Space to Model the Infinite Flux of Social Trends
Modeling Infinite Topics on Social Behavior Data with Spatio-temporal Dependence
The paper introduces Time and Space Dependent Chinese Restaurant Processes (TSD-CRP), a non-parametric Bayesian model for social behavior data. It utilizes a unified spatio-temporal distance measure to capture high-order dependencies and automatically determine the optimal number of latent topics.
TL;DR
Social media behavior isn't just about "what" is said, but "where" and "when." TSD-CRP is a non-parametric Bayesian framework that treats time and space as a unified distance. By doing so, it automatically discovers the number of latent topics—be they global events like the World Cup or localized gatherings—while capturing long-range dependencies that traditional Markovian models miss.
Context: Why "Flat" Topic Models Fail
Traditional Latent Dirichlet Allocation (LDA) and its variants often treat social media posts as independent "bags of words." Even more advanced "dynamic" topic models frequently suffer from two fatal flaws:
- Strict Markov Chains: They assume the current state only depends on the immediate past, ignoring long-term periodicities or distant social influences.
- Fixed Topic Counts: Social media is volatile. Pre-defining topics is impossible when new hashtags and trends emerge every hour.
The authors argue that a user's behavior is tied to Action Inertia (temporal consistency) and Social/Local Influence (spatial consistency).
Methodology: The Geometry of Social Interaction
The core innovation of TSD-CRP is the transformation of user records into a linked graph where the probability of a connection depends on a Spatio-temporal Distance.
1. Defining the Unified Distance
The distance between two behavior records is defined as: Where:
- : The temporal time difference.
- : A combination of online social network shortest paths and offline co-location status.
Exponentially weighting the spatial distance ensures that social proximity acts as a "multiplier" for temporal relevance. A close friend's post from 3 hours ago might be more relevant than a stranger's post from 30 minutes ago.
2. The Generative Process
Unlike a standard CRP where customers sit at tables based on popularity, in TSD-CRP, a record (customer) links to a previous record based on the decay function . This creates a "forest" where each tree represents a topic.
In the figure above, the link determines topic assignment. If a record links to itself, a new topic is born.
Experiments: Superior Predictive Performance
The model was tested against state-of-the-art benchmarks on Weibo and Douban datasets.
Quantifiable Gains
TSD-CRP consistently outperformed Dynamic Topic Models (DTM) and Recurrent CRP (RCRP) in predictive log-likelihood:
| Dataset | TSD-CRP | RCRP | DTM |
|---|---|---|---|
| -14.92 | -15.08 | -15.62 | |
| Douban | -4.474 | -4.504 | -4.781 |
Visualizing Topics
The experiments revealed that topics recovered by TSD-CRP are highly concentrated in time, but the model uniquely identifies periodic topics—recurring trends that typical "sliding window" or Markov models would treat as separate, unrelated events.
The visualization shows how topics (latent behaviors) spike and fade, with TSD-CRP capturing the concentration more accurately than its predecessors.
Critical Insight: The "Social-Spatial" Multiplier
The most striking takeaway is shown in the correlation analysis: the closer the spatial distance (social or physical), the larger the time span over which two records remain correlated.
This proves that "Location" and "Social Ties" expand the temporal horizon of relevance. If you are in the same building as someone, your topics of interest remain synchronized for a significantly longer duration than they would with a random user across the globe.
Conclusion
TSD-CRP provides a mathematically elegant way to merge geography, social graphs, and time into a single non-parametric framework. While the complexity requires careful implementation of the sample space, the model's ability to handle high-order dependencies makes it a robust choice for modern social behavior analysis.
Limitations: The model assumes a fixed social graph . In future iterations, modeling the evolution of the social graph alongside the topics could provide even deeper insights into how influence propagates.
