Decoding the Flow: Statistical Laws of Information Exchange in VKontakte

Analysis of the Implementation of Information Exchange Algorithms in Social Networks

2021-06-07
Mikhail Kuptsov, Vladimir Minaev, Sergey Yablochnikov, Irina Yablochnikova, V. B. Dzobelova
Summary
Problem
Method
Results
Takeaways
Abstract

This study investigates the statistical patterns of information exchange in the "VKontakte" (VK) social network by classifying communities into four categories: marketing, political, blogging, and individual. The authors synthesized regression models and confirmed that "relative views" is the most robust predictor for user activity, achieving determination coefficients () between 0.73 and 0.86.

TL;DR

Social networks are more than just digital spaces; they are complex environments governed by specific statistical laws. By analyzing 150 days of data from the VK network, researchers have discovered that while most engagement metrics are volatile, "Relative Views" serves as the ultimate signal for predicting user activity. They achieved a high prediction accuracy (up to 0.86 ) by factoring in community types and aggressive marketing outliers.

Background: Beyond the 24-Hour Window

Most existing research focuses on viral "bursts"—how info spreads in a day. This paper argues that understanding the evolution of a community requires a long-term lens. By categorizing communities into Marketing, Political, Blogging, and Individual, the authors move beyond a one-size-fits-all approach to social media modeling.

Problem & Motivation: The Chaos of Long-Term Data

Over short periods (50-100 days), data likes to follow a Normal Distribution. However, as time passes, the "mean" and "variance" shift. This makes long-term forecasting difficult.

The authors' core insight: Aggressive Marketing acts as a statistical "shocker." It breaks regular patterns. To prove this, they created their own VK community, "Youth Fever," and intentionally ran aggressive campaigns to observe how they distorted the anticipated normal law distribution.

Methodology: The "Relative Views" Insight

The researchers tracked subscribers, likes, reposts, and comments. Through regression analysis, they found that "Relative Views" acts as an integral reflection of all other activities.

Modeling Political Action

A unique challenge was predicting "active supporters"—those who actually show up to real-world political events. Standard regressions failed here until the authors introduced Fictitious Variables (e.g., whether an event was "concerted" or "unconcerted"). This adjustment transformed "noisy" data into a highly predictive model.

Table 1: Probability Distributions of Indicators Figure 1: Comparison of data aggregates following Normal, Uniform, and Exponential laws across different community types.

Experiments & Results: The Power of Regression

The authors found a massive disparity between community types:

  • Marketing & Political groups behaved similarly because both seek "active supporters" (buyers or voters).
  • Blogging groups were outliers, often lacking the stabilized statistical relationships found in targeted campaigns.

Table 3: Significant Linear Regressions Figure 2: Analysis of determination coefficients over long intervals (150-500 days), showing the dominance of "Relative Views" as a predictor.

Key quantitative takeaway: In 81.4% of cases, the models achieved a determination coefficient (R²) of up to 0.73, proving that social network evolution is far from random—it is mathematically structured.

Critical Insight & Conclusion

The Takeaway

If you want to know how many people will actually buy a product or join a protest, stop looking at subscriber counts. Look at "Views" and "Likes" relative to the previous day’s baseline.

Limitations

While the models are robust, they rely on "fictitious variables" for political context, which suggests that the purely digital data isn't enough—you still need real-world context (like the legal status of an event) to make the math work.

Future Outlook

This work sets the stage for Simulation Models. By knowing that data typically reverts to a normal distribution once a "marketing shock" ends, developers can build AI that detects when a community is being artificially manipulated versus growing organically.

Find Similar Papers

Try Our Examples

  • Search for recent studies that utilize fictitious or binary variables in social network regression analysis to improve predictive accuracy of user behavior.
  • Which paper first established the distinction between "passive followers" and "active supporters" in social media analytics, and how has this study evolved those definitions?
  • Explore how these statistical patterns of "relative views" as an integral metric have been applied to sentiment analysis or political forecasting in non-Slavic social networks like X (Twitter) or Facebook.
Contents
Decoding the Flow: Statistical Laws of Information Exchange in VKontakte
1. TL;DR
2. Background: Beyond the 24-Hour Window
3. Problem & Motivation: The Chaos of Long-Term Data
4. Methodology: The "Relative Views" Insight
4.1. Modeling Political Action
5. Experiments & Results: The Power of Regression
6. Critical Insight & Conclusion
6.1. The Takeaway
6.2. Limitations
6.3. Future Outlook