Digital Devotion: Predicting Religious Identity through Microblogging Networks
On predicting religion labels in microblogging networks
This paper introduces a classification framework for predicting users' religious affiliations (specifically Christian and Muslim) within microblogging networks like Twitter. By leveraging content, structural, and aggregate features from a community-based dataset of over 111k Singaporean users, the authors achieve high-performance results using SVM classifiers with F1-scores exceeding 0.8.
TL;DR
Researchers from Singapore Management University have developed a high-accuracy system to predict whether a Twitter user is Christian or Muslim by looking not just at what they tweet, but who they follow. By introducing the concept of Representative Users, the model achieves F1-scores above 0.88, proving that our social circles often speak louder than our words when it comes to personal faith.
Background & Motivation
Religion is a fundamental driver of human behavior, influencing everything from consumer preferences to social relationships. Historically, capturing this data required massive, slow-moving government censuses. While social media offers a goldmine of real-time data, there is a catch: Linguistic Sparsity. In the studied dataset of 111k Singaporean users, less than 1% explicitly stated their religion. The challenge was to bridge the gap between this tiny "labeled" slippet and the silent majority.
Methodology: The Power of Representation
The core innovation of this research is the Representative User measure. Instead of treating all "popular" accounts (like news agencies) equally, the authors developed RInDeg (Representative Indegree Ratio).
1. Identifying the Influencers
Unlike standard PageRank or Degree Centrality, RInDeg calculates a user's dominance within a specific religious cluster relative to others. For example, while a news account like STcom might have many followers from both religions, a church leader like JosephPrince has a high RInDeg for the Christian community.

2. Feature Engineering
The model uses three types of features:
- Content: TF-IDF scores of words in tweets (e.g., "psalm" for Christians, "insyaallah" for Muslims).
- Structural: Connectivity patterns—who does the user follow? Are these followees representative of a specific faith?
- Aggregated: Statistics like the fraction of tweets containing specific religious terms.
Crucially, they use Social Features, which append the content of the top neighbors to the target user's profile, providing context even if the target user rarely tweets about religion.
Experimental Results
The study evaluated the model using SVM (Support Vector Machines) on a ground-truth set of 1,029 manually labeled users.
Key Findings:
- Social > Self: Using content from a user’s neighborhood (Social Content) significantly outperformed using only the user's own tweets (Self Content).
- RInDeg > Degree: Neighbors selected via the RInDeg metric provided much cleaner signals for classification than neighbors selected by simple popularity (Degree).
- High Precision: The "Self + Social" combined approach reached an F1-score of 0.909 for the Christian category.

Critical Insight: Why it Works
The success of this method relies on the social principle of Homophily—the tendency of individuals to associate with similar others. Even if a user never posts a religious keyword, their "digital neighborhood" (the preachers they follow, the religious news they consume) acts as a mirror of their own identity.
Limitations & Future Outlook
While the results are impressive, the study was limited to Christian and Muslim labels due to the lack of self-declared Buddhist or Hindu users in the Singaporean Twitter sphere. This highlights a cultural "reservation" in certain religious groups regarding online disclosure.
Future research could investigate more complex graph-based embeddings (like Node2Vec or GraphSAGE) to capture even more nuanced community structures, potentially uncovering religious affiliations even in highly reserved populations.
Summary Takeaway
This work demonstrates that identifying representative anchors in a social network is the most effective way to solve the data sparsity problem in latent attribute prediction.
