Unmasking the Invisible: Extracting Social Networks from Noisy SMS Data
Hidden social networks analysis by semantic mining of noisy corpora
This paper introduces a paradigm for Social and Semantic Networks Analysis (SSNA) to extract hidden social ties from noisy, incomplete corpora (e.g., SMS messages) where explicit relationship data is missing. It employs semantic mining, full-text ranking (BM25-based), and noise reduction to construct "hidden knowledge graphs" and community models.
TL;DR
In the world of big data, "hidden" social networks exist where friendship links aren't explicitly recorded. This paper presents a disruptive approach to Social Network Analysis (SNA) by mining the semantic content of noisy SMS corpora. By analyzing how people share specific topics and terms, the author reconstructs social structures and community clusters even from incomplete datasets where recipient information is totally missing.
Background: Beyond the "Follow" Button
We usually think of social networks as graphs of "A follows B." But what happens when you have a database of 88,000 text messages with the recipient IDs deleted for privacy? Traditional SNA fails here. This paper argues that what people say and who says it is just as indicative of a social tie as a formal "friendship" declaration.
The author's core insight is that information flows in social networks behave like physical flows (modeled after electrophysics). By treating text content as a signal and orthographic noise as interference, we can "hear" the social structure hidden in the noise.
Methodology: Ranking, Denoising, and Modeling
The author's workflow transforms raw, messy text into a structured knowledge graph through a three-stage process:
1. The Ranking Engine
To value the "strength" of a link between a user and a keyword, the paper uses an advanced probabilistic ranking model inspired by BM25. This goes beyond simple frequency (TF-IDF) by adjusting for document length and the specific linguistic quirks (inflections) of the French SMS language.
2. Denoising & Semantic Mining
SMS data is incredibly noisy (e.g., 'cat' appearing as 'chat', 'chot', or 'cha').
- Stop-word filtering: Using the Morfalou lexicon to strip out adverbs and verbs, keeping only high-value nouns.
- Thesaurus Expansion: Utilizing the Google Suggest API to map slang and contractions (like "abdos" to "abdominaux") which traditional lexical distance algorithms often miss.
3. Dual-Graph Construction
The methodology branches into two specific visualizations:
- Hidden Social-Semantic Graphs: Nodes represent both people and terms.
- Hidden Knowledge Networks: A more complex model that links term to term based on the probability that they are shared by the same senders.
Fig 1: The distribution of arc weights shows a classic long-tail curve, which the author uses to split the network into "head" (frequent terms) and "tail" (niche topics) for better analysis.
Experimental Analysis: Results from 88milSMS
Using the 88milSMS dataset, the proof of concept identified several fascinating community structures.
- Community Identification: Using modularity clustering, 14 distinct groups were identified.
- Influencer Detection: The graphs revealed "community leaders"—individuals who occupy central positions by using specific vocabulary that bridges different parts of their cluster.
Fig 2: A medium-sized community where specific "leaders" (left) control a cloud of unique terms, influencing the central vocabulary of the group.
One of the most striking outputs is the "Hidden Knowledge Sphere". This visualization (Fig 7 in the paper) resembles a mind map of collective social attention. It isn't a formal ontology, but rather a "belief of crowds" representation—showing which concepts are inextricably linked in the public consciousness.
Fig 3: The Hidden Knowledge Sphere for frequent terms. The spherical structure emerges from the 'knowlink' weights calculated between semantic nodes.
Critical Insight & Conclusion
This work represents a shift from Sociometry (who knows whom) to Semantometry (who shares what).
Key Takeaways:
- Reliability: Semantic sharing is a high-fidelity proxy for social relationships.
- Scalability: The ranking/denoising pipeline allows for the analysis of millions of hidden ties without needing explicit relational databases.
- Future Potential: This approach opens doors for Neuroinformatics and Cognitive Science, moving beyond "Wisdom of Crowds" into analyzing the "Deep Feelings" and collective attention of a society.
Limitations: The reliance on external APIs like Google Suggest for the thesaurus introduces a dependency on third-party semantic logic. Future iterations would benefit from locally-trained LLM-based embeddings to handle the "tail" of noisy orthography more autonomously.
