FamilyID: Unveiling the Privacy Risks of "Oversharing" in the Microblogging Era

FamilyID: A Hybrid Approach to Identify Family Information from Microblogs

2015-01-01
Jamuna Gopal, Shu Huang, Bo Luo
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces FamilyID, a hybrid information retrieval framework designed to identify family-related information within microblogs (Twitter). It combines part-of-speech (POS) tagging, pattern matching, lexical similarity, and semantic similarity to achieve high-precision extraction of sensitive relationship data.

Executive Summary

TL;DR: FamilyID is a hybrid framework that automatically identifies and extracts family-related information from Twitter. By moving beyond simple keywords to a multi-layered lexical and semantic analysis, it achieves 83% precision in pinpointing sensitive personal disclosures.

Background Positioning: This work bridges the gap between traditional Information Retrieval (IR) and Privacy-Preserving Research. It acts as a "white-hat" audit tool, proving that what users think are isolated, harmless posts can be aggregated into a comprehensive map of their private family life.

The "Needle in the Haystack": Motivation

The challenge with microblogs like Twitter is two-fold: extremity of scale and density of noise. The authors observed that while family information is highly sensitive (revealing birthdays, last names, and locations), it is present in less than 1% of total tweets.

Existing SOTA methods often rely on keyword spotting. However, language is nuanced. A user tweeting "My boy is the best" might be a mother talking about her son, or a teacher talking about a student. Purely lexical methods fail to distinguish these nuances, leading to massive false-positive rates.

Methodology: The Three-Stage Sieve

FamilyID avoids the computational cost of processing every tweet with heavy NLP by using a progressively refined "sieve" approach:

1. Pattern Extraction & Matching

The system first uses the Stanford NLP tagger to identify syntactic structures. Instead of just looking for "sister," it looks for POS patterns like PRP$ (Possessive Pronoun) + JJ (Adjective) + NN (Noun). This filters out a huge volume of structurally irrelevant data.

System Architecture

2. Lexical Similarity Assessment

Once a pattern matches, the system evaluates the words themselves. Using the UMBC ebiquity system, it checks if the nouns are actually related to family. This stage effectively prunes "false relatives" like "my dear dog" or "my sweet neighbor," which match the syntax but not the semantics of a family member.

3. Semantic Similarity Analysis

The final and most sophisticated layer uses a sliding window algorithm and the GetStsSim API. It compares the candidate tweet against a "Seed Set" of verified family-related expressions. This is the "brain" of the operation, capable of realizing that "Jesus is great" is semantically distant from "Mother is great," even if they share similar syntax.

Experiments & Results

The authors tested FamilyID on a massive dataset of 450,000 tweets from 150 users.

  • Precision: Human evaluators confirmed an 83% precision rate.
  • Noise Reduction: FamilyID rejected over 62% of tweets that keyword-based systems would have flagged as "family-related."
  • Data Sparsity: The system confirmed that from thousands of tweets, usually only about 30 per user actually reveal family secrets—making manual stalking difficult but automated extraction highly efficient.

Comparison of Total Tweets vs Family Related

Critical Insight & Conclusion

Takeaway: The real danger isn't a single tweet; it's the aggregation. FamilyID demonstrates that an adversary doesn't need to hack a database to build a family tree—they just need a smart enough parser to listen to what you're already saying.

Limitations: The paper notes that urban slang and abbreviations remain a hurdle. While they use a "Term Expansion" table (e.g., "bro" to "brother"), the evolving nature of internet slang means these dictionaries require constant updates.

Future Outlook: As LLMs (Large Language Models) become more accessible, the "Semantic Similarity" phase of FamilyID could be evolved into a zero-shot classification task, potentially pushing precision even higher and making privacy-assessment tools a standard feature for social media dashboards.

Find Similar Papers

Try Our Examples

  • Search for recent papers that use Deep Learning or Transformers to identify interpersonal relationships and kinship from social media text.
  • Which study first introduced the concept of "Information Aggregation Attacks" in social networks, and how does FamilyID build upon that threat model?
  • Explore how the multi-phase filtering approach of FamilyID could be applied to identifying other sensitive attributes like employer information or location history in short-text streams.
Contents
FamilyID: Unveiling the Privacy Risks of "Oversharing" in the Microblogging Era
1. Executive Summary
2. The "Needle in the Haystack": Motivation
3. Methodology: The Three-Stage Sieve
3.1. 1. Pattern Extraction & Matching
3.2. 2. Lexical Similarity Assessment
3.3. 3. Semantic Similarity Analysis
4. Experiments & Results
5. Critical Insight & Conclusion