Precise Privacy: Syntactic Pattern Matching for Natural Temporal Anonymization

Anonymizing Temporal Phrases in Natural Language Text to be Posted on Social Networking Services

2014-01-01
Hoang-Quoc Nguyen-Son, Anh-Tu Hoang, Minh-Triet Tran, Hiroshi Yoshiura, Noboru Sonehara, Isao Echizen
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces a novel pattern-based algorithm for anonymizing temporal phrases in Social Networking Service (SNS) text. By integrating temporal tagging (SUTime) with sentence parsing subtrees, the method identifies and deletes time-related segments to protect user privacy without compromising the grammatical naturalness of the message.

TL;DR

Researchers have developed a pattern-matching algorithm that scrubs time-related info from SNS posts (like "at 9 AM") by targeting specific subtrees in a sentence's parsing structure. Unlike crude deletion, this method maintains the "natural feel" of a tweet, achieving 84.53% precision—a major jump over traditional rule-based taggers.

Context: The Social Media Stalking Vector

When you post "I'm heading to the gym at 3 PM," you aren't just sharing a routine; you're providing a timestamped blueprint for potential attackers. While we have tools to anonymize names or locations, temporal phrases are notoriously difficult to scrub. They are highly variable ("in 20 mins," "last Tuesday," "at dusk") and don't fit into standard generalization dictionaries. If you simply delete them, you're often left with a linguistic "scar" that tells an attacker exactly where information was hidden.

The Core Insight: Syntactic Independence

The authors observed a fundamental linguistic property: in English, temporal phrases often act as adjuncts. They are "bolted onto" the sentence structure rather than being its core.

Consider the sentence: “I go to Tokyo with friends at 9AM.” If you visualize this as a parsing tree, "at 9AM" is a coherent branch (a Prepositional Phrase). Cutting this branch off doesn't collapse the rest of the tree; it leaves a perfectly valid sentence: “I go to Tokyo with friends.”

Image Figure 1: The architecture of a temporal pattern where the bold branch is targeted for deletion.

How It Works: The Two-Phase Pipeline

1. Pattern Extraction

The system doesn't just look for words; it looks for shapes.

  • Normalization: It fixes "SNS-speak" (e.g., "w/" to "with", "g0" to "go").
  • Feature Fusion: It combines the Stanford Parser (to see the tree) with SUTime (to tag leaves as /TIME or /DATE).
  • Abstraction: It creates a generic pattern like (PP (IN/ANY) (NP (NN/TIME))). This allows a single pattern to catch "at night," "on Monday," and "in July."

2. Text Anonymization

When a user prepares to post a message, the system:

  1. Parses the new message.
  2. Matches the tree against the "Known Temporal Patterns" database.
  3. Prunes the matching subtrees.
  4. Re-normalizes: If the original post had intentional typos like "eeats sushiii", the system restores that flavor to the remaining text to ensure the anonymization is invisible to the casual observer.

Image Figure 2: Demonstrating how one learned pattern can catch diverse temporal expressions.

Results and SOTA Comparison

The algorithm was tested on a massive dataset of 4,008 tweets.

  • Accuracy: 84.53% of tweets were anonymized both correctly and naturally.
  • Efficiency: The "Power Law" of language applies here—the top 10 most frequent patterns (out of hundreds) managed to catch nearly 88% of all temporal phrases.
  • Baseline: It outperformed pure SUTime tagging (72.88%) because SUTime often misses the surrounding prepositions, leaving behind awkward fragments like "I go to Tokyo at."

Image Figure 3: Precision curve showing how a few patterns quickly reach high performance.

Critical Analysis: When Deletion Fails

The authors are honest about the "Subject Problem." If the time is the subject—e.g., "Tomorrow is my birthday"—the deletion method produces a nonsensical "is my birthday".

In these cases, the paper suggests that Generalization (replacing "Tomorrow" with "A certain day") must be used instead of suppression. This highlights the boundary of the current work: it is highly effective for adjunct removal but requires a secondary strategy for argument replacement.

Conclusion

This work provides a robust framework for SNS privacy that goes beyond simple keyword blocking. By understanding the "geometry" of a sentence via parsing trees, we can excise sensitive data while leaving the social and linguistic value of the communication intact. Its applicability to other domains (location, religion, or military data) makes it a versatile tool for the future of private digital communication.

Find Similar Papers

Try Our Examples

  • Search for recent papers that use syntactic tree pruning or subtree replacement for text anonymization beyond temporal phrases.
  • Which paper first established the SUTime library, and how do current Transformer-based NER models compare to it in temporal expression recognition?
  • Explore studies investigating the "naturalness" or "readability" of anonymized social media text for preventing adversarial detection.
Contents
Precise Privacy: Syntactic Pattern Matching for Natural Temporal Anonymization
1. TL;DR
2. Context: The Social Media Stalking Vector
3. The Core Insight: Syntactic Independence
4. How It Works: The Two-Phase Pipeline
4.1. 1. Pattern Extraction
4.2. 2. Text Anonymization
5. Results and SOTA Comparison
6. Critical Analysis: When Deletion Fails
7. Conclusion