Identifying the Stranger: Privacy-Preserving Professional Profiling via Mobile Big Data

Identifying unfamiliar callers’ professions from privacy-preserving mobile phone data

2020-12-01
Jiaquan Zhang, Xiaoming Yao, Xiaoming Fu
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a machine learning framework to identify the professions of unfamiliar callers (Normal, Taxi Driver, Delivery, or Telemarketer/Fraudster) using privacy-preserving mobile cellular data. By analyzing anonymized web service requests from 1,282 users in Shanghai, the authors achieved an overall classification accuracy of 75.64% using a Random Forest model.

TL;DR

In an era of rampant telemarketing and "last-mile" delivery services, knowing who is calling is a matter of safety and efficiency. This paper proposes a system that identifies a caller's profession—Taxi Driver, Delivery Staff, or Telemarketer—by analyzing statistical patterns in their mobile data traffic. By focusing on "how" people move and use apps rather than "where" they are or "what" they specifically do, the method achieves over 75% accuracy while remaining strictly GDPR-compliant.

The "Lazy Label" Problem: Why Current Systems Fail

Most caller ID apps (like TrueCaller or Baidu's label service) depend on users manually flagging numbers. This creates significant issues:

  • Inaccuracy: One disgruntled user can mislabel a normal person as "Harassment."
  • Latency: A fraudster can make thousands of calls before enough reports accumulate to trigger a public label.
  • Data Sensitivity: Previous research attempts to automate this often required "raw" data—exact GPS coordinates or full App usage logs—which are now restricted by privacy laws like GDPR.

Methodology: The Art of Statistical Fingerprinting

The authors' core insight is that professions have distinct behavioral signatures that persist even after sensitive details are stripped away. They propose three primary feature sets:

1. Directional Mobility Patterns

Instead of tracking exact locations, the model calculates the Standard Deviation (SD) of a user's position across 12 evenly distributed directions.

  • Taxi Drivers: High SD in all directions (wide, multi-directional coverage).
  • Delivery Staff: High SD in specific regions (hub-and-spoke patterns).
  • Telemarketers: Simple strip-like patterns (basic home-to-office commuting).

Mobility Patterns Figure 1: Comparison of location distributions (a) and 12-direction SD vectors (b) for different professions.

2. Request Volume & Temporal Activeness

By dividing the day into six slices (6:00 to 24:00), the model tracks the volume of web service requests. Telemarketers and drivers show higher activity after 18:00, while delivery staff exhibit a more stable, evenly distributed request pattern throughout the day.

3. App Preference Distribution

To maintain privacy, the model doesn't look at which app is used, but rather the distribution curve of the top 10 most-used domains. Normal users have a "flat" distribution (diverse interests), whereas professional users have a "steep" curve dominated by 1 or 2 work-specific apps (e.g., driver versions of Uber/DiDi).

App Usage Distribution Figure 2: Sorted usage rates demonstrating the concentrated app focus of professionals vs. normal users.

Experiments and Results

The study utilized a real-world dataset from a major Chinese telecom operator in Shanghai, covering 1,282 users.

  • Top Performer: Random Forest (RF) achieved the highest overall accuracy of 75.64%.
  • Professions: Identification was most accurate for Drivers (79.12%) and Delivery staff (78.84%).
  • Efficiency: The researchers found that data from just one day was sufficient to reach these accuracy levels, making real-time identification feasible for newly activated professional numbers.

Performance Comparison Table 1: Accuracy results across different machine learning algorithms.

Critical Insight: Why This Matters

The breakthrough here isn't just the 75% accuracy—it's the inductive bias of the feature engineering. By using sorted directional SDs, the model becomes rotation-invariant. Whether a delivery driver works in a north-south grid or a circular city layout, their "distribution" remains identifiable.

Limitations:

  • The model struggles slightly to distinguish "Normal" users from "Harassment" callers (both show lower mobility compared to drivers).
  • Professional shifts (people changing jobs) can introduce noise into the labels over a 3-month period.

Conclusion

This work bridges the gap between big data utility and user privacy. It proves that we don't need to know exactly where a person is to understand their professional role; their movement and digital rhythms speak for themselves. This paves the way for carrier-level identification services that protect users from fraud without compromising the privacy of the callers.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Differential Privacy or Federated Learning to improve caller identification accuracy without raw data access.
  • What are the seminal works in trajectory-based user profiling, and how does this paper's "sorted direction standard deviation" refine those earlier mobility models?
  • Explore how statistical App preference distributions, rather than raw App logs, have been applied to demographic prediction (age/gender) in recent mobile big data studies.
Contents
Identifying the Stranger: Privacy-Preserving Professional Profiling via Mobile Big Data
1. TL;DR
2. The "Lazy Label" Problem: Why Current Systems Fail
3. Methodology: The Art of Statistical Fingerprinting
3.1. 1. Directional Mobility Patterns
3.2. 2. Request Volume & Temporal Activeness
3.3. 3. App Preference Distribution
4. Experiments and Results
5. Critical Insight: Why This Matters
6. Conclusion