Driving Patterns as Biometrics: Is Your Vehicle CAN Bus Data Leaking Your Identity?

Can driving patterns predict identity and gender?

2019-09-07
Osman Abul, Batuhan Karatas
Summary
Problem
Method
Results
Takeaways
Abstract

This study investigates the predictability of driver identity and gender using Controller Area Network (CAN) bus data through machine learning. Utilizing the Uyanik dataset, the authors demonstrate that driving patterns can achieve up to 0.97 accuracy for gender prediction and 0.98 accuracy for 2-class identity recognition.

TL;DR

Researchers have demonstrated that the "digital exhaust" of your car—the subtle ways you press the gas pedal or turn the steering wheel—is unique enough to be a biometric signature. By analyzing CAN bus data, machine learning models can predict your gender with 97% accuracy and identify specific drivers in a small group with near-perfect precision (98%). This turns driving data into a quasi-identifier, posing a significant privacy risk.

Context & Motivation: The Invisible Privacy Leak

Modern vehicles are sensor-rich environments. The Controller Area Network (CAN) bus serves as the nervous system of the car, broadcasting real-time data on everything from Steering Wheel Angle (SWA) to Engine RPM (ERPM).

While this data is a goldmine for insurance companies and traffic managers, the authors of this study argue that it is being shared too casually. The core "insight" here is that driving is not just a utility but a habitual behavior influenced by personality, mood, and physical traits. If a car's data is shared "anonymously," can it be linked back to you? The answer is a resounding yes.

Methodology: Translating Driving to Data

The researchers used the Uyanik dataset, featuring 105 drivers on a 25km route in Istanbul. The challenge with time-series data like driving is that trip durations vary. To create a standardized input for ML, they used a sophisticated pipeline:

  1. Feature Extraction: Using the Tsfresh library, they converted raw signals into 216 statistical features (mean, variance, skewness, auto-regression coefficients).
  2. Addressing Data Imbalance: With a heavy bias toward male drivers (88 vs 17), they used SMOTE (Synthetic Minority Oversampling Technique) to generate synthetic "female" driving patterns, ensuring the model didn't simply learn to guess the majority class.
  3. The Quasi-Identifier Framework: They treated the CAN bus data as a join-key. If an attacker has an "anonymous" drunk-driving record and a named driving habits database, they can perform a record-linkage attack based on driving similarity.

Model Architecture/Workflow Figure 1: The proposed driver re-identification framework showcasing the record-linkage attack vector.

Experimental Insights: Who is Behind the Wheel?

1. Gender Prediction (The "Attribute-Linkage" Attack)

Using all 10 CAN bus lines (including pedal pressure and gear status), the Support Vector Machine (SVM) and Random Forest (RF) models reached a staggering 97% accuracy. Interestingly, signals like Engine RPM (ERPM) and Vehicle Speed (VS) were the strongest predictors of gender, likely reflecting differing habits in acceleration and smoothness.

2. Identity Prediction (The "Record-Linkage" Attack)

  • The Family Car Scenario: In a 2-driver scenario, the models achieved 98% accuracy. This means an insurance company could easily distinguish between a father and daughter using the same vehicle.
  • Large Scale Identification: In a massive 105-class task, the accuracy was 10%. While 10% sounds low, it is an order of magnitude higher than the random guessing baseline (~0.95%), indicating that driving patterns are highly distinctive even in large crowds.

Accuracy Comparison Figure 2: Performance gains when combining multiple CAN bus lines. More data dimensions lead to nearly perfect classification.

Analysis: Why This Matters for the Future

The "Inductive Bias" of this research is clear: human behavior is repetitive. Even when restricted to a fixed 25km route, drivers leave a unique "vibrational" signature in the car's sensors.

Key Implications:

  • Data Protection: Simply removing a name or license plate (Pseudonymization) is insufficient. The driving pattern itself is the ID.
  • Policy: Regulatory bodies (like those overseeing GDPR) must treat CAN bus streams with the same sensitivity as fingerprint or location data.
  • Future Work: The authors suggest extending this to predict age groups or education levels, further deepening the profile an attacker can build from a simple commute.

Conclusion

This paper serves as a vital warning. As vehicles become "smart devices on wheels," the very act of driving becomes a form of surveillance. Privacy-preserving methods like k-anonymity and l-diversity must be integrated into the vehicle's data transmission architecture to prevent our cars from testifying against our privacy.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Differential Privacy or k-anonymity specifically for protecting vehicle CAN bus time-series data.
  • What are the seminal works on the "Uyanik" dataset (Takeda et al., 2011), and how have feature extraction techniques for this dataset evolved since then?
  • Explore how Deep Learning models like LSTMs or Temporal Convolutional Networks compare to Random Forests for driver fingerprinting in recent smart vehicle research.
Contents
Driving Patterns as Biometrics: Is Your Vehicle CAN Bus Data Leaking Your Identity?
1. TL;DR
2. Context & Motivation: The Invisible Privacy Leak
3. Methodology: Translating Driving to Data
4. Experimental Insights: Who is Behind the Wheel?
4.1. 1. Gender Prediction (The "Attribute-Linkage" Attack)
4.2. 2. Identity Prediction (The "Record-Linkage" Attack)
5. Analysis: Why This Matters for the Future
5.1. Conclusion