PU Learning: Deciphering the Hidden Patterns of Mobile User Behavior in Big Traffic Data
SPECIAL SECTION ON HUMAN-CENTERED SMART SYSTEMS AND TECHNOLOGIES
This paper introduces a framework for mobile user behavior analysis, specifically targeting App usage prediction and video traffic identification using real-world ISP traffic data. The authors propose two Positive and Unlabeled (PU) learning methods—Spy-based and K-means-based—to effectively train models when only positive interactions (usage records) and unlabeled data are available.
TL;DR
Mobile Internet traffic is a goldmine for understanding human behavior, but it suffers from a fundamental data science challenge: we know what users do, but we don't truly know what they dislike (the "unlabeled" problem). This paper introduces a sophisticated PU Learning (Positive and Unlabeled) framework that treats missing interactions not as negatives, but as opportunities. By constructing a User-App Bipartite Network and using Spy-based detection, researchers achieved an F-score of 0.9 in predicting App usage.
The Motivation: The "Missing Negative" Dilemma
In the world of ISP (Internet Service Provider) data, a record exists only when a connection is made. If a user doesn't use "Youku" (a video app) today, is it because they hate it (Negative), or simply because they were busy (Unlabeled)?
Standard supervised learning fails here because it forces the model to treat all "no-interaction" data as negative, leading to massive bias. Prior works often relied on small-scale smartphone sensor logs, which are too localized. The authors of this paper shift the focus to the Core Network, where billions of flows provide a "God's eye view" of user-app interactions.
Methodology: Mining the Bipartite Network
The cornerstone of this research is the User-App Bipartite Network.
- Nodes: Users (identified by phone numbers) and App Servers (identified by IPs).
- Edges: Directed arrows representing traffic flow, carrying attributes like duration, bytes, and packets.
- Features: The authors extracted high-dimensional features (up to 194 dimensions), including node degree (connection frequency) and "strength" (average downlink/uplink volume).
Above: The systematic framework from data collection to behavior prediction.
The PU Learning Secret Sauce
To solve the lack of labels, they proposed two two-step strategies:
- Spy-based PU: A small portion of "known positives" is hidden within the unlabeled set (the "Spies"). A preliminary classifier is trained. The threshold for what constitutes a "Negative" is then determined by how the model reacts to these hidden spies. If the model thinks a spy is negative, it’s being too aggressive; if it thinks an unlabeled sample is more negative than the spies, it's likely a Reliable Negative (RN).
- K-means-based PU: Unlabeled data is clustered. Clusters that are Euclidean-miles away from the "Positive" cluster center are labeled as RN.
Above: The User-App bipartite representation captures complex interaction patterns.*
Experimental Breakthroughs
The model was tested on massive datasets, including 2.2 million records of Tencent QQ traffic and various video apps like Youku and LeTV.
- App Usage Prediction: The Spy-PU method with a Random Forest classifier dominated the field. It maintained a high F-score (0.85-0.90) even when the data was highly unbalanced—a scenario where traditional Logistic Regression failed miserably (as seen in the sharp drop in performance relative to the sampling ratio).
- Video Identification: Using statistical flow properties (variance of packet length, ratio of bytes to packets), the model successfully identified video traffic amidst the "noise" of life-service apps like Meituan.
Experiment results show that Spy-based PU learning (solid lines) consistently outperforms standard classifiers (dotted lines).
Critical Insight: Why it Works
The brilliance of this approach lies in the Reliable Negative extraction. By identifying samples that are statistically distinct from any observed usage, the second-stage classifier learns the boundary between "User would use this" and "User definitely won't use this" much more cleanly than standard binary classifiers.
Limitations & Future Outlook
While powerful, the computational complexity of calculating bipartite features and running multiple stages of classification is high. The authors propose moving towards a streaming framework (like Spark Streaming or Flink) to handle real-world ISP throughput in real-time.
As we move toward 6G and beyond, using PU learning on the network edge could allow for hyper-personalized service optimization without compromising privacy, as the "content" of the packets (DPI) becomes less relevant than the "pattern" of the behavior.
