CVR Prediction in Mobile Marketing: Balancing Latency and Accuracy in Big Data Streams

A Comparison of Data-Driven Approaches for Mobile Marketing User Conversion Prediction

2018-09-01
Luís Miguel Matos, Paulo Cortez, Rui Mendes, Antoine Moreau
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents an exploratory study on user Conversion Rate (CVR) prediction within the mobile performance marketing sector using real-world big data from OLAmobile. It compares various combinations of data preprocessing (categorical encoding and class balancing) and machine learning algorithms (Logistic Regression, Random Forest, XGBoost, and online learners) using a robust rolling window evaluation.

TL;DR

In the high-stakes world of Mobile Performance Marketing, predicting whether a user will convert (buy a product) after clicking an ad is the "holy grail." This study evaluates various ML pipelines on real-world stream data, finding that while Random Forests offer peak accuracy, XGBoost and Logistic Regression provide the best balance for production systems, hitting sub-10ms latency targets while maintaining robust AUC scores.

The "4V" Challenge of CVR Prediction

Demand-Side Platforms (DSPs) operate in an environment defined by Volume, Velocity, Variety, and Value. Every hour, millions of redirects (clicks) occur, but only a tiny fraction—often less than 1%—result in a sale. This creates two massive technical hurdles:

  1. Extreme Class Imbalance: Most models struggle to learn when "success" is such a rare event.
  2. High-Cardinality Categorical Data: Attributes like "Campaign ID" or "Mobile Operator" can have thousands of unique levels, making traditional One-Hot Encoding computationally suicidal.

Methodology: The Stream Engine & Rolling Windows

The authors developed a custom stream processing engine (Fig. 1) to sample data from OLAmobile’s data centers. To mimic real-world deployment, they eschewed the traditional static train-test split in favor of a Rolling Window Evaluation.

Stream Engine Architecture

Feature Engineering: The Power of Pruning

One of the paper's critical insights is the CP10 (Categorical Pruned) method. Instead of encoding 1,500 campaign IDs, CP10 only tracks the top 10 most frequent levels and groups the rest into an "Other" category. This drastically reduces memory usage and prevents the "curse of dimensionality" without sacrificing substantial predictive signal.

Performance vs. Computational Effort

The study highlights a crucial industry reality: An accurate model is useless if it's too slow. In the bidding world, a decision must be made in milliseconds.

Traffic TypeBest Model (AUC)Best Speed (%)Key Takeaway
BEST (Established)Random Forest (77.3%)XGBoost (0.6s)XGBoost is nearly as accurate as RF but 100x faster.
TEST (New Ads)Logistic Regression (63.7%)XGBoost (3.1s)For new campaigns, simple linear models handle noise better.

Experimental Results Comparison

Deep Insight: Why Why Logistic Regression Won the "TEST" Scenario

In "TEST" traffic (new campaigns with little history), the marketing company typically uses random matching. This paper shows that even a simple Logistic Regression, paired with SMOTE balancing, can achieve an AUC of >60%. Why not Deep Learning or Complex Ensembles? In cold-start/low-data scenarios, high-capacity models like Random Forests tend to overfit the sparse "Sales" signal. The linear inductive bias of LR acts as a regularizer, providing more stable predictions for unseen traffic.

Critical Analysis & Real-World Value

The study proves that:

  • Preprocessing matters more than model complexity: Using CP10 or IP10 (IDF-based pruning) makes big data manageable on standard server hardware.
  • XGBoost is the "Silver Bullet": For established traffic (BEST), it achieved ~75% AUC with a prediction time of just 1ms, meeting the strictest DSP requirements.

Limitations: The study relied on sampled data due to hardware constraints. In a full-scale deployment using frameworks like Apache Spark, the authors anticipate even higher throughput, though the relative performance hierarchy between algorithms is likely to remain stable.

Conclusion

This research provides a pragmatic blueprint for mobile marketers. By moving away from random ad matching to a data-driven approach—specifically using XGBoost with pruned categorical encoding—companies can significantly increase conversion rates while ensuring their infrastructure survives the deluge of big data.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Apache Spark or Flink for real-time Conversion Rate prediction in high-throughput advertising environments.
  • What are the latest advancements in "Feature Engineering" for mobile advertising, specifically focusing on user behavioral sequences post-click?
  • Identify comparative studies between XGboost and Deep Interest Networks (DIN) regarding inference latency in production Demand-Side Platforms.
Contents
CVR Prediction in Mobile Marketing: Balancing Latency and Accuracy in Big Data Streams
1. TL;DR
2. The "4V" Challenge of CVR Prediction
3. Methodology: The Stream Engine & Rolling Windows
3.1. Feature Engineering: The Power of Pruning
4. Performance vs. Computational Effort
5. Deep Insight: Why Why Logistic Regression Won the "TEST" Scenario
6. Critical Analysis & Real-World Value
7. Conclusion