Uber's mRMR: Solving the "Redundancy Paradox" in Large-Scale Marketing ML
Maximum Relevance and Minimum Redundancy Feature Selection Methods for a Marketing Machine Learning Platform
The paper presents a framework for feature selection in large-scale marketing applications using Minimum Redundancy and Maximum Relevance (mRMR). It introduces two novel extensions: Randomized Dependence Coefficient (RDC) for non-linear redundancy and Random Forest importance for relevance, achieving SOTA balancing between model interpretability and predictive accuracy at Uber.
TL;DR
In industrial machine learning, more data does not always mean better models. Uber's latest research demonstrates that selecting the "m best features" is fundamentally different from selecting the "best m features." By extending the Minimum Redundancy Maximum Relevance (mRMR) framework with non-linear measures and model-based importance, Uber achieves higher accuracy with a fraction of the features, drastically reducing engineering overhead and improving model interpretability.
The Problem: The Curse of "Relevant but Redundant"
In marketing ML—predicting churn, cross-sell, or app sign-ups—we often have thousands of features ranging from static demographics to dynamic app usage logs. Standard practice often involves ranking features by their individual correlation to the label (Relevance).
However, this leads to a critical failure: the top 10 features might all represent slightly different versions of the same signal (e.g., "logins in 7 days," "logins in 14 days," "active days"). This redundancy adds noise, increases the risk of overfitting, and bloats the data pipeline infrastructure without adding predictive value.
Methodology: Refining the mRMR Framework
The authors pivot from simple ranking to an optimization problem: Maximize Relevance while Minimizing Redundancy.
1. The mRMR Core Formula
The standard mRMR is defined as:
2. Uber’s Extensions: RFCQ and RFRQ
The researchers identified that standard linear correlation (Pearson) misses non-linear dependencies. They introduced two key upgrades:
- Non-linear Redundancy (RDC): Using the Randomized Dependence Coefficient to catch features that correlate in complex, non-obvious ways.
- Model-based Relevance: Using Random Forest (RF) importance scores instead of F-statistics to ensure the feature selection is aligned with the actual downstream model's logic.
- The Quotient Scheme: Instead of subtracting redundancy from relevance (which fails when scales differ), they use a Quotient (FCQ/RFCQ), which makes the trade-off more robust across different data types.

Evaluation: Synthetic and Real-World Impact
The team tested 8 variants across synthetic data and three massive Uber datasets (cross-sell and up-sell scenarios).
Key Findings:
- Efficiency: Top-tier performance (AUC) was often reached with just 10-15 features, even when the original set contained over 1,300.
- Stability: The FCQ (F-test Correlation Quotient) emerged as the "production hero"—it is nearly as accurate as complex model-based versions but significantly faster to compute ( minute vs. hours for MI-based methods).
- Overfitting Prevention: In several real-world datasets, the mRMR-selected feature subset actually outperformed the model using "All Features," proving that removing noise is as important as finding signal.

From Research to Production: The Uber Implementation
Uber integrated this into their Scala-Spark based AutoML platform. To handle the scale of millions of users, they implemented several optimizations:
- Down-sampling: Feature selection is performed on a representative sample before full training.
- Concurrency: Replacing iterative loops with Scala's functional map/reduce operations to speed up the computation of the redundancy matrix.
- Dynamic Selection: The platform can runtime-evaluate whether a model-based (RFCQ) or model-free (FCQ) approach delivers better results for a specific campaign.
Critical Analysis & Future Outlook
The paper provides a roadmap for "pragmatic ML." While researchers often focus on the most complex non-linear kernels, Uber’s findings suggest that for production, a linear-based quotient (FCQ) offers the best ROI between speed and accuracy.
However, a noted limitation is the computational cost of the RDC and Mutual Information methods on extremely high-dimensional datasets. Future work could explore Streaming Feature Selection or Differential Privacy constraints within the mRMR framework to protect user data while maintaining relevance.
Takeaway: In the era of Big Data, the most valuable models are those that can do more with less. Uber's mRMR framework proves that diversity in features is just as important as individual predictive power.
