AdRec: Resolving the Hidden Conflict in Mobile Ad Recommendations
Addressing the Conflict of Negative Feedback and Sampling for Online Ad Recommendation in Mobile Social Networks
AdRec is a neural network-based framework for online ad recommendation in Mobile Social Networks (MSN). It introduces an Auxiliary Output (AO) and a modified loss function to resolve the conflict between explicit negative feedback and traditional negative sampling, consistently outperforming SOTA baselines like NCF and DMF.
TL;DR
In the world of Mobile Social Networks (MSN), recommending the right ad is a high-stakes game. While traditional systems treat "no interaction" and "explicit dislike" as the same thing, the AdRec framework introduces a clever Auxiliary Output mechanism. By mathematically differentiating between random samples and real user feedback, it stabilizes training and achieves SOTA performance on advertising benchmarks.
The Conflict: Why Not All "Zeros" are Equal
In standard Collaborative Filtering, we often use Negative Sampling (NS). Since we only know what a user liked (positive feedback), we randomly pick items they haven't seen and assume they dislike them (labeling them 0).
However, Mobile Social Networks provide something unique: Explicit Negative Feedback. We know exactly which ads were pushed to a user's feed but were not clicked.
- The Problem: If you treat a random "unseen" ad (low confidence) the same as a "seen but ignored" ad (high confidence), your model gets confused.
- The Result: Direct application of SOTA methods like Neural Collaborative Filtering (NCF) leads to biased weights and severe overfitting.
Methodology: The AdRec Architecture
The authors propose AdRec, which leverages a dual-output neural architecture to handle this data nuance.
1. Unified Embedding Layer
The model takes three inputs: the User (), the Ad (), and Anonymous Features (). These are mapped into a dense latent space.
- Features are processed via a Feature Mapping (FM) layer using ReLU activation.
- User and Ad embeddings are combined with these features to form a representative vector.
2. Dual-Branch Output
The core innovation lies in the output structure:
- Interaction Branch (): Predicts the likelihood of a click (Sigmoid).
- Auxiliary Branch (): Predicts whether the current data point is a "real" interaction or a "sampled" one.

3. The Re-weighted Loss Function
By solving both tasks simultaneously, the loss function essentially allows the model to "learn the confidence" of a label. As proven in the paper's Proposition 1, the auxiliary branch acts as a dynamic weight , ensuring the model pays more attention to explicit feedback than to noisy negative samples.
Performance: SOTA Results
AdRec was tested against NCF (Neural Collaborative Filtering) and DMF (Deep Matrix Factorization) on the Avazu dataset.
Quantifiable Superiority
In every metric—from error rates (MSE/MAE) to ranking quality (Average Precision and AUC)—AdRec came out on top.

Why Negative Sampling (NS) Still Matters
Interestingly, even though NS creates conflict, the authors proved it is still necessary. Without NS, the model overfits rapidly. However, AdRec with AO (Auxiliary Output) provides the best of both worlds: the noise-reduction benefits of sampling without the bias of mislabeling.
Figure: Models without Negative Sampling (NS) show a sharp decrease in accuracy over time, confirming that NS is vital for generalization, but only if managed by an auxiliary signal.
Conclusion & Insights
The AdRec framework is a masterclass in principled architecture design. Instead of just adding more layers or "deeper" networks, the authors identified a specific data-level conflict and solved it with a multi-task learning objective.
Key Takeaways for Practitioners:
- Confidence Matters: In ad-tech, a "non-click" is a more powerful signal than a "non-exposure." Your model should treat them differently.
- Auxiliary Tasks as Regularization: Adding an output to predict the "source" or "nature" of data is an effective way to handle mixed-precision labels.
- Generalization: While designed for MSN ads, this approach could be adapted to any recommendation scenario where rich "negative" logs are available.
