DADIN: Bridging the Domain Gap in CTR Prediction via Adversarial Adaptation
DADIN: Domain Adversarial Deep Interest Network for Cross Domain Recommender Systems
This paper introduces DADIN (Domain Adversarial Deep Interest Network), a cross-domain Click-Through Rate (CTR) prediction model. By framing cross-domain recommendation as an adversarial domain adaptation problem, it achieves state-of-the-art (SOTA) performance on Huawei and Amazon datasets.
Executive Summary
TL;DR: DADIN (Domain Adversarial Deep Interest Network) is an end-to-end framework designed to solve the data sparsity and cold-start problems in Click-Through Rate (CTR) prediction. It treats cross-domain recommendation as a domain adaptation problem, using adversarial training to align the feature distributions of source and target domains.
By introducing a specialized Domain Agnostic Layer and Intra-class Confusion Loss, DADIN ensures that the knowledge transferred from a source domain (e.g., news clicks) actually benefits the target domain (e.g., ad clicks) without causing "negative transfer." It achieves SOTA results on major benchmarks, proving particularly robust when domain distributions are vastly different.
The Core Challenge: Why Knowledge Transfer Fails
In real-world recommendation systems, we often have abundant data in a source domain but very limited interactions in a target domain. The goal of Cross-Domain Recommendation (CDR) is to use that source "wealth" to help the target "poverty."
However, existing methods typically face two hurdles:
- Feature Mismatch: Simply concatenating features leads to a "long tail" distribution where the model fails to find common patterns.
- Negative Transfer: If the source domain behavior (e.g., browsing movies) is too different from the target (e.g., buying music), the model might learn misleading correlations that hurt performance.
The researchers behind DADIN observed that most models only align marginal distributions (the general look of the data) but miss conditional distributions (how features relate to labels).
Methodology: The Adversarial Secret Sauce
DADIN’s architecture is a sophisticated evolution of the "Twin Tower" structure.
1. Hierarchical Interest Extraction
Before the adversarial magic happens, DADIN uses a two-level attention mechanism:
- Item-level Attention: Weights historical behaviors based on their relevance to the current candidate item.
- Interest-level Attention: Dynamically weights user profiles, source history, and target history to form a unified interest embedding.
2. The Domain Agnostic Layer (DAL)
The heart of DADIN is the Domain Agnostic Layer. It employs a Gradient Reversal Layer (GRL). During forward propagation, it acts as an identity function. During backpropagation, it multiplies the gradient by -1.
This forces the model to learn features that make it impossible for a discriminator to tell which domain the data came from. The result is a feature space where "news-clicking behavior" and "ad-clicking behavior" are mapped into a domain-neutral representation.

3. Joint Distribution Alignment
Unlike standard DANNs, DADIN introduces Intra-class Domain Confusion Loss. It specifically aligns the distributions of positive samples and negative samples separately, ensuring that the model doesn't just confuse the domains, but aligns the logic of the clicks across them.
Experimental Evidence
The authors tested DADIN against a massive array of baselines, including Wide&Deep, DIN, CoNet, and MiNet.
SOTA Performance
On the Amazon (Movie to Music) dataset—a notoriously difficult transfer task—DADIN outperformed the highly competitive MiNet by 0.71% in AUC. While 0.71% might sound small, in the world of CTR prediction, an improvement of 0.1% can translate to millions of dollars in increased revenue.

Visualizing the Adaptation
The authors used PCA to visualize the feature space. In the figure below, you can see how the Domain Agnostic Layer successfully "merges" the source and target distributions (red and blue dots moving closer), verifying that the model is indeed learning domain-invariant features.

Critical Insights & Takeaways
- Skip-Connections are Essential: One of the most interesting findings from the ablation study was that removing the "Domain Agnostic Layer" caused the model to collapse. The skip-connection allows the model to blend domain-agnostic and domain-specific information ( ), providing the flexibility needed for accurate prediction.
- The "Distance" Matters: DADIN showed a 16.67% AUC boost over its non-adversarial version on the Amazon dataset, but only 2.34% on the Huawei dataset. This suggests that adversarial adaptation is most valuable when the source and target domains are significantly far apart.
- Limitations: The training process involves a "Min-Max" game, which can be unstable. As shown in the loss analysis, the Total Loss often fluctuates as the CTR predictor and Domain Discriminator fight for dominance before eventually converging.
Conclusion
DADIN moves the needle for cross-domain recommendation by shifting the focus from simple feature sharing to rigorous distribution alignment. By making the model "blind" to the domain origin while "sharp" to user intent, it sets a new standard for how we handle cold-start users in industrial-scale recommendation systems.
