DADIN: Bridging the Domain Gap in CTR Prediction via Adversarial Adaptation

DADIN: Domain Adversarial Deep Interest Network for Cross Domain Recommender Systems

2023-01-01
Menglin Kong, Muzhou Hou, Shaojie Zhao, Feng Liu, Ri Su, Yinghao Chen
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces DADIN (Domain Adversarial Deep Interest Network), a cross-domain Click-Through Rate (CTR) prediction model. By framing cross-domain recommendation as an adversarial domain adaptation problem, it achieves state-of-the-art (SOTA) performance on Huawei and Amazon datasets.

Executive Summary

TL;DR: DADIN (Domain Adversarial Deep Interest Network) is an end-to-end framework designed to solve the data sparsity and cold-start problems in Click-Through Rate (CTR) prediction. It treats cross-domain recommendation as a domain adaptation problem, using adversarial training to align the feature distributions of source and target domains.

By introducing a specialized Domain Agnostic Layer and Intra-class Confusion Loss, DADIN ensures that the knowledge transferred from a source domain (e.g., news clicks) actually benefits the target domain (e.g., ad clicks) without causing "negative transfer." It achieves SOTA results on major benchmarks, proving particularly robust when domain distributions are vastly different.


The Core Challenge: Why Knowledge Transfer Fails

In real-world recommendation systems, we often have abundant data in a source domain but very limited interactions in a target domain. The goal of Cross-Domain Recommendation (CDR) is to use that source "wealth" to help the target "poverty."

However, existing methods typically face two hurdles:

  1. Feature Mismatch: Simply concatenating features leads to a "long tail" distribution where the model fails to find common patterns.
  2. Negative Transfer: If the source domain behavior (e.g., browsing movies) is too different from the target (e.g., buying music), the model might learn misleading correlations that hurt performance.

The researchers behind DADIN observed that most models only align marginal distributions (the general look of the data) but miss conditional distributions (how features relate to labels).


Methodology: The Adversarial Secret Sauce

DADIN’s architecture is a sophisticated evolution of the "Twin Tower" structure.

1. Hierarchical Interest Extraction

Before the adversarial magic happens, DADIN uses a two-level attention mechanism:

  • Item-level Attention: Weights historical behaviors based on their relevance to the current candidate item.
  • Interest-level Attention: Dynamically weights user profiles, source history, and target history to form a unified interest embedding.

2. The Domain Agnostic Layer (DAL)

The heart of DADIN is the Domain Agnostic Layer. It employs a Gradient Reversal Layer (GRL). During forward propagation, it acts as an identity function. During backpropagation, it multiplies the gradient by -1.

This forces the model to learn features that make it impossible for a discriminator to tell which domain the data came from. The result is a feature space where "news-clicking behavior" and "ad-clicking behavior" are mapped into a domain-neutral representation.

Model Architecture

3. Joint Distribution Alignment

Unlike standard DANNs, DADIN introduces Intra-class Domain Confusion Loss. It specifically aligns the distributions of positive samples and negative samples separately, ensuring that the model doesn't just confuse the domains, but aligns the logic of the clicks across them.


Experimental Evidence

The authors tested DADIN against a massive array of baselines, including Wide&Deep, DIN, CoNet, and MiNet.

SOTA Performance

On the Amazon (Movie to Music) dataset—a notoriously difficult transfer task—DADIN outperformed the highly competitive MiNet by 0.71% in AUC. While 0.71% might sound small, in the world of CTR prediction, an improvement of 0.1% can translate to millions of dollars in increased revenue.

Performance Comparison

Visualizing the Adaptation

The authors used PCA to visualize the feature space. In the figure below, you can see how the Domain Agnostic Layer successfully "merges" the source and target distributions (red and blue dots moving closer), verifying that the model is indeed learning domain-invariant features.

PCA Visualization


Critical Insights & Takeaways

  1. Skip-Connections are Essential: One of the most interesting findings from the ablation study was that removing the "Domain Agnostic Layer" caused the model to collapse. The skip-connection allows the model to blend domain-agnostic and domain-specific information ( ), providing the flexibility needed for accurate prediction.
  2. The "Distance" Matters: DADIN showed a 16.67% AUC boost over its non-adversarial version on the Amazon dataset, but only 2.34% on the Huawei dataset. This suggests that adversarial adaptation is most valuable when the source and target domains are significantly far apart.
  3. Limitations: The training process involves a "Min-Max" game, which can be unstable. As shown in the loss analysis, the Total Loss often fluctuates as the CTR predictor and Domain Discriminator fight for dominance before eventually converging.

Conclusion

DADIN moves the needle for cross-domain recommendation by shifting the focus from simple feature sharing to rigorous distribution alignment. By making the model "blind" to the domain origin while "sharp" to user intent, it sets a new standard for how we handle cold-start users in industrial-scale recommendation systems.

Find Similar Papers

Try Our Examples

  • Search for the latest cross-domain recommendation papers published in 2024-2025 that address the "negative transfer" problem using adversarial learning.
  • Which paper first proposed the Gradient Reversal Layer (GRL), and how has it been modified for multi-task learning in recommendation systems?
  • Examine recent studies that apply domain adversarial neural networks (DANN) to multimodal recommendation tasks involving both image and text features.
Contents
DADIN: Bridging the Domain Gap in CTR Prediction via Adversarial Adaptation
1. Executive Summary
2. The Core Challenge: Why Knowledge Transfer Fails
3. Methodology: The Adversarial Secret Sauce
3.1. 1. Hierarchical Interest Extraction
3.2. 2. The Domain Agnostic Layer (DAL)
3.3. 3. Joint Distribution Alignment
4. Experimental Evidence
4.1. SOTA Performance
4.2. Visualizing the Adaptation
5. Critical Insights & Takeaways
6. Conclusion