Realigning AI Safety: Why Deepfake Research Fails Victims of Sexual Violence

Position: AI/ML Deepfake Research is Misaligned with AI-Generated Non-Consensual Intimate Imagery (AIG-NCII)

Li Qiwei, Wells Santo, Sarita Schoenebeck, Eric Gilbert
Summary
Problem
Method
Results
Takeaways
Abstract

This position paper identifies a critical "misalignment" in AI/ML research, where technical interventions for deepfakes focus on "epistemic harms" (truth/authenticity) while ignoring AI-Generated Non-Consensual Intimate Imagery (AIG-NCII). The authors argue for a shift toward subject-centric dignity harms and propose safety-aligned metrics and restricted model access as solutions.

TL;DR

The AI research community is currently obsessed with "Truth"—detecting whether an image is real or fake. However, a new position paper reveals that this focus entirely ignores the most prevalent and damaging use of generative AI: AI-Generated Non-Consensual Intimate Imagery (AIG-NCII). By treating "authenticity" as a proxy for "safety," we are failing to protect the dignity of subjects, and in some cases, our "detection tools" actually make things worse for victims.

Academic Positioning: This is a seminal "Call to Action" that shifts the deepfake discourse from Epistemic Harms (can we trust what we see?) to Dignity Harms (is the subject being violated?).

The "Authenticity" Blind Spot

The term "deepfake" originated in a Reddit community dedicated to non-consensual pornography. Yet, as the technology moved into the academic mainstream, the motivation shifted toward political misinformation and financial fraud.

The authors' landscape analysis of the top 100 most-cited AI defense papers between 2020 and 2025 shows a staggering gap: 87% of the papers don't even mention sexual harm, despite reports suggesting that 98% of deepfake videos found online are pornographic.

The Disconnect between Research and Reality

Why "Real vs. Fake" is the Wrong Metric

The core of the paper's argument is that Authentic Safe. Current technical interventions follow three main paradigms:

  1. Detection: Finding artifacts in diffusion models (e.g., DIRE, spectral traces).
  2. Provenance: Using metadata (C2PA) to track a digital history.
  3. Watermarking: Embedding invisible signals (SynthID).

These tools help a viewer know if they are being lied to. But for a subject of AIG-NCII, the harm is the undignified presentation of their likeness. Whether the image is 100% real or 100% synthetic is often irrelevant to the trauma caused by its circulation.

The Danger of Detection Tools

Crucially, the authors highlight that detection tools can be weaponized:

  • Verification for Abusers: Perpetrators in "nudification" communities can use authenticity detectors to verify that a leaked image is "real," thereby maximizing the humiliation of the victim.
  • Plausible Deniability: Currently, the "fake" nature of deepfakes provides a small safety net of doubt. If a tool definitively labels an image as "Synthetic," it rewards the abuser for being "transparent" while the victim's image remains online.

Viewer-centric vs Subject-centric Harms

Methodology: From Diffusion to Dignity

The paper traces the evolution from GAN-based face-swapping to Diffusion-based synthesis (Stable Diffusion, LoRA). The technical barrier to abuse has collapsed. Anyone with a few reference photos can now "clone" an identity.

The authors propose a matrix that every AI safety researcher should memorize. Safety belongs to the axis of Consent, not the axis of Artificiality:

SafeHarmful
SyntheticArtistic self-expressionAIG-NCII
AuthenticConsensual pornographyTraditional NCII

A Call for "High-Friction" Defenses

So, how does the AI/ML community fix this? The paper offers several radical recommendations:

  1. Restrict High-Risk Assets: Stop the "open-release" of models specifically optimized for high-fidelity identity transfer (inpainting and cloning). We need "friction" to stop casual abusers.
  2. Adversarial Immunization: Instead of detecting fakes after they are made, we should develop tools that "poison" our personal photos so that diffusion models cannot recognize or manipulate our faces (e.g., Glaze, Anti-DreamBooth).
  3. Decouple Labeling from Removal: If a system detects AIG-NCII, it should never apply a public "AI-generated" label. It should trigger an immediate backend flag for suppression or removal.

Critical Insight: The Responsibility of the Field

The authors reject the common excuse that "social harms are for policymakers." They point out that the very techniques used for AIG-NCII—Inpainting, LoRA, and DreamBooth—were celebrated at top-tier venues like CVPR and NeurIPS. Since the field created the capabilities, it bears an "ethical red line" responsibility to mitigate the downstream violence.

Conclusion

Current AI Safety is focused on "existential risk" and "truth." This paper is a sobering reminder that for millions of women and girls, the risk is already here, and it is deeply personal. We must shift our focus from protecting the viewer's eyes to protecting the subject's dignity.

Takeaway: Friction is the goal. If we can't make the technology perfectly safe, we must make it harder to use for harm.

Find Similar Papers

Try Our Examples

  • Search for recent papers that propose adversarial immunization or "image cloaking" techniques specifically designed to prevent identity-preserving Generative AI synthesis.
  • Which studies first introduced the distinction between epistemic harms and dignity harms in the context of digital media, and how has this theory evolved with generative AI?
  • Find research evaluating the psychological impact or "secondary trauma" on human moderators and researchers who analyze AIG-NCII and Child Sexual Abuse Material (CSAM).
Contents
Realigning AI Safety: Why Deepfake Research Fails Victims of Sexual Violence
1. TL;DR
2. The "Authenticity" Blind Spot
3. Why "Real vs. Fake" is the Wrong Metric
3.1. The Danger of Detection Tools
4. Methodology: From Diffusion to Dignity
5. A Call for "High-Friction" Defenses
6. Critical Insight: The Responsibility of the Field
6.1. Conclusion