Talkin’ ’Bout AI Generation: Mapping the Copyright Supply Chain

Talkin' 'Bout AI Generation: Copyright and the Generative-AI Supply Chain

2023-01-01
Katherine Lee, A. Feder Cooper, James Grimmelmann
Summary
Problem
Method
Results
Takeaways
Abstract

The article, forthcoming in the Journal of the Copyright Society of the U.S.A. (2024), introduces the "generative-AI supply chain" framework to analyze copyright infringement. It systematically evaluates legal liabilities across eight technical stages—from data curation to model alignment—positioning this holistic taxonomy as the necessary lens for courts to navigate the complex AI ecosystem.

TL;DR

Is Generative AI a copyright infringer? The legal answer is no longer a simple "yes" or "no." This paper argues that to judge AI, we must stop looking at the chatbot and start looking at the supply chain. By breaking the technology down into eight distinct stages, the authors show how a choice made by a data scraper in stage 3 creates a legal time bomb for a developer in stage 6.

The "Generative AI" Monolith Fallacy

The core problem in the current legal "Gold Rush" against AI companies (OpenAI, Meta, Stability AI) is a lack of technical precision. Most litigants treat AI as a "black box" that swallows art and spits out remixes. However, this ignores the sociotechnical complexity of how these models are actually built.

The authors argue that "Generative AI" is a catch-all name for an ecosystem. A coding assistant like GitHub Copilot raises different legal issues than a music generator like Suno, primarily because their training data, architectures (Transformer vs. Diffusion), and deployment methods differ wildly.

Methodology: The Eight Stages of Risk

The paper’s primary contribution is the Generative-AI Supply Chain. By viewing AI through this lens, we can pinpoint exactly where "copying" happens:

  1. Dataset Curation: The first point of direct infringement. Mass-scraping without provenance creates "data laundering" risks.
  2. Model Pre-training: This transforms expressive works into mathematical weights. Is the model itself a "copy"? If it can reproduce a training image (memorization), the authors argue it might be.
  3. Model Alignment (RLHF): A stage often ignored by lawyers. Here, human feedback steers models. If a model is "aligned" to mimic a specific artist’s style, the intentionality increases the risk of "inducing" infringement.

Generative-AI Supply Chain Figure 1: The taxonomy of the Generative-AI Supply Chain, illustrating the feedback loops between generation and curated datasets.

The "Snoopy Effect" and Memorization

One of the most compelling insights in the paper is the analysis of Substantial Similarity. If you prompt a model for an "archaeologist with a whip," you get Indiana Jones—even if you didn't use the name. This "Snoopy Effect" occurs because the character is so prevalent in the training data that the model develops a "latent concept" of the copyrighted identity.

The Snoopy Effect Figure 2: Examples of models generating recognizable characters (Snoopy) despite varying prompts, proving their latent "memorization" of protected expression.

Why "Fair Use" Isn't a Silver Bullet

For years, machine learning lived under the protection of the "Google Books" precedent: copying is okay if it's for "nonexpressive" indexing. Generative AI breaks this. Because the purpose of a generative model is to create something for human consumption, the "transformativeness" of the use is now a high-stakes jury question.

If a model (like GPT-4) can recreate a page from a Harry Potter book, it is no longer just indexing; it is competing in the market for that book. Factor 4 of the Fair Use test (market harm) becomes a major hurdle for AI labs.

Critical Insight: Who "Pushed the Button"?

The authors tackle the "Volitional Conduct" hurdle. Who is the direct infringer?

  • The User? (If they write a prompt specifically to pirate a movie.)
  • The Service? (If the model produces a copyrighted character even when asked for something generic.)
  • Both? (If the system is designed to "hallucinate" copyrighted styles for profit.)

The paper concludes that liability will likely follow control. Closed-source systems like ChatGPT have more liability because they control the "safety filters" and the alignment. Open-source models (like Llama) transfer that risk to the end-user.

Conclusion: A Future of Licensed Commons

The paper warns that courts should beware of "too-easy" metaphors. AI is not just a "library" or "a collage tool." It is a new infrastructure. The authors predict that we are moving toward an "Infringing Model" regime, where the only safe path forward is a fully licensed training pipeline—Adobe Firefly being a primary example of this "clean" supply chain.

Generative AI hasn't killed copyright; it has made every line of code and every training weight a part of the legal record.

Find Similar Papers

Try Our Examples

  • Search for recent court rulings or legal analyses that apply the "generative-AI supply chain" framework to European Union AI Act compliance or copyright exceptions.
  • Which legal papers first debated the concept of "nonexpressive use" in machine learning, and how has the transition to generative models challenged the core assumptions of those earlier works?
  • Identify research exploring technical "unlearning" or "machine unlearning" methods designed specifically to mitigate copyright liability by removing specific protected works from pre-trained model parameters.
Contents
Talkin’ ’Bout AI Generation: Mapping the Copyright Supply Chain
1. TL;DR
2. The "Generative AI" Monolith Fallacy
3. Methodology: The Eight Stages of Risk
4. The "Snoopy Effect" and Memorization
5. Why "Fair Use" Isn't a Silver Bullet
6. Critical Insight: Who "Pushed the Button"?
7. Conclusion: A Future of Licensed Commons