Talkin’ ’Bout AI Generation: Mapping the Copyright Supply Chain
Talkin' 'Bout AI Generation: Copyright and the Generative-AI Supply Chain
The article, forthcoming in the Journal of the Copyright Society of the U.S.A. (2024), introduces the "generative-AI supply chain" framework to analyze copyright infringement. It systematically evaluates legal liabilities across eight technical stages—from data curation to model alignment—positioning this holistic taxonomy as the necessary lens for courts to navigate the complex AI ecosystem.
TL;DR
Is Generative AI a copyright infringer? The legal answer is no longer a simple "yes" or "no." This paper argues that to judge AI, we must stop looking at the chatbot and start looking at the supply chain. By breaking the technology down into eight distinct stages, the authors show how a choice made by a data scraper in stage 3 creates a legal time bomb for a developer in stage 6.
The "Generative AI" Monolith Fallacy
The core problem in the current legal "Gold Rush" against AI companies (OpenAI, Meta, Stability AI) is a lack of technical precision. Most litigants treat AI as a "black box" that swallows art and spits out remixes. However, this ignores the sociotechnical complexity of how these models are actually built.
The authors argue that "Generative AI" is a catch-all name for an ecosystem. A coding assistant like GitHub Copilot raises different legal issues than a music generator like Suno, primarily because their training data, architectures (Transformer vs. Diffusion), and deployment methods differ wildly.
Methodology: The Eight Stages of Risk
The paper’s primary contribution is the Generative-AI Supply Chain. By viewing AI through this lens, we can pinpoint exactly where "copying" happens:
- Dataset Curation: The first point of direct infringement. Mass-scraping without provenance creates "data laundering" risks.
- Model Pre-training: This transforms expressive works into mathematical weights. Is the model itself a "copy"? If it can reproduce a training image (memorization), the authors argue it might be.
- Model Alignment (RLHF): A stage often ignored by lawyers. Here, human feedback steers models. If a model is "aligned" to mimic a specific artist’s style, the intentionality increases the risk of "inducing" infringement.
Figure 1: The taxonomy of the Generative-AI Supply Chain, illustrating the feedback loops between generation and curated datasets.
The "Snoopy Effect" and Memorization
One of the most compelling insights in the paper is the analysis of Substantial Similarity. If you prompt a model for an "archaeologist with a whip," you get Indiana Jones—even if you didn't use the name. This "Snoopy Effect" occurs because the character is so prevalent in the training data that the model develops a "latent concept" of the copyrighted identity.
Figure 2: Examples of models generating recognizable characters (Snoopy) despite varying prompts, proving their latent "memorization" of protected expression.
Why "Fair Use" Isn't a Silver Bullet
For years, machine learning lived under the protection of the "Google Books" precedent: copying is okay if it's for "nonexpressive" indexing. Generative AI breaks this. Because the purpose of a generative model is to create something for human consumption, the "transformativeness" of the use is now a high-stakes jury question.
If a model (like GPT-4) can recreate a page from a Harry Potter book, it is no longer just indexing; it is competing in the market for that book. Factor 4 of the Fair Use test (market harm) becomes a major hurdle for AI labs.
Critical Insight: Who "Pushed the Button"?
The authors tackle the "Volitional Conduct" hurdle. Who is the direct infringer?
- The User? (If they write a prompt specifically to pirate a movie.)
- The Service? (If the model produces a copyrighted character even when asked for something generic.)
- Both? (If the system is designed to "hallucinate" copyrighted styles for profit.)
The paper concludes that liability will likely follow control. Closed-source systems like ChatGPT have more liability because they control the "safety filters" and the alignment. Open-source models (like Llama) transfer that risk to the end-user.
Conclusion: A Future of Licensed Commons
The paper warns that courts should beware of "too-easy" metaphors. AI is not just a "library" or "a collage tool." It is a new infrastructure. The authors predict that we are moving toward an "Infringing Model" regime, where the only safe path forward is a fully licensed training pipeline—Adobe Firefly being a primary example of this "clean" supply chain.
Generative AI hasn't killed copyright; it has made every line of code and every training weight a part of the legal record.
