Fair use is the US copyright doctrine that allows limited use of copyrighted material without the copyright holder's permission for purposes like criticism, comment, news reporting, teaching, scholarship and research. Whether training AI models on copyrighted content qualifies as fair use is one of the most consequential unresolved legal questions of the AI era, with multiple major lawsuits still working through US courts in 2026.
The four-factor fair use test from US copyright law:
- Purpose and character of the use — is it transformative? Commercial or non-commercial? AI training is generally argued to be transformative; opponents argue commercial AI products are commercial and substitutional.
- Nature of the copyrighted work — factual works get less protection than highly creative ones.
- Amount and substantiality — was the whole work used? AI training typically uses entire works.
- Effect on the potential market — does the use harm the market for the original? Plaintiffs argue AI substitutes for licensing markets that would otherwise exist.
The major US lawsuits relevant in 2026:
- NYT v. OpenAI/Microsoft — newspaper publishers seek damages and injunctions for training on their content. Discovery has surfaced training data details; the case is shaping public understanding.
- Authors Guild v. OpenAI / Anthropic / Meta — class actions on behalf of fiction and non-fiction authors.
- Getty Images v. Stability AI — image rights case in both US and UK.
- Various artist class actions — against Stability AI, Midjourney, DeviantArt and others.
- Music industry cases — RIAA against Suno and Udio for training on copyrighted music.
- Dow Jones v. Perplexity — search-and-summary use of news content.
- Concord Music v. Anthropic — song lyrics in training data.
The state of the law as of 2026:
- No definitive Supreme Court ruling yet on AI training as fair use.
- Mixed lower-court signals — some early decisions favourable to AI defendants on transformative use; others sympathetic to rightsholders on market harm.
- Settlements in some cases — both AI labs and rightsholders have shown willingness to license; Anthropic's reported deals with several publishers and the OpenAI–News Corp licensing deals are markers.
- Training data licensing as a market — has emerged in 2024–2026 as both a hedge against legal risk and a quality strategy.
The international dimension:
- EU AI Act and EU copyright law require AI providers to publish summaries of training data and respect rightsholders' opt-outs.
- UK is finalising rules around text and data mining exceptions; outcome will affect global model trainers.
- Japan has a more permissive stance allowing training on copyrighted material under certain conditions.
- China has its own set of rules, less aligned with Western copyright frameworks.
For a US team building products on top of frontier APIs in 2026, the fair-use risk lives mostly with the model provider rather than with you. But the secondary risks are real: outputs that closely resemble copyrighted training material can expose downstream users; certain commercial uses of generative content may be scrutinised; rights to AI-generated content itself is unsettled (the US Copyright Office holds that purely AI-generated works are not copyrightable). Following the case law and avoiding obvious risk patterns (training on clearly-copyrighted material, generating outputs that closely match specific copyrighted works) is the practical posture.