In 2025, a federal judge made one of the first real rulings on a question every AI company had been arguing about for years: is it “fair use” to train a model on copyrighted books without paying for the rights? The answer turned out to depend less on what the AI does with the text and more on how the company got its hands on it.
What Fair Use Actually Means
Fair use is a doctrine in US copyright law that lets someone use a copyrighted work without permission or payment, under specific conditions — quoting a paragraph in a review, using a clip in a classroom lesson, sampling a few seconds of a song for parody. It isn’t a fixed rule; courts weigh four factors together:
- Purpose and character of the use — is it transformative, and is it commercial or nonprofit?
- Nature of the copyrighted work — factual works get less protection than creative ones.
- Amount used — a small excerpt is safer than reproducing the whole work.
- Effect on the market — does the new use replace sales of the original?
No single factor decides the outcome, which is exactly why AI training — a genuinely new kind of “use” that didn’t exist when the law was written — needed a court to weigh in.
How a Court Applied It to AI Training
The test case was Bartz v. Anthropic, filed in 2024 by novelist Andrea Bartz and other authors. In June 2025, Judge William Alsup issued a split ruling. Training Anthropic’s Claude models on books the company had legally bought and scanned, he found, was fair use — the model learns statistical patterns from the text rather than storing or reproducing it, which he considered highly transformative.
But Anthropic hadn’t only used purchased books. To build its training library faster, it had also downloaded millions of titles from shadow libraries like Library Genesis. Alsup ruled that piracy was not fair use, no matter what the books were later used for — the illegal act was the copying itself, not the training.
Rather than let a jury decide damages for the pirated books, Anthropic agreed to settle. In July 2026, a federal judge granted final approval to a $1.5 billion fund — about $3,000 per work across an estimated 500,000 books — the largest copyright recovery in US history. Because Anthropic settled instead of appealing, the fair-use finding on legally acquired books was never tested by a higher court, so it isn’t binding precedent for any other company.
Why It Matters Beyond One Company
The same question is still being litigated elsewhere. The New York Times’ lawsuit against OpenAI and Microsoft, filed in 2023, argues that training on paywalled journalism harms the market for that journalism — the same fourth factor Alsup weighed, applied to a different kind of content. Publishers, artists, and news organizations have filed similar suits against other AI companies, and several have instead struck licensing deals to sidestep the fight entirely.
For now, the practical lesson from the Anthropic case is narrower than it first sounds: training itself has a real shot at being ruled fair use, but how a company sources its training data is a separate and much less forgiving legal question. Buying and scanning a book is one thing; downloading it from a piracy site is another, even if both books end up feeding the same model. That’s a different question from who owns what a model generates afterward — training-data fair use is about the input, not the output.
In the News
The settlement described here received final court approval this week, closing out one of the first major copyright cases against an AI company — though far from the last.
Sources: court filings and reporting summarized by Courthouse News Service and the US Copyright Office Fair Use Index.