Copyright & AI
The settlement priced acquisition, not use. Here is what the split ruling actually holds, the comparison nobody has run, and the four things to change in your position this quarter.
24 July 2026 · 11 min read · StellarStart GLOBAL
Disclaimer
Articles are general commentary as at their publication date. They are not legal advice and should not be relied on for a specific matter. Several of these areas are changing quickly.
The largest publicly reported copyright recovery in history received final approval on 21 July 2026. Almost every summary of it says the same two things: training AI on books is fair use, and piracy is not. Both are true, and neither is the useful part.
The useful part is a number nobody seems to have run, a claims statistic that tells you more about copyright risk than the judgment does, and one term of the settlement that quietly gave Anthropic the thing it most needed.
Three authors, Andrea Bartz, Charles Graeber and Kirk Wallace Johnson, sued Anthropic over the corpus behind Claude. The case produced two separate events, and conflating them is where most analysis goes wrong.
Judge William Alsup granted partial summary judgment. Training large language models on books was fair use, and not marginally: the court called the purpose “transformative, spectacularly so.” Alsup went further and held that buying millions of print books, stripping the bindings, cutting the pages and scanning them into machine-readable form was also fair use. Destructive scanning of lawfully purchased copies passed.
What failed was the library. Anthropic had also downloaded roughly seven million book files from Library Genesis and the Pirate Library Mirror and kept them in a permanent central repository. Alsup refused to let the downstream training use launder the acquisition. Holding pirated copies was not fair use, and that question went to trial.
Rather than try the piracy claim, Anthropic settled for $1.5 billion. Alsup granted preliminary approval on 25 September 2025. Class counsel sought roughly $300 million in fees, about 20 per cent of the fund, in December 2025. Final approval came from Judge Araceli Martínez-Olguín on 21 July 2026 in a 23-page order, after distributions had been calculated in June. Around 500,000 works qualified, at approximately $3,000 each.
So: the ruling was split, and only the losing half was bought out.
Statutory damages under 17 U.S.C. § 504(c) run from $750 to $30,000 per work, rising to $150,000 per work where infringement is wilful. Anthropic downloaded around seven million files. Even confining exposure to the roughly 500,000 works that ultimately qualified for the class, the theoretical wilful maximum was in the region of $75 billion.
Against that ceiling, $1.5 billion is about two per cent. That is why the market read the settlement as cheap, and why Anthropic’s share of the news cycle was so calm.
Now run the comparison that matters more.
Alsup blessed the alternative path in the same ruling. Anthropic could have bought the books. Used and remaindered print copies bought in bulk cost somewhere in the range of $3 to $15 each. Take the midpoint and add generous provision for logistics, cutting and scanning, and acquiring 500,000 titles legitimately lands in the region of $5 million to $15 million.
The piracy route cost $1.5 billion. On these assumptions, taking the free copies cost roughly one hundred to three hundred times what paying for them would have cost, for the same corpus, with the same fair use outcome at the end of it.
That figure is an estimate built on public numbers and stated assumptions rather than a disclosed cost, and the real internal comparison will differ. But the direction is not close, and it reframes the entire compliance conversation. The question for any organisation building a training set is no longer “will fair use protect us.” It is “why would we take the discount that costs two orders of magnitude more.”
Anthropic reports that 91 per cent of eligible rights holders filed a claim, with more than 440,000 works claimed out of roughly 500,000.
For anyone who has been near a consumer class action, that number is astonishing. Claim rates in the low single digits are normal. Rates above 20 per cent are unusual enough to be remarked on. Ninety-one per cent is close to a total muster of an entire class.
It happened because books are the best-catalogued creative works in existence. They have ISBNs. They have registration records at the Copyright Office. They have publishers with rights databases, and they have organised advocacy in the Authors Guild and the Association of American Publishers, both of which pushed authors toward the claims process.
Copyright exposure does not scale with how much you took. It scales with how identifiable and how organised the owners are.
A corpus of books is the highest-risk material on earth to take without permission, not because books are more protected in law, but because every single owner can be found, matched and mobilised. Scraped forum posts, unattributed images, orphaned blog archives and translated text sit at the other end of that spectrum: legally just as protected, practically far less enforceable, because there is no registry, no trade body and no way to build a class list.
For a rights holder, the inversion of that principle is the cheapest insurance in intellectual property, and almost nobody buys it. Section 412 of the Copyright Act conditions statutory damages and attorney’s fees on registration before the infringement began, or within three months of first publication. Unregistered works were not what made this class work. Identifiable, registered works were. Registration is what converts a creator from “no realistic claim” into “$3,000 per work, paid without a trial.”
A settlement resolves a dispute; it does not create law. But notice the asymmetry. Alsup’s summary judgment ruling, the half favourable to AI developers, survives as persuasive authority in the Northern District of California and beyond. The half that went against Anthropic never reached a verdict, so there is no damages award, no appellate opinion and no published finding of wilfulness for anyone to cite. The industry kept the useful ruling and paid to erase the dangerous one. Read that way, $1.5 billion looks less like a defeat and more like the price of a very specific insurance policy.
Anthropic agreed to destroy the pirated files. It did not agree to retrain Claude without whatever those files contributed, and nothing in copyright law currently compels model deletion as a remedy. The settlement implicitly accepts a proposition of enormous commercial value: the copies were unlawful, the learning derived from them stays. Every AI developer watching this case took that inference away.
The Bartz class was built on inputs, on acquisition and retention of files. What a model emits is a separate question, and that is where the frontier has moved. In January 2026 the court in the New York Times litigation ordered OpenAI to produce twenty million output logs. Andersen v. Stability AI is listed for jury trial on 8 September 2026. Disney and Universal v. Midjourney is working through a discovery schedule running into 2027. None of those turn primarily on how the corpus was assembled.
The destruction obligation reads as housekeeping. It is not.
Destroying a pirated corpus properly means locating every derivative artefact: the original download, the deduplicated set, the tokenised shards, the backups, the snapshots held by cloud providers, the copies on researchers’ machines, and any embeddings or indexes derived from the files. Most organisations that ingested large datasets between 2021 and 2023 cannot produce a reliable inventory of where those files propagated, because nobody was tracking lineage at the time.
There is a second problem. Destruction obligations sit awkwardly beside litigation holds. A company facing parallel claims elsewhere may be required to preserve exactly the material it has agreed to destroy. Sequencing that without spoliation exposure is a genuine legal and engineering exercise, not a checkbox.
The headline said $1.5 billion for using books. The record says something narrower and more useful: you may learn from a book you paid for, and you may not keep a copy you stole. Everything commercially consequential in this case follows from that sentence, and from the awkward fact that paying would have been cheaper by two orders of magnitude.
Sources: Bartz v. Anthropic (N.D. Cal.), summary judgment order of 23 June 2025 and final approval order of 21 July 2026; Authors Alliance; Authors Guild; Association of American Publishers; Norton Rose Fulbright AI litigation update 2026; 17 U.S.C. §§ 412, 504.