A judge approves Anthropic's $1.5B author settlement, pricing pirated training data at $3,000 a book
A federal judge gave final approval to Anthropic's $1.5 billion settlement with authors, the largest known payout in U.S. copyright history. It closes the piracy question in Bartz v. Anthropic and puts a number on training data that was acquired the wrong way.
A federal judge in San Francisco has given final approval to Anthropic’s $1.5 billion settlement with a class of authors, the largest known recovery in the history of U.S. copyright law. The deal closes the piracy question at the center of Bartz v. Anthropic and does something the AI industry had mostly avoided pinning down: it attaches a concrete number, roughly $3,000 a book, to training data that a lab acquired the wrong way.
What the court approved
U.S. District Judge Araceli Martinez-Olguin signed off on the settlement this week, rejecting objections that the sum was too small. Complaints about the size, she wrote, were “not grounded in a realistic assessment of the overall risks and rewards of a trial.” The fund compensates the copyright holders behind roughly 500,000 works at about $3,000 each, and the judge trimmed the plaintiffs’ legal bill along the way, awarding just over $101 million of the $187.5 million their attorneys had requested.
The case, brought in 2024 by the novelists Andrea Bartz and Charles Graeber and the writer Kirk Wallace Johnson, turned on how Anthropic built the corpus behind Claude. The authors alleged the company downloaded millions of titles from shadow libraries such as Library Genesis and the Pirate Library Mirror to assemble a central training library. In June 2025, Judge William Alsup drew the line that shaped everything after: training a model on lawfully acquired books is fair use, but obtaining those books through piracy is not. That ruling left Anthropic facing a damages trial over the acquisition, with statutory penalties that could have run far past the settlement figure, which is the pressure that produced the deal.
Both sides claimed a win. Plaintiffs’ counsel called it “the largest known copyright recovery in history,” and Anthropic said it was “pleased that more than 91% of authors and publishers covered by the settlement have claimed their share of the payment.” For a working developer, the practical read is narrow but real: the models many teams build on were trained in part on material a court found was pirated, and the industry now has a price tag for that exposure rather than an open-ended legal question.
Where this lands in the market
The settlement matters less for its dollar figure than for the distinction it hardens. Courts have separated two things the industry often blurred: training on copyrighted text, which Alsup treated as fair use, and how the text is obtained, which remains ordinary copyright law. A lab can win the first argument and still owe a fortune on the second. That split is now backed by the largest payout on record, which makes it a reference point the next case will be measured against.
Anthropic is not the only lab with pending exposure here, and the others are watching. Suits against OpenAI and other model makers are still working through the courts, and a settlement of this scale gives plaintiffs a live benchmark and gives defendants a rough sense of what clean-up costs. The likely industry response is a quieter shift already underway: paying for licensed corpora, documenting provenance, and being able to prove where the data came from. For teams choosing a foundation model, that pushes data provenance and vendor indemnification from a footnote toward a real procurement question, next to latency, price, and capability.
What’s worth watching
- Whether pending AI copyright suits anchor to this number. A $1.5 billion approved settlement is the first hard comparable in the category, and both sides in the OpenAI and other cases now have it to argue from.
- Whether labs can show clean provenance. The cheap defense is no longer “training is fair use.” It is being able to demonstrate the training set was acquired lawfully, which favors licensing deals and auditable data pipelines.
- Whether provenance becomes a buyer’s question. Watch for indemnification and data-sourcing language to start showing up in enterprise model contracts, the way security posture already does.
The lasting signal here is that the fair-use win did not settle the bill. Anthropic established that training on books can be lawful and still had to pay a record sum for how it got the books, which reframes training data as a supply chain a lab has to account for, not a resource it can simply scrape. Stackmaven’s follow-up coverage will revisit how the pending cases move against this benchmark on or around October 19.