The Price of a Training Corpus
The Price of a Training Corpus
In July 2026 a court approved a $1.5 billion payout for pirated books at about $3,000 a title, and an earlier court had already called training on bought books fair use. The number prices a hook. What actually binds a lab is the set of public positions it can occupy together — each one available later as a quote in someone else’s complaint.
What the number bought
The 25 July 2026 conversation works the payout as the first major public dollar figure on this fight. The panel is not neutral. One seat is an administration official on AI policy. The rest of the room is venture capital sitting on the asset class. The tape says twice that everyone is talking their books. Names live in Sources.
What stuck was making an unauthorized copy. The libraries named on the docket are LibGen and PiLiMi. About seven million books were taken from those sites. The class that could recover is the slice of those works that meet registration and ISBN rules — about 500,000 titles. $1.5 billion divided by 500,000 is $3,000 a book. All three figures check against the settlement papers and the 20 July 2026 final-approval item. The stuck fact on the tape is that the defendant did not pay for one legitimate copy.
That $3,000 is the first observed price for the piracy hook. It is not a licensed training-copy rate. A lab that treats it as a rate for learning from a book has overclaimed. Exposure is sized by multiplying the observed number by works the lab cannot show a receipt for. Works it can show a receipt for sit under a different ruling.
On 23 June 2025 a district court in the Northern District of California had already split the two hooks. Training on legally acquired books is fair use — the opinion called it quintessentially transformative. Downloading and keeping pirated library copies is not. The class was certified for piracy only. The settlement priced the second hook. Training on purchased books had already been won, which is why the rest could be settled. Judge and caption are the dated objects; reporter commentary lives in Sources. The on-air line that training was still undecided is a speaker behind the ruling, or a speaker declining to treat a district-court summary judgment as settled law. His hedge stays. The ruling is added.
This is the first major AI-training copyright lawsuit to settle. More are in the pipeline. A panel estimate of how many major cases sit behind it was unsourced on air and is not printed here as a count.
Settling versus doctrine
A hypothesis assembled from this episode: holdings appear only in the fights nobody can buy out of. Whoever must go to court lives under case law. Whoever can write a cheque large enough to end the case lives under the last cheque written. One data point. The case against later prices it.
The usual story is that a well-funded defendant pays, and paying ends the opinion that would have bound the next defendant, so the opinions that exist are the ones nobody could afford to stop. This case is the opposite in one respect. The training ruling preceded the settlement. Purchased-and-digitized copies had already been held fair use. The buyout covered the copies nobody could show a receipt for.
What then governs depends on who is in the chair. A small builder who cannot buy out a class lives under the opinions. A frontier lab lives under the last cheque, plus whatever the other side’s terms of service already forbid.
The episode contradicts itself on whether a court had already spoken. One seat says courts had already held that training on copyrighted books is fair use. Two minutes later another seat calls the question unresolved. A third case is named as a final judgment on AI training copyright. The previous ruling is the June 2025 split. The named database case is Thomson Reuters v. Ross — Thomson, never Thompson — a February 2025 district-court summary judgment against fair use for a competing legal-research tool trained on a professional database’s headnotes. Market harm did the work. It is not generative book-training. It is not appellate. “Final judgment on AI training copyright” overstates.
On the panel’s own summary, IP law does not have the nuance yet, so society will decide what is fair, and the decision route is settlements, lobbying, and legislation. That is their account.
The constraint that actually binds
A policy sentence can be read into a later case. A lab that argues, in court, that learning from the world’s writing is lawful, and argues, in Washington, that learning from its writing is theft, has handed the first set of plaintiffs a quote. The Washington sentence is an admission. The case against already records that the load-bearing fact is asserted and denied by the same speaker. The mechanism still holds.
The cause of action that does not boomerang is the one that never asserts a property right in outputs. Bulk fake accounts, proxies, lying at signup — described as a deceptive business practice — can ask for government help without saying the outputs are owned. That reframe is reusable. It is the same advice as “never say IP theft” written once.
The test that still stands after the hypocrisy charge is displacement. A model that can talk about a book has taken no customer’s order. A product trained on a professional database in order to sell a rival professional database has taken the order. Those are the two poles of factor four, stated as fact about the test. The database case’s work lived in the previous section. Here the test travels.
Weights, outputs, and the forward pipeline
On the panel’s line, the parameter file is property and the text a model emits is contract. Taking the file would be theft. Learning from the emitted text runs into terms of service. That is their seam, not a holding.
Weights are described on air as the numerical parameters inside the software. Taking that file would be theft. One seat says nobody has alleged it, then earlier reports labs telling government that foreign companies can steal the file. Both lines stay. The full seat treatment is in the case against.
Learning from outputs is given an industrial pedigree: carmakers study each other’s cars; early search engines queried rivals to benchmark. Analogies, not findings.
The live distillation question is terms-of-service enforcement, not copyright. A newspaper-versus-lab case is pleaded on the same seam. Nothing on that side has a number yet. A February lab post is said, on the show, to have coined “industrial scale distillation attack,” and “IP theft” is claimed absent from the same post. This page did not open the document. The useful move is the instruction: open it. The detection drill for what that language is doing sits on Applied Critical Thinking - Testing Frames.
Diffusion is the panel’s easier hypothetical. Reviews, catalog data, and other people’s commentary already live in public. A pre-AI study-guide is offered as proof that what a good book generates cannot be contained. The floor everyone grants: text cannot be lifted, reprinted, and claimed as one’s own. The hypothetical is doing lighter work than it appears. It wins an easier question than millions of pirated books.
Collective withholding is, on their account, the only lever a scattered rightsholder still holds, and a single crawler breaks it. The same bot fetches pages for search and pages for training, so refusing the second fetch also refuses the first. That was a live 2025–26 industry demand.
The rightsholder flywheel named on air is litigate until settlement, treat the settlement as admission the use was licensed, carry the admission to the next counterparty. Music is the worked example. What defeats the flywheel, on their view, is a fragmented base that settles cheap and early — book authors, newspapers, magazines, called “very meek.” Value-laden. Their view.
The replayed demand is organize and name terms. The answer on the tape is “You’re going to get steamrolled.” The industry demand is split the two bots. A newspaper–platform standoff collapsed into a negotiated paywall-exclusion because publishers wanted the platform’s users. Compressed account. Theirs.
The durable commercial claim is the forward pipeline. A stream of new work converts a one-time hit into a subscription. A model cut off from the next book and the next wire story ages out. The pipeline is the asset. The back settlement is a sunk cost.
Provenance is circular on the tape: one lab trained on publishers and settled $1.5 billion; another is sued by a newspaper; other labs trained on American model output. The honest hedge, spoken on air, refuses the word stealing on the ground that ownership of the copyright is itself unclear.
Case against
A single buyout is not a market. $3,000 prices that hook, between one defendant holding an indefensible piracy fact and a class whose counsel had a fee. The next buyout can land at a fraction.
The hypothesis that holdings appear only in fights nobody can buy out of is assembled from two data points in one episode. A well-funded defendant that takes a training question all the way to an opinion falsifies it. The June 2025 ruling is already that data point — a training question run to an opinion, before the settlement. Stronger than the conversation knows.
The hypocrisy chapter’s load-bearing fact is asserted and denied by the same speaker. He never sources the government-meeting version. He is an administration official. He names the stake: tainted foreign-model derivatives, and the startup ecosystem. He opens as no fan of the company and names “regulatory capture.” The co-host supplies market-substitution without drawing the hypocrisy. Every clause of that paragraph is the page’s. Thinning it to “the room is biased” would drop the mechanism. A ten-figure liability behaves like the fixed cost in Regulatory Capture via Doom-Marketing step 3, absorbable only by the largest incumbents. That page owns who holds the gate. This page owns how the corpus gets priced. The loop test on Epistemic Exceptionalism — what a framework can output other than “trust me / my group” — is the check to run on the hypocrisy case, subject to that page’s gate.
A proposed pool of ten percent of builder revenue arrives without a derivation, without a defined base, and without a rule for who gets paid. It is countered on air with a prediction that the other side will ask for the whole stack. Take the mechanism — a forward license — and leave the number. The diffusion hypothetical is unfalsifiable on any horizon a reader could use. The ten-percent pool is the cheapest available outcome for every business in the room, including the conversation that produced this page.
The frame is US copyright only. A jurisdiction that answers by statute is outside it. Nothing here is about personality or voice rights. A lab with clean acquisition has no line item on the piracy hook. What remains is the fair-use question as decided at district court for purchased books, not binding on the next defendant. For licensed or open corpora the binding constraint is the licence.
Four checks, each falsifiable.
Rightsholders whose product is the licensed corpus — music, legal databases — win or settle above $3,000 a book. Commentary-level trade-book claims settle at or below. If the newspaper case comes in cheaper than the book figure, displacement is not doing the work the test assigned it.
Whether any frontier lab writes, in a brief or a signed public paper, that training without consent is IP theft. If one does and its fair-use defense survives, discoverability was overpriced.
Whether the bundled crawler is split without a collective demand. If it is, collective withholding was never the leverage.
The next AI-training settlement’s per-work number. The same order as $3,000 supports a rate. An order of magnitude away means one negotiation was mistaken for a market. That last check is the quit signal. Believing the number is a rate costs the next settlement landing at a fraction.
With output-ownership unsettled and no clean hands, an IP enforcement regime is necessarily selective. The question it answers is who gets protected. When the price is set by settlement, exposure is the set of claims a company has made in public about everyone else. The corpus without receipts is only the part that already has a number. The seam is confirmed, not replaced: the public figure prices copies never bought, after a court had already spoken on learning from copies that were.
Related
- Regulatory Capture via Doom-Marketing — a ten-figure liability behaves like the fixed cost in that page’s step 3, absorbable only by the largest incumbents; that page owns who holds the gate, this page owns how the corpus gets priced.
- Epistemic Exceptionalism — its loop test, what a framework can output other than “trust me / my group,” is the check to run on the hypocrisy case, subject to that page’s gate.
- Riding the AGI — carries proprietary data sets as a one-clause barrier to entry; the settlement figure is what that clause costs.
- Prohibition After Diffusion — the ban question from the same episode; this page supplies the liability channel by which an IP-theft framing taints derivative work.
- The Margin Moves to the Serving Layer — where value settles once models commoditize; a licensed forward pipeline is a cost on whoever holds the model layer.
- Applied Critical Thinking - Testing Frames — owns the detection drill: what language is doing emotional work, who benefits if the frame becomes the default; “industrial scale distillation attack” is a live specimen.
Open Questions
Is the per-author claim too small.
Where is this vault’s own displacement line.
What do the terms of service on this stack’s ingestion paths actually say.
What check runs before repeating “they stole it.”
Sources
- All-In, episode 282, 25 July 2026. Settlement segment. Panel unlabeled in the body; seats as used above. Source of the conversation, the seats, the same-speaker contradiction, and the ten-percent pool.
- Reuters, 20 July 2026. “US judge approves Anthropic’s $1.5 billion settlement in copyright lawsuit.” https://www.reuters.com/world/us-judge-approves-anthropics-15-billion-settlement-copyright-lawsuit-2026-07-20/. Final approval. The approving judge on this item is not the judge who wrote the June 2025 training ruling.
- Anthropic Copyright Settlement site. https://www.anthropiccopyrightsettlement.com/. Class size, per-title allocation, claim mechanics. Case: Bartz v. Anthropic, No. 3:24-cv-05417 (N.D. Cal.).
- Alsup, J., 23 June 2025, Bartz v. Anthropic, N.D. Cal. Training on legally acquired books held fair use; downloading and keeping pirated library copies held not. Class later certified for piracy only.
- Thomson Reuters Enterprise Centre GmbH v. Ross Intelligence Inc. (D. Del. Feb. 2025), Bibas, J. District-court summary judgment: not fair use, because a competing legal-research product trained on Westlaw headnotes; factor four did the work. Not generative book-training. Not appellate.
- Anthropic, “Detecting and preventing distillation attacks.” https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks. Cited on the show. Phrase test not re-run here; open the document.