I have been arguing about the wrong half of this for two years.
Every conversation I have had about AI and copyright — and I have had a tiring number of them — was about whether a machine reading a book is like a person reading a book. Whether training is theft or study. Whether a model that has absorbed ten million sentences and can produce an eleventh has copied anything at all. It is a genuinely interesting question, it makes for a good argument over dinner, and it turns out to be close to irrelevant.
The cases are not being won there. They are being won several steps earlier, on a question so unglamorous that almost nobody outside the filings talks about it: where did you get the file?
The same disclosure I owed you two weeks ago: this newsletter is written with AI assistance, and the company being sued here makes the assistant. My correction is to build the plaintiffs' case out of their own filed words and label every allegation as one. The complaint is linked below; check me.
What is actually in the caption
On August 28, 2026, Sony Music Publishing and Warner Chappell Music — thirty-five affiliated entities between them — filed suit in the Northern District of California. The document is worth opening for the first line of the caption alone, because the case is styled against a company and two human beings: Anthropic PBC, Dario Amodei, and Benjamin Mann.
Not "Anthropic and its officers." Not the company alone, with executives mentioned in the body. The pleading lists the two founders the way it would list any private defendant — Amodei as "an individual residing in San Anselmo, California," Mann as "an individual residing in or around San Francisco, California." Then it splits the counts. Three of the four run against the company or against everyone. Count II, contributory infringement, runs against Amodei and Mann and no one else.
The publishers open by calling this "one of the largest and most blatant ongoing thefts of intellectual property in history" — a "brazen campaign of illegally torrenting, scraping, and downloading copyrighted works on a massive scale." That is advocacy, and you should read it as advocacy. But the specificity underneath it is what makes the personal claim pleadable at all.
The complaint alleges that Mann "personally used BitTorrent to unlawfully download and upload via torrenting millions of pirated books from LibGen," that he "expressly directed other Anthropic employees to do the same," and that "Dr. Amodei expressly authorized and directed Mr. Mann's infringements." It alleges he discussed it on internal Slack in real time, shared a screenshot of his own torrent client, and wrote a script to pause the download when his disk filled up — a script the complaint says he described as "a cute little libgen babysitter." It alleges he knew LibGen's reputation before he ever joined the company and called the site "sketchy AF" in internal Anthropic discussions, and that Anthropic's own archive team called it a "blatant violation of copyright."
None of that has been proved. It is a plaintiff's filing in a live case; there is no answer yet, no motion decided, no finding of any kind. Hold it at exactly that distance. But notice what kind of allegations they are. They are not claims about what a model learned. They are claims about what two named people did on a Tuesday, with a laptop, knowing what the website was.
Why the download and not the training
Here is the part that reorganized my thinking, and it comes from a case Anthropic already lost.
In June 2025, Judge William Alsup decided the authors' class action against the company, and the ruling splits cleanly down the middle. On training, Anthropic won, and won big: the use of the books to train Claude "was exceedingly transformative and was a fair use." Alsup went further than he needed to — he called the purpose and character of training "transformative — spectacularly so." If you have been arguing that training is theft, a federal judge disagreed with you in unusually warm language.
And then, in the next breath, he took the company apart on the other half. "Anthropic had no entitlement to use pirated copies for its central library. Creating a permanent, general-purpose library was not itself a fair use excusing Anthropic's piracy." The sentence the whole industry now quotes runs to fifteen words: "the person who copies the textbook from a pirate site has infringed already, full stop."
Read that again with the business model in mind. It does not matter what you did next. The download was a completed act of infringement before the training ever started. Buy the book and you owe nothing further; take it from a pirate site and the wrong is already done, whatever transformative thing you build afterward.
Alsup even flagged that he suspected the rule was broader — that piracy of otherwise-available copies is "inherently, irredeemably infringing even if the pirated copies are immediately used for the transformative use and immediately discarded" — and then, carefully, declined to decide the case on that ground. The hedge is his. I am keeping it. He also left behind the line that has been on a thousand slides since: "There is no carveout, however, from the Copyright Act for AI companies."
That case settled for $1.5 billion, with the company also agreeing to destroy the downloaded files. And the new complaint's theory of why the publishers are in court at all is that the number did not work: Anthropic "clearly considers that to be just the cost of doing business," they write, and "$1.5 billion is obviously not a large enough settlement to deter infringing conduct."
Whatever you think of that as rhetoric, it explains the strategy precisely. If the corporate check is affordable, aim somewhere the check does not reach.
So that settles it. Right?
No — and the honest version of this piece has to spend real time here, because there are at least three good reasons to think the personal claims fail.
First, this is not new, and I nearly told you it was. The idea that naming founders is unprecedented is wrong, and it took about ten minutes of checking to find out. In April, the same law firm, in the parallel case running next door, moved its contributory-infringement claim off the company and onto Amodei and Mann personally — the same two men, four and a half months earlier. In May, five book publishers and the novelist Scott Turow sued Meta and Mark Zuckerberg personally, alleging he "personally authorized and explicitly directed the infringement." Sony and Warner are the third plaintiff group in five months to try this, not the first.
That is less thrilling than a scoop and considerably more important, because a tactic used once is a stunt and a tactic used three times by different plaintiffs is a strategy forming in real time.
Second, there is a specific legal reason it is happening now, and it is not that lawyers got braver. In March, the Supreme Court decided Cox Communications v. Sony Music and tightened contributory infringement to require intent: a service provider is liable "only if it intended that the provided service be used for infringement." Intent is a state of mind, and companies do not have those. People do. If you must prove somebody meant it, you go looking for a somebody. The Slack messages are not colorful detail; they are the whole point.
Third: Anthropic has already argued this, though not as broadly as it first sounds. Amodei's lawyers moved to dismiss the direct-infringement claim against him personally in the parallel case on August 3, with a hearing set for November 4. Their argument is not evasive: "the complaint asserts zero facts of any unlawful copying by Dr. Amodei himself" — the allegations against him "without exception, sound only in contributory infringement." Corporate officers are not ordinarily liable for their company's torts, though an officer who personally directs an infringement is a long-standing exception, which is the whole of the plaintiffs' theory.
Read that quote to the end, though, because the trade press does and most summaries do not. Contributory infringement is Count II — the count aimed at the two founders alone — and it is a claim Amodei's motion pointedly does not ask the court to dismiss.
And the publishers have an answer to "zero facts" that did not come from them. The complaint quotes Mann's own filed answer in the parallel case conceding that he "discussed the potential acquisition of LibGen data with Dr. Amodei," and that "Dr. Amodei approved" the torrenting. A co-defendant's pleading is not plaintiffs' rhetoric.
Anthropic's public answer is of a piece with that. The company says training generative AI models is a transformative fair use "as the court held in Bartz," and it has told reporters this is the third lawsuit from the same lawyers, "recycling allegations from cases already before the courts." On the fair-use half, they are quoting a judge accurately. On the recycling, they have a point about the pattern even if the pattern is the story.
So: a real chance the direct claim against Amodei falls in November. That would narrow the case without ending the personal-liability theory, because the contributory count is not what the motion attacks. I would not bet the house either way. What I would say is that a theory can fail in court and still change behavior, because the discovery it licenses — the Slack logs, the deposition where Amodei reportedly testified that verbatim output of copyrighted material is "against the law" — happens whether or not the count survives.
Japan wrote the opposite rule, and it still says do not do this
Now step outside American law entirely, because the acquisition-versus-training split that governs this fight is not a law of nature. It is a local accident, and there is a country where the accident fell the other way.
Since January 2019, Article 30-4 of Japan's Copyright Act has permitted exploiting a work "in any way and to the extent considered necessary" where it is not the purpose "to personally enjoy or cause another person to enjoy the thoughts or sentiments expressed in that work" — and it names data analysis explicitly. Training a model is the paradigm case.
Now read the provision for what is not in it. There is no lawful-access condition. Nowhere does the text ask where the copy came from. The European and British text-and-data-mining exceptions require lawful access; Singapore's requires it too. Japan's does not mention it. On a plain reading of the statute, the download that cost Anthropic a billion and a half dollars in San Francisco falls inside the exception in Tokyo.
Which sounds like a free-for-all, and here is where the comparison earns its place: it is not one. Japan's Agency for Cultural Affairs tells developers in its official English guidance that collecting training data from sites known to distribute pirated content "should be strictly avoided," and that a developer who does so faces "a high possibility" of being held responsible — but for "any copyright infringement caused by the generative AI developed using the training data," not for the download. Same disapproval, entirely different hinge. America punishes the acquisition and forgives the training. Japan forgives the acquisition and reaches for the output.
I want to be careful not to oversell this. The guidance is administrative interpretation, not statute. And the reading that Japan's exception is sweeping is contested by people who know it well: the copyright analyst Hugh Stephens argues Japan has defined its exception "carefully and narrowly", and notes in the same breath that it "does not make any explicit distinction between legally accessed and non-legally accessed materials." The article's own proviso is a further brake: the exception drops away where the use would "unreasonably prejudice" the copyright owner's interests. Both things are true, which is why the comparison is interesting rather than tidy.
The point is not that Japan is wiser. It is that the question American litigation treats as fundamental — how did you get the file — is one another advanced economy answers by looking downstream instead. Two coherent rules, opposite hinges. That should make anyone confident about the obvious answer here a little less so.
Run it forward
Play out the next two years the boring way.
The personal claims survive in one court and fail in another, because that is what happens to novel theories. But the discovery has already run. Every frontier lab's general counsel has now read a complaint built almost entirely out of a co-founder's own Slack messages, and drawn the obvious conclusion: the exposure is not in the model, it is in the procurement log.
So the behavior changes, and mostly not in public. Data acquisition gets a paper trail and a signature line. Somebody senior stops being cc'd. The engineer who once solved a data problem over a weekend now files a request, and the request is reviewed by someone whose job is to say no — which is good for the people whose work was being taken, and also how a field stops being able to move quickly. Both at once, from the same cause.
And the next wave of models is trained on licensed corpora, which sounds like a happy ending until you notice who can afford one. The publishers' own complaint says they "have entered licenses permitting the authorized use of their musical compositions in connection with AI" — they are not against this technology, they are against not being paid for it. Fair enough. But a world where you need Sony's signature to train a competitive model has fewer people training models, and the survivors are the ones who can write the check.
What the people who follow this closely are saying
The striking thing is where the agreement is.
From digital rights, the Electronic Frontier Foundation says the training half was decided correctly — that fair use protects training precisely because it is "transformative—spectacularly so." From the free-market right, the R Street Institute argues that treating AI training as non-infringing is what keeps copyright from becoming "a barrier to innovation". Different premises, same conclusion about training.
And then both of those camps concede the other half. Public Knowledge, no friend of copyright maximalism, states the split without flinching: using pirated material for training did not negate fair use, but the companies were "potentially liable" for "engaging in book piracy by acquiring and keeping the pirated books that they could have procured through legal means." When the people who most want AI training to be lawful still say the download was a separate wrong, that is not a partisan position anymore. That is just the law as it stands.
The rightsholder side has the sharper rejoinder to my moat worry. In March, eight music-industry bodies including the RIAA and NMPA told the court that a functioning licensing market for AI training already exists — "one that Anthropic has chosen not to participate in." If the market is already there, the moat is one the defendant declined to cross, not one the plaintiffs dug.
One necessary complication from the authors' own side. The Authors Guild notes that Alsup certified the class only for piracy, not for training — so the celebrated fair-use holding "applies only with respect to the three named plaintiffs." And one of those named authors told NPR he is "much poorer for this settlement, ironically."
What does this mean for you?
You are not going to be deposed about a torrent client. Here is what transfers.
Separate the two questions whenever you see this argued. "Should machines be allowed to learn from books?" and "how was this specific file obtained?" are different fights with different answers, and almost every headline blurs them. When someone tells you a court ruled on "AI and copyright," ask which half. The answer is usually that training won and acquisition lost.
If you make things, your leverage is at the source, not the output. Chasing a model that has already learned from your work is close to hopeless. The place the law currently bites is the moment of copying — which means where you host your work, and under what terms, matters more than any argument about style or similarity.
Watch November 4 — and know what it does not decide. The motion being heard that day attacks only the direct infringement claim against Amodei. The contributory count, the one aimed at the founders alone, is not challenged. Whichever way it goes, the theory that executives can be reached survives it.
Do not read a filing as a finding. Everything the publishers allege here is unproven, and some of the most quotable material — the Slack messages especially — is exactly the kind of thing that reads as damning in a complaint and ambiguous in context. I have tried to fence every claim in this piece. Extend the same courtesy to the next complaint you read about a company you like less.
The question I had been ignoring
Here is the part I keep turning over. The most consequential AI legal fight of the decade is not being decided by any new principle about machines. It is being decided by the oldest rule there is: you cannot take the thing without paying for it, and building something clever afterward does not launder the taking. Alsup wrote that the training was spectacular and the library was theft, and both halves of that sentence have survived every appeal to novelty anyone has thrown at them.
Whether that rule should reach past a corporation and touch the people who ran it is the genuinely open question, and I do not know the answer. I want it to be yes when the defendant is a company I distrust — and the wall between a company and its officers protects a great many people I would want protected. Those instincts do not resolve, which is usually the sign of a real question.
November will tell us something. Until then, the thing to hold onto is smaller and firmer than any of it: the fight was never about whether the machine understood the song. It was about how the song got onto the hard drive.
If you have ever argued with someone about whether AI training is theft — and you have — this is the piece to send them, because you were both arguing about the wrong half. More of the same at HAIA and on the Substack.



