Almost nothing in this piece is mine.
I read a court filing, a statute in translation, a hearing report by a journalist who sat through all six hours of it, and a PDF signed by five academics. Then I compressed the lot and handed you the compression. What you are reading is a summary of other people's labor, and the only thing that lets me sleep is the links.
On Thursday, September 3, that question takes its next public step in a courtroom in Luxembourg: when a summary of a press article carries enough of the article inside it, is that a copy? The summarizer in the dock is not me, though. It is a chatbot — and the distinction I would reach for first, that I link back and it mostly does not, may not be the one European law cares about.
The dolphin story that reached Luxembourg
The case that may set Europe's rule for machine reading began with a story about a Hungarian pop singer's plan to put dolphins in a lake-side aquarium near Lake Balaton, on a regional news portal called balatonkornyeke.hu. Somebody asked Google's Gemini to summarize it in Hungarian. It did — in enough detail, the portal's owner says, to include the material it had paid to produce.
That owner is Like Company, a Hungarian press publisher, and its complaint is broader than one dolphin. It alleges that between June 13, 2023 and February 7, 2024, Gemini systematically extracted and displayed significant sections of its protected press publications without permission or payment, and that those publications trained the underlying model. All of that is allegation; nothing has been proved.
The Hungarian judges — the Budapest Környéki Törvényszék, the Budapest Environs Regional Court — did the honest thing with a question their statute book could not answer. They stopped, and in April 2025 sent it to the Court of Justice of the European Union, where specialists spotted it as the CJEU's very first referral on artificial intelligence and copyright.
Three questions went up, and the first is not a lawyer's flourish — the referring court wrote it into the docket itself. It asks whether a chatbot displaying text partially identical to a press publisher's web pages is a "communication to the public," and then whether it matters that the answer came from a process in which the chatbot "merely predicts the next word on the basis of observed patterns".
Read that again. A national court, unprompted, asked Europe's highest court whether next-token prediction is the kind of thing copyright has words for. Questions 2 and 3 go to the other end of the pipe: is training a model an act of reproduction at all, and does the text-and-data-mining exception forgive it?
Two provisions carry the weight. Article 15 of the 2019 Copyright in the Digital Single Market Directive is the "press publishers' right": control over the online use of press publications — except it does not reach "the use of individual words or very short extracts." Nobody has told a court how short is short.
Article 4 of the same directive is the mining exception, which lets anyone copy lawfully accessible material for text and data mining — the industrial reading that makes a model — unless the rightsholder has expressly reserved that use "by machine-readable means". That is the hinge of the whole European settlement: your rights survive the machine only if you told the machine, in machine language, to go away.
Today that means a single line in a text file at the root of your domain: Google-Extended, a product token publishers put in robots.txt to control whether crawled content trains future Gemini models. Google's own documentation notes it does not affect Search ranking.
On March 10, 2026, the fifteen-judge Grand Chamber sat under President Koen Lenaerts, with Octavia Spineanu-Matei as judge-rapporteur. Six hours. The first hearing on generative AI and copyright the Court had ever held. Advocate General Maciej Szpunar — an adviser to the Court, not one of the judges who rule — is expected to deliver his opinion on September 3, the judgment months behind it.
As I write, that opinion is not out. I have no idea what is in it, and neither does anyone telling you they do.
The strongest version of Google's answer, which is stronger than you think
Here is where I argue against my own instincts.
Google's position is not that copying is fine — it is that there is nothing there to copy. Google's answer is that there is no copy to find: the system tokenizes training data and generates text probabilistically rather than storing and retrieving articles; any resemblance is incidental or a hallucination; and no "new public" is reached, since users could have read the original page anyway. Its lawyers put it statistically at the hearing: the model learns linguistic patterns and builds text by predicting the next word, so an output matching a text that exists elsewhere is inherent in the method, and cannot be excluded.
Which produced the day's best moment. Courthouse News, whose reporter sat through the session, records the judge-rapporteur pressing Google directly on the risk of textual memorization: can we infer, she asked, that regardless of a chatbot's intended purpose, you acknowledge that risk? Counsel conceded statistical systems cannot completely rule such outcomes out, while insisting there was no evidence this model had reproduced any specific article. Like Company's lawyer, in the same report, said they were deciding "the future of quality journalism."
There is a second argument I want to state carefully, because it is the one most often repeated wrongly. Google's counsel argued that Article 4's opt-out already protects rightsholders through robots.txt and Google-Extended — and, per Bird & Bird's account of the hearing, Google told the court the applicant had not used those options. That is Google's submission about its opponent, not a finding. I could not verify what was in that publisher's robots.txt, and no one else appears to have, publicly.
Now the uncomfortable part. One of the sharpest academic readings concedes most of the publisher's technical claim and still lands on Google's side. Before the hearing, a European law professor argued that training an LLM on press articles typically does constitute reproduction — it copies protected expression into memory for analysis — but that generative training fits the mining definition, so if the publisher never opted out, the exception covers it. He calls that the bombshell question: if Article 4 does not cover training, AI providers across Europe are in very deep copyright trouble.
Does the memorization worry have legs beyond one judge's question? Some. A peer-reviewed measurement of benign prompts — write a letter, write a tutorial — found up to 15% of the text output by popular conversational language models overlaps with snippets from the Internet, with worst cases where 100% of the content could be found exactly online. Hold it at arm's length: several popular models, generic writing tasks, no adversarial prompting. Not Gemini, and not evidence about anybody's Hungarian dolphin story. The phenomenon exists in the family. That is all it says.
The case that this is the wrong case
Then the objection I did not expect, from people who spend their lives on this law.
A week before the hearing, five copyright scholars writing for the European Copyright Society urged the Grand Chamber to exercise caution. The reference, they said, is factually murky with respect to the technology and services at stake, conflating "chatbot," "large language model" and "search engine" as though those were one thing. They called it convoluted, and warned against drawing rules of wide application from narrow and murky factual settings. The blunt version: "imprecise questions generate unreliable doctrine."
This is not a technicality. Ask a court whether "the chatbot" copied something, when the real system involves a model, an index, a retrieval step and a generation step, and you may get an answer that binds 450 million people and describes no machine that exists. The European Commission — which intervened, and urged the judges to keep training and output apart — went further, suggesting the reference might be partly or wholly inadmissible: the questions dwell on how Gemini works rather than on any specific act of infringement.
A landmark case, in other words, that serious people think should never have been heard.
Tokyo asked a completely different question, eight years ago
Here is what makes Europe's approach look like a choice rather than a law of nature.
Japan got here first, and framed the question with almost nothing in common with Brussels. Article 30-4 of Japan's Copyright Act permits exploiting a work in any way and to the extent necessary — including for data analysis — in any case where "it is not a person's purpose to personally enjoy or cause another person to enjoy the thoughts or sentiments expressed in that work." One proviso: it does not apply where the act would "unreasonably prejudice the interests of the copyright owner."
Notice what the test asks. Japan's statute focuses on the human experience — was anyone there to enjoy the work? — where Europe's asks whether you sent the right signal. Those provisions came into effect on January 1, 2019, enacted as Act No. 30 of 2018, four months before the EU adopted the directive now in front of the Grand Chamber.
And there is no off switch. Waseda's Tatsuhiro Ueno — who traces the provision's ancestor to a 2009 amendment, the first such exception in the world, and who cheerfully calls his country a "Paradise for Machine learning" — writes in a paper hosted by the World Intellectual Property Organization that Japan's version is unusually robust: no restriction on who mines or why, no lawful-access requirement, and rights holders cannot opt out of any TDM activities. On the orthodox reading the "unreasonable prejudice" proviso catches exceptional cases only, since the 2018 process treated these uses as acts that do not normally harm rightsholders.
Japan's own newspaper industry is not thrilled. The Japan Newspaper Publishers & Editors Association complains that the provision governing AI training contains no explicit opt-out provision allowing rights holders to refuse use, and no explicit arrangements regarding technical measures. No line you can add to a file turns it off.
What Japan has instead of a switch is a document: the Agency for Cultural Affairs' general understanding on AI and copyright, an overview of a subcommittee report dated March 15, 2024 that says of itself, in writing, that it is not legally binding. Tokyo's follow-up that year was a checklist rather than a statute.
And here is the detail that turned my head around: robots.txt turns up in the Japanese guidance too, on the opposite side of the line. There it is not a switch that revokes the exception. It is evidence — a site that has taken technical measures against AI-training crawlers may, in certain cases, be assumed to have meant to commercialize the data, so reproducing it anyway can fall inside the proviso and lose the exception. The same two lines of text: in Europe, the instrument that keeps your rights alive; in Japan, a clue about your commercial intent.
Did that buy Japanese publishers peace? No. Last year Nikkei and the Asahi Shimbun sued Perplexity AI in the Tokyo District Court, seeking an injunction and ¥2.2 billion each, alleging the company reproduced and saved their content and ignored coding that indicated content was off-limits. Those are claims in a filed suit, not findings. But note where the fight moved: not to whether the machine could read, but to what came out the other end.
Just imagine the file that ends up governing a continent
Play Europe's answer forward, and take the machine-readable reservation seriously.
Imagine a world where the most consequential legal document a publisher owns is a plain text file most of its staff have never opened, last edited by a contractor who built the site in 2019 and has since changed jobs. A newsroom that reserved its rights keeps them. One that did not, because nobody knew the file existed, has consented to a decade of its own reporting being read industrially. Same law, same journalism, opposite outcomes, decided by whoever had the server password.
Then the maintenance nobody budgeted for. The reservation has to be legible to a specific crawler, and every provider ships its own token, so the line you wrote in 2024 says nothing about the model launched in 2026. Opting out becomes a subscription — a list you keep current forever, against an industry adding entrants faster than a regional weekly can notice. Publishers with a compliance team will manage. The ones who most need the money will not.
And then the map: the same article, the same model, two continents. In one the training was lawful because the publisher failed to object; in the other, because nobody was there to enjoy it. Neither jurisdiction asked whether the writer got paid.
What the smart people are saying, and they are not saying the same thing
The fight over Europe's mining exception splits in ways that map onto no political axis.
From the innovation-and-growth side, the Information Technology and Innovation Foundation told the Commission in June 2026 it should preserve and strengthen the text and data mining exception as the foundational framework for AI training, police output-side harms instead, and resist mandatory licensing. Industry supplies the scary number: CCIA Europe publicized a study by Implement Consulting Group with the Ifo Institute, commissioned by a tech trade association, warning that restricting the framework could cost the EU economy up to €600 billion annually. "Could" and "up to" carry that sentence. It is a projection with a client, not a measurement.
On the day the Grand Chamber sat — March 10, 2026, a coincidence I checked twice because I did not believe it — the European Parliament adopted recommendations protecting copyrighted work from AI by 460 votes to 71, with 88 abstentions. The European Federation of Journalists welcomed it and went further, arguing AI use of journalistic content should be subject to "clear prior authorisation and fair remuneration," managed collectively — not an opt-out at all.
Then a third position that spoils both stories. COMMUNIA, from the digital-commons side, reviewed how the press publishers' right actually performed after 2019 and concluded that an exclusive right does not necessarily translate into revenues — and that although Brussels meant "very short extracts" to be one harmonized concept, three national laws set their own length limits. Europe already ran this experiment. A right did not automatically become money.
One more, carefully. A copyright academic who followed the hearing read the Advocate General's questions from the bench as pointing toward doing away with the prevailing input-output dichotomy — training and output judged as one continuous act — with Google's counsel objecting that the rights must be analyzed separately. That is one observer's reading of questions, published three days later. Questions are not conclusions, and the opinion is not out. I include it as the most interesting thing reported from the room, not as a forecast.
What does this mean for you?
If you publish anything online, open your robots.txt today. Under the European framework, that file is where your rights get exercised, and most site owners have never looked at it. You may be surprised who is welcome.
Do not assume the opt-out is one switch. Every provider ships its own crawler token, and blocking one blocks one. A reservation written two years ago says nothing about a model launched this year.
Read the September 3 document as advice, not a verdict. An Advocate General's opinion is influential and non-binding; the judgment is months out. Anyone saying this week that Europe has "ruled" on AI training is selling something.
Treat a chatbot summary of the news like a colleague's account of an article they skimmed. Broadly right, occasionally carrying whole sentences it did not attribute, no substitute for the page.
Notice which question the fight is about. Nobody senior argues that models cannot read. They argue about who has to be asked, and who has to be paid.
When you summarize somebody's work, link it. Not because a court will make you — because the person who reported it deserves the traffic.
The lesson, as I see it
Two sophisticated legal systems looked at the same machine and asked questions that barely rhyme. Europe's is procedural: did you object, in a format a crawler can parse? Japan's is philosophical: was anyone there to enjoy the work? Europe rewards whoever reads the documentation; Japan rewards whoever can prove harm afterward. Neither is about paying the person who went to the council meeting.
That is the real gap, and no answer on Thursday closes it. If the Court decides training is reproduction and the exception does not cover it, Europe gets a licensing economy — which the copyright scholars themselves warn may end up dominated by non-European multinationals, without solving creator consent or remuneration. If the exception holds, the burden stays on a text file that a regional Hungarian paper covering a dolphin story was never going to have the staff to maintain.
What I want from September 3 is not a winner. It is precision — an opinion that names which machine it describes, so the next case can be argued about facts instead of vocabulary. The scholars are right that you cannot build doctrine on confused terms. But the questions are asked, the Grand Chamber heard them, and "the wrong case" is not an outcome anybody gets to choose.
Somebody wrote about the dolphins. Everything after that is arithmetic about who owes them what.
If this was useful, pass it on. Summarizing it out loud to a friend is, for now, still legal on both continents.





