A few months ago I was writing about a state statute and needed the exact section number. I asked a model. It gave me one — section, subsection, year of amendment, formatted so cleanly it looked lifted straight out of a code book. I dropped it into my draft and moved on, feeling efficient. Two days later I went to pull the actual text. The section existed. It was about something else entirely.
Here is the part that stayed with me. The wrong answer had felt exactly like the right ones — same cadence, same confidence, same tidy formatting. No tremor, no hedge, no tell. I had not been fooled by a bad answer. I had been fooled by a good-sounding one, which is a completely different failure mode and a much harder one to guard against.
I was writing a newsletter. Nobody went to prison. But the profession most fluent in evidence on the planet has been making the identical mistake, in front of judges — and this year the judges stopped being polite about it.
What the Sixth Circuit actually punished
Be precise about the case everyone cited this spring, because the precision is the story. In March, a federal appeals court ruled in three consolidated appeals in which the panel counted over two dozen fake citations and misrepresentations of fact in one side's briefs, listing every one in an appendix. The tone is not weary. It is cold. Of one authority it said the citation simply does not generate a case — five words that should chill anyone who has ever filed anything: we cannot find this case. The two attorneys were each ordered to pay $15,000 into the court registry as punitive sanctions, plus double costs and the opponents' full fees. The legal-technology trade press called it what may be one of the most significant appellate sanctions rulings yet involving fabricated citations. The hedge is theirs, and I am leaving it where they put it.
Now the detail nearly every headline got wrong. The show-cause order had instructed those lawyers to tell the court whether they used generative AI to write the briefs. They did not answer. And the same write-up is careful about something the headlines were not: the court did not expressly find that the fabricated citations came from a model. The lawyers' response was to declare the order void on its face for want of an Article III judge's signature — a theory the Supreme Court had already twice refused to entertain, denying petitions for mandamus from these same two lawyers demanding that the court's clerk stop signing its orders. The second of those denials came four days before the sanctions opinion issued, and the opinion cites it.
So nobody was fined for using AI. They were fined for putting their names on citations that do not exist. The rule the court states is deliberately source-agnostic: no citation, from a generative model or any other source, that a lawyer has not personally read and verified. And the panel explained its number with unusual candor — smaller fines, it wrote, have plainly been inadequate.
That is a court telling the bar the gentle phase is over.
Nor is it isolated. An Oregon federal magistrate judge faced filings with fifteen nonexistent cases and eight fabricated quotations, and imposed roughly $110,000 in fines, fees and costs — a notorious outlier in both degree and volume, as the opinion put it. The fine component alone was $15,500, computed with a fake citation/quotation formula borrowed from the state appeals court. There is now, in other words, a price list. In June a senior federal judge in the Northern District of Mississippi went further, barring two lawyers from appearing in the district for two years for blindly relying on technology. And the detail that stings — sanctions landed on attorneys on both sides of the same case. Both filed fiction. Neither caught the other's.
Zoom out and the shape is unmistakable. One industry analysis counted at least $145,000 in sanctions for AI-generated fake citations in the first quarter of 2026 alone — $5,000 in January, $250 in February, then March detonated — and a columnist watching the same docket had already flagged that the sanctions keep getting more severe. A researcher's tracker listed 1,809 cases identified so far when I checked it on July 25, 2026, 1,250 of them American — and he cautions it is not exhaustive. The real number is larger, and it moves weekly.
Buying the expensive tool does not save you either. Stanford researchers found the flagship products sold to lawyers hallucinate between 17% and 33% of the time, despite marketing that promised hallucination-free citations — and general chatbots asked about random federal cases produce legal hallucinations that are alarmingly prevalent, 58% to 88% of the time.
Now turn around and look at the bench
In March, Northwestern researchers published survey results showing more than 60 percent of judges who responded already use at least one AI tool in their work — mostly legal research and summarizing. Hold that number accurately: 112 replies from a stratified sample of 502 federal judges, in December 2025, voluntary response rather than a census. But the follow-up detail is the one that matters — 45.5% reported that AI training had not been provided by their court. Majority adoption. Minority training.
To be clear, because the distinction is load-bearing: a judge using AI for research is not misconduct. It is what lawyers do too. The trouble arrives where it always does — at the unverified output.
Last autumn the chairman of the Senate Judiciary Committee wrote to two federal judges about error-riddled orders from their courts. Their replies were released publicly, and both judges admitted staff had used generative AI to draft orders that misquoted state law, attributed fake quotes to defendants and referenced people who never appeared in the case. In the Southern District of Mississippi a law clerk had used Perplexity on the docket; in New Jersey, a law school intern used ChatGPT for legal research — acting, on the judge's own account, without authorization and without disclosure.
Both orders were withdrawn. As the reporting stressed, members of their staff used artificial intelligence — chambers, not judges personally typing into a chatbot. That distinction is real and I will not blur it. It is also, to the litigant on the receiving end, irrelevant. An order with a phantom party in it is an order with a phantom party in it.
So here is the asymmetry, stated as carefully as I can. When a lawyer files fiction, the consequence is codified. Rule 11 attaches certification to the signature: sign, and you have represented that your legal contentions are warranted. Violate that and the court may impose an appropriate sanction. Fifteen thousand dollars. Two years out of a district.
When a judge's chambers files fiction, what happens is a withdrawn order and a letter to a senator. The federal judiciary's interim guidance, per the reporting on it, contains a recommendation to consider whether AI use should be disclosed. Consider. That single word does more work than any other in this article. The guidance does say judiciary users remain accountable for all work performed with AI, and warns against delegating core judicial functions to a machine. Not nothing. But no signature, and therefore no Rule 11.
Let me fence that honestly: no court has held that judges owe a verification duty comparable to Rule 11, and nobody has exempted the bench from anything. There is simply no rule that answers the question. The nearest thing to an answer is a bill — in California, legislation would require a judicial officer to disclose whether they relied on generative artificial intelligence in drafting any ruling. It cleared the state Senate unanimously in January and is still working its way through the Assembly. It is not law. Nor is a new Rule 707 for machine-generated evidence, still in rulemaking after its comment period closed in February.
The judiciary's own materials saw this coming, if you read them exactly. A federal judicial-education primer warns that juries may assume machine output has the imprimatur of "science" or "technology", potentially lending it false authority or undue weight. That is a claim about jurors looking at evidence, not about drafting, and there is no jury in chambers. But the mechanism — fluent, technical-looking output borrowing authority it has not earned — does not need one.
Hypocrisy? Honestly, no — and the counter-argument is strong
The lazy version of this piece writes itself: courts punish us, exempt themselves, hypocrites, the end. That does not survive contact with the argument.
The strongest rebuttal is that the asymmetry is not a loophole — it is the architecture. As one technology-law practice group put it, the attorney who signs carries a non-delegable responsibility to ensure the filing makes no false statements of fact. That is the entire function of a signature: you cannot subcontract it to a paralegal, a book or a model, and "the tool did it" was never going to work, because the tool cannot sign. Judges are not adversaries certifying claims to a neutral; they are the neutral, checked by a different mechanism — appeal, recusal, conduct complaints, publication of their errors.
A second, sharper critique comes from the free-market side rather than the civil-liberties side, which is why it is worth hearing. Analysis at the R Street Institute argues AI disclosure rules are largely theater, because knowing whether AI helped draft a document has no bearing on the legal obligations of candor that already bind every lawyer — and that AI use is anyway subject to longstanding legal principles. Novelty exempts nobody.
I find that half-persuasive. Half, because it assumes the duty is enforceable against everyone who now touches a draft. Against a signing lawyer it plainly is — ask the two attorneys who each owe $15,000. Against an unauthorized intern in chambers, the mechanism is a senator writing a letter.
Nor is judicial AI use inherently a scandal. A sitting federal appellate judge has experimented publicly with language models as an interpretive aid, and his framing, quoted in a write-up at the American Enterprise Institute, is deliberately small: that courts consider whether LLMs might provide additional datapoints when pinning down what a word ordinarily means. A law-and-technology review finds the idea useful in determining ordinary meaning yet impractical at scale. And a practitioner reading of the sanctions opinion stresses the obvious: the court did not reject the use of AI — it rejected not reading what you cite.
In England and Wales, somebody wrote the manual first
Now the comparison I cannot stop thinking about. The Courts and Tribunals Judiciary of England and Wales did not wait for a scandal. On December 12, 2023 — before most of the American sanctions docket existed — it published guidance for judicial office holders on artificial intelligence. And it kept rewriting it: the current edition is dated October 31, 2025 and says on its own first page that it updates and replaces the guidance document issued in April 2025. Three editions in under two years, while the American federal answer was still interim.
Note who it reaches. Not just judges — it runs to every judicial office holder under the Lady Chief Justice and the Senior President of Tribunals, and onward to their clerks, judicial assistants, legal advisers/officers and other support staff. Precisely the people who, in New Jersey and Mississippi, put a phantom party into a federal order.
What is striking is not that it is strict. It is that it is practical. A heading in it reads, without ceremony, Tasks not recommended — under which sit legal research and legal analysis — below a list of what the tools genuinely are good for: summarizing large bodies of text, drafting presentations, composing emails. And it gives the reason rather than the rule. Public chatbots do not provide answers from authoritative databases; they generate the most likely combination of words, which is not necessarily the most accurate one, and they may make up fictitious cases, citations or quotes. Anything typed into one is published to all the world. Judges are personally responsible for material which is produced in their name.
And notice what it is not. It is not a rule. Nobody gets fined under it. The American instinct was to wait for someone to do it badly, then bill them for it — Rule 11 sanctions fire after the fake citation reaches the docket, after opposing counsel burns hours chasing a phantom case, after a real dispute has been polluted. The English instinct was to hand every judge a short, readable document explaining the failure modes before they ever opened the thing.
To be fair, the state courts' national center keeps issuing new guidance and resources; America is not a vacuum. But the federal guidance for judges arrived as an interim recommendation to consider disclosure, years in, while England had put a competence document in every judge's hands before the first American six-figure sanction landed. Prevention is not more virtuous than punishment. It is just cheaper, faster and considerably less humiliating.
Now run the tape forward
Give this three or four years and the easy version of the problem disappears: filing systems will refuse a brief whose citations do not resolve against a live database. Trivial engineering, unsolved only because nobody has been embarrassed enough to build it.
What replaces it is harder. The next error is not a case that does not exist. It is a case that exists, is quoted accurately and is characterized wrongly: the holding subtly overstated, a dissent presented as the majority, the distinguishing fact quietly dropped. No database check catches that. Only a human who read the case catches that.
Now put that failure in an order rather than a brief. A litigant loses; on appeal, counsel asks the obvious question — was any part of this ruling drafted with a generative model, and did anyone verify it? Today there is no answer they are entitled to. Just imagine the first appellate argument turning on a discovery request into a judge's chambers; the first standing order demanding prompt logs from counsel; the first motion asking why that rule stops at the bar of the court.
Then imagine the version I actually expect, because institutions are practical: a closed, audited AI system inside the judiciary's own network, every query logged, every citation resolved against the official reporter before it reaches a draft. Not a ban, not a free-for-all — a tool with a flight recorder. And with it, the question nobody wants on the record: if a machine drafted 70% of an order and a human approved it in four minutes, whose reasoning is in that document?
What the people who study this are saying
Those who have stared at this gap longest come from strikingly different directions. Scholars at the Brookings Institution argue the pressing need is to improve judges' ability to understand the technical issues in AI litigation — before worrying what judges do with AI, worry whether courts can competently rule on it. From the civil-liberties side, the ACLU of New Jersey warns that algorithms in the justice system are notoriously opaque, the what, how and why of their use shrouded in mystery.
A retired federal judge and two scholars — one a law professor, one a computer scientist — writing together in Duke's judicial-studies journal go at the mechanism: without testing, nobody can know whether an AI system is more or less biased than the alternative, and the opposing party must get a reasonable opportunity to challenge assertions built on it. That principle is already being written into evidence law — the New York City Bar's Rule 707 comments would gate machine-generated inferential evidence through expert-style reliability screening. In a roundtable in the same journal, one scholar would ban behavioral profiling in courts outright and another declines to rule out any technology per se — the whole spectrum, held together only by a third's insistence on meaningful human control.
An article in the Stanford Technology Law Review argues Rule 11's inadequacy here is what drives individual judges to write their own standing orders, a patchwork substituting for a national answer. And the survey behind all this came out of Northwestern, where the Director of Law and Technology Initiatives ran it; he writes at LegalTech Lever and has been notably unhysterical throughout. The tools are useful. The training is missing. Fix the second thing.
What does this mean for you?
You are probably not filing appellate briefs. You are exposed anyway, because the mechanism that caught those lawyers is pointed at you every day.
Treat fluency as no evidence at all. A model's confidence is a property of its writing, not its knowledge. The moment you think "that sounds right" is the moment to check, not to relax.
Verify anything with a proper noun, a number or a citation. Names, dates, statute sections, dollar figures, study results — precisely what models fabricate, and precisely what you will be held to. Structure, phrasing and summarizing your own text are where they are strong.
If you are hiring a lawyer, ask two questions. Do you use AI in research or drafting, and who verifies citations before filing? Any competent lawyer in 2026 has a crisp answer. Hesitation is the answer.
Understand what you are signing. Filing, expert report, compliance memo, grant application — if your name is on it, the duty is yours. Not an AI rule; the oldest rule there is, which AI has made far easier to break by accident.
Ask your employer for the manual before the incident. If AI tools arrived with no written guidance on what they are good and bad for, you are living in the American model. Ask for the English one.
If you are a litigant, read your own order. Check that the parties named are the parties in your case. Absurd advice two years ago. Not now.
The lesson, as I see it
The system is not broken here, and I want to resist saying it is. Courts caught the problem, named it, priced it. Roughly what you want institutions to do.
But look hard at where the price lands. One lawyer pays $15,000. Another loses two years of practice in a district. One lawyer and their side pay around $110,000 in fines, fees and costs. And the tools sold into this market — marketed on a promise of hallucination-free citations, hallucinating up to a third of the time in testing — pay nothing. (We usually do not even know which tool was involved, or whether one was. No court has made that finding, which is itself part of the problem.) Fining vendors for every bad output ends somewhere silly, so that is not my proposal. But the settlement is unstable: you cannot indefinitely run a system where liability for a technology's failures sits entirely on its least powerful user, the one participant who cannot inspect how it works.
The narrower fix needs no new statute: train the people, publish the manual, log the queries, resolve the citation before it reaches the docket. England worked out that cheap part years ago. The expensive part — deciding what a signature means when the drafting is shared with a machine — is still ahead of all of us.
My vote? Close the training gap before we argue about the disclosure gap. Nearly half the judges who answered that survey said their court administration had never trained them on the tool most of them already use. That is not a scandal. That is a to-do list.
And when a machine tells you something in a beautifully confident voice, do what I should have done with that statute section: go look.
Read what you cite. Cite what you read. Everything else is just very expensive typing. More of this, weekly, from the HAIA Foundation.




