There is a question I avoided asking for most of my adult life, and I avoided it on purpose.
The question is what the words FDA cleared actually mean.
They are printed on a box in my bathroom cabinet — a blood pressure cuff — and on half the health gadgets I have ever bought, and every time I read them I felt the small warm click of reassurance they are engineered to produce. Somebody checked this. Somebody in a lab coat with subpoena power. I never looked further, and I want to be honest about why: I liked the click.
This summer I read a 510(k) file end to end, and I have not been able to unread it.
What is actually in the file
On December 23, 2025, the Food and Drug Administration cleared software called UpDoc under the number K253281 — a prescription tool for people with type 2 diabetes. The patient talks to it, by voice or text, about symptoms and glucose numbers, and it returns insulin instructions.
FDA's record describes a provider portal, a patient app, and a cloud application made of a Conversation Service and a Clinical Service: a clinician sets the parameters, the Clinical Service does the arithmetic, the conversation is the front door. And nowhere in either FDA document — clearance letter or decision summary — do the words artificial intelligence, large language model or generative appear even once. Hold on to that.
The company's framing is different. Six months later, UpDoc announced what it called a first: the first Software as a Medical Device using patient-facing large language models, now running at Cleveland Clinic, Allegheny Health Network and UCSF Health. STAT News put the honest question in its headline — is the LLM an interface or the decision-maker? That "first" is the company's claim, not a regulator's finding.
Now the part that stopped me. The 510(k) route does not ask whether a device works. It asks whether the device is substantially equivalent to something already on the market — a predicate. UpDoc's predicate is Hygieia's d-Nav System, cleared in February 2019 with no change-control plan attached, nearly seven years earlier. d-Nav took numbers on a keypad and returned a dose. It had no mouth. Side by side in the filing, the two differ in one acknowledged way: UpDoc also accepts data by chat and voice. Same intended use, the agency found. The argument, reduced to bones, is that adding a voice does not change what the thing is.
And the evidence for that? The submission's section headed "Summary of Clinical Testing" runs to five words: No clinical testing was performed. The decision summary lists clinical studies as "Not applicable." What was tested was the software, the cybersecurity, and whether people could work the interface.
Careful here. That is a fact about the submission, not an accusation about the product: nothing in the public record shows a patient harmed by this device. The founders separately ran a clinical trial at Stanford Medicine that was not part of the equivalence package — reported by the trade outlet that asked why the equivalence route rather than the one built for novel devices. Whatever that trial showed, the regulator did not weigh it. That is the point.
One more detail, my favorite. The database entry files the device under anesthesiology, regulation 21 CFR 868.1890 — a rule titled "Predictive pulmonary-function value calculator." A talking insulin coach, under a lung-calculator rule. Not a scandal — just a system built for hardware absorbing something it has no shelf for. It also shows the clock: received September 29, 2025, cleared December 23. Eighty-five days.
The test that never asks the obvious question
So how is this legal? Precisely, boringly, because of what the law says.
The statutory test has two prongs: the same intended use as the predicate, plus either the same technological characteristics or different ones that raise no different questions of safety and effectiveness. Does it work is nowhere in there. Not a fringe reading: fifteen years ago the Institute of Medicine, reviewing the pathway at its 35-year mark, wrote that a substantial-equivalence determination does not reflect an FDA evaluation of the safety or effectiveness of either device.
The gate is older than most of the people walking through it. The Medical Device Amendments were enacted on May 28, 1976, and a 2025 chatbot's clearance letter still recites that date, because the chain of equivalence runs back to devices marketed before it. That same guidance states the ratio plainly: FDA requests clinical data for less than 10 percent of 510(k) submissions — a figure it published in July 2014.
Now push an entire industry through that door. As of June 2026, Congress's own analysts reported approximately 1,450 AI-enabled devices authorized for marketing, most cleared through 510(k); in the most precise snapshot published, about 97 percent of the AI-enabled devices on FDA's list had come through 510(k) as of August 2024. Different counts from different years — don't blend them — but the shape holds. One door, overwhelmingly. And the same brief names the design flaw in a sentence: FDA's scheme evaluates a device at a point in time, expecting it to remain largely unchanged, while AI devices are expected to change across their lifecycle. A snapshot standard, for a thing that moves.
The device missing from FDA's own list of AI devices
Here is what I did not expect to find. I will show my work, because this is a live observation, not a document.
FDA maintains a public AI-Enabled Medical Device List — the inventory those congressional counts are drawn from. The agency is candid that it is not comprehensive, and explains why: devices land on it mainly by having AI-related terms in their marketing authorization summaries. Read that mechanism, then remember what I told you two sections ago. Neither FDA document for K253281 uses AI, machine learning, large language model or generative.
When I pulled the list on August 16, 2026 — page content current as of June 16, 2026 — it carried roughly 1,520 entries, and searching it for "UpDoc" or "K253281" returned nothing. The device its maker calls the first cleared patient-facing language model is absent from the government's register of AI-enabled devices, because its paperwork never says AI. The hole is exactly the shape of the mechanism.
That same page says FDA "will explore methods" to identify and tag devices that incorporate foundation models and large language models. Explore. Future tense, still.
The strangest fact follows. Months after the clearance, the authoritative sources were still reporting it had not happened. In March 2026, STAT wrote that the agency had yet to authorize a device that relies on generative AI while covering a breakthrough designation — a promise of future review, not a clearance — for a recovery chatbot. In June 2026 the congressional brief hedged: the agency does not appear to have authorized any generative AI-enabled device. Nobody was lying. They were reading the register, and the register did not know.
Now let me argue the other side, because it is stronger than you think
I have written a thousand words leaning one way. Here is the case against, from people who read the same file.
The deflationary read: this is a narrow, conventional device and the rest of us are projecting — a deterministic, provider-parameterized insulin-dose calculator with a conversational front end, in which the language model does the talking, a separate service does the arithmetic, and the change-control plan requires that modifications maintain deterministic insulin dosing logic without altering core clinical decision-making. The model does not get the keys. And the indications say the device does not interpret or diagnose symptoms at all.
The industry's answer is blunter: in February 2024, AdvaMed argued that AI devices run the same pathways as every other device, and that most were cleared with "locked" algorithms that cannot be modified without further review. Regulate the risk, not the buzzword. There is a mirror-image risk, too: the Paragon Health Institute argues that injury comes not only from faulty AI products but from superior ones not securing market access in a timely fashion. If you have ever titrated insulin by phone tag with a clinic that returns calls in three days, that is not theoretical.
I take all of that seriously. And I still think the pathway is the story, because STAT's question is answered in this file by a promise rather than an experiment. It may well be kept. But it is not idle paperwork: researchers have shown that ordinary chatbots, given ordinary prompts, readily produce device-like clinical decision support — output that looks exactly like the regulated thing, from systems nobody regulates. The line between the model talks and the model decides is not a wall. It is a policy, held in place by a document.
Britain looked at the same problem and decided it did not know the answer yet
Now the comparison — and I want to be scrupulous, because Britain has not solved this either.
The United Kingdom has no equivalence route. What it has instead is an admission. In May 2024 the Medicines and Healthcare products Regulatory Agency launched AI Airlock, a regulatory sandbox for AI as a Medical Device — the agency's first, run with UK Approved Bodies, the NHS and other regulators; the pilot closed in April 2025 and phase two completed in May 2026. A sandbox is not a gate. It is a regulator saying in public that it does not yet know how to do this job.
The pilot report, published in October 2025, is unusually plain about that. Thirty-one percent of applications used generative AI, and the exercise identified regulatory gaps that currently challenge the safe and effective deployment of AI devices — risk management, validating text-based data from large language models, non-determinism, explainability, post-market surveillance. A list of things a regulator could not yet do, written by the regulator.
Phase two took fifty-one applications and produced the sharpest sentence here. For generative AI and language-model products, significant shifts in device behavior can occur through model behavior alone — and without active guardrails, devices may begin to exhibit functions beyond their intended purpose. MHRA's own blog put it plainly in June 2026: "device behaviour can shift significantly without explicit design changes." Set that beside the American file, where the equivalence question was about a keypad. Nor is it hypothetical there: the phase-two cohort includes an NHS England tool using large language models to summarize a patient's hospital stay, in the sandbox so it can be argued over before it is loose.
Then the part that reads, from Washington, like a diplomatic incident conducted in footnotes. Britain is building an international recognition framework so devices approved elsewhere reach its market faster — and its statement of policy intent excludes software as a medical device approved via a route which relies on equivalence to a predicate, naming the US 510(k) in as many words. Fence that properly: policy intent, not law. The draft statutory instrument was published on May 8, 2026, and one consultancy predicts it will not be implemented before early 2027. Nobody in Britain is refusing anything today.
Nor is Britain confident. It stood up a National Commission on the Regulation of AI in Healthcare in September 2025 under Professor Alastair Denniston, and before writing any rule asked the public and the sector first — 760 responses, published in June 2026. The Commission found the framework not well suited to iterative and adaptive AI systems, with 61 percent of respondents calling clinical evidence requirements insufficient and 65 percent pointing to post-market surveillance.
So the honest contrast is posture, not outcome. Britain has authorized no patient-facing language-model dosing device either, and its own answer is still being written. One regulator ran an equivalence test and finished in eighty-five days. The other ran a sandbox and published what it still could not answer. Neither has the answer; only one has said so out loud.
Just imagine the chain, ten years on
None of what follows requires anybody to behave badly.
A cleared device becomes a legal reference point for the next one. So picture 2029: a company builds a conversational tool for insulin and blood pressure, claims equivalence to UpDoc — same intended use, one more parameter, arithmetic still fixed — and clears. In 2031 someone claims equivalence to that one, adding a module that asks about symptoms in a way that shades, gently, toward interpreting them. Each hop is defensible. Each hop is small. Nobody ever runs the trial, because no single hop raises a different question of safety and effectiveness — the statutory phrase, measured against the last device, never the first.
Meanwhile, cleared devices may move without asking again. UpDoc's clearance came with five bounded categories of change the company can ship under its authorized plan, and FDA's guidance on predetermined change control plans exists so such changes need no new submission. I think that mechanism is smart — it is how you regulate something that has to be retrained. But set it against the British finding that a language model can drift outside its stated purpose without anyone changing a line of code, and the gap shows: the plan governs what the company changes, not what the model does.
Then add scale. The tenth device is cleared by reference to the ninth, and at the bottom of the stack sits a keypad from 2019. That is not a failure mode. That is the mechanism working as designed.
What the people who study this actually say
The striking thing is how little ideological daylight there is.
From the center, the Bipartisan Policy Center notes that FDA's postmarket oversight is more limited than its premarket review — thinnest of all for 510(k) devices, where ongoing monitoring obligations are rarer than for the de novo and premarket-approval routes. The lightest door has the lightest follow-up.
From the hospitals that deploy these tools, the American Hospital Association told FDA in December 2025 that the existing reporting variables do not include details on AI-specific risks, and that bias, hallucinations and model drift demonstrate the need for measurement after deployment — from an association whose same letter backed flexibility for market-based innovation. Not a moratorium. A thermometer.
From the academy, the numbers cut both ways. Of 903 AI products, clinical performance studies were available for a little more than half at the time of clearance — 55.9 percent — with 43, or 4.8 percent, later recalled, on average 1.2 years after authorization. A 2026 audit of the public summaries found nearly a quarter providing no evidence for any trustworthiness principle.
And an expert who called this a year early: Sara Gerke, who studies health-AI law at the University of Illinois, argued in July 2025 that FDA must prioritize labeling standards for AI-powered medical devices, because a generative model device was almost certain to arrive — and when it did, clinicians and patients would need to be told the output was AI-generated. Five months later, one arrived whose FDA paperwork never uses the word.
What does this mean for you?
Learn the difference between "cleared" and "approved." Approved means FDA reviewed evidence that the device is safe and effective. Cleared means it resembles something already sold. Most of what you own is cleared.
Look up the number. Every cleared device has a 510(k) number, and FDA's database is public and free. Type it in for the decision date, the regulation it was filed under, and whether a change-control plan was authorized — meaning parts of it may change without new review.
If an app ever gives you a dose, ask two questions. Who set the parameters, and what happens if I type something strange? For this device the answers are good. For the next one, ask anyway.
Do not treat FDA's AI list as a complete inventory. The agency says so itself. A missing product tells you about its paperwork, not its technology.
Tell your prescriber when the software surprises you — not the app store, the human who configured it. Post-market signals are the weakest part of this system, and they depend on somebody speaking up.
The lesson, as I see it
I avoided the question for twenty years because I suspected the answer would be less comforting than the phrase, and I was right. But the answer is not that the system is a fraud, and not that this device is dangerous. There is no evidence it is.
It is narrower and harder to shout: we are letting a new kind of thing into the clinic through a test written to compare old kinds of things, and the test is doing exactly what it was built to do. In 1976, in a world of catheters and hip implants, is this like the last one? was a useful question. Ask it of a system that generates language and it quietly stops meaning anything — because the interesting properties of a language model are not the ones equivalence knows how to see.
My vote? Not a ban, and not a panic. Three unglamorous things. Make devices that contain a language model say so on the label — what Gerke asked for, and what FDA has so far only promised to explore. Fix the register, so a device cannot vanish from the national AI inventory by declining to use the word. And build the post-market monitoring that everyone from the hospitals to the British regulator says is missing, because for a thing that can drift while nobody is touching it, the day of clearance is the least informative day in its life.
I still keep the blood pressure cuff. I have just stopped hearing the click.
If you know someone who manages a chronic condition through an app — or a clinician who has been handed one to configure — pass this along. The question worth asking is not "is it FDA cleared," it's "cleared as equivalent to what?" Keeping that question out loud is the whole job of the HAIA Foundation, and it's what turns up in the Substack every week.




