The first time I bought medicine in a country whose alphabet I could not read, I did what everybody does. I looked for the box that resembled the box at home.
Same rounded shape, same white-and-green, same small cross in the corner. It was the right thing — I checked afterwards with someone who could read the label — and being right is exactly what made the habit stick. Resemblance is fast, free and usually good enough — I have used it on apartments, schools, mechanics and more doctors than I care to admit.
There is a version of that shortcut written into American law, and it is the main door through which medical devices reach you. Substantial equivalence does not ask whether a device works. It asks whether the device is close enough to something already on the market.
On December 23, 2025, that door opened for a piece of software you talk to.
Eight months later, on August 18, 2026, the Food and Drug Administration published a paper asking the public how software like that ought to be regulated.
The sequence is the story: the clearance came first, the framework is still a question, and it closes on October 19 while the cleared device stays where it is.
What the agency published — and what it carefully did not
The document is Considerations for the Regulation of Generative AI-Enabled Medical Devices, out of the FDA's device center. It runs to thirty pages with twenty-six numbered discussion questions.
Start with what it is not. Not a rule, not a proposed rule, not draft guidance. The file it lives in is classified as nonrulemaking — read on September 18, 2026, its fields say type "Nonrulemaking," status "OPEN," comments closing at the end of October 19. A suggestion box with a deadline.
Then the sentence that stopped me. On its first page the paper says it is "not intended to address whether the approaches discussed below are within FDA's existing legal authorities" or whether new legal authorities would be necessary.
Read that twice. The agency floats a way to regulate a whole category of product while stating, up front, that whether it is allowed to sits outside the paper's scope. Not a confession of powerlessness — two serious readings disagree about that, and I will get to both — but a remarkable thing to open with. The candor continues: these devices, with their variable outputs and varying autonomy, present challenges that existing frameworks "were not designed to address."
Michelle Tarver, who runs the device center, called it a transparent process and "a potential model for regulators around the world". The agency's page says you may answer only the questions inside your expertise, and partial responses are fine — which is what you write when you are not yet certain what you are asking. As of September 18, 2026, ninety-eight comments had been posted, most of them from individuals.
The idea: credential the software the way we credential a doctor
To be fair, the central proposal is clever. The regulatory trade press caught it in a line: the FDA is seeking feedback on a competency-based approach — whether a device should have to demonstrate competence much as physicians do before it goes near a patient. The idea is "inspired, at a high level, by how human clinicians are evaluated and credentialed." Nobody examines a doctor on every possible patient; you test the knowledge, then supervise the work.
Risk is sorted on two axes. One is activity: "the risk of a software function depends in part on how independently it directs or takes action." The other is severity — and the agency's illustration of the severe end is a function "directing a patient to increase or decrease their basal insulin dose." A footnote says the examples are illustrative only. Noted. I still cannot unsee it: the hypothetical for the dangerous end describes, almost feature for feature, a kind of product the agency had already cleared.
The exam has rungs. Device benchmarking is high-throughput non-clinical testing — clinical knowledge, safety behavior, communication, and whether the device notices a request has exceeded its scope and refuses under adversarial prompting. Then clinical confirmation, a ladder that climbs through increasingly clinical scenarios ending, potentially, with real patients — including shadow deployment, where the device runs live while its outputs stay hidden from everyone.
Then the trade at the center of the document. Question 18 asks whether it is appropriate to "accept greater premarket uncertainty" about a device's benefit-risk profile "through greater reliance on postmarket monitoring," and asks the public when that would be justified. The last of the twenty-six questions is about devices that act on their own. Like the other twenty-five, it is a question, not an answer.
And one passage I keep going back to: the center is considering whether a patient-facing function becomes any less directive for having a "talk to your doctor" line bolted on. Hold that thought — Britain ran the test.
The device that already walked through the old door
December 23, 2025. Software called UpDoc, cleared under number K253281 for people with type 2 diabetes. It is still cleared and still on the market: substantially equivalent, with a change control plan authorized, the page refreshed on September 14, 2026. Precision matters — there is no recall, no enforcement action, and no public evidence this device has harmed anyone. I looked.
What there is, is a heading in the FDA's own summary. Under "Summary of Clinical Testing," the answer runs to five words: No clinical testing was performed. The difference the filing concedes between this device and the older one it was measured against is "voice and chat interfaces for data entry." A mouth. The company announced in late June that this was the first clearance for medical software using "patient-facing large language models" — its own characterization, reported as such.
Something I checked myself on September 18: the first cleared talking device does not appear on the FDA's public list of AI-enabled medical devices, current as of September 4, 2026 and carrying other devices cleared on that same December day. Not a cover-up — the agency says the list is not a comprehensive resource, built by identifying AI-related terms in authorization paperwork. The tracker reads the words; this filing does not use them.
The pool is large: a market-side tracker counted 1,524 authorization records on that list as of August 20, 2026, 96.2 percent cleared rather than approved (records, not products). And after two years of advisory meetings — including one in November 2025 where Tarver counted more than 1,200 authorized AI-enabled devices and said none yet involved generative AI for mental-health conditions — the center has proposed no new policies specific to generative AI.
One last detail. UpDoc filed a comment in the new docket on September 9, 2026, and it is measured: technologies advance faster than frameworks, closing that gap is a shared responsibility, and what would help is clarity about the boundary between AI that supports a clinician's plan and AI that requires independent clinical judgment — so that line stops being settled, in the company's words, product by product. Which is an accurate description of how it got cleared.
The strongest case against my alarm
I would rather lose an argument than win one unfairly, so: the other side, which is not weak.
One device-law practice read the same paper without reaching for the smelling salts: the design, development and premarket suggestions "appear to fit within existing processes," while the postmarket control ideas "may require new authorities." Those hedges are the author's. The same analysis adds that in its experience the center has been "reluctant to adopt this trade-off" whatever a discussion paper floats.
This is also, straightforwardly, how the FDA works. Another firm notes that the 2019 discussion paper on adaptive algorithms was followed by an action plan and, years later, by guidance — and states the price plainly: stakeholders "may not see formal FDA policy on genAI for several months or years." Whether the framework would take an act of Congress is, per one advisory, simply unanswered. And the broader AI device guidance has been stuck in draft since January 2025, parked on the center's B-list, per an assessment written four months before this paper. Some of what looks like reluctance is a queue.
So: not a scandal. But the trade at the heart of the paper — less certainty at the gate, more watching afterwards — is only good if the watching works. Researchers who audited the system meant to catch problems afterwards — adverse events for roughly 950 AI/ML devices authorized between 2010 and 2023 — concluded it "is insufficient for properly assessing the safety and effectiveness of AI/ML devices." That machinery was built for stents and hip implants; software was bolted on later.
Britain asked the question first and cleared nothing
In May 2024, more than two years before the American paper, the UK's Medicines and Healthcare products Regulatory Agency launched a regulatory sandbox for AI as a medical device called AI Airlock. Manufacturers work out how to generate evidence under the regulator's supervision, in a simulated setting, and selection confers nothing: "being selected for AI Airlock does not constitute a regulatory approval." The Airlock is not a route to market, the agency's blog says. The learning happens without touching patient care.
The pilot report has no American twin. Forty applications, five taken forward, four completed — and 31 percent of applicants were already building generative AI, in 2024. The gaps it found are published by name: validating text output from language models, non-determinism, explainability, post-market surveillance. The cohort is public and specific — a Philips generative-AI function that drafts the impression at the end of a radiology report; an agent that, when there is not enough information, "will not output a response." A firm sharing the name of one of those four, Newton's Tree, later filed in the American docket to argue that benchmark results alone cannot give sufficient assurance of safety and effectiveness.
Phase 2, which completed in May 2026 with seven innovators across three unresolved regulatory challenges — one of them NHS England's own hospital-stay summarizer — produced the finding the American docket never had. It measured out-of-scope behavior: in one company's case study, without active guardrails, "out-of-scope performance was detected in approximately 39% of real-world notes, falling to 20% with guardrails active." One case study inside a sandbox, not a population estimate. The same report settles the disclaimer question the FDA is still considering: disclaimers "function as passive controls" that "do not prevent generative AI systems from producing content that constitute functionality outside of the stated intended purpose." And the British formulation on the central trade turns on one word — "post-market monitoring complements pre-market evidence." Complements. Not replaces. A third Airlock phase is designed but not yet open.
Then, on September 10, 2026 — eight days before I sat down to write this — an independent commission published its recommendations on regulating AI in UK healthcare, drawn from more than 12,000 people over a year: forty-four recommendations, with the evidence behind them. The headline is staged authorization, explained as "L-plates" for learner drivers — deployment inside a tightly controlled scope, widened as evidence arrives, instead of "a single, point-in-time approval." Which is exactly what a 510(k) clearance is.
Britain is not the hero here, though. The Airlock is not a route to market, so nothing reached patients faster, and Britain has authorized no patient-facing language model that adjusts anybody's insulin either. The Commission's flagship idea is a version of the same bargain the FDA floats in Question 18 — and it is a recommendation, not law; a formal government response is still to come. The difference is not the answer. It is the order, and the receipts.
Now run the clock forward
Say the competency idea survives the docket. Picture 2029.
Your pharmacy app has passed its boards. In a settings screen you will never open sits the record of what it was benchmarked on and when it last sat the exam. It shipped with a change control plan, so it updates inside an agreed envelope without new review: the thing instructing you in March is not the thing evaluated in January.
Now the rung above — an agentic version that does not advise but acts, refilling the prescription, flagging the lab value, moving the dose within a range a clinician set months ago. It earned that independence gradually, the way a resident does — and under shadow-deployment logic an earlier build already ran silently on your chart for a year, because its outputs did not affect your care.
Here is the part that keeps me up. The edifice rests on the monitoring half — the half researchers have just called insufficient. A doctor who makes a catastrophic error carries it through a career. A device that produces a harmful output, as a physician writing about the proposal put it, gets re-benchmarked and shipped again. The FDA prints that objection in its own footnotes, citing researchers who argue that analogous mechanisms are needed for flawed devices.
Borrow the structure of credentialing and you get the exam. The part that changes behavior is the part that does not transfer.
What the people who study this are arguing about
The split is not left against right. It is sequence against speed, and it cuts across the usual lines.
The clearance-versus-approval count above comes from a market-oriented health policy institute. An independent, nonpartisan patient-safety organization put AI chatbots at the top of its 2026 list of health technology hazards, and added the detail that matters most: the ones doing the damage are "not regulated as medical devices nor validated for healthcare purposes," yet are used by clinicians and patients anyway.
Brigid DeCoursey Bondoc, a device lawyer at Morrison Foerster, has named the snag in the metaphor: the FDA could regulate the product as a device while a state says the same thing is the practice of medicine — at which point the observation that a device, by nature, is not a person stops being philosophy and becomes a jurisdiction fight.
A cross-party British think tank that welcomed the UK recommendations still named the gap: a fixed algorithm reading a scan has bounded and predictable failure modes; a large language model powering a clinical assistant, by design, does not.
And the most practical objection came from a doctor, not a lawyer: if we trade premarket certainty for postmarket surveillance, somebody has to say who will do it and who will pay for it. As of the September 18 read of the docket, no major trade association or advocacy group had filed an answer.
What does this mean for you?
The comment window is open until October 19, 2026 — and it is open to you. The agency asked for feedback from "consumers" and "the public," and says you may answer only the questions you know something about. If you manage a condition through an app, you know something about Question 18.
Ask what a health app was cleared as, not merely whether it was cleared. Every cleared device has a 510(k) number and a public database entry: decision date, the predicate it was compared against, and whether a change control plan was authorized — meaning the software may change without new review.
Treat "talk to your doctor" as decoration, not protection. The British sandbox tested disclaimers and found them passive controls that do not stop a model wandering out of scope.
Report the strange output to a human, not an app store. Post-market surveillance is the weakest joint in this system and the thinnest part of the proposed trade. It works only if somebody says something.
If you are a clinician, a developer or a researcher, file before October 19. Individuals filed nearly all of the first ninety-eight comments, and a docket that closes that way is an easy one to set aside.
The lesson, as I see it
I still use the resemblance shortcut in foreign pharmacies. It works because the things being compared really are the same kind of thing — a blister pack is a blister pack in any alphabet.
It fails when the new thing only looks like the old one from outside. A calculator that takes numbers on a keypad and a system that holds a conversation share an intended use on paper and almost nothing that decides whether they are safe. The FDA's paper knows this and says as much — and having written it down, the agency's next move was to ask the public what to do.
My vote? Ask first, and ask in public. Not because Britain solved this; it has not, and its own flagship recommendation makes the same uncomfortable trade. But the MHRA stood up a sandbox because it did not know the answer, ran two cohorts of named products, measured a failure rate, published what it could not close, and only then proposed a framework. That order is available to any regulator — and it is cheaper than the alternative, which is what we have: a product on the market, a framework that is still a question, and sixty-two days to answer it.
The window is still open. Until October 19, that matters more than anything else here.
Cleared in December; consulted in August. If that order bothers you as much as it bothers me, send this to whoever handed you the app — then come argue with me about it at the Substack, where the HAIA Foundation keeps a running tally of decisions that arrive before the rules meant to govern them.





