I lost a skill last year, and I still cannot tell you which day it went.
It was mental arithmetic — the running-total kind you use at a market stall, out loud, while somebody weighs things and talks to you at the same time. I grew up doing that, and I was good at it in the way you are good at anything you have done a thousand times. Then one afternoon I stood at a stall with the phone already in my hand, and the number simply did not arrive. Not slowly. Not approximately.
Here is the part that stayed with me: there was no moment of decision. Nobody took the skill away and I never noticed handing it over. Something I had done since childhood quietly stopped being available, and I only found out because the machine doing it for me happened, for once, not to be in my hand.
So let me be honest about where I stand. If a doctor booked me for a colonoscopy tomorrow and offered me the version where a computer watches the video feed alongside her, I would take it instantly, without asking a single question. Two sets of eyes beat one, and the machine does not get bored at 4:40 on a Friday.
That instinct is what this piece is about — because there is now a number attached to it, and it points the wrong way.
First, the number — and exactly how much weight it can carry
On August 12, 2025, a team working across four Polish endoscopy centers published something they had not set out to find.
The centers were taking part in the ACCEPT trial and had introduced AI polyp-detection tools at the end of 2021, after which colonoscopies were assigned to run with or without AI according to the date of the examination. Between September 8, 2021 and March 9, 2022, 1,443 patients had a colonoscopy with the machine switched off — 795 in the three months before the AI arrived, 648 in the three months after.
The metric here is the adenoma detection rate, or ADR: the share of colonoscopies in which the endoscopist finds at least one adenoma, a precancerous growth that can be removed before it becomes cancer. It is the quality metric of the whole field, because a colonoscopy that misses the polyp is a colonoscopy that reassured you for nothing.
In the unassisted procedures, that rate fell from 28.4 percent to 22.4 percent — 226 of 795 procedures found one before, 145 of 648 after. In the paper's own arithmetic that is an absolute difference of 6.0 percentage points (95 percent confidence interval −10.5 to −1.6; p=0.0089). The Lancet's press office put the same fall the other way around, as a 20 percent relative and 6 percent absolute reduction. Keep those two apart, because the six and the twenty get blurred constantly and are not the same claim.
These were not trainees. They were nineteen experienced endoscopists, every one of whom had already performed more than two thousand colonoscopies — and two thousand is the floor there, not the average. These are the people you would want holding the scope.
The drop held up when the researchers controlled for who the patients were: in a multivariable model, exposure to AI came out as an independent factor associated with lower detection, at an odds ratio of 0.69.
STAT put it in the plainest possible terms: before the AI helper arrived they were finding adenomas in 28 percent of colonoscopies; after three months of using it, their unassisted rate fell significantly, to 22 percent. The researchers called it the first documentation of a potential deskilling effect from clinical AI.
Now the part most coverage skipped. Read what the authors themselves were willing to conclude, verbatim: "Continuous exposure to AI might reduce the ADR of standard non-AI assisted colonoscopy, suggesting a negative effect on endoscopist behaviour." Might. Suggesting. That is a hedge doing real work, written by the people closest to the data — and every headline that turned it into "AI made doctors worse" took a liberty they did not.
Two more pieces of housekeeping. This was not a randomized test of the thing everyone is discussing: the team ran a retrospective, observational study, looking backward at four endoscopy centers. The randomization inside ACCEPT decided which colonoscopies got AI after implementation; it did not create the before-and-after comparison that produced this result. And a correction to the paper was published three months later. I could not get past the paywall to read it, so I will not tell you what it changed — only that it exists.
How they found it matters as much as what they found
This was an accident, and the lead analyst, Marcin Romańczyk, has been open about it: his team found it looking for something else — "We accidentally found that ADR in standard, non-AI colonoscopy decreased from 28% to 22% after endoscopists were exposed to routine AI use."
That cuts both ways. An unplanned finding cannot have been fished for, which is a point in its favor; it also carries less weight than a result a study was designed to test, which is a point against. And it has a consequence that is easy to miss: why it happened is not in the paper, because nobody anticipated the outcome and so nobody collected the data that would explain it. No withdrawal times, no eye tracking, no record of how carefully anyone looked. The mechanism everyone reaches for — that the humans relaxed because the machine had been doing the vigilance — is a plausible story, not a finding.
The authors put the limits on it themselves: the observational design means factors other than AI may have influenced the results, and because the study used experienced endoscopists it may not generalize to everyone else. Romańczyk's own framing was a call for work rather than a verdict — to his knowledge the first study to suggest a negative impact of regular AI use on a health professional's ability to do a patient-relevant task in medicine of any kind.
What the American regulator asked about this — and what it never asked
Here is where I stopped feeling clever about my own instinct.
The first colonoscopy AI cleared in the United States was GI Genius, authorized on April 9, 2021 through a De Novo request that created a new device classification. The FDA's own decision file for it says this, in the uncertainty section, four years before the Polish paper appeared. The agency flagged the over-reliance problem in 2021, in writing: "Another area of uncertainty is the long-term usage of AI software devices, and the potential for overreliance on lesion detection software leading to errors in treatment and diagnosis, rather than the intended use of the device as an aid to clinicians."
They saw it. They named it. So what did they require in response?
The mitigation, in full, was a warning on the label plus a usability assessment — warnings to avoid over-reliance, and to state that the device does not replace clinical decision-making. The usability evidence was a nineteen-question survey, scored one to five.
The pivotal evidence came from six endoscopists in Italy with ADRs between 25 and 40 percent, and the file says plainly it is unknown whether the results would have held with more endoscopists, US-based ones, or ones outside that band. The benefit-risk conclusion nonetheless judged a negative effect on standard colonoscopy unlikely: despite the uncertainty, there was "unlikely to be a negative impact on the standard colonoscopy outcomes."
That sentence has aged into something worth staring at.
Commercially, the device went to market sold as an "ever vigilant second observer", with a claimed 14 percent absolute increase in adenoma detection compared with colonoscopy alone. To be fair: that benefit is the reason these systems exist, and nothing in the Polish study touches it.
But notice what the whole apparatus measures. The trial, the label, the marketing, the survey — every one asks what the endoscopist can do with the device. Not one asks what the endoscopist can still do without it. That question was not in the business case, the clearance, or the purchase order. It showed up four years later, by accident, in Poland.
So that settles it? Not remotely — here is the case against my own headline
The strongest arguments against this finding are good ones, and they were made immediately, on the record, by people who study exactly this.
Start with the one that keeps me honest. Venet Osmani, professor of clinical AI at Queen Mary University of London, pointed out that the number of colonoscopies performed nearly doubled after the AI tool arrived, going from 795 to 1,382 — and that it is possible this sharp increase in workload, rather than the AI itself, led to the lower detection rate. Sit with that. The AI was not the only thing that changed in those units. His mechanism needs no reference to skill at all: a more intense schedule means less time and more fatigue. And he names the general problem with before-and-after studies — a new machine arrives with everything else that changes around it.
Allan Tucker, professor of artificial intelligence at Brunel University of London, names two more at once: some major changes were made to the endoscopy department in the middle of the study, and the authors themselves make clear that randomized crossover trials are needed before anyone makes more robust claims. That second half should govern how the result gets used — the people who found it are asking for the study that would settle it.
From the American side, Paul E. Oberstein of NYU Langone's Perlmutter Cancer Center struck the balance I would defend: a relatively small study in one setting, warranting further validation, but findings that raise important questions about how to measure AI's impact on human skills.
And the context that surprised me most, because it cuts against the assumption that AI colonoscopy was a settled good this paper disturbed. On March 20, 2025 — five months before the Polish paper appeared — the American Gastroenterological Association declined to recommend the technology at all: "In adults undergoing colonoscopy, AGA makes no recommendation on the use of CADe-assisted colonoscopy." Its reason was very low certainty of evidence on the outcomes that actually matter — cancer incidence, cancer mortality, post-colonoscopy cancer. That is not a reaction to the deskilling study; it predates it. It is a professional society saying, before any of this was measured, that the long-run evidence was already thin.
The pattern across specialties — and exactly where it frays
This result did not die as a curiosity, because it landed in a literature already circling the same worry.
A 2026 review that went looking across every specialty — radiology, pathology, endoscopy, clinical decision support — concluded that evidence of clinical deskilling, though scarce, is consistent across specialties, and argued that safeguarding clinical expertise should be a central component of AI safety in medicine. It comes from Pierre Heudel and colleagues at Centre Léon Bérard in Lyon, reviewing work published up to August 2025.
One caution: that review's abstract describes the Polish study as a multicenter randomized trial, while its own full text correctly calls it observational. When a review and its abstract disagree, go to the primary studies.
In mammography, an experiment on twenty-seven radiologists reading fifty mammograms found that the correctness of the AI's suggestion significantly moved assessments at every experience level — the very experienced readers, the ones you would expect to be immovable, dropping from 82.3 percent correct to 45.5 percent when the suggestion was wrong. In pathology, eight pathologists reading a hundred and fifteen laryngeal biopsy slides were observed omitting invasive carcinomas they had already diagnosed correctly on their own, because the model said otherwise — a false prediction flagged with a low confidence score the less experienced readers did not weigh.
Both are alarming. Neither is the same thing as the colonoscopy result, and the difference is the whole argument.
Those two studies measure what a clinician does while a machine tells them something wrong. The mammography experiment did not even use a real product: it used a system the radiologists were told was AI, so the researchers could control the suggestions. That is automation bias — a failure in the moment, with the tool present. The Polish study measures what is left of the human after the tool is taken away.
Europe has already written a training answer without waiting for the argument to resolve. The curriculum Europe's endoscopy society published in December 2025 carries six recommendations that read like a direct response — basic competency in standard endoscopy first, AI literacy, explicit attention to cognitive bias in human-AI interaction, avoiding over-reliance in clinical decision-making, and continuous monitoring of key performance indicators. And in June 2026, Annals of Internal Medicine put the question on its pages under a title that does not hedge — "The Deskilling Effect: Is Artificial Intelligence Eroding Clinical Competence?" I could not get past the publisher's block to read that three-page commentary, so I will not summarize its argument. Its existence, in that journal, in that year, is the data point.
One country wrote the human into the price list
Now the part that changed how I think about the policy question. It comes from Japan.
Japan was early. It approved its first colonoscopy AI in 2018, when EndoBRAIN cleared as a Class III medical device — three years ahead of the FDA's De Novo, and the first approval of an AI-related medical device in the country.
Then came the stretch the sales decks never mention: for years nobody anywhere was paying for it. There was no reimbursement for AI in colonoscopy anywhere in the world in 2023 or before, which is precisely why implementation crawled.
Japan broke that. In February 2024, the country's public health insurance body announced an added payment for the use of a computer-aided detection tool — EndoBRAIN-EYE — in colonoscopy. And it did the arithmetic first: a published analysis found CADe colonoscopy cost-effective up to a device cost of ¥6,000, and concluded on that basis that reimbursement was reasonable. Somebody showed the working before the money moved.
The mechanics live in a document dated March 5, 2024, published by Japan's Ministry of Health, Labour and Welfare as the official overview of that year's fee revision. Attached to the fee for endoscopic colorectal polypectomy is a new line item: a lesion detection support program add-on, worth 60 points, or ¥600. That is how Japan put the AI on its national fee schedule.
And then it did the thing nobody else did. Read the billing condition in the original:
なお、本加算は、内視鏡検査に関する専門の知識及び5年以上の経験を有する医師により実施された場合に算定する。
In plain English: the add-on may only be claimed when the procedure has been performed by a physician with specialist knowledge of endoscopy and five or more years of experience. Japan tied the payment for the machine to the experience of the doctor holding it.
It goes one layer deeper. The same document writes the experienced specialist into the evidence standard itself: the software must be shown to raise diagnostic accuracy in the hands of a physician with specialist knowledge and experience of colonoscopy, compared with not using it. The state is not paying for a machine that makes an average operator adequate. It is paying for one that makes an expert better, in the hands of an expert.
I have read a lot of AI procurement language this year, across a lot of jurisdictions. It is the only case I have found where the money for the tool carries a written condition about the human.
So is Japan the hero of this story? No — and I want to be careful here, because that is exactly the sentence a piece like this wants to write.
What Japan built is a credential checked once at the door. Five years of experience on the day you bill is not a measurement of what you can still do unaided eighteen months later — the precise thing the Polish study says we should worry about, and Japan is not measuring it either. The country still has no clinical guidelines for any of it.
Nor is the field pretending otherwise. The Japan Gastroenterological Endoscopy Society, which issued nine position statements in June 2025 — running from examination quality to medical safety and legal responsibility — reports that more than ten AI-assisted endoscopic devices have regulatory approval in Japan and that adoption has still not been smooth, citing unclear cost-effectiveness, missing guidelines and absent reimbursement. That is a society arguing with itself in public, which is what serious institutions do.
Here is the detail that bothers me most. Japan already built the instrument that could answer this, and nobody has pointed it at the question.
Since January 2015 the country has run the Japan Endoscopy Database, built to see exactly this kind of thing: a national registry whose stated purposes include capturing the actual performance of endoscopic practice in Japan, with the explicit potential to collect data automatically on competency, on the evaluation of residents, and on procedure numbers nationwide. Its first colonoscopy phase alone covered 38,497 procedures in 31,395 patients, and behind it stands a society with 21,460 board-certified endoscopists on its rolls.
A registry that already tracks performance across a nation, in a country that has just started paying for the AI, is the best-placed system on earth to find out whether the Polish result is real. So far as the reachable record shows, it has not been asked.
Japan's regulator has also thought harder than most about the moving-target problem, building IDATEN, a pathway for algorithms that keep learning after approval — though its own Science Board report says the system "has yet to be fully used."
That same report reproduces the guiding principles American, Canadian and British regulators agreed on in October 2021, and principle seven is the sentence I would put on the wall of every procurement office: focus is placed on the performance of the human-AI team.
The team. Not the tool. Three regulators wrote that down in 2021, and then everybody went off and measured the tool.
Two final threads, because they complicate the picture rather than tidy it. Japan's National Cancer Center is running the region's confirmatory trial — a randomized controlled trial of CADe in colorectal cancer screening across 13 facilities in six Asian countries and territories, with adenoma detection rate as the primary endpoint. And Japan helped pay for the study that found the problem: the Japan Society for the Promotion of Science is a named funder of the Polish paper, whose senior author holds a post at Showa University Northern Yokohama Hospital and discloses royalty fees from the Japanese maker of EndoBRAIN. The country that put the AI on its fee schedule co-funded the research that raised the alarm. Whatever that is, it is not a captured system.
Though note what the same circle was saying two years earlier. The position statement these people wrote in 2023, under the World Endoscopy Organization, with six Japan-based authors, framed the promise in the opposite direction: CADe "offers promise to reduce unwanted operator-dependent variability in colonoscopy performance." Nobody was yet asking what it might do to the operators.
One industry already wrote this rule, about a different machine
All of this is treated as a novel problem. It is not: the same question is more than a decade old, and aviation got there first.
On January 4, 2013, the Federal Aviation Administration sent airlines a safety alert that said the quiet part out loud: "Unfortunately, continuous use of those systems does not reinforce a pilot's knowledge and skills in manual flight operations" — and continuous use "could lead to degradation of the pilot's ability to quickly recover the aircraft from an undesired state."
Same sentence, different machine, written twelve years earlier by a regulator that had been reading accident reports.
What makes it useful is the remedy, which was operational rather than rhetorical: the FAA told operators to develop or review policies ensuring there are appropriate opportunities for pilots to actually fly by hand, and to reinforce it in training and proficiency checks. Not a label warning: scheduled practice, plus a check that the practice worked. Four years later the agency wrote it into the training rules, adding manual flight maneuvers to the Part 121 requirements so that proficiency is "developed and maintained."
None of which means autopilots are bad — that is the point. Aviation kept the automation and the safety gains it bought, and separately built a regime that guarantees the human still gets reps. Medicine bought the automation and skipped the second half.
Now run it forward ten years
Here is a question I cannot answer, and I do not believe anyone can yet.
Every randomized trial that established the benefit of AI colonoscopy compared endoscopists using the AI against endoscopists working normally. That second group is the control arm. But if routine exposure changes what an endoscopist does when the AI is off — the whole hypothesis on the table — then in a trial run inside a unit that had already adopted these tools, the control arm may not be measuring standard colonoscopy at all. It may be measuring a human who has already been changed. Nobody has tested that, and the Polish paper does not. It is a question, not a result — and one nobody, so far as I can find, is funded to ask.
Now push the timeline out.
Just imagine a hospital in 2032 where nobody on the endoscopy roster has ever done a hundred consecutive colonoscopies without the box switched on. Not because anyone forbade it — because the box was always there, and the department's throughput targets were rebuilt around the assumption that it always would be. Then imagine the week the vendor contract lapses, or the model is pulled for a safety update, or the system opens a site in a town that cannot afford the license. What is the real detection rate on that floor, that week? Nobody knows, because nobody measures it.
Now imagine the version where it goes right. A procurement contract in 2030 that will not release payment unless the vendor's dashboard reports per-operator unassisted performance every quarter. A board certification requiring AI-off sessions the way an airline requires hand-flown legs. An insurer paying a premium not for the AI, but for the unit that can prove its people still hold the skill without it. A national registry — Japan already has one — answering the question for 21,000 endoscopists instead of nineteen.
And then the version that is not about medicine at all, because you are already living in it. Something similar has been measured in software: in a randomized study, fifty-two software engineers were split by whether they used an AI assistant, and the ones who had used it averaged 50 percent on the follow-up quiz against 67 percent for those who had not.
Those futures are not equally likely. The first is the default, because it is what happens when nobody does anything — and nobody doing anything is the historical base rate.
What the smart people are saying
The gastroenterologist who wrote the commentary published alongside the study is Omer Ahmad of University College London. Earlier work had hinted at behavior change after AI exposure, he wrote, but this study provides the first real-world clinical evidence for the phenomenon of deskilling — and while AI continues to offer great promise, "we must also safeguard against the quiet erosion of fundamental skills required for high-quality endoscopy."
The study's senior author is blunter about how little anyone knows. Yuichi Mori, who teaches on exactly this at the University of Oslo while holding a visiting post in Yokohama, put it flatly to reporters covering the research: "There is no established solution against deskilling right now." It should, he said, be a very hot research topic for the next decade.
Away from medicine, the same worry is arriving from three different political directions, which usually means something real sits underneath it.
From the center-left, Molly Kinder at Brookings argues that the pathway by which juniors become experts is the thing AI is eating: hiring junior workers to do routine tasks while they build expertise will not survive when the machine handles those tasks, and if employers, philanthropy or government do not step in, the cost falls on young people themselves. Her proposed model, pointedly, is the medical residency.
From the right, Adam Thierer at R Street would rather we not overcorrect, and he is right to press it: safeguards are needed, but lawmakers must weigh the unintended effects of regulation, and ask whether the policies under consideration would hold back innovations and treatments that could save lives. A rule written in a panic that keeps a detection aid out of an under-resourced clinic is not a win.
From the middle, the Bipartisan Policy Center has been counting the payment codes: a major barrier to clinical AI adoption in America is the lack of a clear Medicare benefit category, and as of January 2026 there are 26 CPT codes for clinical AI, only three of them Category I. America has not yet built the payment lever Japan used. When it does, the question is whether anything gets attached to it.
What does this mean for you?
You are not going to regulate an endoscopy suite this afternoon. But parts of this are genuinely yours.
If you are booked for a colonoscopy, ask about the unit's adenoma detection rate. Not whether they have AI — ask what their ADR is. It is a real, tracked number, it is the metric the field runs on, and a unit that cannot tell you has told you something. My own answer on the AI question is still yes: take the second observer if it is offered.
Do not read this as a reason to refuse the technology. Read it as a reason to ask what else gets measured. The failure here is not that hospitals bought AI; it is that nobody wrote down a baseline of what the humans could do first.
If you use AI in your own work — and you do — schedule the unassisted rep. Not out of nostalgia: the Polish endoscopists were not careless people, and it took about three months. Pick the skill that would hurt most to lose and do it cold, on a real task, on a schedule you keep. Aviation calls that a hand-flown leg; it is not a mood, it is a calendar entry.
Notice that this is invisible from the inside. Nothing in this story depended on anyone noticing, and I certainly did not notice losing the arithmetic. If your only detector is your own sense of competence, you do not have a detector.
If you manage people or buy software, measure the human before the tool arrives. A baseline costs one quarter of ordinary data collection and beats every vendor benchmark you will be shown. Afterward you can never reconstruct it.
If you buy or sell clinical AI, put the operator in the contract. Japan attached a condition about the physician to the payment for the machine; nothing stops a hospital attaching one to a purchase order — quarterly reporting of unassisted performance, or protected time with the tool off.
If you write policy, you have two working templates, neither theoretical. An experience condition tied to a reimbursement code, and a regulator telling operators to schedule manual practice and then checking they did. Both exist, and neither required a new agency.
The lesson, as I see it
We have spent five years asking one question about clinical AI — does the doctor plus the machine beat the doctor alone? It is a good question, and the answer looks like yes.
But it is half a question, and we built an entire apparatus of trials, clearances, marketing claims and purchase orders on top of the half. The other half is what the doctor can still do when the machine is not there: on the night shift, in the clinic that lost the license, in the town the contract never reached, in the ten seconds after the software throws an error. Nobody costed that. Nobody measured it. It surfaced by accident, in a data set collected for another purpose — and the honest state of the evidence is one hedged observational finding, two named confounders, a professional society that never recommended the technology in the first place, and the authors themselves asking for the trial that would settle it.
That is not a reason to panic. It is a reason to run the study — and meanwhile to write down the baseline while we still have people who have one.
My vote? Keep the machine. Keep the second observer, and keep the adenomas it finds. And then spend a fraction of what we spent buying it on finding out what it is doing to the person holding the scope — because the skill you never practice is the one you will need on the day the system is down, and by then it is too late to go looking.
Nobody bills for the thing you quietly stopped being able to do. Watching for that is the whole reason the HAIA Foundation exists, and the weekly version of it lives over at the Substack. Send this to whoever on your team has stopped double-checking.






