I read a forty-page report in about eleven minutes last spring, and I have rarely been so certain I understood something.
I had not read it, obviously. A chatbot walked me through it — my questions answered in the order I happened to think of them, every answer in clean prose pitched at exactly the difficulty I wanted. It felt like comprehension. Better than comprehension usually feels, because comprehension normally involves a stretch of being confused and this did not.
Three days later that report came up in a conversation where it was the entire point, and I could produce its shape and none of its contents. Not a figure. Not the name of the body that wrote it. Not even the joint in the argument where I had, at the time, felt a flicker of doubt.
Here is what stayed with me. Handed a feedback form as I closed the tab, I would have rated that session highly, and meant it. The feeling of learning and the fact of it had come apart, and nothing in the feeling told me so. That gap is not a personal failing — it has now been measured, in teenagers, with a control group.
First, what the trial actually did
A team split between Cambridge University Press & Assessment and Microsoft Research ran a pre-registered randomized controlled experiment in seven secondary schools in England: 405 Year 10 students, aged fourteen and fifteen, studying two history passages and sitting comprehension and retention tests three days later, under three conditions — a large language model alone, note-taking alone, and the model with note-taking alongside it.
One number before the others. Those 405 were the students recruited; after the study's exclusion criteria, 344 were retained for analysis. Every effect below belongs to the 344, not the headline 405.
The paper went online in November 2025 in Computers & Education. Microsoft's own review, written seven weeks earlier from the preprint, summarizes it without flinching: traditional note-taking outperformed the combined condition across all three measures, with the LLM-only condition scoring the lowest. Cambridge University Press & Assessment announced its own study on December 4, 2025 under a headline that gives away the ending: "Note taking more effective than AI for learning."
Now the detail the summaries flatten. Against the chatbot alone, note-taking won on all three outcomes: literal retention, comprehension, and free recall — reproducing the passage cold, with nothing to prompt you. The combination, chatbot plus notes, beat the chatbot on the first two, but on free recall the paper reports no significant difference between LLM + Notes compared to LLM. Adding a notepad did not restore the ability to summon the material unaided.
And the size of it, because this is where a piece like mine is tempted to cheat. Table 3 puts note-taking's advantage at Cohen's d of 0.44 on literal retention, 0.38 on comprehension and 0.21 on free recall — where 0.2 is conventionally small and 0.5 moderate. A real difference in a classroom of thirty. Not a collapse.
Here is the part that is not obvious. Students rated the chatbot-only condition as less difficult than both alternatives, and — against note-taking, though not against notes-plus-chatbot — reported investing less effort in it and perceived it as more helpful for understanding the text. Asked which they would choose, Group 1 — the students who tried the chatbot against note-taking — split like this in the paper's Table 4: 42.0 percent preferred the model, 26.9 percent preferred notes, 22.6 percent had no preference, 8.5 percent were not sure. The authors call that "most." I would call it a plurality — but the direction is not in doubt. Group 2 cuts the other way: offered the chatbot against chatbot-plus-notes, 50.5 percent took the combination and 16.2 percent the chatbot alone. Against note-taking alone, then, the largest group of these teenagers preferred, and rated more helpful, the condition that taught them least.
Before anyone reaches for the obvious explanation — they learned less because they did less — the authors close that door. Students did spend less time on task with the model, but the gap was roughly one to one and a half minutes, and time on task did not significantly predict performance on any outcome (p = .237, .348 and .559 across the three). The effect, they conclude, is not attributable to engagement duration. Whatever happened to those students happened inside the minutes they spent, not because of how few there were.
Two things from inside the logs. The team analyzed all 4,929 prompts the students typed, and only six prompts questioned the LLM's trustworthiness. Six. About 10 percent were off-topic — "what is the meaning to life," "Tell me about Harry Potter" — teenagers being teenagers. Six is a different kind of number. And in the combined condition, 25.63 percent of students lifted more than 70 percent of their notes' three-word sequences straight from the machine's output. Given a notepad and a chatbot side by side, a quarter of them were copying heavily. And none of this comes from AI skeptics: the paper's declaration of interests puts some authors at a company that invests in generative AI, and the rest at an assessment organization that uses AI in its own products.
So chatbots are bad for teenagers? Slow down — here is the case against my own headline
The strongest objections come from the authors themselves. Start with the machine. The model was OpenAI's GPT-3.5 turbo, with each student capped at twenty prompts, and the paper concedes it is not as advanced as some of the more recent models. That is the objection I would press hardest from the other side. The second is scope: the test measured mastery of two set passages, and the authors say plainly it did not capture broader learning. They float the possibility that the outcomes are complementary — notes for mastering assigned material, the model for curiosity beyond it.
Microsoft's Jake Hofman framed it that way publicly, telling the Cambridge student paper that rather than treating note-taking and generative AI as competing alternatives, we should view them as complementary. Cambridge's Martina Kuvalja was blunter: "No pain, no gain – if you make your own notes, you're probably going to remember what you've learned better than if someone – or something – summarises it for you."
Then there is the awkward matter of the trials that found the opposite. At Harvard, a randomized experiment with 194 undergraduates in a large physics course found that students learn significantly more in less time when using the AI tutor than in an in-class active learning session. That is the English result turned inside out — and a different intervention: a tutor purpose-built on pedagogical principles, used by consenting undergraduates.
Its authors also decline to over-claim, writing that we do not presume that structured AI tutoring will always outperform in-class active learning in all contexts.
In Edo State, Nigeria, a World Bank randomized controlled trial of a six-week after-school program measured gains of approximately 0.3 standard deviation, which its authors equate to one and a half to two years of typical progress. Their caveat sits in the same post: the result reflects the full package — structured prompts, curriculum alignment, teacher support, peer learning and the AI — and they cannot isolate the effect of the chatbot alone.
Closest to home, an exploratory preprint from Google's LearnLM team and Eedi, across five UK secondary schools, reports that students on a purpose-built tutor performed at least as well as students chatting with human tutors. "At least as well as" is equivalence, not victory, and a preprint is not peer-reviewed — but it points the other way.
Which leaves the honest position, the one Microsoft's own reviewers wrote down last October: there is a lack of consensus across the research literature on whether generative AI improves learning performance at all. So the defensible claim is narrower than my headline: a general-purpose chatbot handed to a fourteen-year-old is not the same intervention as a tutor designed to make them work.
Which is precisely the decision being made in staffrooms this week
England's position is devolved almost to the individual school: the Department for Education says it's up to schools and colleges to decide whether students can use AI, with final responsibility for anything the software produces resting on the teacher and their school. The department has moved, though quietly. On January 19, 2026 — about eight weeks after this trial went online — it updated its safety expectations for generative AI products sold into education to include new standards on cognitive development. That is a regulator writing down that a learning product can harm the learning.
Meanwhile the students have already decided. The National Literacy Trust's 2025 survey of more than 60,000 young people aged thirteen to eighteen found two in three using generative AI, and 1 in 4 (25.1%) admitted to 'just copying' AI outputs for homework. Hold that against the trial's 25.63 percent copying rate: a controlled experiment and a national survey landed on nearly the same quarter.
Estonia bet the other way — and its own scientists are asking the hard question
In February 2025, Estonia's Ministry of Education and Research announced AI Leap — TI-Hüpe — committing to 20,000 high school students in grades 10-11 and their 3,000 teachers from the start of the next school year. The European Commission's Eurydice network files it as a national program to integrate artificial intelligence (AI) into education, and OpenAI told Euronews it was proud to work with Estonia to bring ChatGPT Edu to students and teachers.
The easy version is: Estonia gave every teenager an AI tutor in September 2025. The record is more specific. The AI Leap Foundation's own announcement notes that teachers were granted premium ChatGPT and Google Gemini access in August 2025, and that only in late January 2026 did it begin creating accounts for approximately 20,000 10th–11th grade students in 154 upper secondary schools — two year groups, not an entire school system. And what they got is not a general-purpose chatbot: the program describes a learning app that supports thinking skills better than a regular large language model precisely because it does not simply provide answers, evaluated against 19 behavioral criteria written by Estonian educational psychologists.
Now the part that made me sit up. The scientific lead is Jaan Aru, a University of Tartu neuroscientist who works at the intersection of cognition and artificial intelligence. His team — the people building the national AI tutor — published this: picture a student running laps to build endurance, and current AI models as an overly supportive friend shouting from the sidelines, "Take five, I'll run the next lap for you!" The lap gets finished in record time, and the kid's fitness won't improve one jot. The same piece records the ministry commissioning the University of Tartu to study whether the app actually helps students learn.
Read that sequence again. The country that made AI in schools a national bet is now paying to find out whether the bet works — and says so out loud: we're not sure whether our AI really is conducive to learning, so we need to investigate it, the program stated in February 2026, adding that its goal is not more AI use but more meaningful use. Estonia is not hedging its bet; it is auditing it in public.
South Korea offers a third posture: it went early, then reversed, its National Assembly reclassifying AI digital textbooks in December 2024 so that the decision on whether to use AI textbooks as supplementary material will be given to school principals rather than the education minister. Three countries, three answers — and the one that went furthest is the one asking the hardest question about itself.
Now run the next ten years
What worries me is not a robot teacher. It is that the preference finding is a market signal. Forty-two percent choosing the tool that taught them least is, to a product team, a satisfaction score — and satisfaction scores get optimized. Nothing sinister has to happen. You ship the version students rate highly, iterate toward the version they rate more highly, and arrive by good-faith increments at a product engineered to produce exactly the feeling I had with that report.
So picture a school in 2032 where every text arrives pre-digested, every question is answered before it becomes uncomfortable, and the dashboards are magnificent. Attendance up, complaints down, homework done. Nobody notices, because what went missing is invisible from the inside — a fourteen-year-old cannot tell you what they failed to encode, and neither could I.
Then picture the other version, equally available and considerably less likely: products built with friction on purpose. Microsoft's own review lists the moves, including limit copy-paste functionality, because the transaction cost is what encourages memory formation, and support for metacognitive calibration, because the software can distort a student's sense of how much they have learned. Schools could grade what a student reproduces three days later rather than what they produced in the lesson. That future is a procurement question, and it is available now.
What the smart people are saying
People who agree on almost nothing else are landing in the same place.
From the center-left, a January 2026 Brookings report says it about as plainly as a think tank ever does: the risks of utilizing generative AI in children's education overshadow its benefits. Note the wording, which is temporal — "at this point in its trajectory" — and note that the same report allows that well-designed tools deliver real benefits if deployed as part of an overall, pedagogically sound approach. That is not "AI is bad for children"; it is a verdict on the state of the deployment.
One of its authors, Rebecca Winthrop, put the mechanism to NPR: when kids use generative AI that tells them the answer, they are not thinking for themselves — not learning to parse truth from fiction, not learning what makes a good argument, because they are not engaging the material. An education outlet covering the same report distilled the distinction: for most young people, AI is not a 'cognitive partner' but a surrogate.
From the free-market right, the American Enterprise Institute is not anti-adoption in the slightest — it argues schools must reinvent educational practice, not merely coordinate. It lands on the same risk anyway, warning that if schools do not evolve we risk a generation fluent in prompting chatbots but lacking the deeper reasoning that defines real learning. Its analysts are equally deflationary about Washington: the Department of Education's AI priority steers grant money, and this is not a regulation, a mandate, or a new program. Nobody is coming to settle this for your school.
And from the nonpartisan middle, Bellwether notes that some studies suggest AI tools may reinforce shortcuts instead of supporting deep learning — while insisting it does not have to be that way. Left, right and center, the worry is the one the Estonian researchers put in their stadium metaphor, and that convergence is worth more than any single trial.
What does this mean for you?
Most of this gets decided at kitchen tables and in individual classrooms, not in a national curriculum.
If you have a teenager, ask the school two questions, not one. Not "do you allow AI" — that one has an easy answer. Ask what the school's tools do when a student asks for the answer, and how anyone will know whether it worked.
Judge any tool by the three-day test, not by the session. Come back three days later and try to reproduce the thing cold. That is the measurement this trial made — and the students who worked least hard found it easiest and rated it most helpful. So did I.
Use it around the material, not instead of it. The authors' generous reading is that these tools help with initial understanding and with curiosity beyond the set text. Nothing here argues against asking a model to explain a hard paragraph; it argues against letting it do the encoding for you.
If you buy the software or teach with it, ask about friction — and take a baseline. Does the product make the user do the work, and does it limit copy-paste? What can your students do unaided, this term? Nobody can reconstruct that number later.
The lesson, as I see it
Every argument about AI in schools is conducted as if the danger were cheating. Cheating is a discipline problem, and schools have handled those for centuries.
This trial points at something harder. The students were not cheating. They did the work as set, in good faith — and came out with less of the material and more confidence about it. The mechanism was not laziness and it was not the clock; the authors checked. The difficulty had been removed, and the difficulty was where the learning lived. That is a hard sell in a world where we build for ease, measure satisfaction, and ship what people prefer.
My vote? Keep the chatbot — it is genuinely useful for getting into a hard text, and the counter-evidence from Harvard, Nigeria and those five UK schools is real. But keep the notepad, keep the confusion, and keep one honest measurement of what your kids, or your team, or you, can still do with the machine switched off. Estonia is spending public money to find out whether its bet paid off. That is the posture worth copying. Not the confidence. The checking.
Somebody you know is about to write a school's AI policy, or sit an exam under one. Send this to them — it travels far better secondhand than as a lecture. That forwarding is how the HAIA Foundation reaches the people making these decisions, and the rest of the week's work is over at the Substack.





