Khanmigo AI tutoring study finds small maths gains in two-year trial
A two-year NBER trial in 18 Tennessee middle schools found Khan Academy’s Khanmigo raised maths scores by 1.26 percentile ranks per term.
Photo by Katerina Holmes on Pexels
Giving middle school students access to Khan Academy’s AI tutor Khanmigo raised maths achievement by 1.26 national percentile ranks per term, according to a two-year randomised trial circulated by the National Bureau of Economic Research on 17 August 2026. The Khanmigo AI tutoring study, registered as NBER working paper w35620 and written by Philip Oreopoulos of the University of Toronto with Nina Low of Charles River Associates, ran across 18 middle schools in Hamilton County Schools, Tennessee, over the 2024-25 and 2025-26 school years. The authors report that the gain “resembles those from Khan Academy practice without AI assistance”, and identify student engagement rather than model capability as the binding constraint.
The paper was one of three tutoring trials Oreopoulos and co-authors circulated through NBER on the same day. Together they cover a school-based AI tutor, a one-week classroom experiment testing how AI should be structured, and a two-year home tutoring programme in Toronto. None of the three has been peer reviewed.
What the Khanmigo AI tutoring study measured
The Khanmigo trial was a cluster randomised experiment. Within each school the researchers randomly assigned whole grades to the treated or control condition, producing 53 grade-within-school clusters, 28 treated and 25 control. Treated students used Khan Academy with Khanmigo, configured to coach rather than supply answers, during the daily Response to Intervention remedial maths block of 25 to 40 minutes. Control students received the usual remedial instruction, described in the paper as “a mix of teacher-led practice and other math software without AI tutoring”. The analysis covers 6,902 student-term observations across five terms.
The pre-registered pooled effect was 1.26 national percentile ranks per term, with a standard error of 0.60, equal to 0.040 population standard deviations. Effects were larger in the second year, at 0.084 population standard deviations with a standard error of 0.041, against 0.020 in the first year and not distinguishable from zero. The implied effect of a full year of active participation in year two reaches 0.142 standard deviations. Programme cost is put at roughly $15 per student per year.
Students opened the tutor but rarely talked to it
The usage data explains the modest result. Across 563 activated second-year treated students at 16 schools, the researchers logged 25,286 platform-active student-days and 100,017 exercise sessions, of which 56,362 involved a mistake. Ninety-six per cent of students messaged Khanmigo at least once. The median student messaged it on 33 per cent of the days they practised, in 14 per cent of exercise sessions, and in only 17 per cent of the sessions where they made a mistake. Students sent messages in 9,362 of the 56,362 mistake sessions.
What students typed also mattered. The authors classified student messages and found most were not mathematical dialogue.
| Message type | Share of student messages |
|---|---|
| Bare answer to the tutor | 39.4% |
| Click on a suggested prompt | 24.2% |
| Low-effort (“idk”, asks for the answer) | 12.6% |
| Mathematical question or step of reasoning | 14.5% |
| Off-task | 9.3% |
Source: Oreopoulos and Low, NBER working paper 35620, August 2026.
The authors are explicit about what the design can and cannot show. Because assignment moved every part of the programme at once, the study “cannot isolate the tutor’s marginal contribution”. Tennessee’s TCAP state assessment was registered as a co-primary outcome and its results are not yet in the paper.
The engagement pattern echoes earlier evidence. Winss reported in similar terms on an AI tutoring study from Stanford that found students barely used the tool when left to work independently, and on the practical classroom question of best practices when using AI in the classroom.
Structure changed what the AI did
The second paper, “Making AI Tutoring Productive” by Oreopoulos, Michael Liut of the University of Toronto, Alp Sungu of the University of Pennsylvania and Low, tested whether the way AI is embedded matters. More than 6,000 Hamilton County students in grades 6 to 8, across 20 schools and just under 100 teachers, were randomised at login across three dimensions: AI support versus conventional computer-assisted learning, mastery versus non-mastery progression, and one of two maths topics. The practice session ran the week of 23 March 2026 and a delayed assessment followed the week of 30 March.
Mastery progression required a student to watch at least a minute of video and answer three questions in a row correctly before advancing. That rule raised three-correct-in-a-row attainment by 28.7 percentage points and added 4.62 attempted questions, but on its own it did not improve delayed learning. AI on its own also did not. The one positive delayed-test result came from the combination: among mastery students, 40.2 per cent of those with AI answered the practised delayed question correctly against 37.0 per cent of those without, a gap of 3.1 percentage points at p = 0.065. The authors describe the effect as marginal and concentrated on practised material, with no transfer to unpractised items.
Oreopoulos told The Hechinger Report on the day of release: “I don’t want to jump out and say we’ve demonstrated that AI is going to be the game changer that we hope it is.” He added that the result “might be the first kind of evidence that shows there’s at least some hints that it has some positive value against no AI at all.”
The paper lists five limitations, including a delayed assessment of only four questions and a single mastery rule that “may be too weak, too gameable, or too disconnected from conceptual understanding”. The authors also state that NUMI, the platform built for the study by University of Toronto computer science students, “was a specific platform, not generic AI access”, and that results should not be generalised to all uses of large language models in education.
Take-up rose sharply when the invitation changed
The third paper, “Virtual Tutoring with Computer-Assisted Learning” by Oreopoulos, Ruochong Dong and Low, ran with the Toronto District School Board across 2023-24 and 2024-25. Teachers nominated struggling grade 4 to 8 students, whose parents were then randomly offered about an hour a week of one-to-one online tutoring at home from volunteer university students, layered on top of a classroom Khan Academy programme both groups received.
Participation was the obstacle. Only 45.3 per cent of assigned students reached a first session in year one. After the researchers reframed the invitation and simplified enrolment, first-session take-up rose to 82.5 per cent, though weekly attendance stayed below half because students attended intermittently rather than dropping out. The offer raised Khan Academy practice by 10.24 minutes a week. Maths assessment scores rose 0.055 standard deviations, but the 95 per cent confidence interval runs from -0.09 to 0.20, and the design had 80 per cent power only for effects of about 0.21 standard deviations, so the result cannot be distinguished from zero. Report-card marks rose about 0.08 standard deviations at p below 0.10. The authors also flag two baseline imbalances favouring treated students, meaning any bias in the achievement estimate “would run upward”.
The three trials side by side
| Trial | Setting | Sample | Duration | Headline result |
|---|---|---|---|---|
| Khanmigo in remedial maths (w35620) | 18 middle schools, Hamilton County, Tennessee | 6,902 student-term observations, 53 clusters | 2024-25 and 2025-26 | 1.26 percentile ranks per term; 0.062 SD per school year, stacked |
| AI plus mastery practice (w35621) | 20 schools, grades 6-8, Hamilton County | More than 6,000 students | One practice week plus delayed test, March 2026 | 3.1 percentage points on the practised delayed question, p = 0.065 |
| Home virtual tutoring (w35622) | Toronto District School Board, grades 4-8 | About 1,388 nominated students | 2023-24 and 2024-25 | Take-up 45.3% to 82.5%; +10.24 practice minutes a week; 0.055 SD, not distinguishable from zero |
Sources: NBER working papers 35620, 35621 and 35622, all issued August 2026.
Background
Randomised evidence on generative AI tutoring in schools has been thin relative to the volume of product launches. Khan Academy released Khanmigo in 2023 as a chatbot layer over its practice platform, and positioned it as a coach that withholds answers rather than supplying them. In April 2026, Chalkbeat reported that Khan Academy’s leadership had described Khanmigo as, for most students, “a non-event”, and the organisation redesigned the student experience that year so that the tutor activates automatically during practice. Philip Oreopoulos has worked on tutoring effectiveness for more than a decade, including a 2020 NBER meta-analysis of tutoring trials with Andre Nickow and Vincent Quan. The Hamilton County Khanmigo trial began in the 2024-25 school year and was pre-registered as AEARCTR-0013519, with state test results still to be added. Taken together, the three August 2026 papers point at the same conclusion from different angles: the measured constraint on AI tutoring in schools is how often and how deeply students use it, not whether they have access to it.
Sources: National Bureau of Economic Research; National Bureau of Economic Research; National Bureau of Economic Research; The Hechinger Report
Featured image: photo by Katerina Holmes on Pexels (free Pexels license).
Become a Sponsor
Our website is the heart of the mission of WINSS – it’s where we share updates, publish research, highlight community impact, and connect with supporters around the world. To keep this essential platform running, updated, and accessible, we rely on the generosity of you, who believe in our work.
We offer the option to sponsor monthly, or just once choosing the amount of your choice. If you run a company, please contact us via info@winssolutions.org.
I specialize in sustainability education, curriculum co-creation, and early-stage project strategy. At WINSS, I craft articles on sustainability, transformative AI, and related topics. When I’m not writing, you’ll find me chasing the perfect sushi roll, exploring cities around the globe, or unwinding with my dog Puffy — the world’s most loyal sidekick.