September 10, 2026

Stanford brief grades AI tutoring by how much human contact it keeps

A Stanford SCALE brief sorts AI tutoring into four models by human involvement and rates AI-only tutoring as emergent evidence, unknown effectiveness.

A tutor sitting beside a boy at a desk, illustrating the AI tutoring evidence on human involvement

Photo by Yan Krukau on Pexels

The phrase “AI tutor” covers four different products, and the AI tutoring evidence behind them is not the same in any two cases. That is the argument of “AI Tutoring is Not a Monolith: What We Actually Know”, a research brief dated 20 August 2026 and published by the SCALE Initiative at Stanford University through its AI Hub for Education and the National Student Support Accelerator. The brief sorts tutoring models along a scale it calls relational intensity, then attaches a separate evidence rating to each point on that scale. Human tutoring has a robust evidence base. AI-only tutoring has, in the brief’s own words, “unknown effectiveness”.

The named writers and editors are Chayne Turano, Valeria Pihl, Chris Agnew, Lauren Ziegler and Susanna Loeb, all of SCALE. Acknowledged contributors include staff from tutoring and edtech providers ThirdSpace Learning, Eedi, Tutor.com, Paper, Goblins, Snorkl and EduRaptor. The document is a synthesis of existing studies rather than a new trial, so it carries no sample of its own.

What the AI tutoring evidence says at each point on the spectrum

The brief defines relational intensity as “the depth and consistency of the human connection between a student and a tutor within a given tutoring model, ranging from human-led models with close, sustained relationships to AI-led models with minimal or no human interaction”. It then states the trade-off plainly: “As direct human relationships decrease, the evidence base becomes thinner, leaving unresolved questions about student safety, developmental impact, and long-term efficacy.”

Four models are rated.

Tutoring model Evidence and effectiveness, as stated in the brief
In-person or remote human tutoring “Robust evidence base; highly effective.” The brief notes the in-person evidence is stronger, while virtual tutoring has also shown positive effects in high-quality studies.
Human tutoring with AI support “Emergent research base; potentially as effective as, or more effective than, in-person and remote tutoring.”
AI tutoring with human support “Emergent evidence; early research indicates implementation determines effectiveness on student outcomes.”
AI-only tutoring “Emergent evidence; unknown effectiveness.”

Source: “AI Tutoring is Not a Monolith: What We Actually Know”, SCALE Initiative at Stanford University, 20 August 2026.

The brief also applies a three-colour key: green for high-intensity human instruction with strong research support, yellow for emerging AI-assisted models that depend on implementation quality, and red for fully automated AI models with evidence gaps. The colour assigned to each model appears only in the brief’s graphic, not in its text.

Where the strongest numbers come from

The clearest gain in the AI tutoring evidence comes from the model that keeps a human tutor in place and points AI at the adult instead of the child. Tutor CoPilot, a Stanford system tested in what its authors describe as the first randomised controlled trial of a human-AI system in live tutoring, worked with 900 tutors and 1,800 K-12 students from Title I schools in a school district in the southern United States, in partnership with the provider FEV Tutor. Students of tutors given the tool were 4 percentage points more likely to pass their exit ticket. The effect was larger for the tutors who had been rated lowest before the trial.

Tutor CoPilot effect on the share of students passing an exit ticket, by prior tutor rating Tutor CoPilot: change in students passing the lesson exit ticket Percentage points, intent-to-treat estimates. 900 tutors, 1,800 K-12 students. All students +4 pp Students of mid-rated tutors +7 pp Students of lower-rated tutors +9 pp 0 3 6 9 percentage points Source: “Tutor CoPilot: A Human-AI Approach for Scaling Real-Time Expertise”, Stanford University. The all-student estimate is reported at p<0.01. The mid-rated group moved from 61% to 68% passing.
The Tutor CoPilot trial reported its largest effect among students working with the tutors rated weakest before the study. The paper puts the tool’s cost at about 20 US dollars per tutor a year, based on usage during the trial.

The paper reports a second, larger estimate of 14 percentage points in a specification based on actual tool usage rather than assignment. The 4 and 9 point figures are the intent-to-treat headline, and are the ones the brief carries.

The usage problem sitting underneath every model

The brief spends as much space on whether students open the software as on whether it works. In a study of 181,359 pupils in grades 3 to 8 across 99 school districts using a supplemental maths platform, only 5 per cent reached the recommended 30 minutes or more a week and 9 per cent reached the 15 to 29 minute band. An intraclass correlation of 0.57 indicated that teacher, school and district factors explained 57 per cent of the variance in how much students used it. Roughly 41 per cent recorded no use at all. That work, by Grimaldi, Weatherholtz and Millwood Hill, was published in the Proceedings of the 15th International Conference on Educational Data Mining in 2022. All three authors worked for the platform’s maker, Khan Academy, and the design was quasi-experimental rather than randomised.

The two-district randomised trials the brief draws on found between 40 and 47 per cent of students never used the AI platform, with those who did logging on for four to five weeks in total across a multi-month intervention. Winss covered that study when it appeared: see students barely used AI tutors when left alone. Adding human check-ins raised engagement but did not lift end-of-year reading achievement.

Benchmark cited by the brief Figure Named source
Learning gains from high-impact tutoring 3 to more than 15 months American Educational Research Journal, DOI 10.3102/00028312231208687
Recommended dosage 3 or more sessions weekly, 90 minutes total, 10 or more weeks NBER working paper 22130 and NSSA design principles
Recommended tutor-to-student ratio 1:4 or lower, with 1:1 strongest NSSA EdResearch one-pager; EEPA 10.3102/01623737251364573
Students reaching 30+ minutes a week on a maths platform 5% of 181,359 Grimaldi et al., EDM 2022
Students recording no use of that platform 41% Grimaldi et al., EDM 2022
Students who never used the AI tutor in two district trials 40% to 47% “Access Is Not Enough”, Stanford NSSA
Absence reduction on scheduled tutoring days, Washington DC 7% less likely to be absent DOI 10.1177/23328584261448132

Chayne Turano, quality and improvement program manager for high-impact tutoring at SCALE and the brief’s first-listed author, told Government Technology on 27 August: “We can’t even get to the question of, ‘is the program good enough?’, because we aren’t seeing the [student] usage.” Turano added that the support and the evidence currently sit “in using AI to support the adults that are implementing the high-impact tutoring”.

Ana Trindade Ribeiro, a senior researcher at SCALE who is not listed as an author of the brief, told the same outlet that “Good AI tutoring is something more akin to a program structure. It doesn’t come as just a tool.” Government Technology also reported cost figures from a separate Khanmigo trial in 18 middle schools in Hamilton County, Tennessee, putting that tool at about 15 US dollars per student a year against several thousand dollars for high-dosage human tutoring. Those figures do not appear in the brief itself. Winss reported the underlying trial separately in the Khanmigo AI tutoring study finding small maths gains.

Why the spectrum matters for procurement

The practical consequence is that a district comparing two products labelled “AI tutor” may be comparing a staffed programme with a chatbot. The brief’s position is that the evidence supports the models that keep an adult in the loop, and that the fully automated end has not yet been tested at a standard that would settle the question. That maps onto the wider pattern set out in earlier coverage of how AI can support rather than replace teachers and the OECD finding that AI can raise grades while hurting learning.

About the National Student Support Accelerator and SCALE

The National Student Support Accelerator was launched in 2021, led by Susanna Loeb, to widen access to high-impact tutoring for K-12 students after pandemic school closures. It is now a programme of the SCALE Initiative at Stanford University, based at the Stanford Graduate School of Education and part of the Stanford Accelerator for Learning. SCALE stands for Systems Change Advancing Learning and Equity, and its projects include the AI Hub for Education, Getting Down to Facts and Tips-by-Text. NSSA reports more than 60 research studies, over 40 tools and briefs, and research partnerships in more than 18 states. Its listed funders include the Overdeck Family Foundation, the Bill and Melinda Gates Foundation, the Walton Family Foundation, Zoom, America Achieves, Charles and Lynn Schusterman Family Philanthropies and Kenneth C. Griffin.

Loeb is a professor and faculty director of the SCALE Initiative at Stanford. She previously led the Annenberg Institute at Brown University, was founding director of the Center for Education Policy at Stanford, and returned to Stanford in 2023. The AI Hub for Education, which Stanford also lists under the title Generative AI for Education Hub, published “The Evidence Base on AI in K-12: A 2026 Review” earlier this year, a survey that found only a small fraction of studies on AI in schools used rigorous designs. The new brief applies that same standard product by product, and concludes that the further a model moves from a human tutor, the less there is to go on.


Sources: Stanford SCALE Initiative; Stanford National Student Support Accelerator; Government Technology; Tutor CoPilot working paper, Stanford University; Stanford National Student Support Accelerator; Grimaldi, Weatherholtz and Millwood Hill, Educational Data Mining 2022; American Educational Research Journal; Stanford National Student Support Accelerator; Stanford SCALE Initiative

Featured image: photo by Yan Krukau on Pexels (free Pexels license).


Become a Sponsor

Our website is the heart of the mission of WINSS – it’s where we share updates, publish research, highlight community impact, and connect with supporters around the world. To keep this essential platform running, updated, and accessible, we rely on the generosity of you, who believe in our work.

We offer the option to sponsor monthly, or just once choosing the amount of your choice. If you run a company, please contact us via info@winssolutions.org.

Select a Donation Option (USD)

Enter Donation Amount (USD)

What do you feel about this?