A Wharton-led field study gave nearly 1,000 high school students plain ChatGPT access for math practice. Their practice scores jumped 48%. Then researchers took the AI away and gave everyone the same exam. Scores fell 17% below a group that never used AI at all. The gains didn't stick, because students had leaned on the tool for answers instead of learning the steps.
Short answer: In a 2025 PNAS study of ~1,000 Turkish high schoolers, unrestricted GPT-4 access raised practice-problem scores 48% but lowered later exam scores 17% versus a no-AI control group. A second, "guardrailed" AI tutor that withheld direct answers still boosted practice scores 127% and erased the exam-score drop, showing the harm came from how the AI was used, not from AI itself.

Last updated: August 22, 2026 · By Vishal Swami, Founder & Lead AI Reviewer, AISagely
I write about AI tools people use to study, and this is the study I keep sending to parents and teachers who ask if ChatGPT homework help actually helps. The short version: it depends on which mode the student is using. The gap between "helps" and "hurts" came out to about 65 percentage points on the same exam.
What the study actually found
The paper is "Generative AI without guardrails can harm learning: Evidence from high school mathematics." The authors are Hamsa Bastani, Osbert Bastani, Alp Sungu, Haosen Ge, Özge Kabakcı, and Rei Mariman. It ran in the Proceedings of the National Academy of Sciences on June 25, 2025, per the PNAS paper. The field experiment involved close to 1,000 high school students in Turkey, working through math practice problems over several weeks. Students were split into three groups:
- Control — no AI access, just the standard practice problems.
- GPT Base — a plain ChatGPT-style interface with no restrictions on how it answered.
- GPT Tutor — the same underlying model, but prompted with pedagogical guardrails that pushed hints and partial steps instead of finished answers.
During practice, both AI groups beat the control group by a wide margin. GPT Base students scored 48% higher. GPT Tutor students scored 127% higher, according to Knowledge at Wharton’s coverage. Then the researchers took the AI away and gave everyone the same unassisted exam. GPT Base students scored 17% worse than the control group, a statistically significant drop. GPT Tutor students, the ones who'd been blocked from getting direct answers, scored about the same as the control group. The negative effect wasn't just smaller for them. It was essentially gone.
Why practice scores went up but exam scores fell
The mechanism is simple once you see it laid out. When GPT Base students hit a problem they didn't understand, the fastest path to a correct-looking answer was to paste it into ChatGPT and copy what came back. That worked for the assignment in front of them. It did nothing for the math skill the assignment was supposed to build. Lead author Hamsa Bastani put it directly in the Wharton writeup: researchers were "really worried that if humans don't learn, if they start using these tools as a crutch and rely on it, then they won't build those fundamental skills."
GPT Tutor's students hit the same wall, but the tool wouldn't just hand them the answer. It walked them through one step, then made them attempt the next. That's slower and more frustrating in the moment, which is exactly why it worked: the students still had to reason it out themselves, and the AI just kept them from getting stuck. In my testing of similar guardrailed setups, like ChatGPT's built-in Study Mode, the friction is the point. A tool that never makes you think isn't teaching you anything. It's doing the assignment for you.
GPT Base vs. GPT Tutor vs. no AI: the numbers
| Group | Practice-problem score vs. control | Unassisted exam score vs. control |
|---|---|---|
| Control (no AI) | Baseline | Baseline |
| GPT Base (unrestricted ChatGPT-style access) | +48% | −17% (statistically significant) |
| GPT Tutor (guardrailed, withholds direct answers) | +127% | ~0% (no significant difference) |
Example prompts that avoid the crutch effect
If you're a student, parent, or teacher trying to get the GPT Tutor result instead of the GPT Base result, the fix isn't switching tools — it's changing the prompt. These are close to what I use when I'm testing an AI study session myself:
- "I'm stuck on [problem]. Don't give me the final answer — ask me a question that helps me figure out the next step."
- "Check my work on this problem and tell me which step has the error, but don't fix it for me."
- "Quiz me on [topic] with one question at a time, and only give me the next question after I explain my reasoning for the last one."
- "Explain the concept behind this problem type using a different, simpler example, then let me try the original on my own."
Tools built specifically for this — like NotebookLM’s study features, which generate quizzes from your own notes instead of just answering questions — tend to enforce this structure automatically, which is one reason I recommend them over an open chat window for exam prep.
Common mistakes people make with AI and homework
The first mistake, and the one this study measured directly, is treating AI chat as an answer key instead of a tutor — pasting in the problem and copying the output without reading it. I've watched this happen with my own testing sessions: it's genuinely hard to resist when the answer is one message away and the deadline is in twenty minutes.
The second mistake is assuming any "AI study tool" label means it has guardrails like GPT Tutor did. Plenty of study apps are just a chatbot with a different skin, and the difference between "hints only" and "answers on demand" often comes down to a system prompt the student never sees. Check what the tool actually withholds before trusting it to prep someone for a closed-book exam.
The third mistake is skipping the unassisted practice entirely. Even GPT Tutor's students needed to attempt each step themselves — the AI removed getting permanently stuck, not the work of thinking. A tool that answers instantly every time removes both.
Tools that make this easier
If you want the GPT Tutor result rather than the GPT Base result, start with a roundup rather than a single default chatbot. My AI tools for students guide compares the study-specific options against plain ChatGPT, and best AI tool for study ranks them by how much they resist the temptation to just hand over the answer. For structured practice, ChatGPT Study Mode is the closest built-in equivalent to the guardrailed arm in this study, since OpenAI designed it to ask questions back rather than solve the problem outright. This crutch effect isn't limited to math homework, either — the same underlying pattern shows up in how schools are responding to AI-written essays, which I covered in Denmark’s oral-defense rule for AI cheating and in a professor’s invisible prompt trap that caught students copy-pasting into ChatGPT.
My take
This study is the clearest evidence I've seen that "AI hurts learning" and "AI helps learning" can both be true about the exact same model. The outcome depends on whether the tool is allowed to just answer. My honest verdict: unrestricted chat is fine for a first draft of an email or a quick fact check. It's a bad default for anything a student will later be tested on without it. If you're picking a tool for homework help, pick or configure one that acts more like GPT Tutor than GPT Base. Treat "it gave me the right answer immediately" as a yellow flag, not a green one.
Frequently Asked Questions
Did the study say AI is bad for learning?
No — it found that unrestricted AI use hurt exam performance, while a guardrailed version of the same model didn't. The researchers' conclusion was that design choices, not the technology itself, determined whether students learned or just copied answers.
How big was the exam score drop?
Students who used the unrestricted GPT Base tool scored 17% worse than a no-AI control group on an unassisted exam, a statistically significant difference, per the PNAS paper published June 25, 2025.
Does ChatGPT Study Mode fix this problem?
It's built around the same idea as the study's guardrailed "GPT Tutor" arm — asking questions and giving hints instead of final answers — which is why I recommend it over a plain chat window for exam prep. See my ChatGPT Study Mode guide for how to turn it on.
Is this specific to math, or does it apply to other subjects?
The field experiment tested high school mathematics specifically, so the numbers above are math-specific. The underlying mechanism — that copying AI output skips the practice that builds a skill — isn't math-specific, which is why similar concerns are showing up around AI-written essays and other schoolwork.
Who ran the study and where?
Hamsa Bastani and Osbert Bastani, of Wharton and Penn Engineering, led the team, with co-authors Alp Sungu, Haosen Ge, Özge Kabakcı, and Rei Mariman. The field experiment involved nearly 1,000 high school students in Turkey. It was published in the Proceedings of the National Academy of Sciences in June 2025.