The Trap Inside the Instant Answer
The Scholé team
The first time an AI writes the tricky email for you, something small and pleasant happens: a task that would have cost twenty minutes costs two. The draft is good. Better, possibly, than yours would have been. It is hard to overstate how persuasive that moment is, and how little it tells you. Because two things happened at once, and only one of them was visible. Your output improved. And the twenty minutes of working out what to say, the part that would have made the next email easier to write yourself, didn’t happen. Researchers who study this describe a learning-performance paradox: the same help that lifts what you produce can quietly lower what you’re able to do. Nobody notices, because the work keeps shipping. This article is about that gap, and about what it would mean to build AI that closes it instead of widening it.
One experiment runs through this article: roughly 1,000 students at a large Turkish high school practiced math with an AI chatbot, then sat a closed-book exam without it (Bastani and colleagues, in the journal PNAS, 2025). What it found:
- +48%
- better on practice problems for students whose AI simply handed over answers, versus classmates without AI
- −17%
- worse on the closed-book exam for those same students, versus classmates without AI
- +127%
- better on practice for a second group, whose AI gave hints instead of answers, with exam scores indistinguishable from classmates without AI
Cognitive offloading: why you can’t learn anymore
The pattern inside that email moment is not a hunch; it has a literature. Psychologists call the move cognitive offloading: handing a mental process to an external tool so you no longer have to run it yourself. Offloading is often rational; no one should be memorizing phone numbers. But it carries a known cost: the process you offload is the process you stop practicing. When that process happens to be the very skill you were trying to build, the help and the harm arrive in the same gesture.
The clearest picture of how this plays out with generative AI comes from one large field study. Bastani and colleagues ran an experiment in Turkish high-school math classes (roughly a thousand students), reported in 2024 and published in 2025 under the title “Generative AI Without Guardrails Can Harm Learning.” Students who practiced with an unguarded GPT-4 chat interface (a standard chatbot, free to hand over the answer when asked) scored 48% better on practice problems than peers working without AI. Then came the closed-book exam, where the same students scored 17% worse than peers who never had the tool. Every visible signal during practice said the AI was working. The ledger that mattered said otherwise.
Hold on to one detail, because the article turns on it. The study had a second group: the same model wrapped in tutoring guardrails that gave hints instead of answers. Those students scored 127% better on practice than peers without AI (guided practice helped even more than handed answers), and the exam harm essentially vanished; their scores were statistically indistinguishable from students who never used AI at all. It is one study, in one setting, and deserves that much caution. But the distance between those two groups is the point: what separated the help that cost from the help that didn’t was not the AI. It was the design.
Two designs, two ledgers
In the study, students practiced math with one of two versions of the same AI: one that simply gave answers when asked, and one built to tutor: hints, never the answer itself. Every bar compares a group’s scores with classmates who worked without AI.

One study, one setting: Bastani, Bastani, Sungu, Ge, Kabakçı & Mariman, “Generative AI Without Guardrails Can Harm Learning: Evidence from High School Mathematics,” PNAS (2025); the paper’s “GPT Base” and “GPT Tutor” groups. The bars compare groups with their classmates; they do not mean any student’s own scores fell. The ≈0 bar is drawn at zero because the hints group’s exam scores were statistically indistinguishable from classmates who never used AI.
The paradox
The uncomfortable part is why the cost shows up at all. The tempting explanation is moral: the students leaned too hard and should have known better. The evidence points somewhere less flattering to the rest of us: the effort the AI removed was not the packaging around learning. It was the learning.
Cognitive science has circled this from several directions. The generation effect: you remember what you produced better than what you merely read, often even when what you produced was worse. Desirable difficulties, in Robert Bjork’s phrase: the conditions that slow you down and multiply your errors during practice (retrieving, struggling, generating) are, in study after study, the ones that make new knowledge stick. The twenty minutes of working out what to say was never overhead. It was the mechanism by which the skill formed.
Which is why the loss is structural, not a failure of willpower. When a tool removes the effort, it removes the mechanism, and no amount of discipline afterward re-runs a process that never ran. The paradox also protects itself, because fluent practice feels like learning, and polished output reads as mastery, to you and to anyone reviewing your work. Nothing visible goes wrong, so no one goes looking.
The right way to use AI
The conclusion this seems to point to (use AI less) is the wrong one. It surrenders the genuine gains, and it misreads the study, whose successful group did not use less AI. Start somewhere else: a good tutor is also a form of help, and no one worries that an hour with a good tutor erodes what you can do on your own.
The difference is what the help is aimed at. The AI that writes your email is aimed at the task; its success is a finished email, and your growth is nobody’s objective. A tutor is aimed at the person. A tutor watches how you’re doing, remembers where you struggled, and chooses the next problem so that it stretches you without losing you. The work stays yours; the tutor’s intelligence goes into choosing it. That is the design answer the paradox is asking for. Not AI that does the work, but AI that does the noticing, the remembering, and the choosing, and then hands the work back.
What the help is aimed at
Two ways to build AI help. Both start with you and a task. They part at what the intelligence is spent on.

Adaptive learning is the right-hand column, made software: the intelligence is spent on choosing the encounter, and each next step lands in the stretch.
Adaptive Learning Companion
We build Scholé, an adaptive learning companion that teaches people how to work with AI. The paradox is our design brief.
Not AI that does the work, but AI that does the noticing, the remembering, and the choosing, and then hands the work back.
First, the suggestion that explains itself. When you finish a category of lessons, your Learning Journey (the path of lessons Scholé lays out ahead of you) suggests one new lesson and states the reason it was picked, drawn from what you’ve told it about your role and goals, and from what it has watched you do. It never adds itself to your path; you accept or decline. The reason is the load-bearing part. A recommendation without one is a small act of offloading in its own right: the system quietly taking over your judgment about what you need next. Showing the reasoning keeps that judgment inspectable, and keeps the deciding yours.
Second, the step that knows your steps. Inside a lesson, each next step (an explanation, a visualization, a drill) is chosen in light of the steps you’ve already taken, so that what comes next is new but within reach. And an idea that didn’t stick doesn’t simply repeat, louder; it comes back later in a different form. Within reach does not mean easy: you still retrieve, still struggle, still produce; the system does not remove the difficulty, it places it where it pays. That is desirable difficulty made operational: the system’s intelligence is spent not on producing your answers but on arranging the encounters in which you produce them.
Watching the estimate move
Straight from the app’s help page: five steps of one lesson. Answers come in, right and wrong, and Scholé’s estimate of where the learner stands climbs from 0.20 to 0.90.

Right and wrong answers alike feed the estimate. A miss is information, not a verdict; the next step is chosen knowing how the last one went.
Third, the no that stays a no. Decline three suggestions in a row and something unusual happens: nothing. The suggestions stop. There is no escalation: no new tone, no manufactured urgency, no nudge dressed up as encouragement. Scholé takes the pattern for what it is, information, and goes quiet until you decide to turn lesson recommendations back on. It is a small mechanic, easy to miss, and it is the part of the system we would give up last. Software that adapts to you is, in effect, making claims about you. The least it owes you is the grace to be wrong, and the decency to hear a no the first few times you say it.
The life of a suggestion
When you finish a category, one lesson is offered, with the reason it was picked. It never adds itself; the yes or no stays yours.

The quiet is designed. After three declines in a row, Scholé stops suggesting until you turn Lesson recommendations back on.
Here is what we can claim, and what we can’t. One large field study says unguarded AI can raise performance while lowering learning, and that guardrails aimed at the learner all but erased the harm (in one subject, in one setting, and among students, not working professionals). Adaptive learning more broadly has promising results to date, not proof. Scholé is built on the bet that the tutoring intuition generalizes: that AI aimed at the person can carry those gains into working life without the quiet debit. That is a bet, and we would rather say so than read a single study as a verdict.
What we ask is that you keep the standard this article has been circling. When an AI helps you, or helps the team you’re training, ask what the help was aimed at, and whether the work came back. Hold every tool that claims to teach to that standard. Ours included.
- Learning science
- Cognitive offloading
- Adaptive learning
- Desirable difficulty
- AI proficiency