The Brown case: when students get the right answer for the wrong reason
In late 2025, Brown professor Andrew Serrano published the numbers from a single semester of an intro course. Take-home midterm: 40 of 86 students scored a perfect 100, class average 96. Same cohort, in-person final with no AI: average 48, nineteen fails, and 22 of the 27 students who dropped out had scored 100 on the take-home.
The scores were real. The learning wasn't. It's the cleanest public data point yet on what unsupervised AI use does to a classroom — and it names a problem the whole category has been dancing around: students are getting the right answer for the wrong reason, and neither the student nor the teacher notices until it's too late.
Why the right answer is the trap
A correct answer feels like understanding. It isn't. The retrieval attempt, the confidence calibration, and the second-pass check are the steps that actually build memory.
A correct answer feels like understanding. It isn't. If a student types a question, pastes the reply, and moves on, three things are missing: the retrieval attempt (which is what actually builds memory), the confidence calibration (do I know this, or did the model?), and the second-pass check (would I get this again in a week, on paper, cold?).
Most AI tutors optimize for the first thing — deliver a correct, well-formatted answer — and quietly punish the other two. The chat surface is frictionless in the wrong direction.
What an honest AI tutor should do instead
Two features have to sit at the center, not on a settings page.
Proof Receipt. Every answer ships with a chip that shows where the answer came from (your notes, the model's general knowledge, or a verified second opinion) and how sure the model actually is. Low-confidence answers are labeled low-confidence — never dressed up. A one-tap Check my work re-runs the prompt on a stronger, different-family model; if they disagree, the disagreement is shown, not hidden.
Practice. The moment a student gets an answer, the next tap turns it into a 3-minute recall drill — same concept, different question, on paper-style prompts with no AI in the loop. This is the retrieval step the take-home midterm skipped. It's also the step that would have caught the 22 students before the final.
The gap a parent digest should surface
Serrano didn't have a signal until the final. A weekly digest that compares AI-assisted work against unassisted recall — take-home vs. cold-recall gap, per topic — is the metric he wishes he'd had in November. That's a small line in an email, but it's the difference between finding out in week 6 and finding out in week 14.
The honest framing
AI that makes homework easier isn't neutral. It shifts effort out of the exact step that produces learning, and it hides the shift behind a good-looking score. The fix isn't to ban the tool — students will use it anyway — it's to build the tool so the honest path is the default: show uncertainty, show sources, and make recall the next tap, not the next tab.
LemonSugar Ai ships Proof Receipt on every answer and puts Practice one tap from the reply. Right answer, right reason — or we say so.
A ready-to-paste post tailored to this article.
- 2026-07-26
Chamath's two flaws describe the chasm. Memory is the bridge.
AI is about to hit its disillusionment phase. The winners will be the ones that remember the user. Here's why Chamath's two flaws map cleanly onto the case for a memory-first study companion.
- 2026-07-19
Small model, right tools, closed loop
Karpathy's line on agents distills the whole AI shift: small model + right tools + closed loop = terrifying capability. That's exactly the shape of a study companion that routes cheaply and remembers everything.
LemonSugar Ai routes each prompt to the cheapest capable model — automatically.