The Streak and the Straw: What Karpathy Gets Right, and What Duolingo Gets Away With
Real learning is slow and painful. So why do I trust my Duolingo owl? A note on load-bearing friction, and what it means for the products we design.
Not long ago I built a small app to help designers rehearse the conversations they dread: the “I disagree with this roadmap” talk, the “I need to renegotiate this deadline” talk. Somewhere in the build I hit a fork I couldn’t design my way around. Do I reward people for showing up to practice, with a streak and a satisfying green check, or for actually doing the uncomfortable reps, the halting, awkward, say-it-out-loud part that makes the real conversation easier later?
Those are not the same product. And I had a bad feeling I knew which one people would open every day.
Andrej Karpathy would have an opinion about my fork. This year he, the person who taught a generation of us neural networks from scratch one hand-derived gradient at a time, quietly joined Anthropic’s pretraining team. Not agents. Not the demo-friendly layer everyone’s pitching. Pretraining: the slowest, least glamorous grind at the very bottom of the stack. It’s a fitting move for someone who keeps insisting that real learning is slow and painful, that you should want the mental equivalent of sweating, close the “Learn X in 10 minutes” tabs, and seek the meal instead of the snack.
I agree with all of it. I’d also just spent weeks building something that, done wrong, would let designers fake the sweat.
So which is it?
The feeling of learning is not the learning
Here’s the uncomfortable part: Karpathy isn’t just being a purist. The cognitive science backs him almost embarrassingly well.
In the early ’90s, the psychologists Robert and Elizabeth Bjork named a paradox they’d been chasing for years: the conditions that make learning feel smooth and productive in the moment are often the ones that produce the weakest long-term memory. They called the good kind of hard “desirable difficulties,” things like spacing your practice out, testing yourself instead of rereading, and mixing topics so your brain has to work to tell them apart. Every one of these feels worse while you do it. Every one of them wins on the delayed test.
Their sharpest insight is that performance during learning and actual learning are not the same thing. They can even move in opposite directions. When material flows easily, you feel fluent, and you mistake that fluency for mastery. Elizabeth Bjork’s line stays with me: forgetting is the friend of learning. The struggle to drag something back out of your memory is the thing that strengthens it.
That’s Karpathy’s “sweating,” dressed in a lab coat. The feeling of learning is a trap.
But Duolingo isn’t the villain in this story
Here’s where the “close the tabs, seek the meal” crowd loses me a little.
If you actually look under Duolingo’s candy shell, its engine is desirable difficulty. Spaced repetition. Retrieval practice: recall, not recognition. Interleaving across lesson types. The green owl smuggled decades of memory science into a format that feels like a game. The bite-sized packaging isn’t the enemy of hard learning; in the best case, it’s the delivery mechanism.
And it solves the one problem Karpathy’s advice quietly assumes away: showing up. Slow, painful learning only compounds if you keep doing it, and the failure mode for most adults isn’t that the textbook is too easy. It’s that the textbook is closed. A streak is a commitment device wearing a cartoon costume.
So where does it actually fail? Right where the reviews land in 2026: learners plateau around A2 to B1, and the app’s progress metrics quietly overstate how much they know. People “game the app,” chasing the streak, tapping through the easy path, mistaking a 1,000-day counter for conversational French. The candy shell starts eating the medicine. The moment the streak becomes the goal instead of the scaffold, it stops measuring learning at all. Goodhart’s law, in a fifteen-second dopamine loop.
What this means for the products we make
I don’t think the design lesson here is “make things harder.” It’s more precise, and more useful, than that:
Know which friction is load-bearing.
Every product decision about learning, and increasingly about work, is a decision about which difficulty to preserve and which to remove. Removing friction is our whole job, right up until we remove the friction that was doing the learning. So before you smooth something out, run a quick audit:
What’s the real outcome, and what’s the proxy you’re actually optimizing? (Fluency, or streak length?)
Which friction produces that outcome, and which is just annoying? Effortful recall is load-bearing. A confusing settings menu is not.
Is your engagement loop serving the outcome, or quietly replacing it?
The honesty test: does your progress metric track the thing, or the feeling of the thing?
If that sounds abstract from your day, look at where AI tooling is right now. Karpathy, again, is the one calling much of “vibe coding” what it is: a fast path to confident slop. The same trap, one level up. These tools make the feeling of productivity nearly free, and the temptation is to optimize for that feeling instead of the competence underneath it. The designer’s edge in that world isn’t speed. It’s judgment about which friction is worth keeping.
That, to me, is the whole game as generation gets commoditized. Anyone can remove friction now; a machine will do it for you. Knowing which friction was holding the roof up is the part that’s still hard, still slow, and still yours.
For what it’s worth, I keep landing on the same answer for my own app: reward the reps that are actually hard, the ones where you stumble, not the days someone showed up to tap through the easy ones. Even knowing the second version is the one that would juice my numbers. I’m not always sure that’s the right business call. I’m fairly sure it’s the right design one.
So here’s my open question, and I’d love a fight about it in the comments: as our tools get better and better at making everything feel effortless, what’s the load-bearing friction you refuse to design away?
References
Andrej Karpathy, “on shortification of ‘learning’.” X, Feb 2024.
Andrej Karpathy. Wikipedia (Anthropic pretraining team, 2026). https://en.wikipedia.org/wiki/Andrej_Karpathy
“Beyond the Hype: 5 Counter-Intuitive Truths About AI from Andrej Karpathy” (summary of the Dwarkesh Patel interview). DEV Community, Oct 2025. https://dev.to/amananandrai/beyond-the-hype-5-counter-intuitive-truths-about-ai-from-andrej-karpathy-afk
Robert & Elizabeth Bjork, “Desirable Difficulties: Why Effortful Learning Outlasts Easy Learning.” Glasp, 2026. https://glasp.co/articles/desirable-difficulties
Bjork & Bjork, “Introducing Desirable Difficulties Into Practice and Instruction” (PDF). University of New Hampshire. https://www.unh.edu/teaching-learning-resource-hub/sites/default/files/media/2023-06/itow-introducing-desirable-difficulties-into-practice-and-instruction-bjork-and-bjork.pdf
“Does Duolingo Work? How Effective It Actually Is.” The Linguist, Mar 2026. https://blog.thelinguist.com/duolingo-review/
“Duolingo French Review 2026: Does It Actually Work?” HelloFrench, Apr 2026. https://www.hellofrench.com/en/tips/duolingo-french-review
“Results from Duolingo efficacy studies.” Duolingo Blog. https://blog.duolingo.com/results-duolingo-efficacy-studies



