A student is partway through a spoken roleplay with our tutor, asking a lecturer for an extension, and they stop. They know roughly what they mean. They cannot get to the words. What happens in the next two seconds is the one part of our product that turns out to have a named theory of learning behind it, and we did not know that when we built it.
The system does not hand over the sentence. The first prompt is the least explicit one it has. If that does not get the student moving, the next is slightly more explicit, and the one after that is more explicit again. Supplying the actual form is the last thing the system does, and it only gets there when everything smaller has already failed.
Written down, that reads as a preference about tone. What it changes is what the interaction produces. If a student needs the full answer handed over, we have learned one thing about where they are. If the smallest prompt was enough, we have learned something quite different. Both students may end up saying the identical sentence out loud, and by the measure most software applies, they performed the same. They did not. The amount of help it took is the actual information, and a system that always helps in the same way at the same moment destroys that information before it exists.
How much nudging it took is the thing worth knowing, and stating that needs no vocabulary at all.
The tradition the ladder turns out to sit in
It has a name anyway. Fusing assessment and instruction instead of running them in sequence, offering the least explicit help that might work, escalating only as far as necessary, and treating responsiveness to that help as the measure of where a learner is, is Dynamic Assessment. In second language work the names attached to it are Matthew Poehner and James Lantolf. It descends from Reuven Feuerstein's work on mediated learning and behind that from Vygotsky, and it turns on his zone of proximal development, the distance between what a learner can manage alone and what they can manage with help.
That phrase is the only piece of borrowed vocabulary in this essay we could not do without, because it is the one place the idea is unintelligible without it. Everything else here can be said in ordinary English, and where it cannot, that is usually a sign the thought is not finished.
We should be honest about how we arrived at it, because the honest account is less flattering than the theoretical one and more useful.
The instinct came first and the reading came later
We did not read Vygotsky. Neither of us has read Poehner and Lantolf. We are a maths student and an engineering student, and we built the hint ladder that way because handing somebody the answer the moment they hesitated felt like doing the work for them. That was an instinct about fairness. It was not a position about learning, because we did not have one. We went looking for the theory only after the ladder was already built, and found a tradition running back to Vygotsky sitting behind the instinct.
The tempting response to that is confidence. We got somewhere real without the reading, so perhaps the reading was optional. We think that is exactly the wrong conclusion. One instinct out of many happened to land on a defensible position, and we had no way of telling in advance that it would, which means we have no way of telling now which of our other instincts did not land. Getting something right by accident is not a method you can rely on twice. If anything it should make a company more careful, because it shows that the difference between the good decision and the bad one was invisible to the people making it.
The rest of the product makes that concrete, because most of it does the opposite.
What the writing feedback does with an error
Our writing feedback works on practice pieces, never on real assessed coursework, and it will not rewrite a piece at essay level. That much we are comfortable with. But when it finds an error, it supplies the corrected form. In the terms the field uses, that is direct corrective feedback, as against indirect feedback, which flags that something is wrong and makes the student produce the fix. The argument for indirect is the argument for the roleplay ladder: the retrieval has to happen inside the student, and if the system performs the retrieval, a student can copy a correct sentence without ever working out why it is correct.
We do not want to overstate the evidence, because it is not as clean as we would like. John Truscott's 1996 paper in Language Learning put the strong case against grammar correction and has not survived in that form. His methodological point has, and it is the one we are most likely to fall for: a student fixing an error we flagged is not evidence that the student learned anything.
How far the correction evidence goes
On the meta-analytic side, what we can honestly report is a disagreement. Scherer, Graham and Busse (2026), across thirty-three studies in Assessing Writing, found that surface-level gains from automated feedback were not maintained at follow-up, with an effect of -0.02, while deeper-level gains persisted and grew, at 0.54, and that second language learners profited more than any other group. Luo and Zhan (2026) put the overall effect at 0.58 across forty-three studies. Kaliisa and colleagues (2026), across forty-one studies and 4,813 students, found no statistically significant difference between machine and human feedback at all.
Anyone quoting one of those three and not the others is choosing their answer in advance, so we will claim only what the one we like is worth. It is a reason to suspect that the fastest thing to build and the most satisfying thing to use is also the thing least likely to survive the term. It is not a finding about our product, which nobody has studied.
So one feature of our product embodies a theory of learning that its own field takes seriously, another feature contradicts it, and until writing this page put the two side by side we had not noticed that they were in tension. We do not have a resolution to offer. The choice is between rebuilding the writing feedback so that it withholds the form the way the roleplay withholds it, and explaining why an argument that governs speaking should stop at the edge of writing. We have not made that choice yet, and we would rather say so than construct a reason after the fact.
What a student picked out unprompted
There is one piece of outside evidence and it is small. We gave the app to two Chinese students. One of them, at IELTS 6.5, used it for about two days and wrote back. She named the roleplay as the most useful thing in it, and the reason she gave was this: "For Chinese students, we only remember and spell but don't know how to use it; it's a great way to tell students how to use it."
That is one user over two days and we are not going to inflate it. What we notice is that the feature we can defend theoretically and the feature she picked out unprompted are the same feature, and that her reason and our reason are recognisably the same reason arrived at from opposite ends. She has no stake in the argument and no vocabulary for it. She described the difference between holding a word and being able to use it, which is the distinction the whole hint ladder exists to detect.
The concession we owe here is that the roleplay is the best thing we have built and it is not the thing we set out to build. The founding idea was vocabulary drawn from a student's own course material, and that is where the engineering attention went. The roleplay was the other feature, and it is not where the thinking went. The part of our product with a named theory behind it is the part we were not concentrating on. The same student, in the same message, named the course material upload as the innovative thing in the app and told us the opening content was pitched too low for somebody at her level, which is a fair description of a company that took its emphasis from its own enthusiasm and not from anybody using the thing.
Mediation without the assessment
The caveat is larger than the concession, and it goes to whether we can claim the theory at all. Dynamic Assessment treats how much help a learner needed as the measure. Our system generates that information every time a student hesitates and then does nothing with it. How far up the hint ladder a student went is not recorded anywhere that student or a tutor could see, and it does not change what the app offers the next day. We have not built that. Until we do, what we have is the mediation without the assessment, which is half of the thing, and calling it Dynamic Assessment would be describing a shape and not a working measure.
The open question we cannot answer from where we are standing is whether the measure survives being automated at all. When a human tutor withholds an answer and then offers a slightly larger prompt, they are reading a person they know, in a room, over weeks. Poehner and Lantolf's mediators are people. Ours carries no record of how much help this student needed last week, and it cannot tell a hesitation caused by not knowing the word from a hesitation caused by tiredness. Nobody has used our app for a hundred hours, so we have no data either way.
The most we would claim for it is a ceiling, which is a mediocre classroom that happens to be open at eleven at night, in the hours when there is nobody there to read anybody. It is possible that the ladder works because the principle is sound and the principle is indifferent to who is holding it. It is also possible that the thing doing the work in the original research is the person, and that we have built the visible half.