Suppose a student writes, in a practice piece, "In the experiment, researchers measured effect of temperature on reaction rate." Our writing feedback flags the missing article and gives back the corrected form: the effect. One exchange and it is done. She sees what was wrong and what it should have been, then moves on. It is the interaction we were most pleased with when we built it. It is also the one we are least able to defend.
Articles are the sort of thing a machine handles well. So are tense, spelling, punctuation and subject-verb agreement. These features share a property that makes them tractable. The error is local. The corrected form can be recovered from the sentence it sits in, and there is usually one right answer, so nothing about the wider piece has to be understood in order to fix it. That is why every automated feedback product on the market, ours included, is strongest here. It follows from the shape of what can be detected reliably, and not from anybody's design decision.
The things that decide whether a piece of academic writing works are not like that. Whether a claim is actually supported by the evidence offered for it. Whether the argument is staged in the order the genre expects, so that a reader who knows the genre can follow it without effort. Answering either of those needs to know what the piece is for and what counts as an adequate move at that point in that discipline.
The pre-LLM literature on automated writing evaluation is consistent about the split: reliable on surface form, unreliable on content, rhetoric and organisation. Dikli and Bleyle (2014), Stevenson and Phakiti (2014) and Ranalli (2018) all land in roughly the same place. Large language models have made the surface layer faster and more comprehensive. What we have read of the work on them suggests the same asymmetry survives, with the weaker performance still on the responses that depend on knowing the situation the writing sits in, though that work is recent enough that we would not want to lean on it hard.
What survives at follow-up
The uncomfortable part is that durability appears to run the other way round.
Scherer, Graham and Busse, in Assessing Writing in 2026, synthesised thirty-three studies of automated feedback. The overall effect was 0.36, which is unremarkable. What happened at follow-up is not. Gains from surface-level feedback were not maintained, at an effect of -0.02. Gains from deeper-level feedback on organisation and argument were maintained and grew, at 0.54. Second-language learners profited more than any other group in the sample.
We should say immediately that this literature is contested and that quoting one meta-analysis is a way of getting the answer you already wanted. Luo and Zhan (2026) put the overall effect at 0.58 across forty-three studies and found it strongest where the automated feedback sat alongside a teacher's. Kaliisa and colleagues (2026), across forty-one studies and 4,813 students, found no statistically significant difference between AI and human feedback. Those three do not agree about size, and anyone citing only one of them is selecting.
On the last of those we should be explicit, because it is the finding a vendor would be tempted to wave about. We are not making that comparison and we do not think it is the interesting one. Barrot (2026), in the Journal of Second Language Writing, argues that the AI-versus-teacher framing has run out of usefulness and that the work now is calibration: human and machine feedback blended rather than ranked. We think he is right. The question this essay is about is what our feedback does on its own terms, not what it does relative to somebody who has read the module handbook and knows the student.
Where the correction argument has settled
The three meta-analyses do not disagree about the shape, and the shape has been argued about for thirty years. John Truscott's 1996 paper in Language Learning, "The case against grammar correction in L2 writing classes", put the strong version: correction is ineffective and may do harm, partly because it consumes time that could go on writing, and partly because a student who expects to be corrected writes simpler sentences that offer less to correct. Dana Ferris replied at length in the Journal of Second Language Writing in 1999 and has gone on replying since. Where it has settled, roughly, is that focused correction of a small set of what Ferris calls treatable features, the ones a student can be pointed back to a rule about, can produce measurable improvement, while correcting everything wrong across a whole text does not produce durable gains.
So surface correction is the fastest thing for us to build and the most satisfying thing for a student to receive. On the follow-up evidence it is also the least likely to change anything. We would rather write that sentence about our own product than have somebody else write it about us.
Who does the work of the correction
There is a more precise version of the problem and it names what we got wrong. Asked where the line falls between helping a student and writing the piece for them, our answer was that we correct the piece where it needs correcting, we do not rewrite it, and the student has to redo the work to get new feedback. We thought the line ran between the sentence and the whole piece.
The distinction the field uses runs somewhere else. Direct corrective feedback supplies the correct form. Indirect feedback signals that something is wrong, sometimes with a code for the kind of error, and requires the student to produce the correction. Our writing feedback is direct. The case for indirect is a case about mechanism: the retrieval and the reasoning have to happen inside the student for the student to have done anything. We should be careful about how hard we push that, because we have read summaries of this literature rather than the studies behind it, and a summary is always more confident than the studies it summarises. What we are confident about is narrower. We drew the line carefully. We drew it at the wrong granularity.
This is a problem a user has already handed us, in her own words. One of the two students who have written to us with feedback told us that the process of making a flashcard is itself a review and recall exercise, and that it slows the rate at which she forgets words. She was describing retrieval practice accurately without using the term. In the same paragraph she asked whether we could let her make flashcards with a single click, because the making is troublesome for somebody who is not very motivated. The effortful step is the step that works, and it is the step she would like removed. Our writing feedback has the same structure, with one difference: nobody asked us. We shipped the frictionless version by default, because supplying the answer was the obvious thing to build and because a student who is given the answer feels helped.
Feeling helped is not incidental. Ranalli, writing in the Journal of Second Language Writing in 2021 on second-language student engagement with automated feedback, made trust the mediating construct and reported engagement that is often shallow: accept or dismiss, with little consideration of why the suggestion was made. The finding cuts both ways. A learner who disbelieves the feedback ignores it, and a learner who believes it completely stops thinking about it. An interface where accepting takes one tap and understanding is optional selects for the second failure. Ours is such an interface.
What we control and what we do not measure
What we refuse is worth stating, because these are controls and not intentions. Asked what should happen when a student asks the app to rewrite their paragraph, our answer was two words: no, it cannot. There is no rewrite function. The writing feedback takes practice pieces only and not real assessed coursework, which is a fact about what the feature is built to accept. We should be honest about the size of that: it governs what the feature is for, not what a determined student could paste into a practice box. Both limits are in the product today. Both are limits on generation, and neither of them answers anything above. Refusing to write the paragraph does not make the correction of the paragraph durable.
The concession is this. We do not know whether our corrections stick, because we have never instrumented for it, and there was no technical reason we could not have. The measurement is not difficult. Take the features we flag and follow a student across submissions, counting how often a flagged error type reappears in the next piece. That gives a re-error rate per feature, and it is close to the only number that would tell us whether the writing feedback is teaching anybody anything. We do not compute it. We do not store what we would need in order to compute it well.
Truscott's methodological objection, which has held up better than his strong claim, is that revision accuracy is not acquisition: a student fixing an error you pointed at is not evidence that the student learned something. We are one step behind even that, because we do not measure the revision accuracy either.
The roleplay and the writing feedback disagree
There is an inconsistency inside our own product that we did not see until we wrote this down. The roleplay gives the least explicit hint that will work and escalates only as far as it has to, which is graduated mediation in the sense Poehner and Lantolf use it, and it is built and running. That is the indirect approach implemented properly, in one feature. In the writing feedback we supply the answer. Two parts of the same product hold opposite positions on the same question, and we arrived at both without arguing for either.
The obvious response is to make the writing feedback indirect, and we are not sure that survives contact with a student. Ranalli's point about trust is the reason. A student working at eleven at night who is told there is something wrong with the article in the second sentence, and then left to find it, may well close the app, and a feature nobody opens teaches nobody anything. That is the same trade the flashcard poses and we have not resolved it there either.
Underneath it sits a harder problem. The feedback the evidence says is durable is feedback about whether the paragraph does the job the section needs, and our writing feedback cannot see the assignment brief or the marking criteria, and does not know which module the piece was written for. Uploaded course material drives vocabulary extraction and does not reach the writing feedback at all. That is buildable and we have not built it. So the machine that would give the feedback worth giving is not the machine we have, and we do not yet know whether it is one we can build or only one we can describe.