A card in our product reads mitigate on the front and to make less severe on the back. A student turns it over, gets it right, and the card goes back into the queue at a longer interval. Suppose the scheduling does what it is designed to do and the pairing is still available a month later. Asked what that student would then have, the most accurate description we have found is our own: it is just word association to an extent. They can pair a word with a gloss. They cannot tell you what usually follows it, or whether it belongs in a methods section or a discussion section. The card addresses the easiest of the several things involved in knowing a word and leaves the others alone.
The tradition that best explains why the spacing on that card works is the same tradition that says the card is the wrong shape. Usage-based accounts of language, associated with Nick Ellis and Michael Tomasello, describe what a speaker knows as an inventory of constructions: form-meaning pairings at any size, from a morpheme up to a clause frame, abstracted statistically from encounters with actual use. On that account spacing sits close to the mechanism of learning rather than being a study trick bolted onto it, and spacing does replicate in second language vocabulary specifically (Kim and Webb, 2022, in Language Learning). The same account is unkind to our unit. If the thing being learned is the construction, then mitigate on its own has had the learnable part stripped out of it. The front of the card should read mitigate the effects of.
We accepted that argument, and accepted it quickly, which is usually a sign that something was already half believed. A chunk carries some of its context with it. It arrives closer to the ground than a bare definition does, and it makes the thing being stored resemble the thing the student will eventually have to do. Karl Maton's Knowledge and Knowers (Routledge, 2014) sets segmental learning, which arrives in disconnected pieces and never joins up, against cumulative learning, where what is learned later builds on what came before. A deck of single words paired with single glosses is close to the purest available form of the segmental version.
It would be overstating things to say that switching the unit fixes it. Teaching a chunk properly still requires bringing it down to a concrete instance in the student's own subject and taking it back up to the general pattern, and no card does that by itself. The unit is where the problem starts. It is not where it ends. We still had the unit wrong.
Why the real cost is not engineering
Asked what changing the unit would cost, our answer was that it would be hard but we could probably do it, and that it would be nothing like other language apps. The second half of that sentence is the real cost, and it is worth stating carefully, because it is not an engineering estimate.
A word on the front and a definition on the back is what a language app looks like. Somebody who has used Duolingo or Quizlet knows within about three seconds what kind of object they are holding, and that recognition does an enormous amount of work in a consumer market. A deck built from multi-word units does not deliver it. The item becomes a frame with slots in it. The count of things learned this week stops being a clean number. What is left does not look like the category it is sitting in, and something that does not look like the category is difficult to sell to a person deciding in three seconds.
The calculation changes when the buyer changes. A university considering a licence is not deciding in three seconds, and the objection that would actually block the purchase comes from academic staff who distrust generic bought-in material. Basil Bernstein's name for what that distrust is about is recontextualisation: knowledge is selected out of the place where it was produced and reshaped on its way into a course, and the reshaping is never neutral. A department that buys material is buying somebody else's selection, and a generalist textbook is the most conspicuous form of that. In that room, failing to resemble a consumer language app is closer to a qualification than a defect. The cost we named is real, and it is denominated in a market we are on our way out of.
Where the flashcard actually came from
The uncomfortable part comes next, and it is uncomfortable in the direction of our own conduct rather than the market's. Lea and Street (1998), in Studies in Higher Education, argued that academic literacy is plural and discipline-specific rather than one transferable skill that can be taught once and carried anywhere, and Wingate (2015) made the case against teaching it as a set of generic techniques detached from a subject. Mainstream EAP coursebook design has mostly not followed either of them, and the reason is at least partly commercial: material written for one discipline sells into a fraction of the market that general material sells into. That is an account of publishers. It describes us precisely.
We did not derive the flashcard from a theory of language and then build it. We inherited it from the category we were building inside, because that is what the products in the category do, and we asked what it was for afterwards. The same is true of the streak, which we lifted from other language apps more or less outright and have defended since on a threshold we imposed later. Our unit is not a position we arrived at and it is not one we defended. It is the shape the market was already in when we got here. On that account of how materials come to look the way they look, we are an instance and not an exception, and we would sooner say so than have it inferred.
What is built and what is not
We have not changed the unit. That needs saying plainly, because the alternative is a sentence that reads as though we had. The decision has been taken in principle and it has not been taken in code. The product today stores the word. Extraction returns single orthographic words, which is exactly where multi-word academic meaning goes missing. Our topic decks also group semantically related words into the same set, which Nakata and Suzuki (2018) found impedes learning instead of assisting it, so the unit is not the only thing we took from the category without examining it. We are not attaching a date to any of this. A dated roadmap item would be easy to write and would be worth less than the admission, since the only person who would hold us to it is us.
One part of the vocabulary side we would keep unchanged. The words come from the student's own uploaded course material and not from a general list, which follows from Hyland and Tse's 2007 finding in TESOL Quarterly that academic vocabulary behaves differently from one discipline to another and that general lists mislead. That finding is contested. Durrant (2016) and Gardner and Davies have argued the other side of it in English for Specific Purposes, so it is a position we hold and not a settled matter we are reporting. It is also the only part of the product the student's own material currently reaches. It drives the vocabulary extraction and it does not yet feed the voice tutor or the writing tutor. That is buildable and it is not built, and it belongs in this essay because a better unit fed by the wrong corpus would not be much of an improvement.
What a better unit would still not do
There is a limit past which none of this helps, and we stated it before we understood it as a limit. A student who does everything we ask will still need some real world experience to pair that with, and that is just the way it is. Changing the unit from the word to the chunk moves what we store closer to what gets used. It does not turn a queue of cards into a seminar, and we have no evidence about what it would do, because nobody has used the product for a hundred hours and we have no efficacy data of any kind.
What we cannot answer is whether the change survives contact with the student. One of the two students who have written to us with feedback used the product for two days and told us that the effortful part of making a flashcard was the part that worked for her. In the next sentence she asked whether we could remove that effort with one-click creation. Two students over two days is not evidence of anything, and we give the scale for that reason. She was right about both halves, and a chunk-based deck makes the trade sharper instead of easier, because a chunk takes longer to enter and longer to review than a word does.
Students also count words. Forty words this week is legible in a way that forty constructions is not, and legible progress is a large part of why anyone opens the app on a Tuesday. So the open question is not whether the chunk is the better unit. It is whether a student trained by every other product to count words will accept one that counts something else, and whether we would hold the position if it turned out they did not.