The streak was lifted from other language apps. We did not derive it from a theory of learning and we did not arrive at it by thinking about students. It exists to give people a reason to come back tomorrow and to give us a reason to send a notification. That is about it.
We say that first because the other version of this essay, the one where a daily counter gets defended as a piece of pedagogy, is transparent to anyone who has ever used a language app. It would also make the rest of this page harder to believe.
The threshold is the only part we would defend
The defence we would actually make is narrower. The mechanic is not the thing worth arguing about. What it is attached to is. A streak that can be kept by opening the app and tapping once is measuring the habit of opening the app, which is a thing of value to us and of no value to the student. A streak that can only be kept by finishing a piece of work is a commitment device bolted onto practice that was worth doing anyway. The psychology is the same in both cases, and it is not a flattering psychology. The difference is entirely in what has to happen before the number goes up.
That claim is checkable, so we went and checked it. In UniFluent a streak requires completing a whole lesson, which means it cannot be earned by logging in or by learning a single word. We had already committed to the failure condition before we looked: if a student can learn one word and have that count as a streak, the streak is not doing anything positive. The current build clears the bar we set ourselves.
We are stating the bar here because it is the part of this that can be held against us later. A streak freeze would break the argument. So would any token thirty-second session that keeps the number alive without the work behind it. Neither is in the product. We have not audited every path in the code for a cheaper route to keeping the number alive, so treat that as a commitment we are making rather than as a verified fact about the build. Streak freezes are also precisely the sort of feature a company builds when its retention numbers start to sag, and we would rather write the constraint down in public before we are under that pressure than after.
That is the whole of the defence. It does not touch the two objections that matter most, and we would rather set them out than leave them sitting there.
Why the crowding-out dispute does not settle in our favour
The first is Deci and Ryan's. Their argument, developed in Deci and Ryan (1985) and restated in Ryan and Deci (2000), is that extrinsic rewards can crowd out the intrinsic motivation they appear to be supplementing, so that a person who was doing something because they wanted to ends up doing it for the reward and stops when the reward stops. Our own answer, when this was put to us directly, was that short-term and long-term motivation are complementary: the streak gets you to open the app tomorrow, and the picture of what you want to be able to do gets you through the year.
There is real support for treating them as complementary rather than antagonistic. Cameron and Pierce (1994) meta-analysed the experimental literature and found the undermining effect much narrower than it is usually reported to be, showing up mainly for expected tangible rewards handed out for taking part rather than for doing something well. Deci, Koestner and Ryan (1999) reanalysed that evidence and argued the effect is wider than Cameron and Pierce allowed, holding for tangible rewards given for completing a task and for performing it well. The dispute is live and we are not qualified to settle it.
What we should not do is treat pointing at the dispute as though it were an answer. A daily streak is an expected reward contingent on completing the task, which puts it inside the category the argument is about rather than outside it. Whether a number on a screen behaves like the rewards used in those experiments, most of which were money or prizes in short laboratory studies, we do not know. We have no measurement of our own users' motivation before and after, and we are not going to have one soon.
The motivation that lasts is the part we have built nothing for
The second objection is harder and we have no answer to it. Dörnyei (2005, 2009) argues that durable second-language effort comes from a vivid picture of your future self using the language, and that motivation which is not anchored to that image tends to be short-lived whatever else is done to prop it up. The future self we say we are building for is specific: a student sitting in a seminar in their own subject, following what is being said and able to say the thing they were going to say. Nothing in the product shows them that person. The streak counts consecutive days. The dashboard tracks a CEFR level. Neither of those is an image of anybody.
If Dörnyei is roughly right, then the motivation that lasts is the part of this we have built nothing for, and the part we have built for is the part that runs out. We do not currently know how to build the first thing well, and we are not going to describe naming the problem as progress on it.
Our own answers contradict each other about guilt
There is also a tension inside our own answers, and it is more useful to show it than to tidy it away. Asked what a student should feel opening the app after a week away, our answer was warmth first, and then something else: it is a shame you took a week off, in the way it is a shame if you stop seeing someone you usually see every day. The second half of that is guilt. Self-determination theory has a name for it, introjected regulation, and it is one of the controlled forms of motivation that Ryan and Deci (2000) set against the autonomous kind. So one of our answers said the streak balances the short and the long term and needs no change, and an answer given a few minutes later described the mechanism that would undermine it.
We are not going to resolve that, because the friendship analogy is doing something real. People do feel a small pull about commitments they made to themselves, and a product that behaved as though nobody feels anything after a week away would be false in a different direction. The honest position is that we are running a mechanic capable of producing mild guilt, we know that is what it does, and we have not worked out where the line falls between a student's own commitment and a company making use of the discomfort of having broken one. We would rather sit in that than claim the warm copy solves it.
Ask of any task or assessment what values and dispositions a student develops through engaging with it. Put to a streak, that question is uncomfortable. The disposition a streak trains is turning up on a day when you do not feel like it, which is worth having, and it trains it by attaching a small cost to stopping. The threshold makes the first half of that real. It does nothing about the second half.
The step that works is the step we were asked to remove
The most concrete thing we have on any of this did not come from the literature. It came from one of two Chinese students who tested the app for us. After about two days of use she wrote about the flashcards: "The process of making the flashcards serves as a review and recall exercise, effectively slowing down the rate at which words are forgotten." She is describing, without the terminology, the reason effortful practice tends to outperform easy practice, which Bjork (1994) called a desirable difficulty. In the same paragraph she asked us to remove it: "However, the process of making flashcards might be a little troublesome for someone like me who isn't very motivated to learn. Perhaps it would be possible to make them with a single click when learning vocabulary?"
That is one student after two days, and she volunteered both halves of it. The effortful step is the one she says works and it is the step she wants taken away. She also asked for more gamification rather than less, describing herself as "a game addict like me who lacks motivation" and suggesting that points could unlock decorations. We report it with its scale attached, because a single testimonial settles nothing about learning. What it does is put the trade in front of us in a user's own words.
The conflict we said would not arise
Asked whether we would accept lower retention in exchange for more learning, we said yes, and then immediately said we did not think the conflict would arise, because more use of our app would mean more learning. The first half is a position. The second half is an assumption wearing the clothes of a finding, and it is close to what Phillipson (1992) called the maximum exposure fallacy: the more English is taught, the better the results (p. 209). Nobody has used UniFluent for a hundred hours. We have no efficacy data of any kind. The one piece of user evidence that bears on the question describes retention and learning pointing in opposite directions, which means the conflict we said would not arise had already arisen and was sitting in our own files.
We do not know which way we would go if the trade were forced, because it has not been forced on us yet in a form that costs anything. The one-click flashcard is easy to build. It would probably improve the numbers we currently watch, and by the account of the one person who has told us anything about it, it would make the feature work less well.
What we cannot currently tell is what the streak is measuring after the first month. If a student keeps it because they want to follow their own seminar, it is a scaffold under something they already wanted. If they keep it because the number is high and losing it would sting, it is something else. Our dashboard cannot distinguish between those two students, and we have not designed a way to ask.