Nobody has used UniFluent for a hundred hours.
We were asked what would happen to a student who did, with our product and nothing else, and the truthful answer was that we could not say. What we offered instead was a guess. They would probably know a good number of words, and have some sense of how certain conversations tend to go, but the English would be formulaic rather than contextual, and they would still need real experience of using it alongside. That is our own estimate of our own ceiling. It is not a finding. Nobody has done the hundred hours, so nobody has checked whether the guess runs optimistic or pessimistic.
So we have no efficacy data. People often say that when they mean thin data, or early data. We mean none. Every sentence we have written that says our product improves a person's English is, as of today, supported by nothing except our expectation of it.
Why we said inference was enough
When the statement was put to us that a company which cannot show evidence should not claim an effect, our answer was partial. True to some degree, but it is very hard to show conclusive evidence about anything, so if something is inferable then we can claim it.
We still think the first half of that is correct, and we want to defend it before conceding the rest. Conclusive evidence about a teaching intervention is close to unobtainable. Effect sizes in education are small and unstable, populations differ, the thing being measured is usually a proxy for the thing anybody cares about, and randomised designs are often impossible or unethical inside a real university. A rule that no claim may be made without conclusive evidence sounds strict. Applied consistently it forbids everything, including a great deal of what universities say about their own teaching. Nobody operates that standard, and a company that said it did would be lying about a second thing.
What that standard produced on our own website
The trouble is with the other half, and the demonstration is on our own website.
It currently says the product will "adapt to your learning style". The idea underneath that phrase, that a student is a visual learner or an auditory one and does better when taught in the matching mode, is among the most examined and least supported in education research. Pashler, McDaniel, Rohrer and Bjork set out in Psychological Science in the Public Interest in 2008 that the evidence required to support it barely exists, and that where the right studies had been run they came out against it. The site also says "proven educational methods". "Proven" is a word academics apply to almost nothing, and a claim that something has been proven is precisely the sort of claim that gets checked.
Nobody at this company set out to mislead anybody with either sentence. Both were inferable. The roleplay does respond to the individual user, so adapting to how they learn felt like a fair summary. Spacing has a real evidence base behind it, and Kim and Webb (2022) in Language Learning is the meta-analysis for second language learning specifically, so "proven" felt like a fair summary too. The standard we described is the standard that produced both. That is the strongest argument against it that we can think of, and it is ours rather than anybody else's.
Separating a capability from an effect
What follows is a proposal rather than a settled policy. We have not tested it, and we have not finished arguing about it between ourselves. We would like it argued with from outside as well.
The proposal is to separate claiming a capability from claiming an effect, and to hold the two to different standards.
A capability claim describes what the product does. "Builds vocabulary lists from the course material you upload" is a capability claim, and it needs no efficacy evidence, because anybody can open the product and look. So is a claim that the roleplay offers the least explicit hint it can and escalates only when the student does not take it. That second one is Dynamic Assessment in the sense Poehner and Lantolf use the term, it is built, and describing it costs us nothing in evidence, because the evidence is the software.
We should add that we did not build it from the theory. We built it because it seemed like the decent way to treat somebody who is stuck, and we learned afterwards that it had a name and a literature behind it. Saying so matters, because otherwise a fortunate design decision gets read as a research programme.
Even inside the capability line it is easy to cross. We had written that the correction standard can be set to British academic English, on the grounds that a setting is a setting and anybody can check it. Then we checked it properly. Three options are selectable, and we are not currently certain whether choosing British academic English changes the standard a piece of writing is corrected against or only the register the feedback is written in. That distinction is the whole of the claim. Until we can answer it, the sentence is a guess about our own software rather than a description of it, which is a fairly humiliating thing to discover in the middle of writing this.
Effect claims, and what an observation carries instead
An effect claim says the product changes something about the person. "Improves your academic English" is an effect claim. So is any promise about marks or about taking part in a seminar. Those need evidence about outcomes, we have none, and we should not be making them until we do.
There is a third kind of claim, and it is the one we have been neglecting. An observation reported with its scale attached claims exactly what it can carry and no more. We have two of them. Two students have written to us about the product. One used it for about two days and said the roleplay was the most useful part, because Chinese students, in her account, remember and spell words without knowing how to use them. The other uploaded a slide deck from her course and said the questions it generated helped her notice points she had missed. The second one is the interesting one, because the benefit she describes is not a language benefit at all. She understood her own subject better. The product does not currently claim that anywhere.
Two students is two students. It does not generalise, and one of them had been using it for two days when she wrote. But an observation with its scale attached is honest in a way a general effect claim never is, because the reader can see how much weight to put on it. "Thirty students used it for a term and this is what they told us" is a sentence we could earn within a year.
We should say plainly that the capability line does not rescue everything. "Adapt to your learning style" was written as a description of a capability, and it is still wrong, because the construct it names does not hold up. A capability claim has to be true about the product, and it also has to avoid smuggling in a theory the field has already discarded. We suspect a fair amount of our copy describes a capability in language that quietly implies an effect, and we have not worked out how to write around that.
Half of what we said about being tested was wrong
Then there is the pilot, which is the part of our own position we most need to correct.
Asked what we would do if the research said our core feature did not work, we said we would not discard it merely because theory research said so. We would run our own study, and in the long run we would look at our own data. Half of that is right and we want to keep it. General findings about flashcards do not settle what our particular implementation does with a particular set of students on a particular course.
The naive form of the anti-flashcard argument, that deliberately learned vocabulary stays inert and never becomes usable knowledge, is also not what the field found. Elgort (2011), in Language Learning, showed deliberately learned words producing both form priming and semantic priming, which is evidence that they become real lexical knowledge. Critics of flashcards routinely leave that out. Insisting on being tested rather than assumed to fail is a reasonable position to hold.
The other half is a problem. Calling the relevant work theory research misdescribes it. The vocabulary and feedback literatures are largely empirical and often meta-analytic, and the Kim and Webb paper we just used to defend ourselves is itself a meta-analysis. More importantly, we had the reliability ordering upside down. A vendor-run pilot, small, without a control group, designed by the people selling the product, with a sale riding on the result, is close to the weakest evidence design that exists. It is not a trump card over a meta-analysis. Our version amounts to a refusal to be evidenced, and it would be read that way, correctly.
Who should choose the outcome measure
This is why a pilot should be designed with the institution that would run it rather than presented to it. If we choose the outcome measure ourselves we will choose one we pass. That is the reason, although it also happens to be a better way to sell something. What a pilot measures should follow from what we claim the product is for, and we have said elsewhere that what we care about is whether a student can take part in a seminar. A study measuring vocabulary retention would establish something we already half know and would leave the claim we actually want to make untouched.
Moore and Morton (2004) found that IELTS writing tasks and the assignments universities set make different demands. Turner (2004), quoted by Ding and Campion (2016), described a quick fix philosophy of "maximum throughput of students with minimum attainment levels in the language in the shortest possible time". Both are arguments about who chooses the measure and what it is allowed to stand for, and we would rather not be answering them on our own behalf.
What saying less costs
The concession is this. We have been claiming more than we can support, in writing, on a public website read by the people we are asking for money. That is the standard we described being applied consistently, which makes it a position rather than a slip. Waiting for a better study does not help, because a study takes a year and the sentences are up now. The only thing available today is saying less until there is one, and that means the rebuilt site carries fewer promises than the current one, which makes it a worse sales page.
Those sentences are on there because they work. We are proposing to take them off with nothing to put in their place except a description of what the software does, and we are aware that is the version of the page a competitor will out-promise.
What we cannot yet answer is where the line falls in the ordinary cases. A page that describes a feature at all implies it is worth using, and implying it is worth using sits very close to claiming it works. We do not know how to write a product page that avoids that entirely, and we are not sure one exists. There is also a question we cannot settle from inside the company, which is whether inviting a university to design the pilot is genuine collaboration or a more sophisticated version of the same sales move, given who benefits if it goes well. We would rather be told which of those it looks like from the other side of the table.