We were asked what would be different about a seminar room at Durham if UniFluent worked perfectly. The answer we gave was that students would speak. They would have the academic vocabulary the discussion needed, and they would take part in it alongside the people sitting next to them. Nothing in that answer was about levels or scores. Nothing in it was about how much of our product anybody had used.
Later in the same set of questions we were asked whether confidence is a language app's job or whether claiming it is mission creep. The answer we gave then was that getting better at the language helps with confidence, but that confidence itself belongs to other people and to real contact, and that our job is to get them ready for those social interactions. We did not connect that answer to the first one at the time. It is a smaller claim than this market usually makes, and we would rather keep it small than widen it.
Then a third question, put more bluntly. If a student used the product for a term and still said nothing in seminars, would that count as a failure? Our answer was yes.
Those three answers were given at different points, to questions that were not about each other, and we did not notice until they were written down together that they define one outcome measure between them. Whether the student speaks.
A measure that can fail
It is worth being exact about what that excludes. It is not time in the app. It is not weekly retention. It is not the number of vocabulary items still available at four weeks, which we could measure and which would tell us about our spaced repetition rather than about a seminar. It is not a CEFR band, and we have written elsewhere about why the single number we currently display is accurate to about one level at best. All of those are things happening inside our software. The measure we have arrived at is a thing happening in a room we will never be in.
That is what makes it worth naming, and it is also the uncomfortable part. A measure that cannot fail is not a measure. Time in the app cannot really fail: if it falls we can say the term got busy, and if it rises we can say the students are committed. Lessons completed cannot fail either, because we control how long a lesson is. Whether a student speaks in a seminar can fail. It can fail after a term of heavy use by a student who liked the product and told us so. A pilot built around it could come back with a result we do not like, and we would then be holding evidence against our own product that we had asked somebody to collect.
The qualification we would insist on is about the starting position, and we think it is a fair one rather than a hedge. A student who arrives able to follow the discussion but not enter it, and who ends the term asking one question, has moved further than a student who was already contributing and now contributes more often. An absolute threshold would reward the second and record nothing for the first. So the honest version of the measure is distance travelled rather than a line everybody has to cross, which is harder to define and considerably harder to collect, and we would rather say that now than discover it halfway through a pilot.
What a dashboard is able to count
Set against that, consider what is easy. Time in the app, cards reviewed, lessons completed, streak length. Every one of those can be counted inside our own systems without asking any institution for anything, and every one of them is what a dashboard wants, because the numbers arrive continuously and they go up. None of them is the thing we said we were for. What they have in common is that they all measure how much our product is being used, which is a fact about our product rather than a fact about a student.
Karl Maton has a name for the habit of talking at length about learning and skills without ever examining the knowledge being taught. He calls it knowledge blindness (Knowledge and Knowers, Routledge, 2014). A dashboard is knowledge-blind by construction. It can report that four hundred cards were reviewed in a week and it cannot say what any of them were for. The question worth asking about a student is what they can now do that they could not do before, which is near to what Michael Young (2008) means by powerful knowledge as distinct from the knowledge that powerful people happen to hold. Engagement figures do not answer that. They do not even fail to answer it interestingly. They answer a different question, which is how sticky the software is.
The same worry is older than any dashboard, and it has been put more sharply inside EAP than we would put it. Moore and Morton (2004), in the Journal of English for Academic Purposes, compared IELTS writing tasks with the writing that degree study actually asks for and found the two less alike than the use made of the score implies. A benchmark can be useful and still not be the thing being taught. Ding and Campion (2016) press the point about provision rather than instruments, writing that a quick fix attitude has persisted particularly in pre-sessional courses and perpetuates a "maximum throughput of students with minimum attainment levels in the language in the shortest possible time" philosophy, a phrase they take from Turner (2004, page 97). Those arguments are about pre-sessional teaching and about IELTS rather than about anything of ours, and we are not going to enlarge either. What we take from them is the shape of the objection. If a field can ask that question about its own reporting, a vendor can certainly be asked it about a dashboard.
We do not measure this yet
Here is the concession, and it is a real one. We do not currently measure whether anybody speaks. We have no instrument for it, no agreement with any institution to collect it, and no baseline. Everything we do collect is a proxy for usage. Worse than that: when we were asked which metric we were most tempted to put on the dashboard and should not, we could not name one, and the metric we volunteered as the big one was retention. That is the answer of a company that had not thought about measurement hard enough, and it was our answer while these essays were being written.
There is a second thing in our own files that we should put next to it. We have said, in as many words, that you would hope more of our product equates to more learning. That is a hope, not a finding, and one of the two students who have written to us about the product has already supplied the counter-example. She had spent about two days with it. She wrote that the process of making a flashcard is the part that produces the memory, because making it works as recall practice and slows down forgetting. In the next sentence she asked us to remove that step and generate the cards in one click, because it is troublesome when she is not motivated.
Learning and usage pointed in opposite directions inside a single paragraph written by somebody with no stake in the argument and no vocabulary for what she was describing. One student over two days tells us nothing about how much anybody learns. It does show that the trade is real for at least one person, which is more than we had before. If we optimise against the metrics we presently collect, we take her second suggestion and lose the thing she told us was working.
The seminar is not ours to observe
The measure we prefer has problems of its own and they are not small. We cannot observe a seminar and we have no right to. No version of this gets collected without an institution doing the collecting and deciding it is worth the effort. Self-report by students is weak evidence and we would treat it as such. The confounds are heavy: the tutor, the composition of the group, the discipline, and whether the seminar in question rewards talking at all.
Some students are quiet in their first language and in every language, and treating speech as a virtue in itself would push us toward telling students what kind of person to be in a room. A product that does that becomes highly normative while sounding neutral, and it sends messages about how to behave that nobody involved would defend if they were written down plainly. The risk is heaviest across cultures, where when it is appropriate to speak and what it costs to disagree with a tutor in public are not settled the same way everywhere. A participation measure is exactly where that risk lives, and we would rather it were designed by people who have thought about it for longer than we have.
What we could not claim, and what we have not read
The last problem is attribution and it may be the one that sinks the whole thing. Suppose the number moves. We could not say we moved it. Students learn academic English over a degree whether or not anybody sells their university anything, and a term is long enough for a great deal else to happen.
There is also a tension we have not resolved between our own account of why students go silent and what their field says about it. Our position has been that silence follows from the language and from not having practised speaking, which is a diagnosis our product happens to treat. The seminar participation literature is mostly affective. The terms in it are willingness to communicate, anxiety, face, and what it costs to be heard getting something wrong in front of people you will sit next to again next week. We are reporting that at second hand. We have not read it, and we would rather say so than produce a citation we could not defend.
Our own strongest piece of material points the same way that literature does. We were told by a friend, a British student, that her seminar group turned and pointed at her to answer a question nobody had answered, without anybody saying anything first, and that as far as she could tell she was the only person at that table whose first language was English. Whatever was happening in that room, it was not a shortage of flashcards. The account also comes from the one person present for whom the language was not the difficulty, so we have the behaviour of that room and not the experience of it, and we have written elsewhere about how much that limits what we can take from it.
So there is a version of a pilot where this measure defeats us for reasons that have nothing to do with us, and a version where it defeats us for precisely the right reason and we learn that getting somebody ready for the interaction leaves the interaction unchanged. From where we sit we cannot tell those two apart in advance, and we have not settled which of them we should be more worried about.