Our dashboard says "Track your CEFR level progress". Five words. The first four assert that there is a thing called your level and that we have found out what yours is. The fifth asks a student to watch it move. We are most exposed on the claim to have found it, we are least confident about the invitation to watch, and the sentence is still on the product today.
The scale was assembled from teacher judgements
Start with where the levels came from, because almost nobody in this market knows and it changes what the number can honestly be asked to do. The illustrative descriptors in the Common European Framework of Reference came out of Brian North's Swiss project in the 1990s. Around two thousand descriptors were harvested from existing scales, sorted by teachers, then scaled from those teacher judgements, with the cut scores placed so that the resulting bands came out equidistant.
Glenn Fulcher's verdict, in "Deluded by artifices? The Common European Framework and harmonization", Language Assessment Quarterly 1(4), 2004, is that the descriptors were selected and assembled on psychometric grounds from data that were "intuitive teacher judgments rather than samples of performance". In a 2010 paper on reification he calls the results Frankenstein scales. Jan Hulstijn, writing in the Modern Language Journal 91(4) in 2007 under the title "The shaky ground beneath the CEFR", puts it more flatly: "the CEFR is not based on empirical evidence taken from L2-learner data", and there are no longitudinal studies showing that learners at any functional level other than A1 arrived there by passing through the level below.
None of that makes the framework useless. What it makes it is a careful description agreed among people who teach languages, which is a real thing to have, and a different thing from a ruler. The equidistance of the bands was a decision rather than a finding, which is why "you are twelve per cent through B2" is not a quantity and should never appear on a progress bar. This is roughly where we already sit. Our own position, stated four separate times while these essays were being put together, is that the levels are not real things, that they were imposed by people, that they are useful anyway, that a student's actual ability and their assigned level can differ quite a lot, and that the difference is something that just has to be acknowledged.
The problem is that acknowledging it is not what our dashboard does.
Our own estimate has never been checked
The weakest link in that sentence is not the CEFR at all. It is us. Our level estimate is generated by a language model. It has never been calibrated against a labelled sample and we have never measured how often it agrees with a human rater, so everything we can honestly say about its accuracy comes from other people's numbers.
Those numbers are not encouraging. UniversalCEFR, presented at EMNLP in 2025, is the largest study of automatic CEFR classification done so far, with 505,807 labelled texts across thirteen languages, and the best average weighted F1 for genuine six-way classification runs at about 60 to 63 per cent. Prompting a language model with the band descriptors, which is the cheapest method and the one closest to what an app like ours would reach for first, was the weakest method tested, averaging 34.2 against a most-frequent-class baseline of 19.3. Duolingo's own researchers found that with no calibration examples, neither GPT-3.5 nor GPT-4 "even outperform the baseline classifier using character length only". Prediction is worst at the C levels, and the C levels are where university students sit. Tack and colleagues reported in 2017 an exact accuracy of 53 per cent alongside an adjacent accuracy of 98 per cent, and reporting only the second of those two numbers is the standard move in this industry.
There is a further absence worth naming, because procurement will eventually ask about it. A real linking study to the CEFR has five stages: familiarisation, specification, standardisation, standard setting and empirical validation. We have done none of them. Fulcher's 2004 description of what providers do instead is that most "are prepared to make an intuitive guess and print this in guides and on web sites", and we are not going to pretend we are the exception.
So the honest description of our indicator is that it is a coarse orientation, good to about one band if it behaves like the systems that have actually been evaluated, and unmeasured if it does not. That is a much smaller claim than "track your progress". We have not changed the sentence. We already thought it was not ideal before any of this reading, and it is still there, which is the sort of thing that is easier to write in an essay than to explain in a meeting.
Being the standard for language apps is not a defence
There is a second admission underneath the first. When we were asked whether it was acceptable to show a number accurate to about one band, part of our answer was that this is completely the standard for language apps. That is true and it is worthless as a defence.
Basil Bernstein's argument in Pedagogy, Symbolic Control and Identity (2000) is that knowledge does not travel unchanged from where it is made to where it is taught. It gets selected and reordered on the way, and whoever does the selecting has settled what will count. Bernstein calls that recontextualisation, and his position is that it is never neutral. He was writing about pedagogic discourse rather than about software, so putting a progress bar under that description is our move and not his. The stretch is short. A band on a dashboard selects one description of a student's language out of every description available and puts it where the student and the tutor will look. When a department buys that in rather than building it, the selection was made somewhere else, by us, partly on the grounds of what is cheap to compute. Saying that everybody in our market does it this way then amounts to saying the selection is sound because every seller made the same one. We would rather retire the defence in writing than have it retired for us.
A scale of our own would be worse
Having said all that, we want to make the strongest case we can for keeping the framework, because we believe it and because dropping it would be the wrong response to everything above. Any metric we invented would inherit the same problems with less work behind it. People have tried very hard to make the best scale possible, and a homegrown fluency score would be a single number with less external referent than a six-band scale that at least has published descriptors and twenty-five years of use behind it. It would also invite exactly the reification Fulcher objects to, only with nobody outside our company having ever looked at it.
North's rebuttal to the critics, "Trolls, unicorns and the CEFR" in the CEFR Journal in 2020, is worth reading against everything above, and his defence of the framework as a descriptive tool rather than a measurement instrument is also what we think its value is. It lets a university and a student's previous language school mean approximately the same thing by B2. That is a genuine service and no rival scale of ours could perform it.
What a single letter hides
Keep the framework, then. Drop the compression.
The two things are separable, and the framework's own custodians separate them more carefully than we have. The Council of Europe states on its own site that "it is not the role of the Council of Europe to verify and validate the quality of the link between language examinations or diplomas and the CEFR's proficiency levels", and that the CEFR "does not tell practitioners what to do, or how to do it". Every claim of CEFR alignment in our market, ours and everyone else's, is unaudited by design.
The 2020 Companion Volume added mediation and plurilingual competence, which are ways of describing what a person can do with the languages they have rather than how far up a single ladder they have climbed. Krumm's contribution to that same 2007 exchange in the Modern Language Journal is titled "Profiles Instead of Levels", and the reason is straightforward: there is no evidence that a learner who is B2 overall is simultaneously B2 on each of the scales underneath it, vocabulary range and phonological control and the rest. Real students are uneven. A single letter is the one representation guaranteed to hide it.
Showing a profile across skills, with the uncertainty stated, is therefore closer to what the CEFR asks for than a single band is. So is replacing the band with can-do statements drawn from the student's own course material, so that what comes back is that you can now follow an extended lecture in your own subject. We have not built that. Uploaded course material currently drives vocabulary extraction and nothing else, and a discipline-specific can-do profile is a piece of work we have described rather than shipped.
We think this is the same objection we have already made somewhere else. Asked what we would say if a university wanted UniFluent output fed into a student's grade, our answer was that it would not make any sense, and the reason given was that incorporating everything into one grade is just pointless. Lots of different measures around lots of different things, possibly. One number, no. That is the identical complaint: compression destroys the information that mattered. It appears to be an instinct we hold about assessment generally, and having held it about grades we cannot coherently abandon it about our own dashboard.
When an external proxy starts to define the work
One caution about borrowing an argument that is not ours. Moore and Morton set IELTS writing tasks beside the writing university students are actually assigned, in "Dimensions of difference: a comparison of university writing and IELTS writing", Journal of English for Academic Purposes 3(1), 2004, and found the test task closer to a piece of public commentary than to the assignments it is used to predict. That is a finding about IELTS writing tasks. It is not a finding about the CEFR, and we are not going to convert it into one.
What carries across is only the shape of the worry. An external proxy gets reported because it is legible to everybody, and once it is the thing being reported it starts to set what the provision is for. Turner named that pattern in "Language as academic purpose", Journal of English for Academic Purposes 3(2), 2004, as a philosophy of "maximum throughput of students with minimum attainment levels in the language in the shortest possible time". Ding and Campion (2016) observe that the attitude behind it has persisted, particularly on pre-sessional courses. A company selling a dashboard is in a much worse position to resist any of that than a university department is.
The change may be smaller than the error
Which leaves the thing we cannot get past. If our indicator is no better than the systems that have been evaluated, it is good to about one band, and a pre-sessional course runs over a single summer. The change a student actually makes in that time is probably smaller than the error on the instrument reporting it. Honest reporting would show a profile that does not visibly move, on a page whose entire purpose is to show movement. We do not know whether a student finding that page useful, or closing it and going back to a number that lies to them more pleasantly, is the more likely outcome. We also do not know what belongs in the space where the number used to be.