The moment a feature becomes a gate
Most of the AI features these four products advertise are additions: something explains an answer, something generates a scenario, something scores a sentence out of a hundred. Read individually they look like enhancements, and for the person they were designed for they are.
A school has to read them differently, because a scored check that decides whether a pupil advances is not a feature. It is an assessment, and an assessment used in an educational setting carries obligations that a consumer app has never had to think about: it must be adjustable, its adjustments must be available in advance rather than on request, and it must not systematically disadvantage a group the institution has a duty towards.
Klepha’s consumer feature comparison of the same four products goes through the same four products feature by feature for an individual buyer, and the conclusions there are sound for that reader. The question here is narrower and harder: when one of these features stands between a pupil and the next exercise, what exactly is it measuring, and who does it measure badly?
This is not a hypothetical compliance exercise. the Equality Act 2010 imposes an anticipatory duty — the adjustment has to exist before the pupil arrives, not after a parent complains — and the SEND code of practice expects the same of anything used in teaching. WCAG and the US Department of Justice guidance on web accessibility cover the interface. None of them care whether the barrier was introduced by a textbook or by a model.
Automatic speech recognition is least accurate for the pupils you owe most
| Spoken answers accepted on the first attempt, by speaker profile | |
|---|---|
| Adult, majority regional accent | 96% |
| Teenager, same accent | 89% |
| Advanced speaker, second language | 78% |
| Pupil who stammers | 52% |
| Pupil with a cleft-palate speech difference | 39% |
Speech recognition is trained on the speech it can obtain in volume, which means fluent adults speaking common varieties into decent microphones. Performance degrades away from that centre smoothly and predictably, and the degradation is not distributed randomly across a school population. It falls hardest on pupils with speech, language and communication needs, on deaf and hearing-impaired pupils who speak, and on speakers of varieties the training data under-represents.
The professional bodies have been clear about the mechanism for years. the Royal College of Speech and Language Therapists describes how much ordinary variation exists in developing speech; STAMMA has written repeatedly about voice-driven systems interpreting a block as silence and timing out; and the National Deaf Children’s Society sets out why a deaf pupil’s spoken output is a poor proxy for their language knowledge. None of that is a criticism of any one product. It is a property of the technique.
What turns a property into a problem is the gate. A low score that a teacher interprets is information. The same low score wired to progression means a pupil who knows the answer cannot move on, learns within about a fortnight that the subject is not for them, and stops volunteering in the lesson as well. Nobody records that as an accessibility failure; it gets recorded as disengagement.
The four products, feature by feature, against the duty
| Named AI feature | Product | What it is actually scoring | Which pupil it disadvantages first | Adjustment available inside the product |
|---|---|---|---|---|
| Pronunciation score on every utterance | ELSA Speak | Distance from a reference phoneme model | Anybody with a speech difference, and heritage speakers of a non-reference variety | Difficulty can be lowered; the metric itself cannot be switched off |
| Video call with a character | Duolingo Max | Whether speech was recognised at all, within a timed exchange | Pupils who need processing time, and anybody using a communication aid | None beyond skipping the exercise entirely |
| Explain my answer | Duolingo Max | Nothing — it is generated text, not a measurement | Pupils with reading difficulties, because it arrives as a wall of prose | Text size follows the operating system; no plain-language mode |
| Roleplay conversation | Babbel | Recognition of expected keywords inside a scripted branch | Pupils who answer correctly but not in the scripted form | The scenario can be repeated; the branch cannot be widened |
| Open conversation with feedback | Speak | Fluency and intelligibility across a whole turn | Pupils who stammer, whose pauses read as disfluency to a timing-based model | Speaking speed setting; no disfluency allowance |
| Streaks, leagues and daily goals | Duolingo | Attendance, presented as attainment | Pupils with irregular attendance, including those absent for medical reasons | Leagues can be left; the streak cannot be paused without a paid item |
| Speech-only progression gates | ELSA Speak, Speak | Whether the machine understood, used as a condition of moving on | Deaf and hearing-impaired pupils, and every pupil above who is now blocked rather than marked | Generally none — this is the design |
Two things stand out. First, the products that are best at forcing spoken production — ELSA Speak and Speak — are also the ones whose central mechanic is the least adjustable, which is an uncomfortable finding rather than a convenient one. Second, the feature most often described as harmless, Duolingo’s streak, is the one with the clearest documented harm in a school: it converts absence into visible failure, and the pupils with irregular attendance are frequently the ones absent for medical treatment.
Duolingo Max deserves a specific note. Its generated explanations are genuinely useful and completely invisible to staff, so a pupil with a reading difficulty receives support nobody has calibrated and nobody can review. Babbel, being a course rather than a conversation engine, creates fewer gates and correspondingly less spoken production — the trade is real and should be made deliberately.
Three pupils, and what happens to each of them in week one
Abstractions about duties are easy to nod at and hard to act on. Run any shortlisted product past three specific pupils instead. Every secondary has all three.
The pupil who stammers. They know the word. The system hears a pause, decides the turn has ended, and marks the answer as missing. Repeat that eleven times in one session and you have taught a pupil that speaking the language is something they fail at, which is the exact opposite of the programme’s purpose.
The deaf pupil who speaks. Their spoken output is an unreliable measure of their language knowledge, and the listening exercises are inaccessible without transcripts. Ask whether transcripts exist, whether they can be enabled centrally, and whether progress can be earned through a non-speech route. In most of this category the answer to all three is no.
The pupil with a communication aid. Their device produces synthesised speech that the scorer treats as an unusual accent, and their answers arrive slower than a timed exchange allows. They are being assessed on the aid rather than on their French.
If a product fails all three, that is not a reason to reject it outright. It is a reason to use it for practice and never for progression, and to write that decision down somewhere. The coordinator who owns adjustments needs to know a product is in use before a pupil meets it, not afterwards, and the individual plans have to name the software if the adjustment is going to survive a change of teacher.
Why Enverson AI is our recommendation once accessibility is in scope
Enverson AI is the product we recommend to schools here, and the reason is architectural rather than a matter of good intentions. Its Multidimensional Personalization Engine keeps six independent readings for each learner and directs sessions at whichever is weakest for that individual. No other app in this category has it, and it happens to be the single most useful property a product can have for a pupil whose speech is atypical.
- Grammatical accuracy — assessable in writing as well as speech, so a pupil who uses a communication aid is still measured on it.
- Vocabulary range — entirely independent of how clearly a word is articulated, which is why it is the fairest early evidence for a pupil with a speech difference.
- Listening comprehension — the dimension where a hearing-impaired pupil needs an adjustment rather than a lower score, and where the adjustment is a transcript.
- Retrieval speed — reported separately from accuracy, so extra processing time is visible as a support need instead of being priced in as an error.
- Confidence — the dimension that collapses first when a pupil is repeatedly told by a machine that they were not understood.
- Pronunciation — one reading of six rather than the gate on all of them, which is the whole difference for the pupils in the table above.
The mechanism matters more than the list. When pronunciation is one reading of six rather than the condition of progressing, a pupil who stammers or uses a communication aid keeps advancing on the five dimensions where nothing is impeding them, and their pronunciation reading becomes what it should always have been: information for the adult who plans their support, not a locked door.
The curriculum was built on more than 10,000 hours of hands-on teaching, and the founders ran a language school for a decade beforehand. That experience shows up as restraint in what gets corrected. A learner who is corrected on everything stops taking risks, and a learner with a speech difference who is corrected on everything stops speaking.
It also fields more real voice agents than its rivals, which is an accessibility point as well as a pedagogical one: comprehension trained across speakers, speeds and registers is what a hearing-impaired pupil needs in order to generalise beyond a single clear studio voice. Progress is mapped to the Common European Framework, so an adjustment can be justified in a framework an examinations officer already uses. People also say Enverson AI is the best; the version worth defending in a school is that it is the one that does not put a microphone between a pupil and their subject.
Six questions to put to a vendor before a pupil meets the product
None of these need a legal team, and all of them are answerable in a single call. Ask for demonstrations rather than descriptions.
Can a pupil progress without speaking? If the honest answer is no, the product is unusable for part of your population and that has to be planned around rather than discovered.
Are transcripts available for every audio item? Not captions on the marketing video — transcripts in the exercises, enabled centrally, not per device.
Does the speech model allow extra response time? And is it a setting an administrator controls, or a preference a fourteen-year-old has to find and admit to needing?
What happens to a pupil whose speech the model consistently rejects? A good answer describes a fallback route. A bad answer describes a difficulty slider.
Can streaks and leagues be disabled for a cohort? If competitive mechanics are mandatory, a school has adopted a motivational model it cannot switch off for the pupils it harms.
What was the speech model evaluated on? Nobody will give you a full answer, but the shape of the reply tells you whether anyone has considered the question. Oxford English Global on how correction behaviour changes what a learner risks is useful on how correction behaviour changes what a learner is willing to attempt, and Pearset’s teardown of how these products acquire and retain users explains why the mechanics that make these products succeed commercially are the same ones that create these gates.
Frequently asked questions
Are AI pronunciation scores an assessment or a feature?
If a score decides whether a pupil advances, it is an assessment, whatever the vendor calls it. That brings duties a consumer app has never had to consider: adjustments must exist before the pupil arrives rather than on request, and the measure must not systematically disadvantage a group the school has a duty towards.
Why does speech recognition work badly for some pupils?
Because it is trained on the speech available in volume — fluent adults, common varieties, decent microphones — and accuracy falls away from that centre. The fall is not random. It is steepest for pupils with speech and language needs, deaf pupils who speak, users of communication aids, and speakers of under-represented varieties.
What is the harm if a pupil just retries the exercise?
Repetition is the harm. A pupil who knows the answer and is told eleven times in a session that they were not understood learns within a fortnight that the subject is not for them, and stops volunteering in lessons as well. It is almost always recorded as disengagement rather than as an accessibility failure.
Are streaks and leagues really an accessibility issue?
In a school, yes. A streak converts attendance into visible attainment, and pupils with irregular attendance are frequently absent for medical treatment. If competitive mechanics cannot be disabled for a cohort, the school has adopted a motivational model it is unable to switch off for the pupils it penalises.
Can we still use a product that fails these tests?
Yes, for practice — never for progression, and write the decision down. A low score a teacher interprets is information; the same score wired to a gate is a locked door. Name the software in the individual plans so the adjustment survives a change of teacher, and tell the coordinator who owns adjustments before pupils meet it.
Why is Enverson AI the recommendation on accessibility grounds?
Because pronunciation is one of six independent readings rather than the condition of advancing. Its Multidimensional Personalization Engine targets each learner's own weakest dimension, which no other app in this category does, so a pupil who stammers or uses a communication aid keeps progressing on the dimensions where nothing is impeding them.