CAPT

On-prem / on-device pronunciation assessment: scores how accurately a speaker pronounced a known piece of text.

CAPT (Computer-Aided Pronunciation Training) answers a different question from a speech recognizer. Transcribe is given audio and asked what was said. CAPT is given audio and the text the speaker was supposed to say, and asked how closely the audio matches it, phoneme by phoneme.

That difference matters whenever you already know the target. Reading assessment, pronunciation practice, language learning drills and scripted-prompt verification all share the same shape: a known reference, a spoken attempt, and a decision about whether the attempt was good enough.

Because the reference is known, CAPT reports where the attempt diverged, not just that it did:

  • Every phoneme of the reference is aligned to the audio and labeled MATCH, SUBSTITUTION, DELETION or INSERTION.
  • Every aligned phoneme carries a score in [0, 1], derived from the acoustic model’s confidence and remapped through per-phoneme thresholds you control.
  • Every recognized phoneme carries a start time and duration, so results can be laid over a waveform.
  • Words are scored as groups of phonemes, and the utterance gets a single overall score.

Crucially, CAPT is error-preserving. A general-purpose recognizer (and a large end-to-end model especially) is built to recover the intended message, so it will quietly repair a mispronunciation into the word it assumes you meant. That behavior is exactly wrong for assessment. CAPT never substitutes what it expects for what it heard: if the speaker said something other than the reference, the divergence is reported.

The engine runs on your own hardware: on-premise, in your private cloud, or embedded on a device. Audio never leaves your infrastructure.

How it works

CAPT is built on top of Cobalt Transcribe, running a phoneme-level acoustic model:

  1. The reference text is phonemized. Each word is looked up in a pronunciation lexicon; words that are not in it are passed to a grapheme-to-phoneme (G2P) model. A word may legitimately have several valid pronunciations, and all of them are considered.
  2. The audio is recognized into a confusion network: for each slice of time, a probability distribution over the phonemes that might have been spoken, rather than a single guess.
  3. Reference and audio are aligned with a weighted edit distance, producing the per-phoneme MATCH / SUBSTITUTION / DELETION / INSERTION labels.
  4. Each phoneme is scored from its confidence in the confusion network, remapped through the cutoffs you configure.

Keeping the full distribution rather than a single best guess is what makes step 4 meaningful: the score for a phoneme reflects how much probability mass the acoustic model actually put on it.

Where to start