Interpreting Results
This is the page worth reading carefully. Getting audio into CAPT is straightforward; deciding what its output means for your application is where the design work is.
The shape of a result
An EvaluationResult is a tree:
EvaluationResult
├── score overall score for the utterance, 0..1
├── is_partial true while more audio may still change this
├── is_speech_endpoint true if the speaker appears to have finished
├── alignments [ ]AlignedWord
│ └── AlignedWord
│ ├── text the reference word
│ ├── start_time_ms when the word starts in the audio
│ ├── duration_ms how long it lasts
│ └── tokens [ ]AlignedToken (the phonemes)
│ └── AlignedToken
│ ├── kind MATCH | SUBSTITUTION | DELETION | INSERTION
│ ├── reference the expected phoneme
│ ├── score 0..1 for this phoneme
│ └── hypotheses [ ]AlignmentHypothesis (what was heard)
│ └── token, confidence, start_time_ms, duration_ms
└── alternative_alignments the same, once per alternative reference text
The reference text drives the structure: your words appear as AlignedWord
entries in order, and within each one there is an AlignedToken per expected
phoneme.
alignments is not a one-to-one list of your reference words. Insertions
are reported as extra AlignedWord entries with an empty text, positioned
where they were heard, so the list can be longer than
reference_text.split() and the indices do not line up. Match on text, or
skip entries whose text is empty, rather than zipping the two lists together.
See Insertions.
The four alignment kinds
Every expected phoneme gets exactly one of four verdicts.
| Kind | Meaning | reference | hypotheses | score |
|---|---|---|---|---|
MATCH | The expected phoneme was heard with enough confidence | set | populated | 0.5 to 1.0 |
SUBSTITUTION | The expected phoneme was not heard with enough confidence: either something else was heard in its place, or the phoneme itself was heard but below its cutoff | set | populated | 0 to below 0.5 |
DELETION | The phoneme was not heard at all | set | empty | 0 |
INSERTION | Something extra was heard that the reference does not account for | empty | populated | 0 |
A SUBSTITUTION does not imply a score of 0. When the expected phoneme
was heard but fell short of its cutoff, the score is its remapped confidence,
which is non-zero but below 0.5. The k scoring 0.49 in the example below
is exactly this case. Only when the phoneme is absent from the hypotheses
entirely does the score reach 0. Decide on kind, or on a threshold, but do
not infer one from the other.
A real example makes this concrete. Below is the result of scoring a recording of “when the sunlight strikes” against a deliberately wrong reference, “WHEN THE MOONLIGHT STRIKES” (phonemes are X-SAMPA):
overall score: 0.5157
WHEN MATCH w (1.00) MATCH E (0.99) MATCH n (1.00)
THE MATCH D (0.87) MATCH @ (0.67)
MOONLIGHT SUB m (0.00) SUB u (0.00) MATCH n (0.65)
MATCH l (0.78) MATCH aI (0.82) MATCH t (0.73)
STRIKES MATCH s (0.79) DEL t (0.00) DEL r\ (0.00)
DEL aI (0.00) SUB k (0.49) DEL s (0.00)
Read that as a diagnosis, not just a number. The first two words were said as
expected. In MOONLIGHT the leading m u was not there (the speaker said
s V n), but the shared tail n l aI t matched, so those phonemes score well.
STRIKES shows what happens when the audio has already run out: the reference
still has phonemes to account for, and they come back as deletions.
The overall score of 0.52 is the average of that mixture. This is why an
overall score alone is rarely the right thing to threshold. 0.52 here does
not mean “roughly half right”, it means “two words right and two words wrong”.
Insertions
Insertions do not belong to any reference word, so they are reported in their
own AlignedWord entries with an empty text, positioned in the sequence
where they were heard. Scoring the same audio against the single word
“ELEPHANT” shows this:
overall score: 0.3966
(insertion) INS (heard: w)
ELEPHANT MATCH E (0.99) SUB l (0.00) MATCH @ (0.67)
SUB f (0.00) SUB n= (0.00) MATCH t (0.73)
(insertion) INS (heard: k) INS (heard: s)
The speaker said far more than the reference accounted for, and the leftover audio at each end is reported as insertions rather than being silently discarded.
Whether phonemes inserted within a word are attached to that word or reported
separately is controlled by the model’s IncludeIntraWordInsertion setting.
See Tuning and Customization.
Where the numbers come from
Phoneme scores
The acoustic model does not commit to a single phoneme per time slice. It produces a distribution, a confusion network, and the score for an expected phoneme is derived from how much probability mass landed on it.
You can see this in the hypotheses list. Here is a single MATCH from a
real result:
{
"kind": "ALIGNMENT_KIND_MATCH",
"reference": "n",
"score": 0.9798913,
"hypotheses": [
{ "token": "n", "confidence": 0.963, "start_time_ms": "570", "duration_ms": "119" },
{ "token": "m", "confidence": 0.029, "start_time_ms": "570", "duration_ms": "119" },
{ "token": "n=", "confidence": 0.006, "start_time_ms": "570", "duration_ms": "119" },
{ "token": "N", "confidence": 0.002, "start_time_ms": "570", "duration_ms": "119" }
]
}
The model heard n with 0.963 confidence, but also considered m, n= and
N. The reported score of 0.98 is that confidence remapped through the
model’s cutoff for n. See Tuning and
Customization
for how that remapping works and how to change it.
The practical consequence: a score is a calibrated quantity, not a raw
probability. A cutoff is chosen so that a score of exactly 0.5 sits at the
boundary between acceptable and unacceptable for that phoneme. That is what
makes 0.5 a meaningful place to threshold, and it is why different phonemes
can be held to different standards.
Word and utterance scores
An AlignedWord does not carry its own score field; a word’s quality is the
scores of its phonemes. Compute whatever summary suits your task: the mean,
the minimum, or the fraction of phonemes that came back MATCH.
The top-level score is the overall score for the utterance: the mean of the
scores of the reference tokens. If alternative reference texts were
supplied, it is the best score across the primary and all alternatives.
Insertions are not counted in the overall score. The score answers “how
well was the reference produced”, not “was anything else said”. A recording
that contains the target words plus a great deal of unrelated speech can still
score 1.0.
If extraneous speech should count against the speaker (an assistant reading a
prompt, or a child answering twice), inspect the insertion entries yourself and
apply your own rule. The INSERTION tokens carry timestamps, so you can also
measure how much of the audio they account for.
Timestamps
Timestamps live on the hypotheses, not on the token. An AlignedToken has
no time fields of its own; start_time_ms and duration_ms are on each
AlignmentHypothesis inside it, and
on the enclosing AlignedWord.
This has one consequence that catches people out:
A DELETION has an empty hypotheses list and therefore no timestamp at
all. There is no audio to point at; that is what a deletion means. Any
timeline, waveform or spectrogram view must handle tokens that have no time
span, rather than assuming every phoneme can be drawn.
Likewise, a word made up entirely of deletions has nothing to anchor to, and
its start_time_ms and duration_ms are reported as 0.
Turning results into a decision
A pass/fail mark for one item
If your application needs a yes/no verdict (did the speaker say the target correctly?), do not reach for the overall score first. Consider what the task actually requires:
Strict, verbatim tasks. Any divergence is a failure. Require every token to be
MATCH:correct = all( t.kind == capt.ALIGNMENT_KIND_MATCH for w in result.alignments for t in w.tokens )This is exactly right when the rule is “any change at all, however minor, is wrong”, because substitutions, deletions and insertions all break it.
Tolerant tasks. Some divergence is acceptable. A slightly indistinct consonant should not fail an otherwise good attempt. Threshold on the proportion of matched phonemes, or on the mean phoneme score, rather than demanding a clean sweep. Note that a tolerant rule built on the overall
score, or on reference tokens alone, will not notice extra speech: decide separately whether insertions should fail the item.Targeted tasks. Only some phonemes matter: the contrast the item is testing. Score only those tokens and ignore the rest.
Choosing between several candidate answers
When you have a known correct answer and a set of known incorrect answers, and you need to know which one was said, evaluate the audio once per candidate and compare the resulting scores:
candidates = ["BED", "LOUNGE", "SOFA"]
scores = {}
for candidate in candidates:
cfg = capt.EvaluationConfig(model_id=modelID, reference_text=candidate)
with open(path, "rb") as audio:
for resp in client.StreamingEvaluate(stream(cfg, audio)):
if not resp.evaluation_result.is_partial:
scores[candidate] = resp.evaluation_result.score
best = max(scores, key=scores.get)
This gives you a score per candidate, so you can see not just the winner but
the margin, so you can reject the result as unclear when two candidates score
alike. Packing the candidates into alternative_reference_text instead would
collapse them into a single best score and lose exactly that information.
Calibrating against human judgement
If you are automating a decision a person currently makes, the metric that matters is agreement with that person, not any internal notion of accuracy.
The recommended approach is to collect audio alongside the human verdict, run CAPT over it, then sweep your decision rule across the collected set and count true positives, true negatives, false positives and false negatives at each setting. That tells you where to set the threshold, and, just as importantly, what the residual disagreement rate is and which way it leans. A rule that is wrong in the safe direction for your use case is often better than one that is wrong less often overall.
Per-phoneme cutoffs can then correct systematic biases that a single global threshold cannot; see Tuning and Customization.
Alternative alignments
If you supplied alternative_reference_text, each alternative comes back in
alternative_alignments as an
AlternativeAlignment carrying its
own reference_text, score, and full alignments tree with the same
structure described above.
The top-level score is the best across the primary and the alternatives, so
if you only care whether any acceptable rendering was produced, the top-level
score is enough. If you care which one, see above.