Interpreting Results

How to read an EvaluationResult and turn it into a decision.

This is the page worth reading carefully. Getting audio into CAPT is straightforward; deciding what its output means for your application is where the design work is.

The shape of a result

An EvaluationResult is a tree:

EvaluationResult
├── score                    overall score for the utterance, 0..1
├── is_partial               true while more audio may still change this
├── is_speech_endpoint       true if the speaker appears to have finished
├── alignments               [ ]AlignedWord
│   └── AlignedWord
│       ├── text             the reference word
│       ├── start_time_ms    when the word starts in the audio
│       ├── duration_ms      how long it lasts
│       └── tokens           [ ]AlignedToken   (the phonemes)
│           └── AlignedToken
│               ├── kind           MATCH | SUBSTITUTION | DELETION | INSERTION
│               ├── reference      the expected phoneme
│               ├── score          0..1 for this phoneme
│               └── hypotheses     [ ]AlignmentHypothesis  (what was heard)
│                   └── token, confidence, start_time_ms, duration_ms
└── alternative_alignments   the same, once per alternative reference text

The reference text drives the structure: your words appear as AlignedWord entries in order, and within each one there is an AlignedToken per expected phoneme.

The four alignment kinds

Every expected phoneme gets exactly one of four verdicts.

KindMeaningreferencehypothesesscore
MATCHThe expected phoneme was heard with enough confidencesetpopulated0.5 to 1.0
SUBSTITUTIONThe expected phoneme was not heard with enough confidence: either something else was heard in its place, or the phoneme itself was heard but below its cutoffsetpopulated0 to below 0.5
DELETIONThe phoneme was not heard at allsetempty0
INSERTIONSomething extra was heard that the reference does not account foremptypopulated0

A real example makes this concrete. Below is the result of scoring a recording of “when the sunlight strikes” against a deliberately wrong reference, “WHEN THE MOONLIGHT STRIKES” (phonemes are X-SAMPA):

overall score: 0.5157

WHEN         MATCH w (1.00)  MATCH E (0.99)  MATCH n (1.00)
THE          MATCH D (0.87)  MATCH @ (0.67)
MOONLIGHT    SUB   m (0.00)  SUB   u (0.00)  MATCH n (0.65)
             MATCH l (0.78)  MATCH aI (0.82) MATCH t (0.73)
STRIKES      MATCH s (0.79)  DEL   t (0.00)  DEL   r\ (0.00)
             DEL   aI (0.00) SUB   k (0.49)  DEL   s (0.00)

Read that as a diagnosis, not just a number. The first two words were said as expected. In MOONLIGHT the leading m u was not there (the speaker said s V n), but the shared tail n l aI t matched, so those phonemes score well. STRIKES shows what happens when the audio has already run out: the reference still has phonemes to account for, and they come back as deletions.

The overall score of 0.52 is the average of that mixture. This is why an overall score alone is rarely the right thing to threshold. 0.52 here does not mean “roughly half right”, it means “two words right and two words wrong”.

Insertions

Insertions do not belong to any reference word, so they are reported in their own AlignedWord entries with an empty text, positioned in the sequence where they were heard. Scoring the same audio against the single word “ELEPHANT” shows this:

overall score: 0.3966

(insertion)  INS  (heard: w)
ELEPHANT     MATCH E (0.99)  SUB l (0.00)  MATCH @ (0.67)
             SUB   f (0.00)  SUB n= (0.00) MATCH t (0.73)
(insertion)  INS  (heard: k)  INS (heard: s)

The speaker said far more than the reference accounted for, and the leftover audio at each end is reported as insertions rather than being silently discarded.

Where the numbers come from

Phoneme scores

The acoustic model does not commit to a single phoneme per time slice. It produces a distribution, a confusion network, and the score for an expected phoneme is derived from how much probability mass landed on it.

You can see this in the hypotheses list. Here is a single MATCH from a real result:

{
  "kind": "ALIGNMENT_KIND_MATCH",
  "reference": "n",
  "score": 0.9798913,
  "hypotheses": [
    { "token": "n",  "confidence": 0.963, "start_time_ms": "570", "duration_ms": "119" },
    { "token": "m",  "confidence": 0.029, "start_time_ms": "570", "duration_ms": "119" },
    { "token": "n=", "confidence": 0.006, "start_time_ms": "570", "duration_ms": "119" },
    { "token": "N",  "confidence": 0.002, "start_time_ms": "570", "duration_ms": "119" }
  ]
}

The model heard n with 0.963 confidence, but also considered m, n= and N. The reported score of 0.98 is that confidence remapped through the model’s cutoff for n. See Tuning and Customization for how that remapping works and how to change it.

The practical consequence: a score is a calibrated quantity, not a raw probability. A cutoff is chosen so that a score of exactly 0.5 sits at the boundary between acceptable and unacceptable for that phoneme. That is what makes 0.5 a meaningful place to threshold, and it is why different phonemes can be held to different standards.

Word and utterance scores

An AlignedWord does not carry its own score field; a word’s quality is the scores of its phonemes. Compute whatever summary suits your task: the mean, the minimum, or the fraction of phonemes that came back MATCH.

The top-level score is the overall score for the utterance: the mean of the scores of the reference tokens. If alternative reference texts were supplied, it is the best score across the primary and all alternatives.

Timestamps

Timestamps live on the hypotheses, not on the token. An AlignedToken has no time fields of its own; start_time_ms and duration_ms are on each AlignmentHypothesis inside it, and on the enclosing AlignedWord.

This has one consequence that catches people out:

Likewise, a word made up entirely of deletions has nothing to anchor to, and its start_time_ms and duration_ms are reported as 0.

Turning results into a decision

A pass/fail mark for one item

If your application needs a yes/no verdict (did the speaker say the target correctly?), do not reach for the overall score first. Consider what the task actually requires:

  • Strict, verbatim tasks. Any divergence is a failure. Require every token to be MATCH:

    correct = all(
        t.kind == capt.ALIGNMENT_KIND_MATCH
        for w in result.alignments
        for t in w.tokens
    )
    

    This is exactly right when the rule is “any change at all, however minor, is wrong”, because substitutions, deletions and insertions all break it.

  • Tolerant tasks. Some divergence is acceptable. A slightly indistinct consonant should not fail an otherwise good attempt. Threshold on the proportion of matched phonemes, or on the mean phoneme score, rather than demanding a clean sweep. Note that a tolerant rule built on the overall score, or on reference tokens alone, will not notice extra speech: decide separately whether insertions should fail the item.

  • Targeted tasks. Only some phonemes matter: the contrast the item is testing. Score only those tokens and ignore the rest.

Choosing between several candidate answers

When you have a known correct answer and a set of known incorrect answers, and you need to know which one was said, evaluate the audio once per candidate and compare the resulting scores:

candidates = ["BED", "LOUNGE", "SOFA"]
scores = {}

for candidate in candidates:
    cfg = capt.EvaluationConfig(model_id=modelID, reference_text=candidate)
    with open(path, "rb") as audio:
        for resp in client.StreamingEvaluate(stream(cfg, audio)):
            if not resp.evaluation_result.is_partial:
                scores[candidate] = resp.evaluation_result.score

best = max(scores, key=scores.get)

This gives you a score per candidate, so you can see not just the winner but the margin, so you can reject the result as unclear when two candidates score alike. Packing the candidates into alternative_reference_text instead would collapse them into a single best score and lose exactly that information.

Calibrating against human judgement

If you are automating a decision a person currently makes, the metric that matters is agreement with that person, not any internal notion of accuracy.

The recommended approach is to collect audio alongside the human verdict, run CAPT over it, then sweep your decision rule across the collected set and count true positives, true negatives, false positives and false negatives at each setting. That tells you where to set the threshold, and, just as importantly, what the residual disagreement rate is and which way it leans. A rule that is wrong in the safe direction for your use case is often better than one that is wrong less often overall.

Per-phoneme cutoffs can then correct systematic biases that a single global threshold cannot; see Tuning and Customization.

Alternative alignments

If you supplied alternative_reference_text, each alternative comes back in alternative_alignments as an AlternativeAlignment carrying its own reference_text, score, and full alignments tree with the same structure described above.

The top-level score is the best across the primary and the alternatives, so if you only care whether any acceptable rendering was produced, the top-level score is enough. If you care which one, see above.