Phoneme Models

Evaluating individual phonemes produced in isolation.

CAPT models come in two kinds, reported as kind by ListModels.

  • Word models (MODEL_KIND_SPEECH_EVALUATION) evaluate phonemes in word context. The reference_text is ordinary text, and this is what you want for reading words, sentences and passages.
  • Phoneme models (MODEL_KIND_PHONEME_EVALUATION) evaluate phonemes produced in isolation. The reference_text is a sequence of phones.

The distinction matters because a sound produced on its own is acoustically quite different from the same sound inside a word. A model trained on connected speech often mis-recognizes an isolated phoneme, because nothing in its training looked like that. If your task asks a speaker to produce a single sound (“say the /sh/ sound”), a phoneme model is substantially more reliable.

Both kinds are used through exactly the same API calls and return the same EvaluationResult. Only the reference_text differs.

Reference text for phoneme models

For a phoneme model, reference_text is one or more phones separated by whitespace:

# A single phone.
cfg = capt.EvaluationConfig(model_id=phonemeModelID, reference_text="AA")

# A sequence of phones.
cfg = capt.EvaluationConfig(model_id=phonemeModelID, reference_text="Z AE P")

Phone groups

Some phones are only meaningfully produced as a pair, and the model treats such a pair as a single unit called a phone group. Phone groups are written with the members joined by a period:

# A phone group.
cfg = capt.EvaluationConfig(model_id=phonemeModelID, reference_text="AO.NG")

# Mixed with ordinary phones.
cfg = capt.EvaluationConfig(model_id=phonemeModelID, reference_text="F AE.NG")

For all practical purposes a phone group behaves as its own distinct phone: it is matched, substituted or deleted as a single token, and receives a single score.

The phone set

The phones a model accepts are specific to that model, both the inventory and which groups exist. Supplying a phone the model does not know is an error, so check against the set for the model you were given. A typical US English phoneme model accepts:

AA     AA.R   AE     AE.NG  AH     AH.NG  AO     AO.NG
AO.R   AW     AY     B      CH     D      DH     EH
EH.R   ER     EY     F      G      HH     IH     IH.NG
IH.R   IY     JH     K      L      M      N      NG
OW     OY     P      R      S      SH     T      TH
UH     UW     V      W      Y      Y.UW   Z      ZH

Reading the results

Results have the same structure as any other evaluation. For each phone in the reference_text there is an aligned token:

  • ALIGNMENT_KIND_MATCH means the phone was found in the audio with a high degree of confidence.
  • ALIGNMENT_KIND_SUBSTITUTION means a different phone was recognized in its place, or the phone was recognized only with low confidence.
  • ALIGNMENT_KIND_DELETION means the phone was not recognized at all.

Anything else recognized in the audio that does not align to the reference is reported as insertions, grouped separately.

Isolated-phoneme audio very often contains more than the phoneme itself: a speaker clearing their throat, a lead-in, or the assessor’s prompt. Those show up as insertions, which is why the example below scores 1.0 despite three extra tokens: the reference phone ZH was produced correctly, and the surrounding material is reported rather than being allowed to affect the score of the phone you asked about.

{
  "score": 1.0,
  "alignments": [
    {
      "start_time_ms": "680",
      "duration_ms": "360",
      "tokens": [
        {
          "kind": "ALIGNMENT_KIND_INSERTION",
          "hypotheses": [
            { "token": "T", "confidence": 1.0, "start_time_ms": "680", "duration_ms": "120" }
          ]
        },
        {
          "kind": "ALIGNMENT_KIND_INSERTION",
          "hypotheses": [
            { "token": "R", "confidence": 1.0, "start_time_ms": "840", "duration_ms": "120" }
          ]
        },
        {
          "kind": "ALIGNMENT_KIND_INSERTION",
          "hypotheses": [
            { "token": "EH", "confidence": 0.994, "start_time_ms": "960", "duration_ms": "80" }
          ]
        }
      ]
    },
    {
      "text": "ZH",
      "start_time_ms": "1000",
      "duration_ms": "240",
      "tokens": [
        {
          "kind": "ALIGNMENT_KIND_MATCH",
          "reference": "ZH",
          "score": 1.0,
          "hypotheses": [
            { "token": "ZH", "confidence": 1.0, "start_time_ms": "1000", "duration_ms": "240" }
          ]
        }
      ]
    }
  ]
}

If extraneous audio should count against the speaker in your task, inspect the insertion entries and apply your own rule; the information is there either way.