This is the multi-page printable view of this section. Click here to print.

Return to the regular view of this page.

Phoneme Models

Evaluating individual phonemes produced in isolation.

    CAPT models come in two kinds, reported as kind by ListModels.

    • Word models (MODEL_KIND_SPEECH_EVALUATION) evaluate phonemes in word context. The reference_text is ordinary text, and this is what you want for reading words, sentences and passages.
    • Phoneme models (MODEL_KIND_PHONEME_EVALUATION) evaluate phonemes produced in isolation. The reference_text is a sequence of phones.

    The distinction matters because a sound produced on its own is acoustically quite different from the same sound inside a word. A model trained on connected speech often mis-recognizes an isolated phoneme, because nothing in its training looked like that. If your task asks a speaker to produce a single sound (“say the /sh/ sound”), a phoneme model is substantially more reliable.

    Both kinds are used through exactly the same API calls and return the same EvaluationResult. Only the reference_text differs.

    Reference text for phoneme models

    For a phoneme model, reference_text is one or more phones separated by whitespace:

    # A single phone.
    cfg = capt.EvaluationConfig(model_id=phonemeModelID, reference_text="AA")
    
    # A sequence of phones.
    cfg = capt.EvaluationConfig(model_id=phonemeModelID, reference_text="Z AE P")
    

    Phone groups

    Some phones are only meaningfully produced as a pair, and the model treats such a pair as a single unit called a phone group. Phone groups are written with the members joined by a period:

    # A phone group.
    cfg = capt.EvaluationConfig(model_id=phonemeModelID, reference_text="AO.NG")
    
    # Mixed with ordinary phones.
    cfg = capt.EvaluationConfig(model_id=phonemeModelID, reference_text="F AE.NG")
    

    For all practical purposes a phone group behaves as its own distinct phone: it is matched, substituted or deleted as a single token, and receives a single score.

    The phone set

    The phones a model accepts are specific to that model, both the inventory and which groups exist. Supplying a phone the model does not know is an error, so check against the set for the model you were given. A typical US English phoneme model accepts:

    AA     AA.R   AE     AE.NG  AH     AH.NG  AO     AO.NG
    AO.R   AW     AY     B      CH     D      DH     EH
    EH.R   ER     EY     F      G      HH     IH     IH.NG
    IH.R   IY     JH     K      L      M      N      NG
    OW     OY     P      R      S      SH     T      TH
    UH     UW     V      W      Y      Y.UW   Z      ZH
    

    Reading the results

    Results have the same structure as any other evaluation. For each phone in the reference_text there is an aligned token:

    • ALIGNMENT_KIND_MATCH means the phone was found in the audio with a high degree of confidence.
    • ALIGNMENT_KIND_SUBSTITUTION means a different phone was recognized in its place, or the phone was recognized only with low confidence.
    • ALIGNMENT_KIND_DELETION means the phone was not recognized at all.

    Anything else recognized in the audio that does not align to the reference is reported as insertions, grouped separately.

    Isolated-phoneme audio very often contains more than the phoneme itself: a speaker clearing their throat, a lead-in, or the assessor’s prompt. Those show up as insertions, which is why the example below scores 1.0 despite three extra tokens: the reference phone ZH was produced correctly, and the surrounding material is reported rather than being allowed to affect the score of the phone you asked about.

    {
      "score": 1.0,
      "alignments": [
        {
          "start_time_ms": "680",
          "duration_ms": "360",
          "tokens": [
            {
              "kind": "ALIGNMENT_KIND_INSERTION",
              "hypotheses": [
                { "token": "T", "confidence": 1.0, "start_time_ms": "680", "duration_ms": "120" }
              ]
            },
            {
              "kind": "ALIGNMENT_KIND_INSERTION",
              "hypotheses": [
                { "token": "R", "confidence": 1.0, "start_time_ms": "840", "duration_ms": "120" }
              ]
            },
            {
              "kind": "ALIGNMENT_KIND_INSERTION",
              "hypotheses": [
                { "token": "EH", "confidence": 0.994, "start_time_ms": "960", "duration_ms": "80" }
              ]
            }
          ]
        },
        {
          "text": "ZH",
          "start_time_ms": "1000",
          "duration_ms": "240",
          "tokens": [
            {
              "kind": "ALIGNMENT_KIND_MATCH",
              "reference": "ZH",
              "score": 1.0,
              "hypotheses": [
                { "token": "ZH", "confidence": 1.0, "start_time_ms": "1000", "duration_ms": "240" }
              ]
            }
          ]
        }
      ]
    }
    

    If extraneous audio should count against the speaker in your task, inspect the insertion entries and apply your own rule; the information is there either way.