CAPT models come in two kinds, reported as kind by
ListModels.
- Word models (
MODEL_KIND_SPEECH_EVALUATION) evaluate phonemes in word context. Thereference_textis ordinary text, and this is what you want for reading words, sentences and passages. - Phoneme models (
MODEL_KIND_PHONEME_EVALUATION) evaluate phonemes produced in isolation. Thereference_textis a sequence of phones.
The distinction matters because a sound produced on its own is acoustically quite different from the same sound inside a word. A model trained on connected speech often mis-recognizes an isolated phoneme, because nothing in its training looked like that. If your task asks a speaker to produce a single sound (“say the /sh/ sound”), a phoneme model is substantially more reliable.
Both kinds are used through exactly the same API calls and return the same
EvaluationResult. Only the reference_text
differs.
Reference text for phoneme models
For a phoneme model, reference_text is one or more phones separated by
whitespace:
# A single phone.
cfg = capt.EvaluationConfig(model_id=phonemeModelID, reference_text="AA")
# A sequence of phones.
cfg = capt.EvaluationConfig(model_id=phonemeModelID, reference_text="Z AE P")
Phone groups
Some phones are only meaningfully produced as a pair, and the model treats such a pair as a single unit called a phone group. Phone groups are written with the members joined by a period:
# A phone group.
cfg = capt.EvaluationConfig(model_id=phonemeModelID, reference_text="AO.NG")
# Mixed with ordinary phones.
cfg = capt.EvaluationConfig(model_id=phonemeModelID, reference_text="F AE.NG")
For all practical purposes a phone group behaves as its own distinct phone: it is matched, substituted or deleted as a single token, and receives a single score.
The phone set
The phones a model accepts are specific to that model, both the inventory and which groups exist. Supplying a phone the model does not know is an error, so check against the set for the model you were given. A typical US English phoneme model accepts:
AA AA.R AE AE.NG AH AH.NG AO AO.NG
AO.R AW AY B CH D DH EH
EH.R ER EY F G HH IH IH.NG
IH.R IY JH K L M N NG
OW OY P R S SH T TH
UH UW V W Y Y.UW Z ZH
Note that word models and phoneme models may use different phone alphabets. A
word model may report X-SAMPA (w E n, r\, aI) while a phoneme model uses
the ARPABET-style symbols above. Do not assume a phone string from one is valid
in the other.
Reading the results
Results have the same structure as any other evaluation. For each phone in the
reference_text there is an aligned token:
ALIGNMENT_KIND_MATCHmeans the phone was found in the audio with a high degree of confidence.ALIGNMENT_KIND_SUBSTITUTIONmeans a different phone was recognized in its place, or the phone was recognized only with low confidence.ALIGNMENT_KIND_DELETIONmeans the phone was not recognized at all.
Anything else recognized in the audio that does not align to the reference is reported as insertions, grouped separately.
Isolated-phoneme audio very often contains more than the phoneme itself: a
speaker clearing their throat, a lead-in, or the assessor’s prompt. Those show
up as insertions, which is why the example below scores 1.0 despite three
extra tokens: the reference phone ZH was produced correctly, and the
surrounding material is reported rather than being allowed to affect the score
of the phone you asked about.
{
"score": 1.0,
"alignments": [
{
"start_time_ms": "680",
"duration_ms": "360",
"tokens": [
{
"kind": "ALIGNMENT_KIND_INSERTION",
"hypotheses": [
{ "token": "T", "confidence": 1.0, "start_time_ms": "680", "duration_ms": "120" }
]
},
{
"kind": "ALIGNMENT_KIND_INSERTION",
"hypotheses": [
{ "token": "R", "confidence": 1.0, "start_time_ms": "840", "duration_ms": "120" }
]
},
{
"kind": "ALIGNMENT_KIND_INSERTION",
"hypotheses": [
{ "token": "EH", "confidence": 0.994, "start_time_ms": "960", "duration_ms": "80" }
]
}
]
},
{
"text": "ZH",
"start_time_ms": "1000",
"duration_ms": "240",
"tokens": [
{
"kind": "ALIGNMENT_KIND_MATCH",
"reference": "ZH",
"score": 1.0,
"hypotheses": [
{ "token": "ZH", "confidence": 1.0, "start_time_ms": "1000", "duration_ms": "240" }
]
}
]
}
]
}
If extraneous audio should count against the speaker in your task, inspect the insertion entries and apply your own rule; the information is there either way.