Every evaluation begins with an
EvaluationConfig. It is the first
message of a StreamingEvaluate stream, and the config field of a unary
Evaluate request. This page describes what goes in it.
Choosing a model
model_id selects the model to evaluate against, and must be one of the id
values returned by ListModels. Each model
also reports the sample_rate it expects and its kind:
MODEL_KIND_SPEECH_EVALUATIONevaluates words and sentences. Thereference_textis ordinary text.MODEL_KIND_PHONEME_EVALUATIONevaluates phonemes produced in isolation, which a word-oriented model may not recognize correctly. Thereference_textis a sequence of phones. See Phoneme Models.
Both kinds return the same EvaluationResult
structure, so only the reference_text you supply differs.
Reference text
reference_text is the text the speaker is expected to have said: the thing
the audio is scored against. This is the field that makes CAPT what it is: the
evaluation is anchored to it.
cfg = capt.EvaluationConfig(
model_id="en_US-16khz",
reference_text="WHEN THE SUNLIGHT STRIKES",
)
A few properties worth knowing:
- Lexicon lookup is case-insensitive, so
when,WhenandWHENbehave identically. The uppercase reference text used throughout these examples is a convention, not a requirement. - Punctuation handling is a per-model setting,
RemovePunctuationin the model’s[Tokenization]config, and it is off in theen_US-16khzmodel shipped today. Punctuation therefore stays attached to the word: a reference ofWHEN, THE SUNLIGHT STRIKEScomes back withtextofWHEN,, and because the token no longer matches the lexicon it is phonemized by G2P, which changed that word’s score in our own testing. Sentence-final marks are usually harmless, but the safe habit is to send the words alone. See Tuning and Customization. - Words not in the lexicon are handled by a G2P model, so invented words, proper nouns and non-words can be used as reference text without any setup. This is what makes non-word decoding tasks work. See Tuning and Customization.
- Words with several valid pronunciations are all considered, and the
variant that best fits the audio is the one reported. A speaker is not
penalized for saying “THE” as
D irather thanD @.
Alternative reference text
alternative_reference_text accepts additional reference texts that would also
be acceptable. Each is aligned independently and returned in
alternative_alignments with its own
score, and the top-level score becomes the best score across the primary and
all alternatives.
cfg = capt.EvaluationConfig(
model_id="en_US-16khz",
reference_text="WHEN THE SUNLIGHT STRIKES",
alternative_reference_text=["WHEN THE SUN LIGHTS STRIKE"],
)
If you need to know which candidate the audio matched, for example when
choosing between a correct answer and a set of known incorrect answers, it is
clearer to issue a separate evaluation per candidate and compare the scores
yourself. alternative_reference_text optimizes for “any of these is
acceptable”, not for classification. See Interpreting
Results.
Audio format
audio_format tells the server how to interpret the bytes you send. Two shapes
are available.
For files that carry their own header, name the container and let the server read the rest:
cfg = capt.EvaluationConfig(
model_id="en_US-16khz",
reference_text="WHEN THE SUNLIGHT STRIKES",
audio_format=capt.AudioFormat(
audio_format_headered=capt.AUDIO_FORMAT_HEADERED_WAV,
),
)
Supported headered formats are WAV, MP3, FLAC and OGG Opus. Headers are
expected once, at the beginning of the stream, not in every Audio message.
audio_format may be omitted entirely for headered audio. The server then
detects the format from the header, which is why several examples in these
pages leave it out. Naming it explicitly is still recommended: it turns an
unreadable or unexpected header into a clear error instead of a detection
attempt.
Raw audio is the exception. It carries no header to detect, so
audio_format_raw must be supplied in full, and encoding must be set to
something other than AUDIO_ENCODING_UNSPECIFIED or the request is
rejected.
For raw samples with no header, describe them fully:
cfg = capt.EvaluationConfig(
model_id="en_US-16khz",
reference_text="WHEN THE SUNLIGHT STRIKES",
audio_format=capt.AudioFormat(
audio_format_raw=capt.AudioFormatRAW(
encoding=capt.AUDIO_ENCODING_SIGNED,
bit_depth=16,
byte_order=capt.BYTE_ORDER_LITTLE_ENDIAN,
sample_rate=16000,
channels=1,
),
),
)
Every server supports raw little-endian 16-bit signed samples at the model’s own sample rate; support for the other formats depends on how the server was built and configured.
For best accuracy use uncompressed or losslessly compressed audio (WAV or
FLAC) recorded at the model’s native sample rate. Sending audio sampled below
what the model expects produces a non-fatal
EvaluationError alongside your results,
warning that accuracy may be reduced. Upsampling low-rate audio does not
recover the missing detail.
Request metadata
metadata is optional and is never used to influence evaluation. The server
may record it in logs or stored results.
cfg = capt.EvaluationConfig(
model_id="en_US-16khz",
reference_text="WHEN THE SUNLIGHT STRIKES",
metadata=capt.EvaluationMetadata(
custom_id="session-42-item-07",
custom_metadata='{"item":"7","form":"A"}',
),
)
custom_idaccepts up to 64 bytes, restricted to letters, digits, hyphens and underscores. Useful as a tracing or correlation ID.custom_metadatais a free-form string; a plain tag or structured data such as JSON.
Because these fields may be written to logs and to stored results, they must not contain personal or otherwise sensitive information. Use an opaque identifier that only your own systems can resolve back to a person.
Once you have a config, continue to Streaming Evaluation.