This is the multi-page printable view of this section. Click here to print.

Return to the regular view of this page.

Evaluation Configurations

The options available when configuring an evaluation request.

    Every evaluation begins with an EvaluationConfig. It is the first message of a StreamingEvaluate stream, and the config field of a unary Evaluate request. This page describes what goes in it.

    Choosing a model

    model_id selects the model to evaluate against, and must be one of the id values returned by ListModels. Each model also reports the sample_rate it expects and its kind:

    • MODEL_KIND_SPEECH_EVALUATION evaluates words and sentences. The reference_text is ordinary text.
    • MODEL_KIND_PHONEME_EVALUATION evaluates phonemes produced in isolation, which a word-oriented model may not recognize correctly. The reference_text is a sequence of phones. See Phoneme Models.

    Both kinds return the same EvaluationResult structure, so only the reference_text you supply differs.

    Reference text

    reference_text is the text the speaker is expected to have said: the thing the audio is scored against. This is the field that makes CAPT what it is: the evaluation is anchored to it.

    cfg = capt.EvaluationConfig(
        model_id="en_US-16khz",
        reference_text="WHEN THE SUNLIGHT STRIKES",
    )
    

    A few properties worth knowing:

    • Lexicon lookup is case-insensitive, so when, When and WHEN behave identically. The uppercase reference text used throughout these examples is a convention, not a requirement.
    • Punctuation handling is a per-model setting, RemovePunctuation in the model’s [Tokenization] config, and it is off in the en_US-16khz model shipped today. Punctuation therefore stays attached to the word: a reference of WHEN, THE SUNLIGHT STRIKES comes back with text of WHEN,, and because the token no longer matches the lexicon it is phonemized by G2P, which changed that word’s score in our own testing. Sentence-final marks are usually harmless, but the safe habit is to send the words alone. See Tuning and Customization.
    • Words not in the lexicon are handled by a G2P model, so invented words, proper nouns and non-words can be used as reference text without any setup. This is what makes non-word decoding tasks work. See Tuning and Customization.
    • Words with several valid pronunciations are all considered, and the variant that best fits the audio is the one reported. A speaker is not penalized for saying “THE” as D i rather than D @.

    Alternative reference text

    alternative_reference_text accepts additional reference texts that would also be acceptable. Each is aligned independently and returned in alternative_alignments with its own score, and the top-level score becomes the best score across the primary and all alternatives.

    cfg = capt.EvaluationConfig(
        model_id="en_US-16khz",
        reference_text="WHEN THE SUNLIGHT STRIKES",
        alternative_reference_text=["WHEN THE SUN LIGHTS STRIKE"],
    )
    

    Audio format

    audio_format tells the server how to interpret the bytes you send. Two shapes are available.

    For files that carry their own header, name the container and let the server read the rest:

    cfg = capt.EvaluationConfig(
        model_id="en_US-16khz",
        reference_text="WHEN THE SUNLIGHT STRIKES",
        audio_format=capt.AudioFormat(
            audio_format_headered=capt.AUDIO_FORMAT_HEADERED_WAV,
        ),
    )
    

    Supported headered formats are WAV, MP3, FLAC and OGG Opus. Headers are expected once, at the beginning of the stream, not in every Audio message.

    audio_format may be omitted entirely for headered audio. The server then detects the format from the header, which is why several examples in these pages leave it out. Naming it explicitly is still recommended: it turns an unreadable or unexpected header into a clear error instead of a detection attempt.

    Raw audio is the exception. It carries no header to detect, so audio_format_raw must be supplied in full, and encoding must be set to something other than AUDIO_ENCODING_UNSPECIFIED or the request is rejected.

    For raw samples with no header, describe them fully:

    cfg = capt.EvaluationConfig(
        model_id="en_US-16khz",
        reference_text="WHEN THE SUNLIGHT STRIKES",
        audio_format=capt.AudioFormat(
            audio_format_raw=capt.AudioFormatRAW(
                encoding=capt.AUDIO_ENCODING_SIGNED,
                bit_depth=16,
                byte_order=capt.BYTE_ORDER_LITTLE_ENDIAN,
                sample_rate=16000,
                channels=1,
            ),
        ),
    )
    

    Every server supports raw little-endian 16-bit signed samples at the model’s own sample rate; support for the other formats depends on how the server was built and configured.

    Request metadata

    metadata is optional and is never used to influence evaluation. The server may record it in logs or stored results.

    cfg = capt.EvaluationConfig(
        model_id="en_US-16khz",
        reference_text="WHEN THE SUNLIGHT STRIKES",
        metadata=capt.EvaluationMetadata(
            custom_id="session-42-item-07",
            custom_metadata='{"item":"7","form":"A"}',
        ),
    )
    
    • custom_id accepts up to 64 bytes, restricted to letters, digits, hyphens and underscores. Useful as a tracing or correlation ID.
    • custom_metadata is a free-form string; a plain tag or structured data such as JSON.

    Once you have a config, continue to Streaming Evaluation.