This is the multi-page printable view of this section. Click here to print.

Return to the regular view of this page.

Tuning and Customization

Pronunciations, per-phoneme thresholds and scoring behavior.

    A general-purpose speech model is a fixed object: you send it audio and accept what it returns. CAPT is deliberately not that. Because the scoring stage is separate from the acoustic model, a great deal of behavior can be adjusted for your task without retraining anything, by changing which pronunciations count as correct, and how strictly each sound is judged.

    Everything on this page lives in the model directory shipped with your release, and takes effect when the server restarts.

    models/en_US-16khz/capt/
    ├── model.config.toml     scoring behavior and paths to the files below
    ├── lexicon.json          the primary pronunciation dictionary
    ├── lexicon_addenda.tsv   your additions and overrides
    ├── cutoffs.yaml          per-phoneme thresholds
    └── g2p.ort               grapheme-to-phoneme model for unknown words
    

    The model config points at each of these by path, so the names above are the convention rather than a requirement. Older model directories may carry the cutoffs as a plain-text cutoffs.txt instead; the meaning is identical and CutoffsPath says which file is in use.

    Pronunciations

    The lexicon and its addenda

    The lexicon maps words to their phoneme sequences. A word may have several valid pronunciations, and CAPT considers all of them, reporting whichever best fits the audio, so a speaker is not penalized for a legitimate variant.

    lexicon_addenda.tsv is loaded after the primary lexicon and is the file you should edit. Entries in it add new words and override existing ones, leaving the shipped lexicon intact so that upgrades stay clean.

    This is the mechanism to reach for when your task defines what counts as an acceptable pronunciation. If an item specifies that a target may be produced two ways, add both and each will be accepted as a match:

    MOSP	m A s p
    MOSP	m oU s p
    

    Those two entries accept mosp pronounced to rhyme with wasp, or with soap.

    Out-of-vocabulary words and G2P

    Words that are in neither the lexicon nor the addenda are passed to a grapheme-to-phoneme model, which predicts a pronunciation from spelling. This happens automatically and requires no configuration.

    It means invented words, proper nouns and non-words can be used as reference_text with no setup at all, which is useful for decoding tasks built on pseudo-words, where the whole point is that the string is not a real word.

    Tokenization

    [Tokenization] in model.config.toml controls how reference text is pre-processed:

    SettingWhat it doesDefaultIn en_US-16khz
    RemovePunctuationStrips punctuation before lexicon lookup. With it off, punctuation stays attached to the word and is phonemized as part of it.offoff
    UseSentenceTokenizerSplits a long reference into sentences before tokenizing.offon
    MaxInputLengthSafeguard on the length of a reference text. Counted in characters, not words; a longer reference is rejected with an error. Applies to each alternative reference text too.20482048

    “Default” is what applies when the key is absent from model.config.toml. Neither boolean has a server-side default beyond false, so what matters in practice is what your model ships with.

    Cutoffs and acceptable confusions

    This is the most useful tuning mechanism CAPT offers, and the least obvious.

    Raw acoustic confidences are biased in ways that are specific to a model and a population of speakers. A model may routinely blur m and n; young children may produce th in a way that reads as s. Left alone, those biases show up as wrong verdicts. Cutoffs correct them by remapping confidence to score.

    Each symbol has a cutoff, the confidence at which it is exactly borderline. The remapping is piecewise linear and pins that point to a score of 0.5:

    • raw confidence in [0, cutoff] maps to a score in [0, 0.5]
    • raw confidence in [cutoff, 1] maps to a score in [0.5, 1]

    So the cutoff is the dial for how strict a given phoneme is. Lowering it is more lenient; raising it demands higher confidence before the phoneme counts as correct. Any symbol not listed uses DefaultCutoffScore from the [Scoring] section.

    A modifier extends this with acceptable confusions. It names alternatives (other phonemes that may also count toward the match) and a multiplier that scales the cutoff when one of them is what was actually heard:

    "A":
      cutoff: 0.50
      modifier:
        multiplier: 1.50
        alternatives: ["O"]
    
    "T":
      cutoff: 0.50
      modifier:
        multiplier: 1.90
        alternatives: ["s"]
    

    Read the first entry as: "A (the vowel of lot) needs 50% confidence. O (the vowel of thought) is close enough to count, but only if the model is much more sure of it: 0.50 × 1.50 = 0.75." The second is stricter still: s may stand in for T (th), but must clear 0.50 × 1.90 = 0.95, which is about right for a contrast that many young speakers have not yet acquired.

    During alignment the confidences of the reference phoneme and any listed alternatives are summed, and the scaled cutoff is applied to that sum.

    The effect is a per-sound tolerance you control. Rather than one global threshold that is too strict for some phonemes and too lax for others, you set the bar sound by sound, and you do it by editing a YAML file, not by collecting data and retraining.

    Scoring behavior

    The [Scoring] section of model.config.toml governs the alignment itself. The settings most likely to matter:

    SettingWhat it does
    DefaultCutoffScoreThe cutoff for any symbol not listed in cutoffs.yaml. Below 0.5 is lenient, above 0.5 is strict.
    SubCost, InsCost, DelCostEdit-distance costs for substitution, insertion and deletion. Raising one makes the aligner prefer explanations that avoid it.
    AlignVowelsPrevents nonsensical vowel↔consonant substitutions, reporting a deletion plus an insertion instead.
    IncludeStressTreat stress markers as distinguishing (AH0 ≠ AH1). Requires both the lexicon and the acoustic model to carry stress.
    IncludeIntraWordInsertionWhether extra phonemes heard inside a word are attached to that word or reported separately.
    NonSpeechTokenConfidenceScalingWeights the confidence of silence and noise tokens. Below 1.0 makes the model less willing to call something non-speech, which is useful in noisy rooms.
    CnetLinkThresholdDiscards time slices where the acoustic model put little probability on any speech token.
    MaxCandidatePronsCaps how many combinations of word pronunciations are enumerated for a reference.

    Non-speech tokens

    NonSpeechTokens in [Vocab] lists the symbols the acoustic model emits for things that are not speech: silence, breath, vocal noise, unknown sounds. These are handled specially: they are excluded from scoring rather than being treated as mispronunciations, and they interact with CnetLinkThreshold and NonSpeechTokenConfidenceScaling above.

    Recordings made in real rooms contain plenty of this material. If you find that noisy recordings score too harshly, these settings, rather than the cutoffs, are usually the right place to look.

    Adapting the acoustic model

    Configuration handles a great deal, but not everything. Where a population is genuinely different from what the model was trained on (young children, a regional accent, a specific recording setup), the acoustic model itself can be adapted to it, which addresses the cause rather than compensating for it downstream.

    Adaptation needs representative audio, ideally paired with the verdicts you want the system to reproduce. If your application already records both, you may have a suitable dataset without any additional collection effort. Contact us to discuss what your data supports.