Tuning and Customization
A general-purpose speech model is a fixed object: you send it audio and accept what it returns. CAPT is deliberately not that. Because the scoring stage is separate from the acoustic model, a great deal of behavior can be adjusted for your task without retraining anything, by changing which pronunciations count as correct, and how strictly each sound is judged.
Everything on this page lives in the model directory shipped with your release, and takes effect when the server restarts.
models/en_US-16khz/capt/
├── model.config.toml scoring behavior and paths to the files below
├── lexicon.json the primary pronunciation dictionary
├── lexicon_addenda.tsv your additions and overrides
├── cutoffs.yaml per-phoneme thresholds
└── g2p.ort grapheme-to-phoneme model for unknown words
The model config points at each of these by path, so the names above are the
convention rather than a requirement. Older model directories may carry the
cutoffs as a plain-text cutoffs.txt instead; the meaning is identical and
CutoffsPath says which file is in use.
Pronunciations
The lexicon and its addenda
The lexicon maps words to their phoneme sequences. A word may have several valid pronunciations, and CAPT considers all of them, reporting whichever best fits the audio, so a speaker is not penalized for a legitimate variant.
lexicon_addenda.tsv is loaded after the primary lexicon and is the file
you should edit. Entries in it add new words and override existing ones,
leaving the shipped lexicon intact so that upgrades stay clean.
This is the mechanism to reach for when your task defines what counts as an acceptable pronunciation. If an item specifies that a target may be produced two ways, add both and each will be accepted as a match:
MOSP m A s p
MOSP m oU s p
Those two entries accept mosp pronounced to rhyme with wasp, or with soap.
Every phone symbol on this page belongs to the model’s own alphabet, and
the files in a model directory all use the same one. The en_US-16khz word
model used in these examples is X-SAMPA, which is why its addenda read
m A s p and its cutoffs are keyed "A" rather than AA. A phoneme model
with an ARPABET inventory would use AA in both files instead. Check the phone
set for the model you were given before editing either file: a symbol the model
does not know matches nothing, and fails silently rather than erroring.
Out-of-vocabulary words and G2P
Words that are in neither the lexicon nor the addenda are passed to a grapheme-to-phoneme model, which predicts a pronunciation from spelling. This happens automatically and requires no configuration.
It means invented words, proper nouns and non-words can be used as
reference_text with no setup at all, which is useful for decoding tasks built on
pseudo-words, where the whole point is that the string is not a real word.
G2P is a prediction, not a definition. Where you know the pronunciations that should be accepted, because your task specifies them, put them in the addenda rather than relying on what G2P infers from the spelling. Reserve G2P for the long tail you have not enumerated.
Tokenization
[Tokenization] in model.config.toml controls how reference text is
pre-processed:
| Setting | What it does | Default | In en_US-16khz |
|---|---|---|---|
RemovePunctuation | Strips punctuation before lexicon lookup. With it off, punctuation stays attached to the word and is phonemized as part of it. | off | off |
UseSentenceTokenizer | Splits a long reference into sentences before tokenizing. | off | on |
MaxInputLength | Safeguard on the length of a reference text. Counted in characters, not words; a longer reference is rejected with an error. Applies to each alternative reference text too. | 2048 | 2048 |
“Default” is what applies when the key is absent from model.config.toml.
Neither boolean has a server-side default beyond false, so what matters in
practice is what your model ships with.
Cutoffs and acceptable confusions
This is the most useful tuning mechanism CAPT offers, and the least obvious.
Raw acoustic confidences are biased in ways that are specific to a model and a
population of speakers. A model may routinely blur m and n; young children
may produce th in a way that reads as s. Left alone, those biases show up
as wrong verdicts. Cutoffs correct them by remapping confidence to score.
Each symbol has a cutoff, the confidence at which it is exactly borderline.
The remapping is piecewise linear and pins that point to a score of 0.5:
- raw confidence in
[0, cutoff]maps to a score in[0, 0.5] - raw confidence in
[cutoff, 1]maps to a score in[0.5, 1]
So the cutoff is the dial for how strict a given phoneme is. Lowering it is
more lenient; raising it demands higher confidence before the phoneme counts as
correct. Any symbol not listed uses DefaultCutoffScore from the [Scoring]
section.
A modifier extends this with acceptable confusions. It names
alternatives (other phonemes that may also count toward the match) and a
multiplier that scales the cutoff when one of them is what was actually
heard:
"A":
cutoff: 0.50
modifier:
multiplier: 1.50
alternatives: ["O"]
"T":
cutoff: 0.50
modifier:
multiplier: 1.90
alternatives: ["s"]
Read the first entry as: "A (the vowel of lot) needs 50% confidence. O
(the vowel of thought) is close enough to count, but only if the model is
much more sure of it: 0.50 × 1.50 = 0.75." The second is stricter still: s
may stand in for T (th), but must clear 0.50 × 1.90 = 0.95, which is about
right for a contrast that many young speakers have not yet acquired.
During alignment the confidences of the reference phoneme and any listed alternatives are summed, and the scaled cutoff is applied to that sum.
The effect is a per-sound tolerance you control. Rather than one global threshold that is too strict for some phonemes and too lax for others, you set the bar sound by sound, and you do it by editing a YAML file, not by collecting data and retraining.
Scoring behavior
The [Scoring] section of model.config.toml governs the alignment itself.
The settings most likely to matter:
| Setting | What it does |
|---|---|
DefaultCutoffScore | The cutoff for any symbol not listed in cutoffs.yaml. Below 0.5 is lenient, above 0.5 is strict. |
SubCost, InsCost, DelCost | Edit-distance costs for substitution, insertion and deletion. Raising one makes the aligner prefer explanations that avoid it. |
AlignVowels | Prevents nonsensical vowel↔consonant substitutions, reporting a deletion plus an insertion instead. |
IncludeStress | Treat stress markers as distinguishing (AH0 ≠ AH1). Requires both the lexicon and the acoustic model to carry stress. |
IncludeIntraWordInsertion | Whether extra phonemes heard inside a word are attached to that word or reported separately. |
NonSpeechTokenConfidenceScaling | Weights the confidence of silence and noise tokens. Below 1.0 makes the model less willing to call something non-speech, which is useful in noisy rooms. |
CnetLinkThreshold | Discards time slices where the acoustic model put little probability on any speech token. |
MaxCandidateProns | Caps how many combinations of word pronunciations are enumerated for a reference. |
MaxCandidateProns is a performance safeguard, and the number of combinations
grows multiplicatively with the number of words that have pronunciation
variants. For long references made of common function words, the cap can be
reached. PriorityWords in the [Vocab] section controls which words get
their variants explored first, and should list your most common
multi-pronunciation function words.
Non-speech tokens
NonSpeechTokens in [Vocab] lists the symbols the acoustic model emits for
things that are not speech: silence, breath, vocal noise, unknown sounds.
These are handled specially: they are excluded from scoring rather than being
treated as mispronunciations, and they interact with CnetLinkThreshold and
NonSpeechTokenConfidenceScaling above.
Recordings made in real rooms contain plenty of this material. If you find that noisy recordings score too harshly, these settings, rather than the cutoffs, are usually the right place to look.
Adapting the acoustic model
Configuration handles a great deal, but not everything. Where a population is genuinely different from what the model was trained on (young children, a regional accent, a specific recording setup), the acoustic model itself can be adapted to it, which addresses the cause rather than compensating for it downstream.
Adaptation needs representative audio, ideally paired with the verdicts you want the system to reproduce. If your application already records both, you may have a suitable dataset without any additional collection effort. Contact us to discuss what your data supports.