FAQ

Common questions about deploying and using CAPT.

How is this different from using a general-purpose speech recognizer?

A recognizer is built to recover the intended message. That makes it actively unsuitable for assessment, because it will repair what it hears into what it assumes you meant: a mispronounced word is transcribed as the word, and the error you were trying to measure disappears. The better the language model, the more thoroughly it hides exactly the signal you need.

CAPT is anchored to a reference you supply and never substitutes expectation for observation. It also gives you things a transcript cannot:

  • Per-phoneme verdicts and scores, not just a word-level transcript.
  • Confidence taken from a full distribution over phonemes, not a single guess.
  • Timestamps for each recognized phoneme.
  • Deterministic, inspectable scoring you can tune, with no risk of a generative model inventing plausible text.
  • Operation entirely on your own hardware, on CPU.

What accuracy should I expect?

It depends on the task, the audio and the speakers, so we do not publish a single figure that would mislead you. What we would recommend instead is measuring it on your own data, against the decisions you actually care about.

If you are automating a judgement a person currently makes, the number that matters is agreement with that person on your items, not word error rate, and not any generic benchmark. See calibrating against human judgement.

Does audio leave my infrastructure?

No. CAPT runs on your own hardware: on-premise, in your private cloud, or embedded on a device. There is no call home, and no dependency on an external service at request time.

Audio and results are not stored unless you explicitly enable it. Storage is off by default; to turn it on, set a storage backend and path in capt-server.cfg.toml:

[storage]
Type = "localfs"
BasePath = "/audio"

With this enabled, each session is written to local disk as two files, the audio and the evaluation result, organized by UTC date. Everything stays on the machine you run.

Note that metadata.custom_id and metadata.custom_metadata may be written to logs and to stored results, so they should carry opaque identifiers rather than personal information.

What hardware does it need?

CAPT is CPU-only; no GPU is required. It runs on x86_64 and Arm64 / aarch64, and a statically linked build is available for minimal or embedded images.

Sizing depends on the model and on how many concurrent evaluations you need. Contact us for guidance against your target hardware.

What languages are supported?

Each model covers one language, and the model’s language is fixed at build time. US English models are available today, and models for other languages can be built. Contact us to discuss a specific language.

What audio should I send?

Use uncompressed or losslessly compressed audio (WAV or FLAC) recorded at the model’s native sample rate, which ListModels reports as attributes.sample_rate.

Audio sampled below what the model expects yields a non-fatal warning and reduced accuracy. Upsampling before sending does not help; the detail is already gone. Lossy formats such as MP3 are accepted but will cost you some accuracy.

Can more than one person be speaking in the recording?

CAPT scores the audio against the reference text; it does not separate speakers. If someone other than the intended speaker is audible, an assistant reading a prompt for instance, that speech is part of the audio being evaluated and can influence the result, usually appearing as insertions.

Where recordings may contain more than one voice, control it at capture time: record only while the intended speaker is expected to be talking, rather than across the whole interaction. If that is not possible for your setup, talk to us about the options.

Can I add my own words and pronunciations?

Yes, and this is a normal thing to do. Add them to lexicon_addenda.tsv in the model directory. You can add words the lexicon does not have, and override the pronunciations of words it does. Words in neither the lexicon nor the addenda are handled automatically by a grapheme-to-phoneme model, so invented words and non-words work without any setup.

See Tuning and Customization.

Can I make scoring stricter or more lenient?

Yes, globally or per phoneme, by editing configuration rather than retraining. DefaultCutoffScore moves the bar for everything; cutoffs.yaml sets it sound by sound, and can additionally name specific confusions that should be tolerated. See Cutoffs and acceptable confusions.

Do I have to use streaming?

No. Evaluate takes the whole audio in one request and returns one result, which is the simplest option for scoring a file. Use StreamingEvaluate when you want results while the speaker is still talking, or for audio longer than about a minute.

Are partial results guaranteed?

No. Servers are not required to produce them, and clients should not depend on them. Drive your logic from the final result, the one with is_partial unset or false, and treat partials as an enhancement for live feedback.

What does a score of 0.5 mean?

It means “exactly borderline” for that phoneme. Scores are not raw probabilities: each phoneme has a cutoff, and the confidence-to-score mapping is built so that the cutoff lands on 0.5. That is what makes 0.5 a meaningful threshold and lets different phonemes be held to different standards.

Overall utterance scores are a summary of many such phoneme scores, so an overall 0.5 does not mean “half correct”. Look at the per-token verdicts before thresholding. See Interpreting Results.