CAPT (Computer-Aided Pronunciation Training) answers a different question from a
speech recognizer. Transcribe is given audio and asked what was said. CAPT is
given audio and the text the speaker was supposed to say, and asked how
closely the audio matches it, phoneme by phoneme.
That difference matters whenever you already know the target. Reading
assessment, pronunciation practice, language learning drills and scripted-prompt
verification all share the same shape: a known reference, a spoken attempt, and a
decision about whether the attempt was good enough.
Because the reference is known, CAPT reports where the attempt diverged, not
just that it did:
Every phoneme of the reference is aligned to the audio and labeled
MATCH, SUBSTITUTION, DELETION or INSERTION.
Every aligned phoneme carries a score in [0, 1], derived from the acoustic
model’s confidence and remapped through per-phoneme thresholds you control.
Every recognized phoneme carries a start time and duration, so results can be
laid over a waveform.
Words are scored as groups of phonemes, and the utterance gets a single
overall score.
Crucially, CAPT is error-preserving. A general-purpose recognizer (and a
large end-to-end model especially) is built to recover the intended message,
so it will quietly repair a mispronunciation into the word it assumes you meant.
That behavior is exactly wrong for assessment. CAPT never substitutes what it
expects for what it heard: if the speaker said something other than the
reference, the divergence is reported.
The engine runs on your own hardware: on-premise, in your private cloud, or
embedded on a device. Audio never leaves your infrastructure.
How it works
CAPT is built on top of Cobalt Transcribe, running a phoneme-level acoustic
model:
The reference text is phonemized. Each word is looked up in a
pronunciation lexicon; words that are not in it are passed to a
grapheme-to-phoneme (G2P) model. A word may legitimately have several valid
pronunciations, and all of them are considered.
The audio is recognized into a confusion network: for each slice of
time, a probability distribution over the phonemes that might have been
spoken, rather than a single guess.
Reference and audio are aligned with a weighted edit distance, producing
the per-phoneme MATCH / SUBSTITUTION / DELETION / INSERTION labels.
Each phoneme is scored from its confidence in the confusion network,
remapped through the cutoffs you configure.
Keeping the full distribution rather than a single best guess is what makes
step 4 meaningful: the score for a phoneme reflects how much probability mass
the acoustic model actually put on it.
A typical CAPT release, provided as a compressed archive, contains a linux
binary (capt-server) for the required native CPU architecture, an
appropriate Dockerfile, and models.
Cobalt CAPT runs either locally on linux or using Docker.
Cobalt CAPT serves the CAPT gRPC API on port 2727. The configuration shipped
with a release also enables an HTTP+JSON gateway on port 8080 and operational
endpoints on port 8081; both are opt-in settings rather than server defaults,
so a config written from scratch has gRPC only.
The cobalt.license.key file is provided separately and must be copied into
the directory resulting from decompressing the archive. Please do this before
running the steps below.
Running CAPT Server Locally on Linux
./capt-server --config capt-server.cfg.toml
By default the binary assumes a configuration file named
capt-server.cfg.toml in the same directory. A different config file may be
specified using the --config argument.
On a successful start the server logs the addresses it is serving on:
The HTTP API is the quickest way to confirm a working server, provided
api.Address is set in your config as the shipped one does. It exposes the
unary calls: Version, ListModels and Evaluate. The bidirectional
StreamingEvaluate is gRPC-only, though browsers can reach it over a
websocket:
curl -s http://localhost:8080/api/capt/v1/version
{"version":"1.4.0"}
The reported version is that of the capt-server release you were given.
The id of a model in this list is what you pass as model_id when
configuring an evaluation, and attributes.sample_rate is the audio sample
rate the model expects. See Evaluation
Configurations for what to do with them.
How to Get a Copy of the CAPT Server and Models
Contact us for a release best suited
to your requirements.
The release you receive is a compressed archive (tar.bz2), generally structured
as follows:
release.tar.bz2
├── COPYING
├── README.md
├── capt-server
├── capt-server.cfg.toml
├── Dockerfile
├── api
│ ├── proto [ cobaltspeech/capt/v1/capt.proto ]
│ └── gen [ pre-generated Go and Python bindings ]
├── models
│ └── en_US-16khz
│ ├── capt [ lexicon, G2P model, cutoffs, model config ]
│ └── transcribe [ acoustic model and decoding graph ]
│
└── cobalt.license.key [ provided separately, needs to be copied over ]
The README.md file contains information about the release and instructions
for starting the server on your system.
The capt-server is the server program, configured using the
capt-server.cfg.toml file.
The Dockerfile can be used to create a container that will let you run CAPT
server on non-linux systems such as macOS and Windows.
The api directory holds the protobuf definition of the API and
pre-generated client bindings. CAPT’s proto is distributed with the release
rather than from a public repository, so this is where it comes from. See
Generating SDKs.
The models directory contains the evaluation models. Each model directory
holds two halves: the transcribe acoustic model that recognizes phonemes,
and the capt resources (pronunciation lexicon, G2P model and scoring
cutoffs) that the evaluation is built from. Both are described in Tuning
and Customization.
System Requirements
Cobalt CAPT runs on Linux, directly as a native application. You can evaluate
the product on Windows or macOS using Docker
Desktop, but we would not recommend that
setup for production.
CAPT runs on x86_64 and Arm64 / aarch64 CPUs, and a statically linked
build is available for deployment onto minimal or embedded images. Because
the engine is CPU-only and the models are small, the same release can be
deployed on a server, in a private cloud, or embedded on a device. Tell us the
target and we will provide a build for it.
Sizing depends on the model and on how many evaluations you need to run
concurrently. Please contact us for
sizing guidance against your target hardware and expected load.
To integrate Cobalt CAPT into your application, follow the next steps to
install or generate the SDK in a language of your choice.
2 - Generating SDKs
How to generate a client SDK for your project from the CAPT proto API definition.
The CAPT API is defined as a protocol buffer specification, a
proto file. This file
allows a developer to auto-generate client SDKs for a number of different
programming languages.
The CAPT proto definition, cobaltspeech/capt/v1/capt.proto, is supplied to
you as part of your release, together with pre-generated Go and Python
bindings. If you only need those two languages you can skip straight to
Using the pre-generated SDKs; to generate
bindings for another language, see Generating
SDKs.
Info
Unlike Cobalt’s other engines, the CAPT proto is not currently published in the
public cobaltspeech/proto repository.
Use the copy included with your release, and contact
us when you upgrade so that your
bindings stay in step with your server.
The relevant part of your release looks like this:
release.tar.bz2
└── api
├── proto
│ └── cobaltspeech
│ └── capt
│ └── v1
│ └── capt.proto
└── gen
├── go
│ └── cobaltspeech/capt/v1/{capt.pb.go, capt_grpc.pb.go}
└── py
└── cobaltspeech/capt/v1/{capt_pb2.py, capt_pb2_grpc.py, capt_pb2.pyi}
Using the pre-generated SDKs
Golang
Copy the api/gen/go tree into your module, or add it as a dependency, and
import the package:
You will also need the gRPC and protobuf runtime libraries:
go get google.golang.org/protobuf
go get google.golang.org/grpc
Python
The Python bindings depend on Python >= 3.8. Place api/gen/py on your
PYTHONPATH, or copy the cobaltspeech package into your project, and install
the runtime dependencies:
To generate bindings for a language we do not ship, generate them yourself from
capt.proto. We recommend using buf, a
command line tool that can generate documentation, schemas and SDK code for
many languages.
Step 1. Installing buf
COBALT="${HOME}/cobalt"mkdir -p "${COBALT}/bin"VERSION="1.72.0"URL="https://github.com/bufbuild/buf/releases/download/v${VERSION}/buf-$(uname -s)-$(uname -m)"curl -L ${URL} -o "${COBALT}/bin/buf"# Give executable permissions and add to $PATH.chmod +x "${COBALT}/bin/buf"exportPATH="${PATH}:${COBALT}/bin"
brew install bufbuild/buf/buf
Step 2. Writing a buf.gen.yaml
Create a buf.gen.yaml
next to the api/proto directory from your release. The example below
generates Go and Python; other plugins can be
added for more languages.
# Removing any previously generated files.rm -rf ./gen
# Generating code for the proto files inside the `proto` directory.buf generate proto
You should now have a gen folder containing the generated code. The latest
version of the CAPT API is v1. Import or copy the generated files into your
project as per the conventions of your language.
First, you need the address where the server is running: e.g.
host:grpc_port. By default this is localhost:2727, and it is logged to the
terminal when you start CAPT server as grpcAddr:
If you are hosting your server with Transport Layer Security (TLS) enabled,
follow the instructions under Connect with TLS. Otherwise,
follow the Default Connection method.
Default Connection
The following snippet connects to the server and queries its version, using an
“insecure” gRPC channel. This would be the case if you have just started a
local instance of CAPT server without TLS enabled.
importgrpcimportcobaltspeech.capt.v1.capt_pb2ascaptimportcobaltspeech.capt.v1.capt_pb2_grpcascapt_grpcserverAddress="localhost:2727"# Using a channel without TLS enabled.channel=grpc.insecure_channel(serverAddress)client=capt_grpc.CAPTServiceStub(channel)# Get server version.versionResp=client.Version(capt.VersionRequest())print(versionResp)# Get the list of models available on the server.modelResp=client.ListModels(capt.ListModelsRequest())formodelinmodelResp.models:print(model)
packagemainimport("context""fmt""os""google.golang.org/grpc""google.golang.org/grpc/credentials/insecure"captpb"github.com/your-org/your-module/gen/go/cobaltspeech/capt/v1")funcmain(){constserverAddress="localhost:2727"ctx,cancel:=context.WithCancel(context.Background())defercancel()// Using a channel without TLS enabled.conn,err:=grpc.NewClient(serverAddress,grpc.WithTransportCredentials(insecure.NewCredentials()))iferr!=nil{fmt.Printf("failed to dial gRPC connection: %v\n",err)os.Exit(1)}deferconn.Close()client:=captpb.NewCAPTServiceClient(conn)// Get server version.versionResp,err:=client.Version(ctx,&captpb.VersionRequest{})iferr!=nil{fmt.Printf("failed to get server version: %v\n",err)os.Exit(1)}fmt.Printf("%v\n",versionResp)// Get the list of models available on the server.modelResp,err:=client.ListModels(ctx,&captpb.ListModelsRequest{})iferr!=nil{fmt.Printf("failed to list models: %v\n",err)os.Exit(1)}for_,m:=rangemodelResp.GetModels(){fmt.Printf("%v\n",m)}}
Connect with TLS
In our recommended setup for deployment, TLS is enabled in the gRPC connection,
and clients validate the server’s SSL certificate to make sure they are talking
to the right party. This is similar to how “https” connections work in web
browsers.
TLS is enabled server-side by providing a certificate and key in
capt-server.cfg.toml:
The following snippets show how to connect to a CAPT server that has TLS
enabled.
importgrpcimportcobaltspeech.capt.v1.capt_pb2ascaptimportcobaltspeech.capt.v1.capt_pb2_grpcascapt_grpcserverAddress="capt.your-org.internal:2727"# Setup a gRPC connection with TLS. You can optionally provide your own# root certificates and private key to grpc.ssl_channel_credentials()# for mutually authenticated TLS.creds=grpc.ssl_channel_credentials()channel=grpc.secure_channel(serverAddress,creds)client=capt_grpc.CAPTServiceStub(channel)# Get server version.versionResp=client.Version(capt.VersionRequest())print(versionResp)
packagemainimport("context""crypto/tls""fmt""os""google.golang.org/grpc""google.golang.org/grpc/credentials"captpb"github.com/your-org/your-module/gen/go/cobaltspeech/capt/v1")funcmain(){constserverAddress="capt.your-org.internal:2727"// Setup a gRPC connection with TLS. You can optionally provide your own// root certificates and private key through tls.Config for mutually// authenticated TLS.tlsCfg:=tls.Config{}creds:=credentials.NewTLS(&tlsCfg)ctx,cancel:=context.WithCancel(context.Background())defercancel()conn,err:=grpc.NewClient(serverAddress,grpc.WithTransportCredentials(creds))iferr!=nil{fmt.Printf("failed to dial gRPC connection: %v\n",err)os.Exit(1)}deferconn.Close()client:=captpb.NewCAPTServiceClient(conn)versionResp,err:=client.Version(ctx,&captpb.VersionRequest{})iferr!=nil{fmt.Printf("failed to get server version: %v\n",err)os.Exit(1)}fmt.Printf("%v\n",versionResp)}
Client Authentication
In some setups it is desirable for the server to validate the clients
connecting to it, and only respond to ones it can verify. If your CAPT server
is configured to do client authentication, you will need to present the
appropriate certificate and key when connecting to it.
Note that in client-authentication mode the client still also verifies the
server’s certificate, so this setup uses mutually authenticated TLS.
creds=grpc.ssl_channel_credentials(root_certificates=root_certificates,# PEM certificate as byte stringprivate_key=private_key,# PEM client key as byte stringcertificate_chain=certificate_chain,# PEM client certificate as byte string)
// Root PEM certificate for validating a self-signed server certificate.varrootCert[]byte// Client PEM certificate and private key.varcertPem,keyPem[]bytecaCertPool:=x509.NewCertPool()ifok:=caCertPool.AppendCertsFromPEM(rootCert);!ok{fmt.Printf("unable to use given caCert\n")os.Exit(1)}clientCert,err:=tls.X509KeyPair(certPem,keyPem)iferr!=nil{fmt.Printf("unable to use given client certificate and key: %v\n",err)os.Exit(1)}tlsCfg:=tls.Config{RootCAs:caCertPool,Certificates:[]tls.Certificate{clientCert},}creds:=credentials.NewTLS(&tlsCfg)
4 - Evaluation Configurations
The options available when configuring an evaluation request.
Every evaluation begins with an
EvaluationConfig. It is the first
message of a StreamingEvaluate stream, and the config field of a unary
Evaluate request. This page describes what goes in it.
Choosing a model
model_id selects the model to evaluate against, and must be one of the id
values returned by ListModels. Each model
also reports the sample_rate it expects and its kind:
MODEL_KIND_SPEECH_EVALUATION evaluates words and sentences. The
reference_text is ordinary text.
MODEL_KIND_PHONEME_EVALUATION evaluates phonemes produced in isolation,
which a word-oriented model may not recognize correctly. The
reference_text is a sequence of phones. See Phoneme
Models.
Both kinds return the same EvaluationResult
structure, so only the reference_text you supply differs.
Reference text
reference_text is the text the speaker is expected to have said: the thing
the audio is scored against. This is the field that makes CAPT what it is: the
evaluation is anchored to it.
cfg=capt.EvaluationConfig(model_id="en_US-16khz",reference_text="WHEN THE SUNLIGHT STRIKES",)
A few properties worth knowing:
Lexicon lookup is case-insensitive, so when, When and WHEN behave
identically. The uppercase reference text used throughout these examples is
a convention, not a requirement.
Punctuation handling is a per-model setting, RemovePunctuation in the
model’s [Tokenization] config, and it is off in the en_US-16khz
model shipped today. Punctuation therefore stays attached to the word: a
reference of WHEN, THE SUNLIGHT STRIKES comes back with text of WHEN,,
and because the token no longer matches the lexicon it is phonemized by G2P,
which changed that word’s score in our own testing. Sentence-final marks are
usually harmless, but the safe habit is to send the words alone. See Tuning
and Customization.
Words not in the lexicon are handled by a G2P model, so invented words,
proper nouns and non-words can be used as reference text without any setup.
This is what makes non-word decoding tasks work. See Tuning and
Customization.
Words with several valid pronunciations are all considered, and the
variant that best fits the audio is the one reported. A speaker is not
penalized for saying “THE” as D i rather than D @.
Alternative reference text
alternative_reference_text accepts additional reference texts that would also
be acceptable. Each is aligned independently and returned in
alternative_alignments with its own
score, and the top-level score becomes the best score across the primary and
all alternatives.
cfg=capt.EvaluationConfig(model_id="en_US-16khz",reference_text="WHEN THE SUNLIGHT STRIKES",alternative_reference_text=["WHEN THE SUN LIGHTS STRIKE"],)
Note
If you need to know which candidate the audio matched, for example when
choosing between a correct answer and a set of known incorrect answers, it is
clearer to issue a separate evaluation per candidate and compare the scores
yourself. alternative_reference_text optimizes for “any of these is
acceptable”, not for classification. See Interpreting
Results.
Audio format
audio_format tells the server how to interpret the bytes you send. Two shapes
are available.
For files that carry their own header, name the container and let the server
read the rest:
cfg=capt.EvaluationConfig(model_id="en_US-16khz",reference_text="WHEN THE SUNLIGHT STRIKES",audio_format=capt.AudioFormat(audio_format_headered=capt.AUDIO_FORMAT_HEADERED_WAV,),)
Supported headered formats are WAV, MP3, FLAC and OGG Opus. Headers are
expected once, at the beginning of the stream, not in every Audio message.
audio_format may be omitted entirely for headered audio. The server then
detects the format from the header, which is why several examples in these
pages leave it out. Naming it explicitly is still recommended: it turns an
unreadable or unexpected header into a clear error instead of a detection
attempt.
Raw audio is the exception. It carries no header to detect, so
audio_format_raw must be supplied in full, and encoding must be set to
something other than AUDIO_ENCODING_UNSPECIFIED or the request is
rejected.
For raw samples with no header, describe them fully:
cfg=capt.EvaluationConfig(model_id="en_US-16khz",reference_text="WHEN THE SUNLIGHT STRIKES",audio_format=capt.AudioFormat(audio_format_raw=capt.AudioFormatRAW(encoding=capt.AUDIO_ENCODING_SIGNED,bit_depth=16,byte_order=capt.BYTE_ORDER_LITTLE_ENDIAN,sample_rate=16000,channels=1,),),)
Every server supports raw little-endian 16-bit signed samples at the model’s
own sample rate; support for the other formats depends on how the server was
built and configured.
Info
For best accuracy use uncompressed or losslessly compressed audio (WAV or
FLAC) recorded at the model’s native sample rate. Sending audio sampled below
what the model expects produces a non-fatal
EvaluationError alongside your results,
warning that accuracy may be reduced. Upsampling low-rate audio does not
recover the missing detail.
Request metadata
metadata is optional and is never used to influence evaluation. The server
may record it in logs or stored results.
cfg=capt.EvaluationConfig(model_id="en_US-16khz",reference_text="WHEN THE SUNLIGHT STRIKES",metadata=capt.EvaluationMetadata(custom_id="session-42-item-07",custom_metadata='{"item":"7","form":"A"}',),)
custom_id accepts up to 64 bytes, restricted to letters, digits, hyphens and
underscores. Useful as a tracing or correlation ID.
custom_metadata is a free-form string; a plain tag or structured data such
as JSON.
Note
Because these fields may be written to logs and to stored results, they must
not contain personal or otherwise sensitive information. Use an opaque
identifier that only your own systems can resolve back to a person.
Describes how to stream audio to CAPT server and receive evaluation results.
CAPT offers two ways to submit audio.
StreamingEvaluate is the primary method. It is bidirectional: you send
audio as it is captured and receive results as the audio is processed. Use it
for anything interactive, and for audio of any length.
Evaluate is a unary convenience method that takes the whole audio in a
single request and returns a single result. It is intended for short audio,
less than a minute, and is the simplest way to score an audio file.
In a StreamingEvaluate stream, the first message must contain the config,
and every subsequent message must contain audio.
An empty Audio message ends the stream. CAPT server treats it as the end
of the audio: it finishes processing, sends the final result, and accepts
nothing further, so a later send fails. Over gRPC you normally do not need it,
because half-closing the stream (CloseSend in Go, exhausting the request
iterator in Python) says the same thing. It matters for
browser clients, where a websocket has no
equivalent half-close.
Do not send one to mean “no audio yet”. If it arrives before the audio is
complete, what you get depends on the format, and a client’s state machine
needs to handle both:
You always receive a final result for the audio sent so far, with
is_partial false. It reflects the truncated input, not the utterance you
intended, so it will usually be full of deletions.
With a headered format the stream then fails. The header promised more
data than arrived, so the server reports a container error such as
wav: unexpected EOF after the final result. This happens whether or not
you send anything afterwards.
With raw audio the stream ends cleanly, because there is no container to
leave incomplete. The result is still only of the audio you sent.
In other words: a final result is not by itself evidence that the whole
utterance was processed. Send the empty message only once the recording is
complete, and treat an error after a final result as “the result is
incomplete”, not “the result is invalid”.
The examples below stream a WAV file in 100 ms chunks and print each result as
it arrives.
importgrpcimportcobaltspeech.capt.v1.capt_pb2ascaptimportcobaltspeech.capt.v1.capt_pb2_grpcascapt_grpcserverAddress="localhost:2727"channel=grpc.insecure_channel(serverAddress)client=capt_grpc.CAPTServiceStub(channel)# Get the list of models on the server and use the first one.modelResp=client.ListModels(capt.ListModelsRequest())modelID=modelResp.models[0].idcfg=capt.EvaluationConfig(model_id=modelID,reference_text="WHEN THE SUNLIGHT STRIKES",audio_format=capt.AudioFormat(audio_format_headered=capt.AUDIO_FORMAT_HEADERED_WAV,),)# The first request must contain only the configuration; subsequent# requests carry audio bytes. A generator is a convenient way to do this.defstream(cfg,audio,bufferSize=3200):yieldcapt.StreamingEvaluateRequest(config=cfg)data=audio.read(bufferSize)whilelen(data)>0:yieldcapt.StreamingEvaluateRequest(audio=capt.Audio(data=data))data=audio.read(bufferSize)withopen("test.wav","rb")asaudio:forrespinclient.StreamingEvaluate(stream(cfg,audio)):ifresp.HasField("error"):print(f"warning: {resp.error.message}")result=resp.evaluation_resultprint(f"partial={result.is_partial} score={result.score:.4f}")# Only act on final results; partials will still change.ifnotresult.is_partial:forwordinresult.alignments:print(f" {word.text}: "+" ".join(f"{t.reference}={t.score:.2f}"fortinword.tokens))
packagemainimport("context""fmt""io""os""google.golang.org/grpc""google.golang.org/grpc/credentials/insecure"captpb"github.com/your-org/your-module/gen/go/cobaltspeech/capt/v1")funcmain(){constserverAddress="localhost:2727"audio,err:=os.ReadFile("test.wav")iferr!=nil{fmt.Printf("failed to read audio: %v\n",err)os.Exit(1)}conn,err:=grpc.NewClient(serverAddress,grpc.WithTransportCredentials(insecure.NewCredentials()))iferr!=nil{fmt.Printf("failed to dial gRPC connection: %v\n",err)os.Exit(1)}deferconn.Close()client:=captpb.NewCAPTServiceClient(conn)stream,err:=client.StreamingEvaluate(context.Background())iferr!=nil{fmt.Printf("failed to open stream: %v\n",err)os.Exit(1)}// The first message must contain the config.iferr:=stream.Send(&captpb.StreamingEvaluateRequest{Request:&captpb.StreamingEvaluateRequest_Config{Config:&captpb.EvaluationConfig{ModelId:"en_US-16khz",ReferenceText:"WHEN THE SUNLIGHT STRIKES",AudioFormat:&captpb.AudioFormat{AudioFormat:&captpb.AudioFormat_AudioFormatHeadered{AudioFormatHeadered:captpb.AudioFormatHeadered_AUDIO_FORMAT_HEADERED_WAV,},},},},});err!=nil{fmt.Printf("failed to send config: %v\n",err)os.Exit(1)}// Send audio in the background while results are read below.gofunc(){constchunk=3200// 100ms of 16kHz 16-bit mono audio.// min() is a builtin from Go 1.21; on older toolchains compute the// bound explicitly.fori:=0;i<len(audio);i+=chunk{end:=min(i+chunk,len(audio))iferr:=stream.Send(&captpb.StreamingEvaluateRequest{Request:&captpb.StreamingEvaluateRequest_Audio{Audio:&captpb.Audio{Data:audio[i:end]},},});err!=nil{return}}_=stream.CloseSend()}()for{resp,err:=stream.Recv()iferr==io.EOF{break}iferr!=nil{fmt.Printf("failed to receive result: %v\n",err)os.Exit(1)}ife:=resp.GetError();e!=nil{fmt.Printf("warning: %v\n",e.GetMessage())}result:=resp.GetEvaluationResult()fmt.Printf("partial=%v score=%.4f\n",result.GetIsPartial(),result.GetScore())// Only act on final results; partials will still change.if!result.GetIsPartial(){for_,word:=rangeresult.GetAlignments(){fmt.Printf(" %s:",word.GetText())for_,t:=rangeword.GetTokens(){fmt.Printf(" %s=%.2f",t.GetReference(),t.GetScore())}fmt.Println()}}}}
Streaming a 1.5 second recording of “when the sunlight strikes” against that
reference produces a short sequence of results:
While audio is still arriving, the server emits partial results, marked
is_partial = true. A partial reflects the audio received so far, and any part
of it may change as more audio arrives. In the sequence above the score rises
as the last word is heard.
When the stream ends, the server emits a final result with is_partial = false. That result is stable.
Use partials for live feedback: highlighting words as they are read,
showing a provisional score, driving a progress indicator.
Use the final result for any decision you record. Scoring an item on a
partial risks scoring it on half a word.
Info
Servers are not required to produce partial results, and clients should not
depend on receiving them. Always drive your logic from the final result and
treat partials as an enhancement.
Speech endpointing
A result may set is_speech_endpoint = true to indicate the server believes
the speaker has finished, typically after a period of silence. This is useful
for closing the microphone automatically once an answer has been given.
Endpoint detection is optional, and clients must be robust to the field never
being set. Do not treat it as the signal that results are final; that is what
is_partial = false is for.
Non-fatal errors
Both response types carry an optional
EvaluationError beside the result. It
reports conditions that degrade quality but do not stop processing, most
commonly audio sampled at a lower rate than the model expects. The server
continues, and you may keep streaming. Log these rather than aborting; they
usually point at a recording pipeline that needs attention.
Evaluating a whole file at once
For short audio, Evaluate avoids the streaming machinery entirely. Over the
HTTP+JSON gateway that is a single POST:
The data field is the audio file, base64 encoded. See Interpreting
Results for what comes back.
6 - Interpreting Results
How to read an EvaluationResult and turn it into a decision.
This is the page worth reading carefully. Getting audio into CAPT is
straightforward; deciding what its output means for your application is where
the design work is.
EvaluationResult
├── score overall score for the utterance, 0..1
├── is_partial true while more audio may still change this
├── is_speech_endpoint true if the speaker appears to have finished
├── alignments [ ]AlignedWord
│ └── AlignedWord
│ ├── text the reference word
│ ├── start_time_ms when the word starts in the audio
│ ├── duration_ms how long it lasts
│ └── tokens [ ]AlignedToken (the phonemes)
│ └── AlignedToken
│ ├── kind MATCH | SUBSTITUTION | DELETION | INSERTION
│ ├── reference the expected phoneme
│ ├── score 0..1 for this phoneme
│ └── hypotheses [ ]AlignmentHypothesis (what was heard)
│ └── token, confidence, start_time_ms, duration_ms
└── alternative_alignments the same, once per alternative reference text
The reference text drives the structure: your words appear as AlignedWord
entries in order, and within each one there is an AlignedToken per expected
phoneme.
Note
alignments is not a one-to-one list of your reference words. Insertions
are reported as extra AlignedWord entries with an empty text, positioned
where they were heard, so the list can be longer than
reference_text.split() and the indices do not line up. Match on text, or
skip entries whose text is empty, rather than zipping the two lists together.
See Insertions.
The four alignment kinds
Every expected phoneme gets exactly one of four verdicts.
Kind
Meaning
reference
hypotheses
score
MATCH
The expected phoneme was heard with enough confidence
set
populated
0.5 to 1.0
SUBSTITUTION
The expected phoneme was not heard with enough confidence: either something else was heard in its place, or the phoneme itself was heard but below its cutoff
set
populated
0 to below 0.5
DELETION
The phoneme was not heard at all
set
empty
0
INSERTION
Something extra was heard that the reference does not account for
empty
populated
0
Note
A SUBSTITUTION does not imply a score of 0. When the expected phoneme
was heard but fell short of its cutoff, the score is its remapped confidence,
which is non-zero but below 0.5. The k scoring 0.49 in the example below
is exactly this case. Only when the phoneme is absent from the hypotheses
entirely does the score reach 0. Decide on kind, or on a threshold, but do
not infer one from the other.
A real example makes this concrete. Below is the result of scoring a recording
of “when the sunlight strikes” against a deliberately wrong reference,
“WHEN THE MOONLIGHT STRIKES” (phonemes are X-SAMPA):
overall score: 0.5157
WHEN MATCH w (1.00) MATCH E (0.99) MATCH n (1.00)
THE MATCH D (0.87) MATCH @ (0.67)
MOONLIGHT SUB m (0.00) SUB u (0.00) MATCH n (0.65)
MATCH l (0.78) MATCH aI (0.82) MATCH t (0.73)
STRIKES MATCH s (0.79) DEL t (0.00) DEL r\ (0.00)
DEL aI (0.00) SUB k (0.49) DEL s (0.00)
Read that as a diagnosis, not just a number. The first two words were said as
expected. In MOONLIGHT the leading m u was not there (the speaker said
s V n), but the shared tail n l aI t matched, so those phonemes score well.
STRIKES shows what happens when the audio has already run out: the reference
still has phonemes to account for, and they come back as deletions.
The overall score of 0.52 is the average of that mixture. This is why an
overall score alone is rarely the right thing to threshold. 0.52 here does
not mean “roughly half right”, it means “two words right and two words wrong”.
Insertions
Insertions do not belong to any reference word, so they are reported in their
own AlignedWord entries with an empty text, positioned in the sequence
where they were heard. Scoring the same audio against the single word
“ELEPHANT” shows this:
overall score: 0.3966
(insertion) INS (heard: w)
ELEPHANT MATCH E (0.99) SUB l (0.00) MATCH @ (0.67)
SUB f (0.00) SUB n= (0.00) MATCH t (0.73)
(insertion) INS (heard: k) INS (heard: s)
The speaker said far more than the reference accounted for, and the leftover
audio at each end is reported as insertions rather than being silently
discarded.
Info
Whether phonemes inserted within a word are attached to that word or reported
separately is controlled by the model’s IncludeIntraWordInsertion setting.
See Tuning and Customization.
Where the numbers come from
Phoneme scores
The acoustic model does not commit to a single phoneme per time slice. It
produces a distribution, a confusion network, and the score for an expected
phoneme is derived from how much probability mass landed on it.
You can see this in the hypotheses list. Here is a single MATCH from a
real result:
The model heard n with 0.963 confidence, but also considered m, n= and
N. The reported score of 0.98 is that confidence remapped through the
model’s cutoff for n. See Tuning and
Customization
for how that remapping works and how to change it.
The practical consequence: a score is a calibrated quantity, not a raw
probability. A cutoff is chosen so that a score of exactly 0.5 sits at the
boundary between acceptable and unacceptable for that phoneme. That is what
makes 0.5 a meaningful place to threshold, and it is why different phonemes
can be held to different standards.
Word and utterance scores
An AlignedWord does not carry its own score field; a word’s quality is the
scores of its phonemes. Compute whatever summary suits your task: the mean,
the minimum, or the fraction of phonemes that came back MATCH.
The top-level score is the overall score for the utterance: the mean of the
scores of the reference tokens. If alternative reference texts were
supplied, it is the best score across the primary and all alternatives.
Note
Insertions are not counted in the overall score. The score answers “how
well was the reference produced”, not “was anything else said”. A recording
that contains the target words plus a great deal of unrelated speech can still
score 1.0.
If extraneous speech should count against the speaker (an assistant reading a
prompt, or a child answering twice), inspect the insertion entries yourself and
apply your own rule. The INSERTION tokens carry timestamps, so you can also
measure how much of the audio they account for.
Timestamps
Timestamps live on the hypotheses, not on the token. An AlignedToken has
no time fields of its own; start_time_ms and duration_ms are on each
AlignmentHypothesis inside it, and
on the enclosing AlignedWord.
This has one consequence that catches people out:
Note
A DELETION has an empty hypotheses list and therefore no timestamp at
all. There is no audio to point at; that is what a deletion means. Any
timeline, waveform or spectrogram view must handle tokens that have no time
span, rather than assuming every phoneme can be drawn.
Likewise, a word made up entirely of deletions has nothing to anchor to, and
its start_time_ms and duration_ms are reported as 0.
Turning results into a decision
A pass/fail mark for one item
If your application needs a yes/no verdict (did the speaker say the target
correctly?), do not reach for the overall score first. Consider what the task
actually requires:
Strict, verbatim tasks. Any divergence is a failure. Require every token
to be MATCH:
This is exactly right when the rule is “any change at all, however minor, is
wrong”, because substitutions, deletions and insertions all break it.
Tolerant tasks. Some divergence is acceptable. A slightly indistinct
consonant should not fail an otherwise good attempt. Threshold on the
proportion of matched phonemes, or on the mean phoneme score, rather than
demanding a clean sweep. Note that a tolerant rule built on the overall
score, or on reference tokens alone, will not notice extra speech: decide
separately whether insertions should fail the item.
Targeted tasks. Only some phonemes matter: the contrast the item is
testing. Score only those tokens and ignore the rest.
Choosing between several candidate answers
When you have a known correct answer and a set of known incorrect answers, and
you need to know which one was said, evaluate the audio once per candidate
and compare the resulting scores:
This gives you a score per candidate, so you can see not just the winner but
the margin, so you can reject the result as unclear when two candidates score
alike. Packing the candidates into alternative_reference_text instead would
collapse them into a single best score and lose exactly that information.
Calibrating against human judgement
If you are automating a decision a person currently makes, the metric that
matters is agreement with that person, not any internal notion of accuracy.
The recommended approach is to collect audio alongside the human verdict, run
CAPT over it, then sweep your decision rule across the collected set and count
true positives, true negatives, false positives and false negatives at each
setting. That tells you where to set the threshold, and, just as importantly,
what the residual disagreement rate is and which way it leans. A rule that is
wrong in the safe direction for your use case is often better than one that is
wrong less often overall.
Per-phoneme cutoffs can then correct systematic biases that a single global
threshold cannot; see Tuning and
Customization.
Alternative alignments
If you supplied alternative_reference_text, each alternative comes back in
alternative_alignments as an
AlternativeAlignment carrying its
own reference_text, score, and full alignments tree with the same
structure described above.
The top-level score is the best across the primary and the alternatives, so
if you only care whether any acceptable rendering was produced, the top-level
score is enough. If you care which one, see above.
7 - Phoneme Models
Evaluating individual phonemes produced in isolation.
CAPT models come in two kinds, reported as kind by
ListModels.
Word models (MODEL_KIND_SPEECH_EVALUATION) evaluate phonemes in word
context. The reference_text is ordinary text, and this is what you want
for reading words, sentences and passages.
Phoneme models (MODEL_KIND_PHONEME_EVALUATION) evaluate phonemes
produced in isolation. The reference_text is a sequence of phones.
The distinction matters because a sound produced on its own is acoustically
quite different from the same sound inside a word. A model trained on connected
speech often mis-recognizes an isolated phoneme, because nothing in its
training looked like that. If your task asks a speaker to produce a single
sound (“say the /sh/ sound”), a phoneme model is substantially more reliable.
Both kinds are used through exactly the same API calls and return the same
EvaluationResult. Only the reference_text
differs.
Reference text for phoneme models
For a phoneme model, reference_text is one or more phones separated by
whitespace:
# A single phone.cfg=capt.EvaluationConfig(model_id=phonemeModelID,reference_text="AA")# A sequence of phones.cfg=capt.EvaluationConfig(model_id=phonemeModelID,reference_text="Z AE P")
Phone groups
Some phones are only meaningfully produced as a pair, and the model treats such
a pair as a single unit called a phone group. Phone groups are written with the
members joined by a period:
# A phone group.cfg=capt.EvaluationConfig(model_id=phonemeModelID,reference_text="AO.NG")# Mixed with ordinary phones.cfg=capt.EvaluationConfig(model_id=phonemeModelID,reference_text="F AE.NG")
For all practical purposes a phone group behaves as its own distinct phone: it
is matched, substituted or deleted as a single token, and receives a single
score.
The phone set
The phones a model accepts are specific to that model, both the inventory
and which groups exist. Supplying a phone the model does not know is an error,
so check against the set for the model you were given. A typical US English
phoneme model accepts:
AA AA.R AE AE.NG AH AH.NG AO AO.NG
AO.R AW AY B CH D DH EH
EH.R ER EY F G HH IH IH.NG
IH.R IY JH K L M N NG
OW OY P R S SH T TH
UH UW V W Y Y.UW Z ZH
Info
Note that word models and phoneme models may use different phone alphabets. A
word model may report X-SAMPA (w E n, r\, aI) while a phoneme model uses
the ARPABET-style symbols above. Do not assume a phone string from one is valid
in the other.
Reading the results
Results have the same structure as any other evaluation. For each phone in the
reference_text there is an aligned token:
ALIGNMENT_KIND_MATCH means the phone was found in the audio with a high degree
of confidence.
ALIGNMENT_KIND_SUBSTITUTION means a different phone was recognized in its
place, or the phone was recognized only with low confidence.
ALIGNMENT_KIND_DELETION means the phone was not recognized at all.
Anything else recognized in the audio that does not align to the reference is
reported as insertions, grouped separately.
Isolated-phoneme audio very often contains more than the phoneme itself: a
speaker clearing their throat, a lead-in, or the assessor’s prompt. Those show
up as insertions, which is why the example below scores 1.0 despite three
extra tokens: the reference phone ZH was produced correctly, and the
surrounding material is reported rather than being allowed to affect the score
of the phone you asked about.
If extraneous audio should count against the speaker in your task, inspect
the insertion entries and apply your own rule; the information is there either
way.
8 - Tuning and Customization
Pronunciations, per-phoneme thresholds and scoring behavior.
A general-purpose speech model is a fixed object: you send it audio and accept
what it returns. CAPT is deliberately not that. Because the scoring stage is
separate from the acoustic model, a great deal of behavior can be adjusted for
your task without retraining anything, by changing which pronunciations
count as correct, and how strictly each sound is judged.
Everything on this page lives in the model directory shipped with your release,
and takes effect when the server restarts.
models/en_US-16khz/capt/
├── model.config.toml scoring behavior and paths to the files below
├── lexicon.json the primary pronunciation dictionary
├── lexicon_addenda.tsv your additions and overrides
├── cutoffs.yaml per-phoneme thresholds
└── g2p.ort grapheme-to-phoneme model for unknown words
The model config points at each of these by path, so the names above are the
convention rather than a requirement. Older model directories may carry the
cutoffs as a plain-text cutoffs.txt instead; the meaning is identical and
CutoffsPath says which file is in use.
Pronunciations
The lexicon and its addenda
The lexicon maps words to their phoneme sequences. A word may have several
valid pronunciations, and CAPT considers all of them, reporting whichever best
fits the audio, so a speaker is not penalized for a legitimate variant.
lexicon_addenda.tsv is loaded after the primary lexicon and is the file
you should edit. Entries in it add new words and override existing ones,
leaving the shipped lexicon intact so that upgrades stay clean.
This is the mechanism to reach for when your task defines what counts as an
acceptable pronunciation. If an item specifies that a target may be produced
two ways, add both and each will be accepted as a match:
MOSP m A s p
MOSP m oU s p
Those two entries accept mosp pronounced to rhyme with wasp, or with
soap.
Note
Every phone symbol on this page belongs to the model’s own alphabet, and
the files in a model directory all use the same one. The en_US-16khz word
model used in these examples is X-SAMPA, which is why its addenda read
m A s p and its cutoffs are keyed "A" rather than AA. A phoneme model
with an ARPABET inventory would use AA in both files instead. Check the phone
set for the model you were given before editing either file: a symbol the model
does not know matches nothing, and fails silently rather than erroring.
Out-of-vocabulary words and G2P
Words that are in neither the lexicon nor the addenda are passed to a
grapheme-to-phoneme model, which predicts a pronunciation from spelling. This
happens automatically and requires no configuration.
It means invented words, proper nouns and non-words can be used as
reference_text with no setup at all, which is useful for decoding tasks built on
pseudo-words, where the whole point is that the string is not a real word.
Info
G2P is a prediction, not a definition. Where you know the pronunciations that
should be accepted, because your task specifies them, put them in the addenda
rather than relying on what G2P infers from the spelling. Reserve G2P for the
long tail you have not enumerated.
Tokenization
[Tokenization] in model.config.toml controls how reference text is
pre-processed:
Setting
What it does
Default
In en_US-16khz
RemovePunctuation
Strips punctuation before lexicon lookup. With it off, punctuation stays attached to the word and is phonemized as part of it.
off
off
UseSentenceTokenizer
Splits a long reference into sentences before tokenizing.
off
on
MaxInputLength
Safeguard on the length of a reference text. Counted in characters, not words; a longer reference is rejected with an error. Applies to each alternative reference text too.
2048
2048
“Default” is what applies when the key is absent from model.config.toml.
Neither boolean has a server-side default beyond false, so what matters in
practice is what your model ships with.
Cutoffs and acceptable confusions
This is the most useful tuning mechanism CAPT offers, and the least obvious.
Raw acoustic confidences are biased in ways that are specific to a model and a
population of speakers. A model may routinely blur m and n; young children
may produce th in a way that reads as s. Left alone, those biases show up
as wrong verdicts. Cutoffs correct them by remapping confidence to score.
Each symbol has a cutoff, the confidence at which it is exactly borderline.
The remapping is piecewise linear and pins that point to a score of 0.5:
raw confidence in [0, cutoff] maps to a score in [0, 0.5]
raw confidence in [cutoff, 1] maps to a score in [0.5, 1]
So the cutoff is the dial for how strict a given phoneme is. Lowering it is
more lenient; raising it demands higher confidence before the phoneme counts as
correct. Any symbol not listed uses DefaultCutoffScore from the [Scoring]
section.
A modifier extends this with acceptable confusions. It names
alternatives (other phonemes that may also count toward the match) and a
multiplier that scales the cutoff when one of them is what was actually
heard:
Read the first entry as: "A (the vowel of lot) needs 50% confidence. O
(the vowel of thought) is close enough to count, but only if the model is
much more sure of it: 0.50 × 1.50 = 0.75." The second is stricter still: s
may stand in for T (th), but must clear 0.50 × 1.90 = 0.95, which is about
right for a contrast that many young speakers have not yet acquired.
During alignment the confidences of the reference phoneme and any listed
alternatives are summed, and the scaled cutoff is applied to that sum.
The effect is a per-sound tolerance you control. Rather than one global
threshold that is too strict for some phonemes and too lax for others, you set
the bar sound by sound, and you do it by editing a YAML file, not by
collecting data and retraining.
Scoring behavior
The [Scoring] section of model.config.toml governs the alignment itself.
The settings most likely to matter:
Setting
What it does
DefaultCutoffScore
The cutoff for any symbol not listed in cutoffs.yaml. Below 0.5 is lenient, above 0.5 is strict.
SubCost, InsCost, DelCost
Edit-distance costs for substitution, insertion and deletion. Raising one makes the aligner prefer explanations that avoid it.
AlignVowels
Prevents nonsensical vowel↔consonant substitutions, reporting a deletion plus an insertion instead.
IncludeStress
Treat stress markers as distinguishing (AH0 ≠ AH1). Requires both the lexicon and the acoustic model to carry stress.
IncludeIntraWordInsertion
Whether extra phonemes heard inside a word are attached to that word or reported separately.
NonSpeechTokenConfidenceScaling
Weights the confidence of silence and noise tokens. Below 1.0 makes the model less willing to call something non-speech, which is useful in noisy rooms.
CnetLinkThreshold
Discards time slices where the acoustic model put little probability on any speech token.
MaxCandidateProns
Caps how many combinations of word pronunciations are enumerated for a reference.
Note
MaxCandidateProns is a performance safeguard, and the number of combinations
grows multiplicatively with the number of words that have pronunciation
variants. For long references made of common function words, the cap can be
reached. PriorityWords in the [Vocab] section controls which words get
their variants explored first, and should list your most common
multi-pronunciation function words.
Non-speech tokens
NonSpeechTokens in [Vocab] lists the symbols the acoustic model emits for
things that are not speech: silence, breath, vocal noise, unknown sounds.
These are handled specially: they are excluded from scoring rather than being
treated as mispronunciations, and they interact with CnetLinkThreshold and
NonSpeechTokenConfidenceScaling above.
Recordings made in real rooms contain plenty of this material. If you find that
noisy recordings score too harshly, these settings, rather than the cutoffs,
are usually the right place to look.
Adapting the acoustic model
Configuration handles a great deal, but not everything. Where a population is
genuinely different from what the model was trained on (young children, a
regional accent, a specific recording setup), the acoustic model itself can be
adapted to it, which addresses the cause rather than compensating for it
downstream.
Adaptation needs representative audio, ideally paired with the verdicts you
want the system to reproduce. If your application already records both, you may
have a suitable dataset without any additional collection effort. Contact
us to discuss what your data supports.
9 - Browser Integration
Calling CAPT from a web browser over HTTP and websockets.
Native applications should use gRPC directly, as shown throughout these docs.
Browsers cannot open a bidirectional gRPC stream, so CAPT server also exposes
an HTTP+JSON gateway and a websocket endpoint for StreamingEvaluate.
The HTTP API is not a property of the server itself: it is served only when
api.Address is set. The configuration shipped with a release sets it to
:8080, which is why the getting started page can curl it
straight away. If you are writing a config from scratch, enable it explicitly
in capt-server.cfg.toml:
[server.http]api.Address=":8080"## Serve the built-in web demo alongside the API.api.EnableWebDemo=true## Only needed when the page is served from a different origin## than the API, e.g. during local development.# api.EnableCORS = true
JSON conventions
The gateway emits JSON using protobuf field names, so:
Field names are snake_case, exactly as in the proto:
evaluation_result, start_time_ms, alternative_reference_text.
Enums are full strings, not numbers: "ALIGNMENT_KIND_MATCH",
"AUDIO_FORMAT_HEADERED_WAV".
Unpopulated fields are emitted, so a field being present does not imply
it was set.
64-bit integers are strings, per the protobuf JSON mapping:
"start_time_ms": "570", not 570. Parse these before doing arithmetic.
bytes fields are base64, so audio is base64-encoded into
{"audio": {"data": "..."}}.
Unary calls over HTTP
The two metadata calls and Evaluate are available as ordinary HTTP requests:
StreamingEvaluate is served over a websocket that carries the same messages
as the gRPC stream, one JSON object per websocket message.
The sequence mirrors the gRPC one: the first message must be the config,
and every message after it carries audio.
constws=newWebSocket(`wss://${location.host}/api/capt/v1/streaming-evaluate`);ws.onopen=()=>{// First message: the configuration.
ws.send(JSON.stringify({config:{model_id:'en_US-16khz',reference_text:'WHEN THE SUNLIGHT STRIKES',audio_format:{audio_format_headered:'AUDIO_FORMAT_HEADERED_WAV'}}}));};ws.onmessage=(event)=>{constmsg=JSON.parse(event.data);// The gateway reports a stream-level failure at the top level; a non-fatal
// EvaluationError is a field of the response, so it sits inside the envelope.
if(msg.error!=null){console.error('stream failed:',msg.error);return;}if(msg.result?.error!=null){console.warn('non-fatal:',msg.result.error.message);}constresult=msg.result?.evaluation_result;if(result==null){return;}if(result.is_partial){showProvisional(result);// live highlighting
}else{recordFinal(result);// the result you act on
}};
Note
Responses over the websocket are wrapped in a result envelope,
{"result": {"evaluation_result": {...}, "error": {...}}}, which the plain
gRPC stream does not have. Read msg.result.evaluation_result, not
msg.evaluation_result.
The envelope holds two different errors, and they are not interchangeable. A
non-fatal EvaluationError is a field of
the response, so it arrives at msg.result.error and processing continues.
A top-level msg.error is the gateway’s stream-level failure, with a
different shape, and the stream is over.
Sending audio means base64-encoding each chunk:
// Spreading a whole buffer into String.fromCharCode overflows the call stack
// on anything but small inputs, so encode in fixed-size blocks.
functiontoBase64(buf){constbytes=newUint8Array(buf);constblock=0x8000;letbinary='';for(leti=0;i<bytes.length;i+=block){binary+=String.fromCharCode.apply(null,bytes.subarray(i,i+block));}returnbtoa(binary);}asyncfunctionsendChunk(blob){ws.send(JSON.stringify({audio:{data:toBase64(awaitblob.arrayBuffer())}}));}// A websocket has no half-close, so an empty audio message is how you say
// "that is all the audio". Send it only once the recording is complete.
functionendStream(){ws.send(JSON.stringify({audio:{data:''}}));}
The same helper is what you want for the whole-recording POST described
below; a complete take is far past the size
at which the one-line spread form fails.
Capturing audio in the browser
The audio you send must match the audio_format you declared. The
MediaRecorder API defaults to a compressed format that CAPT does not accept,
so a browser client generally needs an encoder that can emit WAV chunks while
recording, rather than only at the end.
Two constraints are worth designing around:
Send small, regular chunks. Around 100-250 ms keeps partial results
flowing smoothly. Larger chunks make feedback feel laggy.
Declare the header once. For headered formats the header belongs at the
start of the stream, not on every chunk.
If you do not need live feedback, it is considerably simpler to record the
whole utterance, then POST it to /api/capt/v1/evaluate as a single base64
blob.
The built-in demo
With api.EnableWebDemo = true, the server serves a demo application at the
HTTP address, by default http://localhost:8080. It records from the
microphone or accepts an uploaded file, streams it over the websocket described
above, and renders the per-word and per-phoneme results.
It is a useful way to sanity-check a server, a model and a piece of audio
before writing any client code.
10 - API Reference
Detailed reference for API requests and types.
The API is defined as a protobuf spec, so native bindings can be generated in any language
with gRPC support. We recommend using buf
to generate the bindings.
This section of the documentation is auto-generated from the protobuf spec.
The service contains the methods that can be called, and the “messages” are the data structures
(objects, classes or structs in the generated code, depending on the language) passed to and from the methods.
Performs synchronous speech evaluation by receiving results after all
audio has been sent and processed. It is expected that this request be
typically used for short audio content: less than a minute long. For
longer content, the StreamingEvaluate method should be preferred.
Performs bidirectional streaming for evaluating speech by receiving
results while sending audio. This method is only available via GRPC and
not via HTTP+JSON. However, a web browser may use websockets to use this
service.
Messages
If two or more fields in a message are labeled oneof,
then each method call using that message must have exactly one of the fields populated
If a field is labeled repeated, then the generated code will accept an array (or struct, or list depending on the language).
AlignedToken
AlignedToken contains a single reference token and the alternate hypotheses
recognized from the audio that could be aligned with the reference token.
If kind == ALIGNMENT_KIND_MATCH or ALIGNMENT_KIND_SUBSTITUTION, both the
reference token and hypotheses list will be populated. The score will be
based on the confidence value of the reference token within the hypotheses.
In the case of a substitution, the hypotheses list may not contain the actual
reference token at all, and the score will be 0.
If kind == ALIGNMENT_KIND_DELETION, the hypotheses will be an empty list and
score will be 0.
If kind == ALIGNMENT_KIND_INSERTION, the reference token will be an empty
string and score will be 0.
Fields
kind (AlignmentKind )
The kind of alignment for this token.
score (float )
The score for this token, between 0 and 1 inclusive, based on the
confidence value of the token within the recognized token hypotheses from
the audio.
hypotheses (AlignmentHypothesis repeated)
The recognized tokens from the audio that have been aligned to this
reference token.
AlignedWord
AlignedWord contains the aligned tokens within a single word in the reference
text.
confidence (float )
The confidence with which this token was recognized in the audio, between
0 and 1 inclusive.
start_time_ms (uint64 )
The timestamp at which this token starts in the audio, in milliseconds.
duration_ms (uint64 )
The duration of this token in the audio, in milliseconds.
AlternativeAlignment
AlternativeAlignment contains alignments against alternative reference
text(s) if specified in the EvaluationConfig.
Fields
reference_text (string )
The alternative reference text.
score (float )
Evaluation score, between 0 and 1 inclusive, with 1 indicating a perfect
match between the alternative reference text and what was said in the
audio.
alignments (AlignedWord repeated)
Alignment between the alternative reference text tokens and recognized
tokens.
Depending on how they are configured, server instances of this service may
not support all the formats provided in the API. One format that is
guaranteed to be supported is the RAW format with little-endian 16-bit
signed samples with the sample rate matching that of the model being
requested.
Fields
oneof audio_format.audio_format_raw (AudioFormatRAW )
Audio is raw data without any headers
oneof audio_format.audio_format_headered (AudioFormatHeadered )
Audio has a self-describing header. Headers are expected to be sent
at the beginning of the entire audio file/stream, and not in every
Audio message.
The default value of this type is AUDIO_FORMAT_HEADERED_UNSPECIFIED.
If this value is used, the server may attempt to detect the format of
the audio. However, it is recommended that the exact format be
specified.
AudioFormatRAW
Details of audio in raw format
Fields
encoding (AudioEncoding )
Encoding of the samples. It must be specified explicitly and using the
default value of AUDIO_ENCODING_UNSPECIFIED will result in an error.
bit_depth (uint32 )
Bit depth of each sample (e.g. 8, 16, 24, 32, etc.). This is a required
field.
byte_order (ByteOrder )
Byte order of the samples. This field must be set to a value other than
BYTE_ORDER_UNSPECIFIED when the bit_depth is greater than 8.
sample_rate (uint32 )
Sampling rate in Hz. This is a required field.
channels (uint32 )
Number of channels present in the audio. E.g.: 1 (mono), 2 (stereo), etc.
This is a required field.
EvaluateRequest
The top-level message sent by the client for the Evaluate method. Both the
EvaluationConfig and Audio fields are required. The entire audio data
must be sent in one request. If your audio data is larger, please use the
StreamingEvaluate call.
The message returned by the server for the Evaluate method.
Fields
evaluation_result (EvaluationResult )
Response from the server. The kind of result depends on the kind of the
model that was chosen in the EvaluationConfig.
error (EvaluationError )
A non-fatal error message. If a server encountered a non-fatal error when
processing the request, it will be returned in this message.
The server will continue to process audio and produce further results.
Clients can continue streaming audio even after receiving these messages.
This error message is meant to be informational.
An example of when these errors maybe produced: audio is sampled at a
lower rate than expected by model, producing possibly less accurate
results.
This field will be unset if there is no error to report.
EvaluationConfig
Configuration for a StreamingEvaluateRequest.
Fields
model_id (string )
ID of the model to use. A list of supported IDs can be found using the
ListModels call.
audio_format (AudioFormat )
Format of the audio to be sent.
reference_text (string )
Reference Text that is expected to be contained in the audio.
alternative_reference_text (string repeated)
Alternative reference text(s) that are also acceptable if recognized in
the audio.
metadata (EvaluationMetadata )
This is an optional field. If there is any metadata associated with the
audio being sent, use this field to provide it to the recognizer. The
server may record this metadata when processing the request. The server
does not use this field for any other purpose.
EvaluationError
Developer-facing error message about a non-fatal process issue.
custom_metadata (string )
Any custom metadata that the client wants to associate with the
recording. This could be a simple string (e.g. a tracing ID) or
structured data (e.g. JSON).
custom_id (string )
This is an optional field to specify custom ID to identify the
evaluation request. The custom ID must be a string of upto 64 bytes, and
only alphabets, digits, hyphens and underscores are allowed. This ID may
be recorded by the server in logs or other storage, and should therefore
not include any sensitive information.
EvaluationResult
EvaluationResult contains the result generated by a speech evaluation model.
Fields
is_partial (bool )
If this is set to true, it denotes that the result is an interim partial
result, and could change after more audio is processed. If unset, or set
to false, it denotes that this is a final result and will not change.
Servers are not required to implement support for returning partial
results, and clients should generally not depend on their availability.
score (float )
Overall evaluation score, between 0 and 1 inclusive, with 1 indicating a
perfect match between the reference text and what was said in the audio.
If alternative reference text(s) are specified, then the score will be
take those into account and be based on the reference text that aligns
the best with recognized tokens.
alignments (AlignedWord repeated)
Alignment between the primary expected reference text tokens and
recognized tokens.
alternative_alignments (AlternativeAlignment repeated)
Alignments against alternative reference text tokens (if specified) and
recognized tokens.
is_speech_endpoint (bool )
If true, indicates that a speech endpoint has been detected.
A speech endpoint signifies that the server believes the user has finished
their utterance, typically after detecting a specific duration of silence.
NOTE: Server support for endpoint detection is optional. Clients must be
robust to this field never being set.
ListModelsRequest
The top-level message sent by the client for the ListModels method.
ListModelsResponse
The message returned to the client by the ListModels method.
Fields
models (Model repeated)
List of models available for use that match the request.
Model
Description of a CAPT model.
Fields
id (string )
Unique identifier of the model. This identifier is used to choose the
model the model when configuring a StreamingEvaluate request.
name (string )
Model name. This is a concise name describing the model, and may be
presented to the end-user, for example, to help choose which model to use
for their task.
kind (ModelKind )
The specific kind of model. This determines what type of results
it sends back, and additional config requirements if any.
sample_rate (uint32 )
Audio sample rate (native) supported by the model.
metadata (ModelMetadata )
Metadata associated with the model.
ModelMetadata
Metadata associated with a Capt model.
Fields
version (string )
Model version in semver format (e.g. 1.0.0). If the version is not known,
it will default to “unknown”.
build_date (string )
Date the model was built in YYYY-MM-DD format. If the version is not
known, it will default to “unknown”.
StreamingEvaluateRequest
The top level messages sent by the client for the StreamingEvaluate method.
In this streaming call, multiple StreamingEvaluateRequest messages should
be sent. The first message must contain a EvaluationConfig message, and all
subsequent messages must contain Audio only. All Audio messages must
contain non-empty audio. If audio content is empty, the server may choose to
interpret it as end of stream and stop accepting any further messages.
The message returned by the server for the StreamingEvaluate method.
Fields
evaluation_result (EvaluationResult )
Response from the server. The kind of result depends on the kind of the
model that was chosen in the EvaluationConfig.
error (EvaluationError )
A non-fatal error message. If a server encountered a non-fatal error when
processing the request, it will be returned in this message.
The server will continue to process audio and produce further results.
Clients can continue streaming audio even after receiving these messages.
This error message is meant to be informational.
An example of when these errors maybe produced: audio is sampled at a
lower rate than expected by model, producing possibly less accurate
results.
This field will be unset if there is no error to report.
VersionRequest
The top-level message sent by the client for the Version method.
VersionResponse
The message sent by the server for the Version method.
Fields
version (string )
Version of the server handling these requests.
Enums
AlignmentKind
AlignmentKind represents one of four alignment outcomes possible: a Match,
Substitution, Deletion or Insertion.
Name
Number
Description
ALIGNMENT_KIND_UNSPECIFIED
0
Default value of this type.
ALIGNMENT_KIND_MATCH
1
A match implies that the reference and recognized token from audio are a match with a reasonable amount of confidence.
ALIGNMENT_KIND_SUBSTITUTION
2
A substitution implies that the reference token has been replaced by a different token in the recognized tokens from the audio.
ALIGNMENT_KIND_DELETION
3
A deletion implies that the reference token was not found in the recognized tokens from audio.
ALIGNMENT_KIND_INSERTION
4
A insertion implies that a extraneous token has been recognized in the audio, that cannot be matched to any token in the reference.
AudioEncoding
The encoding of the audio data to be sent for recognition.
Name
Number
Description
AUDIO_ENCODING_UNSPECIFIED
0
AUDIO_ENCODING_UNSPECIFIED is the default value of this type and will result in an error.
AUDIO_ENCODING_SIGNED
1
PCM signed-integer
AUDIO_ENCODING_UNSIGNED
2
PCM unsigned-integer
AUDIO_ENCODING_IEEE_FLOAT
3
PCM IEEE-Float
AUDIO_ENCODING_ULAW
4
G.711 mu-law
AUDIO_ENCODING_ALAW
5
G.711 a-law
AudioFormatHeadered
Name
Number
Description
AUDIO_FORMAT_HEADERED_UNSPECIFIED
0
AUDIO_FORMAT_HEADERED_UNSPECIFIED is the default value of this type.
AUDIO_FORMAT_HEADERED_WAV
1
WAV with RIFF headers
AUDIO_FORMAT_HEADERED_MP3
2
MP3 format with a valid frame header at the beginning of data
AUDIO_FORMAT_HEADERED_FLAC
3
FLAC format
AUDIO_FORMAT_HEADERED_OGG_OPUS
4
Opus format with OGG header
ByteOrder
Byte order of multi-byte data
Name
Number
Description
BYTE_ORDER_UNSPECIFIED
0
BYTE_ORDER_UNSPECIFIED is the default value of this type.
BYTE_ORDER_LITTLE_ENDIAN
1
Little Endian byte order
BYTE_ORDER_BIG_ENDIAN
2
Big Endian byte order
ModelKind
Name
Number
Description
MODEL_KIND_UNSPECIFIED
0
Default value of this type.
MODEL_KIND_SPEECH_EVALUATION
1
Model for evaluating the accuracy of spoken text from audio. This is done by first recognizing what’s said in the audio, and then aligning expected and recognized tokens (phonemes, syllables, etc.). This model returns results in the form of EvaluationResult messages.
MODEL_KIND_PHONEME_EVALUATION
2
Model for evaluating the accuracy of phonemes pronounced in isolation from audio. This is done in a way similar to speech evaluation models, but is more accurate for single phonemes in isolation, which a regular speech model may not recognize correctly. This model returns results in the form of EvaluationResult messages.
How is this different from using a general-purpose speech recognizer?
A recognizer is built to recover the intended message. That makes it actively
unsuitable for assessment, because it will repair what it hears into what it
assumes you meant: a mispronounced word is transcribed as the word, and the
error you were trying to measure disappears. The better the language model, the
more thoroughly it hides exactly the signal you need.
CAPT is anchored to a reference you supply and never substitutes expectation
for observation. It also gives you things a transcript cannot:
Per-phoneme verdicts and scores, not just a word-level transcript.
Confidence taken from a full distribution over phonemes, not a single guess.
Timestamps for each recognized phoneme.
Deterministic, inspectable scoring you can tune, with no risk of a
generative model inventing plausible text.
Operation entirely on your own hardware, on CPU.
What accuracy should I expect?
It depends on the task, the audio and the speakers, so we do not publish a
single figure that would mislead you. What we would recommend instead is
measuring it on your own data, against the decisions you actually care about.
If you are automating a judgement a person currently makes, the number that
matters is agreement with that person on your items, not word error rate, and
not any generic benchmark. See calibrating against human
judgement.
Does audio leave my infrastructure?
No. CAPT runs on your own hardware: on-premise, in your private cloud, or
embedded on a device. There is no call home, and no dependency on an external
service at request time.
Audio and results are not stored unless you explicitly enable it. Storage
is off by default; to turn it on, set a storage backend and path in
capt-server.cfg.toml:
[storage]Type="localfs"BasePath="/audio"
With this enabled, each session is written to local disk as two files, the
audio and the evaluation result, organized by UTC date. Everything stays on
the machine you run.
Note that metadata.custom_id and metadata.custom_metadata may be written to
logs and to stored results, so they should carry opaque identifiers rather than
personal information.
What hardware does it need?
CAPT is CPU-only; no GPU is required. It runs on x86_64 and Arm64 /
aarch64, and a statically linked build is available for minimal or embedded
images.
Sizing depends on the model and on how many concurrent evaluations you need.
Contact us for guidance against your
target hardware.
What languages are supported?
Each model covers one language, and the model’s language is fixed at build
time. US English models are available today, and models for other languages can
be built. Contact us to discuss a
specific language.
What audio should I send?
Use uncompressed or losslessly compressed audio (WAV or FLAC) recorded at the
model’s native sample rate, which ListModels reports as
attributes.sample_rate.
Audio sampled below what the model expects yields a non-fatal warning and
reduced accuracy. Upsampling before sending does not help; the detail is
already gone. Lossy formats such as MP3 are accepted but will cost you some
accuracy.
Can more than one person be speaking in the recording?
CAPT scores the audio against the reference text; it does not separate
speakers. If someone other than the intended speaker is audible, an assistant
reading a prompt for instance, that speech is part of the audio being
evaluated and can influence the result, usually appearing as insertions.
Where recordings may contain more than one voice, control it at capture time:
record only while the intended speaker is expected to be talking, rather than
across the whole interaction. If that is not possible for your setup, talk to
us about the options.
Can I add my own words and pronunciations?
Yes, and this is a normal thing to do. Add them to lexicon_addenda.tsv in the
model directory. You can add words the lexicon does not have, and override the
pronunciations of words it does. Words in neither the lexicon nor the addenda
are handled automatically by a grapheme-to-phoneme model, so invented words and
non-words work without any setup.
Yes, globally or per phoneme, by editing configuration rather than retraining.
DefaultCutoffScore moves the bar for everything; cutoffs.yaml sets it sound
by sound, and can additionally name specific confusions that should be
tolerated. See Cutoffs and acceptable
confusions.
Do I have to use streaming?
No. Evaluate takes the whole audio in one request and returns one result,
which is the simplest option for scoring a file. Use StreamingEvaluate when
you want results while the speaker is still talking, or for audio longer than
about a minute.
Are partial results guaranteed?
No. Servers are not required to produce them, and clients should not depend on
them. Drive your logic from the final result, the one with is_partial unset
or false, and treat partials as an enhancement for live feedback.
What does a score of 0.5 mean?
It means “exactly borderline” for that phoneme. Scores are not raw
probabilities: each phoneme has a cutoff, and the confidence-to-score mapping
is built so that the cutoff lands on 0.5. That is what makes 0.5 a meaningful
threshold and lets different phonemes be held to different standards.
Overall utterance scores are a summary of many such phoneme scores, so an
overall 0.5 does not mean “half correct”. Look at the per-token verdicts before
thresholding. See Interpreting Results.