This is the multi-page printable view of this section. Click here to print.
Speech Assessment
- 1: CAPT
- 1.1: Getting Started
- 1.2: Generating SDKs
- 1.3: Connecting to the Server
- 1.4: Evaluation Configurations
- 1.5: Streaming Evaluation
- 1.6: Interpreting Results
- 1.7: Phoneme Models
- 1.8: Tuning and Customization
- 1.9: Browser Integration
- 1.10: API Reference
- 1.11: FAQ
1 - CAPT
CAPT (Computer-Aided Pronunciation Training) answers a different question from a speech recognizer. Transcribe is given audio and asked what was said. CAPT is given audio and the text the speaker was supposed to say, and asked how closely the audio matches it, phoneme by phoneme.
That difference matters whenever you already know the target. Reading assessment, pronunciation practice, language learning drills and scripted-prompt verification all share the same shape: a known reference, a spoken attempt, and a decision about whether the attempt was good enough.
Because the reference is known, CAPT reports where the attempt diverged, not just that it did:
- Every phoneme of the reference is aligned to the audio and labeled
MATCH,SUBSTITUTION,DELETIONorINSERTION. - Every aligned phoneme carries a score in
[0, 1], derived from the acoustic model’s confidence and remapped through per-phoneme thresholds you control. - Every recognized phoneme carries a start time and duration, so results can be laid over a waveform.
- Words are scored as groups of phonemes, and the utterance gets a single overall score.
Crucially, CAPT is error-preserving. A general-purpose recognizer (and a large end-to-end model especially) is built to recover the intended message, so it will quietly repair a mispronunciation into the word it assumes you meant. That behavior is exactly wrong for assessment. CAPT never substitutes what it expects for what it heard: if the speaker said something other than the reference, the divergence is reported.
The engine runs on your own hardware: on-premise, in your private cloud, or embedded on a device. Audio never leaves your infrastructure.
How it works
CAPT is built on top of Cobalt Transcribe, running a phoneme-level acoustic model:
- The reference text is phonemized. Each word is looked up in a pronunciation lexicon; words that are not in it are passed to a grapheme-to-phoneme (G2P) model. A word may legitimately have several valid pronunciations, and all of them are considered.
- The audio is recognized into a confusion network: for each slice of time, a probability distribution over the phonemes that might have been spoken, rather than a single guess.
- Reference and audio are aligned with a weighted edit distance, producing
the per-phoneme
MATCH/SUBSTITUTION/DELETION/INSERTIONlabels. - Each phoneme is scored from its confidence in the confusion network, remapped through the cutoffs you configure.
Keeping the full distribution rather than a single best guess is what makes step 4 meaningful: the score for a phoneme reflects how much probability mass the acoustic model actually put on it.
Where to start
- Getting Started covers running a server.
- Interpreting Results is the part worth reading
carefully; it explains how to turn an
EvaluationResultinto a decision. - Tuning and Customization covers pronunciations, per-phoneme thresholds and scoring behavior.
1.1 - Getting Started
Using Cobalt CAPT
A typical CAPT release, provided as a compressed archive, contains a linux binary (
capt-server) for the required native CPU architecture, an appropriate Dockerfile, and models.Cobalt CAPT runs either locally on linux or using Docker.
Cobalt CAPT serves the CAPT gRPC API on port 2727. The configuration shipped with a release also enables an HTTP+JSON gateway on port 8080 and operational endpoints on port 8081; both are opt-in settings rather than server defaults, so a config written from scratch has gRPC only.
To quickly try out CAPT, first start the server as shown below and use the SDK in your preferred language to call it from your application.
The cobalt.license.key file is provided separately and must be copied into
the directory resulting from decompressing the archive. Please do this before
running the steps below.
Running CAPT Server Locally on Linux
./capt-server --config capt-server.cfg.toml
By default the binary assumes a configuration file named
capt-server.cfg.toml in the same directory. A different config file may be
specified using the --config argument.
On a successful start the server logs the addresses it is serving on:
2026/09/10 03:57:50 info {"msg":"server initializing"}
2026/09/10 03:57:50 info {"msg":"license verified"}
2026/09/10 03:57:51 info {"source":"transcribe","msg":"formats supported","formats":"[RAW WAV FLAC MP3 Opus]"}
2026/09/10 03:57:51 info {"msg":"runtime initialized","model_count":"1","init_time_taken":"1.485372717s"}
2026/09/10 03:57:51 info {"msg":"server started","grpcAddr":"[::]:2727","httpApiAddr":"[::]:8080","httpOpsAddr":"[::]:8081"}
Running CAPT Server as a Docker Container
To build and run the Docker image for CAPT, run:
docker build -t cobalt-capt .
docker run -p 2727:2727 -p 8080:8080 cobalt-capt
Checking the server is up
The HTTP API is the quickest way to confirm a working server, provided
api.Address is set in your config as the shipped one does. It exposes the
unary calls: Version, ListModels and Evaluate. The bidirectional
StreamingEvaluate is gRPC-only, though browsers can reach it over a
websocket:
curl -s http://localhost:8080/api/capt/v1/version
{"version":"1.4.0"}
The reported version is that of the capt-server release you were given.
curl -s http://localhost:8080/api/capt/v1/list-models
{
"models": [
{
"id": "en_US-16khz",
"name": "en_US CAPT (16khz audio)",
"kind": "MODEL_KIND_SPEECH_EVALUATION",
"attributes": {
"sample_rate": 16000,
"metadata": { "version": "0.0.0", "build_date": "unknown" }
}
}
]
}
The id of a model in this list is what you pass as model_id when
configuring an evaluation, and attributes.sample_rate is the audio sample
rate the model expects. See Evaluation
Configurations for what to do with them.
How to Get a Copy of the CAPT Server and Models
Contact us for a release best suited to your requirements.
The release you receive is a compressed archive (tar.bz2), generally structured as follows:
release.tar.bz2
├── COPYING
├── README.md
├── capt-server
├── capt-server.cfg.toml
├── Dockerfile
├── api
│ ├── proto [ cobaltspeech/capt/v1/capt.proto ]
│ └── gen [ pre-generated Go and Python bindings ]
├── models
│ └── en_US-16khz
│ ├── capt [ lexicon, G2P model, cutoffs, model config ]
│ └── transcribe [ acoustic model and decoding graph ]
│
└── cobalt.license.key [ provided separately, needs to be copied over ]
The
README.mdfile contains information about the release and instructions for starting the server on your system.The
capt-serveris the server program, configured using thecapt-server.cfg.tomlfile.The
Dockerfilecan be used to create a container that will let you run CAPT server on non-linux systems such as macOS and Windows.The
apidirectory holds the protobuf definition of the API and pre-generated client bindings. CAPT’s proto is distributed with the release rather than from a public repository, so this is where it comes from. See Generating SDKs.The
modelsdirectory contains the evaluation models. Each model directory holds two halves: thetranscribeacoustic model that recognizes phonemes, and thecaptresources (pronunciation lexicon, G2P model and scoring cutoffs) that the evaluation is built from. Both are described in Tuning and Customization.
System Requirements
Cobalt CAPT runs on Linux, directly as a native application. You can evaluate the product on Windows or macOS using Docker Desktop, but we would not recommend that setup for production.
CAPT runs on x86_64 and Arm64 / aarch64 CPUs, and a statically linked
build is available for deployment onto minimal or embedded images. Because
the engine is CPU-only and the models are small, the same release can be
deployed on a server, in a private cloud, or embedded on a device. Tell us the
target and we will provide a build for it.
Sizing depends on the model and on how many evaluations you need to run concurrently. Please contact us for sizing guidance against your target hardware and expected load.
To integrate Cobalt CAPT into your application, follow the next steps to install or generate the SDK in a language of your choice.
1.2 - Generating SDKs
The CAPT API is defined as a protocol buffer specification, a
protofile. This file allows a developer to auto-generate client SDKs for a number of different programming languages.The CAPT proto definition,
cobaltspeech/capt/v1/capt.proto, is supplied to you as part of your release, together with pre-generated Go and Python bindings. If you only need those two languages you can skip straight to Using the pre-generated SDKs; to generate bindings for another language, see Generating SDKs.
Unlike Cobalt’s other engines, the CAPT proto is not currently published in the
public cobaltspeech/proto repository.
Use the copy included with your release, and contact
us when you upgrade so that your
bindings stay in step with your server.
The relevant part of your release looks like this:
release.tar.bz2
└── api
├── proto
│ └── cobaltspeech
│ └── capt
│ └── v1
│ └── capt.proto
└── gen
├── go
│ └── cobaltspeech/capt/v1/{capt.pb.go, capt_grpc.pb.go}
└── py
└── cobaltspeech/capt/v1/{capt_pb2.py, capt_pb2_grpc.py, capt_pb2.pyi}
Using the pre-generated SDKs
Golang
Copy the api/gen/go tree into your module, or add it as a dependency, and
import the package:
import captpb "github.com/your-org/your-module/gen/go/cobaltspeech/capt/v1"
You will also need the gRPC and protobuf runtime libraries:
go get google.golang.org/protobuf
go get google.golang.org/grpc
Python
The Python bindings depend on Python >= 3.8. Place api/gen/py on your
PYTHONPATH, or copy the cobaltspeech package into your project, and install
the runtime dependencies:
pip install --upgrade pip
pip install --upgrade protobuf grpcio
Then import the modules:
import cobaltspeech.capt.v1.capt_pb2 as capt
import cobaltspeech.capt.v1.capt_pb2_grpc as capt_grpc
Generating SDKs from the proto
To generate bindings for a language we do not ship, generate them yourself from
capt.proto. We recommend using buf, a
command line tool that can generate documentation, schemas and SDK code for
many languages.
Step 1. Installing buf
COBALT="${HOME}/cobalt"
mkdir -p "${COBALT}/bin"
VERSION="1.72.0"
URL="https://github.com/bufbuild/buf/releases/download/v${VERSION}/buf-$(uname -s)-$(uname -m)"
curl -L ${URL} -o "${COBALT}/bin/buf"
# Give executable permissions and add to $PATH.
chmod +x "${COBALT}/bin/buf"
export PATH="${PATH}:${COBALT}/bin"brew install bufbuild/buf/bufStep 2. Writing a buf.gen.yaml
Create a buf.gen.yaml
next to the api/proto directory from your release. The example below
generates Go and Python; other plugins can be
added for more languages.
version: v1
managed:
enabled: true
go_package_prefix:
default: github.com/your-org/your-module/gen
plugins:
# Golang
- plugin: buf.build/grpc/go
out: gen/go
opt: paths=source_relative
- plugin: buf.build/protocolbuffers/go
out: gen/go
opt: paths=source_relative
# Python
- plugin: buf.build/grpc/python
out: gen/py
- plugin: buf.build/protocolbuffers/python
out: gen/py
- plugin: buf.build/protocolbuffers/pyi
out: gen/py
Step 3. Generating code
# Removing any previously generated files.
rm -rf ./gen
# Generating code for the proto files inside the `proto` directory.
buf generate proto
You should now have a gen folder containing the generated code. The latest
version of the CAPT API is v1. Import or copy the generated files into your
project as per the conventions of your language.
gen
└── py
└── cobaltspeech
└── capt
└── v1
├── capt_pb2_grpc.py
├── capt_pb2.py
└── capt_pb2.pyigen
└── go
└── cobaltspeech
└── capt
└── v1
├── capt_grpc.pb.go
└── capt.pb.goOnce you have an SDK, continue to connecting to the server.
1.3 - Connecting to the Server
Once you have your CAPT server up and running, and have installed or generated the SDK for your project, you can connect to it by “dialing” a gRPC connection.
First, you need the address where the server is running: e.g.
host:grpc_port. By default this is localhost:2727, and it is logged to the
terminal when you start CAPT server as grpcAddr:
2026/09/10 03:57:51 info {"msg":"server started","grpcAddr":"[::]:2727","httpApiAddr":"[::]:8080","httpOpsAddr":"[::]:8081"}
If you are hosting your server with Transport Layer Security (TLS) enabled, follow the instructions under Connect with TLS. Otherwise, follow the Default Connection method.
Default Connection
The following snippet connects to the server and queries its version, using an “insecure” gRPC channel. This would be the case if you have just started a local instance of CAPT server without TLS enabled.
import grpc
import cobaltspeech.capt.v1.capt_pb2 as capt
import cobaltspeech.capt.v1.capt_pb2_grpc as capt_grpc
serverAddress = "localhost:2727"
# Using a channel without TLS enabled.
channel = grpc.insecure_channel(serverAddress)
client = capt_grpc.CAPTServiceStub(channel)
# Get server version.
versionResp = client.Version(capt.VersionRequest())
print(versionResp)
# Get the list of models available on the server.
modelResp = client.ListModels(capt.ListModelsRequest())
for model in modelResp.models:
print(model)package main
import (
"context"
"fmt"
"os"
"google.golang.org/grpc"
"google.golang.org/grpc/credentials/insecure"
captpb "github.com/your-org/your-module/gen/go/cobaltspeech/capt/v1"
)
func main() {
const serverAddress = "localhost:2727"
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
// Using a channel without TLS enabled.
conn, err := grpc.NewClient(serverAddress,
grpc.WithTransportCredentials(insecure.NewCredentials()))
if err != nil {
fmt.Printf("failed to dial gRPC connection: %v\n", err)
os.Exit(1)
}
defer conn.Close()
client := captpb.NewCAPTServiceClient(conn)
// Get server version.
versionResp, err := client.Version(ctx, &captpb.VersionRequest{})
if err != nil {
fmt.Printf("failed to get server version: %v\n", err)
os.Exit(1)
}
fmt.Printf("%v\n", versionResp)
// Get the list of models available on the server.
modelResp, err := client.ListModels(ctx, &captpb.ListModelsRequest{})
if err != nil {
fmt.Printf("failed to list models: %v\n", err)
os.Exit(1)
}
for _, m := range modelResp.GetModels() {
fmt.Printf("%v\n", m)
}
}Connect with TLS
In our recommended setup for deployment, TLS is enabled in the gRPC connection, and clients validate the server’s SSL certificate to make sure they are talking to the right party. This is similar to how “https” connections work in web browsers.
TLS is enabled server-side by providing a certificate and key in
capt-server.cfg.toml:
[server.grpc]
Address = ":2727"
CertFile = "capt-server.crt"
KeyFile = "capt-server.key"
The following snippets show how to connect to a CAPT server that has TLS enabled.
import grpc
import cobaltspeech.capt.v1.capt_pb2 as capt
import cobaltspeech.capt.v1.capt_pb2_grpc as capt_grpc
serverAddress = "capt.your-org.internal:2727"
# Setup a gRPC connection with TLS. You can optionally provide your own
# root certificates and private key to grpc.ssl_channel_credentials()
# for mutually authenticated TLS.
creds = grpc.ssl_channel_credentials()
channel = grpc.secure_channel(serverAddress, creds)
client = capt_grpc.CAPTServiceStub(channel)
# Get server version.
versionResp = client.Version(capt.VersionRequest())
print(versionResp)package main
import (
"context"
"crypto/tls"
"fmt"
"os"
"google.golang.org/grpc"
"google.golang.org/grpc/credentials"
captpb "github.com/your-org/your-module/gen/go/cobaltspeech/capt/v1"
)
func main() {
const serverAddress = "capt.your-org.internal:2727"
// Setup a gRPC connection with TLS. You can optionally provide your own
// root certificates and private key through tls.Config for mutually
// authenticated TLS.
tlsCfg := tls.Config{}
creds := credentials.NewTLS(&tlsCfg)
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
conn, err := grpc.NewClient(serverAddress, grpc.WithTransportCredentials(creds))
if err != nil {
fmt.Printf("failed to dial gRPC connection: %v\n", err)
os.Exit(1)
}
defer conn.Close()
client := captpb.NewCAPTServiceClient(conn)
versionResp, err := client.Version(ctx, &captpb.VersionRequest{})
if err != nil {
fmt.Printf("failed to get server version: %v\n", err)
os.Exit(1)
}
fmt.Printf("%v\n", versionResp)
}Client Authentication
In some setups it is desirable for the server to validate the clients connecting to it, and only respond to ones it can verify. If your CAPT server is configured to do client authentication, you will need to present the appropriate certificate and key when connecting to it.
Note that in client-authentication mode the client still also verifies the server’s certificate, so this setup uses mutually authenticated TLS.
creds = grpc.ssl_channel_credentials(
root_certificates=root_certificates, # PEM certificate as byte string
private_key=private_key, # PEM client key as byte string
certificate_chain=certificate_chain, # PEM client certificate as byte string
)// Root PEM certificate for validating a self-signed server certificate.
var rootCert []byte
// Client PEM certificate and private key.
var certPem, keyPem []byte
caCertPool := x509.NewCertPool()
if ok := caCertPool.AppendCertsFromPEM(rootCert); !ok {
fmt.Printf("unable to use given caCert\n")
os.Exit(1)
}
clientCert, err := tls.X509KeyPair(certPem, keyPem)
if err != nil {
fmt.Printf("unable to use given client certificate and key: %v\n", err)
os.Exit(1)
}
tlsCfg := tls.Config{
RootCAs: caCertPool,
Certificates: []tls.Certificate{clientCert},
}
creds := credentials.NewTLS(&tlsCfg)1.4 - Evaluation Configurations
Every evaluation begins with an
EvaluationConfig. It is the first
message of a StreamingEvaluate stream, and the config field of a unary
Evaluate request. This page describes what goes in it.
Choosing a model
model_id selects the model to evaluate against, and must be one of the id
values returned by ListModels. Each model
also reports the sample_rate it expects and its kind:
MODEL_KIND_SPEECH_EVALUATIONevaluates words and sentences. Thereference_textis ordinary text.MODEL_KIND_PHONEME_EVALUATIONevaluates phonemes produced in isolation, which a word-oriented model may not recognize correctly. Thereference_textis a sequence of phones. See Phoneme Models.
Both kinds return the same EvaluationResult
structure, so only the reference_text you supply differs.
Reference text
reference_text is the text the speaker is expected to have said: the thing
the audio is scored against. This is the field that makes CAPT what it is: the
evaluation is anchored to it.
cfg = capt.EvaluationConfig(
model_id="en_US-16khz",
reference_text="WHEN THE SUNLIGHT STRIKES",
)
A few properties worth knowing:
- Lexicon lookup is case-insensitive, so
when,WhenandWHENbehave identically. The uppercase reference text used throughout these examples is a convention, not a requirement. - Punctuation handling is a per-model setting,
RemovePunctuationin the model’s[Tokenization]config, and it is off in theen_US-16khzmodel shipped today. Punctuation therefore stays attached to the word: a reference ofWHEN, THE SUNLIGHT STRIKEScomes back withtextofWHEN,, and because the token no longer matches the lexicon it is phonemized by G2P, which changed that word’s score in our own testing. Sentence-final marks are usually harmless, but the safe habit is to send the words alone. See Tuning and Customization. - Words not in the lexicon are handled by a G2P model, so invented words, proper nouns and non-words can be used as reference text without any setup. This is what makes non-word decoding tasks work. See Tuning and Customization.
- Words with several valid pronunciations are all considered, and the
variant that best fits the audio is the one reported. A speaker is not
penalized for saying “THE” as
D irather thanD @.
Alternative reference text
alternative_reference_text accepts additional reference texts that would also
be acceptable. Each is aligned independently and returned in
alternative_alignments with its own
score, and the top-level score becomes the best score across the primary and
all alternatives.
cfg = capt.EvaluationConfig(
model_id="en_US-16khz",
reference_text="WHEN THE SUNLIGHT STRIKES",
alternative_reference_text=["WHEN THE SUN LIGHTS STRIKE"],
)
If you need to know which candidate the audio matched, for example when
choosing between a correct answer and a set of known incorrect answers, it is
clearer to issue a separate evaluation per candidate and compare the scores
yourself. alternative_reference_text optimizes for “any of these is
acceptable”, not for classification. See Interpreting
Results.
Audio format
audio_format tells the server how to interpret the bytes you send. Two shapes
are available.
For files that carry their own header, name the container and let the server read the rest:
cfg = capt.EvaluationConfig(
model_id="en_US-16khz",
reference_text="WHEN THE SUNLIGHT STRIKES",
audio_format=capt.AudioFormat(
audio_format_headered=capt.AUDIO_FORMAT_HEADERED_WAV,
),
)
Supported headered formats are WAV, MP3, FLAC and OGG Opus. Headers are
expected once, at the beginning of the stream, not in every Audio message.
audio_format may be omitted entirely for headered audio. The server then
detects the format from the header, which is why several examples in these
pages leave it out. Naming it explicitly is still recommended: it turns an
unreadable or unexpected header into a clear error instead of a detection
attempt.
Raw audio is the exception. It carries no header to detect, so
audio_format_raw must be supplied in full, and encoding must be set to
something other than AUDIO_ENCODING_UNSPECIFIED or the request is
rejected.
For raw samples with no header, describe them fully:
cfg = capt.EvaluationConfig(
model_id="en_US-16khz",
reference_text="WHEN THE SUNLIGHT STRIKES",
audio_format=capt.AudioFormat(
audio_format_raw=capt.AudioFormatRAW(
encoding=capt.AUDIO_ENCODING_SIGNED,
bit_depth=16,
byte_order=capt.BYTE_ORDER_LITTLE_ENDIAN,
sample_rate=16000,
channels=1,
),
),
)
Every server supports raw little-endian 16-bit signed samples at the model’s own sample rate; support for the other formats depends on how the server was built and configured.
For best accuracy use uncompressed or losslessly compressed audio (WAV or
FLAC) recorded at the model’s native sample rate. Sending audio sampled below
what the model expects produces a non-fatal
EvaluationError alongside your results,
warning that accuracy may be reduced. Upsampling low-rate audio does not
recover the missing detail.
Request metadata
metadata is optional and is never used to influence evaluation. The server
may record it in logs or stored results.
cfg = capt.EvaluationConfig(
model_id="en_US-16khz",
reference_text="WHEN THE SUNLIGHT STRIKES",
metadata=capt.EvaluationMetadata(
custom_id="session-42-item-07",
custom_metadata='{"item":"7","form":"A"}',
),
)
custom_idaccepts up to 64 bytes, restricted to letters, digits, hyphens and underscores. Useful as a tracing or correlation ID.custom_metadatais a free-form string; a plain tag or structured data such as JSON.
Because these fields may be written to logs and to stored results, they must not contain personal or otherwise sensitive information. Use an opaque identifier that only your own systems can resolve back to a person.
Once you have a config, continue to Streaming Evaluation.
1.5 - Streaming Evaluation
CAPT offers two ways to submit audio.
StreamingEvaluateis the primary method. It is bidirectional: you send audio as it is captured and receive results as the audio is processed. Use it for anything interactive, and for audio of any length.Evaluateis a unary convenience method that takes the whole audio in a single request and returns a single result. It is intended for short audio, less than a minute, and is the simplest way to score an audio file.
Both accept the same EvaluationConfig and
return the same EvaluationResult.
Streaming from an audio file
In a StreamingEvaluate stream, the first message must contain the config,
and every subsequent message must contain audio.
An empty Audio message ends the stream. CAPT server treats it as the end
of the audio: it finishes processing, sends the final result, and accepts
nothing further, so a later send fails. Over gRPC you normally do not need it,
because half-closing the stream (CloseSend in Go, exhausting the request
iterator in Python) says the same thing. It matters for
browser clients, where a websocket has no
equivalent half-close.
Do not send one to mean “no audio yet”. If it arrives before the audio is complete, what you get depends on the format, and a client’s state machine needs to handle both:
- You always receive a final result for the audio sent so far, with
is_partialfalse. It reflects the truncated input, not the utterance you intended, so it will usually be full of deletions. - With a headered format the stream then fails. The header promised more
data than arrived, so the server reports a container error such as
wav: unexpected EOFafter the final result. This happens whether or not you send anything afterwards. - With raw audio the stream ends cleanly, because there is no container to leave incomplete. The result is still only of the audio you sent.
In other words: a final result is not by itself evidence that the whole utterance was processed. Send the empty message only once the recording is complete, and treat an error after a final result as “the result is incomplete”, not “the result is invalid”.
The examples below stream a WAV file in 100 ms chunks and print each result as it arrives.
import grpc
import cobaltspeech.capt.v1.capt_pb2 as capt
import cobaltspeech.capt.v1.capt_pb2_grpc as capt_grpc
serverAddress = "localhost:2727"
channel = grpc.insecure_channel(serverAddress)
client = capt_grpc.CAPTServiceStub(channel)
# Get the list of models on the server and use the first one.
modelResp = client.ListModels(capt.ListModelsRequest())
modelID = modelResp.models[0].id
cfg = capt.EvaluationConfig(
model_id=modelID,
reference_text="WHEN THE SUNLIGHT STRIKES",
audio_format=capt.AudioFormat(
audio_format_headered=capt.AUDIO_FORMAT_HEADERED_WAV,
),
)
# The first request must contain only the configuration; subsequent
# requests carry audio bytes. A generator is a convenient way to do this.
def stream(cfg, audio, bufferSize=3200):
yield capt.StreamingEvaluateRequest(config=cfg)
data = audio.read(bufferSize)
while len(data) > 0:
yield capt.StreamingEvaluateRequest(audio=capt.Audio(data=data))
data = audio.read(bufferSize)
with open("test.wav", "rb") as audio:
for resp in client.StreamingEvaluate(stream(cfg, audio)):
if resp.HasField("error"):
print(f"warning: {resp.error.message}")
result = resp.evaluation_result
print(f"partial={result.is_partial} score={result.score:.4f}")
# Only act on final results; partials will still change.
if not result.is_partial:
for word in result.alignments:
print(f" {word.text}: " + " ".join(
f"{t.reference}={t.score:.2f}" for t in word.tokens
))package main
import (
"context"
"fmt"
"io"
"os"
"google.golang.org/grpc"
"google.golang.org/grpc/credentials/insecure"
captpb "github.com/your-org/your-module/gen/go/cobaltspeech/capt/v1"
)
func main() {
const serverAddress = "localhost:2727"
audio, err := os.ReadFile("test.wav")
if err != nil {
fmt.Printf("failed to read audio: %v\n", err)
os.Exit(1)
}
conn, err := grpc.NewClient(serverAddress,
grpc.WithTransportCredentials(insecure.NewCredentials()))
if err != nil {
fmt.Printf("failed to dial gRPC connection: %v\n", err)
os.Exit(1)
}
defer conn.Close()
client := captpb.NewCAPTServiceClient(conn)
stream, err := client.StreamingEvaluate(context.Background())
if err != nil {
fmt.Printf("failed to open stream: %v\n", err)
os.Exit(1)
}
// The first message must contain the config.
if err := stream.Send(&captpb.StreamingEvaluateRequest{
Request: &captpb.StreamingEvaluateRequest_Config{
Config: &captpb.EvaluationConfig{
ModelId: "en_US-16khz",
ReferenceText: "WHEN THE SUNLIGHT STRIKES",
AudioFormat: &captpb.AudioFormat{
AudioFormat: &captpb.AudioFormat_AudioFormatHeadered{
AudioFormatHeadered: captpb.AudioFormatHeadered_AUDIO_FORMAT_HEADERED_WAV,
},
},
},
},
}); err != nil {
fmt.Printf("failed to send config: %v\n", err)
os.Exit(1)
}
// Send audio in the background while results are read below.
go func() {
const chunk = 3200 // 100ms of 16kHz 16-bit mono audio.
// min() is a builtin from Go 1.21; on older toolchains compute the
// bound explicitly.
for i := 0; i < len(audio); i += chunk {
end := min(i+chunk, len(audio))
if err := stream.Send(&captpb.StreamingEvaluateRequest{
Request: &captpb.StreamingEvaluateRequest_Audio{
Audio: &captpb.Audio{Data: audio[i:end]},
},
}); err != nil {
return
}
}
_ = stream.CloseSend()
}()
for {
resp, err := stream.Recv()
if err == io.EOF {
break
}
if err != nil {
fmt.Printf("failed to receive result: %v\n", err)
os.Exit(1)
}
if e := resp.GetError(); e != nil {
fmt.Printf("warning: %v\n", e.GetMessage())
}
result := resp.GetEvaluationResult()
fmt.Printf("partial=%v score=%.4f\n", result.GetIsPartial(), result.GetScore())
// Only act on final results; partials will still change.
if !result.GetIsPartial() {
for _, word := range result.GetAlignments() {
fmt.Printf(" %s:", word.GetText())
for _, t := range word.GetTokens() {
fmt.Printf(" %s=%.2f", t.GetReference(), t.GetScore())
}
fmt.Println()
}
}
}
}Streaming a 1.5 second recording of “when the sunlight strikes” against that reference produces a short sequence of results:
partial=true score=0.5514
partial=true score=0.6128
partial=true score=0.6128
partial=false score=0.6128
Partial results
While audio is still arriving, the server emits partial results, marked
is_partial = true. A partial reflects the audio received so far, and any part
of it may change as more audio arrives. In the sequence above the score rises
as the last word is heard.
When the stream ends, the server emits a final result with is_partial = false. That result is stable.
- Use partials for live feedback: highlighting words as they are read, showing a provisional score, driving a progress indicator.
- Use the final result for any decision you record. Scoring an item on a partial risks scoring it on half a word.
Servers are not required to produce partial results, and clients should not depend on receiving them. Always drive your logic from the final result and treat partials as an enhancement.
Speech endpointing
A result may set is_speech_endpoint = true to indicate the server believes
the speaker has finished, typically after a period of silence. This is useful
for closing the microphone automatically once an answer has been given.
Endpoint detection is optional, and clients must be robust to the field never
being set. Do not treat it as the signal that results are final; that is what
is_partial = false is for.
Non-fatal errors
Both response types carry an optional
EvaluationError beside the result. It
reports conditions that degrade quality but do not stop processing, most
commonly audio sampled at a lower rate than the model expects. The server
continues, and you may keep streaming. Log these rather than aborting; they
usually point at a recording pipeline that needs attention.
Evaluating a whole file at once
For short audio, Evaluate avoids the streaming machinery entirely. Over the
HTTP+JSON gateway that is a single POST:
curl -s -X POST http://localhost:8080/api/capt/v1/evaluate \
-H 'Content-Type: application/json' \
-d '{
"config": {
"model_id": "en_US-16khz",
"reference_text": "WHEN THE SUNLIGHT STRIKES",
"audio_format": {"audio_format_headered": "AUDIO_FORMAT_HEADERED_WAV"}
},
"audio": {"data": "<base64-encoded WAV file>"}
}'
The data field is the audio file, base64 encoded. See Interpreting
Results for what comes back.
1.6 - Interpreting Results
This is the page worth reading carefully. Getting audio into CAPT is straightforward; deciding what its output means for your application is where the design work is.
The shape of a result
An EvaluationResult is a tree:
EvaluationResult
├── score overall score for the utterance, 0..1
├── is_partial true while more audio may still change this
├── is_speech_endpoint true if the speaker appears to have finished
├── alignments [ ]AlignedWord
│ └── AlignedWord
│ ├── text the reference word
│ ├── start_time_ms when the word starts in the audio
│ ├── duration_ms how long it lasts
│ └── tokens [ ]AlignedToken (the phonemes)
│ └── AlignedToken
│ ├── kind MATCH | SUBSTITUTION | DELETION | INSERTION
│ ├── reference the expected phoneme
│ ├── score 0..1 for this phoneme
│ └── hypotheses [ ]AlignmentHypothesis (what was heard)
│ └── token, confidence, start_time_ms, duration_ms
└── alternative_alignments the same, once per alternative reference text
The reference text drives the structure: your words appear as AlignedWord
entries in order, and within each one there is an AlignedToken per expected
phoneme.
alignments is not a one-to-one list of your reference words. Insertions
are reported as extra AlignedWord entries with an empty text, positioned
where they were heard, so the list can be longer than
reference_text.split() and the indices do not line up. Match on text, or
skip entries whose text is empty, rather than zipping the two lists together.
See Insertions.
The four alignment kinds
Every expected phoneme gets exactly one of four verdicts.
| Kind | Meaning | reference | hypotheses | score |
|---|---|---|---|---|
MATCH | The expected phoneme was heard with enough confidence | set | populated | 0.5 to 1.0 |
SUBSTITUTION | The expected phoneme was not heard with enough confidence: either something else was heard in its place, or the phoneme itself was heard but below its cutoff | set | populated | 0 to below 0.5 |
DELETION | The phoneme was not heard at all | set | empty | 0 |
INSERTION | Something extra was heard that the reference does not account for | empty | populated | 0 |
A SUBSTITUTION does not imply a score of 0. When the expected phoneme
was heard but fell short of its cutoff, the score is its remapped confidence,
which is non-zero but below 0.5. The k scoring 0.49 in the example below
is exactly this case. Only when the phoneme is absent from the hypotheses
entirely does the score reach 0. Decide on kind, or on a threshold, but do
not infer one from the other.
A real example makes this concrete. Below is the result of scoring a recording of “when the sunlight strikes” against a deliberately wrong reference, “WHEN THE MOONLIGHT STRIKES” (phonemes are X-SAMPA):
overall score: 0.5157
WHEN MATCH w (1.00) MATCH E (0.99) MATCH n (1.00)
THE MATCH D (0.87) MATCH @ (0.67)
MOONLIGHT SUB m (0.00) SUB u (0.00) MATCH n (0.65)
MATCH l (0.78) MATCH aI (0.82) MATCH t (0.73)
STRIKES MATCH s (0.79) DEL t (0.00) DEL r\ (0.00)
DEL aI (0.00) SUB k (0.49) DEL s (0.00)
Read that as a diagnosis, not just a number. The first two words were said as
expected. In MOONLIGHT the leading m u was not there (the speaker said
s V n), but the shared tail n l aI t matched, so those phonemes score well.
STRIKES shows what happens when the audio has already run out: the reference
still has phonemes to account for, and they come back as deletions.
The overall score of 0.52 is the average of that mixture. This is why an
overall score alone is rarely the right thing to threshold. 0.52 here does
not mean “roughly half right”, it means “two words right and two words wrong”.
Insertions
Insertions do not belong to any reference word, so they are reported in their
own AlignedWord entries with an empty text, positioned in the sequence
where they were heard. Scoring the same audio against the single word
“ELEPHANT” shows this:
overall score: 0.3966
(insertion) INS (heard: w)
ELEPHANT MATCH E (0.99) SUB l (0.00) MATCH @ (0.67)
SUB f (0.00) SUB n= (0.00) MATCH t (0.73)
(insertion) INS (heard: k) INS (heard: s)
The speaker said far more than the reference accounted for, and the leftover audio at each end is reported as insertions rather than being silently discarded.
Whether phonemes inserted within a word are attached to that word or reported
separately is controlled by the model’s IncludeIntraWordInsertion setting.
See Tuning and Customization.
Where the numbers come from
Phoneme scores
The acoustic model does not commit to a single phoneme per time slice. It produces a distribution, a confusion network, and the score for an expected phoneme is derived from how much probability mass landed on it.
You can see this in the hypotheses list. Here is a single MATCH from a
real result:
{
"kind": "ALIGNMENT_KIND_MATCH",
"reference": "n",
"score": 0.9798913,
"hypotheses": [
{ "token": "n", "confidence": 0.963, "start_time_ms": "570", "duration_ms": "119" },
{ "token": "m", "confidence": 0.029, "start_time_ms": "570", "duration_ms": "119" },
{ "token": "n=", "confidence": 0.006, "start_time_ms": "570", "duration_ms": "119" },
{ "token": "N", "confidence": 0.002, "start_time_ms": "570", "duration_ms": "119" }
]
}
The model heard n with 0.963 confidence, but also considered m, n= and
N. The reported score of 0.98 is that confidence remapped through the
model’s cutoff for n. See Tuning and
Customization
for how that remapping works and how to change it.
The practical consequence: a score is a calibrated quantity, not a raw
probability. A cutoff is chosen so that a score of exactly 0.5 sits at the
boundary between acceptable and unacceptable for that phoneme. That is what
makes 0.5 a meaningful place to threshold, and it is why different phonemes
can be held to different standards.
Word and utterance scores
An AlignedWord does not carry its own score field; a word’s quality is the
scores of its phonemes. Compute whatever summary suits your task: the mean,
the minimum, or the fraction of phonemes that came back MATCH.
The top-level score is the overall score for the utterance: the mean of the
scores of the reference tokens. If alternative reference texts were
supplied, it is the best score across the primary and all alternatives.
Insertions are not counted in the overall score. The score answers “how
well was the reference produced”, not “was anything else said”. A recording
that contains the target words plus a great deal of unrelated speech can still
score 1.0.
If extraneous speech should count against the speaker (an assistant reading a
prompt, or a child answering twice), inspect the insertion entries yourself and
apply your own rule. The INSERTION tokens carry timestamps, so you can also
measure how much of the audio they account for.
Timestamps
Timestamps live on the hypotheses, not on the token. An AlignedToken has
no time fields of its own; start_time_ms and duration_ms are on each
AlignmentHypothesis inside it, and
on the enclosing AlignedWord.
This has one consequence that catches people out:
A DELETION has an empty hypotheses list and therefore no timestamp at
all. There is no audio to point at; that is what a deletion means. Any
timeline, waveform or spectrogram view must handle tokens that have no time
span, rather than assuming every phoneme can be drawn.
Likewise, a word made up entirely of deletions has nothing to anchor to, and
its start_time_ms and duration_ms are reported as 0.
Turning results into a decision
A pass/fail mark for one item
If your application needs a yes/no verdict (did the speaker say the target correctly?), do not reach for the overall score first. Consider what the task actually requires:
Strict, verbatim tasks. Any divergence is a failure. Require every token to be
MATCH:correct = all( t.kind == capt.ALIGNMENT_KIND_MATCH for w in result.alignments for t in w.tokens )This is exactly right when the rule is “any change at all, however minor, is wrong”, because substitutions, deletions and insertions all break it.
Tolerant tasks. Some divergence is acceptable. A slightly indistinct consonant should not fail an otherwise good attempt. Threshold on the proportion of matched phonemes, or on the mean phoneme score, rather than demanding a clean sweep. Note that a tolerant rule built on the overall
score, or on reference tokens alone, will not notice extra speech: decide separately whether insertions should fail the item.Targeted tasks. Only some phonemes matter: the contrast the item is testing. Score only those tokens and ignore the rest.
Choosing between several candidate answers
When you have a known correct answer and a set of known incorrect answers, and you need to know which one was said, evaluate the audio once per candidate and compare the resulting scores:
candidates = ["BED", "LOUNGE", "SOFA"]
scores = {}
for candidate in candidates:
cfg = capt.EvaluationConfig(model_id=modelID, reference_text=candidate)
with open(path, "rb") as audio:
for resp in client.StreamingEvaluate(stream(cfg, audio)):
if not resp.evaluation_result.is_partial:
scores[candidate] = resp.evaluation_result.score
best = max(scores, key=scores.get)
This gives you a score per candidate, so you can see not just the winner but
the margin, so you can reject the result as unclear when two candidates score
alike. Packing the candidates into alternative_reference_text instead would
collapse them into a single best score and lose exactly that information.
Calibrating against human judgement
If you are automating a decision a person currently makes, the metric that matters is agreement with that person, not any internal notion of accuracy.
The recommended approach is to collect audio alongside the human verdict, run CAPT over it, then sweep your decision rule across the collected set and count true positives, true negatives, false positives and false negatives at each setting. That tells you where to set the threshold, and, just as importantly, what the residual disagreement rate is and which way it leans. A rule that is wrong in the safe direction for your use case is often better than one that is wrong less often overall.
Per-phoneme cutoffs can then correct systematic biases that a single global threshold cannot; see Tuning and Customization.
Alternative alignments
If you supplied alternative_reference_text, each alternative comes back in
alternative_alignments as an
AlternativeAlignment carrying its
own reference_text, score, and full alignments tree with the same
structure described above.
The top-level score is the best across the primary and the alternatives, so
if you only care whether any acceptable rendering was produced, the top-level
score is enough. If you care which one, see above.
1.7 - Phoneme Models
CAPT models come in two kinds, reported as kind by
ListModels.
- Word models (
MODEL_KIND_SPEECH_EVALUATION) evaluate phonemes in word context. Thereference_textis ordinary text, and this is what you want for reading words, sentences and passages. - Phoneme models (
MODEL_KIND_PHONEME_EVALUATION) evaluate phonemes produced in isolation. Thereference_textis a sequence of phones.
The distinction matters because a sound produced on its own is acoustically quite different from the same sound inside a word. A model trained on connected speech often mis-recognizes an isolated phoneme, because nothing in its training looked like that. If your task asks a speaker to produce a single sound (“say the /sh/ sound”), a phoneme model is substantially more reliable.
Both kinds are used through exactly the same API calls and return the same
EvaluationResult. Only the reference_text
differs.
Reference text for phoneme models
For a phoneme model, reference_text is one or more phones separated by
whitespace:
# A single phone.
cfg = capt.EvaluationConfig(model_id=phonemeModelID, reference_text="AA")
# A sequence of phones.
cfg = capt.EvaluationConfig(model_id=phonemeModelID, reference_text="Z AE P")
Phone groups
Some phones are only meaningfully produced as a pair, and the model treats such a pair as a single unit called a phone group. Phone groups are written with the members joined by a period:
# A phone group.
cfg = capt.EvaluationConfig(model_id=phonemeModelID, reference_text="AO.NG")
# Mixed with ordinary phones.
cfg = capt.EvaluationConfig(model_id=phonemeModelID, reference_text="F AE.NG")
For all practical purposes a phone group behaves as its own distinct phone: it is matched, substituted or deleted as a single token, and receives a single score.
The phone set
The phones a model accepts are specific to that model, both the inventory and which groups exist. Supplying a phone the model does not know is an error, so check against the set for the model you were given. A typical US English phoneme model accepts:
AA AA.R AE AE.NG AH AH.NG AO AO.NG
AO.R AW AY B CH D DH EH
EH.R ER EY F G HH IH IH.NG
IH.R IY JH K L M N NG
OW OY P R S SH T TH
UH UW V W Y Y.UW Z ZH
Note that word models and phoneme models may use different phone alphabets. A
word model may report X-SAMPA (w E n, r\, aI) while a phoneme model uses
the ARPABET-style symbols above. Do not assume a phone string from one is valid
in the other.
Reading the results
Results have the same structure as any other evaluation. For each phone in the
reference_text there is an aligned token:
ALIGNMENT_KIND_MATCHmeans the phone was found in the audio with a high degree of confidence.ALIGNMENT_KIND_SUBSTITUTIONmeans a different phone was recognized in its place, or the phone was recognized only with low confidence.ALIGNMENT_KIND_DELETIONmeans the phone was not recognized at all.
Anything else recognized in the audio that does not align to the reference is reported as insertions, grouped separately.
Isolated-phoneme audio very often contains more than the phoneme itself: a
speaker clearing their throat, a lead-in, or the assessor’s prompt. Those show
up as insertions, which is why the example below scores 1.0 despite three
extra tokens: the reference phone ZH was produced correctly, and the
surrounding material is reported rather than being allowed to affect the score
of the phone you asked about.
{
"score": 1.0,
"alignments": [
{
"start_time_ms": "680",
"duration_ms": "360",
"tokens": [
{
"kind": "ALIGNMENT_KIND_INSERTION",
"hypotheses": [
{ "token": "T", "confidence": 1.0, "start_time_ms": "680", "duration_ms": "120" }
]
},
{
"kind": "ALIGNMENT_KIND_INSERTION",
"hypotheses": [
{ "token": "R", "confidence": 1.0, "start_time_ms": "840", "duration_ms": "120" }
]
},
{
"kind": "ALIGNMENT_KIND_INSERTION",
"hypotheses": [
{ "token": "EH", "confidence": 0.994, "start_time_ms": "960", "duration_ms": "80" }
]
}
]
},
{
"text": "ZH",
"start_time_ms": "1000",
"duration_ms": "240",
"tokens": [
{
"kind": "ALIGNMENT_KIND_MATCH",
"reference": "ZH",
"score": 1.0,
"hypotheses": [
{ "token": "ZH", "confidence": 1.0, "start_time_ms": "1000", "duration_ms": "240" }
]
}
]
}
]
}
If extraneous audio should count against the speaker in your task, inspect the insertion entries and apply your own rule; the information is there either way.
1.8 - Tuning and Customization
A general-purpose speech model is a fixed object: you send it audio and accept what it returns. CAPT is deliberately not that. Because the scoring stage is separate from the acoustic model, a great deal of behavior can be adjusted for your task without retraining anything, by changing which pronunciations count as correct, and how strictly each sound is judged.
Everything on this page lives in the model directory shipped with your release, and takes effect when the server restarts.
models/en_US-16khz/capt/
├── model.config.toml scoring behavior and paths to the files below
├── lexicon.json the primary pronunciation dictionary
├── lexicon_addenda.tsv your additions and overrides
├── cutoffs.yaml per-phoneme thresholds
└── g2p.ort grapheme-to-phoneme model for unknown words
The model config points at each of these by path, so the names above are the
convention rather than a requirement. Older model directories may carry the
cutoffs as a plain-text cutoffs.txt instead; the meaning is identical and
CutoffsPath says which file is in use.
Pronunciations
The lexicon and its addenda
The lexicon maps words to their phoneme sequences. A word may have several valid pronunciations, and CAPT considers all of them, reporting whichever best fits the audio, so a speaker is not penalized for a legitimate variant.
lexicon_addenda.tsv is loaded after the primary lexicon and is the file
you should edit. Entries in it add new words and override existing ones,
leaving the shipped lexicon intact so that upgrades stay clean.
This is the mechanism to reach for when your task defines what counts as an acceptable pronunciation. If an item specifies that a target may be produced two ways, add both and each will be accepted as a match:
MOSP m A s p
MOSP m oU s p
Those two entries accept mosp pronounced to rhyme with wasp, or with soap.
Every phone symbol on this page belongs to the model’s own alphabet, and
the files in a model directory all use the same one. The en_US-16khz word
model used in these examples is X-SAMPA, which is why its addenda read
m A s p and its cutoffs are keyed "A" rather than AA. A phoneme model
with an ARPABET inventory would use AA in both files instead. Check the phone
set for the model you were given before editing either file: a symbol the model
does not know matches nothing, and fails silently rather than erroring.
Out-of-vocabulary words and G2P
Words that are in neither the lexicon nor the addenda are passed to a grapheme-to-phoneme model, which predicts a pronunciation from spelling. This happens automatically and requires no configuration.
It means invented words, proper nouns and non-words can be used as
reference_text with no setup at all, which is useful for decoding tasks built on
pseudo-words, where the whole point is that the string is not a real word.
G2P is a prediction, not a definition. Where you know the pronunciations that should be accepted, because your task specifies them, put them in the addenda rather than relying on what G2P infers from the spelling. Reserve G2P for the long tail you have not enumerated.
Tokenization
[Tokenization] in model.config.toml controls how reference text is
pre-processed:
| Setting | What it does | Default | In en_US-16khz |
|---|---|---|---|
RemovePunctuation | Strips punctuation before lexicon lookup. With it off, punctuation stays attached to the word and is phonemized as part of it. | off | off |
UseSentenceTokenizer | Splits a long reference into sentences before tokenizing. | off | on |
MaxInputLength | Safeguard on the length of a reference text. Counted in characters, not words; a longer reference is rejected with an error. Applies to each alternative reference text too. | 2048 | 2048 |
“Default” is what applies when the key is absent from model.config.toml.
Neither boolean has a server-side default beyond false, so what matters in
practice is what your model ships with.
Cutoffs and acceptable confusions
This is the most useful tuning mechanism CAPT offers, and the least obvious.
Raw acoustic confidences are biased in ways that are specific to a model and a
population of speakers. A model may routinely blur m and n; young children
may produce th in a way that reads as s. Left alone, those biases show up
as wrong verdicts. Cutoffs correct them by remapping confidence to score.
Each symbol has a cutoff, the confidence at which it is exactly borderline.
The remapping is piecewise linear and pins that point to a score of 0.5:
- raw confidence in
[0, cutoff]maps to a score in[0, 0.5] - raw confidence in
[cutoff, 1]maps to a score in[0.5, 1]
So the cutoff is the dial for how strict a given phoneme is. Lowering it is
more lenient; raising it demands higher confidence before the phoneme counts as
correct. Any symbol not listed uses DefaultCutoffScore from the [Scoring]
section.
A modifier extends this with acceptable confusions. It names
alternatives (other phonemes that may also count toward the match) and a
multiplier that scales the cutoff when one of them is what was actually
heard:
"A":
cutoff: 0.50
modifier:
multiplier: 1.50
alternatives: ["O"]
"T":
cutoff: 0.50
modifier:
multiplier: 1.90
alternatives: ["s"]
Read the first entry as: "A (the vowel of lot) needs 50% confidence. O
(the vowel of thought) is close enough to count, but only if the model is
much more sure of it: 0.50 × 1.50 = 0.75." The second is stricter still: s
may stand in for T (th), but must clear 0.50 × 1.90 = 0.95, which is about
right for a contrast that many young speakers have not yet acquired.
During alignment the confidences of the reference phoneme and any listed alternatives are summed, and the scaled cutoff is applied to that sum.
The effect is a per-sound tolerance you control. Rather than one global threshold that is too strict for some phonemes and too lax for others, you set the bar sound by sound, and you do it by editing a YAML file, not by collecting data and retraining.
Scoring behavior
The [Scoring] section of model.config.toml governs the alignment itself.
The settings most likely to matter:
| Setting | What it does |
|---|---|
DefaultCutoffScore | The cutoff for any symbol not listed in cutoffs.yaml. Below 0.5 is lenient, above 0.5 is strict. |
SubCost, InsCost, DelCost | Edit-distance costs for substitution, insertion and deletion. Raising one makes the aligner prefer explanations that avoid it. |
AlignVowels | Prevents nonsensical vowel↔consonant substitutions, reporting a deletion plus an insertion instead. |
IncludeStress | Treat stress markers as distinguishing (AH0 ≠ AH1). Requires both the lexicon and the acoustic model to carry stress. |
IncludeIntraWordInsertion | Whether extra phonemes heard inside a word are attached to that word or reported separately. |
NonSpeechTokenConfidenceScaling | Weights the confidence of silence and noise tokens. Below 1.0 makes the model less willing to call something non-speech, which is useful in noisy rooms. |
CnetLinkThreshold | Discards time slices where the acoustic model put little probability on any speech token. |
MaxCandidateProns | Caps how many combinations of word pronunciations are enumerated for a reference. |
MaxCandidateProns is a performance safeguard, and the number of combinations
grows multiplicatively with the number of words that have pronunciation
variants. For long references made of common function words, the cap can be
reached. PriorityWords in the [Vocab] section controls which words get
their variants explored first, and should list your most common
multi-pronunciation function words.
Non-speech tokens
NonSpeechTokens in [Vocab] lists the symbols the acoustic model emits for
things that are not speech: silence, breath, vocal noise, unknown sounds.
These are handled specially: they are excluded from scoring rather than being
treated as mispronunciations, and they interact with CnetLinkThreshold and
NonSpeechTokenConfidenceScaling above.
Recordings made in real rooms contain plenty of this material. If you find that noisy recordings score too harshly, these settings, rather than the cutoffs, are usually the right place to look.
Adapting the acoustic model
Configuration handles a great deal, but not everything. Where a population is genuinely different from what the model was trained on (young children, a regional accent, a specific recording setup), the acoustic model itself can be adapted to it, which addresses the cause rather than compensating for it downstream.
Adaptation needs representative audio, ideally paired with the verdicts you want the system to reproduce. If your application already records both, you may have a suitable dataset without any additional collection effort. Contact us to discuss what your data supports.
1.9 - Browser Integration
Native applications should use gRPC directly, as shown throughout these docs.
Browsers cannot open a bidirectional gRPC stream, so CAPT server also exposes
an HTTP+JSON gateway and a websocket endpoint for StreamingEvaluate.
The HTTP API is not a property of the server itself: it is served only when
api.Address is set. The configuration shipped with a release sets it to
:8080, which is why the getting started page can curl it
straight away. If you are writing a config from scratch, enable it explicitly
in capt-server.cfg.toml:
[server.http]
api.Address = ":8080"
## Serve the built-in web demo alongside the API.
api.EnableWebDemo = true
## Only needed when the page is served from a different origin
## than the API, e.g. during local development.
# api.EnableCORS = true
JSON conventions
The gateway emits JSON using protobuf field names, so:
- Field names are
snake_case, exactly as in the proto:evaluation_result,start_time_ms,alternative_reference_text. - Enums are full strings, not numbers:
"ALIGNMENT_KIND_MATCH","AUDIO_FORMAT_HEADERED_WAV". - Unpopulated fields are emitted, so a field being present does not imply it was set.
- 64-bit integers are strings, per the protobuf JSON mapping:
"start_time_ms": "570", not570. Parse these before doing arithmetic. bytesfields are base64, so audio is base64-encoded into{"audio": {"data": "..."}}.
Unary calls over HTTP
The two metadata calls and Evaluate are available as ordinary HTTP requests:
| Method | Route |
|---|---|
Version | GET /api/capt/v1/version |
ListModels | GET /api/capt/v1/list-models |
Evaluate | POST /api/capt/v1/evaluate |
StreamingEvaluate | GET /api/capt/v1/streaming-evaluate (websocket) |
const models = await fetch('/api/capt/v1/list-models').then((r) => r.json());
const modelID = models.models[0].id;
Streaming over a websocket
StreamingEvaluate is served over a websocket that carries the same messages
as the gRPC stream, one JSON object per websocket message.
The sequence mirrors the gRPC one: the first message must be the config, and every message after it carries audio.
const ws = new WebSocket(`wss://${location.host}/api/capt/v1/streaming-evaluate`);
ws.onopen = () => {
// First message: the configuration.
ws.send(
JSON.stringify({
config: {
model_id: 'en_US-16khz',
reference_text: 'WHEN THE SUNLIGHT STRIKES',
audio_format: { audio_format_headered: 'AUDIO_FORMAT_HEADERED_WAV' }
}
})
);
};
ws.onmessage = (event) => {
const msg = JSON.parse(event.data);
// The gateway reports a stream-level failure at the top level; a non-fatal
// EvaluationError is a field of the response, so it sits inside the envelope.
if (msg.error != null) {
console.error('stream failed:', msg.error);
return;
}
if (msg.result?.error != null) {
console.warn('non-fatal:', msg.result.error.message);
}
const result = msg.result?.evaluation_result;
if (result == null) {
return;
}
if (result.is_partial) {
showProvisional(result); // live highlighting
} else {
recordFinal(result); // the result you act on
}
};
Responses over the websocket are wrapped in a result envelope,
{"result": {"evaluation_result": {...}, "error": {...}}}, which the plain
gRPC stream does not have. Read msg.result.evaluation_result, not
msg.evaluation_result.
The envelope holds two different errors, and they are not interchangeable. A
non-fatal EvaluationError is a field of
the response, so it arrives at msg.result.error and processing continues.
A top-level msg.error is the gateway’s stream-level failure, with a
different shape, and the stream is over.
Sending audio means base64-encoding each chunk:
// Spreading a whole buffer into String.fromCharCode overflows the call stack
// on anything but small inputs, so encode in fixed-size blocks.
function toBase64(buf) {
const bytes = new Uint8Array(buf);
const block = 0x8000;
let binary = '';
for (let i = 0; i < bytes.length; i += block) {
binary += String.fromCharCode.apply(null, bytes.subarray(i, i + block));
}
return btoa(binary);
}
async function sendChunk(blob) {
ws.send(JSON.stringify({ audio: { data: toBase64(await blob.arrayBuffer()) } }));
}
// A websocket has no half-close, so an empty audio message is how you say
// "that is all the audio". Send it only once the recording is complete.
function endStream() {
ws.send(JSON.stringify({ audio: { data: '' } }));
}
The same helper is what you want for the whole-recording POST described below; a complete take is far past the size at which the one-line spread form fails.
Capturing audio in the browser
The audio you send must match the audio_format you declared. The
MediaRecorder API defaults to a compressed format that CAPT does not accept,
so a browser client generally needs an encoder that can emit WAV chunks while
recording, rather than only at the end.
Two constraints are worth designing around:
- Send small, regular chunks. Around 100-250 ms keeps partial results flowing smoothly. Larger chunks make feedback feel laggy.
- Declare the header once. For headered formats the header belongs at the start of the stream, not on every chunk.
If you do not need live feedback, it is considerably simpler to record the
whole utterance, then POST it to /api/capt/v1/evaluate as a single base64
blob.
The built-in demo
With api.EnableWebDemo = true, the server serves a demo application at the
HTTP address, by default http://localhost:8080. It records from the
microphone or accepts an uploaded file, streams it over the websocket described
above, and renders the per-word and per-phoneme results.
It is a useful way to sanity-check a server, a model and a piece of audio before writing any client code.
1.10 - API Reference
The API is defined as a protobuf spec, so native bindings can be generated in any language with gRPC support. We recommend using buf to generate the bindings.
This section of the documentation is auto-generated from the protobuf spec. The service contains the methods that can be called, and the “messages” are the data structures (objects, classes or structs in the generated code, depending on the language) passed to and from the methods.
CAPTService
Service that implements the Cobalt CAPT API.
Version
Version(VersionRequest) VersionResponse
Returns version information from the server.
ListModels
ListModels(ListModelsRequest) ListModelsResponse
Returns information about the models available on the server.
Evaluate
Evaluate(EvaluateRequest) EvaluateResponse
Performs synchronous speech evaluation by receiving results after all
audio has been sent and processed. It is expected that this request be
typically used for short audio content: less than a minute long. For
longer content, the StreamingEvaluate method should be preferred.
StreamingEvaluate
StreamingEvaluate(StreamingEvaluateRequest) StreamingEvaluateResponse
Performs bidirectional streaming for evaluating speech by receiving results while sending audio. This method is only available via GRPC and not via HTTP+JSON. However, a web browser may use websockets to use this service.
Messages
- If two or more fields in a message are labeled oneof, then each method call using that message must have exactly one of the fields populated
- If a field is labeled
repeated, then the generated code will accept an array (or struct, or list depending on the language).
AlignedToken
AlignedToken contains a single reference token and the alternate hypotheses recognized from the audio that could be aligned with the reference token.
If kind == ALIGNMENT_KIND_MATCH or ALIGNMENT_KIND_SUBSTITUTION, both the reference token and hypotheses list will be populated. The score will be based on the confidence value of the reference token within the hypotheses. In the case of a substitution, the hypotheses list may not contain the actual reference token at all, and the score will be 0.
If kind == ALIGNMENT_KIND_DELETION, the hypotheses will be an empty list and score will be 0.
If kind == ALIGNMENT_KIND_INSERTION, the reference token will be an empty string and score will be 0.
Fields
kind (AlignmentKind ) The kind of alignment for this token.
reference (string ) The actual reference token.
score (float ) The score for this token, between 0 and 1 inclusive, based on the confidence value of the token within the recognized token hypotheses from the audio.
hypotheses (AlignmentHypothesis repeated) The recognized tokens from the audio that have been aligned to this reference token.
AlignedWord
AlignedWord contains the aligned tokens within a single word in the reference text.
Fields
text (string ) The actual word.
start_time_ms (uint64 ) The timestamp at which this token starts in the audio, in milliseconds.
duration_ms (uint64 ) The duration of this token in the audio, in milliseconds.
tokens (AlignedToken repeated) The tokens that make up the word.
AlignmentHypothesis
AlignmentHypothesis contains the metadata for a token recognized from audio, including its timestamps and confidence scores.
Fields
token (string ) The actual token.
confidence (float ) The confidence with which this token was recognized in the audio, between 0 and 1 inclusive.
start_time_ms (uint64 ) The timestamp at which this token starts in the audio, in milliseconds.
duration_ms (uint64 ) The duration of this token in the audio, in milliseconds.
AlternativeAlignment
AlternativeAlignment contains alignments against alternative reference text(s) if specified in the EvaluationConfig.
Fields
reference_text (string ) The alternative reference text.
score (float ) Evaluation score, between 0 and 1 inclusive, with 1 indicating a perfect match between the alternative reference text and what was said in the audio.
alignments (AlignedWord repeated) Alignment between the alternative reference text tokens and recognized tokens.
Audio
Audio to be sent to Capt.
Fields
- data (bytes )
AudioFormat
Format of the audio to be sent for recognition.
Depending on how they are configured, server instances of this service may not support all the formats provided in the API. One format that is guaranteed to be supported is the RAW format with little-endian 16-bit signed samples with the sample rate matching that of the model being requested.
Fields
oneof audio_format.audio_format_raw (AudioFormatRAW ) Audio is raw data without any headers
oneof audio_format.audio_format_headered (AudioFormatHeadered ) Audio has a self-describing header. Headers are expected to be sent at the beginning of the entire audio file/stream, and not in every
Audiomessage.The default value of this type is AUDIO_FORMAT_HEADERED_UNSPECIFIED. If this value is used, the server may attempt to detect the format of the audio. However, it is recommended that the exact format be specified.
AudioFormatRAW
Details of audio in raw format
Fields
encoding (AudioEncoding ) Encoding of the samples. It must be specified explicitly and using the default value of
AUDIO_ENCODING_UNSPECIFIEDwill result in an error.bit_depth (uint32 ) Bit depth of each sample (e.g. 8, 16, 24, 32, etc.). This is a required field.
byte_order (ByteOrder ) Byte order of the samples. This field must be set to a value other than
BYTE_ORDER_UNSPECIFIEDwhen thebit_depthis greater than 8.sample_rate (uint32 ) Sampling rate in Hz. This is a required field.
channels (uint32 ) Number of channels present in the audio. E.g.: 1 (mono), 2 (stereo), etc. This is a required field.
EvaluateRequest
The top-level message sent by the client for the Evaluate method. Both the
EvaluationConfig and Audio fields are required. The entire audio data
must be sent in one request. If your audio data is larger, please use the
StreamingEvaluate call.
Fields
config (EvaluationConfig )
audio (Audio )
EvaluateResponse
The message returned by the server for the Evaluate method.
Fields
evaluation_result (EvaluationResult ) Response from the server. The kind of result depends on the kind of the model that was chosen in the EvaluationConfig.
error (EvaluationError ) A non-fatal error message. If a server encountered a non-fatal error when processing the request, it will be returned in this message. The server will continue to process audio and produce further results. Clients can continue streaming audio even after receiving these messages. This error message is meant to be informational.
An example of when these errors maybe produced: audio is sampled at a lower rate than expected by model, producing possibly less accurate results.
This field will be unset if there is no error to report.
EvaluationConfig
Configuration for a StreamingEvaluateRequest.
Fields
model_id (string ) ID of the model to use. A list of supported IDs can be found using the
ListModelscall.audio_format (AudioFormat ) Format of the audio to be sent.
reference_text (string ) Reference Text that is expected to be contained in the audio.
alternative_reference_text (string repeated) Alternative reference text(s) that are also acceptable if recognized in the audio.
metadata (EvaluationMetadata ) This is an optional field. If there is any metadata associated with the audio being sent, use this field to provide it to the recognizer. The server may record this metadata when processing the request. The server does not use this field for any other purpose.
EvaluationError
Developer-facing error message about a non-fatal process issue.
Fields
- message (string )
EvaluationMetadata
Metadata associated with the evaluation request
Fields
custom_metadata (string ) Any custom metadata that the client wants to associate with the recording. This could be a simple string (e.g. a tracing ID) or structured data (e.g. JSON).
custom_id (string ) This is an optional field to specify custom ID to identify the evaluation request. The custom ID must be a string of upto 64 bytes, and only alphabets, digits, hyphens and underscores are allowed. This ID may be recorded by the server in logs or other storage, and should therefore not include any sensitive information.
EvaluationResult
EvaluationResult contains the result generated by a speech evaluation model.
Fields
is_partial (bool ) If this is set to true, it denotes that the result is an interim partial result, and could change after more audio is processed. If unset, or set to false, it denotes that this is a final result and will not change.
Servers are not required to implement support for returning partial results, and clients should generally not depend on their availability.
score (float ) Overall evaluation score, between 0 and 1 inclusive, with 1 indicating a perfect match between the reference text and what was said in the audio. If alternative reference text(s) are specified, then the score will be take those into account and be based on the reference text that aligns the best with recognized tokens.
alignments (AlignedWord repeated) Alignment between the primary expected reference text tokens and recognized tokens.
alternative_alignments (AlternativeAlignment repeated) Alignments against alternative reference text tokens (if specified) and recognized tokens.
is_speech_endpoint (bool ) If true, indicates that a speech endpoint has been detected.
A speech endpoint signifies that the server believes the user has finished their utterance, typically after detecting a specific duration of silence.
NOTE: Server support for endpoint detection is optional. Clients must be robust to this field never being set.
ListModelsRequest
The top-level message sent by the client for the ListModels method.
ListModelsResponse
The message returned to the client by the ListModels method.
Fields
- models (Model repeated) List of models available for use that match the request.
Model
Description of a CAPT model.
Fields
id (string ) Unique identifier of the model. This identifier is used to choose the model the model when configuring a StreamingEvaluate request.
name (string ) Model name. This is a concise name describing the model, and may be presented to the end-user, for example, to help choose which model to use for their task.
kind (ModelKind ) The specific kind of model. This determines what type of results it sends back, and additional config requirements if any.
attributes (ModelAttributes ) Model Attributes.
ModelAttributes
Attributes of a Capt model.
Fields
sample_rate (uint32 ) Audio sample rate (native) supported by the model.
metadata (ModelMetadata ) Metadata associated with the model.
ModelMetadata
Metadata associated with a Capt model.
Fields
version (string ) Model version in semver format (e.g. 1.0.0). If the version is not known, it will default to “unknown”.
build_date (string ) Date the model was built in YYYY-MM-DD format. If the version is not known, it will default to “unknown”.
StreamingEvaluateRequest
The top level messages sent by the client for the StreamingEvaluate method.
In this streaming call, multiple StreamingEvaluateRequest messages should
be sent. The first message must contain a EvaluationConfig message, and all
subsequent messages must contain Audio only. All Audio messages must
contain non-empty audio. If audio content is empty, the server may choose to
interpret it as end of stream and stop accepting any further messages.
Fields
oneof request.config (EvaluationConfig )
StreamingEvaluateResponse
The message returned by the server for the StreamingEvaluate method.
Fields
evaluation_result (EvaluationResult ) Response from the server. The kind of result depends on the kind of the model that was chosen in the EvaluationConfig.
error (EvaluationError ) A non-fatal error message. If a server encountered a non-fatal error when processing the request, it will be returned in this message. The server will continue to process audio and produce further results. Clients can continue streaming audio even after receiving these messages. This error message is meant to be informational.
An example of when these errors maybe produced: audio is sampled at a lower rate than expected by model, producing possibly less accurate results.
This field will be unset if there is no error to report.
VersionRequest
The top-level message sent by the client for the Version method.
VersionResponse
The message sent by the server for the Version method.
Fields
- version (string ) Version of the server handling these requests.
Enums
AlignmentKind
AlignmentKind represents one of four alignment outcomes possible: a Match, Substitution, Deletion or Insertion.
| Name | Number | Description |
|---|---|---|
| ALIGNMENT_KIND_UNSPECIFIED | 0 | Default value of this type. |
| ALIGNMENT_KIND_MATCH | 1 | A match implies that the reference and recognized token from audio are a match with a reasonable amount of confidence. |
| ALIGNMENT_KIND_SUBSTITUTION | 2 | A substitution implies that the reference token has been replaced by a different token in the recognized tokens from the audio. |
| ALIGNMENT_KIND_DELETION | 3 | A deletion implies that the reference token was not found in the recognized tokens from audio. |
| ALIGNMENT_KIND_INSERTION | 4 | A insertion implies that a extraneous token has been recognized in the audio, that cannot be matched to any token in the reference. |
AudioEncoding
The encoding of the audio data to be sent for recognition.
| Name | Number | Description |
|---|---|---|
| AUDIO_ENCODING_UNSPECIFIED | 0 | AUDIO_ENCODING_UNSPECIFIED is the default value of this type and will result in an error. |
| AUDIO_ENCODING_SIGNED | 1 | PCM signed-integer |
| AUDIO_ENCODING_UNSIGNED | 2 | PCM unsigned-integer |
| AUDIO_ENCODING_IEEE_FLOAT | 3 | PCM IEEE-Float |
| AUDIO_ENCODING_ULAW | 4 | G.711 mu-law |
| AUDIO_ENCODING_ALAW | 5 | G.711 a-law |
AudioFormatHeadered
| Name | Number | Description |
|---|---|---|
| AUDIO_FORMAT_HEADERED_UNSPECIFIED | 0 | AUDIO_FORMAT_HEADERED_UNSPECIFIED is the default value of this type. |
| AUDIO_FORMAT_HEADERED_WAV | 1 | WAV with RIFF headers |
| AUDIO_FORMAT_HEADERED_MP3 | 2 | MP3 format with a valid frame header at the beginning of data |
| AUDIO_FORMAT_HEADERED_FLAC | 3 | FLAC format |
| AUDIO_FORMAT_HEADERED_OGG_OPUS | 4 | Opus format with OGG header |
ByteOrder
Byte order of multi-byte data
| Name | Number | Description |
|---|---|---|
| BYTE_ORDER_UNSPECIFIED | 0 | BYTE_ORDER_UNSPECIFIED is the default value of this type. |
| BYTE_ORDER_LITTLE_ENDIAN | 1 | Little Endian byte order |
| BYTE_ORDER_BIG_ENDIAN | 2 | Big Endian byte order |
ModelKind
| Name | Number | Description |
|---|---|---|
| MODEL_KIND_UNSPECIFIED | 0 | Default value of this type. |
| MODEL_KIND_SPEECH_EVALUATION | 1 | Model for evaluating the accuracy of spoken text from audio. This is done by first recognizing what’s said in the audio, and then aligning expected and recognized tokens (phonemes, syllables, etc.). This model returns results in the form of EvaluationResult messages. |
| MODEL_KIND_PHONEME_EVALUATION | 2 | Model for evaluating the accuracy of phonemes pronounced in isolation from audio. This is done in a way similar to speech evaluation models, but is more accurate for single phonemes in isolation, which a regular speech model may not recognize correctly. This model returns results in the form of EvaluationResult messages. |
Scalar Value Types
| .proto Type | C++ Type | C# Type | Go Type | Java Type | PHP Type | Python Type | Ruby Type |
|---|---|---|---|---|---|---|---|
double | double | double | float64 | double | float | float | Float |
float | float | float | float32 | float | float | float | Float |
int32 | int32 | int | int32 | int | integer | int | Bignum or Fixnum (as required) |
int64 | int64 | long | int64 | long | integer/string | int/long | Bignum |
uint32 | uint32 | uint | uint32 | int | integer | int/long | Bignum or Fixnum (as required) |
uint64 | uint64 | ulong | uint64 | long | integer/string | int/long | Bignum or Fixnum (as required) |
sint32 | int32 | int | int32 | int | integer | int | Bignum or Fixnum (as required) |
sint64 | int64 | long | int64 | long | integer/string | int/long | Bignum |
fixed32 | uint32 | uint | uint32 | int | integer | int | Bignum or Fixnum (as required) |
fixed64 | uint64 | ulong | uint64 | long | integer/string | int/long | Bignum |
sfixed32 | int32 | int | int32 | int | integer | int | Bignum or Fixnum (as required) |
sfixed64 | int64 | long | int64 | long | integer/string | int/long | Bignum |
bool | bool | bool | bool | boolean | boolean | boolean | TrueClass/FalseClass |
string | string | string | string | String | string | str/unicode | String (UTF-8) |
bytes | string | ByteString | []byte | ByteString | string | str | String (ASCII-8BIT) |
1.11 - FAQ
How is this different from using a general-purpose speech recognizer?
A recognizer is built to recover the intended message. That makes it actively unsuitable for assessment, because it will repair what it hears into what it assumes you meant: a mispronounced word is transcribed as the word, and the error you were trying to measure disappears. The better the language model, the more thoroughly it hides exactly the signal you need.
CAPT is anchored to a reference you supply and never substitutes expectation for observation. It also gives you things a transcript cannot:
- Per-phoneme verdicts and scores, not just a word-level transcript.
- Confidence taken from a full distribution over phonemes, not a single guess.
- Timestamps for each recognized phoneme.
- Deterministic, inspectable scoring you can tune, with no risk of a generative model inventing plausible text.
- Operation entirely on your own hardware, on CPU.
What accuracy should I expect?
It depends on the task, the audio and the speakers, so we do not publish a single figure that would mislead you. What we would recommend instead is measuring it on your own data, against the decisions you actually care about.
If you are automating a judgement a person currently makes, the number that matters is agreement with that person on your items, not word error rate, and not any generic benchmark. See calibrating against human judgement.
Does audio leave my infrastructure?
No. CAPT runs on your own hardware: on-premise, in your private cloud, or embedded on a device. There is no call home, and no dependency on an external service at request time.
Audio and results are not stored unless you explicitly enable it. Storage
is off by default; to turn it on, set a storage backend and path in
capt-server.cfg.toml:
[storage]
Type = "localfs"
BasePath = "/audio"
With this enabled, each session is written to local disk as two files, the audio and the evaluation result, organized by UTC date. Everything stays on the machine you run.
Note that metadata.custom_id and metadata.custom_metadata may be written to
logs and to stored results, so they should carry opaque identifiers rather than
personal information.
What hardware does it need?
CAPT is CPU-only; no GPU is required. It runs on x86_64 and Arm64 /
aarch64, and a statically linked build is available for minimal or embedded
images.
Sizing depends on the model and on how many concurrent evaluations you need. Contact us for guidance against your target hardware.
What languages are supported?
Each model covers one language, and the model’s language is fixed at build time. US English models are available today, and models for other languages can be built. Contact us to discuss a specific language.
What audio should I send?
Use uncompressed or losslessly compressed audio (WAV or FLAC) recorded at the
model’s native sample rate, which ListModels reports as
attributes.sample_rate.
Audio sampled below what the model expects yields a non-fatal warning and reduced accuracy. Upsampling before sending does not help; the detail is already gone. Lossy formats such as MP3 are accepted but will cost you some accuracy.
Can more than one person be speaking in the recording?
CAPT scores the audio against the reference text; it does not separate speakers. If someone other than the intended speaker is audible, an assistant reading a prompt for instance, that speech is part of the audio being evaluated and can influence the result, usually appearing as insertions.
Where recordings may contain more than one voice, control it at capture time: record only while the intended speaker is expected to be talking, rather than across the whole interaction. If that is not possible for your setup, talk to us about the options.
Can I add my own words and pronunciations?
Yes, and this is a normal thing to do. Add them to lexicon_addenda.tsv in the
model directory. You can add words the lexicon does not have, and override the
pronunciations of words it does. Words in neither the lexicon nor the addenda
are handled automatically by a grapheme-to-phoneme model, so invented words and
non-words work without any setup.
Can I make scoring stricter or more lenient?
Yes, globally or per phoneme, by editing configuration rather than retraining.
DefaultCutoffScore moves the bar for everything; cutoffs.yaml sets it sound
by sound, and can additionally name specific confusions that should be
tolerated. See Cutoffs and acceptable
confusions.
Do I have to use streaming?
No. Evaluate takes the whole audio in one request and returns one result,
which is the simplest option for scoring a file. Use StreamingEvaluate when
you want results while the speaker is still talking, or for audio longer than
about a minute.
Are partial results guaranteed?
No. Servers are not required to produce them, and clients should not depend on
them. Drive your logic from the final result, the one with is_partial unset
or false, and treat partials as an enhancement for live feedback.
What does a score of 0.5 mean?
It means “exactly borderline” for that phoneme. Scores are not raw probabilities: each phoneme has a cutoff, and the confidence-to-score mapping is built so that the cutoff lands on 0.5. That is what makes 0.5 a meaningful threshold and lets different phonemes be held to different standards.
Overall utterance scores are a summary of many such phoneme scores, so an overall 0.5 does not mean “half correct”. Look at the per-token verdicts before thresholding. See Interpreting Results.