1 - CAPT

On-prem / on-device pronunciation assessment: scores how accurately a speaker pronounced a known piece of text.

CAPT (Computer-Aided Pronunciation Training) answers a different question from a speech recognizer. Transcribe is given audio and asked what was said. CAPT is given audio and the text the speaker was supposed to say, and asked how closely the audio matches it, phoneme by phoneme.

That difference matters whenever you already know the target. Reading assessment, pronunciation practice, language learning drills and scripted-prompt verification all share the same shape: a known reference, a spoken attempt, and a decision about whether the attempt was good enough.

Because the reference is known, CAPT reports where the attempt diverged, not just that it did:

  • Every phoneme of the reference is aligned to the audio and labeled MATCH, SUBSTITUTION, DELETION or INSERTION.
  • Every aligned phoneme carries a score in [0, 1], derived from the acoustic model’s confidence and remapped through per-phoneme thresholds you control.
  • Every recognized phoneme carries a start time and duration, so results can be laid over a waveform.
  • Words are scored as groups of phonemes, and the utterance gets a single overall score.

Crucially, CAPT is error-preserving. A general-purpose recognizer (and a large end-to-end model especially) is built to recover the intended message, so it will quietly repair a mispronunciation into the word it assumes you meant. That behavior is exactly wrong for assessment. CAPT never substitutes what it expects for what it heard: if the speaker said something other than the reference, the divergence is reported.

The engine runs on your own hardware: on-premise, in your private cloud, or embedded on a device. Audio never leaves your infrastructure.

How it works

CAPT is built on top of Cobalt Transcribe, running a phoneme-level acoustic model:

  1. The reference text is phonemized. Each word is looked up in a pronunciation lexicon; words that are not in it are passed to a grapheme-to-phoneme (G2P) model. A word may legitimately have several valid pronunciations, and all of them are considered.
  2. The audio is recognized into a confusion network: for each slice of time, a probability distribution over the phonemes that might have been spoken, rather than a single guess.
  3. Reference and audio are aligned with a weighted edit distance, producing the per-phoneme MATCH / SUBSTITUTION / DELETION / INSERTION labels.
  4. Each phoneme is scored from its confidence in the confusion network, remapped through the cutoffs you configure.

Keeping the full distribution rather than a single best guess is what makes step 4 meaningful: the score for a phoneme reflects how much probability mass the acoustic model actually put on it.

Where to start

1.1 - Getting Started

How to get a CAPT server running on your system

Using Cobalt CAPT

  • A typical CAPT release, provided as a compressed archive, contains a linux binary (capt-server) for the required native CPU architecture, an appropriate Dockerfile, and models.

  • Cobalt CAPT runs either locally on linux or using Docker.

  • Cobalt CAPT serves the CAPT gRPC API on port 2727. The configuration shipped with a release also enables an HTTP+JSON gateway on port 8080 and operational endpoints on port 8081; both are opt-in settings rather than server defaults, so a config written from scratch has gRPC only.

  • To quickly try out CAPT, first start the server as shown below and use the SDK in your preferred language to call it from your application.

Running CAPT Server Locally on Linux

./capt-server --config capt-server.cfg.toml

By default the binary assumes a configuration file named capt-server.cfg.toml in the same directory. A different config file may be specified using the --config argument.

On a successful start the server logs the addresses it is serving on:

2026/09/10 03:57:50 info  {"msg":"server initializing"}
2026/09/10 03:57:50 info  {"msg":"license verified"}
2026/09/10 03:57:51 info  {"source":"transcribe","msg":"formats supported","formats":"[RAW WAV FLAC MP3 Opus]"}
2026/09/10 03:57:51 info  {"msg":"runtime initialized","model_count":"1","init_time_taken":"1.485372717s"}
2026/09/10 03:57:51 info  {"msg":"server started","grpcAddr":"[::]:2727","httpApiAddr":"[::]:8080","httpOpsAddr":"[::]:8081"}

Running CAPT Server as a Docker Container

To build and run the Docker image for CAPT, run:

docker build -t cobalt-capt .
docker run -p 2727:2727 -p 8080:8080 cobalt-capt

Checking the server is up

The HTTP API is the quickest way to confirm a working server, provided api.Address is set in your config as the shipped one does. It exposes the unary calls: Version, ListModels and Evaluate. The bidirectional StreamingEvaluate is gRPC-only, though browsers can reach it over a websocket:

curl -s http://localhost:8080/api/capt/v1/version
{"version":"1.4.0"}

The reported version is that of the capt-server release you were given.

curl -s http://localhost:8080/api/capt/v1/list-models
{
  "models": [
    {
      "id": "en_US-16khz",
      "name": "en_US CAPT (16khz audio)",
      "kind": "MODEL_KIND_SPEECH_EVALUATION",
      "attributes": {
        "sample_rate": 16000,
        "metadata": { "version": "0.0.0", "build_date": "unknown" }
      }
    }
  ]
}

The id of a model in this list is what you pass as model_id when configuring an evaluation, and attributes.sample_rate is the audio sample rate the model expects. See Evaluation Configurations for what to do with them.

How to Get a Copy of the CAPT Server and Models

Contact us for a release best suited to your requirements.

The release you receive is a compressed archive (tar.bz2), generally structured as follows:

release.tar.bz2
├── COPYING
├── README.md
├── capt-server
├── capt-server.cfg.toml
├── Dockerfile
├── api
│   ├── proto             [ cobaltspeech/capt/v1/capt.proto ]
│   └── gen               [ pre-generated Go and Python bindings ]
├── models
│   └── en_US-16khz
│       ├── capt          [ lexicon, G2P model, cutoffs, model config ]
│       └── transcribe    [ acoustic model and decoding graph ]
│
└── cobalt.license.key [ provided separately, needs to be copied over ]
  • The README.md file contains information about the release and instructions for starting the server on your system.

  • The capt-server is the server program, configured using the capt-server.cfg.toml file.

  • The Dockerfile can be used to create a container that will let you run CAPT server on non-linux systems such as macOS and Windows.

  • The api directory holds the protobuf definition of the API and pre-generated client bindings. CAPT’s proto is distributed with the release rather than from a public repository, so this is where it comes from. See Generating SDKs.

  • The models directory contains the evaluation models. Each model directory holds two halves: the transcribe acoustic model that recognizes phonemes, and the capt resources (pronunciation lexicon, G2P model and scoring cutoffs) that the evaluation is built from. Both are described in Tuning and Customization.

System Requirements

Cobalt CAPT runs on Linux, directly as a native application. You can evaluate the product on Windows or macOS using Docker Desktop, but we would not recommend that setup for production.

CAPT runs on x86_64 and Arm64 / aarch64 CPUs, and a statically linked build is available for deployment onto minimal or embedded images. Because the engine is CPU-only and the models are small, the same release can be deployed on a server, in a private cloud, or embedded on a device. Tell us the target and we will provide a build for it.

Sizing depends on the model and on how many evaluations you need to run concurrently. Please contact us for sizing guidance against your target hardware and expected load.

To integrate Cobalt CAPT into your application, follow the next steps to install or generate the SDK in a language of your choice.

1.2 - Generating SDKs

How to generate a client SDK for your project from the CAPT proto API definition.
  • The CAPT API is defined as a protocol buffer specification, a proto file. This file allows a developer to auto-generate client SDKs for a number of different programming languages.

  • The CAPT proto definition, cobaltspeech/capt/v1/capt.proto, is supplied to you as part of your release, together with pre-generated Go and Python bindings. If you only need those two languages you can skip straight to Using the pre-generated SDKs; to generate bindings for another language, see Generating SDKs.

The relevant part of your release looks like this:

release.tar.bz2
└── api
    ├── proto
    │   └── cobaltspeech
    │       └── capt
    │           └── v1
    │               └── capt.proto
    └── gen
        ├── go
        │   └── cobaltspeech/capt/v1/{capt.pb.go, capt_grpc.pb.go}
        └── py
            └── cobaltspeech/capt/v1/{capt_pb2.py, capt_pb2_grpc.py, capt_pb2.pyi}

Using the pre-generated SDKs

Golang

Copy the api/gen/go tree into your module, or add it as a dependency, and import the package:

import captpb "github.com/your-org/your-module/gen/go/cobaltspeech/capt/v1"

You will also need the gRPC and protobuf runtime libraries:

go get google.golang.org/protobuf
go get google.golang.org/grpc

Python

The Python bindings depend on Python >= 3.8. Place api/gen/py on your PYTHONPATH, or copy the cobaltspeech package into your project, and install the runtime dependencies:

pip install --upgrade pip
pip install --upgrade protobuf grpcio

Then import the modules:

import cobaltspeech.capt.v1.capt_pb2 as capt
import cobaltspeech.capt.v1.capt_pb2_grpc as capt_grpc

Generating SDKs from the proto

To generate bindings for a language we do not ship, generate them yourself from capt.proto. We recommend using buf, a command line tool that can generate documentation, schemas and SDK code for many languages.

Step 1. Installing buf

COBALT="${HOME}/cobalt"
mkdir -p "${COBALT}/bin"

VERSION="1.72.0"
URL="https://github.com/bufbuild/buf/releases/download/v${VERSION}/buf-$(uname -s)-$(uname -m)"
curl -L ${URL} -o "${COBALT}/bin/buf"

# Give executable permissions and add to $PATH.
chmod +x "${COBALT}/bin/buf"
export PATH="${PATH}:${COBALT}/bin"
brew install bufbuild/buf/buf

Step 2. Writing a buf.gen.yaml

Create a buf.gen.yaml next to the api/proto directory from your release. The example below generates Go and Python; other plugins can be added for more languages.

version: v1

managed:
  enabled: true
  go_package_prefix:
    default: github.com/your-org/your-module/gen

plugins:
  # Golang
  - plugin: buf.build/grpc/go
    out: gen/go
    opt: paths=source_relative

  - plugin: buf.build/protocolbuffers/go
    out: gen/go
    opt: paths=source_relative

  # Python
  - plugin: buf.build/grpc/python
    out: gen/py

  - plugin: buf.build/protocolbuffers/python
    out: gen/py

  - plugin: buf.build/protocolbuffers/pyi
    out: gen/py

Step 3. Generating code

# Removing any previously generated files.
rm -rf ./gen

# Generating code for the proto files inside the `proto` directory.
buf generate proto

You should now have a gen folder containing the generated code. The latest version of the CAPT API is v1. Import or copy the generated files into your project as per the conventions of your language.

gen
└── py
  └── cobaltspeech
    └── capt
      └── v1
        ├── capt_pb2_grpc.py
        ├── capt_pb2.py
        └── capt_pb2.pyi
gen
└── go
  └── cobaltspeech
    └── capt
      └── v1
        ├── capt_grpc.pb.go
        └── capt.pb.go

Once you have an SDK, continue to connecting to the server.

1.3 - Connecting to the Server

Describes how to connect to a running Cobalt CAPT server instance.

Once you have your CAPT server up and running, and have installed or generated the SDK for your project, you can connect to it by “dialing” a gRPC connection.

First, you need the address where the server is running: e.g. host:grpc_port. By default this is localhost:2727, and it is logged to the terminal when you start CAPT server as grpcAddr:

2026/09/10 03:57:51 info  {"msg":"server started","grpcAddr":"[::]:2727","httpApiAddr":"[::]:8080","httpOpsAddr":"[::]:8081"}

Default Connection

The following snippet connects to the server and queries its version, using an “insecure” gRPC channel. This would be the case if you have just started a local instance of CAPT server without TLS enabled.

import grpc
import cobaltspeech.capt.v1.capt_pb2 as capt
import cobaltspeech.capt.v1.capt_pb2_grpc as capt_grpc

serverAddress = "localhost:2727"

# Using a channel without TLS enabled.
channel = grpc.insecure_channel(serverAddress)
client = capt_grpc.CAPTServiceStub(channel)

# Get server version.
versionResp = client.Version(capt.VersionRequest())
print(versionResp)

# Get the list of models available on the server.
modelResp = client.ListModels(capt.ListModelsRequest())
for model in modelResp.models:
    print(model)
package main

import (
	"context"
	"fmt"
	"os"

	"google.golang.org/grpc"
	"google.golang.org/grpc/credentials/insecure"

	captpb "github.com/your-org/your-module/gen/go/cobaltspeech/capt/v1"
)

func main() {
	const serverAddress = "localhost:2727"

	ctx, cancel := context.WithCancel(context.Background())
	defer cancel()

	// Using a channel without TLS enabled.
	conn, err := grpc.NewClient(serverAddress,
		grpc.WithTransportCredentials(insecure.NewCredentials()))
	if err != nil {
		fmt.Printf("failed to dial gRPC connection: %v\n", err)
		os.Exit(1)
	}
	defer conn.Close()

	client := captpb.NewCAPTServiceClient(conn)

	// Get server version.
	versionResp, err := client.Version(ctx, &captpb.VersionRequest{})
	if err != nil {
		fmt.Printf("failed to get server version: %v\n", err)
		os.Exit(1)
	}

	fmt.Printf("%v\n", versionResp)

	// Get the list of models available on the server.
	modelResp, err := client.ListModels(ctx, &captpb.ListModelsRequest{})
	if err != nil {
		fmt.Printf("failed to list models: %v\n", err)
		os.Exit(1)
	}

	for _, m := range modelResp.GetModels() {
		fmt.Printf("%v\n", m)
	}
}

Connect with TLS

In our recommended setup for deployment, TLS is enabled in the gRPC connection, and clients validate the server’s SSL certificate to make sure they are talking to the right party. This is similar to how “https” connections work in web browsers.

TLS is enabled server-side by providing a certificate and key in capt-server.cfg.toml:

[server.grpc]
Address = ":2727"
CertFile = "capt-server.crt"
KeyFile = "capt-server.key"

The following snippets show how to connect to a CAPT server that has TLS enabled.

import grpc
import cobaltspeech.capt.v1.capt_pb2 as capt
import cobaltspeech.capt.v1.capt_pb2_grpc as capt_grpc

serverAddress = "capt.your-org.internal:2727"

# Setup a gRPC connection with TLS. You can optionally provide your own
# root certificates and private key to grpc.ssl_channel_credentials()
# for mutually authenticated TLS.
creds = grpc.ssl_channel_credentials()
channel = grpc.secure_channel(serverAddress, creds)
client = capt_grpc.CAPTServiceStub(channel)

# Get server version.
versionResp = client.Version(capt.VersionRequest())
print(versionResp)
package main

import (
	"context"
	"crypto/tls"
	"fmt"
	"os"

	"google.golang.org/grpc"
	"google.golang.org/grpc/credentials"

	captpb "github.com/your-org/your-module/gen/go/cobaltspeech/capt/v1"
)

func main() {
	const serverAddress = "capt.your-org.internal:2727"

	// Setup a gRPC connection with TLS. You can optionally provide your own
	// root certificates and private key through tls.Config for mutually
	// authenticated TLS.
	tlsCfg := tls.Config{}
	creds := credentials.NewTLS(&tlsCfg)

	ctx, cancel := context.WithCancel(context.Background())
	defer cancel()

	conn, err := grpc.NewClient(serverAddress, grpc.WithTransportCredentials(creds))
	if err != nil {
		fmt.Printf("failed to dial gRPC connection: %v\n", err)
		os.Exit(1)
	}
	defer conn.Close()

	client := captpb.NewCAPTServiceClient(conn)

	versionResp, err := client.Version(ctx, &captpb.VersionRequest{})
	if err != nil {
		fmt.Printf("failed to get server version: %v\n", err)
		os.Exit(1)
	}

	fmt.Printf("%v\n", versionResp)
}

Client Authentication

In some setups it is desirable for the server to validate the clients connecting to it, and only respond to ones it can verify. If your CAPT server is configured to do client authentication, you will need to present the appropriate certificate and key when connecting to it.

Note that in client-authentication mode the client still also verifies the server’s certificate, so this setup uses mutually authenticated TLS.

creds = grpc.ssl_channel_credentials(
  root_certificates=root_certificates,  # PEM certificate as byte string
  private_key=private_key,              # PEM client key as byte string
  certificate_chain=certificate_chain,  # PEM client certificate as byte string
)
// Root PEM certificate for validating a self-signed server certificate.
var rootCert []byte

// Client PEM certificate and private key.
var certPem, keyPem []byte

caCertPool := x509.NewCertPool()
if ok := caCertPool.AppendCertsFromPEM(rootCert); !ok {
	fmt.Printf("unable to use given caCert\n")
	os.Exit(1)
}

clientCert, err := tls.X509KeyPair(certPem, keyPem)
if err != nil {
	fmt.Printf("unable to use given client certificate and key: %v\n", err)
	os.Exit(1)
}

tlsCfg := tls.Config{
	RootCAs:      caCertPool,
	Certificates: []tls.Certificate{clientCert},
}

creds := credentials.NewTLS(&tlsCfg)

1.4 - Evaluation Configurations

The options available when configuring an evaluation request.

Every evaluation begins with an EvaluationConfig. It is the first message of a StreamingEvaluate stream, and the config field of a unary Evaluate request. This page describes what goes in it.

Choosing a model

model_id selects the model to evaluate against, and must be one of the id values returned by ListModels. Each model also reports the sample_rate it expects and its kind:

  • MODEL_KIND_SPEECH_EVALUATION evaluates words and sentences. The reference_text is ordinary text.
  • MODEL_KIND_PHONEME_EVALUATION evaluates phonemes produced in isolation, which a word-oriented model may not recognize correctly. The reference_text is a sequence of phones. See Phoneme Models.

Both kinds return the same EvaluationResult structure, so only the reference_text you supply differs.

Reference text

reference_text is the text the speaker is expected to have said: the thing the audio is scored against. This is the field that makes CAPT what it is: the evaluation is anchored to it.

cfg = capt.EvaluationConfig(
    model_id="en_US-16khz",
    reference_text="WHEN THE SUNLIGHT STRIKES",
)

A few properties worth knowing:

  • Lexicon lookup is case-insensitive, so when, When and WHEN behave identically. The uppercase reference text used throughout these examples is a convention, not a requirement.
  • Punctuation handling is a per-model setting, RemovePunctuation in the model’s [Tokenization] config, and it is off in the en_US-16khz model shipped today. Punctuation therefore stays attached to the word: a reference of WHEN, THE SUNLIGHT STRIKES comes back with text of WHEN,, and because the token no longer matches the lexicon it is phonemized by G2P, which changed that word’s score in our own testing. Sentence-final marks are usually harmless, but the safe habit is to send the words alone. See Tuning and Customization.
  • Words not in the lexicon are handled by a G2P model, so invented words, proper nouns and non-words can be used as reference text without any setup. This is what makes non-word decoding tasks work. See Tuning and Customization.
  • Words with several valid pronunciations are all considered, and the variant that best fits the audio is the one reported. A speaker is not penalized for saying “THE” as D i rather than D @.

Alternative reference text

alternative_reference_text accepts additional reference texts that would also be acceptable. Each is aligned independently and returned in alternative_alignments with its own score, and the top-level score becomes the best score across the primary and all alternatives.

cfg = capt.EvaluationConfig(
    model_id="en_US-16khz",
    reference_text="WHEN THE SUNLIGHT STRIKES",
    alternative_reference_text=["WHEN THE SUN LIGHTS STRIKE"],
)

Audio format

audio_format tells the server how to interpret the bytes you send. Two shapes are available.

For files that carry their own header, name the container and let the server read the rest:

cfg = capt.EvaluationConfig(
    model_id="en_US-16khz",
    reference_text="WHEN THE SUNLIGHT STRIKES",
    audio_format=capt.AudioFormat(
        audio_format_headered=capt.AUDIO_FORMAT_HEADERED_WAV,
    ),
)

Supported headered formats are WAV, MP3, FLAC and OGG Opus. Headers are expected once, at the beginning of the stream, not in every Audio message.

audio_format may be omitted entirely for headered audio. The server then detects the format from the header, which is why several examples in these pages leave it out. Naming it explicitly is still recommended: it turns an unreadable or unexpected header into a clear error instead of a detection attempt.

Raw audio is the exception. It carries no header to detect, so audio_format_raw must be supplied in full, and encoding must be set to something other than AUDIO_ENCODING_UNSPECIFIED or the request is rejected.

For raw samples with no header, describe them fully:

cfg = capt.EvaluationConfig(
    model_id="en_US-16khz",
    reference_text="WHEN THE SUNLIGHT STRIKES",
    audio_format=capt.AudioFormat(
        audio_format_raw=capt.AudioFormatRAW(
            encoding=capt.AUDIO_ENCODING_SIGNED,
            bit_depth=16,
            byte_order=capt.BYTE_ORDER_LITTLE_ENDIAN,
            sample_rate=16000,
            channels=1,
        ),
    ),
)

Every server supports raw little-endian 16-bit signed samples at the model’s own sample rate; support for the other formats depends on how the server was built and configured.

Request metadata

metadata is optional and is never used to influence evaluation. The server may record it in logs or stored results.

cfg = capt.EvaluationConfig(
    model_id="en_US-16khz",
    reference_text="WHEN THE SUNLIGHT STRIKES",
    metadata=capt.EvaluationMetadata(
        custom_id="session-42-item-07",
        custom_metadata='{"item":"7","form":"A"}',
    ),
)
  • custom_id accepts up to 64 bytes, restricted to letters, digits, hyphens and underscores. Useful as a tracing or correlation ID.
  • custom_metadata is a free-form string; a plain tag or structured data such as JSON.

Once you have a config, continue to Streaming Evaluation.

1.5 - Streaming Evaluation

Describes how to stream audio to CAPT server and receive evaluation results.

CAPT offers two ways to submit audio.

  • StreamingEvaluate is the primary method. It is bidirectional: you send audio as it is captured and receive results as the audio is processed. Use it for anything interactive, and for audio of any length.
  • Evaluate is a unary convenience method that takes the whole audio in a single request and returns a single result. It is intended for short audio, less than a minute, and is the simplest way to score an audio file.

Both accept the same EvaluationConfig and return the same EvaluationResult.

Streaming from an audio file

In a StreamingEvaluate stream, the first message must contain the config, and every subsequent message must contain audio.

An empty Audio message ends the stream. CAPT server treats it as the end of the audio: it finishes processing, sends the final result, and accepts nothing further, so a later send fails. Over gRPC you normally do not need it, because half-closing the stream (CloseSend in Go, exhausting the request iterator in Python) says the same thing. It matters for browser clients, where a websocket has no equivalent half-close.

Do not send one to mean “no audio yet”. If it arrives before the audio is complete, what you get depends on the format, and a client’s state machine needs to handle both:

  • You always receive a final result for the audio sent so far, with is_partial false. It reflects the truncated input, not the utterance you intended, so it will usually be full of deletions.
  • With a headered format the stream then fails. The header promised more data than arrived, so the server reports a container error such as wav: unexpected EOF after the final result. This happens whether or not you send anything afterwards.
  • With raw audio the stream ends cleanly, because there is no container to leave incomplete. The result is still only of the audio you sent.

In other words: a final result is not by itself evidence that the whole utterance was processed. Send the empty message only once the recording is complete, and treat an error after a final result as “the result is incomplete”, not “the result is invalid”.

The examples below stream a WAV file in 100 ms chunks and print each result as it arrives.

import grpc
import cobaltspeech.capt.v1.capt_pb2 as capt
import cobaltspeech.capt.v1.capt_pb2_grpc as capt_grpc

serverAddress = "localhost:2727"

channel = grpc.insecure_channel(serverAddress)
client = capt_grpc.CAPTServiceStub(channel)

# Get the list of models on the server and use the first one.
modelResp = client.ListModels(capt.ListModelsRequest())
modelID = modelResp.models[0].id

cfg = capt.EvaluationConfig(
    model_id=modelID,
    reference_text="WHEN THE SUNLIGHT STRIKES",
    audio_format=capt.AudioFormat(
        audio_format_headered=capt.AUDIO_FORMAT_HEADERED_WAV,
    ),
)

# The first request must contain only the configuration; subsequent
# requests carry audio bytes. A generator is a convenient way to do this.
def stream(cfg, audio, bufferSize=3200):
    yield capt.StreamingEvaluateRequest(config=cfg)

    data = audio.read(bufferSize)
    while len(data) > 0:
        yield capt.StreamingEvaluateRequest(audio=capt.Audio(data=data))
        data = audio.read(bufferSize)

with open("test.wav", "rb") as audio:
    for resp in client.StreamingEvaluate(stream(cfg, audio)):
        if resp.HasField("error"):
            print(f"warning: {resp.error.message}")

        result = resp.evaluation_result
        print(f"partial={result.is_partial} score={result.score:.4f}")

        # Only act on final results; partials will still change.
        if not result.is_partial:
            for word in result.alignments:
                print(f"  {word.text}: " + " ".join(
                    f"{t.reference}={t.score:.2f}" for t in word.tokens
                ))
package main

import (
	"context"
	"fmt"
	"io"
	"os"

	"google.golang.org/grpc"
	"google.golang.org/grpc/credentials/insecure"

	captpb "github.com/your-org/your-module/gen/go/cobaltspeech/capt/v1"
)

func main() {
	const serverAddress = "localhost:2727"

	audio, err := os.ReadFile("test.wav")
	if err != nil {
		fmt.Printf("failed to read audio: %v\n", err)
		os.Exit(1)
	}

	conn, err := grpc.NewClient(serverAddress,
		grpc.WithTransportCredentials(insecure.NewCredentials()))
	if err != nil {
		fmt.Printf("failed to dial gRPC connection: %v\n", err)
		os.Exit(1)
	}
	defer conn.Close()

	client := captpb.NewCAPTServiceClient(conn)

	stream, err := client.StreamingEvaluate(context.Background())
	if err != nil {
		fmt.Printf("failed to open stream: %v\n", err)
		os.Exit(1)
	}

	// The first message must contain the config.
	if err := stream.Send(&captpb.StreamingEvaluateRequest{
		Request: &captpb.StreamingEvaluateRequest_Config{
			Config: &captpb.EvaluationConfig{
				ModelId:       "en_US-16khz",
				ReferenceText: "WHEN THE SUNLIGHT STRIKES",
				AudioFormat: &captpb.AudioFormat{
					AudioFormat: &captpb.AudioFormat_AudioFormatHeadered{
						AudioFormatHeadered: captpb.AudioFormatHeadered_AUDIO_FORMAT_HEADERED_WAV,
					},
				},
			},
		},
	}); err != nil {
		fmt.Printf("failed to send config: %v\n", err)
		os.Exit(1)
	}

	// Send audio in the background while results are read below.
	go func() {
		const chunk = 3200 // 100ms of 16kHz 16-bit mono audio.

		// min() is a builtin from Go 1.21; on older toolchains compute the
		// bound explicitly.
		for i := 0; i < len(audio); i += chunk {
			end := min(i+chunk, len(audio))

			if err := stream.Send(&captpb.StreamingEvaluateRequest{
				Request: &captpb.StreamingEvaluateRequest_Audio{
					Audio: &captpb.Audio{Data: audio[i:end]},
				},
			}); err != nil {
				return
			}
		}

		_ = stream.CloseSend()
	}()

	for {
		resp, err := stream.Recv()
		if err == io.EOF {
			break
		}

		if err != nil {
			fmt.Printf("failed to receive result: %v\n", err)
			os.Exit(1)
		}

		if e := resp.GetError(); e != nil {
			fmt.Printf("warning: %v\n", e.GetMessage())
		}

		result := resp.GetEvaluationResult()
		fmt.Printf("partial=%v score=%.4f\n", result.GetIsPartial(), result.GetScore())

		// Only act on final results; partials will still change.
		if !result.GetIsPartial() {
			for _, word := range result.GetAlignments() {
				fmt.Printf("  %s:", word.GetText())

				for _, t := range word.GetTokens() {
					fmt.Printf(" %s=%.2f", t.GetReference(), t.GetScore())
				}

				fmt.Println()
			}
		}
	}
}

Streaming a 1.5 second recording of “when the sunlight strikes” against that reference produces a short sequence of results:

partial=true  score=0.5514
partial=true  score=0.6128
partial=true  score=0.6128
partial=false score=0.6128

Partial results

While audio is still arriving, the server emits partial results, marked is_partial = true. A partial reflects the audio received so far, and any part of it may change as more audio arrives. In the sequence above the score rises as the last word is heard.

When the stream ends, the server emits a final result with is_partial = false. That result is stable.

  • Use partials for live feedback: highlighting words as they are read, showing a provisional score, driving a progress indicator.
  • Use the final result for any decision you record. Scoring an item on a partial risks scoring it on half a word.

Speech endpointing

A result may set is_speech_endpoint = true to indicate the server believes the speaker has finished, typically after a period of silence. This is useful for closing the microphone automatically once an answer has been given.

Endpoint detection is optional, and clients must be robust to the field never being set. Do not treat it as the signal that results are final; that is what is_partial = false is for.

Non-fatal errors

Both response types carry an optional EvaluationError beside the result. It reports conditions that degrade quality but do not stop processing, most commonly audio sampled at a lower rate than the model expects. The server continues, and you may keep streaming. Log these rather than aborting; they usually point at a recording pipeline that needs attention.

Evaluating a whole file at once

For short audio, Evaluate avoids the streaming machinery entirely. Over the HTTP+JSON gateway that is a single POST:

curl -s -X POST http://localhost:8080/api/capt/v1/evaluate \
  -H 'Content-Type: application/json' \
  -d '{
    "config": {
      "model_id": "en_US-16khz",
      "reference_text": "WHEN THE SUNLIGHT STRIKES",
      "audio_format": {"audio_format_headered": "AUDIO_FORMAT_HEADERED_WAV"}
    },
    "audio": {"data": "<base64-encoded WAV file>"}
  }'

The data field is the audio file, base64 encoded. See Interpreting Results for what comes back.

1.6 - Interpreting Results

How to read an EvaluationResult and turn it into a decision.

This is the page worth reading carefully. Getting audio into CAPT is straightforward; deciding what its output means for your application is where the design work is.

The shape of a result

An EvaluationResult is a tree:

EvaluationResult
├── score                    overall score for the utterance, 0..1
├── is_partial               true while more audio may still change this
├── is_speech_endpoint       true if the speaker appears to have finished
├── alignments               [ ]AlignedWord
│   └── AlignedWord
│       ├── text             the reference word
│       ├── start_time_ms    when the word starts in the audio
│       ├── duration_ms      how long it lasts
│       └── tokens           [ ]AlignedToken   (the phonemes)
│           └── AlignedToken
│               ├── kind           MATCH | SUBSTITUTION | DELETION | INSERTION
│               ├── reference      the expected phoneme
│               ├── score          0..1 for this phoneme
│               └── hypotheses     [ ]AlignmentHypothesis  (what was heard)
│                   └── token, confidence, start_time_ms, duration_ms
└── alternative_alignments   the same, once per alternative reference text

The reference text drives the structure: your words appear as AlignedWord entries in order, and within each one there is an AlignedToken per expected phoneme.

The four alignment kinds

Every expected phoneme gets exactly one of four verdicts.

KindMeaningreferencehypothesesscore
MATCHThe expected phoneme was heard with enough confidencesetpopulated0.5 to 1.0
SUBSTITUTIONThe expected phoneme was not heard with enough confidence: either something else was heard in its place, or the phoneme itself was heard but below its cutoffsetpopulated0 to below 0.5
DELETIONThe phoneme was not heard at allsetempty0
INSERTIONSomething extra was heard that the reference does not account foremptypopulated0

A real example makes this concrete. Below is the result of scoring a recording of “when the sunlight strikes” against a deliberately wrong reference, “WHEN THE MOONLIGHT STRIKES” (phonemes are X-SAMPA):

overall score: 0.5157

WHEN         MATCH w (1.00)  MATCH E (0.99)  MATCH n (1.00)
THE          MATCH D (0.87)  MATCH @ (0.67)
MOONLIGHT    SUB   m (0.00)  SUB   u (0.00)  MATCH n (0.65)
             MATCH l (0.78)  MATCH aI (0.82) MATCH t (0.73)
STRIKES      MATCH s (0.79)  DEL   t (0.00)  DEL   r\ (0.00)
             DEL   aI (0.00) SUB   k (0.49)  DEL   s (0.00)

Read that as a diagnosis, not just a number. The first two words were said as expected. In MOONLIGHT the leading m u was not there (the speaker said s V n), but the shared tail n l aI t matched, so those phonemes score well. STRIKES shows what happens when the audio has already run out: the reference still has phonemes to account for, and they come back as deletions.

The overall score of 0.52 is the average of that mixture. This is why an overall score alone is rarely the right thing to threshold. 0.52 here does not mean “roughly half right”, it means “two words right and two words wrong”.

Insertions

Insertions do not belong to any reference word, so they are reported in their own AlignedWord entries with an empty text, positioned in the sequence where they were heard. Scoring the same audio against the single word “ELEPHANT” shows this:

overall score: 0.3966

(insertion)  INS  (heard: w)
ELEPHANT     MATCH E (0.99)  SUB l (0.00)  MATCH @ (0.67)
             SUB   f (0.00)  SUB n= (0.00) MATCH t (0.73)
(insertion)  INS  (heard: k)  INS (heard: s)

The speaker said far more than the reference accounted for, and the leftover audio at each end is reported as insertions rather than being silently discarded.

Where the numbers come from

Phoneme scores

The acoustic model does not commit to a single phoneme per time slice. It produces a distribution, a confusion network, and the score for an expected phoneme is derived from how much probability mass landed on it.

You can see this in the hypotheses list. Here is a single MATCH from a real result:

{
  "kind": "ALIGNMENT_KIND_MATCH",
  "reference": "n",
  "score": 0.9798913,
  "hypotheses": [
    { "token": "n",  "confidence": 0.963, "start_time_ms": "570", "duration_ms": "119" },
    { "token": "m",  "confidence": 0.029, "start_time_ms": "570", "duration_ms": "119" },
    { "token": "n=", "confidence": 0.006, "start_time_ms": "570", "duration_ms": "119" },
    { "token": "N",  "confidence": 0.002, "start_time_ms": "570", "duration_ms": "119" }
  ]
}

The model heard n with 0.963 confidence, but also considered m, n= and N. The reported score of 0.98 is that confidence remapped through the model’s cutoff for n. See Tuning and Customization for how that remapping works and how to change it.

The practical consequence: a score is a calibrated quantity, not a raw probability. A cutoff is chosen so that a score of exactly 0.5 sits at the boundary between acceptable and unacceptable for that phoneme. That is what makes 0.5 a meaningful place to threshold, and it is why different phonemes can be held to different standards.

Word and utterance scores

An AlignedWord does not carry its own score field; a word’s quality is the scores of its phonemes. Compute whatever summary suits your task: the mean, the minimum, or the fraction of phonemes that came back MATCH.

The top-level score is the overall score for the utterance: the mean of the scores of the reference tokens. If alternative reference texts were supplied, it is the best score across the primary and all alternatives.

Timestamps

Timestamps live on the hypotheses, not on the token. An AlignedToken has no time fields of its own; start_time_ms and duration_ms are on each AlignmentHypothesis inside it, and on the enclosing AlignedWord.

This has one consequence that catches people out:

Likewise, a word made up entirely of deletions has nothing to anchor to, and its start_time_ms and duration_ms are reported as 0.

Turning results into a decision

A pass/fail mark for one item

If your application needs a yes/no verdict (did the speaker say the target correctly?), do not reach for the overall score first. Consider what the task actually requires:

  • Strict, verbatim tasks. Any divergence is a failure. Require every token to be MATCH:

    correct = all(
        t.kind == capt.ALIGNMENT_KIND_MATCH
        for w in result.alignments
        for t in w.tokens
    )
    

    This is exactly right when the rule is “any change at all, however minor, is wrong”, because substitutions, deletions and insertions all break it.

  • Tolerant tasks. Some divergence is acceptable. A slightly indistinct consonant should not fail an otherwise good attempt. Threshold on the proportion of matched phonemes, or on the mean phoneme score, rather than demanding a clean sweep. Note that a tolerant rule built on the overall score, or on reference tokens alone, will not notice extra speech: decide separately whether insertions should fail the item.

  • Targeted tasks. Only some phonemes matter: the contrast the item is testing. Score only those tokens and ignore the rest.

Choosing between several candidate answers

When you have a known correct answer and a set of known incorrect answers, and you need to know which one was said, evaluate the audio once per candidate and compare the resulting scores:

candidates = ["BED", "LOUNGE", "SOFA"]
scores = {}

for candidate in candidates:
    cfg = capt.EvaluationConfig(model_id=modelID, reference_text=candidate)
    with open(path, "rb") as audio:
        for resp in client.StreamingEvaluate(stream(cfg, audio)):
            if not resp.evaluation_result.is_partial:
                scores[candidate] = resp.evaluation_result.score

best = max(scores, key=scores.get)

This gives you a score per candidate, so you can see not just the winner but the margin, so you can reject the result as unclear when two candidates score alike. Packing the candidates into alternative_reference_text instead would collapse them into a single best score and lose exactly that information.

Calibrating against human judgement

If you are automating a decision a person currently makes, the metric that matters is agreement with that person, not any internal notion of accuracy.

The recommended approach is to collect audio alongside the human verdict, run CAPT over it, then sweep your decision rule across the collected set and count true positives, true negatives, false positives and false negatives at each setting. That tells you where to set the threshold, and, just as importantly, what the residual disagreement rate is and which way it leans. A rule that is wrong in the safe direction for your use case is often better than one that is wrong less often overall.

Per-phoneme cutoffs can then correct systematic biases that a single global threshold cannot; see Tuning and Customization.

Alternative alignments

If you supplied alternative_reference_text, each alternative comes back in alternative_alignments as an AlternativeAlignment carrying its own reference_text, score, and full alignments tree with the same structure described above.

The top-level score is the best across the primary and the alternatives, so if you only care whether any acceptable rendering was produced, the top-level score is enough. If you care which one, see above.

1.7 - Phoneme Models

Evaluating individual phonemes produced in isolation.

CAPT models come in two kinds, reported as kind by ListModels.

  • Word models (MODEL_KIND_SPEECH_EVALUATION) evaluate phonemes in word context. The reference_text is ordinary text, and this is what you want for reading words, sentences and passages.
  • Phoneme models (MODEL_KIND_PHONEME_EVALUATION) evaluate phonemes produced in isolation. The reference_text is a sequence of phones.

The distinction matters because a sound produced on its own is acoustically quite different from the same sound inside a word. A model trained on connected speech often mis-recognizes an isolated phoneme, because nothing in its training looked like that. If your task asks a speaker to produce a single sound (“say the /sh/ sound”), a phoneme model is substantially more reliable.

Both kinds are used through exactly the same API calls and return the same EvaluationResult. Only the reference_text differs.

Reference text for phoneme models

For a phoneme model, reference_text is one or more phones separated by whitespace:

# A single phone.
cfg = capt.EvaluationConfig(model_id=phonemeModelID, reference_text="AA")

# A sequence of phones.
cfg = capt.EvaluationConfig(model_id=phonemeModelID, reference_text="Z AE P")

Phone groups

Some phones are only meaningfully produced as a pair, and the model treats such a pair as a single unit called a phone group. Phone groups are written with the members joined by a period:

# A phone group.
cfg = capt.EvaluationConfig(model_id=phonemeModelID, reference_text="AO.NG")

# Mixed with ordinary phones.
cfg = capt.EvaluationConfig(model_id=phonemeModelID, reference_text="F AE.NG")

For all practical purposes a phone group behaves as its own distinct phone: it is matched, substituted or deleted as a single token, and receives a single score.

The phone set

The phones a model accepts are specific to that model, both the inventory and which groups exist. Supplying a phone the model does not know is an error, so check against the set for the model you were given. A typical US English phoneme model accepts:

AA     AA.R   AE     AE.NG  AH     AH.NG  AO     AO.NG
AO.R   AW     AY     B      CH     D      DH     EH
EH.R   ER     EY     F      G      HH     IH     IH.NG
IH.R   IY     JH     K      L      M      N      NG
OW     OY     P      R      S      SH     T      TH
UH     UW     V      W      Y      Y.UW   Z      ZH

Reading the results

Results have the same structure as any other evaluation. For each phone in the reference_text there is an aligned token:

  • ALIGNMENT_KIND_MATCH means the phone was found in the audio with a high degree of confidence.
  • ALIGNMENT_KIND_SUBSTITUTION means a different phone was recognized in its place, or the phone was recognized only with low confidence.
  • ALIGNMENT_KIND_DELETION means the phone was not recognized at all.

Anything else recognized in the audio that does not align to the reference is reported as insertions, grouped separately.

Isolated-phoneme audio very often contains more than the phoneme itself: a speaker clearing their throat, a lead-in, or the assessor’s prompt. Those show up as insertions, which is why the example below scores 1.0 despite three extra tokens: the reference phone ZH was produced correctly, and the surrounding material is reported rather than being allowed to affect the score of the phone you asked about.

{
  "score": 1.0,
  "alignments": [
    {
      "start_time_ms": "680",
      "duration_ms": "360",
      "tokens": [
        {
          "kind": "ALIGNMENT_KIND_INSERTION",
          "hypotheses": [
            { "token": "T", "confidence": 1.0, "start_time_ms": "680", "duration_ms": "120" }
          ]
        },
        {
          "kind": "ALIGNMENT_KIND_INSERTION",
          "hypotheses": [
            { "token": "R", "confidence": 1.0, "start_time_ms": "840", "duration_ms": "120" }
          ]
        },
        {
          "kind": "ALIGNMENT_KIND_INSERTION",
          "hypotheses": [
            { "token": "EH", "confidence": 0.994, "start_time_ms": "960", "duration_ms": "80" }
          ]
        }
      ]
    },
    {
      "text": "ZH",
      "start_time_ms": "1000",
      "duration_ms": "240",
      "tokens": [
        {
          "kind": "ALIGNMENT_KIND_MATCH",
          "reference": "ZH",
          "score": 1.0,
          "hypotheses": [
            { "token": "ZH", "confidence": 1.0, "start_time_ms": "1000", "duration_ms": "240" }
          ]
        }
      ]
    }
  ]
}

If extraneous audio should count against the speaker in your task, inspect the insertion entries and apply your own rule; the information is there either way.

1.8 - Tuning and Customization

Pronunciations, per-phoneme thresholds and scoring behavior.

A general-purpose speech model is a fixed object: you send it audio and accept what it returns. CAPT is deliberately not that. Because the scoring stage is separate from the acoustic model, a great deal of behavior can be adjusted for your task without retraining anything, by changing which pronunciations count as correct, and how strictly each sound is judged.

Everything on this page lives in the model directory shipped with your release, and takes effect when the server restarts.

models/en_US-16khz/capt/
├── model.config.toml     scoring behavior and paths to the files below
├── lexicon.json          the primary pronunciation dictionary
├── lexicon_addenda.tsv   your additions and overrides
├── cutoffs.yaml          per-phoneme thresholds
└── g2p.ort               grapheme-to-phoneme model for unknown words

The model config points at each of these by path, so the names above are the convention rather than a requirement. Older model directories may carry the cutoffs as a plain-text cutoffs.txt instead; the meaning is identical and CutoffsPath says which file is in use.

Pronunciations

The lexicon and its addenda

The lexicon maps words to their phoneme sequences. A word may have several valid pronunciations, and CAPT considers all of them, reporting whichever best fits the audio, so a speaker is not penalized for a legitimate variant.

lexicon_addenda.tsv is loaded after the primary lexicon and is the file you should edit. Entries in it add new words and override existing ones, leaving the shipped lexicon intact so that upgrades stay clean.

This is the mechanism to reach for when your task defines what counts as an acceptable pronunciation. If an item specifies that a target may be produced two ways, add both and each will be accepted as a match:

MOSP	m A s p
MOSP	m oU s p

Those two entries accept mosp pronounced to rhyme with wasp, or with soap.

Out-of-vocabulary words and G2P

Words that are in neither the lexicon nor the addenda are passed to a grapheme-to-phoneme model, which predicts a pronunciation from spelling. This happens automatically and requires no configuration.

It means invented words, proper nouns and non-words can be used as reference_text with no setup at all, which is useful for decoding tasks built on pseudo-words, where the whole point is that the string is not a real word.

Tokenization

[Tokenization] in model.config.toml controls how reference text is pre-processed:

SettingWhat it doesDefaultIn en_US-16khz
RemovePunctuationStrips punctuation before lexicon lookup. With it off, punctuation stays attached to the word and is phonemized as part of it.offoff
UseSentenceTokenizerSplits a long reference into sentences before tokenizing.offon
MaxInputLengthSafeguard on the length of a reference text. Counted in characters, not words; a longer reference is rejected with an error. Applies to each alternative reference text too.20482048

“Default” is what applies when the key is absent from model.config.toml. Neither boolean has a server-side default beyond false, so what matters in practice is what your model ships with.

Cutoffs and acceptable confusions

This is the most useful tuning mechanism CAPT offers, and the least obvious.

Raw acoustic confidences are biased in ways that are specific to a model and a population of speakers. A model may routinely blur m and n; young children may produce th in a way that reads as s. Left alone, those biases show up as wrong verdicts. Cutoffs correct them by remapping confidence to score.

Each symbol has a cutoff, the confidence at which it is exactly borderline. The remapping is piecewise linear and pins that point to a score of 0.5:

  • raw confidence in [0, cutoff] maps to a score in [0, 0.5]
  • raw confidence in [cutoff, 1] maps to a score in [0.5, 1]

So the cutoff is the dial for how strict a given phoneme is. Lowering it is more lenient; raising it demands higher confidence before the phoneme counts as correct. Any symbol not listed uses DefaultCutoffScore from the [Scoring] section.

A modifier extends this with acceptable confusions. It names alternatives (other phonemes that may also count toward the match) and a multiplier that scales the cutoff when one of them is what was actually heard:

"A":
  cutoff: 0.50
  modifier:
    multiplier: 1.50
    alternatives: ["O"]

"T":
  cutoff: 0.50
  modifier:
    multiplier: 1.90
    alternatives: ["s"]

Read the first entry as: "A (the vowel of lot) needs 50% confidence. O (the vowel of thought) is close enough to count, but only if the model is much more sure of it: 0.50 × 1.50 = 0.75." The second is stricter still: s may stand in for T (th), but must clear 0.50 × 1.90 = 0.95, which is about right for a contrast that many young speakers have not yet acquired.

During alignment the confidences of the reference phoneme and any listed alternatives are summed, and the scaled cutoff is applied to that sum.

The effect is a per-sound tolerance you control. Rather than one global threshold that is too strict for some phonemes and too lax for others, you set the bar sound by sound, and you do it by editing a YAML file, not by collecting data and retraining.

Scoring behavior

The [Scoring] section of model.config.toml governs the alignment itself. The settings most likely to matter:

SettingWhat it does
DefaultCutoffScoreThe cutoff for any symbol not listed in cutoffs.yaml. Below 0.5 is lenient, above 0.5 is strict.
SubCost, InsCost, DelCostEdit-distance costs for substitution, insertion and deletion. Raising one makes the aligner prefer explanations that avoid it.
AlignVowelsPrevents nonsensical vowel↔consonant substitutions, reporting a deletion plus an insertion instead.
IncludeStressTreat stress markers as distinguishing (AH0 ≠ AH1). Requires both the lexicon and the acoustic model to carry stress.
IncludeIntraWordInsertionWhether extra phonemes heard inside a word are attached to that word or reported separately.
NonSpeechTokenConfidenceScalingWeights the confidence of silence and noise tokens. Below 1.0 makes the model less willing to call something non-speech, which is useful in noisy rooms.
CnetLinkThresholdDiscards time slices where the acoustic model put little probability on any speech token.
MaxCandidatePronsCaps how many combinations of word pronunciations are enumerated for a reference.

Non-speech tokens

NonSpeechTokens in [Vocab] lists the symbols the acoustic model emits for things that are not speech: silence, breath, vocal noise, unknown sounds. These are handled specially: they are excluded from scoring rather than being treated as mispronunciations, and they interact with CnetLinkThreshold and NonSpeechTokenConfidenceScaling above.

Recordings made in real rooms contain plenty of this material. If you find that noisy recordings score too harshly, these settings, rather than the cutoffs, are usually the right place to look.

Adapting the acoustic model

Configuration handles a great deal, but not everything. Where a population is genuinely different from what the model was trained on (young children, a regional accent, a specific recording setup), the acoustic model itself can be adapted to it, which addresses the cause rather than compensating for it downstream.

Adaptation needs representative audio, ideally paired with the verdicts you want the system to reproduce. If your application already records both, you may have a suitable dataset without any additional collection effort. Contact us to discuss what your data supports.

1.9 - Browser Integration

Calling CAPT from a web browser over HTTP and websockets.

Native applications should use gRPC directly, as shown throughout these docs. Browsers cannot open a bidirectional gRPC stream, so CAPT server also exposes an HTTP+JSON gateway and a websocket endpoint for StreamingEvaluate.

The HTTP API is not a property of the server itself: it is served only when api.Address is set. The configuration shipped with a release sets it to :8080, which is why the getting started page can curl it straight away. If you are writing a config from scratch, enable it explicitly in capt-server.cfg.toml:

[server.http]
api.Address = ":8080"

## Serve the built-in web demo alongside the API.
api.EnableWebDemo = true

## Only needed when the page is served from a different origin
## than the API, e.g. during local development.
# api.EnableCORS = true

JSON conventions

The gateway emits JSON using protobuf field names, so:

  • Field names are snake_case, exactly as in the proto: evaluation_result, start_time_ms, alternative_reference_text.
  • Enums are full strings, not numbers: "ALIGNMENT_KIND_MATCH", "AUDIO_FORMAT_HEADERED_WAV".
  • Unpopulated fields are emitted, so a field being present does not imply it was set.
  • 64-bit integers are strings, per the protobuf JSON mapping: "start_time_ms": "570", not 570. Parse these before doing arithmetic.
  • bytes fields are base64, so audio is base64-encoded into {"audio": {"data": "..."}}.

Unary calls over HTTP

The two metadata calls and Evaluate are available as ordinary HTTP requests:

MethodRoute
VersionGET /api/capt/v1/version
ListModelsGET /api/capt/v1/list-models
EvaluatePOST /api/capt/v1/evaluate
StreamingEvaluateGET /api/capt/v1/streaming-evaluate (websocket)
const models = await fetch('/api/capt/v1/list-models').then((r) => r.json());
const modelID = models.models[0].id;

Streaming over a websocket

StreamingEvaluate is served over a websocket that carries the same messages as the gRPC stream, one JSON object per websocket message.

The sequence mirrors the gRPC one: the first message must be the config, and every message after it carries audio.

const ws = new WebSocket(`wss://${location.host}/api/capt/v1/streaming-evaluate`);

ws.onopen = () => {
  // First message: the configuration.
  ws.send(
    JSON.stringify({
      config: {
        model_id: 'en_US-16khz',
        reference_text: 'WHEN THE SUNLIGHT STRIKES',
        audio_format: { audio_format_headered: 'AUDIO_FORMAT_HEADERED_WAV' }
      }
    })
  );
};

ws.onmessage = (event) => {
  const msg = JSON.parse(event.data);

  // The gateway reports a stream-level failure at the top level; a non-fatal
  // EvaluationError is a field of the response, so it sits inside the envelope.
  if (msg.error != null) {
    console.error('stream failed:', msg.error);
    return;
  }

  if (msg.result?.error != null) {
    console.warn('non-fatal:', msg.result.error.message);
  }

  const result = msg.result?.evaluation_result;
  if (result == null) {
    return;
  }

  if (result.is_partial) {
    showProvisional(result); // live highlighting
  } else {
    recordFinal(result); // the result you act on
  }
};

Sending audio means base64-encoding each chunk:

// Spreading a whole buffer into String.fromCharCode overflows the call stack
// on anything but small inputs, so encode in fixed-size blocks.
function toBase64(buf) {
  const bytes = new Uint8Array(buf);
  const block = 0x8000;
  let binary = '';

  for (let i = 0; i < bytes.length; i += block) {
    binary += String.fromCharCode.apply(null, bytes.subarray(i, i + block));
  }

  return btoa(binary);
}

async function sendChunk(blob) {
  ws.send(JSON.stringify({ audio: { data: toBase64(await blob.arrayBuffer()) } }));
}

// A websocket has no half-close, so an empty audio message is how you say
// "that is all the audio". Send it only once the recording is complete.
function endStream() {
  ws.send(JSON.stringify({ audio: { data: '' } }));
}

The same helper is what you want for the whole-recording POST described below; a complete take is far past the size at which the one-line spread form fails.

Capturing audio in the browser

The audio you send must match the audio_format you declared. The MediaRecorder API defaults to a compressed format that CAPT does not accept, so a browser client generally needs an encoder that can emit WAV chunks while recording, rather than only at the end.

Two constraints are worth designing around:

  • Send small, regular chunks. Around 100-250 ms keeps partial results flowing smoothly. Larger chunks make feedback feel laggy.
  • Declare the header once. For headered formats the header belongs at the start of the stream, not on every chunk.

If you do not need live feedback, it is considerably simpler to record the whole utterance, then POST it to /api/capt/v1/evaluate as a single base64 blob.

The built-in demo

With api.EnableWebDemo = true, the server serves a demo application at the HTTP address, by default http://localhost:8080. It records from the microphone or accepts an uploaded file, streams it over the websocket described above, and renders the per-word and per-phoneme results.

It is a useful way to sanity-check a server, a model and a piece of audio before writing any client code.

1.10 - API Reference

Detailed reference for API requests and types.

The API is defined as a protobuf spec, so native bindings can be generated in any language with gRPC support. We recommend using buf to generate the bindings.

This section of the documentation is auto-generated from the protobuf spec. The service contains the methods that can be called, and the “messages” are the data structures (objects, classes or structs in the generated code, depending on the language) passed to and from the methods.

CAPTService

Service that implements the Cobalt CAPT API.

Version

Version(VersionRequest) VersionResponse

Returns version information from the server.

ListModels

ListModels(ListModelsRequest) ListModelsResponse

Returns information about the models available on the server.

Evaluate

Evaluate(EvaluateRequest) EvaluateResponse

Performs synchronous speech evaluation by receiving results after all audio has been sent and processed. It is expected that this request be typically used for short audio content: less than a minute long. For longer content, the StreamingEvaluate method should be preferred.

StreamingEvaluate

StreamingEvaluate(StreamingEvaluateRequest) StreamingEvaluateResponse

Performs bidirectional streaming for evaluating speech by receiving results while sending audio. This method is only available via GRPC and not via HTTP+JSON. However, a web browser may use websockets to use this service.

Messages

  • If two or more fields in a message are labeled oneof, then each method call using that message must have exactly one of the fields populated
  • If a field is labeled repeated, then the generated code will accept an array (or struct, or list depending on the language).

AlignedToken

AlignedToken contains a single reference token and the alternate hypotheses recognized from the audio that could be aligned with the reference token.

If kind == ALIGNMENT_KIND_MATCH or ALIGNMENT_KIND_SUBSTITUTION, both the reference token and hypotheses list will be populated. The score will be based on the confidence value of the reference token within the hypotheses. In the case of a substitution, the hypotheses list may not contain the actual reference token at all, and the score will be 0.

If kind == ALIGNMENT_KIND_DELETION, the hypotheses will be an empty list and score will be 0.

If kind == ALIGNMENT_KIND_INSERTION, the reference token will be an empty string and score will be 0.

Fields

  • kind (AlignmentKind ) The kind of alignment for this token.

  • reference (string ) The actual reference token.

  • score (float ) The score for this token, between 0 and 1 inclusive, based on the confidence value of the token within the recognized token hypotheses from the audio.

  • hypotheses (AlignmentHypothesis repeated) The recognized tokens from the audio that have been aligned to this reference token.

AlignedWord

AlignedWord contains the aligned tokens within a single word in the reference text.

Fields

  • text (string ) The actual word.

  • start_time_ms (uint64 ) The timestamp at which this token starts in the audio, in milliseconds.

  • duration_ms (uint64 ) The duration of this token in the audio, in milliseconds.

  • tokens (AlignedToken repeated) The tokens that make up the word.

AlignmentHypothesis

AlignmentHypothesis contains the metadata for a token recognized from audio, including its timestamps and confidence scores.

Fields

  • token (string ) The actual token.

  • confidence (float ) The confidence with which this token was recognized in the audio, between 0 and 1 inclusive.

  • start_time_ms (uint64 ) The timestamp at which this token starts in the audio, in milliseconds.

  • duration_ms (uint64 ) The duration of this token in the audio, in milliseconds.

AlternativeAlignment

AlternativeAlignment contains alignments against alternative reference text(s) if specified in the EvaluationConfig.

Fields

  • reference_text (string ) The alternative reference text.

  • score (float ) Evaluation score, between 0 and 1 inclusive, with 1 indicating a perfect match between the alternative reference text and what was said in the audio.

  • alignments (AlignedWord repeated) Alignment between the alternative reference text tokens and recognized tokens.

Audio

Audio to be sent to Capt.

Fields

AudioFormat

Format of the audio to be sent for recognition.

Depending on how they are configured, server instances of this service may not support all the formats provided in the API. One format that is guaranteed to be supported is the RAW format with little-endian 16-bit signed samples with the sample rate matching that of the model being requested.

Fields

  • oneof audio_format.audio_format_raw (AudioFormatRAW ) Audio is raw data without any headers

  • oneof audio_format.audio_format_headered (AudioFormatHeadered ) Audio has a self-describing header. Headers are expected to be sent at the beginning of the entire audio file/stream, and not in every Audio message.

    The default value of this type is AUDIO_FORMAT_HEADERED_UNSPECIFIED. If this value is used, the server may attempt to detect the format of the audio. However, it is recommended that the exact format be specified.

AudioFormatRAW

Details of audio in raw format

Fields

  • encoding (AudioEncoding ) Encoding of the samples. It must be specified explicitly and using the default value of AUDIO_ENCODING_UNSPECIFIED will result in an error.

  • bit_depth (uint32 ) Bit depth of each sample (e.g. 8, 16, 24, 32, etc.). This is a required field.

  • byte_order (ByteOrder ) Byte order of the samples. This field must be set to a value other than BYTE_ORDER_UNSPECIFIED when the bit_depth is greater than 8.

  • sample_rate (uint32 ) Sampling rate in Hz. This is a required field.

  • channels (uint32 ) Number of channels present in the audio. E.g.: 1 (mono), 2 (stereo), etc. This is a required field.

EvaluateRequest

The top-level message sent by the client for the Evaluate method. Both the EvaluationConfig and Audio fields are required. The entire audio data must be sent in one request. If your audio data is larger, please use the StreamingEvaluate call.

Fields

EvaluateResponse

The message returned by the server for the Evaluate method.

Fields

  • evaluation_result (EvaluationResult ) Response from the server. The kind of result depends on the kind of the model that was chosen in the EvaluationConfig.

  • error (EvaluationError ) A non-fatal error message. If a server encountered a non-fatal error when processing the request, it will be returned in this message. The server will continue to process audio and produce further results. Clients can continue streaming audio even after receiving these messages. This error message is meant to be informational.

    An example of when these errors maybe produced: audio is sampled at a lower rate than expected by model, producing possibly less accurate results.

    This field will be unset if there is no error to report.

EvaluationConfig

Configuration for a StreamingEvaluateRequest.

Fields

  • model_id (string ) ID of the model to use. A list of supported IDs can be found using the ListModels call.

  • audio_format (AudioFormat ) Format of the audio to be sent.

  • reference_text (string ) Reference Text that is expected to be contained in the audio.

  • alternative_reference_text (string repeated) Alternative reference text(s) that are also acceptable if recognized in the audio.

  • metadata (EvaluationMetadata ) This is an optional field. If there is any metadata associated with the audio being sent, use this field to provide it to the recognizer. The server may record this metadata when processing the request. The server does not use this field for any other purpose.

EvaluationError

Developer-facing error message about a non-fatal process issue.

Fields

EvaluationMetadata

Metadata associated with the evaluation request

Fields

  • custom_metadata (string ) Any custom metadata that the client wants to associate with the recording. This could be a simple string (e.g. a tracing ID) or structured data (e.g. JSON).

  • custom_id (string ) This is an optional field to specify custom ID to identify the evaluation request. The custom ID must be a string of upto 64 bytes, and only alphabets, digits, hyphens and underscores are allowed. This ID may be recorded by the server in logs or other storage, and should therefore not include any sensitive information.

EvaluationResult

EvaluationResult contains the result generated by a speech evaluation model.

Fields

  • is_partial (bool ) If this is set to true, it denotes that the result is an interim partial result, and could change after more audio is processed. If unset, or set to false, it denotes that this is a final result and will not change.

    Servers are not required to implement support for returning partial results, and clients should generally not depend on their availability.

  • score (float ) Overall evaluation score, between 0 and 1 inclusive, with 1 indicating a perfect match between the reference text and what was said in the audio. If alternative reference text(s) are specified, then the score will be take those into account and be based on the reference text that aligns the best with recognized tokens.

  • alignments (AlignedWord repeated) Alignment between the primary expected reference text tokens and recognized tokens.

  • alternative_alignments (AlternativeAlignment repeated) Alignments against alternative reference text tokens (if specified) and recognized tokens.

  • is_speech_endpoint (bool ) If true, indicates that a speech endpoint has been detected.

    A speech endpoint signifies that the server believes the user has finished their utterance, typically after detecting a specific duration of silence.

    NOTE: Server support for endpoint detection is optional. Clients must be robust to this field never being set.

ListModelsRequest

The top-level message sent by the client for the ListModels method.

ListModelsResponse

The message returned to the client by the ListModels method.

Fields

  • models (Model repeated) List of models available for use that match the request.

Model

Description of a CAPT model.

Fields

  • id (string ) Unique identifier of the model. This identifier is used to choose the model the model when configuring a StreamingEvaluate request.

  • name (string ) Model name. This is a concise name describing the model, and may be presented to the end-user, for example, to help choose which model to use for their task.

  • kind (ModelKind ) The specific kind of model. This determines what type of results it sends back, and additional config requirements if any.

  • attributes (ModelAttributes ) Model Attributes.

ModelAttributes

Attributes of a Capt model.

Fields

  • sample_rate (uint32 ) Audio sample rate (native) supported by the model.

  • metadata (ModelMetadata ) Metadata associated with the model.

ModelMetadata

Metadata associated with a Capt model.

Fields

  • version (string ) Model version in semver format (e.g. 1.0.0). If the version is not known, it will default to “unknown”.

  • build_date (string ) Date the model was built in YYYY-MM-DD format. If the version is not known, it will default to “unknown”.

StreamingEvaluateRequest

The top level messages sent by the client for the StreamingEvaluate method. In this streaming call, multiple StreamingEvaluateRequest messages should be sent. The first message must contain a EvaluationConfig message, and all subsequent messages must contain Audio only. All Audio messages must contain non-empty audio. If audio content is empty, the server may choose to interpret it as end of stream and stop accepting any further messages.

Fields

StreamingEvaluateResponse

The message returned by the server for the StreamingEvaluate method.

Fields

  • evaluation_result (EvaluationResult ) Response from the server. The kind of result depends on the kind of the model that was chosen in the EvaluationConfig.

  • error (EvaluationError ) A non-fatal error message. If a server encountered a non-fatal error when processing the request, it will be returned in this message. The server will continue to process audio and produce further results. Clients can continue streaming audio even after receiving these messages. This error message is meant to be informational.

    An example of when these errors maybe produced: audio is sampled at a lower rate than expected by model, producing possibly less accurate results.

    This field will be unset if there is no error to report.

VersionRequest

The top-level message sent by the client for the Version method.

VersionResponse

The message sent by the server for the Version method.

Fields

  • version (string ) Version of the server handling these requests.

Enums

AlignmentKind

AlignmentKind represents one of four alignment outcomes possible: a Match, Substitution, Deletion or Insertion.

NameNumberDescription
ALIGNMENT_KIND_UNSPECIFIED0Default value of this type.
ALIGNMENT_KIND_MATCH1A match implies that the reference and recognized token from audio are a match with a reasonable amount of confidence.
ALIGNMENT_KIND_SUBSTITUTION2A substitution implies that the reference token has been replaced by a different token in the recognized tokens from the audio.
ALIGNMENT_KIND_DELETION3A deletion implies that the reference token was not found in the recognized tokens from audio.
ALIGNMENT_KIND_INSERTION4A insertion implies that a extraneous token has been recognized in the audio, that cannot be matched to any token in the reference.

AudioEncoding

The encoding of the audio data to be sent for recognition.

NameNumberDescription
AUDIO_ENCODING_UNSPECIFIED0AUDIO_ENCODING_UNSPECIFIED is the default value of this type and will result in an error.
AUDIO_ENCODING_SIGNED1PCM signed-integer
AUDIO_ENCODING_UNSIGNED2PCM unsigned-integer
AUDIO_ENCODING_IEEE_FLOAT3PCM IEEE-Float
AUDIO_ENCODING_ULAW4G.711 mu-law
AUDIO_ENCODING_ALAW5G.711 a-law

AudioFormatHeadered

NameNumberDescription
AUDIO_FORMAT_HEADERED_UNSPECIFIED0AUDIO_FORMAT_HEADERED_UNSPECIFIED is the default value of this type.
AUDIO_FORMAT_HEADERED_WAV1WAV with RIFF headers
AUDIO_FORMAT_HEADERED_MP32MP3 format with a valid frame header at the beginning of data
AUDIO_FORMAT_HEADERED_FLAC3FLAC format
AUDIO_FORMAT_HEADERED_OGG_OPUS4Opus format with OGG header

ByteOrder

Byte order of multi-byte data

NameNumberDescription
BYTE_ORDER_UNSPECIFIED0BYTE_ORDER_UNSPECIFIED is the default value of this type.
BYTE_ORDER_LITTLE_ENDIAN1Little Endian byte order
BYTE_ORDER_BIG_ENDIAN2Big Endian byte order

ModelKind

NameNumberDescription
MODEL_KIND_UNSPECIFIED0Default value of this type.
MODEL_KIND_SPEECH_EVALUATION1Model for evaluating the accuracy of spoken text from audio. This is done by first recognizing what’s said in the audio, and then aligning expected and recognized tokens (phonemes, syllables, etc.). This model returns results in the form of EvaluationResult messages.
MODEL_KIND_PHONEME_EVALUATION2Model for evaluating the accuracy of phonemes pronounced in isolation from audio. This is done in a way similar to speech evaluation models, but is more accurate for single phonemes in isolation, which a regular speech model may not recognize correctly. This model returns results in the form of EvaluationResult messages.

Scalar Value Types

.proto TypeC++ TypeC# TypeGo TypeJava TypePHP TypePython TypeRuby Type

double
doubledoublefloat64doublefloatfloatFloat

float
floatfloatfloat32floatfloatfloatFloat

int32
int32intint32intintegerintBignum or Fixnum (as required)

int64
int64longint64longinteger/stringint/longBignum

uint32
uint32uintuint32intintegerint/longBignum or Fixnum (as required)

uint64
uint64ulonguint64longinteger/stringint/longBignum or Fixnum (as required)

sint32
int32intint32intintegerintBignum or Fixnum (as required)

sint64
int64longint64longinteger/stringint/longBignum

fixed32
uint32uintuint32intintegerintBignum or Fixnum (as required)

fixed64
uint64ulonguint64longinteger/stringint/longBignum

sfixed32
int32intint32intintegerintBignum or Fixnum (as required)

sfixed64
int64longint64longinteger/stringint/longBignum

bool
boolboolboolbooleanbooleanbooleanTrueClass/FalseClass

string
stringstringstringStringstringstr/unicodeString (UTF-8)

bytes
stringByteString[]byteByteStringstringstrString (ASCII-8BIT)

1.11 - FAQ

Common questions about deploying and using CAPT.

How is this different from using a general-purpose speech recognizer?

A recognizer is built to recover the intended message. That makes it actively unsuitable for assessment, because it will repair what it hears into what it assumes you meant: a mispronounced word is transcribed as the word, and the error you were trying to measure disappears. The better the language model, the more thoroughly it hides exactly the signal you need.

CAPT is anchored to a reference you supply and never substitutes expectation for observation. It also gives you things a transcript cannot:

  • Per-phoneme verdicts and scores, not just a word-level transcript.
  • Confidence taken from a full distribution over phonemes, not a single guess.
  • Timestamps for each recognized phoneme.
  • Deterministic, inspectable scoring you can tune, with no risk of a generative model inventing plausible text.
  • Operation entirely on your own hardware, on CPU.

What accuracy should I expect?

It depends on the task, the audio and the speakers, so we do not publish a single figure that would mislead you. What we would recommend instead is measuring it on your own data, against the decisions you actually care about.

If you are automating a judgement a person currently makes, the number that matters is agreement with that person on your items, not word error rate, and not any generic benchmark. See calibrating against human judgement.

Does audio leave my infrastructure?

No. CAPT runs on your own hardware: on-premise, in your private cloud, or embedded on a device. There is no call home, and no dependency on an external service at request time.

Audio and results are not stored unless you explicitly enable it. Storage is off by default; to turn it on, set a storage backend and path in capt-server.cfg.toml:

[storage]
Type = "localfs"
BasePath = "/audio"

With this enabled, each session is written to local disk as two files, the audio and the evaluation result, organized by UTC date. Everything stays on the machine you run.

Note that metadata.custom_id and metadata.custom_metadata may be written to logs and to stored results, so they should carry opaque identifiers rather than personal information.

What hardware does it need?

CAPT is CPU-only; no GPU is required. It runs on x86_64 and Arm64 / aarch64, and a statically linked build is available for minimal or embedded images.

Sizing depends on the model and on how many concurrent evaluations you need. Contact us for guidance against your target hardware.

What languages are supported?

Each model covers one language, and the model’s language is fixed at build time. US English models are available today, and models for other languages can be built. Contact us to discuss a specific language.

What audio should I send?

Use uncompressed or losslessly compressed audio (WAV or FLAC) recorded at the model’s native sample rate, which ListModels reports as attributes.sample_rate.

Audio sampled below what the model expects yields a non-fatal warning and reduced accuracy. Upsampling before sending does not help; the detail is already gone. Lossy formats such as MP3 are accepted but will cost you some accuracy.

Can more than one person be speaking in the recording?

CAPT scores the audio against the reference text; it does not separate speakers. If someone other than the intended speaker is audible, an assistant reading a prompt for instance, that speech is part of the audio being evaluated and can influence the result, usually appearing as insertions.

Where recordings may contain more than one voice, control it at capture time: record only while the intended speaker is expected to be talking, rather than across the whole interaction. If that is not possible for your setup, talk to us about the options.

Can I add my own words and pronunciations?

Yes, and this is a normal thing to do. Add them to lexicon_addenda.tsv in the model directory. You can add words the lexicon does not have, and override the pronunciations of words it does. Words in neither the lexicon nor the addenda are handled automatically by a grapheme-to-phoneme model, so invented words and non-words work without any setup.

See Tuning and Customization.

Can I make scoring stricter or more lenient?

Yes, globally or per phoneme, by editing configuration rather than retraining. DefaultCutoffScore moves the bar for everything; cutoffs.yaml sets it sound by sound, and can additionally name specific confusions that should be tolerated. See Cutoffs and acceptable confusions.

Do I have to use streaming?

No. Evaluate takes the whole audio in one request and returns one result, which is the simplest option for scoring a file. Use StreamingEvaluate when you want results while the speaker is still talking, or for audio longer than about a minute.

Are partial results guaranteed?

No. Servers are not required to produce them, and clients should not depend on them. Drive your logic from the final result, the one with is_partial unset or false, and treat partials as an enhancement for live feedback.

What does a score of 0.5 mean?

It means “exactly borderline” for that phoneme. Scores are not raw probabilities: each phoneme has a cutoff, and the confidence-to-score mapping is built so that the cutoff lands on 0.5. That is what makes 0.5 a meaningful threshold and lets different phonemes be held to different standards.

Overall utterance scores are a summary of many such phoneme scores, so an overall 0.5 does not mean “half correct”. Look at the per-token verdicts before thresholding. See Interpreting Results.