Native applications should use gRPC directly, as shown throughout these docs.
Browsers cannot open a bidirectional gRPC stream, so CAPT server also exposes
an HTTP+JSON gateway and a websocket endpoint for StreamingEvaluate.
The HTTP API is not a property of the server itself: it is served only when
api.Address is set. The configuration shipped with a release sets it to
:8080, which is why the getting started page can curl it
straight away. If you are writing a config from scratch, enable it explicitly
in capt-server.cfg.toml:
[server.http]
api.Address = ":8080"
## Serve the built-in web demo alongside the API.
api.EnableWebDemo = true
## Only needed when the page is served from a different origin
## than the API, e.g. during local development.
# api.EnableCORS = true
JSON conventions
The gateway emits JSON using protobuf field names, so:
- Field names are
snake_case, exactly as in the proto:evaluation_result,start_time_ms,alternative_reference_text. - Enums are full strings, not numbers:
"ALIGNMENT_KIND_MATCH","AUDIO_FORMAT_HEADERED_WAV". - Unpopulated fields are emitted, so a field being present does not imply it was set.
- 64-bit integers are strings, per the protobuf JSON mapping:
"start_time_ms": "570", not570. Parse these before doing arithmetic. bytesfields are base64, so audio is base64-encoded into{"audio": {"data": "..."}}.
Unary calls over HTTP
The two metadata calls and Evaluate are available as ordinary HTTP requests:
| Method | Route |
|---|---|
Version | GET /api/capt/v1/version |
ListModels | GET /api/capt/v1/list-models |
Evaluate | POST /api/capt/v1/evaluate |
StreamingEvaluate | GET /api/capt/v1/streaming-evaluate (websocket) |
const models = await fetch('/api/capt/v1/list-models').then((r) => r.json());
const modelID = models.models[0].id;
Streaming over a websocket
StreamingEvaluate is served over a websocket that carries the same messages
as the gRPC stream, one JSON object per websocket message.
The sequence mirrors the gRPC one: the first message must be the config, and every message after it carries audio.
const ws = new WebSocket(`wss://${location.host}/api/capt/v1/streaming-evaluate`);
ws.onopen = () => {
// First message: the configuration.
ws.send(
JSON.stringify({
config: {
model_id: 'en_US-16khz',
reference_text: 'WHEN THE SUNLIGHT STRIKES',
audio_format: { audio_format_headered: 'AUDIO_FORMAT_HEADERED_WAV' }
}
})
);
};
ws.onmessage = (event) => {
const msg = JSON.parse(event.data);
// The gateway reports a stream-level failure at the top level; a non-fatal
// EvaluationError is a field of the response, so it sits inside the envelope.
if (msg.error != null) {
console.error('stream failed:', msg.error);
return;
}
if (msg.result?.error != null) {
console.warn('non-fatal:', msg.result.error.message);
}
const result = msg.result?.evaluation_result;
if (result == null) {
return;
}
if (result.is_partial) {
showProvisional(result); // live highlighting
} else {
recordFinal(result); // the result you act on
}
};
Responses over the websocket are wrapped in a result envelope,
{"result": {"evaluation_result": {...}, "error": {...}}}, which the plain
gRPC stream does not have. Read msg.result.evaluation_result, not
msg.evaluation_result.
The envelope holds two different errors, and they are not interchangeable. A
non-fatal EvaluationError is a field of
the response, so it arrives at msg.result.error and processing continues.
A top-level msg.error is the gateway’s stream-level failure, with a
different shape, and the stream is over.
Sending audio means base64-encoding each chunk:
// Spreading a whole buffer into String.fromCharCode overflows the call stack
// on anything but small inputs, so encode in fixed-size blocks.
function toBase64(buf) {
const bytes = new Uint8Array(buf);
const block = 0x8000;
let binary = '';
for (let i = 0; i < bytes.length; i += block) {
binary += String.fromCharCode.apply(null, bytes.subarray(i, i + block));
}
return btoa(binary);
}
async function sendChunk(blob) {
ws.send(JSON.stringify({ audio: { data: toBase64(await blob.arrayBuffer()) } }));
}
// A websocket has no half-close, so an empty audio message is how you say
// "that is all the audio". Send it only once the recording is complete.
function endStream() {
ws.send(JSON.stringify({ audio: { data: '' } }));
}
The same helper is what you want for the whole-recording POST described below; a complete take is far past the size at which the one-line spread form fails.
Capturing audio in the browser
The audio you send must match the audio_format you declared. The
MediaRecorder API defaults to a compressed format that CAPT does not accept,
so a browser client generally needs an encoder that can emit WAV chunks while
recording, rather than only at the end.
Two constraints are worth designing around:
- Send small, regular chunks. Around 100-250 ms keeps partial results flowing smoothly. Larger chunks make feedback feel laggy.
- Declare the header once. For headered formats the header belongs at the start of the stream, not on every chunk.
If you do not need live feedback, it is considerably simpler to record the
whole utterance, then POST it to /api/capt/v1/evaluate as a single base64
blob.
The built-in demo
With api.EnableWebDemo = true, the server serves a demo application at the
HTTP address, by default http://localhost:8080. It records from the
microphone or accepts an uploaded file, streams it over the websocket described
above, and renders the per-word and per-phoneme results.
It is a useful way to sanity-check a server, a model and a piece of audio before writing any client code.