Browser Integration

Calling CAPT from a web browser over HTTP and websockets.

Native applications should use gRPC directly, as shown throughout these docs. Browsers cannot open a bidirectional gRPC stream, so CAPT server also exposes an HTTP+JSON gateway and a websocket endpoint for StreamingEvaluate.

The HTTP API is not a property of the server itself: it is served only when api.Address is set. The configuration shipped with a release sets it to :8080, which is why the getting started page can curl it straight away. If you are writing a config from scratch, enable it explicitly in capt-server.cfg.toml:

[server.http]
api.Address = ":8080"

## Serve the built-in web demo alongside the API.
api.EnableWebDemo = true

## Only needed when the page is served from a different origin
## than the API, e.g. during local development.
# api.EnableCORS = true

JSON conventions

The gateway emits JSON using protobuf field names, so:

  • Field names are snake_case, exactly as in the proto: evaluation_result, start_time_ms, alternative_reference_text.
  • Enums are full strings, not numbers: "ALIGNMENT_KIND_MATCH", "AUDIO_FORMAT_HEADERED_WAV".
  • Unpopulated fields are emitted, so a field being present does not imply it was set.
  • 64-bit integers are strings, per the protobuf JSON mapping: "start_time_ms": "570", not 570. Parse these before doing arithmetic.
  • bytes fields are base64, so audio is base64-encoded into {"audio": {"data": "..."}}.

Unary calls over HTTP

The two metadata calls and Evaluate are available as ordinary HTTP requests:

MethodRoute
VersionGET /api/capt/v1/version
ListModelsGET /api/capt/v1/list-models
EvaluatePOST /api/capt/v1/evaluate
StreamingEvaluateGET /api/capt/v1/streaming-evaluate (websocket)
const models = await fetch('/api/capt/v1/list-models').then((r) => r.json());
const modelID = models.models[0].id;

Streaming over a websocket

StreamingEvaluate is served over a websocket that carries the same messages as the gRPC stream, one JSON object per websocket message.

The sequence mirrors the gRPC one: the first message must be the config, and every message after it carries audio.

const ws = new WebSocket(`wss://${location.host}/api/capt/v1/streaming-evaluate`);

ws.onopen = () => {
  // First message: the configuration.
  ws.send(
    JSON.stringify({
      config: {
        model_id: 'en_US-16khz',
        reference_text: 'WHEN THE SUNLIGHT STRIKES',
        audio_format: { audio_format_headered: 'AUDIO_FORMAT_HEADERED_WAV' }
      }
    })
  );
};

ws.onmessage = (event) => {
  const msg = JSON.parse(event.data);

  // The gateway reports a stream-level failure at the top level; a non-fatal
  // EvaluationError is a field of the response, so it sits inside the envelope.
  if (msg.error != null) {
    console.error('stream failed:', msg.error);
    return;
  }

  if (msg.result?.error != null) {
    console.warn('non-fatal:', msg.result.error.message);
  }

  const result = msg.result?.evaluation_result;
  if (result == null) {
    return;
  }

  if (result.is_partial) {
    showProvisional(result); // live highlighting
  } else {
    recordFinal(result); // the result you act on
  }
};

Sending audio means base64-encoding each chunk:

// Spreading a whole buffer into String.fromCharCode overflows the call stack
// on anything but small inputs, so encode in fixed-size blocks.
function toBase64(buf) {
  const bytes = new Uint8Array(buf);
  const block = 0x8000;
  let binary = '';

  for (let i = 0; i < bytes.length; i += block) {
    binary += String.fromCharCode.apply(null, bytes.subarray(i, i + block));
  }

  return btoa(binary);
}

async function sendChunk(blob) {
  ws.send(JSON.stringify({ audio: { data: toBase64(await blob.arrayBuffer()) } }));
}

// A websocket has no half-close, so an empty audio message is how you say
// "that is all the audio". Send it only once the recording is complete.
function endStream() {
  ws.send(JSON.stringify({ audio: { data: '' } }));
}

The same helper is what you want for the whole-recording POST described below; a complete take is far past the size at which the one-line spread form fails.

Capturing audio in the browser

The audio you send must match the audio_format you declared. The MediaRecorder API defaults to a compressed format that CAPT does not accept, so a browser client generally needs an encoder that can emit WAV chunks while recording, rather than only at the end.

Two constraints are worth designing around:

  • Send small, regular chunks. Around 100-250 ms keeps partial results flowing smoothly. Larger chunks make feedback feel laggy.
  • Declare the header once. For headered formats the header belongs at the start of the stream, not on every chunk.

If you do not need live feedback, it is considerably simpler to record the whole utterance, then POST it to /api/capt/v1/evaluate as a single base64 blob.

The built-in demo

With api.EnableWebDemo = true, the server serves a demo application at the HTTP address, by default http://localhost:8080. It records from the microphone or accepts an uploaded file, streams it over the websocket described above, and renders the per-word and per-phoneme results.

It is a useful way to sanity-check a server, a model and a piece of audio before writing any client code.