Streaming Evaluation

Describes how to stream audio to CAPT server and receive evaluation results.

CAPT offers two ways to submit audio.

  • StreamingEvaluate is the primary method. It is bidirectional: you send audio as it is captured and receive results as the audio is processed. Use it for anything interactive, and for audio of any length.
  • Evaluate is a unary convenience method that takes the whole audio in a single request and returns a single result. It is intended for short audio, less than a minute, and is the simplest way to score an audio file.

Both accept the same EvaluationConfig and return the same EvaluationResult.

Streaming from an audio file

In a StreamingEvaluate stream, the first message must contain the config, and every subsequent message must contain audio.

An empty Audio message ends the stream. CAPT server treats it as the end of the audio: it finishes processing, sends the final result, and accepts nothing further, so a later send fails. Over gRPC you normally do not need it, because half-closing the stream (CloseSend in Go, exhausting the request iterator in Python) says the same thing. It matters for browser clients, where a websocket has no equivalent half-close.

Do not send one to mean “no audio yet”. If it arrives before the audio is complete, what you get depends on the format, and a client’s state machine needs to handle both:

  • You always receive a final result for the audio sent so far, with is_partial false. It reflects the truncated input, not the utterance you intended, so it will usually be full of deletions.
  • With a headered format the stream then fails. The header promised more data than arrived, so the server reports a container error such as wav: unexpected EOF after the final result. This happens whether or not you send anything afterwards.
  • With raw audio the stream ends cleanly, because there is no container to leave incomplete. The result is still only of the audio you sent.

In other words: a final result is not by itself evidence that the whole utterance was processed. Send the empty message only once the recording is complete, and treat an error after a final result as “the result is incomplete”, not “the result is invalid”.

The examples below stream a WAV file in 100 ms chunks and print each result as it arrives.

import grpc
import cobaltspeech.capt.v1.capt_pb2 as capt
import cobaltspeech.capt.v1.capt_pb2_grpc as capt_grpc

serverAddress = "localhost:2727"

channel = grpc.insecure_channel(serverAddress)
client = capt_grpc.CAPTServiceStub(channel)

# Get the list of models on the server and use the first one.
modelResp = client.ListModels(capt.ListModelsRequest())
modelID = modelResp.models[0].id

cfg = capt.EvaluationConfig(
    model_id=modelID,
    reference_text="WHEN THE SUNLIGHT STRIKES",
    audio_format=capt.AudioFormat(
        audio_format_headered=capt.AUDIO_FORMAT_HEADERED_WAV,
    ),
)

# The first request must contain only the configuration; subsequent
# requests carry audio bytes. A generator is a convenient way to do this.
def stream(cfg, audio, bufferSize=3200):
    yield capt.StreamingEvaluateRequest(config=cfg)

    data = audio.read(bufferSize)
    while len(data) > 0:
        yield capt.StreamingEvaluateRequest(audio=capt.Audio(data=data))
        data = audio.read(bufferSize)

with open("test.wav", "rb") as audio:
    for resp in client.StreamingEvaluate(stream(cfg, audio)):
        if resp.HasField("error"):
            print(f"warning: {resp.error.message}")

        result = resp.evaluation_result
        print(f"partial={result.is_partial} score={result.score:.4f}")

        # Only act on final results; partials will still change.
        if not result.is_partial:
            for word in result.alignments:
                print(f"  {word.text}: " + " ".join(
                    f"{t.reference}={t.score:.2f}" for t in word.tokens
                ))
package main

import (
	"context"
	"fmt"
	"io"
	"os"

	"google.golang.org/grpc"
	"google.golang.org/grpc/credentials/insecure"

	captpb "github.com/your-org/your-module/gen/go/cobaltspeech/capt/v1"
)

func main() {
	const serverAddress = "localhost:2727"

	audio, err := os.ReadFile("test.wav")
	if err != nil {
		fmt.Printf("failed to read audio: %v\n", err)
		os.Exit(1)
	}

	conn, err := grpc.NewClient(serverAddress,
		grpc.WithTransportCredentials(insecure.NewCredentials()))
	if err != nil {
		fmt.Printf("failed to dial gRPC connection: %v\n", err)
		os.Exit(1)
	}
	defer conn.Close()

	client := captpb.NewCAPTServiceClient(conn)

	stream, err := client.StreamingEvaluate(context.Background())
	if err != nil {
		fmt.Printf("failed to open stream: %v\n", err)
		os.Exit(1)
	}

	// The first message must contain the config.
	if err := stream.Send(&captpb.StreamingEvaluateRequest{
		Request: &captpb.StreamingEvaluateRequest_Config{
			Config: &captpb.EvaluationConfig{
				ModelId:       "en_US-16khz",
				ReferenceText: "WHEN THE SUNLIGHT STRIKES",
				AudioFormat: &captpb.AudioFormat{
					AudioFormat: &captpb.AudioFormat_AudioFormatHeadered{
						AudioFormatHeadered: captpb.AudioFormatHeadered_AUDIO_FORMAT_HEADERED_WAV,
					},
				},
			},
		},
	}); err != nil {
		fmt.Printf("failed to send config: %v\n", err)
		os.Exit(1)
	}

	// Send audio in the background while results are read below.
	go func() {
		const chunk = 3200 // 100ms of 16kHz 16-bit mono audio.

		// min() is a builtin from Go 1.21; on older toolchains compute the
		// bound explicitly.
		for i := 0; i < len(audio); i += chunk {
			end := min(i+chunk, len(audio))

			if err := stream.Send(&captpb.StreamingEvaluateRequest{
				Request: &captpb.StreamingEvaluateRequest_Audio{
					Audio: &captpb.Audio{Data: audio[i:end]},
				},
			}); err != nil {
				return
			}
		}

		_ = stream.CloseSend()
	}()

	for {
		resp, err := stream.Recv()
		if err == io.EOF {
			break
		}

		if err != nil {
			fmt.Printf("failed to receive result: %v\n", err)
			os.Exit(1)
		}

		if e := resp.GetError(); e != nil {
			fmt.Printf("warning: %v\n", e.GetMessage())
		}

		result := resp.GetEvaluationResult()
		fmt.Printf("partial=%v score=%.4f\n", result.GetIsPartial(), result.GetScore())

		// Only act on final results; partials will still change.
		if !result.GetIsPartial() {
			for _, word := range result.GetAlignments() {
				fmt.Printf("  %s:", word.GetText())

				for _, t := range word.GetTokens() {
					fmt.Printf(" %s=%.2f", t.GetReference(), t.GetScore())
				}

				fmt.Println()
			}
		}
	}
}

Streaming a 1.5 second recording of “when the sunlight strikes” against that reference produces a short sequence of results:

partial=true  score=0.5514
partial=true  score=0.6128
partial=true  score=0.6128
partial=false score=0.6128

Partial results

While audio is still arriving, the server emits partial results, marked is_partial = true. A partial reflects the audio received so far, and any part of it may change as more audio arrives. In the sequence above the score rises as the last word is heard.

When the stream ends, the server emits a final result with is_partial = false. That result is stable.

  • Use partials for live feedback: highlighting words as they are read, showing a provisional score, driving a progress indicator.
  • Use the final result for any decision you record. Scoring an item on a partial risks scoring it on half a word.

Speech endpointing

A result may set is_speech_endpoint = true to indicate the server believes the speaker has finished, typically after a period of silence. This is useful for closing the microphone automatically once an answer has been given.

Endpoint detection is optional, and clients must be robust to the field never being set. Do not treat it as the signal that results are final; that is what is_partial = false is for.

Non-fatal errors

Both response types carry an optional EvaluationError beside the result. It reports conditions that degrade quality but do not stop processing, most commonly audio sampled at a lower rate than the model expects. The server continues, and you may keep streaming. Log these rather than aborting; they usually point at a recording pipeline that needs attention.

Evaluating a whole file at once

For short audio, Evaluate avoids the streaming machinery entirely. Over the HTTP+JSON gateway that is a single POST:

curl -s -X POST http://localhost:8080/api/capt/v1/evaluate \
  -H 'Content-Type: application/json' \
  -d '{
    "config": {
      "model_id": "en_US-16khz",
      "reference_text": "WHEN THE SUNLIGHT STRIKES",
      "audio_format": {"audio_format_headered": "AUDIO_FORMAT_HEADERED_WAV"}
    },
    "audio": {"data": "<base64-encoded WAV file>"}
  }'

The data field is the audio file, base64 encoded. See Interpreting Results for what comes back.