CAPT offers two ways to submit audio.
StreamingEvaluateis the primary method. It is bidirectional: you send audio as it is captured and receive results as the audio is processed. Use it for anything interactive, and for audio of any length.Evaluateis a unary convenience method that takes the whole audio in a single request and returns a single result. It is intended for short audio, less than a minute, and is the simplest way to score an audio file.
Both accept the same EvaluationConfig and
return the same EvaluationResult.
Streaming from an audio file
In a StreamingEvaluate stream, the first message must contain the config,
and every subsequent message must contain audio.
An empty Audio message ends the stream. CAPT server treats it as the end
of the audio: it finishes processing, sends the final result, and accepts
nothing further, so a later send fails. Over gRPC you normally do not need it,
because half-closing the stream (CloseSend in Go, exhausting the request
iterator in Python) says the same thing. It matters for
browser clients, where a websocket has no
equivalent half-close.
Do not send one to mean “no audio yet”. If it arrives before the audio is complete, what you get depends on the format, and a client’s state machine needs to handle both:
- You always receive a final result for the audio sent so far, with
is_partialfalse. It reflects the truncated input, not the utterance you intended, so it will usually be full of deletions. - With a headered format the stream then fails. The header promised more
data than arrived, so the server reports a container error such as
wav: unexpected EOFafter the final result. This happens whether or not you send anything afterwards. - With raw audio the stream ends cleanly, because there is no container to leave incomplete. The result is still only of the audio you sent.
In other words: a final result is not by itself evidence that the whole utterance was processed. Send the empty message only once the recording is complete, and treat an error after a final result as “the result is incomplete”, not “the result is invalid”.
The examples below stream a WAV file in 100 ms chunks and print each result as it arrives.
import grpc
import cobaltspeech.capt.v1.capt_pb2 as capt
import cobaltspeech.capt.v1.capt_pb2_grpc as capt_grpc
serverAddress = "localhost:2727"
channel = grpc.insecure_channel(serverAddress)
client = capt_grpc.CAPTServiceStub(channel)
# Get the list of models on the server and use the first one.
modelResp = client.ListModels(capt.ListModelsRequest())
modelID = modelResp.models[0].id
cfg = capt.EvaluationConfig(
model_id=modelID,
reference_text="WHEN THE SUNLIGHT STRIKES",
audio_format=capt.AudioFormat(
audio_format_headered=capt.AUDIO_FORMAT_HEADERED_WAV,
),
)
# The first request must contain only the configuration; subsequent
# requests carry audio bytes. A generator is a convenient way to do this.
def stream(cfg, audio, bufferSize=3200):
yield capt.StreamingEvaluateRequest(config=cfg)
data = audio.read(bufferSize)
while len(data) > 0:
yield capt.StreamingEvaluateRequest(audio=capt.Audio(data=data))
data = audio.read(bufferSize)
with open("test.wav", "rb") as audio:
for resp in client.StreamingEvaluate(stream(cfg, audio)):
if resp.HasField("error"):
print(f"warning: {resp.error.message}")
result = resp.evaluation_result
print(f"partial={result.is_partial} score={result.score:.4f}")
# Only act on final results; partials will still change.
if not result.is_partial:
for word in result.alignments:
print(f" {word.text}: " + " ".join(
f"{t.reference}={t.score:.2f}" for t in word.tokens
))package main
import (
"context"
"fmt"
"io"
"os"
"google.golang.org/grpc"
"google.golang.org/grpc/credentials/insecure"
captpb "github.com/your-org/your-module/gen/go/cobaltspeech/capt/v1"
)
func main() {
const serverAddress = "localhost:2727"
audio, err := os.ReadFile("test.wav")
if err != nil {
fmt.Printf("failed to read audio: %v\n", err)
os.Exit(1)
}
conn, err := grpc.NewClient(serverAddress,
grpc.WithTransportCredentials(insecure.NewCredentials()))
if err != nil {
fmt.Printf("failed to dial gRPC connection: %v\n", err)
os.Exit(1)
}
defer conn.Close()
client := captpb.NewCAPTServiceClient(conn)
stream, err := client.StreamingEvaluate(context.Background())
if err != nil {
fmt.Printf("failed to open stream: %v\n", err)
os.Exit(1)
}
// The first message must contain the config.
if err := stream.Send(&captpb.StreamingEvaluateRequest{
Request: &captpb.StreamingEvaluateRequest_Config{
Config: &captpb.EvaluationConfig{
ModelId: "en_US-16khz",
ReferenceText: "WHEN THE SUNLIGHT STRIKES",
AudioFormat: &captpb.AudioFormat{
AudioFormat: &captpb.AudioFormat_AudioFormatHeadered{
AudioFormatHeadered: captpb.AudioFormatHeadered_AUDIO_FORMAT_HEADERED_WAV,
},
},
},
},
}); err != nil {
fmt.Printf("failed to send config: %v\n", err)
os.Exit(1)
}
// Send audio in the background while results are read below.
go func() {
const chunk = 3200 // 100ms of 16kHz 16-bit mono audio.
// min() is a builtin from Go 1.21; on older toolchains compute the
// bound explicitly.
for i := 0; i < len(audio); i += chunk {
end := min(i+chunk, len(audio))
if err := stream.Send(&captpb.StreamingEvaluateRequest{
Request: &captpb.StreamingEvaluateRequest_Audio{
Audio: &captpb.Audio{Data: audio[i:end]},
},
}); err != nil {
return
}
}
_ = stream.CloseSend()
}()
for {
resp, err := stream.Recv()
if err == io.EOF {
break
}
if err != nil {
fmt.Printf("failed to receive result: %v\n", err)
os.Exit(1)
}
if e := resp.GetError(); e != nil {
fmt.Printf("warning: %v\n", e.GetMessage())
}
result := resp.GetEvaluationResult()
fmt.Printf("partial=%v score=%.4f\n", result.GetIsPartial(), result.GetScore())
// Only act on final results; partials will still change.
if !result.GetIsPartial() {
for _, word := range result.GetAlignments() {
fmt.Printf(" %s:", word.GetText())
for _, t := range word.GetTokens() {
fmt.Printf(" %s=%.2f", t.GetReference(), t.GetScore())
}
fmt.Println()
}
}
}
}Streaming a 1.5 second recording of “when the sunlight strikes” against that reference produces a short sequence of results:
partial=true score=0.5514
partial=true score=0.6128
partial=true score=0.6128
partial=false score=0.6128
Partial results
While audio is still arriving, the server emits partial results, marked
is_partial = true. A partial reflects the audio received so far, and any part
of it may change as more audio arrives. In the sequence above the score rises
as the last word is heard.
When the stream ends, the server emits a final result with is_partial = false. That result is stable.
- Use partials for live feedback: highlighting words as they are read, showing a provisional score, driving a progress indicator.
- Use the final result for any decision you record. Scoring an item on a partial risks scoring it on half a word.
Servers are not required to produce partial results, and clients should not depend on receiving them. Always drive your logic from the final result and treat partials as an enhancement.
Speech endpointing
A result may set is_speech_endpoint = true to indicate the server believes
the speaker has finished, typically after a period of silence. This is useful
for closing the microphone automatically once an answer has been given.
Endpoint detection is optional, and clients must be robust to the field never
being set. Do not treat it as the signal that results are final; that is what
is_partial = false is for.
Non-fatal errors
Both response types carry an optional
EvaluationError beside the result. It
reports conditions that degrade quality but do not stop processing, most
commonly audio sampled at a lower rate than the model expects. The server
continues, and you may keep streaming. Log these rather than aborting; they
usually point at a recording pipeline that needs attention.
Evaluating a whole file at once
For short audio, Evaluate avoids the streaming machinery entirely. Over the
HTTP+JSON gateway that is a single POST:
curl -s -X POST http://localhost:8080/api/capt/v1/evaluate \
-H 'Content-Type: application/json' \
-d '{
"config": {
"model_id": "en_US-16khz",
"reference_text": "WHEN THE SUNLIGHT STRIKES",
"audio_format": {"audio_format_headered": "AUDIO_FORMAT_HEADERED_WAV"}
},
"audio": {"data": "<base64-encoded WAV file>"}
}'
The data field is the audio file, base64 encoded. See Interpreting
Results for what comes back.