The API is defined as a protobuf spec, so native bindings can be generated in any language with gRPC support. We recommend using buf to generate the bindings.
This section of the documentation is auto-generated from the protobuf spec. The service contains the methods that can be called, and the “messages” are the data structures (objects, classes or structs in the generated code, depending on the language) passed to and from the methods.
CAPTService
Service that implements the Cobalt CAPT API.
Version
Version(VersionRequest) VersionResponse
Returns version information from the server.
ListModels
ListModels(ListModelsRequest) ListModelsResponse
Returns information about the models available on the server.
Evaluate
Evaluate(EvaluateRequest) EvaluateResponse
Performs synchronous speech evaluation by receiving results after all
audio has been sent and processed. It is expected that this request be
typically used for short audio content: less than a minute long. For
longer content, the StreamingEvaluate method should be preferred.
StreamingEvaluate
StreamingEvaluate(StreamingEvaluateRequest) StreamingEvaluateResponse
Performs bidirectional streaming for evaluating speech by receiving results while sending audio. This method is only available via GRPC and not via HTTP+JSON. However, a web browser may use websockets to use this service.
Messages
- If two or more fields in a message are labeled oneof, then each method call using that message must have exactly one of the fields populated
- If a field is labeled
repeated, then the generated code will accept an array (or struct, or list depending on the language).
AlignedToken
AlignedToken contains a single reference token and the alternate hypotheses recognized from the audio that could be aligned with the reference token.
If kind == ALIGNMENT_KIND_MATCH or ALIGNMENT_KIND_SUBSTITUTION, both the reference token and hypotheses list will be populated. The score will be based on the confidence value of the reference token within the hypotheses. In the case of a substitution, the hypotheses list may not contain the actual reference token at all, and the score will be 0.
If kind == ALIGNMENT_KIND_DELETION, the hypotheses will be an empty list and score will be 0.
If kind == ALIGNMENT_KIND_INSERTION, the reference token will be an empty string and score will be 0.
Fields
kind (AlignmentKind ) The kind of alignment for this token.
reference (string ) The actual reference token.
score (float ) The score for this token, between 0 and 1 inclusive, based on the confidence value of the token within the recognized token hypotheses from the audio.
hypotheses (AlignmentHypothesis repeated) The recognized tokens from the audio that have been aligned to this reference token.
AlignedWord
AlignedWord contains the aligned tokens within a single word in the reference text.
Fields
text (string ) The actual word.
start_time_ms (uint64 ) The timestamp at which this token starts in the audio, in milliseconds.
duration_ms (uint64 ) The duration of this token in the audio, in milliseconds.
tokens (AlignedToken repeated) The tokens that make up the word.
AlignmentHypothesis
AlignmentHypothesis contains the metadata for a token recognized from audio, including its timestamps and confidence scores.
Fields
token (string ) The actual token.
confidence (float ) The confidence with which this token was recognized in the audio, between 0 and 1 inclusive.
start_time_ms (uint64 ) The timestamp at which this token starts in the audio, in milliseconds.
duration_ms (uint64 ) The duration of this token in the audio, in milliseconds.
AlternativeAlignment
AlternativeAlignment contains alignments against alternative reference text(s) if specified in the EvaluationConfig.
Fields
reference_text (string ) The alternative reference text.
score (float ) Evaluation score, between 0 and 1 inclusive, with 1 indicating a perfect match between the alternative reference text and what was said in the audio.
alignments (AlignedWord repeated) Alignment between the alternative reference text tokens and recognized tokens.
Audio
Audio to be sent to Capt.
Fields
- data (bytes )
AudioFormat
Format of the audio to be sent for recognition.
Depending on how they are configured, server instances of this service may not support all the formats provided in the API. One format that is guaranteed to be supported is the RAW format with little-endian 16-bit signed samples with the sample rate matching that of the model being requested.
Fields
oneof audio_format.audio_format_raw (AudioFormatRAW ) Audio is raw data without any headers
oneof audio_format.audio_format_headered (AudioFormatHeadered ) Audio has a self-describing header. Headers are expected to be sent at the beginning of the entire audio file/stream, and not in every
Audiomessage.The default value of this type is AUDIO_FORMAT_HEADERED_UNSPECIFIED. If this value is used, the server may attempt to detect the format of the audio. However, it is recommended that the exact format be specified.
AudioFormatRAW
Details of audio in raw format
Fields
encoding (AudioEncoding ) Encoding of the samples. It must be specified explicitly and using the default value of
AUDIO_ENCODING_UNSPECIFIEDwill result in an error.bit_depth (uint32 ) Bit depth of each sample (e.g. 8, 16, 24, 32, etc.). This is a required field.
byte_order (ByteOrder ) Byte order of the samples. This field must be set to a value other than
BYTE_ORDER_UNSPECIFIEDwhen thebit_depthis greater than 8.sample_rate (uint32 ) Sampling rate in Hz. This is a required field.
channels (uint32 ) Number of channels present in the audio. E.g.: 1 (mono), 2 (stereo), etc. This is a required field.
EvaluateRequest
The top-level message sent by the client for the Evaluate method. Both the
EvaluationConfig and Audio fields are required. The entire audio data
must be sent in one request. If your audio data is larger, please use the
StreamingEvaluate call.
Fields
config (EvaluationConfig )
audio (Audio )
EvaluateResponse
The message returned by the server for the Evaluate method.
Fields
evaluation_result (EvaluationResult ) Response from the server. The kind of result depends on the kind of the model that was chosen in the EvaluationConfig.
error (EvaluationError ) A non-fatal error message. If a server encountered a non-fatal error when processing the request, it will be returned in this message. The server will continue to process audio and produce further results. Clients can continue streaming audio even after receiving these messages. This error message is meant to be informational.
An example of when these errors maybe produced: audio is sampled at a lower rate than expected by model, producing possibly less accurate results.
This field will be unset if there is no error to report.
EvaluationConfig
Configuration for a StreamingEvaluateRequest.
Fields
model_id (string ) ID of the model to use. A list of supported IDs can be found using the
ListModelscall.audio_format (AudioFormat ) Format of the audio to be sent.
reference_text (string ) Reference Text that is expected to be contained in the audio.
alternative_reference_text (string repeated) Alternative reference text(s) that are also acceptable if recognized in the audio.
metadata (EvaluationMetadata ) This is an optional field. If there is any metadata associated with the audio being sent, use this field to provide it to the recognizer. The server may record this metadata when processing the request. The server does not use this field for any other purpose.
EvaluationError
Developer-facing error message about a non-fatal process issue.
Fields
- message (string )
EvaluationMetadata
Metadata associated with the evaluation request
Fields
custom_metadata (string ) Any custom metadata that the client wants to associate with the recording. This could be a simple string (e.g. a tracing ID) or structured data (e.g. JSON).
custom_id (string ) This is an optional field to specify custom ID to identify the evaluation request. The custom ID must be a string of upto 64 bytes, and only alphabets, digits, hyphens and underscores are allowed. This ID may be recorded by the server in logs or other storage, and should therefore not include any sensitive information.
EvaluationResult
EvaluationResult contains the result generated by a speech evaluation model.
Fields
is_partial (bool ) If this is set to true, it denotes that the result is an interim partial result, and could change after more audio is processed. If unset, or set to false, it denotes that this is a final result and will not change.
Servers are not required to implement support for returning partial results, and clients should generally not depend on their availability.
score (float ) Overall evaluation score, between 0 and 1 inclusive, with 1 indicating a perfect match between the reference text and what was said in the audio. If alternative reference text(s) are specified, then the score will be take those into account and be based on the reference text that aligns the best with recognized tokens.
alignments (AlignedWord repeated) Alignment between the primary expected reference text tokens and recognized tokens.
alternative_alignments (AlternativeAlignment repeated) Alignments against alternative reference text tokens (if specified) and recognized tokens.
is_speech_endpoint (bool ) If true, indicates that a speech endpoint has been detected.
A speech endpoint signifies that the server believes the user has finished their utterance, typically after detecting a specific duration of silence.
NOTE: Server support for endpoint detection is optional. Clients must be robust to this field never being set.
ListModelsRequest
The top-level message sent by the client for the ListModels method.
ListModelsResponse
The message returned to the client by the ListModels method.
Fields
- models (Model repeated) List of models available for use that match the request.
Model
Description of a CAPT model.
Fields
id (string ) Unique identifier of the model. This identifier is used to choose the model the model when configuring a StreamingEvaluate request.
name (string ) Model name. This is a concise name describing the model, and may be presented to the end-user, for example, to help choose which model to use for their task.
kind (ModelKind ) The specific kind of model. This determines what type of results it sends back, and additional config requirements if any.
attributes (ModelAttributes ) Model Attributes.
ModelAttributes
Attributes of a Capt model.
Fields
sample_rate (uint32 ) Audio sample rate (native) supported by the model.
metadata (ModelMetadata ) Metadata associated with the model.
ModelMetadata
Metadata associated with a Capt model.
Fields
version (string ) Model version in semver format (e.g. 1.0.0). If the version is not known, it will default to “unknown”.
build_date (string ) Date the model was built in YYYY-MM-DD format. If the version is not known, it will default to “unknown”.
StreamingEvaluateRequest
The top level messages sent by the client for the StreamingEvaluate method.
In this streaming call, multiple StreamingEvaluateRequest messages should
be sent. The first message must contain a EvaluationConfig message, and all
subsequent messages must contain Audio only. All Audio messages must
contain non-empty audio. If audio content is empty, the server may choose to
interpret it as end of stream and stop accepting any further messages.
Fields
oneof request.config (EvaluationConfig )
StreamingEvaluateResponse
The message returned by the server for the StreamingEvaluate method.
Fields
evaluation_result (EvaluationResult ) Response from the server. The kind of result depends on the kind of the model that was chosen in the EvaluationConfig.
error (EvaluationError ) A non-fatal error message. If a server encountered a non-fatal error when processing the request, it will be returned in this message. The server will continue to process audio and produce further results. Clients can continue streaming audio even after receiving these messages. This error message is meant to be informational.
An example of when these errors maybe produced: audio is sampled at a lower rate than expected by model, producing possibly less accurate results.
This field will be unset if there is no error to report.
VersionRequest
The top-level message sent by the client for the Version method.
VersionResponse
The message sent by the server for the Version method.
Fields
- version (string ) Version of the server handling these requests.
Enums
AlignmentKind
AlignmentKind represents one of four alignment outcomes possible: a Match, Substitution, Deletion or Insertion.
| Name | Number | Description |
|---|---|---|
| ALIGNMENT_KIND_UNSPECIFIED | 0 | Default value of this type. |
| ALIGNMENT_KIND_MATCH | 1 | A match implies that the reference and recognized token from audio are a match with a reasonable amount of confidence. |
| ALIGNMENT_KIND_SUBSTITUTION | 2 | A substitution implies that the reference token has been replaced by a different token in the recognized tokens from the audio. |
| ALIGNMENT_KIND_DELETION | 3 | A deletion implies that the reference token was not found in the recognized tokens from audio. |
| ALIGNMENT_KIND_INSERTION | 4 | A insertion implies that a extraneous token has been recognized in the audio, that cannot be matched to any token in the reference. |
AudioEncoding
The encoding of the audio data to be sent for recognition.
| Name | Number | Description |
|---|---|---|
| AUDIO_ENCODING_UNSPECIFIED | 0 | AUDIO_ENCODING_UNSPECIFIED is the default value of this type and will result in an error. |
| AUDIO_ENCODING_SIGNED | 1 | PCM signed-integer |
| AUDIO_ENCODING_UNSIGNED | 2 | PCM unsigned-integer |
| AUDIO_ENCODING_IEEE_FLOAT | 3 | PCM IEEE-Float |
| AUDIO_ENCODING_ULAW | 4 | G.711 mu-law |
| AUDIO_ENCODING_ALAW | 5 | G.711 a-law |
AudioFormatHeadered
| Name | Number | Description |
|---|---|---|
| AUDIO_FORMAT_HEADERED_UNSPECIFIED | 0 | AUDIO_FORMAT_HEADERED_UNSPECIFIED is the default value of this type. |
| AUDIO_FORMAT_HEADERED_WAV | 1 | WAV with RIFF headers |
| AUDIO_FORMAT_HEADERED_MP3 | 2 | MP3 format with a valid frame header at the beginning of data |
| AUDIO_FORMAT_HEADERED_FLAC | 3 | FLAC format |
| AUDIO_FORMAT_HEADERED_OGG_OPUS | 4 | Opus format with OGG header |
ByteOrder
Byte order of multi-byte data
| Name | Number | Description |
|---|---|---|
| BYTE_ORDER_UNSPECIFIED | 0 | BYTE_ORDER_UNSPECIFIED is the default value of this type. |
| BYTE_ORDER_LITTLE_ENDIAN | 1 | Little Endian byte order |
| BYTE_ORDER_BIG_ENDIAN | 2 | Big Endian byte order |
ModelKind
| Name | Number | Description |
|---|---|---|
| MODEL_KIND_UNSPECIFIED | 0 | Default value of this type. |
| MODEL_KIND_SPEECH_EVALUATION | 1 | Model for evaluating the accuracy of spoken text from audio. This is done by first recognizing what’s said in the audio, and then aligning expected and recognized tokens (phonemes, syllables, etc.). This model returns results in the form of EvaluationResult messages. |
| MODEL_KIND_PHONEME_EVALUATION | 2 | Model for evaluating the accuracy of phonemes pronounced in isolation from audio. This is done in a way similar to speech evaluation models, but is more accurate for single phonemes in isolation, which a regular speech model may not recognize correctly. This model returns results in the form of EvaluationResult messages. |
Scalar Value Types
| .proto Type | C++ Type | C# Type | Go Type | Java Type | PHP Type | Python Type | Ruby Type |
|---|---|---|---|---|---|---|---|
double | double | double | float64 | double | float | float | Float |
float | float | float | float32 | float | float | float | Float |
int32 | int32 | int | int32 | int | integer | int | Bignum or Fixnum (as required) |
int64 | int64 | long | int64 | long | integer/string | int/long | Bignum |
uint32 | uint32 | uint | uint32 | int | integer | int/long | Bignum or Fixnum (as required) |
uint64 | uint64 | ulong | uint64 | long | integer/string | int/long | Bignum or Fixnum (as required) |
sint32 | int32 | int | int32 | int | integer | int | Bignum or Fixnum (as required) |
sint64 | int64 | long | int64 | long | integer/string | int/long | Bignum |
fixed32 | uint32 | uint | uint32 | int | integer | int | Bignum or Fixnum (as required) |
fixed64 | uint64 | ulong | uint64 | long | integer/string | int/long | Bignum |
sfixed32 | int32 | int | int32 | int | integer | int | Bignum or Fixnum (as required) |
sfixed64 | int64 | long | int64 | long | integer/string | int/long | Bignum |
bool | bool | bool | bool | boolean | boolean | boolean | TrueClass/FalseClass |
string | string | string | string | String | string | str/unicode | String (UTF-8) |
bytes | string | ByteString | []byte | ByteString | string | str | String (ASCII-8BIT) |