This is the multi-page printable view of this section. Click here to print.

Return to the regular view of this page.

API Reference

Detailed reference for API requests and types.

    The API is defined as a protobuf spec, so native bindings can be generated in any language with gRPC support. We recommend using buf to generate the bindings.

    This section of the documentation is auto-generated from the protobuf spec. The service contains the methods that can be called, and the “messages” are the data structures (objects, classes or structs in the generated code, depending on the language) passed to and from the methods.

    CAPTService

    Service that implements the Cobalt CAPT API.

    Version

    Version(VersionRequest) VersionResponse

    Returns version information from the server.

    ListModels

    ListModels(ListModelsRequest) ListModelsResponse

    Returns information about the models available on the server.

    Evaluate

    Evaluate(EvaluateRequest) EvaluateResponse

    Performs synchronous speech evaluation by receiving results after all audio has been sent and processed. It is expected that this request be typically used for short audio content: less than a minute long. For longer content, the StreamingEvaluate method should be preferred.

    StreamingEvaluate

    StreamingEvaluate(StreamingEvaluateRequest) StreamingEvaluateResponse

    Performs bidirectional streaming for evaluating speech by receiving results while sending audio. This method is only available via GRPC and not via HTTP+JSON. However, a web browser may use websockets to use this service.

    Messages

    • If two or more fields in a message are labeled oneof, then each method call using that message must have exactly one of the fields populated
    • If a field is labeled repeated, then the generated code will accept an array (or struct, or list depending on the language).

    AlignedToken

    AlignedToken contains a single reference token and the alternate hypotheses recognized from the audio that could be aligned with the reference token.

    If kind == ALIGNMENT_KIND_MATCH or ALIGNMENT_KIND_SUBSTITUTION, both the reference token and hypotheses list will be populated. The score will be based on the confidence value of the reference token within the hypotheses. In the case of a substitution, the hypotheses list may not contain the actual reference token at all, and the score will be 0.

    If kind == ALIGNMENT_KIND_DELETION, the hypotheses will be an empty list and score will be 0.

    If kind == ALIGNMENT_KIND_INSERTION, the reference token will be an empty string and score will be 0.

    Fields

    • kind (AlignmentKind ) The kind of alignment for this token.

    • reference (string ) The actual reference token.

    • score (float ) The score for this token, between 0 and 1 inclusive, based on the confidence value of the token within the recognized token hypotheses from the audio.

    • hypotheses (AlignmentHypothesis repeated) The recognized tokens from the audio that have been aligned to this reference token.

    AlignedWord

    AlignedWord contains the aligned tokens within a single word in the reference text.

    Fields

    • text (string ) The actual word.

    • start_time_ms (uint64 ) The timestamp at which this token starts in the audio, in milliseconds.

    • duration_ms (uint64 ) The duration of this token in the audio, in milliseconds.

    • tokens (AlignedToken repeated) The tokens that make up the word.

    AlignmentHypothesis

    AlignmentHypothesis contains the metadata for a token recognized from audio, including its timestamps and confidence scores.

    Fields

    • token (string ) The actual token.

    • confidence (float ) The confidence with which this token was recognized in the audio, between 0 and 1 inclusive.

    • start_time_ms (uint64 ) The timestamp at which this token starts in the audio, in milliseconds.

    • duration_ms (uint64 ) The duration of this token in the audio, in milliseconds.

    AlternativeAlignment

    AlternativeAlignment contains alignments against alternative reference text(s) if specified in the EvaluationConfig.

    Fields

    • reference_text (string ) The alternative reference text.

    • score (float ) Evaluation score, between 0 and 1 inclusive, with 1 indicating a perfect match between the alternative reference text and what was said in the audio.

    • alignments (AlignedWord repeated) Alignment between the alternative reference text tokens and recognized tokens.

    Audio

    Audio to be sent to Capt.

    Fields

    AudioFormat

    Format of the audio to be sent for recognition.

    Depending on how they are configured, server instances of this service may not support all the formats provided in the API. One format that is guaranteed to be supported is the RAW format with little-endian 16-bit signed samples with the sample rate matching that of the model being requested.

    Fields

    • oneof audio_format.audio_format_raw (AudioFormatRAW ) Audio is raw data without any headers

    • oneof audio_format.audio_format_headered (AudioFormatHeadered ) Audio has a self-describing header. Headers are expected to be sent at the beginning of the entire audio file/stream, and not in every Audio message.

      The default value of this type is AUDIO_FORMAT_HEADERED_UNSPECIFIED. If this value is used, the server may attempt to detect the format of the audio. However, it is recommended that the exact format be specified.

    AudioFormatRAW

    Details of audio in raw format

    Fields

    • encoding (AudioEncoding ) Encoding of the samples. It must be specified explicitly and using the default value of AUDIO_ENCODING_UNSPECIFIED will result in an error.

    • bit_depth (uint32 ) Bit depth of each sample (e.g. 8, 16, 24, 32, etc.). This is a required field.

    • byte_order (ByteOrder ) Byte order of the samples. This field must be set to a value other than BYTE_ORDER_UNSPECIFIED when the bit_depth is greater than 8.

    • sample_rate (uint32 ) Sampling rate in Hz. This is a required field.

    • channels (uint32 ) Number of channels present in the audio. E.g.: 1 (mono), 2 (stereo), etc. This is a required field.

    EvaluateRequest

    The top-level message sent by the client for the Evaluate method. Both the EvaluationConfig and Audio fields are required. The entire audio data must be sent in one request. If your audio data is larger, please use the StreamingEvaluate call.

    Fields

    EvaluateResponse

    The message returned by the server for the Evaluate method.

    Fields

    • evaluation_result (EvaluationResult ) Response from the server. The kind of result depends on the kind of the model that was chosen in the EvaluationConfig.

    • error (EvaluationError ) A non-fatal error message. If a server encountered a non-fatal error when processing the request, it will be returned in this message. The server will continue to process audio and produce further results. Clients can continue streaming audio even after receiving these messages. This error message is meant to be informational.

      An example of when these errors maybe produced: audio is sampled at a lower rate than expected by model, producing possibly less accurate results.

      This field will be unset if there is no error to report.

    EvaluationConfig

    Configuration for a StreamingEvaluateRequest.

    Fields

    • model_id (string ) ID of the model to use. A list of supported IDs can be found using the ListModels call.

    • audio_format (AudioFormat ) Format of the audio to be sent.

    • reference_text (string ) Reference Text that is expected to be contained in the audio.

    • alternative_reference_text (string repeated) Alternative reference text(s) that are also acceptable if recognized in the audio.

    • metadata (EvaluationMetadata ) This is an optional field. If there is any metadata associated with the audio being sent, use this field to provide it to the recognizer. The server may record this metadata when processing the request. The server does not use this field for any other purpose.

    EvaluationError

    Developer-facing error message about a non-fatal process issue.

    Fields

    EvaluationMetadata

    Metadata associated with the evaluation request

    Fields

    • custom_metadata (string ) Any custom metadata that the client wants to associate with the recording. This could be a simple string (e.g. a tracing ID) or structured data (e.g. JSON).

    • custom_id (string ) This is an optional field to specify custom ID to identify the evaluation request. The custom ID must be a string of upto 64 bytes, and only alphabets, digits, hyphens and underscores are allowed. This ID may be recorded by the server in logs or other storage, and should therefore not include any sensitive information.

    EvaluationResult

    EvaluationResult contains the result generated by a speech evaluation model.

    Fields

    • is_partial (bool ) If this is set to true, it denotes that the result is an interim partial result, and could change after more audio is processed. If unset, or set to false, it denotes that this is a final result and will not change.

      Servers are not required to implement support for returning partial results, and clients should generally not depend on their availability.

    • score (float ) Overall evaluation score, between 0 and 1 inclusive, with 1 indicating a perfect match between the reference text and what was said in the audio. If alternative reference text(s) are specified, then the score will be take those into account and be based on the reference text that aligns the best with recognized tokens.

    • alignments (AlignedWord repeated) Alignment between the primary expected reference text tokens and recognized tokens.

    • alternative_alignments (AlternativeAlignment repeated) Alignments against alternative reference text tokens (if specified) and recognized tokens.

    • is_speech_endpoint (bool ) If true, indicates that a speech endpoint has been detected.

      A speech endpoint signifies that the server believes the user has finished their utterance, typically after detecting a specific duration of silence.

      NOTE: Server support for endpoint detection is optional. Clients must be robust to this field never being set.

    ListModelsRequest

    The top-level message sent by the client for the ListModels method.

    ListModelsResponse

    The message returned to the client by the ListModels method.

    Fields

    • models (Model repeated) List of models available for use that match the request.

    Model

    Description of a CAPT model.

    Fields

    • id (string ) Unique identifier of the model. This identifier is used to choose the model the model when configuring a StreamingEvaluate request.

    • name (string ) Model name. This is a concise name describing the model, and may be presented to the end-user, for example, to help choose which model to use for their task.

    • kind (ModelKind ) The specific kind of model. This determines what type of results it sends back, and additional config requirements if any.

    • attributes (ModelAttributes ) Model Attributes.

    ModelAttributes

    Attributes of a Capt model.

    Fields

    • sample_rate (uint32 ) Audio sample rate (native) supported by the model.

    • metadata (ModelMetadata ) Metadata associated with the model.

    ModelMetadata

    Metadata associated with a Capt model.

    Fields

    • version (string ) Model version in semver format (e.g. 1.0.0). If the version is not known, it will default to “unknown”.

    • build_date (string ) Date the model was built in YYYY-MM-DD format. If the version is not known, it will default to “unknown”.

    StreamingEvaluateRequest

    The top level messages sent by the client for the StreamingEvaluate method. In this streaming call, multiple StreamingEvaluateRequest messages should be sent. The first message must contain a EvaluationConfig message, and all subsequent messages must contain Audio only. All Audio messages must contain non-empty audio. If audio content is empty, the server may choose to interpret it as end of stream and stop accepting any further messages.

    Fields

    StreamingEvaluateResponse

    The message returned by the server for the StreamingEvaluate method.

    Fields

    • evaluation_result (EvaluationResult ) Response from the server. The kind of result depends on the kind of the model that was chosen in the EvaluationConfig.

    • error (EvaluationError ) A non-fatal error message. If a server encountered a non-fatal error when processing the request, it will be returned in this message. The server will continue to process audio and produce further results. Clients can continue streaming audio even after receiving these messages. This error message is meant to be informational.

      An example of when these errors maybe produced: audio is sampled at a lower rate than expected by model, producing possibly less accurate results.

      This field will be unset if there is no error to report.

    VersionRequest

    The top-level message sent by the client for the Version method.

    VersionResponse

    The message sent by the server for the Version method.

    Fields

    • version (string ) Version of the server handling these requests.

    Enums

    AlignmentKind

    AlignmentKind represents one of four alignment outcomes possible: a Match, Substitution, Deletion or Insertion.

    NameNumberDescription
    ALIGNMENT_KIND_UNSPECIFIED0Default value of this type.
    ALIGNMENT_KIND_MATCH1A match implies that the reference and recognized token from audio are a match with a reasonable amount of confidence.
    ALIGNMENT_KIND_SUBSTITUTION2A substitution implies that the reference token has been replaced by a different token in the recognized tokens from the audio.
    ALIGNMENT_KIND_DELETION3A deletion implies that the reference token was not found in the recognized tokens from audio.
    ALIGNMENT_KIND_INSERTION4A insertion implies that a extraneous token has been recognized in the audio, that cannot be matched to any token in the reference.

    AudioEncoding

    The encoding of the audio data to be sent for recognition.

    NameNumberDescription
    AUDIO_ENCODING_UNSPECIFIED0AUDIO_ENCODING_UNSPECIFIED is the default value of this type and will result in an error.
    AUDIO_ENCODING_SIGNED1PCM signed-integer
    AUDIO_ENCODING_UNSIGNED2PCM unsigned-integer
    AUDIO_ENCODING_IEEE_FLOAT3PCM IEEE-Float
    AUDIO_ENCODING_ULAW4G.711 mu-law
    AUDIO_ENCODING_ALAW5G.711 a-law

    AudioFormatHeadered

    NameNumberDescription
    AUDIO_FORMAT_HEADERED_UNSPECIFIED0AUDIO_FORMAT_HEADERED_UNSPECIFIED is the default value of this type.
    AUDIO_FORMAT_HEADERED_WAV1WAV with RIFF headers
    AUDIO_FORMAT_HEADERED_MP32MP3 format with a valid frame header at the beginning of data
    AUDIO_FORMAT_HEADERED_FLAC3FLAC format
    AUDIO_FORMAT_HEADERED_OGG_OPUS4Opus format with OGG header

    ByteOrder

    Byte order of multi-byte data

    NameNumberDescription
    BYTE_ORDER_UNSPECIFIED0BYTE_ORDER_UNSPECIFIED is the default value of this type.
    BYTE_ORDER_LITTLE_ENDIAN1Little Endian byte order
    BYTE_ORDER_BIG_ENDIAN2Big Endian byte order

    ModelKind

    NameNumberDescription
    MODEL_KIND_UNSPECIFIED0Default value of this type.
    MODEL_KIND_SPEECH_EVALUATION1Model for evaluating the accuracy of spoken text from audio. This is done by first recognizing what’s said in the audio, and then aligning expected and recognized tokens (phonemes, syllables, etc.). This model returns results in the form of EvaluationResult messages.
    MODEL_KIND_PHONEME_EVALUATION2Model for evaluating the accuracy of phonemes pronounced in isolation from audio. This is done in a way similar to speech evaluation models, but is more accurate for single phonemes in isolation, which a regular speech model may not recognize correctly. This model returns results in the form of EvaluationResult messages.

    Scalar Value Types

    .proto TypeC++ TypeC# TypeGo TypeJava TypePHP TypePython TypeRuby Type

    double
    doubledoublefloat64doublefloatfloatFloat

    float
    floatfloatfloat32floatfloatfloatFloat

    int32
    int32intint32intintegerintBignum or Fixnum (as required)

    int64
    int64longint64longinteger/stringint/longBignum

    uint32
    uint32uintuint32intintegerint/longBignum or Fixnum (as required)

    uint64
    uint64ulonguint64longinteger/stringint/longBignum or Fixnum (as required)

    sint32
    int32intint32intintegerintBignum or Fixnum (as required)

    sint64
    int64longint64longinteger/stringint/longBignum

    fixed32
    uint32uintuint32intintegerintBignum or Fixnum (as required)

    fixed64
    uint64ulonguint64longinteger/stringint/longBignum

    sfixed32
    int32intint32intintegerintBignum or Fixnum (as required)

    sfixed64
    int64longint64longinteger/stringint/longBignum

    bool
    boolboolboolbooleanbooleanbooleanTrueClass/FalseClass

    string
    stringstringstringStringstringstr/unicodeString (UTF-8)

    bytes
    stringByteString[]byteByteStringstringstrString (ASCII-8BIT)