This is the multi-page printable view of this section. Click here to print.

Return to the regular view of this page.

Getting Started

How to get a CAPT server running on your system

    Using Cobalt CAPT

    • A typical CAPT release, provided as a compressed archive, contains a linux binary (capt-server) for the required native CPU architecture, an appropriate Dockerfile, and models.

    • Cobalt CAPT runs either locally on linux or using Docker.

    • Cobalt CAPT serves the CAPT gRPC API on port 2727. The configuration shipped with a release also enables an HTTP+JSON gateway on port 8080 and operational endpoints on port 8081; both are opt-in settings rather than server defaults, so a config written from scratch has gRPC only.

    • To quickly try out CAPT, first start the server as shown below and use the SDK in your preferred language to call it from your application.

    Running CAPT Server Locally on Linux

    ./capt-server --config capt-server.cfg.toml
    

    By default the binary assumes a configuration file named capt-server.cfg.toml in the same directory. A different config file may be specified using the --config argument.

    On a successful start the server logs the addresses it is serving on:

    2026/09/10 03:57:50 info  {"msg":"server initializing"}
    2026/09/10 03:57:50 info  {"msg":"license verified"}
    2026/09/10 03:57:51 info  {"source":"transcribe","msg":"formats supported","formats":"[RAW WAV FLAC MP3 Opus]"}
    2026/09/10 03:57:51 info  {"msg":"runtime initialized","model_count":"1","init_time_taken":"1.485372717s"}
    2026/09/10 03:57:51 info  {"msg":"server started","grpcAddr":"[::]:2727","httpApiAddr":"[::]:8080","httpOpsAddr":"[::]:8081"}
    

    Running CAPT Server as a Docker Container

    To build and run the Docker image for CAPT, run:

    docker build -t cobalt-capt .
    docker run -p 2727:2727 -p 8080:8080 cobalt-capt
    

    Checking the server is up

    The HTTP API is the quickest way to confirm a working server, provided api.Address is set in your config as the shipped one does. It exposes the unary calls: Version, ListModels and Evaluate. The bidirectional StreamingEvaluate is gRPC-only, though browsers can reach it over a websocket:

    curl -s http://localhost:8080/api/capt/v1/version
    
    {"version":"1.4.0"}
    

    The reported version is that of the capt-server release you were given.

    curl -s http://localhost:8080/api/capt/v1/list-models
    
    {
      "models": [
        {
          "id": "en_US-16khz",
          "name": "en_US CAPT (16khz audio)",
          "kind": "MODEL_KIND_SPEECH_EVALUATION",
          "attributes": {
            "sample_rate": 16000,
            "metadata": { "version": "0.0.0", "build_date": "unknown" }
          }
        }
      ]
    }
    

    The id of a model in this list is what you pass as model_id when configuring an evaluation, and attributes.sample_rate is the audio sample rate the model expects. See Evaluation Configurations for what to do with them.

    How to Get a Copy of the CAPT Server and Models

    Contact us for a release best suited to your requirements.

    The release you receive is a compressed archive (tar.bz2), generally structured as follows:

    release.tar.bz2
    ├── COPYING
    ├── README.md
    ├── capt-server
    ├── capt-server.cfg.toml
    ├── Dockerfile
    ├── api
    │   ├── proto             [ cobaltspeech/capt/v1/capt.proto ]
    │   └── gen               [ pre-generated Go and Python bindings ]
    ├── models
    │   └── en_US-16khz
    │       ├── capt          [ lexicon, G2P model, cutoffs, model config ]
    │       └── transcribe    [ acoustic model and decoding graph ]
    │
    └── cobalt.license.key [ provided separately, needs to be copied over ]
    
    • The README.md file contains information about the release and instructions for starting the server on your system.

    • The capt-server is the server program, configured using the capt-server.cfg.toml file.

    • The Dockerfile can be used to create a container that will let you run CAPT server on non-linux systems such as macOS and Windows.

    • The api directory holds the protobuf definition of the API and pre-generated client bindings. CAPT’s proto is distributed with the release rather than from a public repository, so this is where it comes from. See Generating SDKs.

    • The models directory contains the evaluation models. Each model directory holds two halves: the transcribe acoustic model that recognizes phonemes, and the capt resources (pronunciation lexicon, G2P model and scoring cutoffs) that the evaluation is built from. Both are described in Tuning and Customization.

    System Requirements

    Cobalt CAPT runs on Linux, directly as a native application. You can evaluate the product on Windows or macOS using Docker Desktop, but we would not recommend that setup for production.

    CAPT runs on x86_64 and Arm64 / aarch64 CPUs, and a statically linked build is available for deployment onto minimal or embedded images. Because the engine is CPU-only and the models are small, the same release can be deployed on a server, in a private cloud, or embedded on a device. Tell us the target and we will provide a build for it.

    Sizing depends on the model and on how many evaluations you need to run concurrently. Please contact us for sizing guidance against your target hardware and expected load.

    To integrate Cobalt CAPT into your application, follow the next steps to install or generate the SDK in a language of your choice.