# API Reference

# Voice Translation WebSocket API

## Overview

The Voice Translation WebSocket API provides real-time audio transcription and translation via a persistent WebSocket connection.

***

## Client ↔︎ WSS Server

### Endpoint

```
wss://streaming.krisp.ai/vt?authorization=Api-Key SESSION_KEY
```

## Session key generation

```javascript
const axios = require('axios');

let config = {
  method: 'get',
  url: 'https://api.developers.krisp.ai/v2/sdk/voice-translation/session/token?expiration_ttl=100'
	// expiration_ttl is in MINUTES. Min: 5, Max: 1440 (24h).
  headers: { 
    'Authorization': 'api-key API_KEY'
  }
};

axios.request(config)
  .then((response) => {
  // response.data.data = {
  //   session_key: "session_key_...",
  //   expires_at: "2026-05-04T12:00:00.000Z",
  //   key_id: 123,
  //   status: "active",
  //   type: "session"
  // }
	const SESSION_KEY = response.data.data.session_key;
})
.catch((error) => {
  console.log(error);
});
```
```python
# 1. Get a short-lived session key (or have your backend mint one).
session_key = get_vt_session_key(API_KEY)["session_key"]

# 2. Build a session config.
config = VtSessionConfig(
    auth_token           = session_key,
    input_language_code  = "en-US",         # BCP-47, source
    output_language_code = "fr-FR",         # BCP-47, target
    voice                = VtVoice.FEMALE,   # "male" | "female"
    # input/output_sample_rate, input_frame_duration default to 16 kHz / 20 ms
    custom_vocabulary    = VtCustomVocabularyData(
        vocabulary = ["Krisp", "AcmeCorp"], # ASR boost for domain terms (optional)
        dictionary = {                      # force specific translations (optional)
            "hello": "bonjour",
            "goodbye": "au revoir",
        },
    ),
    metadata             = VtSessionMetadata(
        reference_id = "your-reference-id",  # optional correlation id for support
    ),
    # Optional opt-in feature flags (default False), e.g.:
    # background_voice_cancellation = True,    # filter other voices server-side
)

# 3. Open the session; register five callbacks. They all fire on the SDK's
#    background thread, so route results through a queue if you need them
#    back on the main thread.
vt = Vt.create(
    config,
    original_transcript_callback   = on_source_text,       # interim + final source text
    translated_transcript_callback = on_target_text,       # interim + final target text
    audio_result_callback          = on_translated_audio,  # raw int16 PCM bytes
    event_callback                 = on_event,             # INPUT_ALLOWED / INPUT_NOT_ALLOWED
    error_callback                 = on_error,             # VtErrorType
)

# 4. Stream audio in real time. One PCM chunk per `input_frame_duration` ms.
#    16 kHz mono s16le => 640 bytes per 20 ms chunk.
for chunk in pcm_chunks:
    vt.process(chunk)
    sleep(0.020)                       # pace in real time — required by the API

# 5. Drain trailing events, then close.
sleep(2.0)                             # let the final transcript / translation / audio land
vt.close()                             # idempotent
```

***

### Initial Client Message

Once the WebSocket connection is established, the client must send a single JSON configuration message as the first message, before sending any audio.

```json
{
  "config": {
    "audio": {
      "format": "pcm_s16le", // pcm_s16le or Opus
      "sample_rate": 16000,
      "channels": 1
    },
    "source_language": "en-US",
    "target_language": "fr-FR",
    "voice": "male" | "female",
    "vocabulary": ["special", "domain", "terms"],
    "translation_dictionary": [
      { "source": "referral", "target": "référence" },
      { "source": "copay", "target": "quote-part" },
      { "source": "email", "target": "courriel" }
    ],
    "transcript": {
      "interim": true,
      "final": true,
      "translate": true
    },
	  "features": {
	    "background_voice_cancellation": true
	  },
    "metadata": {
      "reference_id": "your-reference-id"
    }
  }
}
```

### Parameters

| Parameter                | Type             | Required | Default     | Description                                                                                                                 |
| ------------------------ | ---------------- | -------- | ----------- | --------------------------------------------------------------------------------------------------------------------------- |
| `audio.format`           | string           | No       | `pcm_s16le` | Audio encoding format                                                                                                       |
| `audio.sample_rate`      | integer          | No       | `16000`     | Audio sample rate in Hz                                                                                                     |
| `audio.channels`         | integer          | No       | `1`         | Number of audio channels                                                                                                    |
| `source_language`        | string           | Yes      | —           | Source language locale in [BCP 47](https://www.rfc-editor.org/rfc/rfc5646) format (e.g. `en-US`)                            |
| `target_language`        | string           | Yes      | —           | Target language locale in [BCP 47](https://www.rfc-editor.org/rfc/rfc5646) format (e.g. `fr-FR`)                            |
| `voice`                  | string           | ✅        | —           | Voice ID for translated audio output                                                                                        |
| `vocabulary`             | array of strings | No       | —           | Domain-specific or uncommon words for improved recognition accuracy                                                         |
| `translation_dictionary` | array of objects | No       | —           | Custom source→target term mappings for ambiguous or specialized vocabulary                                                  |
| `transcript.interim`     | boolean          | No       | `true`      | When the entire transcript object is omitted. If a transcript is provided but the interim is missing, it defaults to false. |
| `transcript.final`       | boolean          | No       | `true`      | `true` for the final transcript of an utterance, `false` for interim transcripts                                            |
| `transcript.translate`   | boolean          | No       | `true`      | Whether to emit translated transcript events                                                                                |
| `features`               | json             | No       | -           | Additional features                                                                                                         |
| `metadata`               | json             | No       | —           | Client-supplied metadata                                                                                                    |

***

### &#x20;Binary Audio

After the initial JSON configuration message, audio is exchanged as raw **binary<br />WebSocket frames** in both directions. Each frame carries audio in<br />the encoding declared in `config.audio.format`:

* `pcm_s16le` — raw 16-bit little-endian PCM chunks
* `opus` — one Opus packet per frame

The server streams the synthesized translated speech back in the same encoding.
Regardless of format, audio is 16 kHz, mono.

| Direction       | Description                                                                  |
| --------------- | ---------------------------------------------------------------------------- |
| Client → Server | Raw PCM audio chunks in the format declared in `audio.*` fields              |
| Server → Client | Synthesised translated speech PCM audio, same encoding as the inbound stream |
|                 |                                                                              |

Binary frames and JSON event frames coexist on the same WebSocket connection. The receiver distinguishes them by WebSocket frame opcode: `0x2` (binary) for audio, `0x1` (text) for JSON events.

***

### Server Events (Client ↔︎ WSS Server)

The server emits JSON text frames over the WebSocket. Each event corresponds to one of the types below.

### Transcript

```json
{
  "transcript": {
    "text": "Hello, how are you?",
    "final": false,
    "start": "2026-03-25T19:24:45.370+00:00",
    "duration": 436,
    "utterance_id": "daslkndlkans",
    "reference_id": "askdnl"
  }
}
```

| Field          | Type              | Required | Description                                                            |
| -------------- | ----------------- | -------- | ---------------------------------------------------------------------- |
| `text`         | string            | Yes      | Transcript text for this chunk                                         |
| `final`        | boolean           | Yes      | `false` for interim events                                             |
| `start`        | string (ISO 8601) | Yes      | Chunk start timestamp (UTC)                                            |
| `duration`     | integer           | Yes      | Chunk duration in milliseconds                                         |
| `utterance_id` | string            | Yes      | Links this event to the corresponding final transcript and translation |
| `reference_id` | string            | No       | Echoed client reference ID                                             |

### Translation

```json
{
  "translate": {
    "text": "¿Hola, cómo estás?",
    "final": true,
    "utterance_id": "daslkndlkans",
    "reference_id": "askdnl"
  }
}
```

| Field          | Type    | Required | Description                                      |
| -------------- | ------- | -------- | ------------------------------------------------ |
| `text`         | string  | Yes      | Translated text                                  |
| `final`        | boolean | Yes      | `false` for interim events                       |
| `utterance_id` | string  | Yes      | Links this event to the corresponding transcript |
| `reference_id` | string  | No       | Echoed client reference ID                       |

### Error

Emitted when the server encounters an issue processing the request.

```json
{
  "error": {
    "code": 400,
    "reason": "Bad request",
    "description": "Actual description of the error",
    "reference_id": "askdnl"
  }
}
```

| Field          | Type    | Required | Description                |
| -------------- | ------- | -------- | -------------------------- |
| `code`         | integer | Yes      | HTTP-style status code     |
| `reason`       | string  | Yes      | Short reason string        |
| `description`  | string  | Yes      | Detailed error description |
| `reference_id` | string  | No       | Echoed client reference ID |

***

### Error Codes

| Code  | Reason                | Causes                                                                               |
| ----- | --------------------- | ------------------------------------------------------------------------------------ |
| `400` | Bad Request           | Malformed request · Invalid audio format · Incorrect target language · Invalid voice |
| `401` | Unauthorized          | Missing API key · Invalid API key · Expired API key                                  |
| `402` | Payment Required      | Balance exhausted · Subscription expired                                             |
| `429` | Too Many Requests     | Rate limit exceeded · Max concurrent connections exceeded                            |
| `500` | Internal Server Error | Server failed to process the request                                                 |

<br />

<br />

### List Supported Languages

Returns the languages currently available for voice translation. Use the returned `language_code` values as `source_language` / `target_language` when starting a session.

```curl
GET /voice-translation/languages
```

**Authentication**: `Authorization: API-Key {api-key}`

**Response 200**

```json
{
  "success": true,
  "code": 0,
  "data": [
    {
      "name": "English (United States)",
      "language_code": "en-US"
    },
    {
      "name": "French",
      "language_code": "fr-FR"
    },
    {
      "name": "German",
      "language_code": "de-DE"
    }
  ],
  "req_id": "..."
}
```

Language codes are BCP-47 (e.g. en-US, not en). The list is dynamic — fetch at session start rather than hard-coding.

<br />

<br />

***

## Notes & Known Limitations

* **Audio formats** — Only `pcm_s16le and Opus` are supported.
* **Sample rates & channels** — Only `16000 Hz` mono confirmed.

<br />