Kataleptic
§ Docs Updated 2026-09-25

Realtime voice

Speech-to-speech over WebSocket, wire-compatible with the OpenAI Realtime API. Point an unmodified OpenAI realtime client at api.kataleptic.com and it works — GA dialect by default, beta dialect auto-detected. Three tiers behind one endpoint, selected by the model id.

Endpointwss://api.kataleptic.com/v1/realtime?model=<id>
AuthBearer dg_... · ?token= · subprotocol
ProtocolOpenAI Realtime (GA; beta auto-detected)
AudioPCM16 @ 16/24 kHz · G.711 on Azure tiers

GPT-Live 1: full-duplex voice

gpt-live-1 listens and speaks continuously. Connect to wss://api.kataleptic.com/v1/live/sessions and put the model in session.start. It uses its own GPT-Live protocol; the /v1/realtime examples below apply to the other voice models.

Authenticate the WebSocket upgrade with Authorization: Bearer <key>. For a web application, connect the browser to your own authenticated server and relay the audio from there to Kataleptic. Store the Kataleptic service key only on that server; never embed it in browser code or send it to site visitors. This endpoint does not issue short-lived browser credentials.

{
  "type": "session.start",
  "session": {
    "model": "gpt-live-1",
    "instructions": "You are a helpful receptionist. Keep answers brief.",
    "audio": {
      "format": {"type": "audio/pcm", "rate": 24000},
      "output": {"voice": "cedar"}
    }
  }
}

Wait for session.started and verify its audio settings. Stream base64 mono PCM16 audio using session.input_audio.append with an audio field. Input must flow before the model will speak: send silence at real-time pace while waiting for a greeting, then replace it with microphone audio.

{"type":"session.input_audio.append","audio":"<base64 PCM16>"}
{"type":"session.commentary.append","delegation_id":null,"content":"Greet the caller and ask how you can help."}

Play the delta from session.output_audio.delta using the negotiated format. Audio arrives continuously, including silence, with no audio-done event. Transcript fragments arrive as session.input_transcript.delta and session.output_transcript.delta, carrying delta, start_ms and end_ms. Input and output can overlap.

Voices: marin (default), cedar, alloy, coral, shimmer, verse, ash, sage, ballad and echo. G.711 μ-law is supported with {"type":"audio/pcmu","rate":8000}. Specify the rate whenever you provide a format.

Optional delegation and tools

Omit delegation for voice-only use. To run function tools, add {"type":"responses","responses":{"model":"gpt-5.4-mini","tools":[...]}} as session.delegation. The delegate must be an eligible Azure OpenAI chat model. Only function tools are supported; hosted web search, image generation and other hosted tools are refused.

Delegated responses have a 1,024-token output ceiling, including reasoning. You may set delegation.responses.max_output_tokens to an integer from 16 to 1024; omitting it from an update preserves the accepted value. Very low limits can prevent a tool call from completing.

Delegated events arrive inside response.event. For a response.output_item.done containing a function call, execute the function on your server, send its result using response.item.create with a function_call_output item, then send bare {"type":"response.create"}. That continuation is separately billed. session.update can change delegation settings only; wait for session.updated before relying on the change.

Billing and closing

The session price is $0.09375/min, billed per second of session clock. The clock starts with the first input audio, includes subsequent silence, and can advance faster if you send audio faster than real time. Pre-audio idle time is not billed. Delegated responses additionally bill their input, cached-input and output tokens at the selected chat model's catalogue rates.

This price uses a provisional Azure cost basis: Azure has not yet published a GPT-Live 1 meter. The cost basis is $0.075/min and the standard 1.25× margin gives the price above. Check the catalogue for current prices.

Send {"type":"session.close"} to end a session. The session.closed event reports usage.seconds. The gateway uses an estimate of the same clock if Azure never supplies that report. Sessions last at most one hour and may end sooner under the gateway's configured limit.

An insufficient_reservation error rejects the specific event while leaving the connection open: pace audio in real time, or wait and retry a delegated request. Exhausted credit closes the session with code 4402; a duration limit uses 4408, and upstream capacity refusals use 1013. WebRTC and session-attach endpoints are not exposed.

Quickstart

Open a WebSocket, configure the session, send a bare response.create to make the agent speak first, then stream microphone audio in and play audio deltas out. That is the whole loop.

// Browser / Cloudflare Workers — auth via subprotocol
const ws = new WebSocket(
  "wss://api.kataleptic.com/v1/realtime?model=kataleptic-realtime",
  ["realtime", "openai-insecure-api-key." + KATALEPTIC_API_KEY],
);

ws.onopen = () => {
  ws.send(JSON.stringify({
    type: "session.update",
    session: {
      instructions: "You are the booking agent for a small hotel. " +
                    "Open by greeting the caller and asking how you can help.",
      turn_detection: { type: "server_vad", silence_duration_ms: 400 },
    },
  }));
  // Bare response.create → the agent speaks its opening line.
  ws.send(JSON.stringify({ type: "response.create" }));
};

ws.onmessage = (e) => {
  const ev = JSON.parse(e.data);
  if (ev.type === "response.output_audio.delta") {
    playPcm16(atob(ev.delta));            // PCM16 mono @ 24 kHz
  } else if (ev.type === "response.output_audio_transcript.done") {
    console.log("agent said:", ev.transcript);
  }
};

// Stream microphone audio as base64-encoded PCM16:
function sendChunk(base64Pcm16) {
  ws.send(JSON.stringify({ type: "input_audio_buffer.append", audio: base64Pcm16 }));
}

Already on the OpenAI SDK? Unmodified OpenAI realtime clients work as-is — change the host to api.kataleptic.com and keep your code. We speak the GA dialect by default and switch to the beta dialect automatically when your client sends the OpenAI-Beta: realtime=v1 header or the openai-beta.realtime-v1 subprotocol.

Authentication

Three ways to present your dg_… key, in order of preference:

  • Header — Authorization: Bearer dg_…. Use this from servers.
  • Query parameter — ?token=dg_… appended to the WebSocket URL, for clients that cannot set headers.
  • Subprotocol — openai-insecure-api-key.dg_… in the WebSocket subprotocol list, the same convention OpenAI uses for browser and Workers clients. As the name says: only use this with short-lived keys you are comfortable exposing to the client.

The three tiers

One endpoint, three engines. The model id in ?model= selects the engine; everything else about the protocol stays the same.

Model idEngineFirst audioTranscriptsTypical price
kataleptic-realtime Cascade: Whisper STT → chat model → Piper TTS ~250 ms Exact ≈$0.0133/min
kataleptic-realtime-hd Azure Voice Live ~1.2 s Exact ≈$0.03/min
gpt-realtime-2.1 Native speech-to-speech ~1.0 s Model approximation ≈$0.07/min
gpt-realtime-2.1-mini Native speech-to-speech ~1.0 s Model approximation ≈$0.02/min
gpt-realtime-2 Native speech-to-speech (previous) ~1.0 s Model approximation ≈$0.07/min

Retiring 2026-10-23: kataleptic-realtime (the cascade), its Piper voices, and selecting it with ?model=<chat model> or with the compatibility names gpt-realtime, gpt-4o-realtime* and gpt-audio*. From that date those sessions are refused with a model_retired error, and a session that names no model gets gpt-realtime-2.1-mini. Move to kataleptic-realtime-hd or gpt-realtime-2.1 / gpt-realtime-2.1-mini now — see retiring models.

kataleptic-realtime — the default

A cascade on our own fleet: streaming Whisper speech-to-text, a catalogue chat model in the middle, Piper text-to-speech on the way out. The brain is swappable per session — pass any catalogue chat model id in ?model= (default mistral-nemo-12b) and the cascade uses it. Ten languages are auto-detected per utterance — EN, DE, FR, ES, NL, SV, DA, IT, FI, RU — and the TTS voice follows the detected language. Server-side VAD with barge-in; speech recognition is noise-gated (Silero VAD plus no-speech and language-probability thresholds), so breathing and background noise do not become turns.

kataleptic-realtime-hd — premium voices

The same WebSocket, served by Azure Voice Live: 600+ HD neural voices, deep noise suppression, echo cancellation, and semantic turn detection (the model judges whether the caller is done, not just the silence timer). Exact transcripts.

gpt-realtime-2.1 — native speech-to-speech

No cascade — one model hears audio and speaks audio. Best prosody and expressiveness of the tiers; it responds to tone, hesitation, and emphasis, not just words. gpt-realtime-2.1 is the current model (better instruction following and turn-taking than gpt-realtime-2, same token price, plus image input); gpt-realtime-2.1-mini is the same protocol and voices at roughly a third the audio-token cost, sized for call volume. gpt-realtime-2 stays available unchanged. The trade-off is real and listed in caveats: transcripts are the model's own approximation of what was said.

session.update reference

Send session.update as your first message to configure the conversation. The supported subset:

FieldTypeWhat it does
instructionsstringThe system prompt. Persona, opening line, guardrails.
voicestringVoice selection. On the default tier the voice follows the detected language; on the Azure tiers pick from their voice catalogues.
turn_detection.type"server_vad"Server-side voice activity detection. The server decides when the caller's turn ends.
turn_detection.thresholdnumberVAD sensitivity. Higher = needs louder/clearer speech to open a turn.
turn_detection.prefix_padding_msnumberAudio retained from before speech onset, so first syllables are not clipped.
turn_detection.silence_duration_msnumberTrailing silence that ends the turn. Lower = snappier, more interruptions.
turn_detection.create_responsebooleanAuto-respond when a turn ends. Set false to drive responses yourself with response.create.
turn_detection.interrupt_responsebooleanBarge-in: caller speech cancels the agent's in-flight reply.

Protocol subset

Client events we accept, on every tier:

  • session.update — configure instructions, voice, turn detection (see above).
  • input_audio_buffer.append / .commit / .clear — stream caller audio; commit manually if you run your own VAD.
  • conversation.item.create / .delete / .truncate — edit conversation history, including previous_item_id placement and root insertion.
  • response.create / response.cancel — request or cancel an agent reply.

Audio is PCM16 at 16 or 24 kHz in both directions on all tiers. The two Azure tiers additionally accept G.711 for telephony — see below.

Transcripts & call logging

Both directions of the conversation arrive as text events, which is all you need to build a call log:

  • Caller side — conversation.item.input_audio_transcription.completed fires once per caller utterance with the final transcript.
  • Agent side — response.output_audio_transcript.delta streams the agent's words as it speaks; response.output_audio_transcript.done carries the full utterance.

Native tiers only (gpt-realtime-2*): caller transcripts are on by default (we enable input_audio_transcription with whisper-1 for you; override or disable it in session.update). The agent-side transcript is the model's approximation of its own speech, not an exact STT transcript; if your call logs have compliance weight, use the standard or HD tier. On the standard tier, transcription events also carry language and language_probability fields.

Greeting pattern

Phone agents should speak first. Put the opening line in instructions, then send a bare response.create — no conversation items needed:

{ "type": "session.update",
  "session": { "instructions": "Greet the caller: 'Grüß Gott, Hotel Sacher reception.' Then assist." } }

{ "type": "response.create" }

The agent speaks the greeting per its instructions, and the normal turn-taking loop begins from there.

Telephony / G.711

SIP trunks and most PSTN gateways hand you G.711. On kataleptic-realtime-hd and the gpt-realtime-2* tiers you can pass it straight through without transcoding:

  • Beta-dialect flat fields: "input_audio_format": "g711_ulaw" (or "g711_alaw"), same for output.
  • GA-dialect format objects: {"type": "audio/pcmu"} / {"type": "audio/pcma"}.

The default kataleptic-realtime tier is PCM16-only — transcode at your media gateway if you bridge it to a trunk.

Caveats per tier

kataleptic-realtime

  • Cascade voices are functional, not studio-grade — if voice quality is the product, use HD.
  • One voice per language; the voice field has limited effect because the voice follows the detected language.

kataleptic-realtime-hd

  • First audio ~1.2 s — noticeably slower to open than the default tier's ~250 ms.

gpt-realtime-2 / 2.1 / 2.1-mini

  • Agent transcripts are model approximations, not exact STT output (caller transcripts use whisper-1, on by default).

Function calling

All three tiers support OpenAI Realtime function calling. Define tools in session.update (flat realtime shape: {"type": "function", "name", "description", "parameters"}); when the model decides to call one you receive response.function_call_arguments.delta events, a final response.function_call_arguments.done with the JSON arguments, and a function_call item in response.done. Send the result back as a conversation.item.create with {"type": "function_call_output", "call_id", "output"} followed by response.create.

On the standard tier the cascade brain executes the tool call; small models occasionally write a call as prose instead of invoking it — the server strips text that exactly matches a defined tool-call pattern from the spoken audio, so the agent never says "end_call()" aloud.

Session limits & lifecycle

  • Max session duration: 60 minutes. One minute before the cutoff the server emits a vendor-extension event {"type": "session.expiring", "reason": "max_session_duration", "expires_in_seconds": …} so bridges can reconnect gracefully. Clients that ignore unknown events lose nothing.
  • Idle timeout: 5 minutes without any WebSocket message (continuous audio streaming counts as activity).
  • Server deploys can terminate live sessions; production bridges should reconnect on unexpected close and re-send session.update.

Voice catalog per tier

  • kataleptic-realtime — voice follows the detected caller language automatically across all ten languages. Before the first caller utterance, the initial voice seeds from input_audio_transcription.language when set, or from the language of your instructions — so instruction-driven greetings come out in the right voice with zero configuration. The language field is a seed and STT-accuracy hint, not a cage: once real speech arrives, per-utterance detection overrides it — in the reported language field, the voice, and what the model is told — even when it contradicts the seed. To pin a voice explicitly, pass a Piper id as voice: en_US-lessac-medium, de_DE-thorsten-medium, fr_FR-siwis-medium, es_ES-sharvard-medium, nl_NL-mls-medium, sv_SE-nst-medium, da_DK-talesyntese-medium, it_IT-paola-medium, fi_FI-harri-medium, ru_RU-irina-medium. OpenAI voice names are accepted and ignored in favor of language-matching.
  • kataleptic-realtime-hd — any Azure neural voice name passes through (e.g. de-DE-SeraphinaMultilingualNeural, 600+ voices); OpenAI voice names (alloy, marin, …) map to Azure multilingual voices; Piper ids map to the closest Azure voice. Default is a multilingual voice, so language-follow works with no configuration.
  • gpt-realtime-2 / 2.1 / 2.1-mini — OpenAI voices only (marin, cedar, alloy, …); non-OpenAI names coerce to marin. Voices are natively multilingual.

The full machine-readable catalog (including the live per-language Piper map) is served at GET /v1/realtime/voices — no auth required. One namespace, three tiers: OpenAI names work everywhere; engine-native names (Piper ids, Azure voice names) work on their own tier and degrade gracefully elsewhere. Unknown transcription.model values return an error event; supported values are whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe (→ turbo) and whisper-large-v3 (full model, ~+110 ms, better on noisy audio).

Choosing the cascade brain

  • mistral-nemo-12b (default) — fastest replies (~0.3 s first audio), solid small-talk and form-filling; weaker at multi-step reasoning (dates, arithmetic) and occasionally imperfect language adherence on long prompts.
  • llama-3.3-70b — strong reasoning and reliable multilingual replies at ~1–1.4 s first audio. Recommended for production receptionists that must reason about schedules.
  • gpt-5.4-mini / other catalogue models — pick any chat model via ?model=; latency is dominated by that model's time-to-first-token.

Streaming transcription

The same WebSocket also runs transcription-only sessions: you stream audio in and get text back, with no spoken reply. Open wss://api.kataleptic.com/v1/realtime?intent=transcription, optionally with &model=<id> (default gpt-live-transcribe). It speaks the GA OpenAI Realtime transcription-session protocol; beta transcription_session.update clients are translated.

Model idPrice
gpt-live-transcribe (default)$0.02125 / min
gpt-realtime-whisper$0.02125 / min
  1. Send session.update with session.type: "transcription" and the model in audio.input.transcription.model, optionally with language and prompt. If you also passed ?model=, the two must match, or you get a model_mismatch error.
  2. Stream audio with input_audio_buffer.append — base64 PCM16, mono, 24 kHz.
  3. Send input_audio_buffer.commit to close each utterance. These models have no turn detection: a commit is the only thing that ends an item, and any turn_detection you send is set to null.
  4. conversation.item.input_audio_transcription.delta events stream text while you are still sending audio; conversation.item.input_audio_transcription.completed carries the final transcript after each commit.
import asyncio, base64, json, os, wave
import websockets

URL = "wss://api.kataleptic.com/v1/realtime?intent=transcription"

async def transcribe(path, model="gpt-live-transcribe", language=None):
    async with websockets.connect(
        URL,
        additional_headers={"Authorization": f"Bearer {os.environ['KATALEPTIC_API_KEY']}"},
    ) as ws:
        await ws.recv()                                   # session.created
        tr = {"model": model, **({"language": language} if language else {})}
        await ws.send(json.dumps({"type": "session.update", "session": {
            "type": "transcription",
            "audio": {"input": {"format": {"type": "audio/pcm", "rate": 24000},
                                "transcription": tr}},
        }}))
        with wave.open(path) as w:                        # PCM16 mono 24 kHz
            pcm = w.readframes(w.getnframes())
        for i in range(0, len(pcm), 4800):                # 100 ms chunks
            await ws.send(json.dumps({"type": "input_audio_buffer.append",
                                      "audio": base64.b64encode(pcm[i:i + 4800]).decode()}))
            await asyncio.sleep(0.1)
        await ws.send(json.dumps({"type": "input_audio_buffer.commit"}))
        async for raw in ws:
            ev = json.loads(raw)
            if ev["type"] == "conversation.item.input_audio_transcription.delta":
                print(ev["delta"], end="", flush=True)
            elif ev["type"] == "conversation.item.input_audio_transcription.completed":
                return ev["transcript"]
            elif ev["type"] == "error":
                raise RuntimeError(ev["error"])

asyncio.run(transcribe("call.wav", language="en"))

Billing: each buffer bills the seconds of audio you appended to it, rounded up, silence included, when you commit or clear it. Many tiny commits therefore round up many times. For recorded audio the file endpoint, POST /v1/audio/transcriptions, is far cheaper — see transcription.

Replacing /v1/listen: the two streaming models on /v1/listen retire on 2026-10-23, and the endpoint with them. Move to a transcription session now. It costs more than those streams did ($0.0033/min).

Pricing & billing

  • kataleptic-realtime — $0.0033/min audio in + $0.01/min audio out, plus the chat model's tokens at its catalogue rate. ≈$0.0133/min all-in with the default brain.
  • kataleptic-realtime-hd — billed per token at Azure Voice Live rates with a service margin; ≈$0.03/min typical.
  • gpt-realtime-2.1 (and gpt-realtime-2) — billed per text + audio token; ≈$0.07/min typical.
  • gpt-realtime-2.1-mini — same, at roughly a third the audio rate; ≈$0.02/min typical.
  • gpt-live-transcribe / gpt-realtime-whisper — streaming transcription; $0.02125 per minute of audio, billed per buffer and rounded up to the second.

Usage shows up on your key under the model ids kataleptic-realtime, kataleptic-realtime-hd and the gpt-realtime-2* ids — same GET /v1/auth/key surface as everything else.

Voice minutes are cheap to try: the $2 of free signup credit lets you try the available voice models. Get a key and say hello to it.