Realtime voice
Speech-to-speech over WebSocket, wire-compatible with the OpenAI
Realtime API. Point an unmodified OpenAI realtime client at
api.kataleptic.com and it works — GA dialect by default,
beta dialect auto-detected. Three tiers behind one endpoint,
selected by the model id.
wss://api.kataleptic.com/v1/realtime?model=<id>Bearer dg_... · ?token= · subprotocolOpenAI Realtime (GA; beta auto-detected)PCM16 @ 16/24 kHz · G.711 on Azure tiersGPT-Live 1: full-duplex voice
gpt-live-1 listens and speaks continuously. Connect to
wss://api.kataleptic.com/v1/live/sessions and put the model in
session.start. It uses its own GPT-Live protocol; the
/v1/realtime examples below apply to the other voice models.
Authenticate the WebSocket upgrade with Authorization: Bearer <key>.
For a web application, connect the browser to your own authenticated server and
relay the audio from there to Kataleptic. Store the Kataleptic service key only
on that server; never embed it in browser code or send it to site visitors.
This endpoint does not issue short-lived browser credentials.
{
"type": "session.start",
"session": {
"model": "gpt-live-1",
"instructions": "You are a helpful receptionist. Keep answers brief.",
"audio": {
"format": {"type": "audio/pcm", "rate": 24000},
"output": {"voice": "cedar"}
}
}
}
Wait for session.started and verify its audio settings. Stream base64
mono PCM16 audio using session.input_audio.append with an
audio field. Input must flow before the model will speak: send silence
at real-time pace while waiting for a greeting, then replace it with microphone audio.
{"type":"session.input_audio.append","audio":"<base64 PCM16>"}
{"type":"session.commentary.append","delegation_id":null,"content":"Greet the caller and ask how you can help."}
Play the delta from session.output_audio.delta using the
negotiated format. Audio arrives continuously, including silence, with no audio-done
event. Transcript fragments arrive as session.input_transcript.delta and
session.output_transcript.delta, carrying delta,
start_ms and end_ms. Input and output can overlap.
Voices: marin (default), cedar, alloy,
coral, shimmer, verse, ash,
sage, ballad and echo.
G.711 μ-law is supported with {"type":"audio/pcmu","rate":8000}.
Specify the rate whenever you provide a format.
Optional delegation and tools
Omit delegation for voice-only use. To run function tools, add
{"type":"responses","responses":{"model":"gpt-5.4-mini","tools":[...]}}
as session.delegation. The delegate must be an eligible Azure OpenAI
chat model. Only function tools are supported; hosted web search, image generation
and other hosted tools are refused.
Delegated responses have a 1,024-token output ceiling, including reasoning.
You may set delegation.responses.max_output_tokens to an integer
from 16 to 1024; omitting it from an update preserves the accepted value.
Very low limits can prevent a tool call from completing.
Delegated events arrive inside response.event. For a
response.output_item.done containing a function call, execute the
function on your server, send its result using response.item.create
with a function_call_output item, then send bare
{"type":"response.create"}. That continuation is separately billed.
session.update can change delegation settings only; wait for
session.updated before relying on the change.
Billing and closing
The session price is $0.09375/min, billed per second of session clock. The clock starts with the first input audio, includes subsequent silence, and can advance faster if you send audio faster than real time. Pre-audio idle time is not billed. Delegated responses additionally bill their input, cached-input and output tokens at the selected chat model's catalogue rates.
This price uses a provisional Azure cost basis: Azure has not yet published a GPT-Live 1 meter. The cost basis is $0.075/min and the standard 1.25× margin gives the price above. Check the catalogue for current prices.
Send {"type":"session.close"} to end a session. The
session.closed event reports usage.seconds. The gateway
uses an estimate of the same clock if Azure never supplies that report.
Sessions last at most one hour and may end sooner under the gateway's configured limit.
An insufficient_reservation error rejects the specific event while
leaving the connection open: pace audio in real time, or wait and retry a delegated
request. Exhausted credit closes the session with code 4402; a duration
limit uses 4408, and upstream capacity refusals use 1013.
WebRTC and session-attach endpoints are not exposed.
Quickstart
Open a WebSocket, configure the session, send a bare
response.create to make the agent speak first, then
stream microphone audio in and play audio deltas out. That is the
whole loop.
// Browser / Cloudflare Workers — auth via subprotocol
const ws = new WebSocket(
"wss://api.kataleptic.com/v1/realtime?model=kataleptic-realtime",
["realtime", "openai-insecure-api-key." + KATALEPTIC_API_KEY],
);
ws.onopen = () => {
ws.send(JSON.stringify({
type: "session.update",
session: {
instructions: "You are the booking agent for a small hotel. " +
"Open by greeting the caller and asking how you can help.",
turn_detection: { type: "server_vad", silence_duration_ms: 400 },
},
}));
// Bare response.create → the agent speaks its opening line.
ws.send(JSON.stringify({ type: "response.create" }));
};
ws.onmessage = (e) => {
const ev = JSON.parse(e.data);
if (ev.type === "response.output_audio.delta") {
playPcm16(atob(ev.delta)); // PCM16 mono @ 24 kHz
} else if (ev.type === "response.output_audio_transcript.done") {
console.log("agent said:", ev.transcript);
}
};
// Stream microphone audio as base64-encoded PCM16:
function sendChunk(base64Pcm16) {
ws.send(JSON.stringify({ type: "input_audio_buffer.append", audio: base64Pcm16 }));
}
import asyncio, base64, json, os
import websockets
URL = "wss://api.kataleptic.com/v1/realtime?model=kataleptic-realtime"
async def main():
async with websockets.connect(
URL,
additional_headers={
"Authorization": f"Bearer {os.environ['KATALEPTIC_API_KEY']}",
},
) as ws:
await ws.send(json.dumps({
"type": "session.update",
"session": {
"instructions": (
"You are the booking agent for a small hotel. "
"Open by greeting the caller and asking how you can help."
),
"turn_detection": {"type": "server_vad", "silence_duration_ms": 400},
},
}))
# Bare response.create → the agent speaks its opening line.
await ws.send(json.dumps({"type": "response.create"}))
async for raw in ws:
ev = json.loads(raw)
if ev["type"] == "response.output_audio.delta":
pcm16 = base64.b64decode(ev["delta"]) # PCM16 mono @ 24 kHz
elif ev["type"] == "response.output_audio_transcript.done":
print("agent said:", ev["transcript"])
asyncio.run(main())
Already on the OpenAI SDK? Unmodified OpenAI realtime
clients work as-is — change the host to
api.kataleptic.com and keep your code. We speak the GA
dialect by default and switch to the beta dialect automatically when
your client sends the OpenAI-Beta: realtime=v1 header
or the openai-beta.realtime-v1 subprotocol.
Authentication
Three ways to present your dg_… key, in order of preference:
-
Header —
Authorization: Bearer dg_…. Use this from servers. -
Query parameter —
?token=dg_…appended to the WebSocket URL, for clients that cannot set headers. -
Subprotocol —
openai-insecure-api-key.dg_…in the WebSocket subprotocol list, the same convention OpenAI uses for browser and Workers clients. As the name says: only use this with short-lived keys you are comfortable exposing to the client.
The three tiers
One endpoint, three engines. The model id in
?model= selects the engine; everything else about the
protocol stays the same.
| Model id | Engine | First audio | Transcripts | Typical price |
|---|---|---|---|---|
kataleptic-realtime |
Cascade: Whisper STT → chat model → Piper TTS | ~250 ms | Exact | ≈$0.0133/min |
kataleptic-realtime-hd |
Azure Voice Live | ~1.2 s | Exact | ≈$0.03/min |
gpt-realtime-2.1 |
Native speech-to-speech | ~1.0 s | Model approximation | ≈$0.07/min |
gpt-realtime-2.1-mini |
Native speech-to-speech | ~1.0 s | Model approximation | ≈$0.02/min |
gpt-realtime-2 |
Native speech-to-speech (previous) | ~1.0 s | Model approximation | ≈$0.07/min |
Retiring 2026-10-23: kataleptic-realtime (the
cascade), its Piper voices, and selecting it with
?model=<chat model> or with the
compatibility names gpt-realtime,
gpt-4o-realtime* and gpt-audio*. From that
date those sessions are refused with a model_retired
error, and a session that names no model gets
gpt-realtime-2.1-mini. Move to
kataleptic-realtime-hd or
gpt-realtime-2.1 / gpt-realtime-2.1-mini
now — see retiring models.
kataleptic-realtime — the default
A cascade on our own fleet: streaming Whisper
speech-to-text, a catalogue chat model in the middle, Piper
text-to-speech on the way out. The brain is swappable per session —
pass any catalogue chat model id in ?model=
(default mistral-nemo-12b) and the cascade uses it.
Ten languages are auto-detected per utterance — EN, DE, FR, ES, NL,
SV, DA, IT, FI, RU — and the TTS voice follows the detected
language. Server-side VAD with barge-in; speech recognition is
noise-gated (Silero VAD plus no-speech and language-probability
thresholds), so breathing and background noise do not become turns.
kataleptic-realtime-hd — premium voices
The same WebSocket, served by Azure Voice Live: 600+ HD neural voices, deep noise suppression, echo cancellation, and semantic turn detection (the model judges whether the caller is done, not just the silence timer). Exact transcripts.
gpt-realtime-2.1 — native speech-to-speech
No cascade — one model hears audio and speaks audio. Best prosody
and expressiveness of the tiers; it responds to tone, hesitation,
and emphasis, not just words. gpt-realtime-2.1 is the
current model (better instruction following and turn-taking than
gpt-realtime-2, same token price, plus image input);
gpt-realtime-2.1-mini is the same protocol and voices
at roughly a third the audio-token cost, sized for call volume.
gpt-realtime-2 stays available unchanged. The
trade-off is real and listed in
caveats: transcripts are the
model's own approximation of what was said.
session.update reference
Send session.update as your first message to configure
the conversation. The supported subset:
| Field | Type | What it does |
|---|---|---|
instructions | string | The system prompt. Persona, opening line, guardrails. |
voice | string | Voice selection. On the default tier the voice follows the detected language; on the Azure tiers pick from their voice catalogues. |
turn_detection.type | "server_vad" | Server-side voice activity detection. The server decides when the caller's turn ends. |
turn_detection.threshold | number | VAD sensitivity. Higher = needs louder/clearer speech to open a turn. |
turn_detection.prefix_padding_ms | number | Audio retained from before speech onset, so first syllables are not clipped. |
turn_detection.silence_duration_ms | number | Trailing silence that ends the turn. Lower = snappier, more interruptions. |
turn_detection.create_response | boolean | Auto-respond when a turn ends. Set false to drive responses yourself with response.create. |
turn_detection.interrupt_response | boolean | Barge-in: caller speech cancels the agent's in-flight reply. |
Protocol subset
Client events we accept, on every tier:
session.update— configure instructions, voice, turn detection (see above).input_audio_buffer.append/.commit/.clear— stream caller audio; commit manually if you run your own VAD.conversation.item.create/.delete/.truncate— edit conversation history, includingprevious_item_idplacement and root insertion.response.create/response.cancel— request or cancel an agent reply.
Audio is PCM16 at 16 or 24 kHz in both directions on all tiers. The two Azure tiers additionally accept G.711 for telephony — see below.
Transcripts & call logging
Both directions of the conversation arrive as text events, which is all you need to build a call log:
-
Caller side —
conversation.item.input_audio_transcription.completedfires once per caller utterance with the final transcript. -
Agent side —
response.output_audio_transcript.deltastreams the agent's words as it speaks;response.output_audio_transcript.donecarries the full utterance.
Native tiers only (gpt-realtime-2*): caller transcripts are on by default
(we enable input_audio_transcription with
whisper-1 for you; override or disable it in
session.update). The agent-side transcript is
the model's approximation of its own speech, not an exact STT
transcript; if your call logs have compliance weight, use the
standard or HD tier. On the standard tier, transcription events
also carry language and
language_probability fields.
Greeting pattern
Phone agents should speak first. Put the opening line in
instructions, then send a bare
response.create — no conversation items needed:
{ "type": "session.update",
"session": { "instructions": "Greet the caller: 'Grüß Gott, Hotel Sacher reception.' Then assist." } }
{ "type": "response.create" }
The agent speaks the greeting per its instructions, and the normal turn-taking loop begins from there.
Telephony / G.711
SIP trunks and most PSTN gateways hand you G.711. On
kataleptic-realtime-hd and the
gpt-realtime-2* tiers
you can pass it straight through without transcoding:
- Beta-dialect flat fields:
"input_audio_format": "g711_ulaw"(or"g711_alaw"), same for output. - GA-dialect format objects:
{"type": "audio/pcmu"}/{"type": "audio/pcma"}.
The default kataleptic-realtime tier is PCM16-only —
transcode at your media gateway if you bridge it to a trunk.
Caveats per tier
kataleptic-realtime
- Cascade voices are functional, not studio-grade — if voice quality is the product, use HD.
- One voice per language; the
voicefield has limited effect because the voice follows the detected language.
kataleptic-realtime-hd
- First audio ~1.2 s — noticeably slower to open than the default tier's ~250 ms.
gpt-realtime-2 / 2.1 / 2.1-mini
- Agent transcripts are model approximations, not exact STT output (caller transcripts use whisper-1, on by default).
Function calling
All three tiers support OpenAI Realtime function calling. Define tools in session.update (flat realtime shape: {"type": "function", "name", "description", "parameters"}); when the model decides to call one you receive response.function_call_arguments.delta events, a final response.function_call_arguments.done with the JSON arguments, and a function_call item in response.done. Send the result back as a conversation.item.create with {"type": "function_call_output", "call_id", "output"} followed by response.create.
On the standard tier the cascade brain executes the tool call; small models occasionally write a call as prose instead of invoking it — the server strips text that exactly matches a defined tool-call pattern from the spoken audio, so the agent never says "end_call()" aloud.
Session limits & lifecycle
- Max session duration: 60 minutes. One minute before the cutoff the server emits a vendor-extension event
{"type": "session.expiring", "reason": "max_session_duration", "expires_in_seconds": …}so bridges can reconnect gracefully. Clients that ignore unknown events lose nothing. - Idle timeout: 5 minutes without any WebSocket message (continuous audio streaming counts as activity).
- Server deploys can terminate live sessions; production bridges should reconnect on unexpected close and re-send
session.update.
Voice catalog per tier
- kataleptic-realtime — voice follows the detected caller language automatically across all ten languages. Before the first caller utterance, the initial voice seeds from
input_audio_transcription.languagewhen set, or from the language of yourinstructions— so instruction-driven greetings come out in the right voice with zero configuration. The language field is a seed and STT-accuracy hint, not a cage: once real speech arrives, per-utterance detection overrides it — in the reportedlanguagefield, the voice, and what the model is told — even when it contradicts the seed. To pin a voice explicitly, pass a Piper id asvoice:en_US-lessac-medium,de_DE-thorsten-medium,fr_FR-siwis-medium,es_ES-sharvard-medium,nl_NL-mls-medium,sv_SE-nst-medium,da_DK-talesyntese-medium,it_IT-paola-medium,fi_FI-harri-medium,ru_RU-irina-medium. OpenAI voice names are accepted and ignored in favor of language-matching. - kataleptic-realtime-hd — any Azure neural voice name passes through (e.g.
de-DE-SeraphinaMultilingualNeural, 600+ voices); OpenAI voice names (alloy,marin, …) map to Azure multilingual voices; Piper ids map to the closest Azure voice. Default is a multilingual voice, so language-follow works with no configuration. - gpt-realtime-2 / 2.1 / 2.1-mini — OpenAI voices only (
marin,cedar,alloy, …); non-OpenAI names coerce tomarin. Voices are natively multilingual.
The full machine-readable catalog (including the live per-language Piper map) is served at GET /v1/realtime/voices — no auth required. One namespace, three tiers: OpenAI names work everywhere; engine-native names (Piper ids, Azure voice names) work on their own tier and degrade gracefully elsewhere. Unknown transcription.model values return an error event; supported values are whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe (→ turbo) and whisper-large-v3 (full model, ~+110 ms, better on noisy audio).
Choosing the cascade brain
mistral-nemo-12b(default) — fastest replies (~0.3 s first audio), solid small-talk and form-filling; weaker at multi-step reasoning (dates, arithmetic) and occasionally imperfect language adherence on long prompts.llama-3.3-70b— strong reasoning and reliable multilingual replies at ~1–1.4 s first audio. Recommended for production receptionists that must reason about schedules.gpt-5.4-mini/ other catalogue models — pick any chat model via?model=; latency is dominated by that model's time-to-first-token.
Streaming transcription
The same WebSocket also runs transcription-only sessions: you
stream audio in and get text back, with no spoken reply. Open
wss://api.kataleptic.com/v1/realtime?intent=transcription,
optionally with &model=<id>
(default gpt-live-transcribe). It speaks the GA OpenAI
Realtime transcription-session protocol; beta
transcription_session.update clients are translated.
| Model id | Price |
|---|---|
gpt-live-transcribe (default) | $0.02125 / min |
gpt-realtime-whisper | $0.02125 / min |
- Send
session.updatewithsession.type: "transcription"and the model inaudio.input.transcription.model, optionally withlanguageandprompt. If you also passed?model=, the two must match, or you get amodel_mismatcherror. - Stream audio with
input_audio_buffer.append— base64 PCM16, mono, 24 kHz. - Send
input_audio_buffer.committo close each utterance. These models have no turn detection: a commit is the only thing that ends an item, and anyturn_detectionyou send is set tonull. conversation.item.input_audio_transcription.deltaevents stream text while you are still sending audio;conversation.item.input_audio_transcription.completedcarries the final transcript after each commit.
import asyncio, base64, json, os, wave
import websockets
URL = "wss://api.kataleptic.com/v1/realtime?intent=transcription"
async def transcribe(path, model="gpt-live-transcribe", language=None):
async with websockets.connect(
URL,
additional_headers={"Authorization": f"Bearer {os.environ['KATALEPTIC_API_KEY']}"},
) as ws:
await ws.recv() # session.created
tr = {"model": model, **({"language": language} if language else {})}
await ws.send(json.dumps({"type": "session.update", "session": {
"type": "transcription",
"audio": {"input": {"format": {"type": "audio/pcm", "rate": 24000},
"transcription": tr}},
}}))
with wave.open(path) as w: # PCM16 mono 24 kHz
pcm = w.readframes(w.getnframes())
for i in range(0, len(pcm), 4800): # 100 ms chunks
await ws.send(json.dumps({"type": "input_audio_buffer.append",
"audio": base64.b64encode(pcm[i:i + 4800]).decode()}))
await asyncio.sleep(0.1)
await ws.send(json.dumps({"type": "input_audio_buffer.commit"}))
async for raw in ws:
ev = json.loads(raw)
if ev["type"] == "conversation.item.input_audio_transcription.delta":
print(ev["delta"], end="", flush=True)
elif ev["type"] == "conversation.item.input_audio_transcription.completed":
return ev["transcript"]
elif ev["type"] == "error":
raise RuntimeError(ev["error"])
asyncio.run(transcribe("call.wav", language="en"))
Billing: each buffer bills the seconds of audio you
appended to it, rounded up, silence included, when you commit or
clear it. Many tiny commits therefore round up many times. For
recorded audio the file endpoint,
POST /v1/audio/transcriptions, is far cheaper — see
transcription.
Replacing /v1/listen: the two streaming
models on /v1/listen retire on 2026-10-23, and the
endpoint with them. Move to a transcription session now. It costs
more than those streams did ($0.0033/min).
Pricing & billing
-
kataleptic-realtime— $0.0033/min audio in + $0.01/min audio out, plus the chat model's tokens at its catalogue rate. ≈$0.0133/min all-in with the default brain. -
kataleptic-realtime-hd— billed per token at Azure Voice Live rates with a service margin; ≈$0.03/min typical. -
gpt-realtime-2.1(andgpt-realtime-2) — billed per text + audio token; ≈$0.07/min typical. -
gpt-realtime-2.1-mini— same, at roughly a third the audio rate; ≈$0.02/min typical. -
gpt-live-transcribe/gpt-realtime-whisper— streaming transcription; $0.02125 per minute of audio, billed per buffer and rounded up to the second.
Usage shows up on your key under the model ids
kataleptic-realtime, kataleptic-realtime-hd
and the gpt-realtime-2* ids — same
GET /v1/auth/key surface as everything else.
Voice minutes are cheap to try: the $2 of free signup credit lets you try the available voice models. Get a key and say hello to it.