
Speech and live video frames in, speech out over one socket, with 56 voices, tool calling during the conversation, and separate text and audio token rates.
Speech and live video frames in, speech out over one socket, with 56 voices, tool calling during the conversation, and separate text and audio token rates.
Session is configured with session.update events over the socket rather than request parameters. Text and audio tokens bill at separate rates.
Also known as Alibaba Cloud Qwen3.8 Omni Flash Realtime
qwen3-8-omni-flash-realtime/v1/realtimeqwen3.8-omni-flash-realtimealibaba/qwen3-8-omni-flash-realtimeLive pay-as-you-go rates from the EmpirioLabs catalog. You are billed only for what you use, with no monthly minimum.
Qwen3.8 Omni Flash Realtime holds a live conversation over a WebSocket at wss://api.empiriolabs.ai/v1/realtime?model=qwen3-8-omni-flash-realtime, not over the HTTP endpoints. Connect with the model id qwen3-8-omni-flash-realtime and send your EmpirioLabs API key as an Authorization: Bearer header on the handshake. Audio travels both ways as base64 16-bit PCM. A browser cannot set headers on a WebSocket, so open this connection from your server and relay audio to the browser over your own socket. Get an API key from the EmpirioLabs dashboard.
import asyncio, json, os, websockets
async def main():
async with websockets.connect(
"wss://api.empiriolabs.ai/v1/realtime?model=qwen3-8-omni-flash-realtime",
additional_headers=[
("Authorization", f"Bearer {os.environ['EMPIRIOLABS_API_KEY']}"),
],
max_size=None,
) as ws:
print(json.loads(await ws.recv())["type"]) # session.created
await ws.send(json.dumps({
"type": "conversation.item.create",
"item": {
"type": "message",
"role": "user",
"content": [{"type": "input_text", "text": "Say hello."}],
},
}))
await ws.send(json.dumps({"type": "response.create"}))
async for raw in ws:
event = json.loads(raw)
if event["type"] == "response.audio.delta":
... # base64 audio chunk, append to your playback buffer
elif event["type"] == "response.done":
break
asyncio.run(main())Request parameters supported by the Qwen3.8 Omni Flash Realtime API on EmpirioLabs. Defaults apply when a field is omitted.
| Parameter | Type | Default | Range / values | Description |
|---|---|---|---|---|
| voice | enum | Tina | Tina, Cindy, Liora Mira, Raymond, Zane, Katerina, Ryan, Mia, ... | Speaking voice for the session. Set it with session.update before the model produces any audio. These voices are specific to this model. |
| instructions | string | - | - | System guidance for how the model should behave and speak during the conversation. |
| modalities | string | ["text","audio"] | - | Which output types the model returns for a turn. Drop audio for a text-only reply. |
| input_audio_format | enum | pcm16 | pcm16 | Encoding of the audio you append to the input buffer: base64 16-bit PCM at 16 kHz. |
| output_audio_format | enum | pcm24 | pcm24 | Encoding of the audio the model streams back: 16-bit PCM at 24 kHz. |
| turn_detection | string | {"type":"server_vad"} | - | Server-side voice activity detection. Decides when you have stopped speaking and the model should reply. It is on by default on this model. |
Full-duplex voice over a WebSocket at wss://api.empiriolabs.ai/v1/realtime?model=qwen3-8-omni-flash-realtime, authenticated with the ordinary Authorization Bearer header. Configure the session with session.update over the socket: voice, instructions, modalities, and turn detection.
The voices on this model are not the same set as the Qwen3.5 Omni realtime models. Pick one from the voice list rather than carrying a voice across, or leave it unset and the default is supplied for you.
Audio and text tokens are priced separately in both directions, and each completed turn is billed on its own from the usage the model reports.
On EmpirioLabs, Qwen3.8 Omni Flash Realtime is billed pay as you go: Input: audio $1.86 per 1M audio input tokens; Input $0.46 per 1M prompt tokens; Output $1.40 per 1M generated tokens; Output: audio $3.74 per 1M generated audio tokens. The live rate card on this page always matches what the API charges.
Qwen3.8 Omni Flash Realtime supports a 192K-token context window with up to 65,536 output tokens per response.
Qwen3.8 Omni Flash Realtime is served through WEBSOCKET /v1/realtime on api.empiriolabs.ai with standard bearer-token authentication.
Open the EmpirioLabs playground and press Start session to try it in your browser. Qwen3.8 Omni Flash Realtime runs over a WebSocket rather than a request and response, so to build with it start from the realtime voice quickstart. Connect from your server with your EmpirioLabs API key and the model id qwen3-8-omni-flash-realtime, then stream audio in both directions.
Create an EmpirioLabs account, then generate a key under API Keys in the dashboard. Billing is pay-as-you-go credits, so you only pay for the requests you make.
Check out our pricing or reach out if you want your own model deployed on our stack.