
Live transcription over a socket, streaming partial words while the speaker talks and a settled sentence at each pause, priced per token.
Live transcription over a socket, streaming partial words while the speaker talks and a settled sentence at each pause, priced per token.
Stream 16 kHz pcm16 audio over the socket. Partial results arrive while the speaker is talking.
Also known as Qwen Audio ASR Stream, Alibaba Cloud Qwen Audio 3.1 ASR Stream
qwen-audio-3-1-asr-stream/v1/realtimeqwen-audio-3.1-asr-streamalibaba/qwen-audio-3-1-asr-streamLive pay-as-you-go rates from the EmpirioLabs catalog. You are billed only for what you use, with no monthly minimum.
Qwen Audio 3.1 ASR Stream transcribes a live audio stream over a WebSocket at wss://api.empiriolabs.ai/v1/realtime?model=qwen-audio-3-1-asr-stream, not over the HTTP endpoints. Connect with the model id qwen-audio-3-1-asr-stream and send your EmpirioLabs API key as an Authorization: Bearer header on the handshake. Append base64 16-bit PCM audio and read transcript events back while the speaker is still talking. A browser cannot set headers on a WebSocket, so open this connection from your server and relay audio to the browser over your own socket. Get an API key from the EmpirioLabs dashboard.
import asyncio, json, os, websockets
async def main():
async with websockets.connect(
"wss://api.empiriolabs.ai/v1/realtime?model=qwen-audio-3-1-asr-stream",
additional_headers=[
("Authorization", f"Bearer {os.environ['EMPIRIOLABS_API_KEY']}"),
],
max_size=None,
) as ws:
print(json.loads(await ws.recv())["type"]) # session.created
import base64
# 16-bit PCM microphone audio, base64 encoded, in small chunks.
for chunk in read_microphone_chunks():
await ws.send(json.dumps({
"type": "input_audio_buffer.append",
"audio": base64.b64encode(chunk).decode(),
}))
await ws.send(json.dumps({"type": "input_audio_buffer.commit"}))
async for raw in ws:
event = json.loads(raw)
if event["type"].endswith("input_audio_transcription.delta"):
# "text" is settled transcript; "stash" is still in progress.
if event.get("text"):
print(event["text"], end="", flush=True)
asyncio.run(main())Request parameters supported by the Qwen Audio 3.1 ASR Stream API on EmpirioLabs. Defaults apply when a field is omitted.
| Parameter | Type | Default | Range / values | Description |
|---|---|---|---|---|
| input_audio_format | enum | pcm16 | pcm16 | Encoding of the audio you append to the input buffer: base64 16-bit PCM, mono, at 16 kHz. |
Live transcription over a WebSocket at wss://api.empiriolabs.ai/v1/realtime?model=qwen-audio-3-1-asr-stream, authenticated with the ordinary Authorization Bearer header.
Stream 16 kHz mono pcm16 audio with input_audio_buffer.append. Voice activity detection decides where each sentence ends, so there is no commit to send.
conversation.item.input_audio_transcription.delta events carry the sentence so far while the speaker is still talking, and conversation.item.input_audio_transcription.completed carries the settled sentence at each pause.
Input and output tokens are priced per 1M, and each completed sentence is billed on its own from the usage the model reports.
On EmpirioLabs, Qwen Audio 3.1 ASR Stream is billed pay as you go: Input $1.86 per 1M audio input tokens; Output $1.40 per 1M generated tokens. The live rate card on this page always matches what the API charges.
Qwen Audio 3.1 ASR Stream is served through WEBSOCKET /v1/realtime on api.empiriolabs.ai with standard bearer-token authentication.
Open the EmpirioLabs playground and press Start session to try it in your browser. Qwen Audio 3.1 ASR Stream runs over a WebSocket rather than a request and response, so to build with it start from the realtime voice quickstart. Connect from your server with your EmpirioLabs API key and the model id qwen-audio-3-1-asr-stream, then stream audio in both directions.
Create an EmpirioLabs account, then generate a key under API Keys in the dashboard. Billing is pay-as-you-go credits, so you only pay for the requests you make.
Check out our pricing or reach out if you want your own model deployed on our stack.