Qwen Audio 3.1 Realtime Plus API

Duplex speech conversation over one socket, with 27 voices, native web search, tool calling, and separate text and audio token rates.

Alibaba CloudAudio Generation256K contextReleased Sep 20, 2026SingaporeProprietary EndpointNew

About Qwen Audio 3.1 Realtime Plus

Duplex speech conversation over one socket, with 27 voices, native web search, tool calling, and separate text and audio token rates.

Session is configured with session.update events over the socket rather than request parameters. Web search is off by default and cannot be combined with function calling. Text and audio tokens bill at separate rates.

Also known as Qwen Audio Realtime Plus, Qwen Audio Realtime, Alibaba Cloud Qwen Audio 3.1 Realtime Plus

realtimespeech to speechaudio inaudio outfunction callingweb searchmultilingual

Qwen Audio 3.1 Realtime Plus specs

Model ID
qwen-audio-3-1-realtime-plus
Author
Alibaba Cloud
Category
Audio Generation
Released
Sep 20, 2026
Context window
256K tokens
Max output
16,384 tokens
Input
AudioText
Output
AudioText
Region
Singapore
Endpoints
WEBSOCKET/v1/realtime
Alternate model IDs
qwen-audio-3.1-realtime-plusalibaba/qwen-audio-3-1-realtime-plus

Qwen Audio 3.1 Realtime Plus API pricing

Live pay-as-you-go rates from the EmpirioLabs catalog. You are billed only for what you use, with no monthly minimum.

Type
Spec
Rate
Input: audio
per 1M audio input tokens
$12.80
Input
per 1M prompt tokens
$1.60
Output
per 1M generated tokens
$12.80
Output: audio
per 1M generated audio tokens
$48.00
Web search
per call when invoked
$0.00
Compare on the full pricing page

How to call the Qwen Audio 3.1 Realtime Plus API

Qwen Audio 3.1 Realtime Plus holds a live conversation over a WebSocket at wss://api.empiriolabs.ai/v1/realtime?model=qwen-audio-3-1-realtime-plus, not over the HTTP endpoints. Connect with the model id qwen-audio-3-1-realtime-plus and send your EmpirioLabs API key as an Authorization: Bearer header on the handshake. Audio travels both ways as base64 16-bit PCM. A browser cannot set headers on a WebSocket, so open this connection from your server and relay audio to the browser over your own socket. Get an API key from the EmpirioLabs dashboard.

Python (websockets)
import asyncio, json, os, websockets

async def main():
    async with websockets.connect(
        "wss://api.empiriolabs.ai/v1/realtime?model=qwen-audio-3-1-realtime-plus",
        additional_headers=[
            ("Authorization", f"Bearer {os.environ['EMPIRIOLABS_API_KEY']}"),
        ],
        max_size=None,
    ) as ws:
        print(json.loads(await ws.recv())["type"])  # session.created

        await ws.send(json.dumps({
            "type": "conversation.item.create",
            "item": {
                "type": "message",
                "role": "user",
                "content": [{"type": "input_text", "text": "Say hello."}],
            },
        }))
        await ws.send(json.dumps({"type": "response.create"}))

        async for raw in ws:
            event = json.loads(raw)
            if event["type"] == "response.audio.delta":
                ...  # base64 audio chunk, append to your playback buffer
            elif event["type"] == "response.done":
                break

asyncio.run(main())
Full Qwen Audio 3.1 Realtime Plus API reference

Qwen Audio 3.1 Realtime Plus API parameters

Request parameters supported by the Qwen Audio 3.1 Realtime Plus API on EmpirioLabs. Defaults apply when a field is omitted.

ParameterTypeDefaultRange / valuesDescription
voiceenumlonganqian_v3.1longanqian_v3.1, longanhuan_v3.1, longanlingxin_v3.1, longanf...Speaking voice for the session. Set it with session.update before the model produces any audio. These voices are specific to this model.
instructionsstring--System guidance for how the model should behave and speak during the conversation.
enable_searchbooleanfalse-Search the web for real-time information. Cannot be used in the same session as function calling. Search results are added to the conversation and count as input tokens.
modalitiesstring["text","audio"]-Which output types the model returns for a turn. Drop audio for a text-only reply.
input_audio_formatenumpcm16pcm16Encoding of the audio you append to the input buffer: base64 16-bit PCM at 16 kHz.
output_audio_formatenumpcm24pcm24Encoding of the audio the model streams back: 16-bit PCM at 24 kHz.
turn_detectionstring{"type":"server_vad"}-Server-side voice activity detection. Decides when you have stopped speaking and the model should reply. It is on by default on this model.

Good to know

Connecting

Full-duplex voice over a WebSocket at wss://api.empiriolabs.ai/v1/realtime?model=qwen-audio-3-1-realtime-plus, authenticated with the ordinary Authorization Bearer header. Configure the session with session.update over the socket: voice, instructions, modalities, and turn detection.

Sending audio

  • Append microphone audio as input_audio_buffer.append events.
  • Voice activity detection decides when you have stopped speaking, so keep streaming for about a second after the speech ends.

Voices

The voices on this model are its own set. Pick one from the voice list rather than carrying a voice across from another model, or leave it unset and the default is supplied for you.

Web search

Web search is off by default. Turn it on with enable_search. It cannot be enabled in the same session as function calling, and search results are added to the conversation and count as input tokens.

Billing

Audio and text tokens are priced separately in both directions, and each completed turn is billed on its own from the usage the model reports.

Qwen Audio 3.1 Realtime Plus API: common questions

How much does the Qwen Audio 3.1 Realtime Plus API cost?

On EmpirioLabs, Qwen Audio 3.1 Realtime Plus is billed pay as you go: Input: audio $12.80 per 1M audio input tokens; Input $1.60 per 1M prompt tokens; Output $12.80 per 1M generated tokens; Output: audio $48.00 per 1M generated audio tokens; Web search $0.00 per call when invoked. The live rate card on this page always matches what the API charges.

What is the context window of Qwen Audio 3.1 Realtime Plus?

Qwen Audio 3.1 Realtime Plus supports a 256K-token context window with up to 16,384 output tokens per response.

Which endpoint does Qwen Audio 3.1 Realtime Plus use?

Qwen Audio 3.1 Realtime Plus is served through WEBSOCKET /v1/realtime on api.empiriolabs.ai with standard bearer-token authentication.

Can I try Qwen Audio 3.1 Realtime Plus in the browser before integrating?

Open the EmpirioLabs playground and press Start session to try it in your browser. Qwen Audio 3.1 Realtime Plus runs over a WebSocket rather than a request and response, so to build with it start from the realtime voice quickstart. Connect from your server with your EmpirioLabs API key and the model id qwen-audio-3-1-realtime-plus, then stream audio in both directions.

How do I get a Qwen Audio 3.1 Realtime Plus API key?

Create an EmpirioLabs account, then generate a key under API Keys in the dashboard. Billing is pay-as-you-go credits, so you only pay for the requests you make.

Ready to use better endpoints?

Check out our pricing or reach out if you want your own model deployed on our stack.