Qwen Audio 3.1 ASR Message API

Transcription tuned for voice messages and voice input over a socket, returning each utterance as one complete transcript when the speaker pauses.

Alibaba CloudTranscriptionReleased Sep 21, 2026SingaporeProprietary EndpointNew

About Qwen Audio 3.1 ASR Message

Transcription tuned for voice messages and voice input over a socket, returning each utterance as one complete transcript when the speaker pauses.

Stream 16 kHz pcm16 audio over the socket. Each utterance is returned whole when the speaker pauses.

Also known as Qwen Audio ASR Message, Alibaba Cloud Qwen Audio 3.1 ASR Message

transcriptionspeech to textrealtimemultilingual

Qwen Audio 3.1 ASR Message specs

Model ID
qwen-audio-3-1-asr-message
Author
Alibaba Cloud
Category
Transcription
Released
Sep 21, 2026
Input
Audio
Output
Text
Region
Singapore
Endpoints
WEBSOCKET/v1/realtime
Alternate model IDs
qwen-audio-3.1-asr-messagealibaba/qwen-audio-3-1-asr-message

Qwen Audio 3.1 ASR Message API pricing

Live pay-as-you-go rates from the EmpirioLabs catalog. You are billed only for what you use, with no monthly minimum.

Type
Spec
Rate
Input
per 1M audio input tokens
$1.86
Output
per 1M generated tokens
$1.40
Compare on the full pricing page

How to call the Qwen Audio 3.1 ASR Message API

Qwen Audio 3.1 ASR Message transcribes a live audio stream over a WebSocket at wss://api.empiriolabs.ai/v1/realtime?model=qwen-audio-3-1-asr-message, not over the HTTP endpoints. Connect with the model id qwen-audio-3-1-asr-message and send your EmpirioLabs API key as an Authorization: Bearer header on the handshake. Append base64 16-bit PCM audio and read transcript events back while the speaker is still talking. A browser cannot set headers on a WebSocket, so open this connection from your server and relay audio to the browser over your own socket. Get an API key from the EmpirioLabs dashboard.

Python (websockets)
import asyncio, json, os, websockets

async def main():
    async with websockets.connect(
        "wss://api.empiriolabs.ai/v1/realtime?model=qwen-audio-3-1-asr-message",
        additional_headers=[
            ("Authorization", f"Bearer {os.environ['EMPIRIOLABS_API_KEY']}"),
        ],
        max_size=None,
    ) as ws:
        print(json.loads(await ws.recv())["type"])  # session.created

        import base64

        # 16-bit PCM microphone audio, base64 encoded, in small chunks.
        for chunk in read_microphone_chunks():
            await ws.send(json.dumps({
                "type": "input_audio_buffer.append",
                "audio": base64.b64encode(chunk).decode(),
            }))

        await ws.send(json.dumps({"type": "input_audio_buffer.commit"}))

        async for raw in ws:
            event = json.loads(raw)
            if event["type"].endswith("input_audio_transcription.delta"):
                # "text" is settled transcript; "stash" is still in progress.
                if event.get("text"):
                    print(event["text"], end="", flush=True)

asyncio.run(main())
Full Qwen Audio 3.1 ASR Message API reference

Qwen Audio 3.1 ASR Message API parameters

Request parameters supported by the Qwen Audio 3.1 ASR Message API on EmpirioLabs. Defaults apply when a field is omitted.

ParameterTypeDefaultRange / valuesDescription
input_audio_formatenumpcm16pcm16Encoding of the audio you append to the input buffer: base64 16-bit PCM, mono, at 16 kHz.

Good to know

Connecting

Voice message transcription over a WebSocket at wss://api.empiriolabs.ai/v1/realtime?model=qwen-audio-3-1-asr-message, authenticated with the ordinary Authorization Bearer header.

Sending audio

Stream 16 kHz mono pcm16 audio with input_audio_buffer.append. Voice activity detection decides where each utterance ends, so there is no commit to send.

What comes back

Each utterance arrives whole, as one conversation.item.input_audio_transcription.completed event once the speaker pauses. This model does not send partial results while the speaker is talking; for live captions, use qwen-audio-3-1-asr-stream.

Billing

Input and output tokens are priced per 1M, and each completed utterance is billed on its own from the usage the model reports.

Qwen Audio 3.1 ASR Message API: common questions

How much does the Qwen Audio 3.1 ASR Message API cost?

On EmpirioLabs, Qwen Audio 3.1 ASR Message is billed pay as you go: Input $1.86 per 1M audio input tokens; Output $1.40 per 1M generated tokens. The live rate card on this page always matches what the API charges.

Which endpoint does Qwen Audio 3.1 ASR Message use?

Qwen Audio 3.1 ASR Message is served through WEBSOCKET /v1/realtime on api.empiriolabs.ai with standard bearer-token authentication.

Can I try Qwen Audio 3.1 ASR Message in the browser before integrating?

Open the EmpirioLabs playground and press Start session to try it in your browser. Qwen Audio 3.1 ASR Message runs over a WebSocket rather than a request and response, so to build with it start from the realtime voice quickstart. Connect from your server with your EmpirioLabs API key and the model id qwen-audio-3-1-asr-message, then stream audio in both directions.

How do I get a Qwen Audio 3.1 ASR Message API key?

Create an EmpirioLabs account, then generate a key under API Keys in the dashboard. Billing is pay-as-you-go credits, so you only pay for the requests you make.

Ready to use better endpoints?

Check out our pricing or reach out if you want your own model deployed on our stack.