Disclosure: This article was written with AI assistance and reviewed by EmpirioLabs AI.
Qwen Audio 3.1 ASR, Qwen Audio 3.1 ASR Stream and Qwen Audio 3.1 ASR Message are available on EmpirioLabs. ASR transcribes recorded audio through POST /v1/audio/transcriptions. ASR Stream and ASR Message transcribe live audio over wss://api.empiriolabs.ai/v1/realtime. All three work in the API and in the Playground.
The three models
qwen-audio-3-1-asrtranscribes short clips and long recordings, and returns the transcript with sentence segments and word timestamps. The language is detected automatically, including Chinese dialects.qwen-audio-3-1-asr-streamis built for live captions. It sends the sentence so far while the speaker is talking, then the settled sentence at each pause.qwen-audio-3-1-asr-messageis built for voice messages and voice input. It returns each utterance as one complete transcript once the speaker pauses, with no partial results.
Transcribing a recording
curl https://api.empiriolabs.ai/v1/audio/transcriptions \
-H "Authorization: Bearer $EMPIRIOLABS_API_KEY" \
-F model=qwen-audio-3-1-asr \
-F file=@meeting.mp3
A JSON body with audio_url or audio_base64 works in place of the file upload. The response is a job:
{
"job_id": "3b5e1c2a-9f1d-4e7a-8c61-2d4f0a9b7e10",
"status": "processing",
"poll_url": "/v1/jobs/3b5e1c2a-9f1d-4e7a-8c61-2d4f0a9b7e10"
}
Poll the job until status is completed:
curl https://api.empiriolabs.ai/v1/jobs/3b5e1c2a-9f1d-4e7a-8c61-2d4f0a9b7e10 \
-H "Authorization: Bearer $EMPIRIOLABS_API_KEY"
The transcript is in result.text, with sentence segments in result.segments and word timestamps in result.words. Short clips finish in seconds; long recordings take longer.
Live transcription over a WebSocket
Open the realtime endpoint with the model in the query string and your API key on the handshake:
wss://api.empiriolabs.ai/v1/realtime?model=qwen-audio-3-1-asr-stream
Authorization: Bearer YOUR_EMPIRIOLABS_API_KEY
Stream 16 kHz, 16-bit mono PCM audio as base64 in input_audio_buffer.append events. There is no commit to send: the model detects where each sentence ends from the pause that follows it.
{"type": "input_audio_buffer.append", "audio": "<base64 PCM>"}
Transcripts arrive as events:
- Partial results.
conversation.item.input_audio_transcription.deltacarries the sentence so far intextwhile the speaker is still talking. ASR Stream only. - Settled results.
conversation.item.input_audio_transcription.completedcarries the finished sentence or utterance intranscript. Both models.
Operational notes
- Every transcription request returns a job, even a short clip. The first response carries a
job_id, not the transcript. Readresult.textfromGET /v1/jobs/<job_id>oncestatusiscompleted. - Display
textfrom each partial event instead of joiningdeltavalues. On ASR Stream,textis always the whole sentence so far, so replacing the displayed line with it on every event keeps live captions correct. - ASR Message is silent until the speaker pauses. It sends no partial results, so no events arrive while someone is talking. Use ASR Stream for live captions.
- Stream a moment of silence after the last words. A sentence settles when the model hears the speaker stop. Keep sending audio briefly after the final words and wait for the
completedevent before closing the socket, or the last sentence is lost.
Pricing
All three models are pay as you go, with no subscription required. Each is priced per token, with one rate for the audio input and one for the transcript output, and there is no cached input tier. ASR bills each request; ASR Stream and ASR Message bill each completed sentence or utterance. The live models have higher rates than ASR, so use ASR for recordings that do not need results while the speaker is talking. Current rates are on the model pages for Qwen Audio 3.1 ASR, ASR Stream and ASR Message, and on the pricing page.
Start building
Try Qwen Audio 3.1 ASR in the Playground by uploading a recording, or ASR Stream and ASR Message by clicking Start session and speaking. The API reference for each model is at ASR, ASR Stream and ASR Message, and the full event tables are in the Realtime Voice API guide.
For voice conversations from the same family, see How to Use Qwen3.8 Omni Flash Realtime and Qwen Audio 3.1.



