Home Blog

How to Use Qwen3.8 Omni Flash Realtime and Qwen Audio 3.1

Qwen3.8 Omni Flash Realtime and Qwen Audio 3.1 Realtime Plus via API

Sep 21, 2026

EmpirioLabs AI

Two new Alibaba voice models are live on EmpirioLabs, and both are conversations rather than requests. You open a WebSocket, stream microphone audio up, and audio comes back while the speaker is still talking.

Qwen3.8 Omni Flash Realtime is the one that can also see. Alongside the audio you are already sending, you can push live video frames, and it will answer questions about what is on screen in the same turn. It carries 56 voices and calls your functions mid-conversation.

Qwen Audio 3.1 Realtime Plus is the one that can look things up. It has native web search you can switch on per session, 27 voices of its own, and a 256K context.

Try either in the playground: Omni Flash Realtime or Audio 3.1 Realtime Plus. Click Start session and talk.

Connecting

There is one realtime endpoint for every model on the platform, and you pick the model with a query parameter. Authenticate the handshake with the same API key you use everywhere else:

wss://api.empiriolabs.ai/v1/realtime?model=qwen3-8-omni-flash-realtime
Authorization: Bearer YOUR_EMPIRIOLABS_API_KEY

The protocol follows the widely used realtime event shape, so a client written against that convention works unchanged. Configure the session over the socket with session.update, append microphone audio as input_audio_buffer.append, and read audio back from response.audio.delta with the matching transcript on response.audio_transcript.delta.

{"type": "session.update", "session": {
    "voice": "Tina",
    "modalities": ["text", "audio"]
}}

Audio travels both ways as 16-bit PCM. These two models take 16 kHz going up and return 24 kHz coming back. The session.created event you receive on connect reports the formats for that session, so read them there rather than assuming.

Video frames on Omni Flash Realtime

Frames go in their own buffer, next to the audio buffer:

{"type": "input_image_buffer.append", "image": "<base64 JPEG or PNG>"}

Send them as you would frames from a camera, then ask about what you are showing it in the same spoken turn. It keeps up to 50 turns and 240 seconds of video history, and up to 100 turns and 600 seconds of audio. Older history falls out of context as those limits are passed.

Web search on Audio 3.1

Search is off unless you ask for it, and you ask for it in the session:

{"type": "session.update", "session": {"enable_search": true}}

With it on, the model can answer questions about things that happened after it was trained. Results are added to the conversation, so a turn that searches reports a noticeably larger input token count than one that does not.

Pricing

Both models bill each completed turn from the usage the model reports, with separate rates for audio tokens and text tokens in both directions. That split matters: a spoken reply is mostly audio tokens, and asking for a text-only reply by dropping audio from modalities genuinely costs less. Current rates live on the Omni Flash Realtime and Audio 3.1 Realtime Plus model pages and on the pricing page, which stay in sync with what you are charged. Neither model has a cached input tier, so every input token bills at the ordinary rate.

Things worth knowing before your first session

  • Do not send back the voice the session reports. Both models announce a voice on connect that their own generator then refuses, which kills the first turn. Leave the voice unset and the platform sets a working default for that model, or pick one from the voice list yourself. The two lists share no names, so a voice that works on one will not work on the other.
  • Voice activity detection needs a moment of silence to fire. These models decide you have stopped talking by hearing you stop. If your client cuts the audio stream the instant the speaker finishes the last word, nothing is committed and no reply arrives, which looks exactly like a broken session. Keep streaming for about a second after the speech ends.
  • Video frames are not message content. They go through input_image_buffer.append. Putting an image inside a conversation.item.create content array is rejected, and on some models it leaves the turn with no user message at all.
  • On Audio 3.1, web search and function calling are mutually exclusive. Enable one per session, not both.
  • A new socket is a new conversation. There is no server-side memory across connections, so a client that reconnects should replay whatever context it wants the model to still have.

Full event tables, voice lists and audio formats are in the Realtime Voice API docs. Both models are available now.

Ready to use better endpoints?

Explore our models, or contact us about business inquiries, custom deployments, or anything else.