Home Blog

How to Use Gemini 3.8 Live and Live Extended Thinking

Gemini 3.8 Live and Live Extended Thinking via API

Sep 24, 2026

EmpirioLabs AI

Disclosure: This article was written with AI assistance and reviewed by EmpirioLabs AI.

Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are available on EmpirioLabs. Both are realtime voice models: you open a WebSocket, stream microphone audio, and the spoken reply streams back on the same connection.

Gemini 3.8 Live is the low-latency option for voice agents and live dialogue. It accepts live video frames alongside the audio, calls your functions during the conversation, and replies in the language the caller speaks.

Gemini 3.8 Live Extended Thinking reasons in the background while it talks, for questions that need more than a quick answer. You set how much it reasons with reasoning_effort.

Try either in the Playground: Gemini 3.8 Live or Live Extended Thinking. Pick a voice, click Start session and talk.

Connecting

Every realtime model on the platform uses one endpoint, selected with the model query parameter. Authenticate the handshake with your EmpirioLabs API key:

wss://api.empiriolabs.ai/v1/realtime?model=gemini-3-8-live
Authorization: Bearer YOUR_EMPIRIOLABS_API_KEY

The protocol follows the widely used realtime event format. Configure the session with session.update before you send the first audio, append microphone audio with input_audio_buffer.append, and read the reply from response.audio.delta, with its transcript on response.audio_transcript.delta.

{"type": "session.update", "session": {
    "voice": "Kore",
    "instructions": "You are a concise travel assistant."
}}

Audio is 16-bit PCM in both directions: 16 kHz from your microphone, or 24 kHz with input_audio_format set to pcm24, and 24 kHz for the reply. Voice activity detection is on from the start, so the model answers when the caller stops speaking, and speaking over a reply interrupts it.

Video frames

Both models can see what the caller shows them. Send frames from a camera or a screen as base64 JPEG or PNG images in their own buffer, and ask about them in the same turn:

{"type": "input_image_buffer.append", "image": "<base64 JPEG or PNG>"}

Function calling

Declare functions in session.update with JSON Schema parameters. When the model calls one, you receive response.function_call_arguments.done with the call_id, the function name and its arguments. Return the result as a function_call_output item and the model continues speaking with it.

Reasoning on Extended Thinking

{"type": "session.update", "session": {"reasoning_effort": "high"}}

reasoning_effort accepts low, medium (the default) or high. The model often acknowledges a question first and then answers in a second response, so keep listening after the first response.done.

Pricing

Both models bill each completed turn from the usage the model reports, with separate rates for audio tokens and text tokens in both directions. Video frames bill as input tokens, and reasoning bills as output tokens. Extended Thinking reports more input tokens on every turn, so the same conversation costs more on it than on Live. Current rates are on the Gemini 3.8 Live and Live Extended Thinking model pages and on the pricing page.

Things worth knowing before your first session

  • Set the session before you speak. The voice, instructions, turn detection, tools, temperature and reasoning effort are fixed once the conversation starts. Changing them later returns an error; open a new session instead.
  • Use one of the 30 listed voices. Puck is the default. Any other value is refused with an error.
  • There is no language setting. The model replies in the language the caller speaks.
  • Keep streaming briefly after the caller stops. Voice activity detection needs a short silence to decide the turn is over, so a client that cuts the audio on the last word waits for a reply.
  • A new socket is a new conversation. Keep your own transcript if a reconnecting client needs the earlier context.

Event tables, the voice list and audio formats are in the Realtime Voice API docs. For text-to-speech from the same family, see How to Use Gemini 3.8 Flash TTS and Flash-Lite TTS. Both Live models are available now.

Ready to use better endpoints?

Explore our models, or contact us about business inquiries, custom deployments, or anything else.