Gemini 3.8 Live Extended Thinking API

Voice conversation that keeps talking while it reasons in the background, with 30 voices, live video input, tool calling, and adjustable reasoning effort.

GoogleAudio Generation128K contextReleased Sep 15, 2026Proprietary EndpointNew

About Gemini 3.8 Live Extended Thinking

Voice conversation that keeps talking while it reasons in the background, with 30 voices, live video input, tool calling, and adjustable reasoning effort.

Session is configured with session.update events over the socket rather than request parameters. Text and audio tokens bill at separate rates.

Also known as Gemini Live Extended Thinking, Google Gemini 3.8 Live Extended Thinking

realtimespeech to speechaudio inaudio outvideofunction callingreasoningmultilingual

Gemini 3.8 Live Extended Thinking specs

Model ID
gemini-3-8-live-extended-thinking
Author
Google
Category
Audio Generation
Released
Sep 15, 2026
Context window
128K tokens
Max output
65,536 tokens
Input
AudioTextVideo
Output
AudioText
Endpoints
WEBSOCKET/v1/realtime
Alternate model IDs
gemini-3.8-live-extended-thinkinggoogle/gemini-3.8-live-extended-thinking

Gemini 3.8 Live Extended Thinking API pricing

Live pay-as-you-go rates from the EmpirioLabs catalog. You are billed only for what you use, with no monthly minimum.

Type
Spec
Rate
Input: audio
per 1M audio input tokens
$7.80
Input
per 1M prompt tokens
$2.60
Output
per 1M generated tokens
$11.70
Output: audio
per 1M generated audio tokens
$31.20
Compare on the full pricing page

How to call the Gemini 3.8 Live Extended Thinking API

Gemini 3.8 Live Extended Thinking holds a live conversation over a WebSocket at wss://api.empiriolabs.ai/v1/realtime?model=gemini-3-8-live-extended-thinking, not over the HTTP endpoints. Connect with the model id gemini-3-8-live-extended-thinking and send your EmpirioLabs API key as an Authorization: Bearer header on the handshake. Audio travels both ways as base64 16-bit PCM. A browser cannot set headers on a WebSocket, so open this connection from your server and relay audio to the browser over your own socket. Get an API key from the EmpirioLabs dashboard.

Python (websockets)
import asyncio, json, os, websockets

async def main():
    async with websockets.connect(
        "wss://api.empiriolabs.ai/v1/realtime?model=gemini-3-8-live-extended-thinking",
        additional_headers=[
            ("Authorization", f"Bearer {os.environ['EMPIRIOLABS_API_KEY']}"),
        ],
        max_size=None,
    ) as ws:
        print(json.loads(await ws.recv())["type"])  # session.created

        await ws.send(json.dumps({
            "type": "conversation.item.create",
            "item": {
                "type": "message",
                "role": "user",
                "content": [{"type": "input_text", "text": "Say hello."}],
            },
        }))
        await ws.send(json.dumps({"type": "response.create"}))

        async for raw in ws:
            event = json.loads(raw)
            if event["type"] == "response.audio.delta":
                ...  # base64 audio chunk, append to your playback buffer
            elif event["type"] == "response.done":
                break

asyncio.run(main())
Full Gemini 3.8 Live Extended Thinking API reference

Gemini 3.8 Live Extended Thinking API parameters

Request parameters supported by the Gemini 3.8 Live Extended Thinking API on EmpirioLabs. Defaults apply when a field is omitted.

ParameterTypeDefaultRange / valuesDescription
voiceenumPuckZephyr, Puck, Charon, Kore, Fenrir, Leda, Orus, Aoede, Callir...Speaking voice for the session. Set it with session.update before the conversation starts; it cannot change afterwards.
instructionsstring--System guidance for how the model behaves and speaks. Set it with session.update before the conversation starts.
reasoning_effortenummediumlow, medium, highHow much the model reasons in the background while it talks. Higher effort uses more output tokens. Set it before the conversation starts.
temperaturenumber-0 to 2Sampling temperature from 0 to 2. Lower values make replies more consistent. Set it before the conversation starts.
input_audio_formatenumpcm16pcm16, pcm24Encoding of the audio you append to the input buffer: pcm16 is base64 16-bit PCM at 16 kHz, and pcm24 is the same at 24 kHz.
output_audio_formatenumpcm24pcm24Encoding of the audio the model streams back: 16-bit PCM at 24 kHz.
turn_detectionstring{"type":"server_vad"}-Server-side voice activity detection. Decides when you have stopped speaking and the model should reply. It is on by default; set it to null to end each turn...

Good to know

Connecting

Full-duplex voice over a WebSocket at wss://api.empiriolabs.ai/v1/realtime?model=gemini-3-8-live-extended-thinking, authenticated with the ordinary Authorization Bearer header. Configure the session with session.update before you send the first audio: voice, instructions, turn detection and tools. These settings are fixed once the conversation starts.

Sending audio and video

  • Append microphone audio as input_audio_buffer.append events: 16-bit PCM at 16 kHz, or at 24 kHz with input_audio_format set to pcm24.
  • Append live video frames as input_image_buffer.append events, as base64 JPEG or PNG images.
  • Voice activity detection is on by default, so keep streaming for about a second after the speech ends. Set turn_detection to null to end each turn yourself with input_audio_buffer.commit.

Voices and language

  • Use one of the 30 voices in the voice list, for example Puck, Kore or Charon. Any other value is refused.
  • The model replies in the language you speak.

Replies

  • Speech streams as response.audio.delta events, 16-bit PCM at 24 kHz, with the words as response.audio_transcript.delta events. Your own speech is transcribed as well.
  • Speaking while the model replies interrupts it.

Reasoning

  • reasoning_effort sets how much the model reasons in the background while it talks: low, medium (the default) or high. Set it before the conversation starts.
  • The model often acknowledges a question first and answers in a second turn. Each turn is billed on its own.

Tools

  • Function calling with JSON Schema parameters. Return each result as a function_call_output item with its call_id.

Billing

Audio and text tokens are priced separately in both directions, video frames bill as input tokens, and reasoning tokens bill as output tokens. Each completed turn is billed on its own from the usage the model reports, including a reply you interrupt.

Gemini 3.8 Live Extended Thinking API: common questions

How much does the Gemini 3.8 Live Extended Thinking API cost?

On EmpirioLabs, Gemini 3.8 Live Extended Thinking is billed pay as you go: Input: audio $7.80 per 1M audio input tokens; Input $2.60 per 1M prompt tokens; Output $11.70 per 1M generated tokens; Output: audio $31.20 per 1M generated audio tokens. The live rate card on this page always matches what the API charges.

What is the context window of Gemini 3.8 Live Extended Thinking?

Gemini 3.8 Live Extended Thinking supports a 128K-token context window with up to 65,536 output tokens per response.

Which endpoint does Gemini 3.8 Live Extended Thinking use?

Gemini 3.8 Live Extended Thinking is served through WEBSOCKET /v1/realtime on api.empiriolabs.ai with standard bearer-token authentication.

Can I try Gemini 3.8 Live Extended Thinking in the browser before integrating?

Open the EmpirioLabs playground and press Start session to try it in your browser. Gemini 3.8 Live Extended Thinking runs over a WebSocket rather than a request and response, so to build with it start from the realtime voice quickstart. Connect from your server with your EmpirioLabs API key and the model id gemini-3-8-live-extended-thinking, then stream audio in both directions.

How do I get a Gemini 3.8 Live Extended Thinking API key?

Create an EmpirioLabs account, then generate a key under API Keys in the dashboard. Billing is pay-as-you-go credits, so you only pay for the requests you make.

Ready to use better endpoints?

Check out our pricing or reach out if you want your own model deployed on our stack.