Disclosure: This article was written with AI assistance and reviewed by EmpirioLabs AI.
Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are available on EmpirioLabs. Both turn text into speech through POST /v1/audio/speech, stream audio through POST /v1/audio/speech:stream, and draw on the same library of more than 2,000 voices. Both work in the API and in the Playground.
The two models
gemini-3-8-flash-ttsis built for creative direction: narration, audiobooks, characters, and two-speaker scenes. It covers 130 languages.gemini-3-8-flash-lite-ttsis built for high-volume work such as voice agents, dubbing, and bulk audio, with a lower rate for generated audio. It covers 101 languages.
Both models take the same request body, so switching between them is a change of model.
Your first request
curl https://api.empiriolabs.ai/v1/audio/speech \
-H "Authorization: Bearer $EMPIRIOLABS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-3-8-flash-tts",
"input": "Welcome back. <short pause> Here is what changed this week.",
"voice": "Kore",
"style_prompt": "warm and unhurried"
}'
The response carries a signed URL to the generated audio. The default format is WAV at 24,000 Hz. Set output_format to MP3, OGG, ALAW, or MULAW, and sample_rate to any of 8,000, 16,000, 22,050, 24,000, 44,100, or 48,000 Hz.
Directing the performance
- Style direction.
style_promptsets the tone for the whole passage, for example "calm and reassuring" or "excited sports commentator". - Vocal tags. Tags such as
<laugh>,<sigh>,<breath>, and<short pause>go inline where the sound should happen. The full list is on theinputparameter in the model reference. Use the English tags even when the text is in another language. - Emphasis. Capitalize a word to stress it.
Two-speaker dialogue
Set mode to multi, choose a voice for each speaker with voice and voice2, and start every line with the speaker's name and a colon. The names must match speaker1_name and speaker2_name, which default to Speaker1 and Speaker2.
{
"model": "gemini-3-8-flash-lite-tts",
"mode": "multi",
"speaker1_name": "Maya",
"speaker2_name": "Theo",
"voice": "Aoede",
"voice2": "Puck",
"input": "Maya: Did the shipment arrive?\nTheo: It did. <laugh> Every box, on time."
}
Choosing a voice
The 30 voices in the voice list, such as Kore, Puck, Charon, and Aoede, speak every supported language. The full library holds more than 2,000 voices across 30 locales, each with a language, accent, gender, pitch, and description. List it with GET /v1/voices and filter by language, gender, pitch, or a search term with q:
curl "https://api.empiriolabs.ai/v1/voices?model=gemini-3-8-flash-tts&language=fr&gender=female" \
-H "Authorization: Bearer $EMPIRIOLABS_API_KEY"
Any voice id the listing returns works in voice and voice2.
Streaming
POST /v1/audio/speech:stream takes the same body and returns Server-Sent Events: audio.chunk events carrying base64 16-bit PCM at 24,000 Hz as the audio is generated, then one audio.done event with a link to the complete file in the requested format.
Operational notes
- Scripts written for Gemini 3.1 Flash TTS need their tags rewritten. The 3.8 models use angle-bracket tags such as
<whispers>and<laugh>instead of the square-bracket tags 3.1 uses. Update the tags when you move a script across. - A library voice keeps its own language. When you set
languagetogether with a library voice, it must be that voice's locale:fr-fr-advisor-1withen-USis rejected. The 30 multilingual voices accept any language. - Dialogue labels must match the speaker names. In
multimode, text before the first labelled line is rejected, and a line whose label matches neitherspeaker1_namenorspeaker2_nameis read as part of the previous turn. speedapplies to complete files. The streaming endpoint rejects aspeedother than 1.0. Request the file fromPOST /v1/audio/speechwhen you need a different speed.
Pricing
Both models are pay as you go, with no subscription. A request is billed for its input text tokens and its generated audio, at 32 audio tokens per second of speech, and Flash-Lite has the lower rate for generated audio. Current rates are on the model pages for Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS and on the pricing page.
Start building
Try Gemini 3.8 Flash TTS or Gemini 3.8 Flash-Lite TTS in the Playground, and read the API reference for Flash and Flash-Lite.
For realtime voice conversations from the same family, see How to Use Gemini 3.8 Live and Live Extended Thinking.



