Gemini 3.8 Flash TTS API

Expressive speech with inline vocal tags, style direction, and two-speaker dialogue, across 130 languages and a library of more than 2,000 voices.

GoogleAudio GenerationReleased Sep 22, 2026Proprietary EndpointNew

About Gemini 3.8 Flash TTS

Expressive speech with inline vocal tags, style direction, and two-speaker dialogue, across 130 languages and a library of more than 2,000 voices.

Also known as Gemini Flash TTS, Google Gemini 3.8 Flash TTS

text to speechmulti speakermultilingualvoice tagsvoice controlstreamingvoice cloning

Gemini 3.8 Flash TTS specs

Model ID
gemini-3-8-flash-tts
Author
Google
Category
Audio Generation
Released
Sep 22, 2026
Input
Text
Output
Audio
Endpoints
POST/v1/audio/speechPOST/v1/audio/speech:streamGET/v1/voices
Alternate model IDs
gemini-3.8-flash-ttsgoogle/gemini-3.8-flash-tts

Gemini 3.8 Flash TTS API pricing

Live pay-as-you-go rates from the EmpirioLabs catalog. You are billed only for what you use, with no monthly minimum.

Type
Spec
Rate
Input
per 1M prompt tokens
$2.60
Output
per 1M generated tokens
$46.80
Compare on the full pricing page

How to call the Gemini 3.8 Flash TTS API

Gemini 3.8 Flash TTS serves speech through POST /v1/audio/speech and returns playable audio. Send the text to speak as input with the model id gemini-3-8-flash-tts. Get an API key from the EmpirioLabs dashboard.

cURL
curl https://api.empiriolabs.ai/v1/audio/speech \
  -H "Authorization: Bearer $EMPIRIOLABS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemini-3-8-flash-tts",
    "input": "Welcome to EmpirioLabs. Your build just finished."
  }' \
  --output speech.mp3
Python
import requests

response = requests.post(
    "https://api.empiriolabs.ai/v1/audio/speech",
    headers={"Authorization": "Bearer YOUR_EMPIRIOLABS_API_KEY"},
    json={"model": "gemini-3-8-flash-tts", "input": "Welcome to EmpirioLabs."},
)
with open("speech.mp3", "wb") as f:
    f.write(response.content)
Full Gemini 3.8 Flash TTS API reference

Gemini 3.8 Flash TTS API parameters

Request parameters supported by the Gemini 3.8 Flash TTS API on EmpirioLabs. Defaults apply when a field is omitted.

ParameterTypeDefaultRange / valuesDescription
inputstring-max 5000Text to convert to speech. For multi-speaker mode, prefix lines with Speaker1: / Speaker2:. Up to 5,000 characters. Put delivery direction in style_prompt, and place...
modeenumsinglesingle, multisingle = one voice, multi = two-voice dialogue (uses voice + voice2 + speaker names).
languagestring--Optional BCP-47 language tag (en-US, es-ES, etc.). The language of the text is detected automatically. The 30 listed voices speak every supported language; a library...
voiceenumCharonZephyr, Puck, Charon, Kore, Fenrir, Leda, Orus, Aoede, Callir...Primary voice name (e.g. Kore, Puck, Aoede). Leave blank for the default. These 30 voices speak every supported language. GET /v1/voices lists the full library of...
voice_audio_urlstring--voice=custom only: an https URL to a 10 to 30 second recording of clean, natural speech from the voice to clone (WAV, MP3, M4A, OGG, WEBM or FLAC). Record it in a...
voice_consent_audio_urlstring--voice=custom only: an https URL to the same speaker reading this statement aloud: "I am the owner of this voice and I consent to Google using this voice to create a...
voice2enumKoreZephyr, Puck, Charon, Kore, Fenrir, Leda, Orus, Aoede, Callir...Second voice name for multi-speaker mode.
speaker1_namestringSpeaker1-Display name used in the input prefix for speaker 1 (default: Speaker1).
speaker2_namestringSpeaker2-Display name used in the input prefix for speaker 2 (default: Speaker2).
output_formatenumWAVWAV, MP3, OGG, ALAW, MULAWAudio file format: WAV, MP3, OGG (Opus), or ALAW / MULAW for telephony.
speednumber10.25 to 2Playback rate. 1.0 = natural; <1 slower, >1 faster.
volume_gainnumber0-96 to 16Output gain in dB. 0 = unchanged.
sample_rateenum240008000, 16000, 22050, 24000, 44100, 48000Output sample rate in Hz (8000, 16000, 22050, 24000, 44100, or 48000).
style_promptstring--Natural-language style direction (e.g. "warm, conversational" or "newscaster, serious").

Good to know

Limits

  • Text: up to 5,000 characters per request
  • Audio billing: 32 output tokens per second of generated audio
  • 130 languages. The language of the text is detected automatically.

Voices

  • The 30 voices in the voice list speak every supported language.
  • GET /v1/voices?model=gemini-3-8-flash-tts lists the full library of more than 2,000 voices across 30 locales, with the language, accent, gender, pitch, and a description of each. Filter it with language, gender, pitch, or q. Any id it returns works as voice or voice2; when you also set language, use the voice's own language.

Custom voices

  • Set voice to custom and pass two recordings of the same adult speaker as https URLs: voice_audio_url, 10 to 30 seconds of natural speech, and voice_consent_audio_url, the speaker reading "I am the owner of this voice and I consent to Google using this voice to create a synthetic voice model." The statement is accepted in 30 languages; consent_statements on that parameter lists each wording.
  • Record both in a quiet room on the same microphone. WAV, MP3, M4A, OGG, WEBM and FLAC are accepted, up to 15 MB each.
  • The response includes voice_id and voice_id_expires_at. Pass the voice_id as voice or voice2 to reuse the voice for 7 days without new recordings. Anyone with the id can use the voice until it expires, so keep it private.

Delivery direction

  • Put direction for the whole passage in style_prompt rather than in the text, for example "warm, unhurried, and reassuring".
  • Place vocal tags inline where the sound should happen. Use these English tags even when the text is in another language:
<argh> <breath> <heavy breath> <exhales> <cackle> <cheer> <chuckle> <chuckles> <cough> <cry> <gasp> <giggle> <groan> <growl> <grunt> <grr> <hiss> <laugh> <laughter> <moan> <pant> <pff> <phew> <scream> <shout> <shriek> <sigh> <sighs> <sneeze> <snicker> <snort> <sob> <throat-clearing> <tsk> <whimper> <whispers> <whispering> <yawn> <short pause> <long pause>
  • Capitalize a word to stress it. Wrap a listener reaction in pipes, for example |mm-hmm|, to add it without starting a new turn.

Multi-speaker

  • Up to two speakers per request. Start each line with a speaker name and a colon, for example Speaker1: and Speaker2:. The names must match speaker1_name and speaker2_name.

Output

  • WAV, MP3, OGG (Opus), ALAW, or MULAW, at 8,000 to 48,000 Hz.
  • POST /v1/audio/speech:stream sends 16-bit PCM chunks at 24,000 Hz while the audio is generated, then the complete file in the requested format. speed is available on POST /v1/audio/speech only.

Gemini 3.8 Flash TTS API: common questions

How much does the Gemini 3.8 Flash TTS API cost?

On EmpirioLabs, Gemini 3.8 Flash TTS is billed pay as you go: Input $2.60 per 1M prompt tokens; Output $46.80 per 1M generated tokens. The live rate card on this page always matches what the API charges.

Which endpoint does Gemini 3.8 Flash TTS use?

Gemini 3.8 Flash TTS is served through POST /v1/audio/speech on api.empiriolabs.ai with standard bearer-token authentication.

Can I try Gemini 3.8 Flash TTS in the browser before integrating?

Yes. The EmpirioLabs playground runs Gemini 3.8 Flash TTS in the browser with the same parameters the API exposes, so you can test prompts before writing code.

How do I get a Gemini 3.8 Flash TTS API key?

Create an EmpirioLabs account, then generate a key under API Keys in the dashboard. Billing is pay-as-you-go credits, so you only pay for the requests you make.

Ready to use better endpoints?

Check out our pricing or reach out if you want your own model deployed on our stack.