
Expressive speech with inline vocal tags, style direction, and two-speaker dialogue, across 130 languages and a library of more than 2,000 voices.
Expressive speech with inline vocal tags, style direction, and two-speaker dialogue, across 130 languages and a library of more than 2,000 voices.
Also known as Gemini Flash TTS, Google Gemini 3.8 Flash TTS
gemini-3-8-flash-tts/v1/audio/speechPOST/v1/audio/speech:streamGET/v1/voicesgemini-3.8-flash-ttsgoogle/gemini-3.8-flash-ttsLive pay-as-you-go rates from the EmpirioLabs catalog. You are billed only for what you use, with no monthly minimum.
Gemini 3.8 Flash TTS serves speech through POST /v1/audio/speech and returns playable audio. Send the text to speak as input with the model id gemini-3-8-flash-tts. Get an API key from the EmpirioLabs dashboard.
curl https://api.empiriolabs.ai/v1/audio/speech \
-H "Authorization: Bearer $EMPIRIOLABS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-3-8-flash-tts",
"input": "Welcome to EmpirioLabs. Your build just finished."
}' \
--output speech.mp3import requests
response = requests.post(
"https://api.empiriolabs.ai/v1/audio/speech",
headers={"Authorization": "Bearer YOUR_EMPIRIOLABS_API_KEY"},
json={"model": "gemini-3-8-flash-tts", "input": "Welcome to EmpirioLabs."},
)
with open("speech.mp3", "wb") as f:
f.write(response.content)Request parameters supported by the Gemini 3.8 Flash TTS API on EmpirioLabs. Defaults apply when a field is omitted.
| Parameter | Type | Default | Range / values | Description |
|---|---|---|---|---|
| input | string | - | max 5000 | Text to convert to speech. For multi-speaker mode, prefix lines with Speaker1: / Speaker2:. Up to 5,000 characters. Put delivery direction in style_prompt, and place... |
| mode | enum | single | single, multi | single = one voice, multi = two-voice dialogue (uses voice + voice2 + speaker names). |
| language | string | - | - | Optional BCP-47 language tag (en-US, es-ES, etc.). The language of the text is detected automatically. The 30 listed voices speak every supported language; a library... |
| voice | enum | Charon | Zephyr, Puck, Charon, Kore, Fenrir, Leda, Orus, Aoede, Callir... | Primary voice name (e.g. Kore, Puck, Aoede). Leave blank for the default. These 30 voices speak every supported language. GET /v1/voices lists the full library of... |
| voice_audio_url | string | - | - | voice=custom only: an https URL to a 10 to 30 second recording of clean, natural speech from the voice to clone (WAV, MP3, M4A, OGG, WEBM or FLAC). Record it in a... |
| voice_consent_audio_url | string | - | - | voice=custom only: an https URL to the same speaker reading this statement aloud: "I am the owner of this voice and I consent to Google using this voice to create a... |
| voice2 | enum | Kore | Zephyr, Puck, Charon, Kore, Fenrir, Leda, Orus, Aoede, Callir... | Second voice name for multi-speaker mode. |
| speaker1_name | string | Speaker1 | - | Display name used in the input prefix for speaker 1 (default: Speaker1). |
| speaker2_name | string | Speaker2 | - | Display name used in the input prefix for speaker 2 (default: Speaker2). |
| output_format | enum | WAV | WAV, MP3, OGG, ALAW, MULAW | Audio file format: WAV, MP3, OGG (Opus), or ALAW / MULAW for telephony. |
| speed | number | 1 | 0.25 to 2 | Playback rate. 1.0 = natural; <1 slower, >1 faster. |
| volume_gain | number | 0 | -96 to 16 | Output gain in dB. 0 = unchanged. |
| sample_rate | enum | 24000 | 8000, 16000, 22050, 24000, 44100, 48000 | Output sample rate in Hz (8000, 16000, 22050, 24000, 44100, or 48000). |
| style_prompt | string | - | - | Natural-language style direction (e.g. "warm, conversational" or "newscaster, serious"). |
<argh> <breath> <heavy breath> <exhales> <cackle> <cheer> <chuckle> <chuckles> <cough> <cry> <gasp> <giggle> <groan> <growl> <grunt> <grr> <hiss> <laugh> <laughter> <moan> <pant> <pff> <phew> <scream> <shout> <shriek> <sigh> <sighs> <sneeze> <snicker> <snort> <sob> <throat-clearing> <tsk> <whimper> <whispers> <whispering> <yawn> <short pause> <long pause>On EmpirioLabs, Gemini 3.8 Flash TTS is billed pay as you go: Input $2.60 per 1M prompt tokens; Output $46.80 per 1M generated tokens. The live rate card on this page always matches what the API charges.
Gemini 3.8 Flash TTS is served through POST /v1/audio/speech on api.empiriolabs.ai with standard bearer-token authentication.
Yes. The EmpirioLabs playground runs Gemini 3.8 Flash TTS in the browser with the same parameters the API exposes, so you can test prompts before writing code.
Create an EmpirioLabs account, then generate a key under API Keys in the dashboard. Billing is pay-as-you-go credits, so you only pay for the requests you make.
Check out our pricing or reach out if you want your own model deployed on our stack.