Qwen Audio 3.0 TTS API

Tiered speech synthesis with over 1,000 voices, 16 languages, 20 Chinese dialects, natural-language delivery direction, and inline emotion tags.

Alibaba CloudGeração de ÁudioLançado 20 de jul. de 2026SingaporeEndpoint proprietárioNovo

Sobre Qwen Audio 3.0 TTS

Tiered speech synthesis with over 1,000 voices, 16 languages, 20 Chinese dialects, natural-language delivery direction, and inline emotion tags.

Voice selection is tier aware: base voice ids work on both tiers and are automatically matched to the selected model_tier, while the six system voices are locked to one tier. Pitch also shifts the pace of the rendered audio, so pair it with speed when you want to preserve the original duration.

Também conhecido como Qwen Audio TTS, Alibaba Cloud Qwen Audio 3.0 TTS, Qwen-Audio-3.0-TTS, qwen-audio-3-0-tts

text to speechmultilingualexpressive prosodyemotion controlvoice tagsdialectlow latency

Especificações de Qwen Audio 3.0 TTS

ID do modelo
qwen-audio-3-0-tts
Provedor
Alibaba Cloud
Categoria
Geração de Áudio
Lançado
20 de jul. de 2026
Entrada
Texto
Saída
Áudio
Região
Singapore
Endpoints
POST/v1/audio/speechPOST/v1/audio/speech:streamGET/v1/voices
IDs de modelo alternativos
qwen-audio-3.0-ttsalibaba/qwen-audio-3-0-ttsqwen-audio-3.0-tts-plusqwen-audio-3.0-tts-flash

Preços da API Qwen Audio 3.0 TTS

Tarifas pay-as-you-go ao vivo do catálogo EmpirioLabs. Você paga só pelo que usa, sem mínimo mensal.

Tipo
Especificação
Tarifa
Plus synthesis
per 10,000 characters
$0.20
Flash synthesis
per 10,000 characters
$0.15
Comparar na página completa de preços

Como chamar a API Qwen Audio 3.0 TTS

Qwen Audio 3.0 TTS gera fala por POST /v1/audio/speech e retorna áudio reproduzível. Envie o texto a ser falado em input com o id de modelo qwen-audio-3-0-tts. Obtenha uma chave de API no painel EmpirioLabs.

cURL
curl https://api.empiriolabs.ai/v1/audio/speech \
  -H "Authorization: Bearer $EMPIRIOLABS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen-audio-3-0-tts",
    "input": "Welcome to EmpirioLabs. Your build just finished."
  }' \
  --output speech.mp3
Python
import requests

response = requests.post(
    "https://api.empiriolabs.ai/v1/audio/speech",
    headers={"Authorization": "Bearer YOUR_EMPIRIOLABS_API_KEY"},
    json={"model": "qwen-audio-3-0-tts", "input": "Welcome to EmpirioLabs."},
)
with open("speech.mp3", "wb") as f:
    f.write(response.content)
Referência completa da API Qwen Audio 3.0 TTS

Parâmetros da API Qwen Audio 3.0 TTS

Parâmetros de requisição suportados pela API Qwen Audio 3.0 TTS na EmpirioLabs. Os padrões valem quando um campo é omitido.

ParâmetroTipoPadrãoIntervalo / valoresDescrição
inputstring-máx. 20000Text to synthesize. Up to 20,000 characters per request. Supports inline expression tags placed directly in the text, for example [excited], [laughing], [whispers],...
model_tierenumplusplus, flashplus: highest audio quality and expressiveness, best for content creation, audiobooks, dubbing, and brand voice work. flash: tuned for real-time interaction with...
voiceenumloongjameszhaoloongjameszhao, loongolivialin, loonglunawang, loongnorahu, l...Voice preset. The base voices work on BOTH tiers, so changing model_tier keeps your selection. Six flagship voices are tier locked: longanlingxin and longanlufeng on...
voice_idstring--Free-form voice id, which overrides voice when set. Accepts any base voice from GET /v1/voices, either as the bare id (recommended, works on both tiers) or fully...
response_formatenummp3mp3, wav, pcm, opusAudio container. mp3 and opus are compressed, wav is uncompressed PCM in a RIFF header, and pcm is headerless raw samples for chunked playback.
sample_rateenum240008000, 16000, 22050, 24000, 44100, 48000Output sample rate in Hz. 24000 suits speech playback, and 48000 gives broadcast-quality output at a larger file size.
speednumber10.5 a 2Speaking rate multiplier. 0.5 is half speed and 2.0 is double speed.
pitchnumber10.5 a 2Pitch multiplier. Values below 1.0 lower the voice and values above raise it. Pitch also shifts the pace of the rendered audio, so pair it with speed when you want...
volumeinteger500 a 100Output loudness, where 50 is the reference level.
instructionstring-máx. 128Natural-language direction for delivery, up to 128 characters. Controls emotion, tone, character, pace, and speaking style, for example 'Speak quickly in an excited,...
language_hintsarray-zh, enLanguage codes that bias pronunciation for mixed-language text, for example ["zh", "en"]. Leave unset to let the model detect the language.
seedinteger00 a 65535Sampling seed. Reuse a seed with identical input and settings for a repeatable render.
bit_rateinteger3200016000 a 64000Encoder bitrate in bps. Applies to the opus format only, and is ignored for mp3, wav, and pcm.
pronunciationobject--Pronunciation overrides keyed by the written form, for example {"重要": "zhong4 yao4"}. Use it to fix names, acronyms, and homographs.
2 parâmetros a mais na documentação

Bom saber

Tiers

  • Plus: highest audio quality and expressiveness, for content creation, audiobooks, dubbing, and brand voice work
  • Flash: tuned for real-time interaction with lower first-packet latency, for voice assistants and live agents
  • Switch with the model_tier parameter, or send qwen-audio-3.0-tts-plus or qwen-audio-3.0-tts-flash as the model id

Voices

  • Over 1,000 base voices, all available on both tiers, so changing tier keeps your selected voice
  • Six flagship system voices are tier specific: longanlingxin and longanlufeng on Plus, and longanhuan_v3.6, longjielidou_v3.6, loongeva_v3.6, loongjohn on Flash
  • Browse the full catalog with GET /v1/voices, and pass any voice through voice_id

Input limits

  • Up to 20,000 characters per request
  • instruction accepts up to 128 characters
  • SSML input is not supported. Use instruction for delivery direction, and inline tags such as [excited] or [whispers] inside the text

Billing

  • Billed per character of input text at the selected tier's rate

API Qwen Audio 3.0 TTS: perguntas frequentes

Quanto custa a API Qwen Audio 3.0 TTS?

Na EmpirioLabs, Qwen Audio 3.0 TTS é cobrado por uso: Plus synthesis $0.20 per 10,000 characters; Flash synthesis $0.15 per 10,000 characters. A tabela de tarifas ao vivo desta página sempre corresponde ao que a API cobra.

Qual endpoint Qwen Audio 3.0 TTS usa?

Qwen Audio 3.0 TTS é servido por POST /v1/audio/speech em api.empiriolabs.ai com autenticação padrão por token Bearer.

Posso testar Qwen Audio 3.0 TTS no navegador antes de integrar?

Sim. O playground da EmpirioLabs executa Qwen Audio 3.0 TTS no navegador com os mesmos parâmetros que a API expõe, para você testar prompts antes de escrever código.

Como consigo uma chave de API Qwen Audio 3.0 TTS?

Crie uma conta EmpirioLabs e gere uma chave em API Keys no painel. A cobrança usa créditos pay-as-you-go, então você paga apenas pelas requisições que faz.

Pronto para usar endpoints melhores?

Confira nossos preços ou entre em contato se quiser que seu próprio modelo seja implementado em nossa pilha.