
Tiered speech synthesis with over 1,000 voices, 16 languages, 20 Chinese dialects, natural-language delivery direction, and inline emotion tags.
Tiered speech synthesis with over 1,000 voices, 16 languages, 20 Chinese dialects, natural-language delivery direction, and inline emotion tags.
Voice selection is tier aware: base voice ids work on both tiers and are automatically matched to the selected model_tier, while the six system voices are locked to one tier. Pitch also shifts the pace of the rendered audio, so pair it with speed when you want to preserve the original duration.
يعرف أيضا باسم Qwen Audio TTS, Alibaba Cloud Qwen Audio 3.0 TTS, Qwen-Audio-3.0-TTS, qwen-audio-3-0-tts
qwen-audio-3-0-tts/v1/audio/speechPOST/v1/audio/speech:streamGET/v1/voicesqwen-audio-3.0-ttsalibaba/qwen-audio-3-0-ttsqwen-audio-3.0-tts-plusqwen-audio-3.0-tts-flashأسعار الدفع حسب الاستخدام مباشرة من كتالوج EmpirioLabs. تدفع فقط مقابل ما تستخدمه، بدون حد أدنى شهري.
يولّد Qwen Audio 3.0 TTS الكلام عبر POST /v1/audio/speech ويعيد صوتًا قابلًا للتشغيل. أرسل النص المطلوب نطقه في input مع معرّف النموذج qwen-audio-3-0-tts. احصل على مفتاح API من لوحة تحكم EmpirioLabs.
curl https://api.empiriolabs.ai/v1/audio/speech \
-H "Authorization: Bearer $EMPIRIOLABS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen-audio-3-0-tts",
"input": "Welcome to EmpirioLabs. Your build just finished."
}' \
--output speech.mp3import requests
response = requests.post(
"https://api.empiriolabs.ai/v1/audio/speech",
headers={"Authorization": "Bearer YOUR_EMPIRIOLABS_API_KEY"},
json={"model": "qwen-audio-3-0-tts", "input": "Welcome to EmpirioLabs."},
)
with open("speech.mp3", "wb") as f:
f.write(response.content)معلمات الطلب التي تدعمها واجهة Qwen Audio 3.0 TTS API على EmpirioLabs. تُطبق القيم الافتراضية عند حذف أي حقل.
| المعامل | النوع | الافتراضي | النطاق / القيم | الوصف |
|---|---|---|---|---|
| input | string | - | الحد الأقصى 20000 | Text to synthesize. Up to 20,000 characters per request. Supports inline expression tags placed directly in the text, for example [excited], [laughing], [whispers],... |
| model_tier | enum | plus | plus, flash | plus: highest audio quality and expressiveness, best for content creation, audiobooks, dubbing, and brand voice work. flash: tuned for real-time interaction with... |
| voice | enum | loongjameszhao | loongjameszhao, loongolivialin, loonglunawang, loongnorahu, l... | Voice preset. The base voices work on BOTH tiers, so changing model_tier keeps your selection. Six flagship voices are tier locked: longanlingxin and longanlufeng on... |
| voice_id | string | - | - | Free-form voice id, which overrides voice when set. Accepts any base voice from GET /v1/voices, either as the bare id (recommended, works on both tiers) or fully... |
| response_format | enum | mp3 | mp3, wav, pcm, opus | Audio container. mp3 and opus are compressed, wav is uncompressed PCM in a RIFF header, and pcm is headerless raw samples for chunked playback. |
| sample_rate | enum | 24000 | 8000, 16000, 22050, 24000, 44100, 48000 | Output sample rate in Hz. 24000 suits speech playback, and 48000 gives broadcast-quality output at a larger file size. |
| speed | number | 1 | 0.5 إلى 2 | Speaking rate multiplier. 0.5 is half speed and 2.0 is double speed. |
| pitch | number | 1 | 0.5 إلى 2 | Pitch multiplier. Values below 1.0 lower the voice and values above raise it. Pitch also shifts the pace of the rendered audio, so pair it with speed when you want... |
| volume | integer | 50 | 0 إلى 100 | Output loudness, where 50 is the reference level. |
| instruction | string | - | الحد الأقصى 128 | Natural-language direction for delivery, up to 128 characters. Controls emotion, tone, character, pace, and speaking style, for example 'Speak quickly in an excited,... |
| language_hints | array | - | zh, en | Language codes that bias pronunciation for mixed-language text, for example ["zh", "en"]. Leave unset to let the model detect the language. |
| seed | integer | 0 | 0 إلى 65535 | Sampling seed. Reuse a seed with identical input and settings for a repeatable render. |
| bit_rate | integer | 32000 | 16000 إلى 64000 | Encoder bitrate in bps. Applies to the opus format only, and is ignored for mp3, wav, and pcm. |
| pronunciation | object | - | - | Pronunciation overrides keyed by the written form, for example {"重要": "zhong4 yao4"}. Use it to fix names, acronyms, and homographs. |
على EmpirioLabs، تتم فوترة Qwen Audio 3.0 TTS حسب الاستخدام: Plus synthesis $0.20 per 10,000 characters; Flash synthesis $0.15 per 10,000 characters. جدول الأسعار المباشر في هذه الصفحة يطابق دائمًا ما تحتسبه الواجهة.
يُقدَّم Qwen Audio 3.0 TTS عبر POST /v1/audio/speech على api.empiriolabs.ai مع مصادقة Bearer القياسية.
نعم. يشغّل ملعب EmpirioLabs نموذج Qwen Audio 3.0 TTS في المتصفح بنفس المعلمات التي تتيحها الواجهة، لتجربة المطالبات قبل كتابة الكود.
أنشئ حساب EmpirioLabs ثم أنشئ مفتاحًا من API Keys في لوحة التحكم. الفوترة برصيد الدفع حسب الاستخدام، فلا تدفع إلا مقابل الطلبات التي تنفذها.
تحقق من تسعيرنا أو تواصل إذا كنت ترغب في نشر نموذجك الخاص على مجموعتنا.