
Generates dialogue, sound effects, ambience, and background music together in one clip from a written scene.
Generates dialogue, sound effects, ambience, and background music together in one clip from a written scene.
Character counting follows StepFun pricing: one Chinese character counts as one character, two English letters count as one character, and two punctuation marks count as one character.
Also known as StepFun StepAudio 3 Gen
stepaudio-3-gen/v1/audio/generationsstepaudio-3-gen-previewstepfun/stepaudio-3-genLive pay-as-you-go rates from the EmpirioLabs catalog. You are billed only for what you use, with no monthly minimum.
StepAudio 3 Gen runs through POST /v1/audio/generations. The request returns a job_id right away; poll GET /v1/jobs/{job_id} until the job completes and read the output URLs from the result. Get an API key from the EmpirioLabs dashboard.
curl https://api.empiriolabs.ai/v1/audio/generations \
-H "Authorization: Bearer $EMPIRIOLABS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "stepaudio-3-gen",
"prompt": "Describe what you want StepAudio 3 Gen to generate."
}'curl https://api.empiriolabs.ai/v1/jobs/JOB_ID \
-H "Authorization: Bearer $EMPIRIOLABS_API_KEY"import requests
response = requests.post(
"https://api.empiriolabs.ai/v1/audio/generations",
headers={"Authorization": "Bearer YOUR_EMPIRIOLABS_API_KEY"},
json={
"model": "stepaudio-3-gen",
"prompt": "Describe what you want StepAudio 3 Gen to generate.",
},
)
job = response.json()
# Generation runs as an async job. Poll until it completes.
import time
while True:
status = requests.get(
f"https://api.empiriolabs.ai/v1/jobs/{job['job_id']}",
headers={"Authorization": "Bearer YOUR_EMPIRIOLABS_API_KEY"},
).json()
if status.get("status") in ("completed", "failed"):
print(status)
break
time.sleep(5)Request parameters supported by the StepAudio 3 Gen API on EmpirioLabs. Defaults apply when a field is omitted.
| Parameter | Type | Default | Range / values | Description |
|---|---|---|---|---|
| prompt | string | - | - | The scene to generate. Wrap a sound effect or music cue in square brackets and a delivery direction in parentheses. |
| instruction | string | - | max 500 | Global guidance for the setting, background music, and emotional tone. Maximum 500 characters. |
| roles | string | - | - | Optional speakers, as objects with a name and a voice description. All names and descriptions together are capped at 500 characters. |
| scripts | string | - | - | Optional ordered lines, as objects with a speaker and text. Capped at 1,000 characters in total. Use this instead of prompt for multi-speaker scenes. |
| response_format | enum | mp3 | mp3, wav, flac, opus, pcm | Output audio format. |
| sample_rate | enum | 24000 | 8000, 16000, 22050, 24000, 48000 | Output sample rate in Hz. |
| speed | number | 1 | 0.5 to 2 | Speech speed. |
| volume | number | 1 | 0.1 to 2 | Output volume. |
| text_normalization | enum | standard | standard, enhanced | How numbers, dates, and symbols are read aloud. |
Describe a scene and the model performs it as one finished clip: spoken lines, sound effects, ambience, and background music together. Send a prompt, or use roles and scripts for multi-speaker control. Wrap a sound effect or music cue in square brackets and a delivery direction in parentheses. Roles and instruction are capped at 500 characters each and scripts at 1,000. Output formats are mp3, wav, flac, opus, and pcm. Reference voices and voice cloning are not available on this model. This model is in preview and its capabilities may change.
On EmpirioLabs, StepAudio 3 Gen is billed pay as you go: Generation $0.36 per 10,000 characters. The live rate card on this page always matches what the API charges.
StepAudio 3 Gen is served through POST /v1/audio/generations on api.empiriolabs.ai with standard bearer-token authentication.
Yes. The EmpirioLabs playground runs StepAudio 3 Gen in the browser with the same parameters the API exposes, so you can test prompts before writing code.
Create an EmpirioLabs account, then generate a key under API Keys in the dashboard. Billing is pay-as-you-go credits, so you only pay for the requests you make.
Check out our pricing or reach out if you want your own model deployed on our stack.