StepAudio 3 Gen API

Generates dialogue, sound effects, ambience, and background music together in one clip from a written scene.

StepFunAudio GenerationReleased Sep 9, 2026InternationalProprietary EndpointNew

About StepAudio 3 Gen

Generates dialogue, sound effects, ambience, and background music together in one clip from a written scene.

Character counting follows StepFun pricing: one Chinese character counts as one character, two English letters count as one character, and two punctuation marks count as one character.

Also known as StepFun StepAudio 3 Gen

audio generationtext to speechsound effectsvoice control

StepAudio 3 Gen specs

Model ID
stepaudio-3-gen
Author
StepFun
Category
Audio Generation
Released
Sep 9, 2026
Input
Text
Output
Audio
Region
International
Endpoints
POST/v1/audio/generations
Alternate model IDs
stepaudio-3-gen-previewstepfun/stepaudio-3-gen

StepAudio 3 Gen API pricing

Live pay-as-you-go rates from the EmpirioLabs catalog. You are billed only for what you use, with no monthly minimum.

Type
Spec
Rate
Generation
per 10,000 characters
$0.36
Compare on the full pricing page

How to call the StepAudio 3 Gen API

StepAudio 3 Gen runs through POST /v1/audio/generations. The request returns a job_id right away; poll GET /v1/jobs/{job_id} until the job completes and read the output URLs from the result. Get an API key from the EmpirioLabs dashboard.

cURL: submit the job
curl https://api.empiriolabs.ai/v1/audio/generations \
  -H "Authorization: Bearer $EMPIRIOLABS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "stepaudio-3-gen",
    "prompt": "Describe what you want StepAudio 3 Gen to generate."
  }'
cURL: poll for the result
curl https://api.empiriolabs.ai/v1/jobs/JOB_ID \
  -H "Authorization: Bearer $EMPIRIOLABS_API_KEY"
Python
import requests

response = requests.post(
    "https://api.empiriolabs.ai/v1/audio/generations",
    headers={"Authorization": "Bearer YOUR_EMPIRIOLABS_API_KEY"},
    json={
        "model": "stepaudio-3-gen",
        "prompt": "Describe what you want StepAudio 3 Gen to generate.",
    },
)
job = response.json()

# Generation runs as an async job. Poll until it completes.
import time
while True:
    status = requests.get(
        f"https://api.empiriolabs.ai/v1/jobs/{job['job_id']}",
        headers={"Authorization": "Bearer YOUR_EMPIRIOLABS_API_KEY"},
    ).json()
    if status.get("status") in ("completed", "failed"):
        print(status)
        break
    time.sleep(5)
Full StepAudio 3 Gen API reference

StepAudio 3 Gen API parameters

Request parameters supported by the StepAudio 3 Gen API on EmpirioLabs. Defaults apply when a field is omitted.

ParameterTypeDefaultRange / valuesDescription
promptstring--The scene to generate. Wrap a sound effect or music cue in square brackets and a delivery direction in parentheses.
instructionstring-max 500Global guidance for the setting, background music, and emotional tone. Maximum 500 characters.
rolesstring--Optional speakers, as objects with a name and a voice description. All names and descriptions together are capped at 500 characters.
scriptsstring--Optional ordered lines, as objects with a speaker and text. Capped at 1,000 characters in total. Use this instead of prompt for multi-speaker scenes.
response_formatenummp3mp3, wav, flac, opus, pcmOutput audio format.
sample_rateenum240008000, 16000, 22050, 24000, 48000Output sample rate in Hz.
speednumber10.5 to 2Speech speed.
volumenumber10.1 to 2Output volume.
text_normalizationenumstandardstandard, enhancedHow numbers, dates, and symbols are read aloud.

Good to know

Describe a scene and the model performs it as one finished clip: spoken lines, sound effects, ambience, and background music together. Send a prompt, or use roles and scripts for multi-speaker control. Wrap a sound effect or music cue in square brackets and a delivery direction in parentheses. Roles and instruction are capped at 500 characters each and scripts at 1,000. Output formats are mp3, wav, flac, opus, and pcm. Reference voices and voice cloning are not available on this model. This model is in preview and its capabilities may change.

StepAudio 3 Gen API: common questions

How much does the StepAudio 3 Gen API cost?

On EmpirioLabs, StepAudio 3 Gen is billed pay as you go: Generation $0.36 per 10,000 characters. The live rate card on this page always matches what the API charges.

Which endpoint does StepAudio 3 Gen use?

StepAudio 3 Gen is served through POST /v1/audio/generations on api.empiriolabs.ai with standard bearer-token authentication.

Can I try StepAudio 3 Gen in the browser before integrating?

Yes. The EmpirioLabs playground runs StepAudio 3 Gen in the browser with the same parameters the API exposes, so you can test prompts before writing code.

How do I get a StepAudio 3 Gen API key?

Create an EmpirioLabs account, then generate a key under API Keys in the dashboard. Billing is pay-as-you-go credits, so you only pay for the requests you make.

Ready to use better endpoints?

Check out our pricing or reach out if you want your own model deployed on our stack.