StepAudio 3 Realtime API

Speech in, speech out, over one socket. Interruptible mid-reply, with voice activity detection deciding when to answer.

StepFunAudio Generation256K contextReleased Aug 6, 2026InternationalProprietary EndpointNew

About StepAudio 3 Realtime

Speech in, speech out, over one socket. Interruptible mid-reply, with voice activity detection deciding when to answer.

Session is configured with session.update events over the socket rather than request parameters.

Also known as StepFun StepAudio 3 Realtime

realtimespeech to speechaudio inaudio outvoice control

StepAudio 3 Realtime specs

Model ID
stepaudio-3-realtime
Author
StepFun
Category
Audio Generation
Released
Aug 6, 2026
Context window
256K tokens
Max output
131,072 tokens
Input
AudioText
Output
AudioText
Region
International
Endpoints
WEBSOCKET/v1/realtime
Alternate model IDs
stepaudio-3-realtime-previewstepfun/stepaudio-3-realtime

StepAudio 3 Realtime API pricing

Live pay-as-you-go rates from the EmpirioLabs catalog. You are billed only for what you use, with no monthly minimum.

Type
Spec
Rate
Input
per 1M prompt tokens
$1.50
Output
per 1M generated tokens
$10.00
Implicit cache read
per 1M cached input tokens
$0.30
Compare on the full pricing page

How to call the StepAudio 3 Realtime API

StepAudio 3 Realtime holds a live conversation over a WebSocket at wss://api.empiriolabs.ai/v1/realtime?model=stepaudio-3-realtime, not over the HTTP endpoints. Connect with the model id stepaudio-3-realtime and send your EmpirioLabs API key as an Authorization: Bearer header on the handshake. Audio travels both ways as base64 16-bit PCM. A browser cannot set headers on a WebSocket, so open this connection from your server and relay audio to the browser over your own socket. Get an API key from the EmpirioLabs dashboard.

Python (websockets)
import asyncio, json, os, websockets

async def main():
    async with websockets.connect(
        "wss://api.empiriolabs.ai/v1/realtime?model=stepaudio-3-realtime",
        additional_headers=[
            ("Authorization", f"Bearer {os.environ['EMPIRIOLABS_API_KEY']}"),
        ],
        max_size=None,
    ) as ws:
        print(json.loads(await ws.recv())["type"])  # session.created

        await ws.send(json.dumps({
            "type": "conversation.item.create",
            "item": {
                "type": "message",
                "role": "user",
                "content": [{"type": "input_text", "text": "Say hello."}],
            },
        }))
        await ws.send(json.dumps({"type": "response.create"}))

        async for raw in ws:
            event = json.loads(raw)
            if event["type"] == "response.audio.delta":
                ...  # base64 audio chunk, append to your playback buffer
            elif event["type"] == "response.done":
                break

asyncio.run(main())
Full StepAudio 3 Realtime API reference

StepAudio 3 Realtime API parameters

Request parameters supported by the StepAudio 3 Realtime API on EmpirioLabs. Defaults apply when a field is omitted.

ParameterTypeDefaultRange / valuesDescription
voiceenumsoft-spoken-gentlemansoft-spoken-gentleman, magnetic-voiced-male, vibrant-youth, l...Session voice, set with session.update before the model produces audio. It cannot be changed once the model has spoken, and any other value is rejected.
instructionsstring--System instructions for the session: persona, speaking style, and boundaries.
modalitiesstring["text","audio"]-Response modalities for the session.
volume_rationumber10.1 to 2Loudness of the spoken reply, relative to the default. Accepted range is 0.1 to 2.0.
input_audio_formatenumpcm16pcm16Encoding of the audio you send. 16-bit PCM.
output_audio_formatenumpcm16pcm16Encoding of the audio the model returns. 16-bit PCM.
turn_detectionstring{"type":"server_vad"}-Server-side voice activity detection. Decides when you have stopped speaking and the model should reply. Without it the model listens but never answers, so leave it...

Good to know

Full-duplex voice over a WebSocket at wss://api.empiriolabs.ai/v1/realtime?model=stepaudio-3-realtime, authenticated with the ordinary Authorization Bearer header. Audio is pcm16 in and out. The model listens while it speaks, so it can be interrupted mid-reply, and server-side voice activity detection decides when to answer. Set voice with session.update before the first audio; it cannot be changed once the model has spoken, and only the seven listed voices are accepted. Conversation is Chinese and English only. Each completed turn is billed on its own from the usage the model reports. This model is in preview and its capabilities may change.

StepAudio 3 Realtime API: common questions

How much does the StepAudio 3 Realtime API cost?

On EmpirioLabs, StepAudio 3 Realtime is billed pay as you go: Input $1.50 per 1M prompt tokens; Output $10.00 per 1M generated tokens; Implicit cache read $0.30 per 1M cached input tokens. The live rate card on this page always matches what the API charges.

What is the context window of StepAudio 3 Realtime?

StepAudio 3 Realtime supports a 256K-token context window with up to 131,072 output tokens per response.

Which endpoint does StepAudio 3 Realtime use?

StepAudio 3 Realtime is served through WEBSOCKET /v1/realtime on api.empiriolabs.ai with standard bearer-token authentication.

Can I try StepAudio 3 Realtime in the browser before integrating?

Open the EmpirioLabs playground and press Start session to try it in your browser. StepAudio 3 Realtime runs over a WebSocket rather than a request and response, so to build with it start from the realtime voice quickstart. Connect from your server with your EmpirioLabs API key and the model id stepaudio-3-realtime, then stream audio in both directions.

How do I get a StepAudio 3 Realtime API key?

Create an EmpirioLabs account, then generate a key under API Keys in the dashboard. Billing is pay-as-you-go credits, so you only pay for the requests you make.

Ready to use better endpoints?

Check out our pricing or reach out if you want your own model deployed on our stack.