Aplomb 1, our first model, is now available. Read the launch post
Browse our latest posts to learn more about us and what our platform has to offer.
API Access
Seedream 5.0 Flash is available on EmpirioLabs for text-to-image generation, edits with up to 10 reference images, coordinate and marker edits, and layer decomposition.
Three new Qwen speech-to-text models: ASR for recorded audio, ASR Stream for live captions, and ASR Message for voice input. Here is how to call them on EmpirioLabs.
Two Gemini realtime voice models on one WebSocket endpoint: Live for low-latency conversation, and Live Extended Thinking for questions that need deeper reasoning. Here is how to call them on EmpirioLabs.
Two new Gemini text-to-speech models: Flash for narration, characters, and dialogue, and Flash-Lite for high-volume voice agents and dubbing. Here is how to call them on EmpirioLabs.
Two new Qwen voice models you talk to over a WebSocket. One of them watches live video while it listens, the other searches the web mid-conversation.
Xiaomi's MiMo V2.6 series is on EmpirioLabs: three models with a 1M token context, native image, video, and audio input, and reasoning on by default.
Step 5 Preview is StepFun's frontier reasoning model, with a 1,024,000 token context, image and video input, parallel tool calling, and strict JSON Schema output. Here is how to call it on EmpirioLabs.
Qwen3.8 Omni Flash reads text, images, audio, and video and answers in text. Qwen3.8 LiveTranslate Flash Realtime is the spoken-translation socket that ships alongside it.
StepAudio 3 ASR Max is StepFun's largest speech recognition model, built for proper nouns, specialist vocabulary, and audio that ordinary transcription struggles with. Here is how to call it on EmpirioLabs.
StepAudio 3 TTS turns text into human-level speech across six languages, with twelve system voices and delivery you describe in plain language. Here is how to call it on EmpirioLabs.
DeepSeek V4.1 Flash is available on EmpirioLabs with native image understanding, hybrid thinking, function calling, and a 1M token context window.
Fugu Ultra v2.0 is Sakana's flagship multi-agent conductor, coordinating its strongest pool of expert models on hard, multi-step problems. Here is how to call it, how the context tier actually works, and what changed from v1.1.
Showing 1–12 of 49 articles
Explore our models, or contact us about business inquiries, custom deployments, or anything else.