Aplomb 1, our first model, is now available. Read the launch post
Browse our latest posts to learn more about us and what our platform has to offer.
API Access
Two new Gemini text-to-speech models: Flash for narration, characters, and dialogue, and Flash-Lite for high-volume voice agents and dubbing. Here is how to call them on EmpirioLabs.
Two new Qwen voice models you talk to over a WebSocket. One of them watches live video while it listens, the other searches the web mid-conversation.
Xiaomi's MiMo V2.6 series is on EmpirioLabs: three models with a 1M token context, native image, video, and audio input, and reasoning on by default.
Step 5 Preview is StepFun's frontier reasoning model, with a 1,024,000 token context, image and video input, parallel tool calling, and strict JSON Schema output. Here is how to call it on EmpirioLabs.
Qwen3.8 Omni Flash reads text, images, audio, and video and answers in text. Qwen3.8 LiveTranslate Flash Realtime is the spoken-translation socket that ships alongside it.
StepAudio 3 ASR Max is StepFun's largest speech recognition model, built for proper nouns, specialist vocabulary, and audio that ordinary transcription struggles with. Here is how to call it on EmpirioLabs.
StepAudio 3 TTS turns text into human-level speech across six languages, with twelve system voices and delivery you describe in plain language. Here is how to call it on EmpirioLabs.
DeepSeek V4.1 Flash is available on EmpirioLabs with native image understanding, hybrid thinking, function calling, and a 1M token context window.
Fugu Ultra v2.0 is Sakana's flagship multi-agent conductor, coordinating its strongest pool of expert models on hard, multi-step problems. Here is how to call it, how the context tier actually works, and what changed from v1.1.
Fugu Max is Sakana's cost-efficient multi-agent conductor, orchestrating a wide pool of open and specialist models behind one OpenAI-compatible endpoint. Here is how to call it, what it bills for, and the three things callers get wrong.
TTS 2 Flash is available on EmpirioLabs: latency-first speech synthesis that holds one voice identity across 200+ languages, with streaming audio and word-level timestamps for captioning.
Muse Spark 1.3 is available on EmpirioLabs with a 1,048,576-token context, image, video, and PDF understanding, always-on reasoning, web search, and tool calling.
Showing 13–24 of 112 articles
Explore our models, or contact us about business inquiries, custom deployments, or anything else.