
Text-to-video and image-to-video with synchronized native audio, at 720p or 1080p for 3 to 15 seconds, with aspect ratio and prompt control.
Pay as you go, or choose a plan with weekly allowances.
Pricing catalog
EmpirioLabs pricing varies by model and unit: tokens, images, seconds of audio or video, messages, 3D assets, and search requests. The interactive pricing table loads the current catalog, while these representative rates remain readable without client JavaScript.
Image input $0.05 per image. Video output starts at $0.096 per second for 480p and $0.168 per second for 720p.
Input starts at $0.40 per 1M prompt tokens for a cost-effective long-context multimodal API.
Input starts at $0.30 per 1M prompt tokens for multimodal reasoning, coding, and agent workloads.
Image-to-3D generation is priced per generated 3D asset, with docs for format and resolution controls.
Prices are listed and billed in USD. Checkout may show local payment methods when eligible.
Weekly allowances across models, media, search and tasks, on top of pay as you go. Usage inside your allowances is included and never touches your credit balance.
3 of these are free to run, so they never draw from your allowance.
Qwen3.8 27B
Muse Glimmer 30B
DeepSeek V4 Flash 0731
Qwen3.7 Flash
Qwen3.7 Flash (Variant 1)
MiniMax M3
MiniMax M3 (Priority)
Step 3.7 Flash
DeepSeek V4 Flash (Variant 2)
DeepSeek V4 Flash
DeepSeek V4 Flash (Variant 1)
Qwen3.6 Flash
Qwen3.6 Flash (Variant 1)
Qwen3.6 35B A3B
Step 3.5 Flash 2603
Gemma 4 26B-A4B
MiniMax M2.7 Highspeed
MiniMax M2.7
Mistral Small 4
Qwen3.5 4B
Qwen3.5 122B-A10B
Qwen3.5 35B-A3B
Qwen3.5 27B
Qwen3.5 Flash (Variant 1)
Qwen3.5 Flash
Qwen3.5 9B
Qwen3.5 397B-A17B
Seed 2.0 Mini
Step 3.5 Flash
GLM 4.7 FlashFreeThis model is free to run, so using it never draws from your weekly allowance.
GLM 4.6V FlashFreeThis model is free to run, so using it never draws from your weekly allowance.
GLM 4.5 FlashFreeThis model is free to run, so using it never draws from your weekly allowance.
Mistral Medium 3
Mistral Small 3.1
Gemma 3 27B
Nova Lite 1.0
Nova Micro 1.0
DeepSeek V4 Pro 0813
Muse Spark 1.2
Qwen3.8 Max
Qwen3.8 Max (Variant 1)
Fugu Ultra v1.1
Kimi K3
Seed 2.1 Turbo
Muse Spark 1.1
Fugu Ultra v1.0
Kimi K2.7 Code
GLM 5.2
Kimi K2.7 Code Highspeed
GLM 5.2 (Variant 1)
Kimi K2.7 Code (Variant 1)
Qwen3.7 Plus
Qwen3.7 Plus (Variant 1)
Qwen3.7 Max
Qwen3.7 Max (Variant 1)
MiMo V2.5 Pro
DeepSeek V4 Pro (Variant 2)
DeepSeek V4 Pro
DeepSeek V4 Pro (Variant 1)
Qwen3.6 27B
MiMo V2.5
Kimi K2.6
Qwen3.6 Max Preview
GLM 5.1
Qwen3.6 Plus (Variant 1)
Qwen3.6 Plus
Qwen3.5 Omni Flash
Qwen3.5 Omni Plus
Qwen3.5 Plus (Variant 1)
Qwen3.5 Plus
Seed 2.0 Code
Seed 2.0 Lite
Seed 2.0 Pro
Nova Lite 2
DeepSeek V3.2
Qwen3 Max
Qwen3 Max Thinking
Qwen3 Max Preview
Mistral Medium 3.1
Nova Premier 1.0
DeepReasoning
Nova Pro 1.0
Perplexity Sonar
StepAudio 2.5 ChatAllowances cover models, media, search and tasks, including generation templates and Compose. Past them, usage bills at normal pay-as-you-go rates. Personal use only, no reselling. Allowances refresh weekly and do not roll over.

Text-to-video and image-to-video with synchronized native audio, at 720p or 1080p for 3 to 15 seconds, with aspect ratio and prompt control.

Reasoning and coding model with a 1M token context, 128K output, adjustable reasoning effort, native web search, and tool calling.
Reasoning and coding model with a 1M token context, 128K output, adjustable reasoning effort, native web search, and tool calling.

Kimi K3 is Moonshot's flagship reasoning model with a 1M token context, always-on thinking, native web search, and text, image, and video inputs.

Meta's updated frontier reasoning model with a 1,048,576-token context, image, video, audio, and PDF understanding, web search, and tool calling.

Moonshot AIInternationalReleased Jun 16, 2026Ctx 256KText GenerationKimi K2.7 Code is Moonshot's trillion-parameter agentic coding model with 256K context, always-on reasoning, and text, image, and video inputs.

Kimi K2.7 Code is Moonshot's trillion-parameter agentic coding model with 256K context, always-on reasoning, and text, image, and video inputs.

Meta frontier reasoning model with a 1,048,576-token context plus image, video, audio, and PDF understanding, web search, and tool calling.

Updated multi-agent conductor for hard reasoning, coding, and research, with distinct max effort, 1M context, image input, and web search.

Cost-effective Qwen3.7 vision-language model for text, image, video, coding, tool use, GUI understanding, and 1M-context workflows.

Cost-effective Qwen3.7 vision-language model for text, image, video, coding, tool use, GUI understanding, and 1M-context workflows.

Moonshot AIInternationalReleased Jun 16, 2026Ctx 256KText GenerationKimi K2.7 Code Highspeed is the faster-serving tier of Moonshot's agentic coding model, with 256K context, always-on reasoning, and image and video input.

Original multi-agent conductor for hard reasoning, coding, and research, with 1M context, image input, function calling, and web search.

Fast Qwen3.7 vision-language model for text, image, video, tool use, and agentic tasks, with implicit caching and a 1M token context.

Fast Qwen3.7 vision-language model for text, image, video, tool use, and agentic tasks, with implicit caching and a 1M token context.

MiniMax M3 is a multimodal reasoning model for coding, agents, and long-context analysis with text, image, and video input.

MiniMax M3 is a multimodal reasoning model for coding, agents, and long-context analysis with text, image, and video input.

Qwen3.7 Max is a flagship text model for coding, productivity, long-running agents, deep thinking, tools, and 1M-token context.

Qwen3.7 Max is a flagship text model for coding, productivity, long-running agents, deep thinking, tools, and 1M-token context.

Trillion-scale MoE flagship for coding, long-horizon agents, and professional work, with image and video understanding across a 1M-token context.

Trillion-scale MoE flagship for coding, long-horizon agents, and professional work, with image and video understanding across a 1M-token context.
Open-source music generation model for text-to-song and lyric-guided audio, with fast 8-step XL Turbo inference for controllable song iteration.

Apache-licensed 4B FLUX.2 Klein image generation and editing model with text-to-image, reference-image editing, and creative workflow support.

MiniMaxSingaporeReleased Mar 18, 2026Ctx 200KText GenerationHigh-speed M2.7 variant tuned for fast inference with strong general-purpose performance with strong agentic capabilities.
Realtime voice model with plain-English voice direction, one voice identity across 100+ languages, and sub-200ms streaming time-to-first-audio.
Sub-130ms TTFB voice synthesis with 271+ voices across 15 languages, expressive prosody, and real-time SSE streaming for low-latency voice agents.
Broadcast-quality voice synthesis with rich expressive prosody, 271+ voices across 15 languages, and real-time SSE streaming with per-word timestamps.
Long-context Zhipu AI reasoning model with 202K context, 128K output, tool calling, structured output, and cache support.

Kimi K2.6 is a Moonshot multimodal reasoning model with 256K context, strong coding, and text, image, and video inputs.

DeepSeekGermanyReleased Jul 31, 2026Ctx 1MText GenerationPost-trained 0731 release with major gains across coding, repository work, tool use, and full-stack tasks, plus a 1M context window.
Generates complete stereo songs from lyrics and a style prompt, with duration, seed, and format controls up to 5 minutes.

DeepSeekSingaporeReleased Aug 13, 2026Ctx 1MText GenerationOfficial 0813 Pro release with major gains across coding, repository work, tool use, and agent tasks, plus hybrid thinking and a 1M context window.

MiniMax M2.7 is a general-purpose reasoning chat model with interleaved thinking, function calling, and prompt caching.
TRELLIS.2 image-to-3D model that turns a reference image into a textured GLB asset with resolution, seed, mesh, texture, and export controls.

Alibaba CloudChinaReleased Feb 24, 2026Ctx 256KText GenerationQwen3.5 122B-A10B is a multimodal reasoning model with 256K context, efficient sparse MoE inference, and text, image, and video input.

Alibaba CloudChinaReleased Feb 16, 2026Ctx 256KText GenerationQwen3.5 397B-A17B is a flagship multimodal reasoning model for language, code, agents, GUI tasks, and image and video understanding.

Qwen3.5 35B-A3B is an efficient native vision-language model with sparse MoE routing, deep thinking, and text, image, and video input.

Qwen3.5 27B is a dense multimodal reasoning model with fast responses, 256K context, and text, image, and video understanding.

Multimodal reasoner with 256K context, image and video input, function tools, structured JSON, and thinking on by default.

Qwen3.6 27B improves agentic coding, STEM reasoning, spatial vision, OCR, and text, image, and video understanding on 256K context.

Fast Qwen3.6 vision-language model for agentic coding, math reasoning, spatial understanding, OCR, and text, image, and video input.

Fast Qwen3.6 vision-language model for agentic coding, math reasoning, spatial understanding, OCR, and text, image, and video input.
Gemma 4 26B A4B is a Google open multimodal model with 256K context, text, image, and video input, tools, and structured output.

Meta open 30B agentic model with image understanding, 128K context, tool calling, structured output, and controllable reasoning strength.

Qwen3.5 9B is a compact multimodal reasoning model with 256K context, image and video input, function tools, and structured output.

Qwen3.5 4B is a low-cost multimodal reasoning model with 256K context, image and video input, function tools, and structured output.

Qwen3.6 35B A3B is a 256-expert mixture-of-experts reasoning model with 128K context, function tools, and strict structured JSON output.

Free lightweight GLM-4.7 text model for coding, reasoning, long-context writing, and general chat.

Free lightweight GLM-4.5 text model for reasoning, coding, long-form chat, and general language tasks.

Free multimodal GLM-4.6V model for image, video, file, and text understanding with native function calling.

Image generation and editing model creating and modifying images from text or image inputs, with inpainting, virtual try-on, and style controls.

Video generation model producing up to 2-minute multi-shot videos from text and optional image prompts with improved quality and consistency.

Speech-to-text transcription using the Nova-3 model with multi-language support and advanced customizable settings for production workloads.

Pairs DeepSeek R1 chain-of-thought reasoning with Anthropic Claude creative and code generation behind a unified, data-controlled interface.

Open-source Mixture-of-Experts LLM tuned for high-efficiency reasoning, coding, and general language tasks across long-form prompts.

Lightweight MoE model with 284B total / 13B active parameters and native 1M context, tuned for low-latency, cost-effective high-concurrency use.

Lightweight MoE model with 284B total / 13B active parameters and native 1M context, tuned for low-latency, cost-effective high-concurrency use.

Lightweight MoE model with 284B total / 13B active parameters and native 1M context, tuned for low-latency, cost-effective high-concurrency use.

Flagship MoE LLM with 1.6T total / 49B active parameters and native 1M context for advanced math, logical inference, and specialized coding.

Flagship MoE LLM with 1.6T total / 49B active parameters and native 1M context for advanced math, logical inference, and specialized coding.

Flagship MoE LLM with 1.6T total / 49B active parameters and native 1M context for advanced math, logical inference, and specialized coding.

Quick LLM-style answer to a natural-language question, grounded in fresh Exa web search results with inline citations and source links.

Web search engine for finding pages, retrieving similar pages, crawling, and dedicated code search across the open web for AI agents.

Low-latency text-to-speech with single- and multi-speaker voices and controllable style, accent, and expressive tone for production apps.

High-quality TTS preview for podcasts, audiobooks, and customer support, with expressive multi-speaker voices across 23+ languages.

Highly controllable TTS with new Audio Tags for precise style, tone, pace, and delivery across narration, assistants, and voice apps.

Open-source vision-language model with 128K context, 140+ languages, improved math/reasoning, structured outputs, and function calling.

LLM-based text-to-speech with zero-shot voice cloning from 3-10s of audio and emotion-expressive, controllable output via multi-reward RL.

Deep-learning detector that flags portions of text likely generated by AI versus human, classifying content as entirely human, AI, or mixed.

Video model offering Text-to-Video, Image-to-Video, Reference-to-Video, and Video Edit modes with high-fidelity, motion-smooth output.

Open-source text-to-image model on a multimodal Mixture-of-Experts architecture with photorealistic detail and strong multilingual text rendering.
8.3B-parameter video model with native 720p output (upscalable to 1080p), strong motion coherence, and bilingual prompt understanding up to 10s.

Autoregressive framework on the Janus Pro 7B model that unifies multimodal understanding and image generation in one architecture.

Video model in Standard or Pro modes with Text-to-Video, Image-to-Video, Reference-to-Video, editing, native sound, and multi-scene transitions.

Kling 3.0 model that transfers motion from a reference video onto a character from a reference image, with Standard 720p and Pro 1080p tiers.

Iterative AI search that keeps querying when initial results are insufficient, returning more comprehensive answers than Standard mode.

AI-powered web search with detailed overviews and answers, faster than Deep Search. Ranks #1 on OpenAI SimpleQA benchmark.

Cost-efficient language model offering strong reasoning and multimodal performance for general production workloads at competitive latency.

Enterprise-grade model with strong reasoning, coding, and STEM performance, supporting hybrid, on-prem, and in-VPC deployments.

24B-parameter multimodal model with 128K context for image analysis, programming, math, and multilingual tasks, tuned for efficient local inference.

Hybrid model unifying Instruct, Reasoning (Magistral), and Devstral families: 40% lower completion time and 3x throughput vs Small 3.

Open-source 32B MoE foundation model that generates synchronized video and audio in one inference step with precise dual-tower lip-sync.

Low-cost multimodal foundation model for text, images, and video on a 300K context (up to ~30 min video), tuned for speed and affordability.

Fast, cost-effective multimodal reasoning model for text, images, documents, and video on a 1M context (long docs and ~90 min clips).

Text-only foundation model tuned for ultra-low latency and cost on 128K context. Strong for summarization, translation, and chat with 44% cache discount.

Most capable model in the family. Multimodal text/image/video on a 1M context with chain-of-thought reasoning across tools and data sources.

Multimodal foundation model balancing accuracy, speed, and cost for text, images, and video on 300K context (up to ~30 min video).

Whisper-1 speech-to-text transcription trained on multilingual supervised audio, with a 25 MB upload limit per file.

Institutional-grade research powered by Claude Opus 4.6 reasoning, with maximum depth, enhanced tool access, and extensive source coverage.

Research model for multi-step retrieval, synthesis, and reasoning, autonomously searching, reading, and evaluating sources across complex topics.

Sonar Pro as an agentic researcher: chains web searches, fetches full pages, and streams live reasoning, adapting strategy for complex queries.

Real-time web search with filtering by domain, language, date, and more. Returns search results, not LLM responses; no file uploads.

Real-time web-connected search with accurate citations and customizable sources for up-to-date AI search integration in production apps.

Search-grounded model with double the citations and a larger context window, tuned for complex queries needing in-depth, nuanced answers.

Reasoning model on the uncensored open-source R1-1776 with web search, outperforming leading search engines and LLMs on the SimpleQA benchmark.

Cinematic video generation in Text-to-Video, Image-to-Video, and Transition modes with high detail, fluid motion, and lifelike animations.

Generates videos from text or 1-2 frame image prompts up to 1080p, multiple aspect ratios, 5-10s durations, with optional synchronized audio.

Unified image generation and editing model with class-leading complex Chinese/English text rendering, realistic textures, and multi-image fusion.

Qwen image generation and editing with Base and Pro variants, multilingual typography, multi-image references, and controllable 1K or 2K output.

Vision-language model with hybrid linear-attention plus sparse MoE, 1M context, and fast multimodal text/image/video inference.

Vision-language model with hybrid linear-attention plus sparse MoE, 1M context, and fast multimodal text/image/video inference.

Alibaba CloudSingaporeReleased Mar 30, 2026Ctx 256KText GenerationCost-efficient omni-modal model handling text, image, audio, and video, with up to 3 hours of audio and 1 hour of video across 90+ languages.

Alibaba CloudSingaporeReleased Mar 30, 2026Ctx 256KText GenerationFlagship omni-modal model for text, image, audio, and video. 3h audio, 1h video, 90+ input and 30+ output languages, 55 voice timbres.

Multimodal model with hybrid architecture for efficient deep thinking and visual understanding across text, image, and video on a 1M context.

Multimodal model with hybrid architecture for efficient deep thinking and visual understanding across text, image, and video on a 1M context.

Alibaba CloudSingaporeReleased Apr 20, 2026Ctx 256KText GenerationLargest preview variant in the 3.6 series (text-only): improved coding agent execution, stronger front-end skills, and broader long-tail knowledge.

Vision-language model with major upgrades over 3.5: agentic and front-end coding, multimodal recognition, OCR, and object localization.

Vision-language model with major upgrades over 3.5: agentic and front-end coding, multimodal recognition, OCR, and object localization.

256K-context flagship with major improvements in reasoning, instruction following, and multilingual support, plus higher coding/math accuracy.

Alibaba CloudSingaporeReleased Sep 5, 2025Ctx 256KText GenerationPreview release with major gains over the 2.5 series in Chinese-English understanding, complex instructions, multilingual ability, and tool use.

Alibaba CloudSingaporeReleased Sep 23, 2025Ctx 256KText GenerationReasoning model with adaptive tool use (search, memory, code interpreter) and test-time scaling for higher accuracy on complex tasks.

Semantic document reranker. Sorts up to 500 candidates per query by relevance, supports 100+ languages, and accepts a custom sorting instruction.

Coding-tuned 256K-context model with strong front-end results and multilingual programming support for AI coding tools and agents.

Balanced general-purpose model for high-frequency enterprise workloads: information processing, content, search, and data analysis.

Latency-focused multimodal model with 256K context, four reasoning effort modes, and image/video understanding for high-concurrency use.

Flagship general model with 256K context for complex reasoning, multimodal understanding, structured generation, and tool-augmented execution.

Speed-optimized 2.0 video variant for cinematic clips with native audio sync, camera control, and stable motion at lower cost per render.

Multimodal video model for cinematic output from text, image, audio, or video inputs, with stable motion and consistent characters.

Unified multimodal image model that reasons through prompts before rendering, producing high-resolution and consistent edits and brand visuals.

Premium Seedream image model that splits one image into editable layers, edits by coordinate or sketch marker, and fuses up to 10 references.

Open-source voice model for long-form, multi-speaker podcast dialogue with paralinguistic control (laughter, sighs) and zero-shot voice cloning.

Generates audio up to 3 minutes from text prompts, supporting text-to-audio and audio-to-audio with adjustable duration, steps, and CFG scale.

Up-to-3-minute audio from text with text-to-audio, audio-to-audio, and audio inpainting for music production, sound design, and remixing.

Stable Video Infinity 2.0 Pro on WAN 2.2: extends still images into theoretically infinite-length video while keeping consistent character IDs.

Multi-search research assistant that explores a topic, analyzes sources, and produces a detailed research report with citations.

Web search with crawl, extract, and URL mapping for fast, structured retrieval across pages and domains for downstream pipelines.

Multilingual text embedding with selectable output dimensions (64–2048). Up to 8,192 tokens per input.

Alibaba CloudSingaporeReleased Sep 23, 2025Ctx 1KEmbeddingsSpeed-optimised multimodal embedding, same shape as Vision-Plus, 3× cheaper image/video tokens.

Alibaba CloudSingaporeReleased Sep 23, 2025Ctx 1KEmbeddingsMultimodal embedding producing independent vectors for text, image, and video inputs.

Multimodal video generation model for cinematic, multi-shot stories with native audio-visual sync (lip-sync, dialogue, music, SFX).

Multimodal video model supporting T2V, I2V, video editing, and reference-to-video, with high-fidelity output from text, image, or video inputs.

Image generation and editing companion model: text-to-image, bounding-box edits, and cohesive image sets, with up to 4K output on Pro.
Controlled Whisper Large v3 Turbo transcription with multilingual ASR, translation, VAD, timestamps, subtitles, hotwords, and decoder controls.

Autonomous AI agent that turns a high-level prompt into subtasks, calls tools and APIs, and delivers end-to-end results without manual orchestration.

Text-to-video, image-to-video, and reference-to-video generation with up to seven reference images for consistent characters, up to 15 seconds at 1080p.

Top-tier model for agentic workflows, complex software engineering, and long-horizon tasks, sustaining work across 1000+ tool calls on 1M context.

Multimodal model with native visual and audio understanding on a 1M context, designed to reason and act across modalities in agentic workflows.

Text, image, and reference-to-video in one model. Cinematic motion, character consistency across up to 9 references, and synchronized native audio.

The fastest, most affordable Seedance 2.0 tier for short cinematic clips with native audio, camera control, and image or video inputs at 480p and 720p.

StepFunInternationalReleased May 28, 2026Ctx 256KText GenerationStepFun multimodal reasoning model with image and video input, tool calling, adjustable reasoning effort, and 256K context.

StepFunInternationalReleased Feb 12, 2026Ctx 256KText GenerationStepFun text reasoning model for agents, coding, tool calling, and long-context analysis.

StepFunInternationalReleased Apr 2, 2026Ctx 256KText GenerationAgent-optimized Step 3.5 Flash variant with low and high reasoning effort modes.

StepFun audio and text conversation model with text output and paralinguistic understanding.

StepFun image generation and image editing model for text-to-image and single-image edits.

Contextual StepFun text-to-speech model with natural-language voice direction and expressive delivery.

StepFun text-to-speech model with official voices, custom cloned voices, and voice tag controls.

StepFun streaming speech recognition model for Chinese and English audio transcription.

Next-generation coding and agent model with engineering-grade code delivery, long-horizon autonomy, and 256K multimodal understanding.

Tiered speech synthesis with over 1,000 voices, 16 languages, 20 Chinese dialects, natural-language delivery direction, and inline emotion tags.

Generates 4 to 15 second clips at up to 2K with native stereo audio, following text, image, video, and audio references in one request.

OpenAI's flagship image model with strong prompt fidelity, crisp text rendering, and instruction-based editing across up to 16 reference images.

Kuaishou's Kling 3.0 video generator with text-to-video, first and last frame image-to-video, native audio, and 720p, 1080p, or 4K output.

Alibaba's Wan 2.5 video model with text-to-video and image-to-video, native audio, optional custom audio tracks, and 480p to 1080p output.

Alibaba's Wan 2.2 video family with standard and flash tiers, first and last frame interpolation, and 480p to 1080p output at low cost.

Alibaba's Wan 2.1 video model with turbo and plus tiers for text-to-video, image-to-video, and first and last frame clips at 480p or 720p.

Alibaba's Wan 2.5 image model with text-to-image and multi-reference editing of up to 3 images at about 1.7 megapixels.

Alibaba's Wan 2.2 text-to-image model with plus and flash tiers and up to 1440x1440 output at low per-image pricing.

Alibaba's Wan 2.1 text-to-image model with turbo and plus tiers and up to 1440x1440 output at low per-image pricing.

Long-form video model for coherent clips up to 30 seconds, with up to 50 reference images, videos, and audio clips, native audio, editing, and extension.

Text-to-image plus multi-reference editing with up to three source images, selectable quality tiers, and 1K or 2K output resolution.

Generates video from text, images, first and last frames, or up to seven references, with native audio, multi-shot cuts, and clip extension.

Cinema-grade video model for action and effects work, from text, images, first and last frames, or references, with native synchronized audio.

Aligns mouth movement in an existing video to an uploaded audio track or to text spoken by one of fourteen built-in voices.
Turns a single portrait into a talking avatar video, driven by an uploaded audio track or by text spoken with a built-in voice.

Six video transformations in one model: prompt editing, restyling, subject swap, motion transfer, upscaling, and generated sound effects.
150 of 165 models