# EmpirioLabs

> Specialized AI model hosting for open, proprietary, and custom stacks. 163+ models across chat, image, video, audio, embeddings, rerankers, and 3D, all behind one OpenAI-compatible API with one bill and one shared rate-limit pool.

Public API base URL: https://api.empiriolabs.ai
Auth: send `Authorization: Bearer <EMPIRIOLABS_API_KEY>` or `x-api-key: <EMPIRIOLABS_API_KEY>` on every request.
Get an API key: https://platform.empiriolabs.ai/dashboard/api-keys

Live machine-readable model catalog: https://api.empiriolabs.ai/v1/models
Per-model parameter schema: https://api.empiriolabs.ai/v1/models/{model_id}
Deeper docs index for AI agents: https://docs.empiriolabs.ai/llms.txt

## Docs

- [Welcome](https://docs.empiriolabs.ai/welcome): overview and quickstart
- [Getting Started](https://docs.empiriolabs.ai/getting-started): first request walkthrough
- [Authentication](https://docs.empiriolabs.ai/authentication): API key usage
- [Models and Pricing](https://docs.empiriolabs.ai/models-pricing): full per-model price table
- [OpenAI and Anthropic Compatibility](https://docs.empiriolabs.ai/compatibility): drop-in replacement notes
- [Integrations](https://docs.empiriolabs.ai/integrations): Continue, Cursor, Cline, Open WebUI, OpenCode, and others
- [Generation Templates](https://docs.empiriolabs.ai/generation-templates): pre-curated effect recipes plus Compose full-video productions
- [Rate Limits and API Keys](https://docs.empiriolabs.ai/rate-limits-api-keys): account-level limits
- [Billing and Credits](https://docs.empiriolabs.ai/billing-credits): pay-as-you-go credit model
- [Account Usage API](https://docs.empiriolabs.ai/account-usage-api): usage and saved Playground conversation APIs

## Model categories

- Text Generation: 73 models
- Image Generation: 17 models
- Video Generation: 30 models
- Audio Generation: 16 models
- Transcription: 4 models
- Research & Search: 13 models
- 3D Generation: 3 models
- Embeddings: 4 models
- Rerankers: 1 model
- Tools & Agents: 2 models

## Subscription plans

Optional weekly usage allowances on top of pay as you go. Usage inside an allowance does not draw from credits; past it, requests bill at normal per-use rates.

- Lite ($19.90/month, $6.90/week, $218.90/year): Weekly: 12M standard tokens, 2M premium tokens, 20s video, 7 images.
- Pro ($49.90/month, $16.90/week, $548.90/year): Weekly: 36M standard tokens, 3M premium tokens, 60s video, 25 images.
- Max ($199.90/month, $64.90/week, $2198.90/year): Weekly: 190M standard tokens, 12M premium tokens, 240s video, 70 images.

Full plan comparison: https://empiriolabs.ai/pricing

## FAQ

### How do payments work?

Use pay-as-you-go credits, or choose an optional plan with weekly allowances. Usage inside your allowances is included, and anything beyond them uses your credit balance at the listed rates. Auto top-up and volume bonuses are available where supported.

### What payment methods are supported?

All major payment methods are supported, including credit and debit cards, PayPal, Apple Pay, Google Pay, Cash App, Alipay, WeChat Pay, local bank transfers, cash vouchers, and more. The methods shown at checkout depend on your country and currency.

### Do you support purchases with crypto?

Yes. We support top-ups with crypto via our payment processor.

### Do I have to be a developer to use your platform?

No. You can use everything through the dashboard with no code required. API access is there when you want to connect EmpirioLabs to your own app or workflow.

### Is my data private?

Yes. We do not train on, sell, or share your prompts, files, or outputs, and we do not log your prompt or response content. Anything you choose to save, like playground chat history, is stored securely and can be deleted anytime, and generated media is removed automatically after a limited time.

## Marketing

- [Home](https://empiriolabs.ai/): platform overview
- [Models](https://empiriolabs.ai/models): browse the full catalog
- [Pricing](https://empiriolabs.ai/pricing): per-model price tables
- [Generation Templates](https://empiriolabs.ai/generation-templates): one-click video and image effects plus Compose recipes, each with its own detail page
- [GPU Cloud](https://empiriolabs.ai/gpu-cloud): hourly cloud GPU rates with a detail page per GPU
- [Hosted Agents](https://empiriolabs.ai/hosted-agents): managed always-on agent runtimes and plans
- [Blog](https://empiriolabs.ai/blog): launches and product updates
- [Contact](https://empiriolabs.ai/contact-us)

## Agent-friendly endpoints

- https://empiriolabs.ai/sitemap.xml: sitemap index covering marketing, docs, and platform
- https://empiriolabs.ai/llms-full.txt: this overview plus the full inline model catalog
- https://docs.empiriolabs.ai/ai-agent-full-context.md: entire docs site rendered as one markdown file
- https://docs.empiriolabs.ai/ai-agent-docs-context.md: docs tab only as one markdown file
- https://docs.empiriolabs.ai/ai-agent-api-reference-context.md: API reference tab only as one markdown file

## Full model catalog

### Text Generation

- `glm-5-3` (zhipu): Coding and agentic model with a 1M token context, 128K output, always-on reasoning at low, high, or max effort, native web search, and tool calling. Pricing: Input $1.40 per 1M prompt tokens; Output $4.40 per 1M generated tokens; Web search $0.033 per request.
- `glm-5-3-flash` (zhipu): Natively multimodal flash model with a 1M token context, image and video input, always-on reasoning, native web search, and tool calling. Pricing: Input $0.075 per 1M prompt tokens; Output $0.25 per 1M generated tokens; Web search $0.033 per request.
- `glm-5-2` (zhipu): Reasoning and coding model with a 1M token context, 128K output, adjustable reasoning effort, native web search, and tool calling. Pricing: Input $1.40 per 1M prompt tokens; Output $4.40 per 1M generated tokens; Web search $0.033 per request.
- `kimi-k3` (moonshot): Kimi K3 is Moonshot's flagship reasoning model with a 1M token context, always-on thinking, native web search, and text, image, and video inputs. Pricing: Input $3.00 per 1M prompt tokens; Output $15.00 per 1M generated tokens; Web search $0.015 per call when invoked.
- `muse-spark-1-2` (meta): Meta's updated frontier reasoning model with a 1,048,576-token context, image, video, audio, and PDF understanding, web search, and tool calling. Pricing: Input $1.25 per 1M prompt tokens; Output $4.25 per 1M generated tokens; Implicit cache read $1.00 per 1M cached input tokens.
- `kimi-k2-7-code` (moonshot): Kimi K2.7 Code is Moonshot's trillion-parameter agentic coding model with 256K context, always-on reasoning, and text, image, and video inputs. Pricing: Input $0.95 per 1M prompt tokens; Output $4.00 per 1M generated tokens; Web search $0.015 per call when invoked.
- `muse-spark-1-1` (meta): Meta frontier reasoning model with a 1,048,576-token context plus image, video, audio, and PDF understanding, web search, and tool calling. Pricing: Input $1.25 per 1M prompt tokens; Output $4.25 per 1M generated tokens; Implicit cache read $1.00 per 1M cached input tokens.
- `fugu-ultra-v1-1` (sakana): Updated multi-agent conductor for hard reasoning, coding, and research, with distinct max effort, 1M context, image input, and web search. Pricing: Input <=272K $5.00; >272K $10.00 per 1M prompt tokens; Output <=272K $30.00; >272K $45.00 per 1M generated tokens; Implicit cache read <=272K $0.50; >272K $1.00 per 1M cached input tokens.
- `qwen3-7-plus` (alibaba): Cost-effective Qwen3.7 vision-language model for text, image, video, coding, tool use, GUI understanding, and 1M-context workflows. Pricing: Input <=256K $0.40; 256K-1M $1.20 per 1M prompt tokens; Output <=256K $1.60; 256K-1M $4.80 per 1M generated tokens; Web search $0.03 per call when invoked.
- `kimi-k2-7-code-highspeed` (moonshot): Kimi K2.7 Code Highspeed is the faster-serving tier of Moonshot's agentic coding model, with 256K context, always-on reasoning, and image and video input. Pricing: Input $1.90 per 1M prompt tokens; Output $8.00 per 1M generated tokens; Web search $0.015 per call when invoked.
- `fugu-ultra-v1-0` (sakana): Original multi-agent conductor for hard reasoning, coding, and research, with 1M context, image input, function calling, and web search. Pricing: Input <=272K $7.50; >272K $15.00 per 1M prompt tokens; Output <=272K $45.00; >272K $67.50 per 1M generated tokens; Implicit cache read <=272K $1.50; >272K $3.00 per 1M cached input tokens.
- `qwen3-7-flash` (alibaba): Fast Qwen3.7 vision-language model for text, image, video, tool use, and agentic tasks, with implicit caching and a 1M token context. Pricing: Input <=32K $0.03; 32K-256K $0.10; 256K-1M $0.20 per 1M prompt tokens; Output <=32K $0.13; 32K-256K $0.40; 256K-1M $0.80 per 1M generated tokens; Implicit cache read <=32K $0.006; 32K-256K $0.02; 256K-1M $0.04 per 1M cached input tokens.
- `minimax-m3` (minimax): MiniMax M3 is a multimodal reasoning model for coding, agents, and long-context analysis with text, image, and video input. Pricing: Input <=512K $0.225 (was $0.30); >512K $0.45 (was $0.60) per 1M prompt tokens; Output <=512K $0.90 (was $1.20); >512K $1.80 (was $2.40) per 1M generated tokens; Implicit cache read <=512K $0.045 (was $0.06); >512K $0.09 (was $0.12) per 1M cached input tokens.
- `qwen3-7-max` (alibaba): Qwen3.7 Max is a flagship text model for coding, productivity, long-running agents, deep thinking, tools, and 1M-token context. Pricing: Input $2.50 per 1M prompt tokens; Output $7.50 per 1M generated tokens; Web search $0.02 per call when invoked.
- `qwen3-8-max` (alibaba): Trillion-scale MoE flagship for coding, long-horizon agents, and professional work, with image and video understanding across a 1M-token context. Pricing: Input $2.00 per 1M prompt tokens; Output $6.00 per 1M generated tokens; Web search $0.02 per call when invoked.
- `qwen3-8-flash` (alibaba): Fast Qwen3.8 vision-language model for coding, agents, and visual understanding, with image and video input, five built-in tools, and a 1M token context. Pricing: Input $0.16 per 1M prompt tokens; Output $0.47 per 1M generated tokens; Web search $0.02 per call when invoked.
- `qwen3-8-max-0902` (alibaba): Upgraded Qwen3.8 Max snapshot with stronger coding, steadier multi-tool agent runs, and sharper chart and document vision across a 1M-token context. Pricing: Input $2.00 per 1M prompt tokens; Output $6.00 per 1M generated tokens; Web search $0.02 per call when invoked.
- `minimax-m2-7-highspeed` (minimax): High-speed M2.7 variant tuned for fast inference with strong general-purpose performance with strong agentic capabilities. Pricing: Input $0.30 (was $0.60) per 1M prompt tokens; Output $1.20 (was $2.40) per 1M generated tokens; Implicit cache read $0.03 (was $0.06) per 1M cached input tokens.
- `glm-5-1` (zhipu): Long-context Zhipu AI reasoning model with 202K context, 128K output, tool calling, structured output, and cache support. Pricing: Input <=32K $0.825 (was $1.40); 32K-200K $1.10 (was $1.40) per 1M prompt tokens; Output <=32K $3.301 (was $4.40); 32K-200K $3.851 (was $4.40) per 1M generated tokens; Implicit cache read <=32K $0.165 (was $0.26); 32K-200K $0.22 (was $0.26) per 1M cached input tokens.
- `kimi-k2-6` (moonshot): Kimi K2.6 is a Moonshot multimodal reasoning model with 256K context, strong coding, and text, image, and video inputs. Pricing: Input $0.8939 (was $0.95) per 1M prompt tokens; Output $3.7131 (was $4.00) per 1M generated tokens; Implicit cache read $0.1788 per 1M cached input tokens.
- `deepseek-v4-flash-0731` (deepseek): Post-trained 0731 release with major gains across coding, repository work, tool use, and full-stack tasks, plus a 1M context window. Pricing: Input $0.424 per 1M prompt tokens; Output $1.272 per 1M generated tokens; Web Search (Linkup) $0.013 per call when invoked.
- `deepseek-v4-pro-0813` (deepseek): Official 0813 Pro release with major gains across coding, repository work, tool use, and agent tasks, plus hybrid thinking and a 1M context window. Pricing: Input $1.32 per 1M prompt tokens; Output $3.96 per 1M generated tokens; Web Search (Linkup) $0.013 per call when invoked.
- `muse-spark-1-3` (meta): Meta's long-horizon agentic reasoning model with a 1,048,576-token context, image, video, and PDF understanding, web search, and tool calling. Pricing: Input $1.25 per 1M prompt tokens; Output $4.25 per 1M generated tokens; Implicit cache read $1.00 per 1M cached input tokens.
- `deepseek-v4-1-flash` (deepseek): Lightweight flagship of DeepSeek's new architecture with native image understanding, hybrid thinking, 1M context, and up to 384K output tokens. Pricing: Input $0.30 per 1M prompt tokens; Output $1.20 per 1M generated tokens; Web Search (Linkup) $0.013 per call when invoked.
- `minimax-m2-7` (minimax): MiniMax M2.7 is a general-purpose reasoning chat model with interleaved thinking, function calling, and prompt caching. Pricing: Input $0.15 (was $0.30) per 1M prompt tokens; Output $0.60 (was $1.20) per 1M generated tokens; Implicit cache read $0.03 (was $0.06) per 1M cached input tokens.
- `qwen3-5-122b-a10b` (alibaba): Qwen3.5 122B-A10B is a multimodal reasoning model with 256K context, efficient sparse MoE inference, and text, image, and video input. Pricing: Input <=128K $0.115 (was $0.40); 128K-256K $0.287 (was $0.40) per 1M prompt tokens; Output <=128K $0.917 (was $3.20); 128K-256K $2.294 (was $3.20) per 1M generated tokens; Web search $0.01 per call when invoked.
- `qwen3-5-397b-a17b` (alibaba): Qwen3.5 397B-A17B is a flagship multimodal reasoning model for language, code, agents, GUI tasks, and image and video understanding. Pricing: Input <=128K $0.172 (was $0.60); 128K-256K $0.43 (was $0.60) per 1M prompt tokens; Output <=128K $1.032 (was $3.60); 128K-256K $2.58 (was $3.60) per 1M generated tokens; Web search $0.01 per call when invoked.
- `qwen3-5-35b-a3b` (alibaba): Qwen3.5 35B-A3B is an efficient native vision-language model with sparse MoE routing, deep thinking, and text, image, and video input. Pricing: Input <=128K $0.057 (was $0.25); 128K-256K $0.229 (was $0.25) per 1M prompt tokens; Output <=128K $0.459 (was $2.00); 128K-256K $1.835 (was $2.00) per 1M generated tokens; Web search $0.01 per call when invoked.
- `qwen3-5-27b` (alibaba): Qwen3.5 27B is a dense multimodal reasoning model with fast responses, 256K context, and text, image, and video understanding. Pricing: Input <=128K $0.086 (was $0.30); 128K-256K $0.258 (was $0.30) per 1M prompt tokens; Output <=128K $0.688 (was $2.40); 128K-256K $2.064 (was $2.40) per 1M generated tokens; Web search $0.01 per call when invoked.
- `qwen3-8-27b` (alibaba): Multimodal reasoner with 256K context, image and video input, function tools, structured JSON, and thinking on by default. Pricing: Input $0.17 (was $0.45) per 1M prompt tokens; Output $0.50 (was $3.20) per 1M generated tokens; Implicit cache read $0.08 per 1M cached input tokens.
- `qwen3-6-27b` (alibaba): Qwen3.6 27B improves agentic coding, STEM reasoning, spatial vision, OCR, and text, image, and video understanding on 256K context. Pricing: Input $0.412564 (was $0.60) per 1M prompt tokens; Output $2.475384 (was $3.60) per 1M generated tokens; Web search $0.01 per call when invoked.
- `qwen3-6-flash` (alibaba): Fast Qwen3.6 vision-language model for agentic coding, math reasoning, spatial understanding, OCR, and text, image, and video input. Pricing: Input <=256K $0.25; 256K-1M $1.00 per 1M prompt tokens; Output <=256K $1.50; 256K-1M $4.00 per 1M generated tokens; Web search $0.02 per call when invoked.
- `gemma-4-26b-a4b` (google): Gemma 4 26B A4B is a Google open multimodal model with 256K context, text, image, and video input, tools, and structured output. Pricing: Input $0.05 (was $0.15) per 1M prompt tokens; Output $0.29 (was $0.50) per 1M generated tokens; Implicit cache read $0.025 (was $0.15) per 1M cached input tokens.
- `muse-glimmer-30b` (meta): Meta open 30B agentic model with image understanding, 128K context, tool calling, structured output, and controllable reasoning strength. Pricing: Input $0.20 (was $0.35) per 1M prompt tokens; Output $0.80 (was $1.50) per 1M generated tokens; Implicit cache read $0.05 per 1M cached input tokens.
- `qwen3-5-9b` (alibaba): Qwen3.5 9B is a compact multimodal reasoning model with 256K context, image and video input, function tools, and structured output. Pricing: Input $0.09 (was $0.10) per 1M prompt tokens; Output $0.13 (was $0.15) per 1M generated tokens; Implicit cache read $0.045 per 1M cached input tokens.
- `qwen3-5-4b` (alibaba): Qwen3.5 4B is a low-cost multimodal reasoning model with 256K context, image and video input, function tools, and structured output. Pricing: Input $0.04 per 1M prompt tokens; Output $0.07 per 1M generated tokens; Implicit cache read $0.02 per 1M cached input tokens.
- `qwen3-6-35b-a3b` (alibaba): Qwen3.6 35B A3B is a 256-expert mixture-of-experts reasoning model with 128K context, function tools, and strict structured JSON output. Pricing: Input $0.07 (was $0.248) per 1M prompt tokens; Output $0.42 (was $1.485) per 1M generated tokens; Implicit cache read $0.035 per 1M cached input tokens.
- `glm-4-7-flash` (zhipu): Free lightweight GLM-4.7 text model for coding, reasoning, long-context writing, and general chat. Pricing: Input Free per 1M prompt tokens; Output Free per 1M generated tokens; Implicit cache read Free per 1M cached input tokens.
- `glm-4-5-flash` (zhipu): Free lightweight GLM-4.5 text model for reasoning, coding, long-form chat, and general language tasks. Pricing: Input Free per 1M prompt tokens; Output Free per 1M generated tokens; Implicit cache read Free per 1M cached input tokens.
- `glm-4-6v-flash` (zhipu): Free multimodal GLM-4.6V model for image, video, file, and text understanding with native function calling. Pricing: Input Free per 1M prompt tokens; Output Free per 1M generated tokens; Implicit cache read Free per 1M cached input tokens.
- `deepreasoning` (winfunc): Pairs DeepSeek R1 chain-of-thought reasoning with Anthropic Claude creative and code generation behind a unified, data-controlled interface. Pricing: R1-0528 + Claude Sonnet 4.5 (Default) In $0.012 / Out $0.058 per 1K tokens; R1-0528 + Claude Haiku 4.5 In $0.0048 / Out $0.023 per 1K tokens; R1-0528 + Claude Opus 4.5 In $0.019 / Out $0.092 per 1K tokens.
- `deepseek-v3-2` (deepseek): Open-source Mixture-of-Experts LLM tuned for high-efficiency reasoning, coding, and general language tasks across long-form prompts. Pricing: Input $0.57 per 1M prompt tokens; Output $1.71 per 1M generated tokens; Web search $0.015 per request when enabled.
- `deepseek-v4-flash` (deepseek): Lightweight MoE model with 284B total / 13B active parameters and native 1M context, tuned for low-latency, cost-effective high-concurrency use. Pricing: Input $0.14 per 1M prompt tokens; Output $0.28 per 1M generated tokens; Web Search (Linkup) $0.013 per call when invoked.
- `deepseek-v4-pro` (deepseek): Flagship MoE LLM with 1.6T total / 49B active parameters and native 1M context for advanced math, logical inference, and specialized coding. Pricing: Input $1.65 (was $1.74) per 1M prompt tokens; Output $3.30 (was $3.48) per 1M generated tokens; Web Search (Linkup) $0.013 per call when invoked.
- `gemma-3-27b` (google): Open-source vision-language model with 128K context, 140+ languages, improved math/reasoning, structured outputs, and function calling. Pricing: Per Message $0.0040 fixed; Web Search (Linkup) $0.013 per call when invoked.
- `mistral-small-3-1` (mistral): 24B-parameter multimodal model with 128K context for image analysis, programming, math, and multilingual tasks, tuned for efficient local inference. Pricing: Per Message $0.0019 fixed; Web Search (Linkup) $0.013 per call when invoked.
- `mistral-small-4` (mistral): Hybrid model unifying Instruct, Reasoning (Magistral), and Devstral families: 40% lower completion time and 3x throughput vs Small 3. Pricing: Input $0.15 per 1M prompt tokens; Output $0.60 per 1M generated tokens; Standard Web Search $0.084 per call.
- `nova-lite-1-0` (amazon): Low-cost multimodal foundation model for text, images, and video on a 300K context (up to ~30 min video), tuned for speed and affordability. Pricing: Input $0.069 per 1M prompt tokens; Output $0.28 per 1M generated tokens; Cached input $0.0386 per 1M tokens.
- `nova-lite-2` (amazon): Fast, cost-effective multimodal reasoning model for text, images, documents, and video on a 1M context (long docs and ~90 min clips). Pricing: Input $0.38 per 1M prompt tokens; Output $3.16 per 1M generated tokens; Cached input $0.2128 per 1M tokens.
- `nova-micro-1-0` (amazon): Text-only foundation model tuned for ultra-low latency and cost on 128K context. Strong for summarization, translation, and chat with 44% cache discount. Pricing: Input $0.040 per 1M prompt tokens; Output $0.16 per 1M generated tokens; Cached input $0.0224 per 1M tokens.
- `nova-pro-1-0` (amazon): Multimodal foundation model balancing accuracy, speed, and cost for text, images, and video on 300K context (up to ~30 min video). Pricing: Input $2.40 per 1M prompt tokens; Output $9.60 per 1M generated tokens; Latency Optimized Input $3.00 per 1M prompt tokens.
- `qwen3-5-flash` (alibaba): Vision-language model with hybrid linear-attention plus sparse MoE, 1M context, and fast multimodal text/image/video inference. Pricing: Input $0.090 (was $0.10) per 1M prompt tokens; Output $0.368 (was $0.40) per 1M generated tokens; Web search $0.015 per call when invoked.
- `qwen3-5-omni-flash` (alibaba): Cost-efficient omni-modal model handling text, image, audio, and video, with up to 3 hours of audio and 1 hour of video across 90+ languages. Pricing: Input per 1M prompt tokens $0.40; per 1M prompt tokens $3.00 per 1M prompt tokens; Output per 1M generated tokens $2.20; per 1M generated tokens $11.90 per 1M generated tokens; Web search $0.015 per call when invoked.
- `qwen3-5-omni-plus` (alibaba): Flagship omni-modal model for text, image, audio, and video. 3h audio, 1h video, 90+ input and 30+ output languages, 55 voice timbres. Pricing: Input per 1M prompt tokens $1.40; per 1M prompt tokens $11.00 per 1M prompt tokens; Output per 1M generated tokens $8.30; per 1M generated tokens $44.00 per 1M generated tokens; Web search $0.015 per call when invoked.
- `qwen3-5-plus` (alibaba): Multimodal model with hybrid architecture for efficient deep thinking and visual understanding across text, image, and video on a 1M context. Pricing: Input <=256K $0.36 (was $0.40); 256K-1M $1.08 (was $1.20) per 1M prompt tokens; Output <=256K $2.21 (was $2.40); 256K-1M $6.62 (was $7.20) per 1M generated tokens; Web search $0.015 per call when invoked.
- `qwen3-6-max-preview` (alibaba): Largest preview variant in the 3.6 series (text-only): improved coding agent execution, stronger front-end skills, and broader long-tail knowledge. Pricing: Input <=128K $1.31; 128K-256K $1.97 per 1M prompt tokens; Output <=128K $7.88; 128K-256K $11.82 per 1M generated tokens; Web search $0.020 per call when invoked.
- `qwen3-6-plus` (alibaba): Vision-language model with major upgrades over 3.5: agentic and front-end coding, multimodal recognition, OCR, and object localization. Pricing: Input <=256K $0.50; 256K-1M $2.00 per 1M prompt tokens; Output <=256K $3.00; 256K-1M $6.00 per 1M generated tokens; Web search $0.026 per call when invoked.
- `qwen3-max` (alibaba): 256K-context flagship with major improvements in reasoning, instruction following, and multilingual support, plus higher coding/math accuracy. Pricing: Input <=32K $1.08 (was $1.20); 32K-128K $2.16 (was $2.40); 128K-256K $2.70 (was $3.00) per 1M prompt tokens; Output <=32K $5.52 (was $6.00); 32K-128K $11.04 (was $12.00); 128K-256K $13.80 (was $15.00) per 1M generated tokens; Web search $0.015 per call when invoked.
- `qwen3-max-preview` (alibaba): Preview release with major gains over the 2.5 series in Chinese-English understanding, complex instructions, multilingual ability, and tool use. Pricing: Input <=32K $1.08 (was $1.20); 32K-128K $2.16 (was $2.40); 128K-256K $2.70 (was $3.00) per 1M prompt tokens; Output <=32K $4.80 (was $6.00); 32K-128K $9.60 (was $12.00); 128K-256K $12.00 (was $15.00) per 1M generated tokens; Web search $0.015 per call when invoked.
- `qwen3-max-thinking` (alibaba): Reasoning model with adaptive tool use (search, memory, code interpreter) and test-time scaling for higher accuracy on complex tasks. Pricing: Input <=32K $1.08 (was $1.20); 32K-128K $2.16 (was $2.40); 128K-256K $2.70 (was $3.00) per 1M prompt tokens; Output <=32K $5.52 (was $6.00); 32K-128K $11.04 (was $12.00); 128K-256K $13.80 (was $15.00) per 1M generated tokens; Web search $0.015 per call when invoked.
- `seed-2-0-code` (bytedance): Coding-tuned 256K-context model with strong front-end results and multilingual programming support for AI coding tools and agents. Pricing: Input <=128K $0.40; 128K-256K $0.80 per 1M prompt tokens; Output <=128K $2.40; 128K-256K $4.80 per 1M generated tokens.
- `seed-2-0-lite` (bytedance): Balanced general-purpose model for high-frequency enterprise workloads: information processing, content, search, and data analysis. Pricing: Input <=128K $0.31; 128K-256K $0.62 per 1M prompt tokens; Output <=128K $2.50; 128K-256K $5.00 per 1M generated tokens.
- `seed-2-0-mini` (bytedance): Latency-focused multimodal model with 256K context, four reasoning effort modes, and image/video understanding for high-concurrency use. Pricing: Input <=128K $0.12; 128K-256K $0.24 per 1M prompt tokens; Output <=128K $0.50; 128K-256K $1.00 per 1M generated tokens.
- `seed-2-0-pro` (bytedance): Flagship general model with 256K context for complex reasoning, multimodal understanding, structured generation, and tool-augmented execution. Pricing: Input <=128K $0.63; 128K-256K $1.26 per 1M prompt tokens; Output <=128K $3.79; 128K-256K $7.58 per 1M generated tokens.
- `mimo-v2-5-pro` (xiaomi): Top-tier model for agentic workflows, complex software engineering, and long-horizon tasks, sustaining work across 1000+ tool calls on 1M context. Pricing: Input $2.175 per 1M prompt tokens; Output $4.35 per 1M generated tokens; Implicit cache read $0.018 per 1M cached input tokens.
- `mimo-v2-5` (xiaomi): Multimodal model with native visual and audio understanding on a 1M context, designed to reason and act across modalities in agentic workflows. Pricing: Input $0.70 per 1M prompt tokens; Output $1.40 per 1M generated tokens; Implicit cache read $0.014 per 1M cached input tokens.
- `step-3-7-flash` (stepfun): StepFun multimodal reasoning model with image and video input, tool calling, adjustable reasoning effort, and 256K context. Pricing: Input $0.20 per 1M prompt tokens; Output $1.15 per 1M generated tokens; Implicit cache read $0.04 per 1M cached input tokens.
- `step-3-5-flash` (stepfun): StepFun text reasoning model for agents, coding, tool calling, and long-context analysis. Pricing: Input $0.10 per 1M prompt tokens; Output $0.30 per 1M generated tokens; Implicit cache read $0.02 per 1M cached input tokens.
- `step-3-5-flash-2603` (stepfun): Agent-optimized Step 3.5 Flash variant with low and high reasoning effort modes. Pricing: Input $0.10 per 1M prompt tokens; Output $0.30 per 1M generated tokens; Implicit cache read $0.02 per 1M cached input tokens.
- `stepaudio-2-5-chat` (stepfun): StepFun audio and text conversation model with text output and paralinguistic understanding. Pricing: Input $1.43 per 1M prompt tokens; Output $3.57 per 1M generated tokens; Implicit cache read $0.29 per 1M cached input tokens.
- `seed-2-1-turbo` (bytedance): Next-generation coding and agent model with engineering-grade code delivery, long-horizon autonomy, and 256K multimodal understanding. Pricing: Input $0.63 per 1M prompt tokens; Output $3.13 per 1M generated tokens.
- `fugu-max` (sakana): Cost-efficient multi-agent conductor that assembles a right-sized expert team per task, with 1M context, image input, and web search. Pricing: Input $2.00 per 1M prompt tokens; Output $6.00 per 1M generated tokens; Implicit cache read $0.25 per 1M cached input tokens.
- `fugu-ultra-v2-0` (sakana): Flagship multi-agent conductor for the hardest reasoning, coding, and research work, with 1M context, image input, and web search. Pricing: Input <=272K $5.00; >272K $10.00 per 1M prompt tokens; Output <=272K $30.00; >272K $45.00 per 1M generated tokens; Implicit cache read <=272K $0.50; >272K $1.00 per 1M cached input tokens.

### Image Generation

- `flux-2-klein-4b` (black-forest-labs): Apache-licensed 4B FLUX.2 Klein image generation and editing model with text-to-image, reference-image editing, and creative workflow support. Pricing: Image generation $0.0085 (was $0.014) per image.
- `amazon-nova-canvas` (amazon): Image generation and editing model creating and modifying images from text or image inputs, with inpainting, virtual try-on, and style controls. Pricing: Small Standard (≤1024×1024) $0.12 per image; Small Premium (≤1024×1024) $0.18 per image; Large Standard (≤2048×2048) $0.18 per image.
- `hunyuan-image-3` (tencent): Open-source text-to-image model on a multimodal Mixture-of-Experts architecture with photorealistic detail and strong multilingual text rendering. Pricing: Standard $0.13 per image.
- `janus-pro-deepseek` (deepseek): Autoregressive framework on the Janus Pro 7B model that unifies multimodal understanding and image generation in one architecture. Pricing: Image Generation $0.030 per image; Image Analysis $0.030 per uploaded image.
- `qwen-image-2-0` (alibaba): Unified image generation and editing model with class-leading complex Chinese/English text rendering, realistic textures, and multi-image fusion. Pricing: Standard $0.035 per image; Pro $0.075 per image.
- `qwen-image-3-0` (alibaba): Qwen image generation and editing with Base and Pro variants, multilingual typography, multi-image references, and controllable 1K or 2K output. Pricing: Base Output (1K or 2K) $0.030 per image; Pro Output up to 2.25MP (1K) $0.040 per image; Pro Output above 2.25MP (2K) $0.075 per image.
- `seedream-5-0-lite` (bytedance): Unified multimodal image model that reasons through prompts before rendering, producing high-resolution and consistent edits and brand visuals. Pricing: Standard $0.0350 per image.
- `seedream-5-0-pro` (bytedance): Premium Seedream image model that splits one image into editable layers, edits by coordinate or sketch marker, and fuses up to 10 references. Pricing: Output up to 2.61MP $0.075 per image; Output above 2.61MP $0.150 per image; Layer decomposition up to 2.61MP $0.0375 per image.
- `wan2-7-image` (alibaba): Image generation and editing companion model: text-to-image, bounding-box edits, and cohesive image sets, with up to 4K output on Pro. Pricing: Standard $0.030 per image; Pro $0.075 per image.
- `step-image-edit-2` (stepfun): StepFun image generation and image editing model for text-to-image and single-image edits. Pricing: Output image $0.003 per generated image.
- `gpt-image-2` (openai): OpenAI's flagship image model with strong prompt fidelity, crisp text rendering, and instruction-based editing across up to 16 reference images. Pricing: Low $0.012 per image; Medium $0.106 per image; High $0.422 per image.
- `wan2-5-image` (alibaba): Alibaba's Wan 2.5 image model with text-to-image and multi-reference editing of up to 3 images at about 1.7 megapixels. Pricing: Standard $0.030 per image.
- `wan2-2-image` (alibaba): Alibaba's Wan 2.2 text-to-image model with plus and flash tiers and up to 1440x1440 output at low per-image pricing. Pricing: Plus $0.05 per image; Flash $0.025 per image.
- `wan2-1-image` (alibaba): Alibaba's Wan 2.1 text-to-image model with turbo and plus tiers and up to 1440x1440 output at low per-image pricing. Pricing: Turbo $0.025 per image; Plus $0.05 per image.
- `grok-imagine-image-2-0` (xai): Text-to-image plus multi-reference editing with up to five source images, automatic or pinned quality tiers, and 1K or 2K output resolution. Pricing: Low quality, 1K $0.048 per image; Low quality, 2K $0.072 per image; Medium quality, 1K $0.072 per image.
- `seedream-4-5` (bytedance): High-fidelity image generation and editing with up to 14 reference images, strong text rendering, and 2K or 4K output at one flat price. Pricing: Standard $0.080 per image.
- `seedream-4-0` (bytedance): Unified image generation and editing with up to 14 reference images, batch sets of up to 15, and 1K to 4K output at one flat per-image price. Pricing: Standard $0.060 per image.

### Video Generation

- `kling-3-0-turbo` (kling): Text-to-video and image-to-video with synchronized native audio, at 720p or 1080p for 3 to 15 seconds, with aspect ratio and prompt control. Pricing: 720p $0.18 per second; 1080p $0.225 per second.
- `amazon-nova-reel-1-1` (amazon): Video generation model producing up to 2-minute multi-shot videos from text and optional image prompts with improved quality and consistency. Pricing: Per Second $0.14 per second.
- `happyhorse-1-0` (alibaba): Video model offering Text-to-Video, Image-to-Video, Reference-to-Video, and Video Edit modes with high-fidelity, motion-smooth output. Pricing: All Modes 720P $0.14 per second; All Modes 1080P $0.24 per second.
- `hunyuan-video-1-5` (tencent): 8.3B-parameter video model with native 720p output (upscalable to 1080p), strong motion coherence, and bilingual prompt understanding up to 10s. Pricing: 480p $0.061 (was $0.075) per second; 720p $0.29 per second; 1080p (upscaled) $0.67 per second.
- `kling-o3` (kling): Video model in Standard or Pro modes with Text-to-Video, Image-to-Video, Reference-to-Video, editing, native sound, and multi-scene transitions. Pricing: Standard T2V/I2V $0.168 per second; Standard T2V/I2V Sound $0.224 per second; Standard Video Input $0.252 per second.
- `kling-v3-motion-control` (kling): Kling 3.0 model that transfers motion from a reference video onto a character from a reference image, with Standard 720p and Pro 1080p tiers. Pricing: Standard (720p) $0.14 per second; Pro (1080p) $0.18 per second.
- `moss-video-and-audio` (openmoss): Open-source 32B MoE foundation model that generates synchronized video and audio in one inference step with precise dual-tower lip-sync. Pricing: 360p Video $0.17 per video; 720p Video $2.82 per video; T2V Fast $0.065 additional fee.
- `pixverse-v5` (pixverse): Cinematic video generation in Text-to-Video, Image-to-Video, and Transition modes with high detail, fluid motion, and lifelike animations. Pricing: 360p/540p 5s $0.45 per video; 360p/540p 8s $0.90 per video; 720p 5s $0.60 per video.
- `pixverse-v5-6` (pixverse): Generates videos from text or 1-2 frame image prompts up to 1080p, multiple aspect ratios, 5-10s durations, with optional synchronized audio. Pricing: 360p/540p 5s no audio $0.40 per video; 360p/540p 5s audio $0.80 per video; 360p/540p 8s no audio $0.80 per video.
- `seedance-2-0-fast` (bytedance): Speed-optimized 2.0 video variant for cinematic clips with native audio sync, camera control, and stable motion at lower cost per render. Pricing: T2V/I2V 480P $0.122 per second; T2V/I2V 720P $0.260 per second; Video Input 480P $0.284 per second.
- `seedance-2-0-pro` (bytedance): Multimodal video model for cinematic output from text, image, audio, or video inputs, with stable motion and consistent characters. Pricing: T2V/I2V 480P $0.139 per second; T2V/I2V 720P $0.300 per second; T2V/I2V 1080P $0.749 per second.
- `svi-2-0-pro` (vita-epfl): Stable Video Infinity 2.0 Pro on WAN 2.2: extends still images into theoretically infinite-length video while keeping consistent character IDs. Pricing: 480p Video $0.057 per second; 720p Video $0.17 per second; T2V Fast $0.065 additional fee.
- `wan-2-6` (alibaba): Multimodal video generation model for cinematic, multi-shot stories with native audio-visual sync (lip-sync, dialogue, music, SFX). Pricing: Standard 720P $0.10 per second; Standard 1080P $0.15 per second; Flash 720P (audio) $0.050 per second.
- `wan-2-7` (alibaba): Multimodal video model supporting T2V, I2V, video editing, and reference-to-video, with high-fidelity output from text, image, or video inputs. Pricing: All Modes 720P $0.10 per second; All Modes 1080P $0.150 per second.
- `grok-imagine-video-1-5` (xai): Text-to-video, image-to-video, and reference-to-video generation with up to seven reference images for consistent characters, up to 15 seconds at 1080p. Pricing: 480p $0.096 per second; 720p $0.168 per second; 1080p $0.300 per second.
- `happyhorse-1-1` (alibaba): Text, image, and reference-to-video in one model. Cinematic motion, character consistency across up to 9 references, and synchronized native audio. Pricing: 720p $0.14 per second; 1080p $0.18 per second.
- `seedance-2-0-mini` (bytedance): The fastest, most affordable Seedance 2.0 tier for short cinematic clips with native audio, camera control, and image or video inputs at 480p and 720p. Pricing: T2V/I2V 480P $0.070 per second; T2V/I2V 720P $0.150 per second; Video Input 480P $0.167 per second.
- `minimax-h3` (minimax): Generates 4 to 15 second clips at up to 2K with native stereo audio, following text, image, video, and audio references in one request. Pricing: 768p $0.18 per second; 2K $0.26 per second; 2K regeneration $0.10 per second.
- `kling-v3` (kling): Kuaishou's Kling 3.0 video generator with text-to-video, first and last frame image-to-video, native audio, and 720p, 1080p, or 4K output. Pricing: Standard T2V/I2V $0.168 per second; Standard T2V/I2V Sound $0.252 per second; Pro T2V/I2V $0.224 per second.
- `wan-2-5` (alibaba): Alibaba's Wan 2.5 video model with text-to-video and image-to-video, native audio, optional custom audio tracks, and 480p to 1080p output. Pricing: 480P $0.05 per second; 720P $0.10 per second; 1080P $0.15 per second.
- `wan-2-2` (alibaba): Alibaba's Wan 2.2 video family with standard and flash tiers, first and last frame interpolation, and 480p to 1080p output at low cost. Pricing: Standard 480P $0.02 per second; Standard 1080P $0.10 per second; Flash 480P $0.015 per second.
- `wan-2-1` (alibaba): Alibaba's Wan 2.1 video model with turbo and plus tiers for text-to-video, image-to-video, and first and last frame clips at 480p or 720p. Pricing: Turbo $0.036 per second; Plus $0.10 per second.
- `seedance-2-5` (bytedance): Long-form video model for coherent clips up to 30 seconds, with up to 50 reference images, videos, and audio clips, native audio, editing, and extension. Pricing: T2V/I2V 480P $0.206 per second; T2V/I2V 720P $0.462 per second; T2V/I2V 1080P $1.137 per second.
- `pixverse-v6` (pixverse): Generates video from text, images, first and last frames, or up to seven references, with native audio, multi-shot cuts, and clip extension. Pricing: Standard 360p no audio $0.10 per second; Standard 360p with audio $0.14 per second; Standard 540p no audio $0.14 per second.
- `pixverse-c1` (pixverse): Cinema-grade video model for action and effects work, from text, images, first and last frames, or references, with native synchronized audio. Pricing: Standard 360p no audio $0.12 per second; Standard 360p with audio $0.16 per second; Standard 540p no audio $0.16 per second.
- `pixverse-lipsync` (pixverse): Aligns mouth movement in an existing video to an uploaded audio track or to text spoken by one of fourteen built-in voices. Pricing: Speech $0.08 per second.
- `pixverse-avatar` (pixverse): Turns a single portrait into a talking avatar video, driven by an uploaded audio track or by text spoken with a built-in voice. Pricing: 360p $0.10 per second; 540p $0.20 per second; 720p $0.30 per second.
- `pixverse-edit` (pixverse): Six video transformations in one model: prompt editing, restyling, subject swap, motion transfer, upscaling, and generated sound effects. Pricing: Modify 360p $0.16 per second; Modify 540p $0.20 per second; Modify 720p $0.24 per second.
- `seedance-1-5-pro` (bytedance): Joint audio and video generation with millisecond lip sync, dialogue in six languages, cinematic camera moves, and up to 1080p output. Pricing: With Audio 480P $0.048 per second; With Audio 720P $0.104 per second; With Audio 1080P $0.233 per second.
- `wan-3-0` (alibaba): All-in-one video generation up to 30 seconds with native audio, omni-modal references (images, video, audio, files, web links), and a fast Prime tier. Pricing: Standard 480P $0.07 per second; Standard 720P $0.14 per second; Standard 1080P $0.28 per second.

### Audio Generation

- `ace-step-1-5-xl` (ace-step): Open-source music generation model for text-to-song and lyric-guided audio, with fast 8-step XL Turbo inference for controllable song iteration. Pricing: Music generation $0.00025 (was $0.0003) per generated second.
- `tts-2` (inworld): Realtime voice model that takes plain-English delivery direction, holds one voice identity across 200+ languages, and streams with word timestamps. Pricing: Synthesis $22.00 (was $25.00) per 1M characters.
- `tts-1-5-mini` (inworld): Sub-130ms TTFB voice synthesis with 271+ voices across 15 languages, expressive prosody, and real-time SSE streaming for low-latency voice agents. Pricing: Synthesis $17.50 (was $25.00) per 1M characters.
- `tts-1-5-max` (inworld): Broadcast-quality voice synthesis with rich expressive prosody, 271+ voices across 15 languages, and real-time SSE streaming with per-word timestamps. Pricing: Synthesis $29.75 (was $35.00) per 1M characters.
- `minimax-music-3` (minimax): Generates complete stereo songs from lyrics and a style prompt, with duration, seed, and format controls up to 5 minutes. Pricing: Music generation $0.0018 (was $0.002) per generated second.
- `tts-2-flash` (inworld): Latency-first realtime voice synthesis holding one voice identity across 200+ languages, tuned for high-volume streaming workloads. Pricing: Synthesis $10.50 (was $15.00) per 1M characters.
- `gemini-2-5-flash-tts` (google): Low-latency text-to-speech with single- and multi-speaker voices and controllable style, accent, and expressive tone for production apps. Pricing: Input $1.50 per 1M prompt tokens; Output $30.00 per 1M generated tokens.
- `gemini-2-5-pro-tts` (google): High-quality TTS preview for podcasts, audiobooks, and customer support, with expressive multi-speaker voices across 23+ languages. Pricing: Input $3.00 per 1M prompt tokens; Output $60.00 per 1M generated tokens.
- `gemini-3-1-flash-tts` (google): Highly controllable TTS with new Audio Tags for precise style, tone, pace, and delivery across narration, assistants, and voice apps. Pricing: Input $2.60 per 1M prompt tokens; Output $52.00 per 1M generated tokens.
- `glm-tts` (zhipu): LLM-based text-to-speech with zero-shot voice cloning from 3-10s of audio and emotion-expressive, controllable output via multi-reward RL. Pricing: Fast (INT8) $0.20 per 1k characters; Quality (FP16) $0.21 per 1k characters.
- `soulx-podcast` (soul-ai-lab): Open-source voice model for long-form, multi-speaker podcast dialogue with paralinguistic control (laughter, sighs) and zero-shot voice cloning. Pricing: Base $0.015 per 1k characters; Dialect $0.015 per 1k characters.
- `stable-audio-2-0` (stability): Generates audio up to 3 minutes from text prompts, supporting text-to-audio and audio-to-audio with adjustable duration, steps, and CFG scale. Pricing: Base Cost $0.58 per generation; Per Step Cost $0.00 per step.
- `stable-audio-2-5` (stability): Up-to-3-minute audio from text with text-to-audio, audio-to-audio, and audio inpainting for music production, sound design, and remixing. Pricing: Generation $0.68 per generation.
- `stepaudio-2-5-tts` (stepfun): Contextual StepFun text-to-speech model with natural-language voice direction and expressive delivery. Pricing: Synthesis $0.85 per 10,000 characters.
- `step-tts-2` (stepfun): StepFun text-to-speech model with official voices, custom cloned voices, and voice tag controls. Pricing: Synthesis $0.40 per 10,000 characters.
- `qwen-audio-3-0-tts` (alibaba): Tiered speech synthesis with over 1,000 voices, 16 languages, 20 Chinese dialects, natural-language delivery direction, and inline emotion tags. Pricing: Plus synthesis $0.20 per 10,000 characters; Flash synthesis $0.15 per 10,000 characters.

### Transcription

- `deepgram-nova-3` (deepgram): Speech-to-text transcription using the Nova-3 model with multi-language support and advanced customizable settings for production workloads. Pricing: Transcription $0.014 per minute of audio.
- `openai-whisper-1` (openai): Whisper-1 speech-to-text transcription trained on multilingual supervised audio, with a 25 MB upload limit per file. Pricing: Per Minute of Audio $0.030 per minute.
- `whisper-large-v3-turbo` (openai): Controlled Whisper Large v3 Turbo transcription with multilingual ASR, translation, VAD, timestamps, subtitles, hotwords, and decoder controls. Pricing: Controlled transcription $0.005 (was $0.006) per minute of audio.
- `stepaudio-2-5-asr` (stepfun): StepFun streaming speech recognition model for Chinese and English audio transcription. Pricing: Transcription $0.022 per hour of audio.

### Research & Search

- `exa-answer` (exa): Quick LLM-style answer to a natural-language question, grounded in fresh Exa web search results with inline citations and source links. Pricing: Answer $0.01 per request.
- `exa-search` (exa): Web search engine for finding pages, retrieving similar pages, crawling, and dedicated code search across the open web for AI agents. Pricing: Search (1-25 results) $0.0060 per search; Search (26-100 results) $0.030 per search; Content (Text/Highlights/Summary) $0.0060 per page/feature.
- `linkup-deep-search` (linkup): Iterative AI search that keeps querying when initial results are insufficient, returning more comprehensive answers than Standard mode. Pricing: Per Message $0.13 fixed.
- `linkup-standard` (linkup): AI-powered web search with detailed overviews and answers, faster than Deep Search. Ranks #1 on OpenAI SimpleQA benchmark. Pricing: Per Message $0.013 fixed.
- `perplexity-advanced-deep-research` (perplexity): Institutional-grade research powered by Claude Opus 4.6 reasoning, with maximum depth, enhanced tool access, and extensive source coverage. Pricing: Input $12.00 per 1M prompt tokens; Output $60.00 per 1M generated tokens; Web Search Call $0.012 per call.
- `perplexity-deep-research` (perplexity): Research model for multi-step retrieval, synthesis, and reasoning, autonomously searching, reading, and evaluating sources across complex topics. Pricing: Input $4.80 per 1M prompt tokens; Output $19.00 per 1M generated tokens; Citation Tokens $4.80 per 1M tokens.
- `perplexity-pro-search` (perplexity): Sonar Pro as an agentic researcher: chains web searches, fetches full pages, and streams live reasoning, adapting strategy for complex queries. Pricing: Input $7.80 per 1M prompt tokens; Output $39.00 per 1M generated tokens; Base Fee (Low Context) $0.036 per request.
- `perplexity-search` (perplexity): Real-time web search with filtering by domain, language, date, and more. Returns search results, not LLM responses; no file uploads. Pricing: Search Request $0.0060 per request.
- `perplexity-sonar` (perplexity): Real-time web-connected search with accurate citations and customizable sources for up-to-date AI search integration in production apps. Pricing: Input $2.40 per 1M prompt tokens; Output $2.40 per 1M generated tokens; Base Fee (Low Context) $0.012 per request.
- `perplexity-sonar-pro` (perplexity): Search-grounded model with double the citations and a larger context window, tuned for complex queries needing in-depth, nuanced answers. Pricing: Input $7.20 per 1M prompt tokens; Output $36.00 per 1M generated tokens; Base Fee (Low Context) $0.014 per request.
- `perplexity-sonar-reasoning-pro` (perplexity): Reasoning model on the uncensored open-source R1-1776 with web search, outperforming leading search engines and LLMs on the SimpleQA benchmark. Pricing: Input $4.80 per 1M prompt tokens; Output $19.00 per 1M generated tokens; Base Fee (Low Context) $0.014 per request.
- `tavily-research` (tavily): Multi-search research assistant that explores a topic, analyzes sources, and produces a detailed research report with citations. Pricing: Mini ~$1.19 average per task; Pro ~$2.75 average per task.
- `tavily-search` (tavily): Web search with crawl, extract, and URL mapping for fast, structured retrieval across pages and domains for downstream pipelines. Pricing: Search (Basic/Fast/Ultra-Fast) $0.0096 per search; Search (Advanced) $0.019 per search; Search (Advanced + Answer) $0.029 per search.

### 3D Generation

- `trellis-2-4b` (microsoft): TRELLIS.2 image-to-3D model that turns a reference image into a textured GLB asset with resolution, seed, mesh, texture, and export controls. Pricing: 512 asset $0.025 (was $0.25) per request; 1024 asset $0.249 (was $0.30) per request; 1536 asset $0.499 per request.
- `hyper3d-gen2` (deemos): Rodin Gen-2 engine that turns a text prompt or up to five reference images into production 3D assets with PBR materials and quad or raw meshes. Pricing: Per generation $0.80 per request.
- `hitem3d-2-0` (hitem3d): Production-focused image-to-3D with structure-aware textures, 1536 and 1536 Pro precision tiers, meshes up to two million faces, and five export formats. Pricing: 1536 textured $2.80 per request; 1536 Pro textured $3.60 per request; 1536 white model $1.60 per request.

### Embeddings

- `text-embedding-v4` (alibaba): Multilingual text embedding with selectable output dimensions (64–2048). Up to 8,192 tokens per input. Pricing: Input $0.07 per 1M prompt tokens.
- `tongyi-embedding-vision-flash` (alibaba): Speed-optimised multimodal embedding, same shape as Vision-Plus, 3× cheaper image/video tokens. Pricing: Text input $0.09 per 1M tokens; Image / video input $0.03 per 1M tokens.
- `tongyi-embedding-vision-plus` (alibaba): Multimodal embedding producing independent vectors for text, image, and video inputs. Pricing: Text input $0.09 per 1M tokens; Image / video input $0.09 per 1M tokens.
- `skylark-embedding-vision` (bytedance): Multimodal embedding that fuses text, images, and video into one 1024 or 2048 dimension vector for cross-modal search and retrieval. Pricing: Text input $0.25 per 1M tokens; Image / video input $0.65 per 1M tokens.

### Rerankers

- `qwen3-rerank` (alibaba): Semantic document reranker. Sorts up to 500 candidates per query by relevance, supports 100+ languages, and accepts a custom sorting instruction. Pricing: Input $0.10 per 1M prompt tokens.

### Tools & Agents

- `gptzero` (gptzero): Deep-learning detector that flags portions of text likely generated by AI versus human, classifying content as entirely human, AI, or mixed. Pricing: Text Scan $0.39 per 1,000 words.
- `manus` (manus): Autonomous AI agent that turns a high-level prompt into subtasks, calls tools and APIs, and delivers end-to-end results without manual orchestration. Pricing: Adaptive - Manus 1.6 Lite $1.44 - $2.63 per task; Adaptive - Manus 1.6 $2.89 - $5.25 per task; Adaptive - Manus 1.6 Max $5.25 - $9.19 per task.

## Recent blog posts

- [How to Use the DeepSeek V4.1 Flash API](https://empiriolabs.ai/blog/deepseek-v4-1-flash-api): DeepSeek V4.1 Flash is available on EmpirioLabs with native image understanding, hybrid thinking, function calling, and a 1M token context window.
- [How to Use the Fugu Ultra v2.0 API](https://empiriolabs.ai/blog/fugu-ultra-v2-0-api): Fugu Ultra v2.0 is Sakana's flagship multi-agent conductor, coordinating its strongest pool of expert models on hard, multi-step problems. Here is how to call it, how the context tier actually works, and what changed from v1.1.
- [How to Use the Fugu Max API](https://empiriolabs.ai/blog/fugu-max-api): Fugu Max is Sakana's cost-efficient multi-agent conductor, orchestrating a wide pool of open and specialist models behind one OpenAI-compatible endpoint. Here is how to call it, what it bills for, and the three things callers get wrong.
- [How to Use the TTS 2 Flash API](https://empiriolabs.ai/blog/tts-2-flash-api): TTS 2 Flash is available on EmpirioLabs: latency-first speech synthesis that holds one voice identity across 200+ languages, with streaming audio and word-level timestamps for captioning.
- [How to Use the Muse Spark 1.3 API](https://empiriolabs.ai/blog/muse-spark-1-3-api): Muse Spark 1.3 is available on EmpirioLabs with a 1,048,576-token context, image, video, and PDF understanding, always-on reasoning, web search, and tool calling.
- [How to Use the Qwen3.8 Max 0902 API](https://empiriolabs.ai/blog/qwen3-8-max-0902-api): Qwen3.8 Max 0902 is the upgraded snapshot of Alibaba's flagship Qwen3.8 Max, with stronger coding and agent work, sharper chart and document vision, a 1M token context, deep thinking, and five built-in tools.
- [How to Use the Qwen3.8 Flash API](https://empiriolabs.ai/blog/qwen3-8-flash-api): Qwen3.8 Flash is Alibaba's fast multimodal Qwen3.8 model, with a 1M token context, image and video input, deep thinking, five built-in tools, and strict JSON Schema output.
- [EmpirioLabs AI is now available on Opper](https://empiriolabs.ai/blog/empiriolabs-on-opper): EmpirioLabs AI is now a listed provider on Opper, the Stockholm-based European AI gateway. Opper users can call our Native Inference text models with European data safeguards in place.
- [How to Use the GLM 5.3 Flash API](https://empiriolabs.ai/blog/glm-5-3-flash-api): GLM 5.3 Flash is Z.ai's natively multimodal flash model, with a 1M token context, image and video input, always-on reasoning, native web search, and tool calling.
- [How to Use the Wan 3.0 API](https://empiriolabs.ai/blog/wan-3-0-api): Wan 3.0 generates clips up to 30 seconds with synchronized native audio, animates first and last frames, and builds scenes from omni-modal references: images, videos, audio, documents, and web links.
- [How to Use the GLM 5.3 API](https://empiriolabs.ai/blog/glm-5-3-api): GLM 5.3 is Z.ai's flagship coding and agentic model, with a 1M token context, always-on reasoning at three effort levels, native web search, and tool calling.
- [How to Use the DeepSeek V4 Pro 0813 API](https://empiriolabs.ai/blog/deepseek-v4-pro-0813-api): DeepSeek V4 Pro 0813 is available on EmpirioLabs with major gains across coding, repository work, tool use, and agent tasks.
- [How to Use the MiniMax Music 3 API](https://empiriolabs.ai/blog/minimax-music-3-api): MiniMax Music 3 is live on EmpirioLabs: lyric-to-song generation up to five minutes, with prompt, lyrics, seed, and format controls on the same audio endpoint as the rest of the catalog.
- [How to Use the Qwen3.8 27B API](https://empiriolabs.ai/blog/qwen3-8-27b-api): Qwen3.8 27B is live on EmpirioLabs native inference with 256K context, image and video input, tools, structured JSON, and thinking on by default.
- [How to Use the Grok Imagine Image 2.0 API](https://empiriolabs.ai/blog/grok-imagine-image-2-0-api): Grok Imagine Image 2.0 is available on EmpirioLabs for text-to-image and multi-reference editing, with selectable quality, 1K or 2K output, and up to three source images per edit.
- [Muse Glimmer 30B vs Qwen3.6 27B: five one-shot coding tests](https://empiriolabs.ai/blog/muse-glimmer-vs-qwen3-6-27b-coding-test): We gave Meta's Muse Glimmer 30B and Qwen3.6 27B the same five coding prompts at max reasoning, one shot each, and rendered every result live.
- [How to Use the Muse Glimmer 30B API](https://empiriolabs.ai/blog/muse-glimmer-30b-api): Muse Glimmer 30B is available on EmpirioLabs native inference with vision input, a 128K context, tool calling, strict structured output, and adjustable reasoning strength.
- [Subscription plans are here: Lite, Pro and Max](https://empiriolabs.ai/blog/subscription-plans-lite-pro-max): Lite, Pro and Max add optional weekly allowances across models, media, search and tasks on top of normal pay as you go.
- [How to Use the Qwen Image 3.0 API](https://empiriolabs.ai/blog/qwen-image-3-0-api): Qwen Image 3.0 combines fast Base generation with a Pro variant for dense layouts, fine multilingual typography, and photorealistic edits.
- [How to Use the Muse Spark 1.2 API](https://empiriolabs.ai/blog/muse-spark-1-2-api): Muse Spark 1.2 is available on EmpirioLabs with a 1,048,576-token context, multimodal input, always-on reasoning, web search, and tool calling.
- [How to Use the Seedance 2.5 API](https://empiriolabs.ai/blog/seedance-2-5-api): Seedance 2.5 generates coherent video up to 30 seconds in a single request, accepts up to 50 reference images, videos, and audio clips, and edits or extends footage you already have.
- [How to Use the MiniMax H3 API](https://empiriolabs.ai/blog/minimax-h3-api): MiniMax H3 is available on EmpirioLabs: 4 to 15 second clips at up to 2K with a native stereo audio track, from text, image, video, and audio references.
- [How to Use the Qwen3.8 Max API](https://empiriolabs.ai/blog/qwen3-8-max-api): Call Qwen3.8 Max, Alibaba's trillion-scale Qwen3.8 flagship with a 1M token context, deep thinking, five built-in tools, and text, image, and video input.
- [How to Use the DeepSeek V4 Flash 0731 API](https://empiriolabs.ai/blog/deepseek-v4-flash-0731-api): DeepSeek V4 Flash 0731 is available on EmpirioLabs with major gains across coding, repository work, tool use, and full-stack agent tasks.
- [How to Use the Qwen3.7 Flash API](https://empiriolabs.ai/blog/qwen3-7-flash-api): Call Qwen3.7 Flash, Alibaba's fast Qwen3.7 vision-language model with a 1M token context, switchable thinking, built-in tools, and text, image, and video input.
- [How to Use the Fugu Ultra v1.1 API](https://empiriolabs.ai/blog/fugu-ultra-v1-1-api): Fugu Ultra v1.1 is available on EmpirioLabs with distinct maximum reasoning, 1M context, image input, function calling, strict structured output, and web search.
- [How to Use the Qwen Audio 3.0 TTS API](https://empiriolabs.ai/blog/qwen-audio-3-0-tts-api): Qwen Audio 3.0 TTS is available on EmpirioLabs as one model with a Plus and Flash tier, over 1,000 shared voices, 16 languages, and inline emotion tags.
- [Kimi K3 vs GLM 5.2 vs DeepSeek V4 Pro](https://empiriolabs.ai/blog/kimi-k3-vs-glm-5-2-vs-deepseek-v4-pro): Compare Kimi K3, GLM 5.2, and DeepSeek V4 Pro on EmpirioLabs: three 1M token reasoning models, side by side on context, price, image and video input, web search, and structured output.
- [How to Use the Kimi K3 API](https://empiriolabs.ai/blog/kimi-k3-api): Call Kimi K3, Moonshot AI's flagship reasoning model with a 1M token context, always-on thinking, native web search, and image and video input, through an OpenAI-compatible API on EmpirioLabs.
- [Seed 2.1 Turbo vs Kimi K2.7 Code vs DeepSeek V4 Pro](https://empiriolabs.ai/blog/seed-2-1-turbo-vs-kimi-k2-7-code-vs-deepseek-v4-pro): A factual comparison of three top coding and agent models on EmpirioLabs: context, pricing, image and video input, web search, and structured output, plus which one to pick.
