
Qwen3.8 27B
Alibaba CloudMultimodal reasoner with 256K context, image and video input, function tools, structured JSON, and thinking on by default.
Browse the full catalog of models across text, image, audio, video, 3D, and more.
Model catalog
Browse text, image, video, audio, 3D, search, and agent endpoints with pay-as-you-go pricing. The interactive catalog loads current availability from EmpirioLabs, and these model docs are crawlable without client JavaScript.
xAI image-to-video with prompt-guided motion, native audio, 480p or 720p output, and up to 15 second clips.
Multimodal video generation for cinematic clips from text, image, audio, or video inputs.
Unified image generation and editing for high-resolution creative, brand, and product visuals.
Cost-effective vision-language model for text, image, video, coding, tools, and 1M-context workflows.
Flagship long-context model for coding, productivity, long-running agents, deep thinking, and tool use.
Multimodal reasoning for coding, agents, long-context analysis, and text, image, and video input.
Moonshot multimodal reasoning with strong coding support, 256K context, and image and video input.
Long-context reasoning with tool calling, structured output, cache support, and 128K output.
Image-to-3D generation that turns a reference image into a textured GLB asset.

Alibaba CloudMultimodal reasoner with 256K context, image and video input, function tools, structured JSON, and thinking on by default.

MiniMaxGenerates complete stereo songs from lyrics and a style prompt, with duration, seed, and format controls up to 5 minutes.

DeepSeekOfficial 0813 Pro release with major gains across coding, repository work, tool use, and agent tasks, plus hybrid thinking and a 1M context window.
Text-to-image plus multi-reference editing with up to three source images, selectable quality tiers, and 1K or 2K output resolution.

Meta AIMeta open 30B agentic model with image understanding, 128K context, tool calling, structured output, and controllable reasoning strength.

ByteDanceLong-form video model for coherent clips up to 30 seconds, with up to 50 reference images, videos, and audio clips, native audio, editing, and extension.

Z.aiReasoning and coding model with a 1M token context, 128K output, adjustable reasoning effort, native web search, and tool calling.

Moonshot AIKimi K3 is Moonshot's flagship reasoning model with a 1M token context, always-on thinking, native web search, and text, image, and video inputs.

Meta AIMeta's updated frontier reasoning model with a 1,048,576-token context, image, video, audio, and PDF understanding, web search, and tool calling.

Moonshot AIKimi K2.7 Code is Moonshot's trillion-parameter agentic coding model with 256K context, always-on reasoning, and text, image, and video inputs.

Meta AIMeta frontier reasoning model with a 1,048,576-token context plus image, video, audio, and PDF understanding, web search, and tool calling.

Sakana AIUpdated multi-agent conductor for hard reasoning, coding, and research, with distinct max effort, 1M context, image input, and web search.

Black Forest LabsApache-licensed 4B FLUX.2 Klein image generation and editing model with text-to-image, reference-image editing, and creative workflow support.

AmazonImage generation and editing model creating and modifying images from text or image inputs, with inpainting, virtual try-on, and style controls.

TencentOpen-source text-to-image model on a multimodal Mixture-of-Experts architecture with photorealistic detail and strong multilingual text rendering.

DeepSeekAutoregressive framework on the Janus Pro 7B model that unifies multimodal understanding and image generation in one architecture.

Alibaba CloudUnified image generation and editing model with class-leading complex Chinese/English text rendering, realistic textures, and multi-image fusion.

Alibaba CloudQwen image generation and editing with Base and Pro variants, multilingual typography, multi-image references, and controllable 1K or 2K output.

Kling AIText-to-video and image-to-video with synchronized native audio, at 720p or 1080p for 3 to 15 seconds, with aspect ratio and prompt control.

AmazonVideo generation model producing up to 2-minute multi-shot videos from text and optional image prompts with improved quality and consistency.

Alibaba CloudVideo model offering Text-to-Video, Image-to-Video, Reference-to-Video, and Video Edit modes with high-fidelity, motion-smooth output.

Tencent8.3B-parameter video model with native 720p output (upscalable to 1080p), strong motion coherence, and bilingual prompt understanding up to 10s.

Kling AIVideo model in Standard or Pro modes with Text-to-Video, Image-to-Video, Reference-to-Video, editing, native sound, and multi-scene transitions.

Kling AIKling 3.0 model that transfers motion from a reference video onto a character from a reference image, with Standard 720p and Pro 1080p tiers.

ACE-StepOpen-source music generation model for text-to-song and lyric-guided audio, with fast 8-step XL Turbo inference for controllable song iteration.

InworldRealtime voice model with plain-English voice direction, one voice identity across 100+ languages, and sub-200ms streaming time-to-first-audio.

InworldSub-130ms TTFB voice synthesis with 271+ voices across 15 languages, expressive prosody, and real-time SSE streaming for low-latency voice agents.

InworldBroadcast-quality voice synthesis with rich expressive prosody, 271+ voices across 15 languages, and real-time SSE streaming with per-word timestamps.

MiniMaxGenerates complete stereo songs from lyrics and a style prompt, with duration, seed, and format controls up to 5 minutes.

GoogleLow-latency text-to-speech with single- and multi-speaker voices and controllable style, accent, and expressive tone for production apps.

DeepgramSpeech-to-text transcription using the Nova-3 model with multi-language support and advanced customizable settings for production workloads.

OpenAIWhisper-1 speech-to-text transcription trained on multilingual supervised audio, with a 25 MB upload limit per file.

OpenAIControlled Whisper Large v3 Turbo transcription with multilingual ASR, translation, VAD, timestamps, subtitles, hotwords, and decoder controls.

StepFunStepFun streaming speech recognition model for Chinese and English audio transcription.

ExaQuick LLM-style answer to a natural-language question, grounded in fresh Exa web search results with inline citations and source links.

ExaWeb search engine for finding pages, retrieving similar pages, crawling, and dedicated code search across the open web for AI agents.

LinkupIterative AI search that keeps querying when initial results are insufficient, returning more comprehensive answers than Standard mode.

LinkupAI-powered web search with detailed overviews and answers, faster than Deep Search. Ranks #1 on OpenAI SimpleQA benchmark.

PerplexityInstitutional-grade research powered by Claude Opus 4.6 reasoning, with maximum depth, enhanced tool access, and extensive source coverage.

PerplexityResearch model for multi-step retrieval, synthesis, and reasoning, autonomously searching, reading, and evaluating sources across complex topics.

MicrosoftTRELLIS.2 image-to-3D model that turns a reference image into a textured GLB asset with resolution, seed, mesh, texture, and export controls.

Alibaba CloudMultilingual text embedding with selectable output dimensions (64–2048). Up to 8,192 tokens per input.

Alibaba CloudSpeed-optimised multimodal embedding, same shape as Vision-Plus, 3× cheaper image/video tokens.

Alibaba CloudMultimodal embedding producing independent vectors for text, image, and video inputs.

Alibaba CloudSemantic document reranker. Sorts up to 500 candidates per query by relevance, supports 100+ languages, and accepts a custom sorting instruction.

GPTZeroDeep-learning detector that flags portions of text likely generated by AI versus human, classifying content as entirely human, AI, or mixed.

ManusAutonomous AI agent that turns a high-level prompt into subtasks, calls tools and APIs, and delivers end-to-end results without manual orchestration.
Check out our pricing or reach out if you want your own model deployed on our stack.