
Fast Qwen3.8 vision-language model for coding, agents, and visual understanding, with image and video input, five built-in tools, and a 1M token context.
Fast Qwen3.8 vision-language model for coding, agents, and visual understanding, with image and video input, five built-in tools, and a 1M token context.
Supports text, image, and video input, thinking mode with enable_thinking and thinking_budget up to 262144 tokens, function calling, structured output including strict JSON Schema, and five built-in tools: tool_web_search, tool_web_extractor, tool_code_interpreter, tool_web_search_image, and tool_image_search. Thinking tokens are billed as output tokens.
Also known as Alibaba Cloud Qwen3.8 Flash
qwen3-8-flash/v1/chat/completionsPOST/v1/responsesPOST/v1/messagesPOST/v1beta/models/qwen3-8-flash:generateContentqwen3.8-flashLive pay-as-you-go rates from the EmpirioLabs catalog. You are billed only for what you use, with no monthly minimum.
Qwen3.8 Flash serves the OpenAI-compatible Chat Completions API. Point any OpenAI SDK at https://api.empiriolabs.ai/v1 with your EmpirioLabs API key and use the model id qwen3-8-flash. Get an API key from the EmpirioLabs dashboard.
curl https://api.empiriolabs.ai/v1/chat/completions \
-H "Authorization: Bearer $EMPIRIOLABS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3-8-flash",
"messages": [
{"role": "user", "content": "Write a haiku about the ocean."}
]
}'from openai import OpenAI
client = OpenAI(
base_url="https://api.empiriolabs.ai/v1",
api_key="YOUR_EMPIRIOLABS_API_KEY",
)
response = client.chat.completions.create(
model="qwen3-8-flash",
messages=[{"role": "user", "content": "Write a haiku about the ocean."}],
)
print(response.choices[0].message.content)Request parameters supported by the Qwen3.8 Flash API on EmpirioLabs. Defaults apply when a field is omitted.
| Parameter | Type | Default | Range / values | Description |
|---|---|---|---|---|
| temperature | number | 0.7 | 0 to 2 | Sampling temperature. 0 is deterministic and 2 is maximum randomness. |
| top_p | number | 0.9 | 0 to 1 | Nucleus sampling probability mass. Lower values make outputs more focused. |
| max_tokens | number | 4096 | 1 to 131072 | Maximum output tokens. |
| stop | string | - | - | Up to 4 strings where the model will stop generating further tokens. |
| enable_thinking | boolean | true | - | Enable reasoning before answering. |
| reasoning_effort | enum | medium | none, low, medium, high, max | Reasoning effort level. none disables thinking. low, medium, high, and max set bounded thinking budgets sized to the selected model. |
| thinking_budget | number | 32768 | 1 to 262144 | Maximum tokens reserved for reasoning when thinking is enabled. |
| vl_high_resolution_images | boolean | true | - | Use higher resolution processing for image inputs. |
| max_pixels | number | 2621440 | 4096 to 16777216 | Maximum pixel count per image when high resolution processing is disabled. |
| video_fps | number | 2 | 0.1 to 10 | Frames per second to sample from video inputs. |
| treat_images_as_video | boolean | false | - | Treat a sequence of images as video frames. |
| tool_web_search | boolean | false | - | Search the web for real-time information. Adds $0.02 to the request cost for each invoked call. |
| tool_web_extractor | boolean | false | - | Extract and read content from URLs. Requires Web Search and Thinking. |
| tool_code_interpreter | boolean | false | - | Run Python code in a sandbox. Requires Thinking. |
Text, image, and video input are supported. Web search, web extractor, code interpreter, text-to-image search, and image-to-image search are optional built-in tools exposed through tool_* parameters. Web search, text-to-image search, and image-to-image search add $0.02 for each invoked call; web extractor and code interpreter run at no extra cost. Web extractor requires web search, and both web extractor and code interpreter require thinking. A single request can invoke a tool more than once, and each invoked call is billed. Thinking tokens are billed as output tokens.
When this model invokes built-in tools inside a single request, the response carries a normalized usage.tool_usage map alongside the token counts:
"usage": {
"prompt_tokens": 123,
"completion_tokens": 456,
"cost_usd": 0.0042,
"tool_usage": {"web_search": 3, "code_interpreter": 1}
}Tool counts are already factored into cost_usd and are surfaced for transparency so you can audit per-tool billing. The field is omitted when no tools were invoked.
On EmpirioLabs, Qwen3.8 Flash is billed pay as you go: Input $0.16 per 1M prompt tokens; Output $0.47 per 1M generated tokens; Web search $0.02 per call when invoked. The live rate card on this page always matches what the API charges.
Qwen3.8 Flash supports a 1M-token context window with up to 131,072 output tokens per response.
Yes. Qwen3.8 Flash serves the OpenAI-compatible Chat Completions API, so existing OpenAI SDKs work by pointing base_url at https://api.empiriolabs.ai/v1 and setting the model id to qwen3-8-flash.
Yes. The EmpirioLabs playground runs Qwen3.8 Flash in the browser with the same parameters the API exposes, so you can test prompts before writing code.
Create an EmpirioLabs account, then generate a key under API Keys in the dashboard. Billing is pay-as-you-go credits, so you only pay for the requests you make.
Check out our pricing or reach out if you want your own model deployed on our stack.