
Fast Qwen3.7 vision-language model for text, image, video, tool use, and agentic tasks, with implicit caching and a 1M token context.
Fast Qwen3.7 vision-language model for text, image, video, tool use, and agentic tasks, with implicit caching and a 1M token context.
Supports text, image, and video inputs, thinking mode with enable_thinking and thinking_budget up to 131072 tokens, function calling, structured output, web search, web extractor, code interpreter, image search tools, prefix continuation, and implicit context caching. Web Search and image-search tools are billed per invoked call. Thinking tokens are billed as output tokens.
Auch bekannt als Alibaba Cloud Qwen3.7 Flash, Qwen3.7-Flash, qwen3-7-flash
qwen3-7-flash/v1/chat/completionsPOST/v1/responsesPOST/v1/messagesPOST/v1beta/models/qwen3-7-flash:generateContentqwen3.7-flashLive Pay-as-you-go-Preise aus dem EmpirioLabs-Katalog. Du zahlst nur für das, was du nutzt, ohne monatliches Minimum.
Qwen3.7 Flash bedient die OpenAI-kompatible Chat Completions API. Richte ein beliebiges OpenAI SDK mit deinem EmpirioLabs API-Schlüssel auf https://api.empiriolabs.ai/v1 und verwende die Modell-ID qwen3-7-flash. Hol dir einen API-Schlüssel im EmpirioLabs Dashboard.
curl https://api.empiriolabs.ai/v1/chat/completions \
-H "Authorization: Bearer $EMPIRIOLABS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3-7-flash",
"messages": [
{"role": "user", "content": "Write a haiku about the ocean."}
]
}'from openai import OpenAI
client = OpenAI(
base_url="https://api.empiriolabs.ai/v1",
api_key="YOUR_EMPIRIOLABS_API_KEY",
)
response = client.chat.completions.create(
model="qwen3-7-flash",
messages=[{"role": "user", "content": "Write a haiku about the ocean."}],
)
print(response.choices[0].message.content)Request-Parameter, die die Qwen3.7 Flash API auf EmpirioLabs unterstützt. Standardwerte gelten, wenn ein Feld weggelassen wird.
| Parameter | Typ | Standard | Bereich / Werte | Beschreibung |
|---|---|---|---|---|
| temperature | number | 0.7 | 0 bis 2 | Sampling temperature. 0 is deterministic and 2 is maximum randomness. |
| top_p | number | 0.9 | 0 bis 1 | Nucleus sampling probability mass. Lower values make outputs more focused. |
| max_tokens | number | 4096 | 1 bis 65536 | Maximum output tokens. |
| stop | string | - | - | Up to 4 strings where the model will stop generating further tokens. |
| enable_thinking | boolean | true | - | Enable reasoning before answering. |
| reasoning_effort | enum | medium | none, low, medium, high, max | Reasoning effort level. none disables thinking. low, medium, high, and max set bounded thinking budgets sized to the selected model. |
| thinking_budget | number | 32768 | 1 bis 131072 | Maximum tokens reserved for reasoning when thinking is enabled. |
| vl_high_resolution_images | boolean | true | - | Use higher resolution processing for image inputs. |
| max_pixels | number | 2621440 | 4096 bis 16777216 | Maximum pixel count per image when high resolution processing is disabled. |
| video_fps | number | 2 | 0.1 bis 10 | Frames per second to sample from video inputs. |
| treat_images_as_video | boolean | false | - | Treat a sequence of images as video frames. |
| tool_web_search | boolean | true | - | Search the web for real-time information. Adds $0.03 to the request cost for each invoked call. |
| tool_web_extractor | boolean | true | - | Extract and read content from URLs. Requires Web Search and Thinking. |
| tool_code_interpreter | boolean | true | - | Run Python code in a sandbox. Requires Thinking. |
Token pricing is tiered by prompt size and steps up above 32K and again above 256K input tokens. Cached prompt tokens are billed at the implicit cache rate for the matching tier. Web Search, Text-to-Image Search, and Image-to-Image Search are billed only when invoked.
Text-to-Image Search and Image-to-Image Search use the Image Search pricing row. Thinking tokens are billed as output tokens.
When this model invokes built-in tools inside a single request, the response carries a normalized usage.tool_usage map alongside the token counts. Tool counts are already factored into cost_usd and are surfaced for transparency.
:variant1
China pricing is discounted versus Singapore. Token pricing is tiered by prompt size and steps up above 32K and again above 256K input tokens. Implicit cache input uses the cached-token row. Web Search, Text-to-Image Search, and Image-to-Image Search are billed only when invoked.
Text-to-Image Search and Image-to-Image Search use the Image Search pricing row. China paid tool calls are $0.01 each. Thinking tokens are billed as output tokens.
Varianten sind alternative Versionen von Qwen3.7 Flash mit eigener Modell-ID. Je nach Variante können sich Serving-Region, Preise oder unterstützte Parameter unterscheiden; alles andere funktioniert gleich.
qwen3-7-flash:variant1Auf EmpirioLabs wird Qwen3.7 Flash nach Verbrauch abgerechnet. Die Live-Preistabelle auf dieser Seite entspricht immer dem, was die API berechnet.
Qwen3.7 Flash unterstützt ein Kontextfenster von 1M Token mit bis zu 65.536 Ausgabe-Token pro Antwort.
Ja. Qwen3.7 Flash bedient die OpenAI-kompatible Chat Completions API. Bestehende OpenAI SDKs funktionieren, indem du base_url auf https://api.empiriolabs.ai/v1 setzt und als Modell-ID qwen3-7-flash verwendest.
Qwen3.7 Flash ist als 2 Modell-IDs verfügbar: der Standard qwen3-7-flash plus qwen3-7-flash:variant1 (China). Varianten können sich in Serving-Region, Preisen oder unterstützten Parametern unterscheiden; die Preistabellen stehen auf dieser Seite.
Ja. Der EmpirioLabs Playground führt Qwen3.7 Flash im Browser mit denselben Parametern aus, die die API bietet. So kannst du Prompts testen, bevor du Code schreibst.
Erstelle ein EmpirioLabs-Konto und generiere dann einen Schlüssel unter API Keys im Dashboard. Die Abrechnung erfolgt über Pay-as-you-go-Guthaben, du zahlst also nur für deine Requests.
Schauen Sie sich unsere Preise an oder kontaktieren Sie uns, wenn Sie Ihr eigenes Modell auf unserem Stack implementieren möchten.