Disclosure: This article was written with AI assistance and reviewed by EmpirioLabs AI.
Xiaomi's MiMo V2.6 series is available on EmpirioLabs. Three models ship together, and all three share the same shape: a 1,048,576 token context window, native image, video, and audio understanding with text output, reasoning on by default, function calling, and structured JSON output up to 131,072 output tokens.
The three models
mimo-v2-6-prois the trillion-parameter flagship, built for agentic workflows, long-horizon tasks, and research.mimo-v2-6-flashis the low-cost member, aimed at high-frequency calls and large-scale batch work.mimo-v2-6-pro-ultraspeedserves the same flagship quality at substantially higher output speed, for latency-sensitive work.
What the models do
All three read text, images, video, and audio in the same request and answer in text, so you can send a screenshot, a short clip, and a voice note alongside a prompt. Thinking is on by default and can be turned off per request. Function calling and both JSON modes work on the standard chat completions surface: a response_format of type json_object for any JSON object, or a strict json_schema when you need an exact response shape.
All three can search the web and return a Sources section with the answer: MiMo V2.6 Pro and MiMo V2.6 Flash with native web search, and MiMo V2.6 Pro UltraSpeed with optional Linkup web search.
How to call MiMo V2.6
Use the OpenAI-compatible chat completions endpoint and set model to the id you want. Image, video, and audio parts use the standard content blocks:
curl https://api.empiriolabs.ai/v1/chat/completions \
-H 'Authorization: Bearer $EMPIRIOLABS_API_KEY' \
-H 'Content-Type: application/json' \
-d '{
"model": "mimo-v2-6-pro",
"enable_thinking": true,
"messages": [
{
"role": "user",
"content": [
{"type": "text", "text": "What does this chart show? Answer in two sentences."},
{"type": "image_url", "image_url": {"url": "https://example.com/chart.png"}}
]
}
]
}'
Operational notes
- Very small images are rejected. A 14 pixel square test image comes back as a generic invalid-request error that does not mention the size, while the same request with a 64 pixel square succeeds. Send icons and swatches at a reasonable resolution rather than their native size.
- Web search uses a different control on UltraSpeed. MiMo V2.6 Pro and MiMo V2.6 Flash take
tool_web_search; MiMo V2.6 Pro UltraSpeed takesweb_search_linkup. When you move a request between models, send the flag that model accepts. - Reasoning is on by default and reasoning tokens bill as output tokens. Turn
enable_thinkingoff for short, latency-sensitive replies. - Output is capped at 131,072 tokens. Asking for more returns an error naming the limit.
Pricing
Token usage is billed at one published input rate and one published output rate per model, with no separate cached-input rate. Web search adds a small per-call fee only when a search actually runs, and native search on MiMo V2.6 Pro and MiMo V2.6 Flash can run more than one search in one request. See the live model page and pricing page for current rates.
Start building
Try the series in the Playground, read the API documentation, or compare it with other models in the model catalog.



