Home Blog

How to Use the MiMo V2.6 API

MiMo V2.6 API blog cover

Sep 21, 2026

EmpirioLabs AI

Disclosure: This article was written with AI assistance and reviewed by EmpirioLabs AI.

Xiaomi's MiMo V2.6 series is available on EmpirioLabs. Three models ship together, and all three share the same shape: a 1,048,576 token context window, native image, video, and audio understanding with text output, reasoning on by default, function calling, and structured JSON output up to 131,072 output tokens.

The three models

  • mimo-v2-6-pro is the trillion-parameter flagship, built for agentic workflows, long-horizon tasks, and research.
  • mimo-v2-6-flash is the low-cost member, aimed at high-frequency calls and large-scale batch work.
  • mimo-v2-6-pro-ultraspeed serves the same flagship quality at substantially higher output speed, for latency-sensitive work.

What the models do

All three read text, images, video, and audio in the same request and answer in text, so you can send a screenshot, a short clip, and a voice note alongside a prompt. Thinking is on by default and can be turned off per request. Function calling and both JSON modes work on the standard chat completions surface: a response_format of type json_object for any JSON object, or a strict json_schema when you need an exact response shape.

All three can search the web and return a Sources section with the answer: MiMo V2.6 Pro and MiMo V2.6 Flash with native web search, and MiMo V2.6 Pro UltraSpeed with optional Linkup web search.

How to call MiMo V2.6

Use the OpenAI-compatible chat completions endpoint and set model to the id you want. Image, video, and audio parts use the standard content blocks:

curl https://api.empiriolabs.ai/v1/chat/completions \
  -H 'Authorization: Bearer $EMPIRIOLABS_API_KEY' \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "mimo-v2-6-pro",
    "enable_thinking": true,
    "messages": [
      {
        "role": "user",
        "content": [
          {"type": "text", "text": "What does this chart show? Answer in two sentences."},
          {"type": "image_url", "image_url": {"url": "https://example.com/chart.png"}}
        ]
      }
    ]
  }'

Operational notes

  • Very small images are rejected. A 14 pixel square test image comes back as a generic invalid-request error that does not mention the size, while the same request with a 64 pixel square succeeds. Send icons and swatches at a reasonable resolution rather than their native size.
  • Web search uses a different control on UltraSpeed. MiMo V2.6 Pro and MiMo V2.6 Flash take tool_web_search; MiMo V2.6 Pro UltraSpeed takes web_search_linkup. When you move a request between models, send the flag that model accepts.
  • Reasoning is on by default and reasoning tokens bill as output tokens. Turn enable_thinking off for short, latency-sensitive replies.
  • Output is capped at 131,072 tokens. Asking for more returns an error naming the limit.

Pricing

Token usage is billed at one published input rate and one published output rate per model, with no separate cached-input rate. Web search adds a small per-call fee only when a search actually runs, and native search on MiMo V2.6 Pro and MiMo V2.6 Flash can run more than one search in one request. See the live model page and pricing page for current rates.

Start building

Try the series in the Playground, read the API documentation, or compare it with other models in the model catalog.

Ready to use better endpoints?

Explore our models, or contact us about business inquiries, custom deployments, or anything else.