Home Blog

Qwen3.8 Flash vs DeepSeek V4.1 Flash vs MiMo V2.6 Flash vs Step 3.7 Flash: four coding tests

Qwen3.8 Flash vs DeepSeek V4.1 Flash vs MiMo V2.6 Flash vs Step 3.7 Flash: four coding tests

Oct 7, 2026

EmpirioLabs AI

We put Qwen3.8 Flash, DeepSeek V4.1 Flash, MiMo V2.6 Flash and Step 3.7 Flash against each other across four one-shot coding tests, all run on EmpirioLabs. Same prompt each, one shot, no edits, and every result rendered live in a real browser.

Watch the four way comparison

Specs at a glance

Qwen3.8 FlashDeepSeek V4.1 FlashMiMo V2.6 FlashStep 3.7 Flash
MakerAlibabaDeepSeekXiaomiStepFun
Context window1,000,000 tokens1,000,000 tokens1,000,000 tokens256,000 tokens
InputText, images and videoText and imagesText, images, video and audioText, images and video
ReasoningEffort up to maxEffort up to maxThinking on or offEffort up to max
Structured outputStrict JSON SchemaJSON modeStrict JSON SchemaJSON mode
Input price$0.16 per 1M tokens$0.30 per 1M tokens$0.14 per 1M tokens$0.20 per 1M tokens
Output price$0.47 per 1M tokens$1.20 per 1M tokens$0.28 per 1M tokens$1.15 per 1M tokens

How we ran it

Each model received the identical prompt for four tasks: build a Ferris wheel at night in a single HTML file; build a pizza baking in a wood-fired oven in a single HTML file; build a toy train set in a single HTML file; build kites flying over a beach in a single HTML file. Every task asked for a single self-contained HTML file with no external libraries that animates on its own. Qwen3.8 Flash, DeepSeek V4.1 Flash and Step 3.7 Flash ran at reasoning_effort: "max"; MiMo V2.6 Flash, which has no effort setting, ran with thinking on. Every model had a 65,536 token output budget and one shot per task. The line counts and tokens-per-second readouts on each panel are measured from the real API calls, and each result is the file the model returned, rendered as-is.

What to look for

DeepSeek V4.1 Flash returned its files fastest, at about 191 tokens per second. MiMo V2.6 Flash wrote the longest files, about 930 lines on average, and Step 3.7 Flash the most compact, about 490. Qwen3.8 Flash used the most output tokens, about 48,000 per task including its reasoning. Watch how each model brings its scene to life: the motion, the detail and how it keeps running on its own. We are not declaring a winner. Run the clip and judge the outputs for your own use case.

Run the same test on EmpirioLabs

curl https://api.empiriolabs.ai/v1/chat/completions \
  -H "Authorization: Bearer $EMPIRIOLABS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3-8-flash",
    "reasoning_effort": "max",
    "messages": [{"role": "user", "content": "Build a Ferris wheel at night in a single HTML file."}]
  }'

Swap model to deepseek-v4-1-flash, mimo-v2-6-flash or step-3-7-flash to run the same request against the others, or try them interactively in the playground.

Frequently asked questions

Were the results edited or retried?

No. Each panel is the first file the model returned for that prompt, rendered exactly as returned. A request that came back without a file was sent again.

Why the highest reasoning setting?

A fair head to head shows each model at its best, so each ran at its highest reasoning setting: max effort for Qwen3.8 Flash, DeepSeek V4.1 Flash and Step 3.7 Flash, thinking on for MiMo V2.6 Flash.

Which model should I use?

MiMo V2.6 Flash has the lowest output price of the four here. Run your own workload against each before deciding: the same request works for all four with only the model name changed.

Try it

Playground | All models | Pricing

Ready to use better endpoints?

Explore our models, or contact us about business inquiries, custom deployments, or anything else.