We put Kimi K3, DeepSeek V4 Pro 0813, MiniMax M3 and MiMo V2.6 Pro against each other across three one-shot coding tests, all run on EmpirioLabs. Same prompt each, one shot, no edits, and every result rendered live in a real browser.
Watch the four way comparison
Specs at a glance
| Kimi K3 | DeepSeek V4 Pro 0813 | MiniMax M3 | MiMo V2.6 Pro | |
|---|---|---|---|---|
| Maker | Moonshot AI | DeepSeek | MiniMax | Xiaomi |
| Context window | 1,000,000 tokens | 1,000,000 tokens | 1,000,000 tokens | 1,000,000 tokens |
| Input | Text, images and video | Text | Text, images and video | Text, images, video and audio |
| Reasoning | Effort up to max | Effort up to max | Thinking on or off | Thinking on or off |
| Structured output | Strict JSON Schema | Strict JSON Schema | JSON mode | Strict JSON Schema |
| Input price | $3.00 per 1M tokens | $1.32 per 1M tokens | $0.225 per 1M tokens up to 512K context, $0.45 above | $0.435 per 1M tokens |
| Output price | $15.00 per 1M tokens | $3.96 per 1M tokens | $0.90 per 1M tokens up to 512K context, $1.80 above | $0.87 per 1M tokens |
How we ran it
Each model received the identical prompt for three tasks: build a washing machine on a spin cycle in a single HTML file; build a plant growing time-lapse in a single HTML file; build popcorn popping in a pot in a single HTML file. Every task asked for a single self-contained HTML file with no external libraries that animates on its own. Kimi K3 and DeepSeek V4 Pro 0813 ran at reasoning_effort: "max"; MiniMax M3 and MiMo V2.6 Pro, which have no effort setting, ran with thinking on. Every model had a 65,536 token output budget and one shot per task. The line counts and tokens-per-second readouts on each panel are measured from the real API calls, and each result is the file the model returned, rendered as-is.
What to look for
MiniMax M3 returned its files fastest, at about 114 tokens per second. MiMo V2.6 Pro wrote the longest files, about 1090 lines on average, and MiniMax M3 the most compact, about 500. MiMo V2.6 Pro used the most output tokens, about 42,000 per task including its reasoning. Watch how each model brings its scene to life: the motion, the detail and how it keeps running on its own. We are not declaring a winner. Run the clip and judge the outputs for your own use case.
Run the same test on EmpirioLabs
curl https://api.empiriolabs.ai/v1/chat/completions \
-H "Authorization: Bearer $EMPIRIOLABS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "kimi-k3",
"reasoning_effort": "max",
"messages": [{"role": "user", "content": "Build a washing machine on a spin cycle in a single HTML file."}]
}'
Swap model to deepseek-v4-pro-0813, minimax-m3 or mimo-v2-6-pro to run the same request against the others, or try them interactively in the playground.
Frequently asked questions
Were the results edited or retried?
No. Each panel is the first file the model returned for that prompt, rendered exactly as returned. A request that came back without a file was sent again.
Why the highest reasoning setting?
A fair head to head shows each model at its best, so each ran at its highest reasoning setting: max effort for Kimi K3 and DeepSeek V4 Pro 0813, thinking on for MiniMax M3 and MiMo V2.6 Pro.
Which model should I use?
MiMo V2.6 Pro has the lowest output price of the four here. Run your own workload against each before deciding: the same request works for all four with only the model name changed.



