We put Kimi K3, GLM 5.3, Qwen3.8 Max 0902 and DeepSeek V4 Pro 0813 against each other across two one-shot coding tests, all run on EmpirioLabs. Same prompt each, one attempt, no edits or retries, and every result rendered live in a real browser.
Watch the four way comparison
Specs at a glance
| Kimi K3 | GLM 5.3 | Qwen3.8 Max 0902 | DeepSeek V4 Pro 0813 | |
|---|---|---|---|---|
| Maker | Moonshot AI | Z.ai | Alibaba | DeepSeek |
| Context window | 1,000,000 tokens | 1,000,000 tokens | 1,000,000 tokens | 1,000,000 tokens |
| Input | Text, images and video | Text | Text, images and video | Text |
| Reasoning effort | Up to max | Up to max | Up to max | Up to max |
| Structured output | Strict JSON Schema | JSON mode | Strict JSON Schema | Strict JSON Schema |
| Input price | $3.00 per 1M tokens | $1.40 per 1M tokens | $2.00 per 1M tokens | $1.32 per 1M tokens |
| Output price | $15.00 per 1M tokens | $4.40 per 1M tokens | $6.00 per 1M tokens | $3.96 per 1M tokens |
How we ran it
Each model received the identical prompt for two tasks: build a Flappy Bird game that plays itself in a single HTML file; build a Missile Command game that plays itself in a single HTML file. Every task asked for a single self-contained HTML file with no external libraries that animates on its own. All four models ran at reasoning_effort: "max" with a 65,536 token output budget, one shot, no retries. The line counts and tokens-per-second readouts on each panel are measured from the real API calls, and each result is the file the model returned, rendered as-is.
What to look for
DeepSeek V4 Pro 0813 returned its files fastest, at about 115 tokens per second. DeepSeek V4 Pro 0813 wrote the longest files, about 920 lines on average, and Kimi K3 the most compact, about 490. GLM 5.3 used the most output tokens, about 59,000 per task including its reasoning. Watch how each model brings its scene to life: the motion, the detail and how it keeps running on its own. We are not declaring a winner. Run the clip and judge the outputs for your own use case.
Run the same test on EmpirioLabs
curl https://api.empiriolabs.ai/v1/chat/completions \
-H "Authorization: Bearer $EMPIRIOLABS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "kimi-k3",
"reasoning_effort": "max",
"messages": [{"role": "user", "content": "Build a Flappy Bird game that plays itself in a single HTML file."}]
}'
Swap model to glm-5-3, qwen3-8-max-0902 or deepseek-v4-pro-0813 to run the same request against the others, or try them interactively in the playground.
Frequently asked questions
Were the results edited or retried?
No. Each model got one attempt per task with the identical prompt, and the rendered result is exactly the file it returned.
Why max reasoning?
A fair head to head shows each model at its best. All four expose a reasoning_effort control on EmpirioLabs, so all four ran at the highest setting.
Which model should I use?
DeepSeek V4 Pro 0813 has the lowest output price of the four here. Run your own workload against each before deciding: the same request works for all four with only the model name changed.



