Home Blog

Kimi K3 vs GLM 5.3 vs Qwen3.8 Max vs DeepSeek V4 Pro: four one-shot coding tests

Kimi K3 vs GLM 5.3 vs Qwen3.8 Max vs DeepSeek V4 Pro: four one-shot coding tests

Oct 4, 2026

EmpirioLabs AI

We put Kimi K3, GLM 5.3, Qwen3.8 Max and DeepSeek V4 Pro 0813 against each other across four one-shot coding tests, all run on EmpirioLabs. Same prompt each, one attempt, no edits or retries, and every result rendered live in a real browser.

Watch the four way comparison

Specs at a glance

Kimi K3GLM 5.3Qwen3.8 MaxDeepSeek V4 Pro 0813
MakerMoonshot AIZ.aiAlibabaDeepSeek
Context window1,000,000 tokens1,000,000 tokens1,000,000 tokens1,000,000 tokens
InputText, images and videoTextText, images and videoText
Reasoning controlreasoning_effort, low to maxreasoning_effort, low to maxreasoning_effort, none to maxreasoning_effort, none to max
Structured outputStrict JSON SchemaJSON modeStrict JSON SchemaStrict JSON Schema
Input price$3.00 per 1M tokens$1.40 per 1M tokens$2.00 per 1M tokens$1.32 per 1M tokens
Output price$15.00 per 1M tokens$4.40 per 1M tokens$6.00 per 1M tokens$3.96 per 1M tokens

How we ran it

Each model received the identical prompt for four tasks: a maze that generates and then solves itself with an animated A* search, a falling-sand simulation, an ocean sunset with a sailboat riding the waves, and rain running down a window at night. Every task asked for a single self-contained HTML file with no external libraries. All four models ran at reasoning_effort: "max" with a 65,536 token output budget, one shot, no retries. The line counts and tokens-per-second readouts on each panel are measured from the real API calls, and each result is the file the model returned, rendered as-is.

What to look for

DeepSeek V4 Pro 0813 returned its files fastest, at about 100 tokens per second. Qwen3.8 Max wrote the longest files, about 540 lines on average, and Kimi K3 the most compact, about 360. GLM 5.3 used the most output tokens, about 47,000 per task including its reasoning. Watch how each model animates the maze search, piles the sand, moves the waves and merges the raindrops. We are not declaring a winner. Run the clip and judge the outputs for your own use case.

Run the same test on EmpirioLabs

curl https://api.empiriolabs.ai/v1/chat/completions \
  -H "Authorization: Bearer $EMPIRIOLABS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "kimi-k3",
    "reasoning_effort": "max",
    "messages": [{"role": "user", "content": "Build a maze that generates and solves itself in a single HTML file."}]
  }'

Swap model to glm-5-3, qwen3-8-max or deepseek-v4-pro-0813 to run the same request against the others, or try them interactively in the playground.

Frequently asked questions

Were the results edited or retried?

No. Each model got one attempt per task with the identical prompt, and the rendered result is exactly the file it returned.

Why max reasoning?

A fair head to head shows each model at its best. All four expose a reasoning_effort control on EmpirioLabs, so all four ran at the highest setting.

Which model should I use?

DeepSeek V4 Pro 0813 has the lowest per-token price of the four and answered the quickest here. Kimi K3 and Qwen3.8 Max accept images and video as well as text. GLM 5.3 sits close to DeepSeek on price. Run your own workload against each before deciding.

Try it

Playground | All models | Pricing

Ready to use better endpoints?

Explore our models, or contact us about business inquiries, custom deployments, or anything else.