Fugu Max is available on EmpirioLabs. Fugu is not a single model. It is a conductor: one API call fans the request out across a pool of expert models and composes their work into a single answer. Fugu Max is the cost-efficient member of that family, drawing on the widest pool Sakana has assembled, including open-weight and specialist models, and assembling a right-sized team for each task instead of always reaching for the most expensive option.
What Fugu Max supports
Fugu Max takes text and image input with a 1M token context window. It supports function calling, strict JSON Schema structured output, and built-in web search with page fetch. Reasoning is always on.
Output is capped at 131,072 tokens per request. Image input is accepted through the normal OpenAI-compatible content parts, and EmpirioLabs handles the upload plumbing for you.
How to call Fugu Max
Use the OpenAI-compatible chat completions endpoint and set model to fugu-max:
curl https://api.empiriolabs.ai/v1/chat/completions \
-H 'Authorization: Bearer $EMPIRIOLABS_API_KEY' \
-H 'Content-Type: application/json' \
-d '{
"model": "fugu-max",
"messages": [
{
"role": "user",
"content": "Compare three approaches to rate limiting a public API, and recommend one for a 5000 req/s service."
}
]
}'
The same model is available through the Responses, Messages, and Google-compatible content-generation endpoints. Switching from another Fugu model is a one-line change, since the request and response contract is identical.
Web search and structured output
Set tool_web_search to true when an answer needs current information. The model runs its own searches and page fetches, cites sources when it has them, and reports exactly what it used in usage.tool_usage, for example {"web_search": 3, "web_extractor": 1}. Those counts are billable and are already included in the reported cost.
For machine-readable output, pass a strict JSON Schema in response_format. Fugu Max returns exactly conformant JSON, so you can parse the result without a repair step.
Three things callers get wrong
1. The token counts are bigger than your prompt. Because the conductor works by fanning out to other models, the orchestration it performs is real token usage and is billed as ordinary input and output tokens. A one-sentence prompt with web search enabled can report tens of thousands of input tokens. This is expected, not a bug, and it is why the returned usage block is the number to budget against rather than the length of your own message.
2. Streaming does not arrive token by token. stream: true is accepted and is useful for keeping long requests alive, but the conductor buffers the whole orchestration and delivers the reasoning and the answer in one burst at the end. If your UI shows a typing effect, expect it to sit idle and then fill in at once.
3. Only three reasoning levels exist. reasoning_effort accepts high, xhigh, and max, and nothing else. Sending low, medium, or none returns a 400. On Fugu Max, max is accepted as an alias of xhigh. Reasoning cannot be turned off. Separately, max_tokens must be at least 16, and small values can truncate or empty the answer because the conductor needs room to work.
Pricing
Fugu Max is pay as you go with no subscription, and it is the only Fugu model with a single flat token rate at every context length, so a long prompt never moves you to a higher tier. Built-in web search and page fetch are billed per executed call on top of tokens, and a single request can run more than one. Current rates for input, output, cached input, and the per-call web tools are on the Fugu Max model page and the pricing page, which always render the live catalog.
Try it
Run a prompt in the playground, or read the full parameter reference in the API docs.



