Fugu Ultra v2.0 is available on EmpirioLabs. Fugu is not a single model. It is a conductor: one API call fans the request out across a pool of expert models and composes their work into a single answer. Ultra is the flagship of that family, and v2.0 is tuned for peak capability on complex, multi-step work such as autonomous reasoning, research, and software development.
What Fugu Ultra v2.0 supports
Fugu Ultra v2.0 takes text and image input with a 1M token context window. It supports function calling, strict JSON Schema structured output, and built-in web search that cites its sources when they are available. Reasoning is always on, and output is capped at 131,072 tokens per request.
How to call Fugu Ultra v2.0
Use the OpenAI-compatible chat completions endpoint and set model to fugu-ultra-v2-0:
curl https://api.empiriolabs.ai/v1/chat/completions \
-H 'Authorization: Bearer $EMPIRIOLABS_API_KEY' \
-H 'Content-Type: application/json' \
-d '{
"model": "fugu-ultra-v2-0",
"reasoning_effort": "xhigh",
"messages": [
{
"role": "user",
"content": "Design a migration plan to move a busy write-heavy service onto a new database, and flag what could go wrong."
}
]
}'
The same model is available through the Responses, Messages, and Google-compatible content-generation endpoints. When the conductor returns a reasoning summary, it arrives as message.reasoning_content alongside the answer.
Moving from v1.0 or v1.1
Version ids are explicit and pinned, so nothing moves under you. fugu-ultra-v1-0 and fugu-ultra-v1-1 keep serving their own snapshots, and the legacy fugu-ultra id stays pinned to v1.0. Switching to v2.0 is a one-line change.
One real behaviour difference to check before you migrate: v1.1 treats max as a distinct maximum reasoning level above xhigh. On v2.0, max is accepted as an alias of xhigh. Code that deliberately reached for v1.1's max will still run on v2.0, but it now selects the same effort as xhigh.
Three things callers get wrong
1. The token counts are bigger than your prompt. The conductor's fan-out to other models is real token usage, billed as ordinary input and output tokens. A short prompt with web search enabled can report a six-figure input token count. Budget against the returned usage block, not the length of your own message.
2. The context tier is chosen by the billed input, not by your prompt. Fugu Ultra v2.0 has a higher rate above 272K input tokens, and that threshold is measured on the full billed input including orchestration. So a modest prompt can still land in the higher tier on a heavy request. If tier stability matters more to you than peak capability, Fugu Max has one flat rate at every context length.
3. Streaming does not arrive token by token. stream: true is accepted and keeps long requests alive, but the conductor buffers the whole orchestration and delivers the reasoning and the answer in one burst at the end. Requests can take from a few seconds to a few minutes, so set generous client timeouts. Also note max_tokens must be at least 16, and small values can truncate or empty the answer.
Pricing
Fugu Ultra v2.0 is pay as you go with no subscription. Input, output, and cached input are context-tiered, with a higher rate above 272K input tokens, and cached input is billed at a steep discount to ordinary input. Built-in web search carries no separate fee on Ultra, because its cost is already inside the orchestration tokens. Current rates are on the Fugu Ultra v2.0 model page and the pricing page, which always render the live catalog.
Try it
Run a prompt in the playground, or read the full parameter reference in the API docs.



