Home Blog

Introducing Aplomb 1

Introducing Aplomb 1, a decision model for text, JSON, images, video and audio

Oct 5, 2026

EmpirioLabs AI

Aplomb 1 is our first model. It takes text, JSON, images, audio and video in a single request, reads up to 1M tokens, and makes a decision on a 1M-token document in about 3 seconds. A short question takes about 15 ms of model time.

It runs on an inference runtime we built for decision models. On our API and in the Playground, long documents go through fast long-context mode: a 1M-token document takes about 3 seconds instead of 111, and fast mode answered all 525 decisions in our long-context tests correctly.

For agents, Aplomb 1 returns a probability for every tool and for each enum and boolean argument in one request, so an agent acts when the model is sure and escalates when it is not. Every answer is calibrated, and it can say when the input does not hold the answer instead of guessing. It is built for the decisions software makes thousands of times a day, from routing a ticket to picking an agent's next tool.

As of October 6, 2026, it is #1 among 4B models on the published Decision Index board, with 44.86 on our run of the official kit, and #1 among models up to 5.3B parameters on 8 of the 38 benchmarks. It is available today in the EmpirioLabs API and Playground, and its weights are on Hugging Face under the EmpirioLabs Model License.

Aplomb 1 at a glance
InputsText, JSON, images, video and audio, in any mix
Window1,000,000 tokens of state plus question
QuestionsYes or no, choice (2 to 255 options), score (2 to 10 levels), tool selection
Not in the stateA probability that the input does not hold the answer
Tool selectionThe next tool to call, with probabilities for its enum and boolean arguments
SpeedAbout 15 ms of model time for a short question; about 3 seconds for a 1M-token document in fast long-context mode
API formatsThe Decisions API, plus the OpenAI, Anthropic and Gemini formats
BillingInput tokens only; output tokens are always zero
Data retentionZero data retention by default
Size5.3 billion parameters, built on Qwen3.5-4B; open weights on Hugging Face
Decision Index 0.2.144.86 on our run of the official kit, #1 among 4B models on the published board as of October 6, 2026

What it answers

Four question types. noul answers yes or no with the probability of yes. choice picks one of 2 to 255 options and returns the full distribution. score places the state on an ordered scale of 2 to 10 levels. tool picks which of your function tools to call and returns distributions over its enum and boolean arguments.

Not in the state. Add "abstain": true to a question and the answer also carries the probability that the state does not contain what the question needs. A record with no carrier field gets a high abstain probability instead of a confident guess.

Independent answers. Up to 128 questions share one state, and each is answered on its own, so adding a question never changes another answer.

Tool selection for agents

Give Aplomb 1 your agent's function tools and the current state (a ticket, a conversation, a record), and it returns the tool to call next with a probability for every tool and a distribution for each enum and boolean argument, in one request. A general LLM writes a tool call out token by token and does not say how sure it is about the tool or any argument. The agent acts when the probability is high and hands the step to a larger model when it is not, so routine steps skip a generated answer. We found no other decision model that returns argument probabilities.

A ticket that reads "Order B-44120 arrived with a cracked screen. I want my money back." with three tools, issue_refund (a reason enum and a full_refund boolean), track_package and escalate_to_human, comes back as (abridged, rounded):

{"tool": "issue_refund",
 "probabilities": {"issue_refund": 0.969, "none": 0.025, "escalate_to_human": 0.005, "track_package": 0.002},
 "arguments": {"issue_refund": {
   "reason": {"value": "damaged", "probabilities": {"damaged": 0.993, "not_as_described": 0.007, "late": 0.001}},
   "full_refund": {"value": true, "probability_true": 0.761}}},
 "open_arguments": {"track_package": ["order_id"]}}

Free-text arguments, such as an order id, are listed under open_arguments for the agent to fill.

Mixed inputs, up to 1M tokens

Text, JSON objects and arrays, images, video and audio go into the same state, in any combination: a support call recording next to the account record, a dashcam clip next to the claim form, a product photo next to the listing. The state plus the longest question can run to 1,000,000 tokens, enough for a long contract bundle, a year of tickets or hours of transcripts in one request.

How it scores

We ran Aplomb 1 and three open 4B decision models on the same items, and set the numbers Intern-Decision published for Jev, JevK5 and SemIf next to them. On the seven-suite bundle Intern-Decision reports, Aplomb 1 averages 88.7%. Against Jev's published row it is ahead on JevBench Hard and typed decisions, level on JevBench Easy, and behind on JevBench Original, ToolACE, AG News and WildJailbreak. Its largest lead is on JevBench Hard, +5.4 points.

Accuracy of Aplomb 1 and other decision models on eleven benchmarks
Figure 1. Accuracy on decision benchmarks. Solid markers are our runs on the same items with 95% bootstrap intervals; hollow markers are numbers Intern-Decision published.
BenchmarkAplomb 1Jev (p)JevK5 (p)SemIf (p)Intern-Decision 4BKev 4BOmniJev 4B
JevBench Easy100.0100.0100.0100.0100.0100.087.5
JevBench Original95.898.697.298.6100.093.172.2
JevBench Hard77.572.173.961.371.254.152.3
Typed decisions74.073.364.562.880.867.061.3
ToolACE89.491.381.085.296.587.488.1
AG News88.889.689.189.290.889.789.6
WildJailbreak95.696.390.592.590.293.592.3
Average88.788.785.284.289.983.577.6

(p) As published by Intern-Decision for the same items. Intern-Decision reports 90.02 for its 4B model on this bundle; the column shows our run.

BenchmarkAplomb 1Intern-Decision 4BKev 4BOmniJev 4B
Calibration pilot70.858.369.853.1
ANLI R354.054.250.8n/a
Banking77 (77 options)74.6n/a84.068.4
CLINC150 (150 options)87.0n/a77.866.8

Long inputs

We measured lookups in order ledgers from 16K to 1M tokens, with the questions about one order placed at a random depth.

Aplomb 1 accuracy from 16K to 1M tokens
Figure 2. Accuracy as the input grows, measured on order ledgers of the stated length.
Input length (tokens)16K128K240K512K1M
Aplomb 1100.0100.0100.0100.096.8

Long documents in seconds

On the EmpirioLabs API and in the Playground, a state longer than 131,072 tokens is read in fast long-context mode by default. A 1M-token document takes about 3 seconds instead of about 111 seconds read in full, and on our long-context test sets fast mode answered all 525 decisions correctly, against 95.2% for a full read.

Decisions correct by document length for fast mode and a full read
Figure 8. Decisions correct by document length on our long-context test sets (164 documents of 131K to 1M tokens, 525 decisions), fast mode and a full read.
Seconds per decision by document length for fast mode and a full read
Figure 9. Median server time per request through the API by document length, fast mode and a full read.
Document length (tokens)228K513K797K979K
Fast mode1.9 s2.4 s2.8 s3.0 s
Full read8.8 s34 s76 s111 s

Send "long_context": "full" to read every token. A full read is the better choice when an answer depends on the whole document, such as a count, the overall tone, or confirming that something never appears: on 24 documents built to test those questions, fast mode was right on 50% of decisions and a full read on 61%. The response's long_context field says which mode was used, and usage counts the whole state in both modes. Fast long-context mode is available only on the EmpirioLabs API and in the Playground; the open weights read every token.

Calibration

A probability is only useful if it means what it says. Across nine text suites, Aplomb 1's stated confidence tracks how often it is right: its pooled calibration error is 3.6, against 5.0 for Intern-Decision 4B and 4.0 for Kev 4B (lower is better).

Reliability diagrams for Aplomb 1 and two open decision models
Figure 3. Stated confidence against observed accuracy, pooled over nine text suites.

Images, video and audio

Benchmark (first 500 items)Aplomb 1Intern-Decision 4BOmniJev 4B
POPE89.089.288.4
MMStar68.665.264.4
AI2D87.081.083.4
MMMU57.257.056.0
HallusionBench79.879.471.2
Aplomb 1 and OmniJev accuracy on three video benchmarks
Figure 4. Decisions about video. Intern-Decision, Kev and Jev do not accept video.
BenchmarkAplomb 1OmniJev 4B
TempCompass79.373.6
Video-MME (short)80.772.2
NExT-QA82.679.8

We added audio to Aplomb 1 ourselves; none of the other decision models above accepts it. It is strongest on speech and vocal sounds; environmental sounds and open-ended questions about audio are harder for it:

BenchmarkAplomb 1
VocalSound (vocal sounds)90.0
CREMA-D (emotion in speech)65.0
ESC-50 (environmental sounds)8.4
MMAU (audio questions)49.1
MMSU (spoken language)43.3

51 languages

Aplomb 1 decides about text in many languages, not only English. On MASSIVE, each request to a voice assistant is written in one of 51 languages while the question and the 59 intent labels stay in English. Without any training on MASSIVE, Aplomb 1 averages 75.0% across the 51 languages when choosing among all 59 intents, 91.0% in English, and 70% or more in 39 of them.

Aplomb 1 intent accuracy in each of 51 languages
Figure 5. Zero-shot MASSIVE intent accuracy, 100 test requests per language, one choice among 59 intents.

For reference, the Laya project reports 40.1% for its multilingual model across the same 51 languages when choosing among 20 intents, and the JevK5 project reports 73.8% in English with all 60 intents.

Decision Index

Decision Index 0.2.1 averages 38 public benchmarks in five areas: knowledge and reasoning, language understanding, retrieval and classification, tools and automation, and arts and human taste. Each benchmark is chance-corrected, so 0 means random guessing and 100 means perfect.

We ran the official kit on Aplomb 1 on October 5, 2026, and it answered all 150,317 scoreable requests. Aplomb 1 scores 44.86. This is our own run of the kit, not an entry on the published board.

Bar chart of Decision Index 0.2.1 scores: Aplomb 1 44.86 on our run of the official kit, ahead of every 4B model on the published board; next are JPT-4B at 43.04 and Jet v6.2 at 42.60.
Figure 6. Decision Index 0.2.1 scores of Aplomb 1 and the 4B models on the published board. The cyan bar is Aplomb 1, our run of the official kit. Gray bars are entries on the published board, data of September 28, 2026.

As of October 6, 2026, Aplomb 1 is #1 among 4B models on the published board, and #1 among models up to 5.3B parameters on 8 of the 38 benchmarks, including MMLU-Pro, GPQA Diamond and BBH. On the published board (data of September 28, 2026), the highest score for a 4B model is JPT-4B's 43.04, 1.82 below Aplomb 1.

The 8 Decision Index benchmarks where Aplomb 1 has the top score among models up to 5.3B on the published board: BBH, MMLU-Pro, GPQA Diamond, NLI4CT, PhishNChips, HoVer, Humicroedit and POP909
The 8 benchmarks where Aplomb 1 has the top score among models up to 5.3B on the published board, from science and reasoning to clinical trials, fact-checking, phishing, music and humor.
Area scores for Aplomb 1 and JPT-4B across knowledge and reasoning, language understanding, retrieval and classification, tools and automation, and arts and human taste.
Figure 7. Area scores, chance-corrected from 0 to 100, for Aplomb 1 and JPT-4B, the best 4B model on the published board.

By area, Aplomb 1 scores highest in tools and automation (62.4) and retrieval and classification (49.5). Compared with JPT-4B, it is ahead in four of the five areas and behind in language understanding (48.3 against 52.5).

Disclosure. Aplomb 1's training data included the public train splits of WinoGrande and ContractNLI, two of the 38 benchmarks in the index. We could not fully verify that index items were excluded from the earliest training data.

How it compares

What each decision model or API accepts and answers, from its own published documentation as of September 29, 2026, with prices as of October 6, 2026:

Aplomb 1 compared with Jev 1.13, the OpenAI Decisions API preview, Intern-Decision 4B and Laya on question types, tool calls with argument probabilities, the not-in-the-input answer, inputs, context window, 1M-token speed, API formats, price per 1M input tokens and open weights
How Aplomb 1 compares, from each model's published documentation as of September 29, 2026, with prices as of October 6, 2026.
Aplomb 1Jev 1.13OpenAI Decisions API (preview)Intern-Decision 4BLaya
Question typesYes or no, choice, score, toolYes or no, choice, scoreChoice among listed answersYes or no, choice, scoreYes or no, choice, score
Tool selection with argument probabilitiesYesNoNot documentedNoNo
Not in the state answerYesNoNot documentedNoNo
InputsText, JSON, images, video, audioTextText, imagesText, imagesText
Window1,000,000 tokens64K tokensNot published8,192 tokens512 to 8,192 tokens
Price per 1M input tokens$0.02$0.042Not announcedSelf-hosted$0.02 on Runware
Open weightsYesNoNoYesYes

Speed and billing

Aplomb 1 runs on EmpirioLabs' own inference runtime for decision models. A short question takes about 15 ms of model time, a question with an image about 35 ms, a 15,000-token document about 0.3 seconds, and a 1M-token state about 3 seconds in fast long-context mode, the default above 131,072 tokens, or about 111 seconds read in full.

You pay for input tokens only: the state once per request and each question's own text, never a prompt template. Output tokens are always zero. The current rate is on the model page and the pricing page.

Zero data retention is on by default: we do not retain the content of your requests or responses.

Open weights

Aplomb 1 is built on Qwen3.5-4B. We extended its window from 262K to 1M tokens, added audio input with the audio encoder of Qwen3-Omni, added our own decision head, which returns every answer as calibrated probabilities instead of generated text, and trained the model for decisions, including tool selection and the not-in-the-input answer. For the API and the Playground, we built and optimized our own inference runtime.

The weights and a reference script for running them yourself are at huggingface.co/empiriolabsai/aplomb-1. The script's answers can differ slightly from the API's.

Research, evaluation, personal use and internal use by organizations with annual revenue under US$1 million are free under the EmpirioLabs Model License; hosting Aplomb 1 as a service or shipping it in a product sold to others needs a commercial license.

Get started

curl https://api.empiriolabs.ai/v1/decisions \
  -H "Authorization: Bearer $EMPIRIOLABS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "aplomb-1",
    "state": {"order": {"id": "B-44120", "status": "shipped", "payment": "unpaid"}},
    "questions": {
      "paid": {"type": "noul", "instructions": "Has order B-44120 been paid?", "abstain": true},
      "carrier": {"type": "choice", "instructions": "Which carrier is delivering order B-44120?",
                  "criteria": {"ups": null, "fedex": null, "dhl": null}, "abstain": true}
    }
  }'

paid comes back as a confident no. carrier comes back with a high abstain probability, because the record does not name a carrier. The full schema, media input and limits are in the Decisions API guide. Aplomb 1 also accepts requests in the OpenAI, Anthropic and Gemini formats, so you can call it from their SDKs; the guide's chat formats section shows how a chat request maps to a decision. Try it without code in the Playground.

Acknowledgements

We built Aplomb 1 with help from our AI agent harness across data, training, evaluation and deployment. Thanks to the Qwen team for Qwen3.5 and Qwen3-Omni, which Aplomb 1 builds on, to Lambda for the GPUs we trained and serve it on, and to the NVIDIA Inception program, the Z.ai startup program and the StepFun startup program for their support.

Ready to use better endpoints?

Explore our models, or contact us about business inquiries, custom deployments, or anything else.