Aplomb is a decision model. You send it a state, which can be text, a JSON record, images, video, audio or any mix of them, together with typed questions. It answers every question as a probability distribution with a confidence value, in one request, and it never generates text.
It is built for the decisions software makes thousands of times a day: whether a ticket is urgent, which team owns it, whether a claim is covered by a policy, what a caller is upset about, which tool an agent should call next. Each answer is a number you can put a threshold on.
Aplomb is available today in the EmpirioLabs API and in the Playground, and its weights are on Hugging Face under the EmpirioLabs Model License.
| Aplomb at a glance | |
|---|---|
| Inputs | Text, JSON, images, video and audio, in any mix |
| Window | 1,000,000 tokens of state plus question |
| Questions | Yes or no, choice (2 to 255 options), score (2 to 10 levels), tool selection |
| Speed | About 15 ms of server time for a short question |
| Billing | Input tokens only; output tokens are always zero |
| Size | 5.3 billion parameters, open weights on Hugging Face |
What it answers
Four question types. noul answers yes or no with the probability of yes. choice picks one of 2 to 255 options and returns the full distribution. score places the state on an ordered scale of 2 to 10 levels. tool picks which of your function tools to call and returns distributions over its enum and boolean arguments.
Not in the state. Add "abstain": true to a question and the answer also carries the probability that the state does not contain what the question needs. A record with no carrier field gets a high abstain probability instead of a confident guess.
Independent answers. Up to 128 questions share one state, and each is answered on its own, so adding a question never changes another answer.
Every input, one million tokens
Text, JSON objects and arrays, images, video and audio go into the same state, in any combination: a support call recording next to the account record, a dashcam clip next to the claim form, a product photo next to the listing. The state plus the longest question can run to 1,000,000 tokens, enough for a long contract bundle, a year of tickets or hours of transcripts in one request.
How it scores
We ran Aplomb and three open 4B decision models on the same items, and set the numbers Intern-Decision published for Jev, JevK5 and SemIf next to them. On the seven-suite bundle Intern-Decision reports, Aplomb averages 87.2%. Against Jev's published row it is ahead on typed decisions, ToolACE and AG News, level on JevBench Easy and JevBench Original, and behind on JevBench Hard and WildJailbreak. Its weakest suite is JevBench Hard.

| Benchmark | Aplomb | Jev (p) | JevK5 (p) | SemIf (p) | Intern-Decision 4B | Kev 4B | OmniJev 4B |
|---|---|---|---|---|---|---|---|
| JevBench Easy | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 87.5 |
| JevBench Original | 98.6 | 98.6 | 97.2 | 98.6 | 100.0 | 93.1 | 72.2 |
| JevBench Hard | 57.7 | 72.1 | 73.9 | 61.3 | 71.2 | 54.1 | 52.3 |
| Typed decisions | 79.4 | 73.3 | 64.5 | 62.8 | 80.8 | 67.0 | 61.3 |
| ToolACE | 91.9 | 91.3 | 81.0 | 85.2 | 96.5 | 87.4 | 88.1 |
| AG News | 89.7 | 89.6 | 89.1 | 89.2 | 90.8 | 89.7 | 89.6 |
| WildJailbreak | 93.2 | 96.3 | 90.5 | 92.5 | 90.2 | 93.5 | 92.3 |
| Average | 87.2 | 88.7 | 85.2 | 84.2 | 89.9 | 83.5 | 77.6 |
(p) As published by Intern-Decision for the same items. Intern-Decision reports 90.02 for its 4B model on this bundle; the column shows our run.
| Benchmark | Aplomb | Intern-Decision 4B | Kev 4B | OmniJev 4B |
|---|---|---|---|---|
| Calibration pilot | 69.8 | 58.3 | 69.8 | 53.1 |
| ANLI R3 | 59.2 | 54.2 | 50.8 | n/a |
| Banking77 (77 options) | 75.4 | n/a | 84.0 | 68.4 |
| CLINC150 (150 options) | 83.8 | n/a | 77.8 | 66.8 |
Long inputs
We measured lookups in order ledgers from 16 thousand to one million tokens, with the questions about one order placed at a random depth.

| Input length (tokens) | 16K | 128K | 240K | 512K | 1M |
|---|---|---|---|---|---|
| Aplomb | 100.0 | 100.0 | 100.0 | 93.3 | 69.8 |
Calibration
A probability is only useful if it means what it says. Across nine text suites, Aplomb's stated confidence tracks how often it is right: its pooled calibration error is 3.1, against 5.0 for Intern-Decision 4B and 4.0 for Kev 4B (lower is better).

Images, video and audio
| Benchmark (first 500 items) | Aplomb | Intern-Decision 4B | OmniJev 4B |
|---|---|---|---|
| POPE | 88.6 | 89.2 | 88.4 |
| MMStar | 68.0 | 65.2 | 64.4 |
| AI2D | 86.8 | 81.0 | 83.4 |
| MMMU | 58.0 | 57.0 | 56.0 |
| HallusionBench | 81.2 | 79.4 | 71.2 |

| Benchmark | Aplomb | OmniJev 4B |
|---|---|---|
| TempCompass | 76.8 | 73.6 |
| Video-MME (short) | 80.6 | 72.2 |
| NExT-QA | 84.0 | 79.8 |
None of the other decision models above accepts audio. Aplomb is strongest on speech and vocal sounds; environmental sounds and open-ended questions about audio are harder for it:
| Benchmark | Aplomb |
|---|---|
| VocalSound (vocal sounds) | 90.4 |
| CREMA-D (emotion in speech) | 76.7 |
| ESC-50 (environmental sounds) | 10.1 |
| MMAU (audio questions) | 45.9 |
| MMSU (spoken language) | 43.3 |
Speed and billing
Aplomb runs on EmpirioLabs' own inference runtime for decision models. A short question takes about 15 ms of server time, and a state of a million tokens with several questions about 90 seconds.
You pay for input tokens only: the state once per request and each question's own text, never a prompt template. Output tokens are always zero. The current rate is on the model page and the pricing page.
Open weights
The weights and a reference script that returns the same answers as the API are at huggingface.co/empiriolabsai/aplomb. Research, evaluation, personal use and internal use by organizations with annual revenue under US$1 million are free under the EmpirioLabs Model License; hosting Aplomb as a service or shipping it in a product sold to others needs a commercial license.
Get started
curl https://api.empiriolabs.ai/v1/decisions \
-H "Authorization: Bearer $EMPIRIOLABS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "aplomb",
"state": {"order": {"id": "B-44120", "status": "shipped", "payment": "unpaid"}},
"questions": {
"paid": {"type": "noul", "instructions": "Has order B-44120 been paid?", "abstain": true},
"carrier": {"type": "choice", "instructions": "Which carrier is delivering order B-44120?",
"criteria": {"ups": null, "fedex": null, "dhl": null}, "abstain": true}
}
}'
paid comes back as a confident no. carrier comes back with a high abstain probability, because the record does not name a carrier. The full schema, media input and limits are in the Decisions API guide. Try it without code in the Playground.
Acknowledgements
Thanks to Lambda for the GPUs we trained and serve Aplomb on, to the Z.ai startup program and the StepFun startup program for the models behind the agent harness we built Aplomb with, and to the NVIDIA Inception program.



