Introducing Aplomb

Introducing Aplomb, a decision model for text, JSON, images, video and audio

Sep 29, 2026

EmpirioLabs AI

Aplomb is a decision model. You send it a state, which can be text, a JSON record, images, video, audio or any mix of them, together with typed questions. It answers every question as a probability distribution with a confidence value, in one request, and it never generates text.

It is built for the decisions software makes thousands of times a day: whether a ticket is urgent, which team owns it, whether a claim is covered by a policy, what a caller is upset about, which tool an agent should call next. Each answer is a number you can put a threshold on.

Aplomb is available today in the EmpirioLabs API and in the Playground, and its weights are on Hugging Face under the EmpirioLabs Model License.

Aplomb at a glance
InputsText, JSON, images, video and audio, in any mix
Window1,000,000 tokens of state plus question
QuestionsYes or no, choice (2 to 255 options), score (2 to 10 levels), tool selection
SpeedAbout 15 ms of server time for a short question
BillingInput tokens only; output tokens are always zero
Size5.3 billion parameters, open weights on Hugging Face

What it answers

Four question types. noul answers yes or no with the probability of yes. choice picks one of 2 to 255 options and returns the full distribution. score places the state on an ordered scale of 2 to 10 levels. tool picks which of your function tools to call and returns distributions over its enum and boolean arguments.

Not in the state. Add "abstain": true to a question and the answer also carries the probability that the state does not contain what the question needs. A record with no carrier field gets a high abstain probability instead of a confident guess.

Independent answers. Up to 128 questions share one state, and each is answered on its own, so adding a question never changes another answer.

Every input, one million tokens

Text, JSON objects and arrays, images, video and audio go into the same state, in any combination: a support call recording next to the account record, a dashcam clip next to the claim form, a product photo next to the listing. The state plus the longest question can run to 1,000,000 tokens, enough for a long contract bundle, a year of tickets or hours of transcripts in one request.

How it scores

We ran Aplomb and three open 4B decision models on the same items, and set the numbers Intern-Decision published for Jev, JevK5 and SemIf next to them. On the seven-suite bundle Intern-Decision reports, Aplomb averages 87.2%. Against Jev's published row it is ahead on typed decisions, ToolACE and AG News, level on JevBench Easy and JevBench Original, and behind on JevBench Hard and WildJailbreak. Its weakest suite is JevBench Hard.

Accuracy of Aplomb and other decision models on eleven benchmarks
Figure 1. Accuracy on decision benchmarks. Solid markers are our runs on the same items with 95% bootstrap intervals; hollow markers are numbers Intern-Decision published.
BenchmarkAplombJev (p)JevK5 (p)SemIf (p)Intern-Decision 4BKev 4BOmniJev 4B
JevBench Easy100.0100.0100.0100.0100.0100.087.5
JevBench Original98.698.697.298.6100.093.172.2
JevBench Hard57.772.173.961.371.254.152.3
Typed decisions79.473.364.562.880.867.061.3
ToolACE91.991.381.085.296.587.488.1
AG News89.789.689.189.290.889.789.6
WildJailbreak93.296.390.592.590.293.592.3
Average87.288.785.284.289.983.577.6

(p) As published by Intern-Decision for the same items. Intern-Decision reports 90.02 for its 4B model on this bundle; the column shows our run.

BenchmarkAplombIntern-Decision 4BKev 4BOmniJev 4B
Calibration pilot69.858.369.853.1
ANLI R359.254.250.8n/a
Banking77 (77 options)75.4n/a84.068.4
CLINC150 (150 options)83.8n/a77.866.8

Long inputs

We measured lookups in order ledgers from 16 thousand to one million tokens, with the questions about one order placed at a random depth.

Aplomb accuracy from 16K to one million tokens
Figure 2. Accuracy as the input grows, measured on order ledgers of the stated length.
Input length (tokens)16K128K240K512K1M
Aplomb100.0100.0100.093.369.8

Calibration

A probability is only useful if it means what it says. Across nine text suites, Aplomb's stated confidence tracks how often it is right: its pooled calibration error is 3.1, against 5.0 for Intern-Decision 4B and 4.0 for Kev 4B (lower is better).

Reliability diagrams for Aplomb and two open decision models
Figure 3. Stated confidence against observed accuracy, pooled over nine text suites.

Images, video and audio

Benchmark (first 500 items)AplombIntern-Decision 4BOmniJev 4B
POPE88.689.288.4
MMStar68.065.264.4
AI2D86.881.083.4
MMMU58.057.056.0
HallusionBench81.279.471.2
Aplomb and OmniJev accuracy on three video benchmarks
Figure 4. Decisions about video. Intern-Decision, Kev and Jev do not accept video.
BenchmarkAplombOmniJev 4B
TempCompass76.873.6
Video-MME (short)80.672.2
NExT-QA84.079.8

None of the other decision models above accepts audio. Aplomb is strongest on speech and vocal sounds; environmental sounds and open-ended questions about audio are harder for it:

BenchmarkAplomb
VocalSound (vocal sounds)90.4
CREMA-D (emotion in speech)76.7
ESC-50 (environmental sounds)10.1
MMAU (audio questions)45.9
MMSU (spoken language)43.3

Speed and billing

Aplomb runs on EmpirioLabs' own inference runtime for decision models. A short question takes about 15 ms of server time, and a state of a million tokens with several questions about 90 seconds.

You pay for input tokens only: the state once per request and each question's own text, never a prompt template. Output tokens are always zero. The current rate is on the model page and the pricing page.

Open weights

The weights and a reference script that returns the same answers as the API are at huggingface.co/empiriolabsai/aplomb. Research, evaluation, personal use and internal use by organizations with annual revenue under US$1 million are free under the EmpirioLabs Model License; hosting Aplomb as a service or shipping it in a product sold to others needs a commercial license.

Get started

curl https://api.empiriolabs.ai/v1/decisions \
  -H "Authorization: Bearer $EMPIRIOLABS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "aplomb",
    "state": {"order": {"id": "B-44120", "status": "shipped", "payment": "unpaid"}},
    "questions": {
      "paid": {"type": "noul", "instructions": "Has order B-44120 been paid?", "abstain": true},
      "carrier": {"type": "choice", "instructions": "Which carrier is delivering order B-44120?",
                  "criteria": {"ups": null, "fedex": null, "dhl": null}, "abstain": true}
    }
  }'

paid comes back as a confident no. carrier comes back with a high abstain probability, because the record does not name a carrier. The full schema, media input and limits are in the Decisions API guide. Try it without code in the Playground.

Acknowledgements

Thanks to Lambda for the GPUs we trained and serve Aplomb on, to the Z.ai startup program and the StepFun startup program for the models behind the agent harness we built Aplomb with, and to the NVIDIA Inception program.

より良いエンドポイントを使う準備はできていますか?

当社のモデルをご覧いただくか、ビジネスの問い合わせ、カスタム展開、その他何でもご連絡ください。