NOMA · OPEN DECISION MODEL

Decisions in milliseconds.

Noma reads a state and a few typed questions and returns a probability for every option, a separate abstain signal, and a measure of its own uncertainty. It never generates text.

GitHub ↗Hugging Face ↗PyPI ↗pip install blackdrome-noma
MEASURED

Fast,calibrated,andcheaptorun.

16 msMedian per decision

End to end over HTTP on one H100, one question per request.

$0.023Per 1,000 decisions

JevBench method. $0.026 self-hosted on the measured H100.

100%JevBench easy

48 of 48. 98.6% on the original tier (71 of 72).

82.6%Sealed set

386 human-reviewed decisions across 12 families, never used in training.

0.011Calibration error

Easy tier. 0.042 on the sealed set.

18 / 32Layers used

Full depth and a 9B backbone were no more accurate.

WHERE IT FITS

Thedecisionlayerofanagent.

01

Route

Which queue, which model tier, does this need tools or a person.

02

Verify

Did the last step succeed, is the next action safe, is the task done.

03

Decide

Classify, score and check, with a probability for every option.

Noma routes requests before an agent runs and verifies each step after
EVIDENCE

Everynumber,withhowitwasmeasured.

Median latency per decision: Noma 16 ms, decider-4b v2 17 ms, Cygnet 35 ms, NInfer Flash-Next 79 ms, JevOne 87 ms, Nimble 9B 389 ms, OpenJev 463 ms, Jev 1.13.0 652 ms
Cost per 1,000 decisions: decider-4b v2 $0.020, Noma $0.023, Cygnet $0.037, Jev 1.13.0 $0.040, Nimble 9B $0.166
Accuracy: 100% on JevBench easy, 98.6% on JevBench original, 82.6% on the sealed set
Ablations: depth and size, training set size, and targeted multi-step data

Noma's figures are our own measurements with the JevBench client and have not yet been submitted to the leaderboard. Other systems' figures are the published leaderboard values. The full evaluation, including the multi-step reasoning tier, is in the repository.

HOW IT WORKS

Oneforwardpass,nodecoding.

Request, fact channel, cut backbone, decision heads, answer
A decision head on a measured cut

Listwise option scoring on the first 18 of 32 backbone layers. Abstain is its own calibrated output, and four bootstrap-trained heads give an uncertainty estimate from one backbone pass.

Fast serving for a hybrid backbone

The state is processed once and its cache, including linear-attention recurrent state, is forked across questions. Length buckets and CUDA graphs take model time from about 2.3 s to 14 ms.

A fact channel

Deterministic preprocessing turns dates, durations, running totals and thresholds into short fact lines. They are hints to the model, never overrides.

Blind, agreement-gated labelling

Two different frontier models label each item blind; a third judges only their disagreements and a random audit sample.

Decision head: hidden states at marked positions feed four listwise scorers; their mean is the answer and their disagreement is the uncertainty
USE IT

Onefile,onecommand.

pip install blackdrome-noma
noma serve        # API on /v1/systemone, playground on /

from noma import Noma

model = Noma.from_pretrained("BlackdromeAILabs/noma")
answers, _ = model.decide(
    state="Deploy 4/6 finished. 3 of 12 pods failing readiness.",
    questions={
        "step_ok": {"type": "noul", "instructions": "Did the rollout succeed?"},
    },
)

The weights are one 5.4 GB file with the backbone and adapter already merged. It runs in 5.2 GB of GPU memory and speaks the same /v1/systemone API as Jev, so existing clients work unchanged.

A local playground ships with the server: paste a state, build questions, and copy the request as code.

The Noma playground answering two questions about a contract clause
SCOPE

Builtforsingle-passdecisions.

Noma classifies, routes, scores and verifies. Questions that need several chained steps of arithmetic or date reasoning belong with a reasoning model, and Noma's abstain and uncertainty signals are there to hand them off. Text only, states up to 4,096 tokens, evaluated in English. Open weights under MPL-2.0.