01 / Start with the output your application needs

Write down one real decision: choose a support queue, judge whether a passage supports an answer, or rate urgency against a defined rubric. Include an unknown outcome. A decision model is a candidate when the allowed answers are bounded; a generated explanation or a long investigation may need a different step.

This guide is an editorial selection framework based on publisher documentation checked on October 5, 2026. We have not run a common evaluation of these models. Model detail pages keep the reviewed version and sources together.

02 / Choose a deployment path

An API reduces the serving work you own. Local deployment gives you control over the runtime and where inputs are processed, but you also own capacity, updates and evaluation. Make that choice before comparing a hosted token rate with a downloadable checkpoint. The table describes documented options, not equal-quality replacements.

Choose a deployment path
ModelAccess documented in this catalogHosted input price
JevHosted endpoint$0.042 / 1M input tokens
LayaSelf-hosted weightsHosted price unverified
Kev-4BSelf-hosted weightsHosted price unverified
GLiNER2.5-DecideSelf-hosted weightsHosted price unverified
Bespoke NimbleSelf-hosted weightsHosted price unverified
Tev1Hosted endpoint; Self-hosted weightsHosted price unverified

03 / Compare what the probabilities mean

Use the chosen answer to measure accuracy. Use the probability distribution to inspect uncertainty. A distribution-concentration score, a token logprob and an observed success rate measure different things; a matching numeric range does not make them interchangeable.

Build a labelled validation set from your own requests. Check whether cases assigned similar probabilities succeed at similar rates. Choose a fallback rule there, then freeze it before evaluating on held-out examples. Revisit the rule after changing model versions or the task.

04 / Build a shortlist by constraint

For a hosted starting point, inspect Jev and its versioned API terms. For local deployment, compare the Laya and Kev model cards against your language, document length and hardware. For combined extraction and decisions, investigate GLiNER2.5-Decide. Nimble and Tev1 provide useful material for studying a training workflow. These are starting points for evaluation, not a quality ranking.

A runtime that serves an existing model belongs in a tooling list. An API-compatible implementation should not be counted as a distinct trained model unless its weights or training establish that distinction.

05 / Use one evaluation contract

Keep labels, input records and acceptance criteria fixed. Record model/checkpoint, language, input length, number of options, hardware or endpoint, concurrency and test date. Report per-task accuracy, calibration, p50/p95 end-to-end latency and total cost. Keep warm and cold runs separate.

Report coverage too: a system can improve the accuracy of automatic decisions simply by sending more cases for review. Compare candidates at a similar error budget and disclose how many requests still need another step. Unknown measurements stay unknown.