SIDE BY SIDE
Compare selected AI models side by side.
Compare uses, prices, speed evidence, and version details for up to three models. Language, video, decision models and historical evaluations are shown separately.
Decision & classification models
| Decision criteria | Jev ↗TypeSafe AI | Laya ↗Convai Innovations | Kev-4B ↗Jared Palmer |
|---|---|---|---|
| Version reviewed | jev-1.13.0 | Laya family · model card checked 2026-10-05 | Kev 1.0 · Kev-4B |
| Judgment types | Choose an option · Yes / no probability · Rubric score | Choose an option · Yes / no probability · Rubric score | Choose an option · Yes / no probability · Rubric score |
| Access & hardware | TypeSafe API. Use a versioned ID when evaluating; aliases can move. Customer-specific fine-tuning is not offered. | CPU or GPU self-hosting, with a Jev-compatible server. This entry covers the downloadable checkpoints; hosted service terms are not verified here. | Self-host on an NVIDIA GPU or Apple Silicon via MLX. Training and serving code are available; CPU deployment is not verified here. |
| API price · USD | $0.042 / 1M input tokensTypeSafe direct API, USD per million input tokens; output is not billed. Per-decision cost depends on input size, questions and retries.TypeSafe · model versions and pricing ↗ | Hosted price unverifiedNo hosted rate verified for this entry. Self-hosting needs compute, capacity and maintenance; downloadable weights do not make inference free. | Hosted price unverifiedNo hosted rate verified for this entry. Self-hosting needs compute, capacity and maintenance; downloadable weights do not make inference free. |
| Languages | English is the primary training language; evaluate other languages separately. | Separate English and multilingual checkpoints; the publisher reports 100+ languages for the multilingual model. | English; the model card places other languages outside its intended scope. |
| Context conditions | 64k tokens across a request; state plus the longest question must fit 32k. Text input only. | English checkpoint: 512 tokens. Multilingual: default 1,024, configurable up to 8,192; long-input quality must be tested. | Validated context: 8,192 tokens. The server accepts longer states, but that is not equivalent to validated accuracy. |
| Probability interpretation | Choice and Score include option distributions and a derived confidence value. Noul returns a yes probability without a separate confidence field. Validate calibration on your workload. | Publisher reports calibration training. Select and record the checkpoint, runtime and any fitted temperature before using its probabilities. | The release applies a fitted temperature. Probability thresholds need rechecking after model, domain or calibration changes. |
| Our accuracy / calibration | Not measured | Not measured | Not measured |
| Our p50 / p95 latency | Not measured | Not measured | Not measured |
| Our cost / 1,000 decisions | Not measured | Not measured | Not measured |
| Limits | A typed result can still be wrong. Define an unknown outcome and a fallback. Serving limits and alias targets can change; record the actual model returned. | English and multilingual checkpoints have different limits and behavior. Published short-input latency is hardware-specific. We have not reproduced it or compared Chinese accuracy. | API compatibility does not establish equal accuracy or transferable thresholds. Keep the checkpoint and precision with any benchmark; this catalog does not combine its different evaluation suites. |
| Sources | TypeSafe · model versions and pricing ↗TypeSafe · decision primitives ↗TypeSafe · interpreting confidence ↗Checked 2026-10-05 | Convai Innovations · Laya model card ↗Checked 2026-10-05 | Jared Palmer · Kev-4B model card ↗Checked 2026-10-05 |
Source-reviewed capabilities, not a leaderboard. Hosted prices and self-hosting costs use different bases. Unknown measurements are not zero. Measure the complete task →