Model benchmark · JevBench v1.3.0

Jev models, compared on the same benchmark.

Explore published results for System One decision models. Compare overall scores, decision quality, calibration, speed, and cost to help choose a model for your workload.

Selected ranking

Jev model benchmark scores

A selection of systems from the public board. The composite score weights intelligence, calibration, speed, and cost equally; see the source for the full ranking and interactive comparisons.

Open the full leaderboard
RankModel / systemScoreIntelligenceCalibrationSpeedCost / 1,000
#1Jev 1.13.0TypeSafe AI74.485.782.783.3$0.040Published price
#2SemIfQwen3.5-4B · open rebuild73.179.072.683.7~$0.022Estimated
#3djevMaisa · DiffusionGemma73.082.765.491.4~$0.026Announced price
#11OpenJevDiffusionGemma 26B-A4B66.479.264.883.2~$0.066Estimated
#33LayaModernBERT-large54.445.862.571.1~$0.003Estimated

Costs are USD per 1,000 complete decisions. Estimates and announced but uncharged prices are labelled; these third-party benchmark figures are not Jev AI Model API prices or a promise of cost for your workload.

Source: Benchmark Heaven · JevBench v1.3.0 · Snapshot: September 21, 2026

How to read the scores

What does the benchmark tell you?

The overall score is a quick view of relative performance on one shared task set. Before production, validate accuracy, confidence thresholds, latency, and cost on your own examples.

Intelligence

How often the system answers correctly.

Calibration

Whether confidence reflects observed outcomes.

Speed & cost

Latency and the price of complete decisions.

Try Jev AI Model yourself

Test a decision on your own example.

The online playground is free after sign in. Buy credits for API calls.

Jev is a model released by TypeSafe AI. Jev AI Model is an independent online playground and API service. This page cites a third-party benchmark for reference and is not an endorsement by the model provider or benchmark publisher.