Leaderboard
Typed-decision accuracy and calibration on held-out test sets. Every number below is computed from the score files in the repository's results/ directory. Click a column header to sort.
JevBench
A 1,700-case subset of the JevBench test set: 100 families per text slice, sampled with seed 20261004 (build_subset.py). Jev 1.13 and Laya 421M are the predictions published with the dataset, filtered to the same cases. Mean accuracy is over the 11 choice and yes/no slices.
| System | Mean accuracy 11 slices, higher is better | Mean ECE lower is better | sst5 MAE lower is better |
|---|---|---|---|
| Jev 1.13 (TypeSafe AI, closed) closed | 0.835 | 0.081 | 0.549 |
| s1 (Gemma 4 26B-A4B + s1 fine-tune), bf16 s1 | 0.808 | 0.083 | 0.611 |
| s1, int4 via Ollama s1 | 0.801 | 0.074 | 0.633 |
| Gemma 4 26B-A4B, untuned (with <bos>) untuned | 0.770 | 0.218 | 0.522 |
| Gemma 4 12B, untuned untuned | 0.757 | 0.225 | 0.599 |
| Gemma 4 31B, untuned untuned | 0.756 | 0.236 | 0.644 |
| gpt-oss-120b, untuned untuned | 0.670 | 0.222 | 0.782 |
| s1 encoder (ModernBERT-large 396M, earlier approach) earlier s1 | 0.492 | 0.161 | 0.715 |
| Laya 421M | 0.485 | 0.224 | 0.817 |
| Von 1.3 | 0.465 | 0.201 | 0.694 |
Per-slice results
Each cell is headline / ECE. The headline is accuracy, except on score slices (MAE of the expected level, lower is better; no ECE). Hover a cell for its 95% bootstrap interval.
| Slice | Type | Jev 1.13 | s1 bf16 | s1 int4 | Gemma 4 26B-A4B | Gemma 4 12B | Gemma 4 31B | gpt-oss-120b | s1 encoder | Laya 421M | Von 1.3 |
|---|---|---|---|---|---|---|---|---|---|---|---|
aegis2 | yes/no | 0.790 / 0.063 | 0.810 / 0.092 | 0.820 / 0.053 | 0.820 / 0.182 | 0.780 / 0.212 | 0.830 / 0.170 | 0.820 / 0.044 | 0.690 / 0.046 | 0.570 / 0.244 | 0.560 / 0.330 |
aegis2_response | yes/no | 0.840 / 0.056 | 0.820 / 0.068 | 0.810 / 0.079 | 0.720 / 0.276 | 0.800 / 0.196 | 0.780 / 0.221 | 0.720 / 0.171 | 0.580 / 0.169 | 0.490 / 0.227 | 0.560 / 0.120 |
atbench500 | yes/no | 0.910 / 0.122 | 0.830 / 0.086 | 0.820 / 0.058 | 0.690 / 0.313 | 0.730 / 0.246 | 0.730 / 0.257 | 0.340 / 0.475 | 0.180 / 0.545 | 0.680 / 0.127 | 0.290 / 0.442 |
banking77 | choice | 0.860 / 0.081 | 0.785 / 0.155 | 0.780 / 0.141 | 0.795 / 0.195 | 0.775 / 0.196 | 0.760 / 0.230 | 0.800 / 0.160 | 0.665 / 0.125 | 0.430 / 0.505 | 0.800 / 0.165 |
jailbreak_classification | yes/no | 0.950 / 0.031 | 0.950 / 0.091 | 0.960 / 0.121 | 0.950 / 0.050 | 0.960 / 0.042 | 0.970 / 0.032 | 0.990 / 0.113 | 0.830 / 0.092 | 0.920 / 0.027 | 0.520 / 0.132 |
medmcqa | choice | 0.825 / 0.047 | 0.715 / 0.082 | 0.720 / 0.048 | 0.705 / 0.265 | 0.645 / 0.309 | 0.560 / 0.422 | 0.655 / 0.231 | 0.345 / 0.178 | 0.255 / 0.232 | 0.270 / 0.229 |
medqa_usmle | choice | 0.845 / 0.045 | 0.790 / 0.053 | 0.755 / 0.059 | 0.770 / 0.215 | 0.685 / 0.280 | 0.790 / 0.208 | 0.655 / 0.274 | 0.270 / 0.236 | 0.280 / 0.191 | 0.260 / 0.087 |
mmlu_pro | choice | 0.755 / 0.067 | 0.600 / 0.094 | 0.585 / 0.069 | 0.610 / 0.334 | 0.530 / 0.401 | 0.460 / 0.497 | 0.410 / 0.400 | 0.190 / 0.072 | 0.145 / 0.210 | 0.220 / 0.314 |
prompt_injections | yes/no | 0.720 / 0.177 | 0.920 / 0.080 | 0.930 / 0.073 | 0.820 / 0.178 | 0.830 / 0.170 | 0.860 / 0.145 | 0.800 / 0.044 | 0.590 / 0.103 | 0.780 / 0.055 | 0.550 / 0.280 |
pubmedqa | choice | 0.750 / 0.183 | 0.710 / 0.082 | 0.670 / 0.073 | 0.650 / 0.328 | 0.680 / 0.335 | 0.660 / 0.336 | 0.370 / 0.432 | 0.390 / 0.144 | 0.290 / 0.498 | 0.490 / 0.046 |
scienceqa_text | choice | 0.945 / 0.024 | 0.955 / 0.033 | 0.960 / 0.035 | 0.935 / 0.062 | 0.910 / 0.087 | 0.920 / 0.075 | 0.805 / 0.098 | 0.680 / 0.066 | 0.490 / 0.146 | 0.590 / 0.067 |
sst5 | MAE | 0.549 | 0.611 | 0.633 | 0.522 | 0.599 | 0.644 | 0.782 | 0.715 | 0.817 | 0.694 |
Calibrated
With per-task calibrators fitted on the dev splits only (calibrate_from_dev.py), the JevBench mean accuracy is 0.817 for s1 and 0.856 for Jev 1.13. Yes/no calibration (Platt scaling) can move the decision threshold, which is why accuracy changes; choice calibration (a temperature) never changes the answer. Figures from the project README.
Structured-decision test set
Rule-labelled decisions over JSON state and slide layouts, generated by this project (generate.py); no model is involved in labelling and every family is evaluation-only. Mean accuracy is over the 12 choice and yes/no slices; the 3 score slices are reported as MAE. Jev 1.13, Laya and Von have no predictions on this set.
| System | Mean accuracy 12 slices, higher is better | Mean ECE lower is better | ppt_density MAE lower is better | state_config_risk MAE lower is better | state_incident_severity MAE lower is better |
|---|---|---|---|---|---|
| Gemma 4 31B, untuned untuned | 0.978 | 0.021 | 0.739 | 0.011 | 0.062 |
| s1 (Gemma 4 26B-A4B + s1 fine-tune), bf16 s1 | 0.971 | 0.033 | 0.272 | 0.166 | 0.085 |
| s1, int4 via Ollama s1 | 0.966 | 0.043 | 0.317 | 0.184 | 0.115 |
| Gemma 4 26B-A4B, untuned (with <bos>) untuned | 0.948 | 0.052 | 0.218 | 0.116 | 0.102 |
| gpt-oss-120b, untuned untuned | 0.853 | 0.129 | 0.850 | 0.650 | 0.544 |
| s1 encoder (ModernBERT-large 396M, earlier approach) earlier s1 | 0.610 | 0.284 | 0.906 | 0.739 | 0.599 |
Per-slice results
| Slice | Type | Gemma 4 31B | s1 bf16 | s1 int4 | Gemma 4 26B-A4B | gpt-oss-120b | s1 encoder |
|---|---|---|---|---|---|---|---|
ppt_density | MAE | 0.739 | 0.272 | 0.317 | 0.218 | 0.850 | 0.906 |
ppt_layout | choice | 1.000 / 0.000 | 1.000 / 0.003 | 1.000 / 0.005 | 0.988 / 0.016 | 0.925 / 0.067 | 0.725 / 0.155 |
ppt_needs_image | yes/no | 1.000 / 0.000 | 1.000 / 0.006 | 1.000 / 0.011 | 1.000 / 0.000 | 0.975 / 0.079 | 0.350 / 0.594 |
ppt_split_slide | yes/no | 0.963 / 0.032 | 0.988 / 0.041 | 0.988 / 0.069 | 0.988 / 0.012 | 0.850 / 0.148 | 0.637 / 0.135 |
ppt_visual_type | choice | 1.000 / 0.000 | 1.000 / 0.004 | 1.000 / 0.006 | 1.000 / 0.000 | 1.000 / 0.011 | 0.988 / 0.242 |
state_access_request | yes/no | 1.000 / 0.001 | 1.000 / 0.004 | 1.000 / 0.007 | 0.988 / 0.014 | 0.588 / 0.281 | 0.537 / 0.378 |
state_agent_loop | yes/no | 1.000 / 0.000 | 1.000 / 0.031 | 1.000 / 0.045 | 1.000 / 0.000 | 0.825 / 0.082 | 0.650 / 0.198 |
state_agent_next_tool | choice | 1.000 / 0.000 | 1.000 / 0.008 | 1.000 / 0.008 | 1.000 / 0.000 | 0.912 / 0.048 | 0.125 / 0.505 |
state_api_retry | yes/no | 1.000 / 0.000 | 0.963 / 0.027 | 0.975 / 0.037 | 0.950 / 0.051 | 0.938 / 0.079 | 0.487 / 0.340 |
state_ci_failure | choice | 1.000 / 0.000 | 1.000 / 0.001 | 1.000 / 0.001 | 1.000 / 0.000 | 1.000 / 0.004 | 0.887 / 0.234 |
state_config_risk | MAE | 0.011 | 0.166 | 0.184 | 0.116 | 0.650 | 0.739 |
state_incident_severity | MAE | 0.062 | 0.085 | 0.115 | 0.102 | 0.544 | 0.599 |
state_log_anomaly | choice | 1.000 / 0.000 | 1.000 / 0.005 | 1.000 / 0.009 | 1.000 / 0.000 | 1.000 / 0.008 | 0.825 / 0.080 |
state_refund_eligibility | yes/no | 0.912 / 0.082 | 0.800 / 0.206 | 0.750 / 0.232 | 0.600 / 0.397 | 0.500 / 0.491 | 0.500 / 0.300 |
state_ticket_routing | choice | 0.863 / 0.134 | 0.900 / 0.066 | 0.875 / 0.087 | 0.863 / 0.138 | 0.725 / 0.254 | 0.613 / 0.250 |
Methodology notes
- Noise
- The 1,700-case subset resolves differences of about ±1.5 points; smaller differences are within noise. Compare systems with paired tests, not by rank.
- Score slices
- Ordinal slices (for example sst5) report the MAE of the expected level, where lower is better. They are excluded from mean accuracy so that an error is never averaged with accuracies.
- ECE
- Expected calibration error with 10 equal-width bins on the top probability, averaged over the accuracy slices. Raw (uncalibrated) probabilities unless stated.
- Dev and test
- Checkpoints, data mixes and calibrators are chosen on dev splits only. The JevBench dev split is sampled with its own seed from families disjoint from the test subset, and cases whose text also appears in the test subset or the training targets are dropped (build_dev.py). The structured-decision dev split comes from the same rules with a different seed. Test sets are used only for reporting.
- Scorer
- All systems go through the upstream JevBench scorer on identical cases (score.py). Raw tables: JevBench summary, structured-decision summary.
More: methodology · data card · model card · lineage.