s1 decision model

Leaderboard

Typed-decision accuracy and calibration on held-out test sets. Every number below is computed from the score files in the repository's results/ directory. Click a column header to sort.

JevBench

A 1,700-case subset of the JevBench test set: 100 families per text slice, sampled with seed 20261004 (build_subset.py). Jev 1.13 and Laya 421M are the predictions published with the dataset, filtered to the same cases. Mean accuracy is over the 11 choice and yes/no slices.

JevBench test subset, 1,700 cases
SystemMean accuracy
11 slices, higher is better
Mean ECE
lower is better
sst5 MAE
lower is better
Jev 1.13 (TypeSafe AI, closed) closed0.8350.0810.549
s1 (Gemma 4 26B-A4B + s1 fine-tune), bf16 s10.8080.0830.611
s1, int4 via Ollama s10.8010.0740.633
Gemma 4 26B-A4B, untuned (with <bos>) untuned0.7700.2180.522
Gemma 4 12B, untuned untuned0.7570.2250.599
Gemma 4 31B, untuned untuned0.7560.2360.644
gpt-oss-120b, untuned untuned0.6700.2220.782
s1 encoder (ModernBERT-large 396M, earlier approach) earlier s10.4920.1610.715
Laya 421M 0.4850.2240.817
Von 1.3 0.4650.2010.694
Per-slice results

Each cell is headline / ECE. The headline is accuracy, except on score slices (MAE of the expected level, lower is better; no ECE). Hover a cell for its 95% bootstrap interval.

JevBench per-slice results: headline / ECE
SliceTypeJev 1.13s1 bf16s1 int4Gemma 4 26B-A4BGemma 4 12BGemma 4 31Bgpt-oss-120bs1 encoderLaya 421MVon 1.3
aegis2yes/no0.790 / 0.0630.810 / 0.0920.820 / 0.0530.820 / 0.1820.780 / 0.2120.830 / 0.1700.820 / 0.0440.690 / 0.0460.570 / 0.2440.560 / 0.330
aegis2_responseyes/no0.840 / 0.0560.820 / 0.0680.810 / 0.0790.720 / 0.2760.800 / 0.1960.780 / 0.2210.720 / 0.1710.580 / 0.1690.490 / 0.2270.560 / 0.120
atbench500yes/no0.910 / 0.1220.830 / 0.0860.820 / 0.0580.690 / 0.3130.730 / 0.2460.730 / 0.2570.340 / 0.4750.180 / 0.5450.680 / 0.1270.290 / 0.442
banking77choice0.860 / 0.0810.785 / 0.1550.780 / 0.1410.795 / 0.1950.775 / 0.1960.760 / 0.2300.800 / 0.1600.665 / 0.1250.430 / 0.5050.800 / 0.165
jailbreak_classificationyes/no0.950 / 0.0310.950 / 0.0910.960 / 0.1210.950 / 0.0500.960 / 0.0420.970 / 0.0320.990 / 0.1130.830 / 0.0920.920 / 0.0270.520 / 0.132
medmcqachoice0.825 / 0.0470.715 / 0.0820.720 / 0.0480.705 / 0.2650.645 / 0.3090.560 / 0.4220.655 / 0.2310.345 / 0.1780.255 / 0.2320.270 / 0.229
medqa_usmlechoice0.845 / 0.0450.790 / 0.0530.755 / 0.0590.770 / 0.2150.685 / 0.2800.790 / 0.2080.655 / 0.2740.270 / 0.2360.280 / 0.1910.260 / 0.087
mmlu_prochoice0.755 / 0.0670.600 / 0.0940.585 / 0.0690.610 / 0.3340.530 / 0.4010.460 / 0.4970.410 / 0.4000.190 / 0.0720.145 / 0.2100.220 / 0.314
prompt_injectionsyes/no0.720 / 0.1770.920 / 0.0800.930 / 0.0730.820 / 0.1780.830 / 0.1700.860 / 0.1450.800 / 0.0440.590 / 0.1030.780 / 0.0550.550 / 0.280
pubmedqachoice0.750 / 0.1830.710 / 0.0820.670 / 0.0730.650 / 0.3280.680 / 0.3350.660 / 0.3360.370 / 0.4320.390 / 0.1440.290 / 0.4980.490 / 0.046
scienceqa_textchoice0.945 / 0.0240.955 / 0.0330.960 / 0.0350.935 / 0.0620.910 / 0.0870.920 / 0.0750.805 / 0.0980.680 / 0.0660.490 / 0.1460.590 / 0.067
sst5MAE0.5490.6110.6330.5220.5990.6440.7820.7150.8170.694

Calibrated

With per-task calibrators fitted on the dev splits only (calibrate_from_dev.py), the JevBench mean accuracy is 0.817 for s1 and 0.856 for Jev 1.13. Yes/no calibration (Platt scaling) can move the decision threshold, which is why accuracy changes; choice calibration (a temperature) never changes the answer. Figures from the project README.

Structured-decision test set

Rule-labelled decisions over JSON state and slide layouts, generated by this project (generate.py); no model is involved in labelling and every family is evaluation-only. Mean accuracy is over the 12 choice and yes/no slices; the 3 score slices are reported as MAE. Jev 1.13, Laya and Von have no predictions on this set.

Structured-decision test set
SystemMean accuracy
12 slices, higher is better
Mean ECE
lower is better
ppt_density MAE
lower is better
state_config_risk MAE
lower is better
state_incident_severity MAE
lower is better
Gemma 4 31B, untuned untuned0.9780.0210.7390.0110.062
s1 (Gemma 4 26B-A4B + s1 fine-tune), bf16 s10.9710.0330.2720.1660.085
s1, int4 via Ollama s10.9660.0430.3170.1840.115
Gemma 4 26B-A4B, untuned (with <bos>) untuned0.9480.0520.2180.1160.102
gpt-oss-120b, untuned untuned0.8530.1290.8500.6500.544
s1 encoder (ModernBERT-large 396M, earlier approach) earlier s10.6100.2840.9060.7390.599
Per-slice results
Structured-decision per-slice results: headline / ECE
SliceTypeGemma 4 31Bs1 bf16s1 int4Gemma 4 26B-A4Bgpt-oss-120bs1 encoder
ppt_densityMAE0.7390.2720.3170.2180.8500.906
ppt_layoutchoice1.000 / 0.0001.000 / 0.0031.000 / 0.0050.988 / 0.0160.925 / 0.0670.725 / 0.155
ppt_needs_imageyes/no1.000 / 0.0001.000 / 0.0061.000 / 0.0111.000 / 0.0000.975 / 0.0790.350 / 0.594
ppt_split_slideyes/no0.963 / 0.0320.988 / 0.0410.988 / 0.0690.988 / 0.0120.850 / 0.1480.637 / 0.135
ppt_visual_typechoice1.000 / 0.0001.000 / 0.0041.000 / 0.0061.000 / 0.0001.000 / 0.0110.988 / 0.242
state_access_requestyes/no1.000 / 0.0011.000 / 0.0041.000 / 0.0070.988 / 0.0140.588 / 0.2810.537 / 0.378
state_agent_loopyes/no1.000 / 0.0001.000 / 0.0311.000 / 0.0451.000 / 0.0000.825 / 0.0820.650 / 0.198
state_agent_next_toolchoice1.000 / 0.0001.000 / 0.0081.000 / 0.0081.000 / 0.0000.912 / 0.0480.125 / 0.505
state_api_retryyes/no1.000 / 0.0000.963 / 0.0270.975 / 0.0370.950 / 0.0510.938 / 0.0790.487 / 0.340
state_ci_failurechoice1.000 / 0.0001.000 / 0.0011.000 / 0.0011.000 / 0.0001.000 / 0.0040.887 / 0.234
state_config_riskMAE0.0110.1660.1840.1160.6500.739
state_incident_severityMAE0.0620.0850.1150.1020.5440.599
state_log_anomalychoice1.000 / 0.0001.000 / 0.0051.000 / 0.0091.000 / 0.0001.000 / 0.0080.825 / 0.080
state_refund_eligibilityyes/no0.912 / 0.0820.800 / 0.2060.750 / 0.2320.600 / 0.3970.500 / 0.4910.500 / 0.300
state_ticket_routingchoice0.863 / 0.1340.900 / 0.0660.875 / 0.0870.863 / 0.1380.725 / 0.2540.613 / 0.250

Methodology notes

Noise
The 1,700-case subset resolves differences of about ±1.5 points; smaller differences are within noise. Compare systems with paired tests, not by rank.
Score slices
Ordinal slices (for example sst5) report the MAE of the expected level, where lower is better. They are excluded from mean accuracy so that an error is never averaged with accuracies.
ECE
Expected calibration error with 10 equal-width bins on the top probability, averaged over the accuracy slices. Raw (uncalibrated) probabilities unless stated.
Dev and test
Checkpoints, data mixes and calibrators are chosen on dev splits only. The JevBench dev split is sampled with its own seed from families disjoint from the test subset, and cases whose text also appears in the test subset or the training targets are dropped (build_dev.py). The structured-decision dev split comes from the same rules with a different seed. Test sets are used only for reporting.
Scorer
All systems go through the upstream JevBench scorer on identical cases (score.py). Raw tables: JevBench summary, structured-decision summary.

More: methodology · data card · model card · lineage.