s1 decision model

Open weights · Apache-2.0

s1: a calibrated System One decision model

Give it a text or JSON state and typed questions. s1 returns a typed answer (a choice among keyed options, yes/no, or an ordinal level) with a calibrated probability, from one forward pass and no generation.

  • baseGemma 4 26B-A4B (mixture of experts)
  • tuningLoRA, merged, trained with <bos>
  • weightsbf16
  • licenceApache-2.0

Key results

Held-out test sets; every system is scored on the same cases by the same scorer, and s1 runs through vLLM in bf16. Differences under about ±1.5 points on JevBench are within noise. Full leaderboard.

0.808
JevBench mean accuracy, raw
Jev 1.13: 0.835. Untuned base with <bos>: 0.770. 11 slices, 1,700 held-out cases.
0.817
JevBench, calibrated
Jev 1.13: 0.856. Calibrators fitted on dev splits only.
0.083
Mean calibration error (ECE), raw
Jev 1.13: 0.081. Untuned base: 0.218. 10 bins, lower is better.
0.971
Structured-decision set, raw
Untuned base: 0.948. Calibrated: 0.981. JSON-state and slide-layout rules; Jev has no predictions on this set.

Try it

Pull the int4 build (about 17 GB) and run decisions on your own machine, or call it through the repository's calibrated HTTP API. A hosted demo is ready in space/ and will be switched on when GPU hosting is in place.

Run it locally with Ollama

Quick start

Single-pass scoring: build the prompt, prefill Answer:, and read the option-letter logits at the last position. Gemma 4's tokenizer does not add <bos> on its own, and the model loses accuracy without it, so prepend it.

import json, torch
from transformers import AutoTokenizer, Gemma4ForConditionalGeneration

REPO = "j-raghavan/s1-gemma4-26b-decision"
tok = AutoTokenizer.from_pretrained(REPO, subfolder="bf16")
model = Gemma4ForConditionalGeneration.from_pretrained(
    REPO, subfolder="bf16", dtype=torch.bfloat16, device_map="auto")

state = {"ticket": "I was charged twice for order 4471."}
options = {"billing": "Payments and refunds", "shipping": "Deliveries", "tech": "Bugs"}
letters = "ABCDEFGHIJKLMNOPQRSTUVWXYZ"[: len(options)]
lines = "\n".join(f"{L}) {k}: {d}" for L, (k, d) in zip(letters, options.items()))
user = ("You are a decision model. Read the state and answer the question by choosing one option.\n\n"
        f"STATE:\n{json.dumps(state, ensure_ascii=False, indent=1)}\n\n"
        "QUESTION: Which team should handle this?\n\n"
        f"OPTIONS:\n{lines}\n\nAnswer with the option letter only.")
# Gemma 4's tokenizer does not add <bos> itself: prepend it, then prefill "Answer:".
prompt = (tok.bos_token + "<|turn>user\n" + user + "<turn|>\n"
          "<|turn>model\n<|channel>thought\n<channel|>Answer:")

ids = tok(prompt, return_tensors="pt", add_special_tokens=False).to(model.device)
with torch.no_grad():
    logits = model(**ids).logits[0, -1]          # one forward pass, no generation
letter_ids = [tok.encode(" " + L, add_special_tokens=False)[0] for L in letters]
probs = torch.softmax(logits[letter_ids].float(), -1)
print(dict(zip(options, probs.tolist())))         # raw (uncalibrated) probabilities

How it works

  1. One forward pass

    The state, the question and lettered options go into a fixed prompt, and the reply is prefilled with Answer:. The model runs once; nothing is generated or parsed.

  2. Option-letter logits

    The answer distribution is a softmax over the option-letter tokens at the last position. Yes/no questions are two options; ordinal scores are read as an expected level. More than 26 options use two-digit codes, read in two passes.

  3. Calibration

    Per answer type, a temperature (choice, score) or Platt scaling (yes/no) maps raw probabilities to calibrated ones. Calibrators are fitted on dev splits only; the test sets are never used.

Details: methodology · model card · data card · lineage of every model and dataset.

Limitations