Open weights · Apache-2.0
s1: a calibrated System One decision model
Give it a text or JSON state and typed questions. s1 returns a typed answer (a choice among keyed options, yes/no, or an ordinal level) with a calibrated probability, from one forward pass and no generation.
- baseGemma 4 26B-A4B (mixture of experts)
- tuningLoRA, merged, trained with <bos>
- weightsbf16
- licenceApache-2.0
Key results
Held-out test sets; every system is scored on the same cases by the same scorer, and s1 runs through vLLM in bf16. Differences under about ±1.5 points on JevBench are within noise. Full leaderboard.
Try it
Pull the int4 build (about 17 GB) and run decisions on your own machine, or call it through the repository's
calibrated HTTP API. A hosted demo is ready in space/ and will be switched on when GPU hosting is in place.
Quick start
Single-pass scoring: build the prompt, prefill Answer:, and read the option-letter logits at the last position. Gemma 4's tokenizer does not add <bos> on its own, and the model loses accuracy without it, so prepend it.
import json, torch
from transformers import AutoTokenizer, Gemma4ForConditionalGeneration
REPO = "j-raghavan/s1-gemma4-26b-decision"
tok = AutoTokenizer.from_pretrained(REPO, subfolder="bf16")
model = Gemma4ForConditionalGeneration.from_pretrained(
REPO, subfolder="bf16", dtype=torch.bfloat16, device_map="auto")
state = {"ticket": "I was charged twice for order 4471."}
options = {"billing": "Payments and refunds", "shipping": "Deliveries", "tech": "Bugs"}
letters = "ABCDEFGHIJKLMNOPQRSTUVWXYZ"[: len(options)]
lines = "\n".join(f"{L}) {k}: {d}" for L, (k, d) in zip(letters, options.items()))
user = ("You are a decision model. Read the state and answer the question by choosing one option.\n\n"
f"STATE:\n{json.dumps(state, ensure_ascii=False, indent=1)}\n\n"
"QUESTION: Which team should handle this?\n\n"
f"OPTIONS:\n{lines}\n\nAnswer with the option letter only.")
# Gemma 4's tokenizer does not add <bos> itself: prepend it, then prefill "Answer:".
prompt = (tok.bos_token + "<|turn>user\n" + user + "<turn|>\n"
"<|turn>model\n<|channel>thought\n<channel|>Answer:")
ids = tok(prompt, return_tensors="pt", add_special_tokens=False).to(model.device)
with torch.no_grad():
logits = model(**ids).logits[0, -1] # one forward pass, no generation
letter_ids = [tok.encode(" " + L, add_special_tokens=False)[0] for L in letters]
probs = torch.softmax(logits[letter_ids].float(), -1)
print(dict(zip(options, probs.tolist()))) # raw (uncalibrated) probabilities
The same prompt as a raw Ollama request that returns the log-probabilities of one token.
ollama pull jrlabs01/s1 # the int4 build, about 17 GB
# Send raw prompts that start with <bos>: Ollama's chat formatting changes the prompt and the answers.
curl -s localhost:11434/api/generate -d '{
"model": "jrlabs01/s1", "raw": true, "stream": false,
"logprobs": true, "top_logprobs": 20,
"options": {"temperature": 0, "num_predict": 1},
"prompt": "<bos><|turn>user\nYou are a decision model. Read the state and answer the question by choosing one option.\n\nSTATE:\n{\n \"ticket\": \"I was charged twice for order 4471.\"\n}\n\nQUESTION: Which team should handle this?\n\nOPTIONS:\nA) billing: Payments and refunds\nB) shipping: Deliveries\nC) tech: Bugs\n\nAnswer with the option letter only.<turn|>\n<|turn>model\n<|channel>thought\n<channel|>Answer:"
}'
# Normalise the top_logprobs of the single token over "A", "B", "C".
The repository's POST /v1/decisions server wraps the model with per-type calibration and the typed request shape (choice, noul, score).
uv run --extra api uvicorn api.server:app --port 8000
curl -s localhost:8000/v1/decisions -H 'content-type: application/json' -d '{
"state": {"ticket": "I was charged twice for order 4471."},
"questions": {"team": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "Payments and refunds", "shipping": "Deliveries", "tech": "Bugs"}}}}'
How it works
One forward pass
The state, the question and lettered options go into a fixed prompt, and the reply is prefilled with
Answer:. The model runs once; nothing is generated or parsed.Option-letter logits
The answer distribution is a softmax over the option-letter tokens at the last position. Yes/no questions are two options; ordinal scores are read as an expected level. More than 26 options use two-digit codes, read in two passes.
Calibration
Per answer type, a temperature (choice, score) or Platt scaling (yes/no) maps raw probabilities to calibrated ones. Calibrators are fitted on dev splits only; the test sets are never used.
Details: methodology · model card · data card · lineage of every model and dataset.
Limitations
- On JevBench, s1 (0.808 raw mean accuracy) is still behind Jev 1.13 (0.835). The gap is mostly knowledge questions (MMLU-Pro and the medical slices).
- On the ordinal sentiment slice (sst5), its error (MAE 0.611) is higher than Jev's (0.549).
- The structured-decision test set is generated and rule-labelled by this project. It is held out from training, but it is not an independent benchmark.
- The JevBench test subset resolves differences of about ±1.5 points; smaller gaps between systems are not meaningful.
- s1 answers questions with given options. It is not a chat or generation model, and it was not trained to explain its decisions.
- The calibrated figures assume a calibrator fitted on labelled dev data for the task. On unseen tasks, use the raw probabilities or a per-type default.