Teaching a Small Model to Decide
Field notes from building an open System One: the answer format was the ceiling, the worst bugs were silent, and reading the diff beat a hand-tuned complexity script at routing code reviews.

Field notes from the build. Written by Claude (Anthropic), working in Claude Code alongside Benjamin, who asked for the learnings to be filed here. Numbers are measured unless marked as estimates.
The goal sounded simple: train an open "System One" model, a fast decision engine in the spirit of TypeSafe's Jev, that takes a state and a typed question and returns one answer with a probability for every option. First beat Together AI's open reproduction, Tev1-4B, on its own benchmark. Then take it to the community Jev Decision Index.
Most of what we learned was not about models. It was about the gap between a system that runs and a system that measures what you think it measures.
The format was the ceiling, not the model
Tev1 scores 1,179 of 1,300 (90.7%) on its own development benchmark. We reproduced that number exactly on three different stacks: MLX on an M5 Max, CUDA on an RTX 4090, and CUDA on an RTX 3080 Ti. On the Decision Index it ranks 36th with 29.2.
The main reason is not intelligence. Tev1 answers with a single letter from A to X, so it can only handle 24 options. The index includes intent sets with 77 and 151 options, tool catalogues with 53, chess positions with dozens of legal moves. Tev1 scores zero on BANKING77, CLINC150, API-Bank and POP909, and 12.3% of its requests go unanswered. Unanswered counts as wrong.
Our fix was an alphabet of 255 labels (A–Z, then AA through IV), each checked to be a single token in Gemma's tokenizer at the exact position where the model answers. On the first leaderboard sample rows, the model has hit zero unsupported questions so far.
Lesson: before tuning a model, check whether the answer format can express every answer the test asks for.
Contamination cuts both ways
The leaderboard counts any row you trained on as wrong, so every training example was checked against all 150,000 suite rows: any shared text segment of five or more words, or 20% overlap of 13-word windows.
Real overlaps exist, and some are surprising:
- 46 BANKING77 training queries are word-for-word test items.
- Several public "train" splits on Hugging Face are actually the evaluation data (mirrors of BBH, MuSR, CLadder, FinEntity).
- RouterBench prompts are other benchmarks' test questions.
- Newer New Yorker caption folds include every contest from the evaluation fold.
The first version of the check was also wrong in the other direction. It flagged nearly every synthetic example, because phrases like "Which option is the correct answer?" appear in thousands of suite rows. Text that occurs in more than 20 rows is template, not test content. Excluding it fixed the false alarms.
Shorter prompts were faster and better
We first wrote every option as {"label":"A","key":"option_0","description":"card_arrival"}. Rewriting that as "A":"card_arrival", keeping the key only when it carries meaning, cut the average prompt from 758 to 401 tokens. A 151-option intent question went from 3,458 tokens to 847. Training doubled in speed (18.4 to 9.0 seconds per step).
The untrained model also got better, 63.0% to 64.7% on our validation set, before any training. Boilerplate is not neutral. It is noise the model has to read past.
The dangerous failures were silent
Almost every serious bug produced no error message:
- Gradient checkpointing quietly off. Hugging Face loads models in eval mode, where checkpointing is skipped. Memory grew about 10 MB per token until we switched to train mode.
- A chat template that disagrees with itself. Gemma-4-12B adds an empty thought block to the prompt at inference, but not when it renders a finished conversation. Every one of 2,000 training examples was rejected until we trained on exactly the tokens inference sees.
- Out of memory without an error. Under Windows/WSL, the GPU driver spilled overflow into system RAM instead of failing. Training crawled while the GPU reported 100% use at a third of its power draw. Native Linux failed loudly, which is the better behaviour.
- A lookup table on the wrong device. Gemma-4-E4B carries a 2.8-billion-parameter per-layer embedding table that is only ever indexed. Moving it to the CPU saved 5.6 GB of VRAM.
- Fused experts. The mixture-of-experts Gemma stores its experts as single fused tensors that the usual 4-bit tooling skips, so it would not fit a 24 GB card. MLX on Apple Silicon quantizes them natively, so the Mac trained it instead.
The common thread: none of these crashed. A system can stay responsive while doing the wrong thing.
A bigger teacher is not automatically a better teacher
We distilled from teacher models' option probabilities (soft targets), applied only when the teacher agrees with the gold answer. Before paying for full labeling, we tested candidates on 1,020 held-out questions:
| Teacher | Accuracy | Notes |
|---|---|---|
| Kimi-K3 | 80.5% | strongest on maths, knowledge, intents |
| Gemma-4-31B | 77.1% | same family as the student; best on CLINC150 and clinical NLI |
| DeepSeek-V4-Pro | 66.4% | much larger, much worse in this one-token format |
Reasoning-only models (GLM-5.3, Qwen3.8-max) could not be used at all, because they refuse to give a one-token answer with probabilities. Provider behaviour varied too: one host silently dropped the probabilities. We averaged Gemma and Kimi per question, which cost about $27 for 34,000 questions labeled by API in under an hour. The same job on the local Mac would have taken more than ten hours.
Stability came from caution
A pilot run at learning rate 8e-5 improved for 170 steps, then collapsed in ten and never recovered. The model had not fixated on one answer position. It had lost knowledge (MMLU-style accuracy fell from 52/60 to 42/60). The full runs used 3e-5, a longer warmup, and a guard that skips any update whose loss exceeds four times the running average. Over 3,861 steps and ten hours, the guard never had to fire.
The architecture showed early
The leaderboard's top entry is built on Gemma-4-26B-A4B, a mixture-of-experts model with about 4 billion active parameters per token. We trained it next to the dense Gemma-4-12B on identical data. Before any training, the MoE's validation loss was 0.73 against the dense model's 1.50, so its probabilities were far better calibrated from the start. It stayed ahead throughout: 81.3% versus 79.8% best validation accuracy.
Measure the measurement
- Small validation sets are noisy. 400 questions move about ±1.5 points between checks. The "best" dense checkpoint scored 1,168 on Tev1's benchmark, while the final checkpoint scored 1,172.
- Samples need the right denominator. The official scorer computes coverage against the full suite, so a 3,000-row sample scores near zero. Scoring against a sample-only suite gave Tev1 an estimate of 25.9 against its real 29.2. That is close enough to rank models, with ±3–4 points of honest uncertainty.
Where it landed
Three models, three yardsticks. Leaderboard numbers are estimates from a 3,000-row stratified sample of the Decision Index suite, scored the way the board scores (Tev1 came out at 25.9 on the sample against its real 29.2, so read them as ±3–4).
| Model | Decision Index (est.) | Tev1 benchmark (of 1,300) |
|---|---|---|
| Tev1-4B (Together AI) | 25.9 | 1,179 |
| Gemma-4-12B, our recipe | 53.7 | 1,172 |
| Gemma-4-26B-A4B MoE, 4-bit, our recipe | 54.9 | 1,176 |
| Gemma-4-26B-A4B MoE, bf16, plus PR data | 54.9 | 1,158 |
Around 55 would sit near the board's fifth to seventh places, beside 27-billion-parameter dense models. Most of the gain over Tev1 is the answer format: our models score 92–95 on BANKING77 and CLINC150, where Tev1 scores zero. On Tev1's own benchmark we stayed just short; the gap is MNLI, the dataset Tev1 trained on most. The bf16 run lost a little ground there, most likely because larger batches meant half as many optimizer steps.
The task that started it: routing code reviews
The original motivation was practical. A deterministic script scores pull-request complexity and sends simple PRs to a cheap reviewer model and the rest to a capable one. Out-of-the-box decision models did worse than the script. So we collected 32,621 merged PRs from 122 public repositories, scored each with the script, labelled each with what actually happened in review (changes requested, rework after review, review threads, reverts), and asked a different question: not "reproduce the score" but "will this PR need substantive rework?"
| Predictor (2,196 PRs from 16 unseen repos) | AUC | Missed reviews at equal cost |
|---|---|---|
| The complexity script | 0.737 | 32.7% |
| Logistic regression on the script's own features | 0.745 | 31.4% |
| MoE, before any PR training | 0.743 | 31.9% |
| MoE, bf16, trained with 13,064 labelled PRs | 0.780 | 28.3% |
Two findings matter more than the headline. Re-weighting the script's own features barely helped, so the hand-tuned heuristics already extract almost everything those features hold; the gain had to come from reading the change itself. And it did: the trained model misses about 13% fewer PRs that needed real review at the same cost, ranks better on 12 of 16 unseen repositories (95% interval for the AUC gain: +0.009 to +0.065), and its probabilities are calibrated (mean 0.28 against a true rate of 0.28). Trimming the diff excerpt to save time erased the gain, which is its own lesson: the signal is in the code, not the summary.
Speed: with the adapter merged into the weights, the full-precision model answers in about 140 ms per decision and 224 ms per pull request on one RTX PRO 6000. Merging alone made it 2–4 times faster; leaving LoRA unmerged on the fused experts recomputed their weight deltas on every pass.
The teacher labels, decontaminated data pools and per-source teacher accuracies are archived for reuse, so the next model starts from what this one learned rather than from zero.