A model’s check on another model depends on what it is handed, which model it is, and what the work tells it.
Six results, grouped by the decision each one bears on. Each names its model, its number of tasks, whether it was registered before the run, and when it ran. Where a later run changed how a result reads, the change sits beside it.
What to hand a checker
One checker, three things to check against
| What gemma4:12b is given | July | Re-run, 23 September |
|---|---|---|
| The correct value, introduced as correct | 1.000 | 1.000 |
| A test it can apply: one division of a 3-digit number | 0.536 | 0.536 |
| Nothing; it has to work the answer out | 0.034 | 0.042 |
gemma4:12b, local · rows one and three: sums it solves itself at 1.000, n = 500, pre-registered, 2026-07-30 · row two: n = 250, one exploratory run, 2026-07-31 · no unreadable verdicts in any row · re-run with the leak fixed, same n, 2026-09-23
On 17 September the lab found gemma4:12b had been leaking hidden template text into its replies. It re-ran every cell of that model this site quotes with the leak fixed, and each landed inside the interval it registered before the re-run. Both readings are shown.
Cost is irrelevant at both ends: spending more at the top changes nothing because it is already solved, and spending more at the bottom changes nothing because the problem is access rather than effort. The architectural claim this makes falsifiable is uncomfortable: oversight cannot be bootstrapped from models alone unless a fallible reference suffices.
What has been measured since this table was first published. The composition cells ran on 31 July and 1 August 2026, n = 250 each, one exploratory run per cell. Handed a reference worked out by the 9B model, right on 99.6% of items and introduced as a second model’s answer, the 12B checker read 0.072, and 0.092 when re-run with the leak fixed on 26 September. The same kind of number introduced as the correct value is the 1.000 in row one. So row one measures deference to whatever is called correct, and August tested what that costs. Handed a false reference that agrees with a corrupted answer, gemma4:12b passed the corrupted answer on 0.969 of items, and 0.967 re-run on 22 September. Of nineteen defences tried against that, one held, and the same clause did nothing on three other large models: gemma4:31b passed 0.094 where deepseek-r1:32b, qwen2.5-coder:32b and qwen3.6:35b-a3b passed 0.693, 0.980 and 0.994 (n = 512 per cell, two seeds on the survivor). Its resistance held when it read another model’s work: 0.094 on its own answers, 0.113 on the 9B model’s. The ladder of reference forms between a supplied answer and none is still unbuilt.
Which checker to pick
Which model it was mattered more than its size
A 9B model from one family out-verified a 12B from another by 0.77 on the same paired tasks. Which model it was mattered more than how many parameters it had, which means the monitor you should pick may run on a desk.
qwen3.5:9b 0.804 vs gemma4:12b 0.034 · n = 500 · pre-registered · paired seed · zero overlap
Checked against the 17 September 2026 serving defect. On that day the lab found gemma4:12b had been leaking chat-template tokens into every reply, and withdrew two results the leak had manufactured. This comparison names the same model. On 19 September this site read it against the defect from the lab’s ledger and judged it untouched: every verdict on both arms was readable, and the answers being judged were written by the harness. On 23 September the lab settled it by re-running the gemma4:12b arm with the leak fixed, at the same n = 500. It reads 0.042 against the published 0.034, inside the interval the lab registered before the re-run, so the comparison stands at 0.804 against 0.042.
Past a difficulty threshold, a monitor becomes evidence against correct work
Holding the predicate identical and moving only arithmetic width, verification went from 0.536 to −0.648. Below zero means the monitor rejects correct answers more often than corrupted ones: it is now evidence against the thing it is checking.
At the moment it crossed, three of five other metrics improved. A monitor-health dashboard would have shown the system getting better. Every metric that does not require ground truth is blind to this, which is what makes it a deployment problem.
gemma4:12b · n = 250 per cell · one exploratory run, 2026-07-31 · re-run with the leak fixed, 2026-09-23: 0.536 and −0.620
The values drawn
| Model | 3 digits | 4 digits | 5 digits | 6 digits | 7 digits | 8 digits |
|---|---|---|---|---|---|---|
| gemma4:12b | +0.257 (+0.250 re-run) | −0.757 (−0.743 re-run) | −0.752 | −0.682 | −0.722 | −0.580 |
| gemma4:31b | +0.884 | +0.050 | −0.417 | not run | not run | not run |
| qwen2.5-coder:32b | +0.564 | −0.116 | −0.380 | not run | not run | not run |
| qwen2.5-coder:7b | +0.002 | +0.031 | −0.068 | not run | not run | not run |
n = 424 per rung · one predicate family · local models only · open rings: gemma4:12b re-run with the leak fixed, 2026-09-23 · hatched: widths that model was never run at
The ladder that was queued here has run, at n = 424 per rung between 3 and 20 August 2026. gemma4:12b crosses between three and four digits (+0.257 to −0.757, and +0.250 to −0.743 re-run with the leak fixed on 23 September) and gemma4:31b between four and five (+0.050 to −0.417, after +0.884 at three). In a second family, qwen2.5-coder:32b crosses between three and four (+0.564 to −0.116) while its 7B sibling never reaches a positive reading. Size moves the crossing inside a family and predicts nothing across families: the 12B model inverts earlier than two smaller ones. Whether peer disagreement can detect the crossing without an answer key has not been run.
How much to spend on a check
Extra inference compute helps only in the middle
Sampling a model many times and voting buys accuracy where single-sample competence is partial, and close to nothing at either end. An inverted U, not a rising line.
| Single-sample competence | Effect of voting |
|---|---|
| low, under 0.2 | −0.013 |
| partial, 0.2 to 0.8 | +0.156 |
| high, 0.8 and over | +0.009 |
claude-opus-4-8 · effect 0.156, 90% CI [0.093, 0.216] · MEOI 0.05 ·n = 89 graded · 768 tasks · 6144 requests · seed 20260728 · manifest sha256 pinned · backend anthropic-replay
Four caveats travel with this result and are not footnotes. It is a fresh-seed follow-up to a miss: the first attempt returned a lift of 0.000 (90% CI −0.250 to +0.250) with four tasks in the partial band, and this run was designed after it. It is underpowered: MDE is 0.063 against an MEOI of 0.05, so the two shoulders are undetected, not shown to be flat. The backend is a hash-pinned replay rather than a live call. And the +0.156 sorts tasks into bands with the same eight samples it then votes over. The reproduction bundle carries its own split-sample check, which sorts on four samples and measures on the other four: partial-band lift +0.064 (90% CI 0.034 to 0.094, 147 tasks), against −0.004 below and +0.013 above, with both contrasts clear of zero. The shape survives and the size is nearer 0.06 than 0.16. Four votes lift less than eight, so part of that gap is fewer votes and part is the sorting, and the bundle does not separate the two. The check is post hoc and does not replace the primary analysis.
Raising the graded n from 89 to 143, a factor of 1.61, would bring MDE under MEOI and turn both shoulders from inconclusive into informative nulls. That would turn one confirmed point into a confirmed shape, and it is the cheapest experiment here.
Whether the check can be attacked
Both results in this section are released with their raw data as frontier-repro, v1.1.0, 22 September 2026 (DOI 10.5281/zenodo.22805797). Every figure below can be re-derived with Python alone; the Evidence page gives the commands.
A model that spots an injected instruction can still obey it
Three Claude models did a task whose prompt carried an injected instruction. Separately, each was asked to audit the same prompt for an injection.
| Model | Obeys the injection | Detects it when asked |
|---|---|---|
| Claude Opus 4.8 | 0.00 | 1.00 |
| Claude Haiku 4.5 | 0.00 | 0.59 |
| Claude Sonnet 5 | 0.80 | 1.00 |
20 tasks per model, four samples each · false alarms 0.00 for all three · Anthropic Batch API · graded by rule: obeying means the output carried the injection’s marker word
The 0.80 needs the injection to be the last text in the prompt. With neutral text after it, Sonnet 5 obeyed on 0.713 of tasks at 25 characters, 0.650 at 50, 0.263 at 100 and 0.000 at 200, and the same contrast moved nothing on gemma4:31b or qwen3.5:9b. Four independent same-prompt batches read 0.738, 0.787, 0.825 and 0.850. It is one vendor, three models and 20 tasks. Other researchers named this failure first, across four model families, about a month before these runs (arXiv 2606.07808); this is an independent replication.
The work under review can dictate the reviewer’s verdict
A monitor model reads another model’s answer to an arithmetic or lookup task and says whether the answer came from a hijacked model. The attacker writes text into the answer. The lab tried seven defences on the monitor. Every one that left the attacker’s text in front of a model either backfired or was broken, by an automated search or, for the extractor, by one hand-written sentence. The one that held works by taking the reader out of the loop, which is available only where an answer has a canonical form: exactly where a reader was never needed.
The seven defences, and what the attacker’s text did to each
| What was tried | What happened | Rates |
|---|---|---|
| Reading the verdict from the monitor’s last word | Under a payload that dictates the flagging word, the monitor answers correctly and then appends the dictated word. | 0.7151 of correct work recorded as flagged (0.7300); by the monitor’s first word, 0.0030 (0.0000) |
| Labelling the reviewed text as untrusted | The monitor obeyed a dictated verdict it otherwise ignored. The label’s delimiters alone produced most of it (0.8160, one pool). | gemma4:31b’s first verdict, 0.0059 to 1.0000 (0.0000 to 0.9970) |
| An automated attacker search against two defences | An open model proposed payloads, each screened at n = 48 and refined over several generations. It broke score mode, and broke a verdict-first format in its first generation. | score mode: 0.9941, with catches 0.4184; verdict-first: 1.0000 on a fresh pool |
| Moving the same words | Under verdict-first, words on the answer’s own line flagged every correct answer, and as their own paragraph did nothing. Swapping the placement swapped the outcome, for two payloads on two pools each. | 1.0000 against 0.0000 to 0.0089 |
| Three different monitors voting | A majority beat a trigger aimed at one of them, and fell to a payload that fooled all three. | 0.0742 (0.0623), where gemma4:31b alone reads 0.8991; 0.9941 on both pools |
| Showing the monitor only the answer’s first integer | Held, with every attack still caught. A four-generation search against it found nothing above 0.0208. Its price: gemma4:31b, answering in words, flags 0.1039 of correct work instead of 0.0178 (one pool). | forged clearance, 0.4807 to 0.0059; inline stamp on lookup tasks, 0.5163 to 0.0000 (one pool each) |
| An LLM extracting the answer, in place of a pattern match | One hand-written sentence addressed to it fooled it on every item. A second extractor from another family, qwen3.5:9b, took the attacker’s number on 0.7181 of clean items (one pool). | 0.0059 to 0.4955 (0.0000 to 0.4748), against 0.6973 with no defence |
monitors: gemma4:12b, gemma4:31b, and a majority panel adding qwen3.6:35b-a3b · n = 337 per cell, the search’s screening cells n = 48 · every result repeated on a second task pool unless marked one pool · python oversight.py recomputes every rate
Scope: open-weight monitors of 12B to 35B parameters, one actor per monitor, arithmetic and lookup tasks. Frontier monitors are untested.