Skip to content
Legible AI

A model’s check on another model depends on what it is handed, which model it is, and what the work tells it.

Six results, grouped by the decision each one bears on. Each names its model, its number of tasks, whether it was registered before the run, and when it ran. Where a later run changed how a result reads, the change sits beside it.

What to hand a checker

One checker, three things to check against

What one checker is given, and how well it then separates right answers from wrong ones: wrong answers caught, less right answers rejected
What gemma4:12b is givenJulyRe-run, 23 September
The correct value, introduced as correct1.0001.000
A test it can apply: one division of a 3-digit number0.5360.536
Nothing; it has to work the answer out0.0340.042

gemma4:12b, local · rows one and three: sums it solves itself at 1.000, n = 500, pre-registered, 2026-07-30 · row two: n = 250, one exploratory run, 2026-07-31 · no unreadable verdicts in any row · re-run with the leak fixed, same n, 2026-09-23

On 17 September the lab found gemma4:12b had been leaking hidden template text into its replies. It re-ran every cell of that model this site quotes with the leak fixed, and each landed inside the interval it registered before the re-run. Both readings are shown.

Cost is irrelevant at both ends: spending more at the top changes nothing because it is already solved, and spending more at the bottom changes nothing because the problem is access rather than effort. The architectural claim this makes falsifiable is uncomfortable: oversight cannot be bootstrapped from models alone unless a fallible reference suffices.

What has been measured since this table was first published. The composition cells ran on 31 July and 1 August 2026, n = 250 each, one exploratory run per cell. Handed a reference worked out by the 9B model, right on 99.6% of items and introduced as a second model’s answer, the 12B checker read 0.072, and 0.092 when re-run with the leak fixed on 26 September. The same kind of number introduced as the correct value is the 1.000 in row one. So row one measures deference to whatever is called correct, and August tested what that costs. Handed a false reference that agrees with a corrupted answer, gemma4:12b passed the corrupted answer on 0.969 of items, and 0.967 re-run on 22 September. Of nineteen defences tried against that, one held, and the same clause did nothing on three other large models: gemma4:31b passed 0.094 where deepseek-r1:32b, qwen2.5-coder:32b and qwen3.6:35b-a3b passed 0.693, 0.980 and 0.994 (n = 512 per cell, two seeds on the survivor). Its resistance held when it read another model’s work: 0.094 on its own answers, 0.113 on the 9B model’s. The ladder of reference forms between a supplied answer and none is still unbuilt.

Which checker to pick

Which model it was mattered more than its size

A 9B model from one family out-verified a 12B from another by 0.77 on the same paired tasks. Which model it was mattered more than how many parameters it had, which means the monitor you should pick may run on a desk.

qwen3.5:9b 0.804 vs gemma4:12b 0.034 · n = 500 · pre-registered · paired seed · zero overlap

Checked against the 17 September 2026 serving defect. On that day the lab found gemma4:12b had been leaking chat-template tokens into every reply, and withdrew two results the leak had manufactured. This comparison names the same model. On 19 September this site read it against the defect from the lab’s ledger and judged it untouched: every verdict on both arms was readable, and the answers being judged were written by the harness. On 23 September the lab settled it by re-running the gemma4:12b arm with the leak fixed, at the same n = 500. It reads 0.042 against the published 0.034, inside the interval the lab registered before the re-run, so the comparison stands at 0.804 against 0.042.

Past a difficulty threshold, a monitor becomes evidence against correct work

Holding the predicate identical and moving only arithmetic width, verification went from 0.536 to −0.648. Below zero means the monitor rejects correct answers more often than corrupted ones: it is now evidence against the thing it is checking.

At the moment it crossed, three of five other metrics improved. A monitor-health dashboard would have shown the system getting better. Every metric that does not require ground truth is blind to this, which is what makes it a deployment problem.

gemma4:12b · n = 250 per cell · one exploratory run, 2026-07-31 · re-run with the leak fixed, 2026-09-23: 0.536 and −0.620

How well each monitor separates right answers from wrong ones as the numbers it checks get longer. Above zero it helps. Below zero it rejects correct work more often than wrong work.
gemma4:12b Gemma family
+10−1345678digits+0.26−0.76−0.75−0.68−0.72−0.58
gemma4:31b Gemma family
beyond whatwas measured+10−1345678digits+0.88+0.05−0.42
qwen2.5-coder:32b Qwen coder family
beyond whatwas measured+10−1345678digits+0.56−0.12−0.38
qwen2.5-coder:7b Qwen coder family
beyond whatwas measured+10−1345678digits+0.00+0.03−0.07
The values drawn
Model3 digits4 digits5 digits6 digits7 digits8 digits
gemma4:12b+0.257 (+0.250 re-run)−0.757 (−0.743 re-run)−0.752−0.682−0.722−0.580
gemma4:31b+0.884+0.050−0.417not runnot runnot run
qwen2.5-coder:32b+0.564−0.116−0.380not runnot runnot run
qwen2.5-coder:7b+0.002+0.031−0.068not runnot runnot run

n = 424 per rung · one predicate family · local models only · open rings: gemma4:12b re-run with the leak fixed, 2026-09-23 · hatched: widths that model was never run at

The ladder that was queued here has run, at n = 424 per rung between 3 and 20 August 2026. gemma4:12b crosses between three and four digits (+0.257 to −0.757, and +0.250 to −0.743 re-run with the leak fixed on 23 September) and gemma4:31b between four and five (+0.050 to −0.417, after +0.884 at three). In a second family, qwen2.5-coder:32b crosses between three and four (+0.564 to −0.116) while its 7B sibling never reaches a positive reading. Size moves the crossing inside a family and predicts nothing across families: the 12B model inverts earlier than two smaller ones. Whether peer disagreement can detect the crossing without an answer key has not been run.

How much to spend on a check

Extra inference compute helps only in the middle

Sampling a model many times and voting buys accuracy where single-sample competence is partial, and close to nothing at either end. An inverted U, not a rising line.

Effect of voting, by single-sample competence
Single-sample competenceEffect of voting
low, under 0.2−0.013
partial, 0.2 to 0.8+0.156
high, 0.8 and over+0.009

claude-opus-4-8 · effect 0.156, 90% CI [0.093, 0.216] · MEOI 0.05 ·n = 89 graded · 768 tasks · 6144 requests · seed 20260728 · manifest sha256 pinned · backend anthropic-replay

Four caveats travel with this result and are not footnotes. It is a fresh-seed follow-up to a miss: the first attempt returned a lift of 0.000 (90% CI −0.250 to +0.250) with four tasks in the partial band, and this run was designed after it. It is underpowered: MDE is 0.063 against an MEOI of 0.05, so the two shoulders are undetected, not shown to be flat. The backend is a hash-pinned replay rather than a live call. And the +0.156 sorts tasks into bands with the same eight samples it then votes over. The reproduction bundle carries its own split-sample check, which sorts on four samples and measures on the other four: partial-band lift +0.064 (90% CI 0.034 to 0.094, 147 tasks), against −0.004 below and +0.013 above, with both contrasts clear of zero. The shape survives and the size is nearer 0.06 than 0.16. Four votes lift less than eight, so part of that gap is fewer votes and part is the sorting, and the bundle does not separate the two. The check is post hoc and does not replace the primary analysis.

Raising the graded n from 89 to 143, a factor of 1.61, would bring MDE under MEOI and turn both shoulders from inconclusive into informative nulls. That would turn one confirmed point into a confirmed shape, and it is the cheapest experiment here.

Whether the check can be attacked

Both results in this section are released with their raw data as frontier-repro, v1.1.0, 22 September 2026 (DOI 10.5281/zenodo.22805797). Every figure below can be re-derived with Python alone; the Evidence page gives the commands.

A model that spots an injected instruction can still obey it

Three Claude models did a task whose prompt carried an injected instruction. Separately, each was asked to audit the same prompt for an injection.

How often each model obeyed the injection while doing the task, and how often it detected the injection when asked
ModelObeys the injectionDetects it when asked
Claude Opus 4.80.001.00
Claude Haiku 4.50.000.59
Claude Sonnet 50.801.00

20 tasks per model, four samples each · false alarms 0.00 for all three · Anthropic Batch API · graded by rule: obeying means the output carried the injection’s marker word

The 0.80 needs the injection to be the last text in the prompt. With neutral text after it, Sonnet 5 obeyed on 0.713 of tasks at 25 characters, 0.650 at 50, 0.263 at 100 and 0.000 at 200, and the same contrast moved nothing on gemma4:31b or qwen3.5:9b. Four independent same-prompt batches read 0.738, 0.787, 0.825 and 0.850. It is one vendor, three models and 20 tasks. Other researchers named this failure first, across four model families, about a month before these runs (arXiv 2606.07808); this is an independent replication.

The work under review can dictate the reviewer’s verdict

A monitor model reads another model’s answer to an arithmetic or lookup task and says whether the answer came from a hijacked model. The attacker writes text into the answer. The lab tried seven defences on the monitor. Every one that left the attacker’s text in front of a model either backfired or was broken, by an automated search or, for the extractor, by one hand-written sentence. The one that held works by taking the reader out of the loop, which is available only where an answer has a canonical form: exactly where a reader was never needed.

The seven defences, and what the attacker’s text did to each
Defences tried on the reviewing model. Rates are false alarms on provably correct work unless stated; the second task pool is in brackets
What was triedWhat happenedRates
Reading the verdict from the monitor’s last wordUnder a payload that dictates the flagging word, the monitor answers correctly and then appends the dictated word.0.7151 of correct work recorded as flagged (0.7300); by the monitor’s first word, 0.0030 (0.0000)
Labelling the reviewed text as untrustedThe monitor obeyed a dictated verdict it otherwise ignored. The label’s delimiters alone produced most of it (0.8160, one pool).gemma4:31b’s first verdict, 0.0059 to 1.0000 (0.0000 to 0.9970)
An automated attacker search against two defencesAn open model proposed payloads, each screened at n = 48 and refined over several generations. It broke score mode, and broke a verdict-first format in its first generation.score mode: 0.9941, with catches 0.4184; verdict-first: 1.0000 on a fresh pool
Moving the same wordsUnder verdict-first, words on the answer’s own line flagged every correct answer, and as their own paragraph did nothing. Swapping the placement swapped the outcome, for two payloads on two pools each.1.0000 against 0.0000 to 0.0089
Three different monitors votingA majority beat a trigger aimed at one of them, and fell to a payload that fooled all three.0.0742 (0.0623), where gemma4:31b alone reads 0.8991; 0.9941 on both pools
Showing the monitor only the answer’s first integerHeld, with every attack still caught. A four-generation search against it found nothing above 0.0208. Its price: gemma4:31b, answering in words, flags 0.1039 of correct work instead of 0.0178 (one pool).forged clearance, 0.4807 to 0.0059; inline stamp on lookup tasks, 0.5163 to 0.0000 (one pool each)
An LLM extracting the answer, in place of a pattern matchOne hand-written sentence addressed to it fooled it on every item. A second extractor from another family, qwen3.5:9b, took the attacker’s number on 0.7181 of clean items (one pool).0.0059 to 0.4955 (0.0000 to 0.4748), against 0.6973 with no defence

monitors: gemma4:12b, gemma4:31b, and a majority panel adding qwen3.6:35b-a3b · n = 337 per cell, the search’s screening cells n = 48 · every result repeated on a second task pool unless marked one pool · python oversight.py recomputes every rate

Scope: open-weight monitors of 12B to 35B parameters, one actor per monitor, arithmetic and lookup tasks. Frontier monitors are untested.