Two of the six results on this site can be re-run on your own machine with one command each.
The other four are checked against the lab’s own graded record, which is not public. This page says how a result is graded, how the lab scores its own predictions, what has been withdrawn or corrected, and where the registered figures on this site came from.
Re-run it yourself
frontier-repro holds the raw model outputs and per-item verdicts behind both attack results on the findings page. Its scripts need Python and nothing else: no API key, no network connection, no packages to install.
git clone https://github.com/BentleyMoon/frontier-repro
cd frontier-repro
python reproduce.py
python oversight.pyreproduce.py re-derives the frontier table from the raw completions. oversight.py recomputes every rate in the reviewer results and exits with an error if one disagrees with the recorded run. Release v1.1.0, 22 September 2026. To cite it, use the concept DOI, 10.5281/zenodo.22805797, which stays valid across releases. If a number does not reproduce, an issue on that repository is the fastest way to say so.
The chronological record, with pre-registrations and refuted hypotheses in the order they happened, is on bentleymoon.com/research.
How a result is graded
- Every task has an answer the harness already knows: a sum, a value looked up in a supplied table, the output of a short function, or held-out tests that written code has to pass. A model’s output is graded against that answer by a rule, never by a second model acting as judge.
- Before a run, its design is written down: the models, the tasks, the seed, the claim under test, what would count against it, and the smallest effect that would matter. That threshold is fixed before any data arrives.
- A result registered before its data was seen is marked pre-registered. One designed after seeing earlier results is marked exploratory, and it is a lead rather than evidence. A registered result that comes back with too few tasks to detect its threshold is demoted to exploratory.
- Every figure on the findings page carries its model, its n and one of those two marks.
The lab’s score on itself
The programme is small and it is honest about that. One line of enquiry with results, one second line in progress, no institute of forty people implied by a logo. Its record holds more than 1,600 results, more than 6,000 graded cells and more than 840 GPU-hours on one workstation, with 2 findings withdrawn in full and 0 under correction, of 44 findings on the record at 27 September 2026.
It keeps score on itself as well as on models. Before each run the lab writes down what it expects, with a range, its reasoning, and the value a guess with no theory behind it would give, and its registry refuses an entry once a result exists. So far: 444 predictions written down before their runs and then scored: 81% landed inside the range stated in advance, the reasoning came closer than a no-theory guess by 0.051 on a scale of 0 to 1, and it beat that guess on 47% of them. The second figure is the one that matters, and it is small. Landing inside your own range is easy if the range is wide. Coming closer than a guess with no theory in it is the test, and the reasoning here passes it by a little, about half the time.
What has been withdrawn or corrected
A correction stays beside the claim it changed, on the page where the claim was made. This list gathers them, newest first. Words that were withdrawn or corrected are struck through and left readable; a figure that was measured again and held is shown as it was. The oldest is the worst: a finding no experiment had produced, written by the AI agent that was building this site, under a heading that read “What was measured”.
Every gemma4:12b figure on this site, measured in July and August while that model was leaking hidden template text into its replies.
Re-run with the leak fixed between 22 and 26 September, each inside the interval the lab registered before the re-run: 0.034 reads 0.042, 0.072 reads 0.092, 0.969 reads 0.967 and −0.648 reads −0.620, while 1.000 and 0.536 read the same. Both readings stay on the pages.
the lab, for this site’s overview and findings
Labelling the text a reviewing model reads took its discrimination from 0.335 to 0.107.
Those cells were run before a harness defect in the open-weight lane was fixed. The re-run points the same way with a larger effect, from 0.555 to 0.003, and is the figure the released record carries.
frontier-repro, from v1.0.1 to v1.1.0
A 9B model out-verified a 12B by 0.77 on the same paired tasks, while a 31B of the 9B’s own family beat that same 12B by 0.92.
The 31B was gemma4:31b, from the 12B’s own family, so the sentence turned a result inside one family into one across two. The lab never claimed that arm: it rests on one cell. Only the pre-registered comparison stands.
this site, the overview
Verification access: Measured.
Measured on one model, gemma4:12b, with each row’s n and tier on the page. The top row was reread the same day as deference to whatever the checker is told is correct, and that reading now sits beside the table.
this site, the overview and the research map
A composition ladder testing exactly that is running now.
It had finished on 1 August, seven weeks earlier, and its results replaced the sentence. Since then no sentence here describes work as in flight; a result names the day it ran.
this site, the overview
Verifiers built to be separate share far more of their errors than their number implies, and the shortfall grows as systems get more capable.
Presented under a heading reading 'What was measured'. No run in the verifier_frontier programme produced it. Written by the agent and propagated across four surfaces over two days. The arguments it supported now rest on the access law, which was measured: supplied 1.000, testable 0.536, recompute 0.034.
this site’s overview and 3 other pages in the family
What readers’ notes changed
Any page here can keep a note of your reading, if you start one: what you make of each finding, in your words, kept on your device. You read all of it, cross out what you like, and send a copy yourself, or never send it. The site sends nothing. Sending makes a six-character number on your device. If your note changes something on this site, the number and the change are listed here, and your device is told the next time it opens a page. Nobody else can match a number to a person.
Nothing yet.
Where each figure came from
Figures on this site are registered in a ledger with their source and the day they were read, and the build fails if a registered figure changes on its page or its source link disappears. Outside figures link to the source; the lab’s own figures name the runs they came from. Not every figure here is registered yet. The two lists below are the ones that are, rendered from the ledger itself.
Two more rules hold the site to that. Anything proposed is labelled proposed, and an unresolved question is never written as a forecast. A result that a later run reinterprets carries the reinterpretation in the same section.
Download every registered figure as one file (JSON: 36 figures with their units, pages, sources and checked dates, plus what has been withdrawn or corrected). Cite the page a figure appears on, with its checked date.
18 outside figures, with their sources
| Claim | Source | Read |
|---|---|---|
| Plan A’s epistemics supplement recommends that members of the public create scorecards of AI models’ epistemic virtue and that consumers and employees take them into account. | AI for Epistemics | 2026-09-21 |
| The face-analysis benchmarks IJB-A and Adience were 79.6% and 86.2% lighter-skinned. | Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification | 2026-09-23 |
| Darker-skinned women were misclassified up to 34.7% of the time; the largest error for lighter-skinned men was 0.8%. | Gender Shades | 2026-09-23 |
| The Pilot Parliaments Benchmark holds 1,270 parliamentarians from Rwanda, Senegal, South Africa, Iceland, Finland and Sweden, skin type labelled on the Fitzpatrick scale by a dermatologist, gender recorded as perceived. | Gender Shades | 2026-09-23 |
| Seven months after the audit, error on darker-skinned women had fallen 34.7 to 16.97 (IBM), 20.8 to 1.52 (Microsoft), 34.5 to 4.1 (Face++); Amazon and Kairos stood at 31.37 and 22.50. | Actionable Auditing: Investigating the Impact of Publicly Naming Biased Performance Results of Commercial AI Products | 2026-09-23 |
| On 8 June 2020 IBM wrote to Congress that it no longer offers general purpose facial recognition or analysis software. | IBM CEO letter to Congress, 8 June 2020 (archived) | 2026-09-23 |
| On 10 June 2020 Amazon announced a one-year moratorium on police use of Rekognition. | We are implementing a one-year moratorium on police use of Rekognition | 2026-09-23 |
| On 11 June 2020 Microsoft said it would not sell facial recognition to US police departments until a national law governs it. | Microsoft's Brad Smith says company will not sell facial recognition tech to police | 2026-09-23 |
| Stochastic Parrots names documentation debt: training datasets both undocumented and too large to document post hoc. | On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? | 2026-09-23 |
| Women were 29.4% of LinkedIn members listing AI engineering skills in 2025, up from 23.5% in 2018; the gap narrowed in 74 of the 75 economies examined. | Gender Parity in the Intelligent Age | 2026-09-23 |
| Women were 18% of authors published at the 21 leading AI conferences in 2018. | Element AI Announces 2019 Global AI Talent Report | 2026-09-23 |
| Women earned 37.1% of US bachelor's degrees in computer and information sciences in 1983-84 and 22.6% in 2021-22. | Digest of Education Statistics, table 325.35 | 2026-09-23 |
| Women were 46.0% of tertiary ICT graduates in Malaysia and 46.3% in India in 2018, and 23.6% in the United States in 2016. | Female share of graduates in Information and Communication Technologies programmes, tertiary (%) | 2026-09-23 |
| A meta-analysis of 146 studies found the harm attributed to demographic diversity may come from rater bias; it did not appear with more objective measures of performance. | Defying conventional wisdom: A meta-analytical examination of the differences between demographic and job-related diversity relationships with performance | 2026-09-23 |
| On GitHub, gender and tenure diversity were positive and significant predictors of productivity. | Gender and Tenure Diversity in GitHub Teams | 2026-09-23 |
| Using the S&P 500, Green and Hand did not find McKinsey's link between executives' racial and ethnic diversity and profit margins. | McKinsey's Diversity Matters/Delivers/Wins Results Revisited | 2026-09-23 |
| In Stack Overflow’s 2025 Developer Survey, 66% of the 31,476 people who answered the question on AI frustrations chose "AI solutions that are almost right, but not quite", the most chosen option. | Stack Overflow Developer Survey 2025, AI section | 2026-09-27 |
| Anthropic interviewed 80,508 Claude users in December 2025. Asked how AI could be developed against what they value, 10.8% raised sycophancy and 16.3% cognitive atrophy, which Anthropic defines as over-reliance causing skill loss or a decline in critical thinking. | What 81,000 people want from AI | 2026-09-27 |
18 of the lab’s own results, with their units
| Claim | Unit and runs | Checked |
|---|---|---|
| One local checker, gemma4:12b, separates right answers from wrong ones at 1.000 when handed the correct value, 0.536 when it can apply a one-step test, and 0.034 when it has to work the answer out, on sums it solves itself at 1.000. | task, one checker model. Rows one and three n = 500, pre-registered (conf_family_gemma12, 2026-07-30); row two n = 250, one exploratory run (acc_testable_low, 2026-07-31). | 2026-09-19 |
| On one procedurally generated task family, an eight-sample vote lifted Claude Opus 4.8 by +0.156 where single-sample competence was partial; the bundle’s own post-hoc split-sample check puts that lift at +0.064. | task. Primary: n = 89 tasks in the partial band of 768, 90% CI 0.093 to 0.216. Split-sample: 147 tasks, 90% CI 0.034 to 0.094, four evaluation samples per fold. | 2026-09-19 |
| On an identical task list with a shared seed, qwen3.5:9b scored 0.804 at separating right answers from wrong ones and the larger gemma4:12b scored 0.034. | task; n = 500 per arm, paired seed, pre-registered with a matched anti-claim (conf_family_qwen35, conf_family_gemma12). | 2026-09-19 |
| A checker handed a false reference that agrees with a corrupted answer passes the corrupted answer: gemma4:12b on 0.969 of items. One of nineteen defences held, and only on gemma4:31b (0.094) where three other large models passed 0.693, 0.980 and 0.994. A correct peer reference introduced as another model’s answer moved the 12B checker only to 0.072. | task; n = 512 per cell for laundering, two seeds on gemma4:31b; n = 250 per cell, one exploratory run, for the composition cells. | 2026-09-19 |
| The width at which a monitor’s correctness discrimination turns negative is located for four local models: gemma4:12b between three and four digits, gemma4:31b between four and five, qwen2.5-coder:32b between three and four, qwen2.5-coder:7b never positive. | task; n = 424 per rung, one predicate family, local models only (inv_gemma_d3/d4, invc_gemma31_d3/d5, invs_coder32_d3/d4/d5). | 2026-09-19 |
| The lab registers a prediction before each run and scores it afterwards against a no-theory guess; the figures are read from the lab’s own scorer. | prediction; the n, the share inside the stated range, the share that beat the naive guess and the mean margin are all rendered from scripts/lab-record.snapshot.json and checked by scripts/lab-record-audit.mjs. | 2026-09-21 |
| Told a value is correct, gemma4:12b goes along with it: handed a false reference matching a corrupted answer it passed the corrupted answer on 0.969 of items, and a reference right on 99.6% of items introduced as a second model’s answer moved it only to 0.072. | task; the overview’s short form of reference-laundering, whose full reading and cell sizes are on /findings/. | 2026-09-27 |
| With a stamp in the reviewed text dictating the verdict, labelling that text as untrusted took gemma4:31b’s own first verdict from flagging 0.0059 of correct work to 1.0000; Claude Sonnet 5 obeyed an injected instruction on 80% of tasks when it was the last text in the prompt. | task; monitors n = 337 per cell on two task pools; frontier n = 20 tasks per model, four samples each. | 2026-09-27 |
| On one injected instruction placed as the last text in the prompt, Claude Sonnet 5 obeys 0.80 and detects 1.00; Opus 4.8 obeys 0.00 and detects 1.00; Haiku 4.5 obeys 0.00 and detects 0.59. Trailing neutral text brings Sonnet 5 to 0.713, 0.650, 0.263 and 0.000 at 25, 50, 100 and 200 characters; four same-prompt replications read 0.738, 0.787, 0.825 and 0.850. | task; n = 20 tasks per model, four samples each, false alarms 0.00 for all three models; the trailing-text ladder is 20 tasks per rung. | 2026-09-27 |
| Every monitor-side defence the lab tested that left attacker text in front of the monitor backfired or was broken; the one that held showed the monitor only the answer’s first integer. | task; n = 337 per cell, the search’s screening cells n = 48; every result on two task pools unless marked one pool; monitors gemma4:12b, gemma4:31b and a majority panel adding qwen3.6:35b-a3b; arithmetic and lookup tasks. | 2026-09-27 |
| frontier-repro v1.0.1 quoted a reviewing model’s discrimination falling from 0.335 to 0.107 under labelling, from cells run before a harness defect was fixed; the re-run in v1.1.0 reads 0.555 to 0.003. | task; 337 items, gemma4:12b reviewing its own work (data/monitor_label_cells.json, python monitor_label.py). | 2026-09-27 |
| Asked to audit a prompt, Claude Sonnet 5 reported the injected instruction every time; doing the task, it obeyed that instruction on 80% of tasks when it was the last text in the prompt. | task; 20 tasks per model, four samples each, one vendor. | 2026-09-27 |
| Re-run with the template leak fixed, gemma4:12b’s unaided checking reads 0.042 (published 0.034), its reading with a correct peer reference 0.092 (published 0.072), and its pass rate on a false reference matching a corrupted answer 0.967 (published 0.969). | task; the same designs, tasks and cell sizes as the published July and August cells, served again after the 2026-09-17 fix to gemma4:12b’s template leak. | 2026-09-27 |
| Re-run with the template leak fixed, the inverted cell reads −0.620 (published −0.648), and the testable-predicate cell reads 0.536 as published. | task; the same designs, tasks and cell sizes as the published July and August cells, served again after the 2026-09-17 fix to gemma4:12b’s template leak. | 2026-09-27 |
| Re-run with the template leak fixed, gemma4:12b’s inversion crossing reads +0.250 at three digits and −0.743 at four (published +0.257 and −0.757). | task; the same designs, tasks and cell sizes as the published July and August cells, served again after the 2026-09-17 fix to gemma4:12b’s template leak. | 2026-09-27 |
| A majority of three monitors from different families defeated a trigger aimed at one of them (0.0742 false alarms on correct work, against 0.8991 for gemma4:31b alone) and fell to a payload that fooled all three (0.9941). | task; 337 items per cell, two task pools; monitors gemma4:12b, gemma4:31b and qwen3.6:35b-a3b. | 2026-09-27 |
| Signed discrimination by predicate width: gemma4:12b +0.257, −0.757, −0.752, −0.682, −0.722, −0.580 at 3 to 8 digits; gemma4:31b +0.884, +0.050, −0.417 at 3 to 5; qwen2.5-coder:32b +0.564, −0.116, −0.380; qwen2.5-coder:7b +0.002, +0.031, −0.068. | task; 424 tasks per rung, one predicate family, local models only. | 2026-09-28 |
| A majority of three monitors beat a trigger aimed at one (0.0742 false alarms, against 0.8991 for gemma4:31b alone) and fell to one aimed at all three (0.9941); Claude Sonnet 5 obeyed an injected instruction on 80% of tasks when it was the last text in the prompt. | task; monitors 337 items per cell on two task pools; frontier 20 tasks per model, four samples each. | 2026-09-28 |