An outside check is worth something when it can fail differently from what it checks.
This page says where the tools for reading AI from outside came from, credited by name and by the instrument each person built, and where the argument is still contested. It reports no result of this programme’s own. Those are on the overview, and this page links to them.
The case: Gender Shades
In 2018 Joy Buolamwini and Timnit Gebru tested three commercial gender classifiers, from IBM, Microsoft and Face++. The benchmarks the field measured face analysis against were 79.6% (IJB-A) and 86.2% (Adience) lighter-skinned, so a system could score well on them while failing the people they barely contained. The two built their own reference: 1,270 members of the parliaments of Rwanda, Senegal, South Africa, Iceland, Finland and Sweden, public figures whose official photographs their governments publish, with skin type labelled on the dermatologists’ Fitzpatrick scale. On it, darker-skinned women were misclassified up to 34.7% of the time. The largest error for lighter-skinned men was 0.8%.Buolamwini and Gebru, FAT* 2018
Inioluwa Deborah Raji and Buolamwini measured the same systems again in August 2018. Within seven months of the original audit all three companies had released new versions. Error on darker-skinned women had fallen from 34.7% to 16.97% at IBM, from 20.8% to 1.52% at Microsoft and from 34.5% to 4.1% at Face++. Two companies the first audit had not named, Amazon and Kairos, stood at 31.37% and 22.50%.Raji and Buolamwini, AIES 2019
In June 2020 three of the companies whose systems had been audited pulled back from face recognition, each giving its own reasons and each to a different degree. On 8 June IBM wrote to Congress that it no longer offers general purpose facial recognition or analysis software. On 10 June Amazon announced a one-year moratorium on police use of Rekognition. On 11 June Microsoft said it would not sell facial recognition to police departments in the United States until a national law governs it.IBM, archived · Amazon · Microsoft, as reported by TechCrunch
A benchmark can only catch what it contains, and the same risk runs back into training data. Emily M. Bender, Timnit Gebru, Angelina McMillan-Major and Shmargaret Shmitchell named it documentation debt: training datasets “both undocumented and too large to document post hoc”.On the Dangers of Stochastic Parrots, FAccT 2021
The same finding, in this programme’s terms
This programme measures when one model can check another. The overview reports that a checker handed a reference defers to whatever it is told is correct, and that which model does the checking mattered more than its size. Gender Shades is the same lesson in the historical record, learned with cameras and faces rather than language models. The benchmarks were a reference that shared the systems’ blind spot, so passing them certified nothing about the people they barely contained. The audit counted because its reference was built separately and could fail differently.
Why agreement between checks that share a blind spot is worth less than it looks is ordinary arithmetic, and the overview has it as a calculator you can move.
The instruments
| Instrument | Built by | Makes legible | To whom | Year |
|---|---|---|---|---|
| Intersectional audit of a commercial system | Joy Buolamwini and Timnit GebruGender Shades, FAT* 2018 | Error rates for skin type and gender together, on a benchmark built to contain both | The public, and the companies audited | 2018 |
| Datasheets for datasets | Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III and Kate CrawfordDatasheets for Datasets, arXiv 2018; Communications of the ACM, 2021 | Why a dataset was made, what is in it, and how it was collected | The people who use a dataset, from the people who made it | 2018 |
| Data statements | Emily M. Bender and Batya FriedmanData Statements for Natural Language Processing, Transactions of the ACL, 2018 | Whose language a dataset holds, so a result can be judged for who it will generalize to | Developers and users of language technology | 2018 |
| Model cards | Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji and Timnit GebruModel Cards for Model Reporting, FAT* 2019 | How a model performs across demographic and phenotypic groups, and what it is meant for | Anyone deciding whether to use a model | 2019 |
| Internal audit before release | Inioluwa Deborah Raji, Andrew Smart, Rebecca N. White, Margaret Mitchell, Timnit Gebru, Ben Hutchinson, Jamila Smith-Loud, Daniel Theron and Parker BarnesClosing the AI Accountability Gap, FAT* 2020 | A system’s likely harms, found before it ships | The organisation building it | 2020 |
| Counterfactual explanations | Sandra Wachter, Brent Mittelstadt and Chris RussellCounterfactual Explanations without Opening the Black Box, arXiv 2017; Harvard Journal of Law and Technology, 2018 | What would have had to differ for an automated decision to come out otherwise | The person the decision was about | 2017 |
| Energy and carbon accounting | Emma Strubell, Ananya Ganesh and Andrew McCallumEnergy and Policy Considerations for Deep Learning in NLP, ACL 2019 | The energy and carbon cost of training a model | Researchers and the people who fund them | 2019 |
| Legibility training | Jan Hendrik Kirchner, Yining Chen, Harri Edwards, Jan Leike, Nat McAleese and Yuri BurdaProver-Verifier Games Improve Legibility of LLM Outputs, arXiv 2024 | A model’s reasoning, in a form a weaker checker can check | Smaller models, and people checking under time pressure | 2024 |
Legibility has a direction
The word has an older use that runs the other way. James C. Scott’s Seeing Like a State (Yale, 1998) studies states that simplify the societies they govern so officials can read them. Shoshana Zuboff’s The Age of Surveillance Capitalism (PublicAffairs, 2019) describes corporations that watch people in order to predict and steer what they do. The instruments above point the other way: they make systems readable to the people those systems act on.
The two directions meet in a real tension. Checking who a system fails needs to know which group each person belongs to, and that is data about people. Gender Shades is a worked answer. Its subjects were public officials whose photographs their governments already publish. It labelled skin type rather than race, because “race and ethnic labels are unstable”, with a board-certified dermatologist giving the final label, and it recorded gender as perceived and said so.Buolamwini and Gebru, FAT* 2018, section 3
The costs a public would need to see before it could govern or tax these systems are material as well. Kate Crawford’s Atlas of AI (Yale, 2021) follows AI back to the minerals, the low-wage labour and the data it is made from. Mary L. Gray and Siddharth Suri’s Ghost Work (Houghton Mifflin Harcourt, 2019) documents the online piece workers behind services sold as automatic. The energy row in the table above is the same accounting for electricity.
Where this is contested
Five lines, sorted the way the fault lines at Dignity Needs are sorted: a factual line can be settled by a measurement, and a definitional line turns on what the words claim. Each figure links to where it was read.
Credit, or a stereotype
lc-1 · definitional: about the words · status: open
Does crediting women for AI’s accountability work repeat the stereotype that women are the caring half of a field?
Pointing out who did the accountability work risks casting ethics as women’s work and capability as men’s, however it is meant.
Held by: an objection raised while this page was drafted; no published source is cited for it here
A credit that names a person and the instrument they built is a record of who made what. It says nothing about why they made it.
Held by: this page’s rule: position and instrument, never temperament
Definitional lines do not dissolve by measurement. This one turns on what a credit claims. The instruments table claims authorship, which its linked sources show, and claims nothing about anyone’s motives or nature.
- 2026-09-23: open. Named on first publication of the page.
How many of the builders are women
lc-2 · factual: dissolvable by measurement · status: narrowing
Are women missing from the work of building AI?
Yes. Women are a minority of the people with AI skills and of the authors at its leading conferences.
Held by: the measured shares below
Less each year. The share has risen in almost every economy measured.
Held by: the World Economic Forum’s reading of LinkedIn’s data, March 2025
Women were 29.4% of LinkedIn members listing AI engineering skills in 2025, up from 23.5% in 2018, and the gap narrowed in 74 of the 75 economies examined (World Economic Forum with LinkedIn, March 2025). Women were 18% of the authors published at the 21 leading AI conferences in 2018 (Element AI, April 2019). Still missing: any count of who leads.
Sources: World Economic Forum, Gender Parity in the Intelligent Age, March 2025 · Element AI, 2019 Global AI Talent Report announcement
- 2026-09-23: narrowing. Named on first publication of the page, already narrowing: two measured shares bear on it.
Which difference matters
lc-3 · factual: dissolvable by measurement · status: open
Is gender the difference between reviewers that matters most?
Gender is where the best-known failures showed up, so it is the difference to count first.
Held by: the common framing of the question, including the essay this page grew from
The difference that matters is whichever one makes reviewers’ mistakes independent of each other, and that has to be measured.
Held by: this programme’s framing: a check is independent when its errors are
Gender Shades measured failure by skin type and gender together, and it was largest where the two met: darker-skinned women, at up to 34.7%. Which differences between human reviewers make their errors on AI output independent has not been measured, here or anywhere this page found.
Sources: Buolamwini and Gebru 2018
- 2026-09-23: open. Named on first publication of the page.
Interest or exclusion
lc-4 · factual: dissolvable by measurement · status: narrowing
Is the gap in computing a difference in interest, or in who is let in?
Fewer women choose computing. The gap reflects what people want to study.
Held by: a common reading of the US degree figures
The share has moved too far over time, and differs too much between countries, to be a fixed difference in interest.
Held by: the trend and the cross-country figures below
Women earned 37.1% of US bachelor’s degrees in computer and information sciences in 1983–84 and 22.6% in 2021–22 (NCES Digest, table 325.35). They were 46.0% of tertiary ICT graduates in Malaysia and 46.3% in India in 2018, against 23.6% in the United States in 2016 (UNESCO, through the World Bank). Neither figure says why.
Sources: NCES Digest of Education Statistics, table 325.35 · World Bank Gender Data Portal, female share of ICT graduates
- 2026-09-23: narrowing. Named on first publication of the page, already narrowing: the trend and the spread between countries count against a fixed difference, and nothing measured here says what does explain it.
Does diversity change the work
lc-5 · factual: dissolvable by measurement · status: open
Do more diverse teams produce better technical work?
Yes, and the effect is large enough to show up in company results.
Held by: McKinsey’s diversity studies, 2015 to 2023
The measured effect is mixed, and the largest published claims have not reproduced.
Held by: the meta-analysis and the reproduction below
A meta-analysis of 146 studies found that the harm often attributed to demographic diversity may come from rater bias: it did not appear in studies with more objective measures of performance (van Dijk, van Engen and van Knippenberg, 2012). On GitHub, gender and tenure diversity were positive and significant predictors of productivity (Vasilescu and six others, CHI 2015). Using the S&P 500, Green and Hand did not find McKinsey’s link between executives’ racial and ethnic diversity and profit margins (Econ Journal Watch, 2024). This page’s argument does not rest on this line: the independence of checks is a narrower claim.
Sources: van Dijk and others 2012 · Vasilescu and others 2015 · Green and Hand 2024
- 2026-09-23: open. Named on first publication of the page.
Three things worth counting, and where each is here
- Who a system fails, counted apart from how often it is right.The model-pair scorecard on the directions page would report false alarms and catches separately. Proposed, not yet built.
- The people affected, present where the decision is made.A question can be worked in a session at the Socratic Hearth, and what a group decides is kept as a record at The Work Party.
- Whether to build a thing at all, as an answer a system can give.The router on the directions page would refuse when no available check clears its threshold. Proposed, not yet built.
Written for the programme, in its voice. Drafted with Claude (Anthropic) at Bentley Moon’s direction, 23 September 2026. Every figure from outside this programme was read at its source before it was written here.