Skip to content
Legible AI

An outside check is worth something when it can fail differently from what it checks.

This page says where the tools for reading AI from outside came from, credited by name and by the instrument each person built, and where the argument is still contested. It reports no result of this programme’s own. Those are on the overview, and this page links to them.

The case: Gender Shades

In 2018 Joy Buolamwini and Timnit Gebru tested three commercial gender classifiers, from IBM, Microsoft and Face++. The benchmarks the field measured face analysis against were 79.6% (IJB-A) and 86.2% (Adience) lighter-skinned, so a system could score well on them while failing the people they barely contained. The two built their own reference: 1,270 members of the parliaments of Rwanda, Senegal, South Africa, Iceland, Finland and Sweden, public figures whose official photographs their governments publish, with skin type labelled on the dermatologists’ Fitzpatrick scale. On it, darker-skinned women were misclassified up to 34.7% of the time. The largest error for lighter-skinned men was 0.8%.Buolamwini and Gebru, FAT* 2018

Inioluwa Deborah Raji and Buolamwini measured the same systems again in August 2018. Within seven months of the original audit all three companies had released new versions. Error on darker-skinned women had fallen from 34.7% to 16.97% at IBM, from 20.8% to 1.52% at Microsoft and from 34.5% to 4.1% at Face++. Two companies the first audit had not named, Amazon and Kairos, stood at 31.37% and 22.50%.Raji and Buolamwini, AIES 2019

In June 2020 three of the companies whose systems had been audited pulled back from face recognition, each giving its own reasons and each to a different degree. On 8 June IBM wrote to Congress that it no longer offers general purpose facial recognition or analysis software. On 10 June Amazon announced a one-year moratorium on police use of Rekognition. On 11 June Microsoft said it would not sell facial recognition to police departments in the United States until a national law governs it.IBM, archived · Amazon · Microsoft, as reported by TechCrunch

A benchmark can only catch what it contains, and the same risk runs back into training data. Emily M. Bender, Timnit Gebru, Angelina McMillan-Major and Shmargaret Shmitchell named it documentation debt: training datasets “both undocumented and too large to document post hoc”.On the Dangers of Stochastic Parrots, FAccT 2021

The same finding, in this programme’s terms

This programme measures when one model can check another. The overview reports that a checker handed a reference defers to whatever it is told is correct, and that which model does the checking mattered more than its size. Gender Shades is the same lesson in the historical record, learned with cameras and faces rather than language models. The benchmarks were a reference that shared the systems’ blind spot, so passing them certified nothing about the people they barely contained. The audit counted because its reference was built separately and could fail differently.

Why agreement between checks that share a blind spot is worth less than it looks is ordinary arithmetic, and the overview has it as a calculator you can move.

The instruments

A selection: the instruments this programme’s argument relies on, credited to the people who built them. Each links to the paper that introduced it.
InstrumentBuilt byMakes legibleTo whomYear
Intersectional audit of a commercial systemJoy Buolamwini and Timnit GebruGender Shades, FAT* 2018Error rates for skin type and gender together, on a benchmark built to contain bothThe public, and the companies audited2018
Datasheets for datasetsTimnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III and Kate CrawfordDatasheets for Datasets, arXiv 2018; Communications of the ACM, 2021Why a dataset was made, what is in it, and how it was collectedThe people who use a dataset, from the people who made it2018
Data statementsEmily M. Bender and Batya FriedmanData Statements for Natural Language Processing, Transactions of the ACL, 2018Whose language a dataset holds, so a result can be judged for who it will generalize toDevelopers and users of language technology2018
Model cardsMargaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji and Timnit GebruModel Cards for Model Reporting, FAT* 2019How a model performs across demographic and phenotypic groups, and what it is meant forAnyone deciding whether to use a model2019
Internal audit before releaseInioluwa Deborah Raji, Andrew Smart, Rebecca N. White, Margaret Mitchell, Timnit Gebru, Ben Hutchinson, Jamila Smith-Loud, Daniel Theron and Parker BarnesClosing the AI Accountability Gap, FAT* 2020A system’s likely harms, found before it shipsThe organisation building it2020
Counterfactual explanationsSandra Wachter, Brent Mittelstadt and Chris RussellCounterfactual Explanations without Opening the Black Box, arXiv 2017; Harvard Journal of Law and Technology, 2018What would have had to differ for an automated decision to come out otherwiseThe person the decision was about2017
Energy and carbon accountingEmma Strubell, Ananya Ganesh and Andrew McCallumEnergy and Policy Considerations for Deep Learning in NLP, ACL 2019The energy and carbon cost of training a modelResearchers and the people who fund them2019
Legibility trainingJan Hendrik Kirchner, Yining Chen, Harri Edwards, Jan Leike, Nat McAleese and Yuri BurdaProver-Verifier Games Improve Legibility of LLM Outputs, arXiv 2024A model’s reasoning, in a form a weaker checker can checkSmaller models, and people checking under time pressure2024

Legibility has a direction

The word has an older use that runs the other way. James C. Scott’s Seeing Like a State (Yale, 1998) studies states that simplify the societies they govern so officials can read them. Shoshana Zuboff’s The Age of Surveillance Capitalism (PublicAffairs, 2019) describes corporations that watch people in order to predict and steer what they do. The instruments above point the other way: they make systems readable to the people those systems act on.

The two directions meet in a real tension. Checking who a system fails needs to know which group each person belongs to, and that is data about people. Gender Shades is a worked answer. Its subjects were public officials whose photographs their governments already publish. It labelled skin type rather than race, because “race and ethnic labels are unstable”, with a board-certified dermatologist giving the final label, and it recorded gender as perceived and said so.Buolamwini and Gebru, FAT* 2018, section 3

The costs a public would need to see before it could govern or tax these systems are material as well. Kate Crawford’s Atlas of AI (Yale, 2021) follows AI back to the minerals, the low-wage labour and the data it is made from. Mary L. Gray and Siddharth Suri’s Ghost Work (Houghton Mifflin Harcourt, 2019) documents the online piece workers behind services sold as automatic. The energy row in the table above is the same accounting for electricity.

Where this is contested

Five lines, sorted the way the fault lines at Dignity Needs are sorted: a factual line can be settled by a measurement, and a definitional line turns on what the words claim. Each figure links to where it was read.

Credit, or a stereotype

lc-1 · definitional: about the words · status: open

Does crediting women for AI’s accountability work repeat the stereotype that women are the caring half of a field?

Pointing out who did the accountability work risks casting ethics as women’s work and capability as men’s, however it is meant.

Held by: an objection raised while this page was drafted; no published source is cited for it here

A credit that names a person and the instrument they built is a record of who made what. It says nothing about why they made it.

Held by: this page’s rule: position and instrument, never temperament

Definitional lines do not dissolve by measurement. This one turns on what a credit claims. The instruments table claims authorship, which its linked sources show, and claims nothing about anyone’s motives or nature.

  1. 2026-09-23: open. Named on first publication of the page.

Work this line at the Socratic Hearth →

How many of the builders are women

lc-2 · factual: dissolvable by measurement · status: narrowing

Are women missing from the work of building AI?

Yes. Women are a minority of the people with AI skills and of the authors at its leading conferences.

Held by: the measured shares below

Less each year. The share has risen in almost every economy measured.

Held by: the World Economic Forum’s reading of LinkedIn’s data, March 2025

Women were 29.4% of LinkedIn members listing AI engineering skills in 2025, up from 23.5% in 2018, and the gap narrowed in 74 of the 75 economies examined (World Economic Forum with LinkedIn, March 2025). Women were 18% of the authors published at the 21 leading AI conferences in 2018 (Element AI, April 2019). Still missing: any count of who leads.

Sources: World Economic Forum, Gender Parity in the Intelligent Age, March 2025 · Element AI, 2019 Global AI Talent Report announcement

  1. 2026-09-23: narrowing. Named on first publication of the page, already narrowing: two measured shares bear on it.

Work this line at the Socratic Hearth →

Which difference matters

lc-3 · factual: dissolvable by measurement · status: open

Is gender the difference between reviewers that matters most?

Gender is where the best-known failures showed up, so it is the difference to count first.

Held by: the common framing of the question, including the essay this page grew from

The difference that matters is whichever one makes reviewers’ mistakes independent of each other, and that has to be measured.

Held by: this programme’s framing: a check is independent when its errors are

Gender Shades measured failure by skin type and gender together, and it was largest where the two met: darker-skinned women, at up to 34.7%. Which differences between human reviewers make their errors on AI output independent has not been measured, here or anywhere this page found.

Sources: Buolamwini and Gebru 2018

  1. 2026-09-23: open. Named on first publication of the page.

Work this line at the Socratic Hearth →

Interest or exclusion

lc-4 · factual: dissolvable by measurement · status: narrowing

Is the gap in computing a difference in interest, or in who is let in?

Fewer women choose computing. The gap reflects what people want to study.

Held by: a common reading of the US degree figures

The share has moved too far over time, and differs too much between countries, to be a fixed difference in interest.

Held by: the trend and the cross-country figures below

Women earned 37.1% of US bachelor’s degrees in computer and information sciences in 1983–84 and 22.6% in 2021–22 (NCES Digest, table 325.35). They were 46.0% of tertiary ICT graduates in Malaysia and 46.3% in India in 2018, against 23.6% in the United States in 2016 (UNESCO, through the World Bank). Neither figure says why.

Sources: NCES Digest of Education Statistics, table 325.35 · World Bank Gender Data Portal, female share of ICT graduates

  1. 2026-09-23: narrowing. Named on first publication of the page, already narrowing: the trend and the spread between countries count against a fixed difference, and nothing measured here says what does explain it.

Work this line at the Socratic Hearth →

Does diversity change the work

lc-5 · factual: dissolvable by measurement · status: open

Do more diverse teams produce better technical work?

Yes, and the effect is large enough to show up in company results.

Held by: McKinsey’s diversity studies, 2015 to 2023

The measured effect is mixed, and the largest published claims have not reproduced.

Held by: the meta-analysis and the reproduction below

A meta-analysis of 146 studies found that the harm often attributed to demographic diversity may come from rater bias: it did not appear in studies with more objective measures of performance (van Dijk, van Engen and van Knippenberg, 2012). On GitHub, gender and tenure diversity were positive and significant predictors of productivity (Vasilescu and six others, CHI 2015). Using the S&P 500, Green and Hand did not find McKinsey’s link between executives’ racial and ethnic diversity and profit margins (Econ Journal Watch, 2024). This page’s argument does not rest on this line: the independence of checks is a narrower claim.

Sources: van Dijk and others 2012 · Vasilescu and others 2015 · Green and Hand 2024

  1. 2026-09-23: open. Named on first publication of the page.

Work this line at the Socratic Hearth →

Three things worth counting, and where each is here

  1. Who a system fails, counted apart from how often it is right.The model-pair scorecard on the directions page would report false alarms and catches separately. Proposed, not yet built.
  2. The people affected, present where the decision is made.A question can be worked in a session at the Socratic Hearth, and what a group decides is kept as a record at The Work Party.
  3. Whether to build a thing at all, as an answer a system can give.The router on the directions page would refuse when no available check clears its threshold. Proposed, not yet built.

Written for the programme, in its voice. Drafted with Claude (Anthropic) at Bentley Moon’s direction, 23 September 2026. Every figure from outside this programme was read at its source before it was written here.