Skip to content
Legible AI

Six questions separate the current results from a dependable oversight system.

These are not predictions dressed as findings. Each question names the missing evidence and the experiment that could answer it. A negative result would still narrow the design space.

  1. How much reference access does useful oversight require?The observed gap between applying a supplied test and recomputing an answer is large. The missing object is a ladder between those endpoints.
    What would answer it: Hold tasks and models fixed while varying only the reference supplied: answer, executable predicate, worked intermediate state, retrieval, partial trace, and no reference.
  2. Do competence and shared error predict oversight value across vendors?The local pilot supports that joint model, but vendor and task diversity are still too narrow for a deployment rule.
    What would answer it: Run paired generators and verifiers from several model families on externally graded tasks, measuring false alarms, catches, calibration, cost, and error overlap on the same cells.
  3. Where does a monitor cross from weak to inverted?A monitor that rejects correct work more often than corrupted work is worse than an uninformative monitor. The crossing point matters more than its average score.
    What would answer it: Sweep difficulty one controlled operation at a time, keep the predicate fixed, and locate the first point at which signed discrimination falls below zero.
  4. Can inversion be detected without a hidden answer key?Ground truth is often unavailable at deployment time. Health metrics that do not depend on it may continue improving while the monitor becomes harmful.
    What would answer it: Test disagreement, abstention shape, sentinel tasks, cross-family divergence, and delayed executable checks against the known inversion boundary.
  5. When does repair discover information, and when does it entrench a mistake?Repair can cross a failure boundary where a model has a foothold and regress where it does not. The active mechanism remains unresolved.
    What would answer it: Compare execution feedback with information-matched, information-free, and adversarial feedback while recording near misses and approach changes per task.
  6. What may an AI system honestly report about its own reasoning?A self-report is produced by the system it describes. It may be useful evidence without being an independent account.
    What would answer it: Compare self-reports with execution traces, interventions, hidden-state probes, and human-readable predictions of later behavior. Grade correspondence, not eloquence.

What should happen next

The reference-composition ladder comes first because it tests the programme's central architectural claim directly. The cross-vendor scorecard follows because it tests whether the local competence-and-overlap model travels.

See the technical sequence