The Governability Index

How governable is an AI system — not by policy, but by evidence?

A system cannot be governed if its behaviour cannot be reconstructed. Each profile measures what an independent reviewer could actually establish about how a real AI system behaved — from public evidence alone.

Each profile reconstructs the observable history of a real AI system from public evidence, checks declared-vs-observed, and reports what is confirmed, contradicted, and honestly not assessable — each graded by how independent the evidence is. Reproducible from the public source; assessed independently, not affiliated with the systems' authors.

Profiles

GOVERNED LLM · CLINICAL QAReal data · independently reviewed · NeoMundi ControlTower × GPT-4o
GPT-4o governed via NeoMundi ControlTower (PubMedQA)  ·  PARTIALLY RECONSTRUCTED
A profile built from an independent partner's runtime measurements — not the system's own logs. NeoMundi measures; FactNotebook reconstructs; the assurance decision stays with a human. Reviewed by both parties.
Evidence state PartialObservability 33%Independent evidence 89%Highest provenance A3NA debt 31
See what an independent witness makes reconstructible →
CODING AGENTReal public data · reproducible · SWE-bench trajectory
OpenHands on creachadair/jrpc2  ·  CONTRADICTION OBSERVED
100 tool calls, 24 file edits, a declared fix — and the repository's own tests say it never worked. So where did the confidence come from?
Benchmark FAILEDEvidence state Contradiction observedObservability 29%Highest provenance A3NA debt 20
See what the trajectory actually shows →
CODING AGENTReal public data · reproducible · SWE-bench trajectory
OpenHands on litestar-org/polyfactory  ·  NO CONTRADICTION OBSERVED
It passed every test — and changed production code with zero human review. "Clean" is not the same as "governed." Here is exactly what could, and could not, be verified.
Benchmark PASSEDEvidence state No contradiction observedObservability 29%Highest provenance A3NA debt 20
See what's confirmed — and what can't be →

Three profiles today. The index grows one system at a time.

Principles

1
No evidence is invented.A profile reports only what was observed. It never fills a gap with a guess.
2
NOT ASSESSABLE is a valid outcome.When no channel recorded the answer, that is the honest result — and it names what the system would need to emit to become answerable.
3
Evidence is preserved independently of system execution.A system cannot preserve its own evidence with independence. It is inside its own trust boundary.
4
Trust is assigned by authorities — not by evidence infrastructure.The index exposes properties. An auditor or a regulator decides admissibility. Those are different powers.
5
Governability is limited by observability.A system cannot be governed beyond what it can honestly observe.
Share