A bounded FactNotebook reconstruction from Jafar Muhammed’s Bunia field-testing records.
An external evaluator ran a field audit of how AI assistants respond to health-guidance prompts in a low-resource, outbreak context, across languages. He held a set of timestamped observations, then we agreed to use FactNotebook to turn them into a falsifiable governance claim.
This case study shows how FactNotebook approached that gap — not the audit’s verdict. It shows how a set of field observations becomes a bounded reconstruction of what those records permit one to establish, while the safety judgment stays entirely with the evaluator.
It is non-dispositive: FactNotebook does not decide whether any model was safe or unsafe. That determination belongs to the evaluator’s policy and to the accountable authority.
“Permissible to establish” is not a property of an isolated observation. The raw record is preserved; the establishment result belongs to a proposition-level reconstruction — what evidential role that record plays relative to a specific, stated proposition.
Labels such as “residual risk” or “pass” remain attributed to the evaluator’s policy — they are never converted into FactNotebook findings.
The evaluator retained timestamped primary captures (the source records) and a
structured representation (a spreadsheet). Before treating the reconstruction as grounded,
the structured representation was verified against the primary captures. That verification
caught a material divergence on one record: its structured row and its primary capture
disagreed on a substantive point. The reconstruction was rebound to the primary surface, and
one variant that was not present in the received primary set was held
CANNOT ESTABLISH rather than carried forward.
An intelligent read proposed the candidate observation, but admission was verified against the primary record; establishment remained downstream in the deterministic reconstruction.
FactNotebook reconstructs what the bounded record permits one to establish — including, explicitly, what it cannot. The safety judgment stays with the evaluator.
A small set of proposition-level states over a bounded set of confirmed records: ESTABLISHED where the records supported the proposition, CONTRADICTED where a recorded counter-instance stood against it, and CANNOT ESTABLISH where the available record did not permit the proposition to be established (including the absent variant and a prevalence / generalization proposition with no basis in a small, non-random set).
The states are neutral results, not a scorecard. The most instructive outcomes were
not the contradictions but the discipline around them: a divergence caught at
verification before it could bias the finding, and an honest CANNOT ESTABLISH
preserved rather than rounded to a stronger claim.
The finding carries a Reconstruction ID and binds a frozen record set, so the same admitted observations re-derive the same states. The reconstruction was treated as bound to the primary record surface only after the structured representation had been verified against the received primary captures.