What the evidence supports — per claim, with its sources. Judgment of sufficiency is yours.
The evidence layer is the product; the verdict is only its projection under a policy.
This document reconstructs observable facts, their provenance, their custody, their contradictions, and the observation gaps. It does not determine whether the available evidence is sufficient. Sufficiency remains an assurance decision performed by the designated authority.
Evidence is reconstructed. Assurance is conferred. This document performs the former and explicitly leaves the latter to the designated assurance authority.
0.9.0 · Question set : engine-builtins v1 (2026-07-03) · Control mapping (proposed pack, versioned) : factnotebook-proposed-pack v0 (EU-AI-Act-flavoured, editable)default-strict v2 — governance states only (CONFIRMED / CONTRADICTED / NOT ASSESSABLE); every NA carries a runtime reason (not_executed / no_channel / channel_broken / skipped) — no reward for not looking is structural; contradiction-dominant on INDISPENSABLE roles only; CONFIRMED requires ALL declared questions confirmed; mixity → PARTIAL; all-NA → NOT ASSESSABLE · unsigned (authored, not counter-signed)| Source class | Records |
|---|---|
| Declared observations (A1) | 2 |
| System observations (A2) | 5 |
| Independent observations (A3+) | 2 |
| Derived observations | 0 |
| Missing observations | 20 |
Provenance, custody-and-integrity, and assurance are independent — properties, not confidence. A self-attested source (A2) can be tamper-evident (custody) yet unaccepted (assurance): each axis is established by a different party and never inflates another.
| Evidence axis | Established by | Current |
|---|---|---|
| Source provenance | System | A3 |
| Integrity | FactNotebook | self-hash + RFC3161 anchor at an external TSA (DigiCert) — independently verifiable by anyone |
| Custody | FactNotebook | single-hop, unsigned — integrity externally anchored; counter-signature still: — (that step belongs to an authority) |
| Assurance | Authority | None — no counter-signature |
Custody from ingest (A2). An independent counter-signature raises it — a custody event, not a change to the evidence. The empty slots below are reserved, not omitted.
FactNotebook establishes a custody record from ingestion onward.
| Evidence package | swebench:creachadair__jrpc2-81 |
Integrity (content unchanged)
| Ingest hash (SHA256) | f402a66964d45ab79e9a87753e9703869855f5f030b47099eb40c0125f0783d5 — verifiable since ingest |
| RFC3161 timestamp | ingest artifact anchored at an external time-stamp authority (DigiCert) — token published as ingest.tsr; verify with `openssl ts -verify` |
Custody (handling chain)
| Custody established | FactNotebook (from ingest) |
| Collected (ingest) | 2026-07-31T20:17:29+00:00 |
| Connector | SWEBenchTrajectoryConnector (unversioned) |
| Storage | .factdna |
| Digital signature | — (custody from ingest, unsigned) |
| Counter-signature | — (none yet) |
| Immutable ledger | — (none yet) |
| Output manifest | manifest.json — SHA256 of every artifact in this package, published alongside the dossier (itself unsigned) |
Declared question set fit to the artifact (a coding-agent trajectory). Enterprise governance controls (Access Policy, Change Governance, Decision Workflow, Evidence Integrity) are retained in the enumeration but marked NOT ASSESSABLE with reason not_applicable: they have no referent in a benchmark trajectory (no deployer / change board / decision workflow / post-market monitoring). This is distinct from no_channel (a referent exists but is not instrumented) and not_executed (could have run, did not).
A control's state is a roll-up, not an average: it reads CONTRADICTED when an indispensable claim is contradicted — even alongside confirmed and partial claims (the strongest signal dominates); CONFIRMED requires every declared claim confirmed; a mix is PARTIAL; all-unobservable is NOT ASSESSABLE.
| scope | criterion: per-role outcomes within coded tolerance of each other events_examined: 303 sources: ["SWEBenchTrajectoryConnector (A2, the execution's own transcript — real)", "SWE-bench evaluation harness (A3, independent: runs the repo's own tests)", "GitHub (A3, independent public source: the repo's own review norm)", 'declared autonomous-agent governance constraints (accuracy; oversight)'] |
| scope | criterion: transitive parent resolution reaches a human principal (no cycle, no orphan) events_examined: 303 sources: ["SWEBenchTrajectoryConnector (A2, the execution's own transcript — real)", "SWE-bench evaluation harness (A3, independent: runs the repo's own tests)", "GitHub (A3, independent public source: the repo's own review norm)", 'declared autonomous-agent governance constraints (accuracy; oversight)'] |
| declaration | constraint_id: SWE-GOV-01 article: Art.15 source: declared governance requirement |
| harness_result | instance_id: creachadair__jrpc2-81 repo: creachadair/jrpc2 resolved: 0 num_modified_files: 3 num_modified_lines: 189 label_source: SWE-bench published evaluation label (resolved) citation: {'dataset': 'nvidia/Open-SWE-Traces', 'config': 'openhands', 'split': 'qwen35_122b', 'instance_id': 'creachadair__jrpc2-81'} attribution: third-party publication (cited), not a local re-run |
| declaration | constraint_id: SWE-GOV-02 article: Art.14 source: declared governance requirement |
| runtime_property | approval_gates: 0 mutating_calls: 24 tool_errors: 0 autonomous: True acceptance_event: False source: SWEBenchTrajectoryConnector (observed run) |
| github | repo: creachadair/jrpc2 merged_prs: 55 review_events: 3 captured_at: 2026-07-31T20:17:28.373922+00:00 window: ['2021-08-01T20:16:53.155677+00:00', '2026-07-31T20:16:53.155677+00:00'] review_pointers: [{'pr': 114, 'reviewer': '2opremio', 'state': 'APPROVED', 'submitted_at': '2024-04-01T16:01:36+00:00', 'permalink': 'https://github.com/creachadair/jrpc2/pull/114#pullrequestreview-1971589357'}, {'pr': 90, 'reviewer': 'radeksimko', 'state': 'APPROVED', 'submitted_at': '2022-12-12T11:49:53+00:00', 'permalink': 'https://github.com/creachadair/jrpc2/pull/90#pullrequestreview-1213376105'}, {'pr': 79, 'reviewer': 'radeksimko', 'state': 'APPROVED', 'submitted_at': '2022-02-08T20:49:20+00:00', 'permalink': 'https://github.com/creachadair/jrpc2/pull/79#pullrequestreview-876631641'}] source: GitHubConnector (independent public API) |
| scope | criterion: actor identity != approver identity for the same resource/decision events_examined: 303 sources: ["SWEBenchTrajectoryConnector (A2, the execution's own transcript — real)", "SWE-bench evaluation harness (A3, independent: runs the repo's own tests)", "GitHub (A3, independent public source: the repo's own review norm)", 'declared autonomous-agent governance constraints (accuracy; oversight)'] |
| scope | criterion: overlapping access windows on the same resource flagged events_examined: 303 sources: ["SWEBenchTrajectoryConnector (A2, the execution's own transcript — real)", "SWE-bench evaluation harness (A3, independent: runs the repo's own tests)", "GitHub (A3, independent public source: the repo's own review norm)", 'declared autonomous-agent governance constraints (accuracy; oversight)'] |
Both sides are reported with their provenance; the engine never picks a winner. 'Kind' only locates the disagreement — between channels (cross-channel) or within one (intra-channel). Whether a cross-channel disagreement is a true governance conflict, a mapping error, or a false positive is the auditor's call, not the engine's.
| control | claim | kind | n | provenance ceiling | detail |
|---|---|---|---|---|---|
| FactNotebook reconstruction — proposed coding-agent control pack | A produced fix must pass the repository's tests before it is presented as resolving the issue [SWE-GOV-01] | cross-channel | 1 | A3 (third-party-published ground truth: SWE-bench's published `resolved` label for this instance — cited, not re-run) | SWE-bench published resolved=0 (FAILED); declared patch touched 3 file(s), 189 line(s); cited from nvidia/Open-SWE-Traces [openhands/qwen35_122b] instance creachadair__jrpc2-81 (not a local harness run) ; reliability: the label attests the patch passed/failed this benchmark's own test suite, not that the issue was correctly solved — public benchmark tasks are themselves fallible, so a defective task can yield a false pass or false fail |
A question with no observation channel is not a failure of the system — it is the map of where you are not set up to know. Each row names the minimum observable event that would make the question answerable: an evidence contract, not a verdict.
| control | claim | reason | missing observation contract |
|---|---|---|---|
| Human Oversight | An independent oversight role was present | no_channel | agent_id agent_role |
| Human Oversight | Critical actions were independently attested | not_executed | agent_id approval_event_present |
| Human Oversight | Approvals reference the action they authorize | not_executed | process_reference_present resource_ref |
| Human Oversight | Oversight did not silently degrade over time | not_executed | agent_role run_id |
| Mission Containment | Actions stayed within the declared mission zone | not_executed | resource_ref |
Showing 5 representative evidence contracts of 20 · full map in the review package.
The projection of the observed states above under one named policy — rendered last on purpose. Blocking is determined by the declared evaluation policy, not by the evidence itself.
Applied policy
| Evidence package | swebench:creachadair__jrpc2-81 |
| Evaluation policy | default-strict v2 |
| Policy status | unsigned · author-defined · not counter-signed |
| Outcome under this policy | CONTRADICTED |
The engine never picks which contradictions block — that is a declared policy. Here the same evidence is re-rolled under alternative, equally defensible policies. A stable verdict is robustness; a flip is disclosed, not hidden.
default-strict v2 (indispensable-blocking) | CONTRADICTED reference |
strict-any v1 (any contradiction blocks) | CONTRADICTED |
consensus v1 (a claim blocks only if no channel confirms it) | CLEAN flips |