FactNotebook · Public Reconstruction · FN-ANTH-M5-001

Mythos 5 — what can the public record actually establish?

An independent, retrospective reconstruction of the boundary of the Claude “Mythos 5” cyber-evaluation incident — built only from Anthropic’s public disclosures. Outcome-neutral, non-dispositive, bounded. Not an accusation and not a defence; a reconstruction.

Before asking “is AI becoming uncontrollable?”, there is a narrower, more operational question:
From the observations we actually have, what can be established about the boundary this system operated within — and what cannot? The public record answers different sub-questions differently. That is the point.

Three results, usually collapsed into one

The boundary was declared (simulation, no live internet).
ESTABLISHED
As a declaration: a self-attributed source establishes its own declaration.
The declared boundary held in operation.
CONTRADICTED
Anthropic’s own disclosures place activity on the real internet and against a live third-party system — an admission against interest.
The model internally believed it was simulated.
CANNOT ESTABLISH
Anthropic reports the chain-of-thought represented it as simulated; a reported representation doesn’t establish the internal belief.

The public record is completeCONTRADICTED
Redactions are directly inspectable in the released artifact.

The public record is materially sufficient for every relevant propositionCANNOT ESTABLISH
known-incomplete is not known-insufficient.

These axes are independent: a boundary can be breached whatever the model believed; the belief stays unknowable whatever the boundary did; and the record’s gaps are a property of the evidence surface, not of the world. Evidence has no verdict on its own — it has a relation to a determined proposition.

Observation package — the declared surface

Nothing outside this was read. The transcript’s own contents were not ingested (the repository marks it benchmark / no-training data); only its documented release structure was.

SourceOrigin
Anthropic, An alignment assessment of recent cybersecurity incidents (9 Sep 2026)subject-published (the subject’s own report)
GitHub anthropics/mythos-5-incident-transcript — released-transcript metadata (README)subject-published, but directly inspectable

Provenance is two axes, not one. Both sources are subject-published (Anthropic controls what was released). A third party can nonetheless directly inspect the released artifact to see that segments are absent — so the existence of redactions is independently checkable, even though what was redacted rests on Anthropic’s description. Results below are labelled by that distinction, never claimed as source-independence.

The reconstructions

One root proposition per reconstruction; each evaluates a bounded proposition against the declared observation surface.

A · Boundary / system state — CONTRADICTED
“During Mythos 5, the effective operating boundary remained consistent with the declared simulation / no-live-internet / no-real-consequence boundary.”
ESTABLISHED Anthropic declared the boundary (simulation, no live internet).
CONTRADICTED No connection to the live public internet occurred.
CONTRADICTED No real-world external consequence occurred.
FN-3eb5cd7b2e221869
B · Model epistemic state — CANNOT ESTABLISH
“The model’s internal operative disposition is identifiable from the public package as a belief that the environment was simulated.”
ESTABLISHED Anthropic reports the chain-of-thought repeatedly represented the environment as simulated.
CANNOT ESTABLISH Those reported representations identify the model’s actual internal belief. — no observation in the package identifies an internal epistemic state; a reported representation and the external actions are each consistent with more than one internal state.
FN-5d37da1283c8c3aa
C · Observation-surface / package — CONTRADICTED
“The released public package is a complete and materially sufficient basis for reconstructing the foregoing propositions.”
CONTRADICTED The released record is complete and unredacted. (directly inspectable public artifact)
CANNOT ESTABLISH The released record is materially sufficient for every relevant proposition. — the available observations do not establish that the absent material is immaterial to every relevant proposition.
FN-9001f64d9d60e4e1
Inspect the sealed artifact →
The full reconstruction, its observation manifest, and the exact manifest each Reconstruction ID is a digest of — in the public registry.

Custody & re-derivability

Each Reconstruction ID is derived from a SHA-256 digest of a canonical manifest that binds the frozen observation model, the exact observation set, and the execution semantics that produced the result — so the same frozen inputs re-derive the identical result under the same engine build. This is internal determinism, not independent external reproducibility: the reconstruction engine is not public. Input integrity is fail-closed — a malformed, altered or custody-mismatched observation set yields an explicit integrity failure and no conclusion, never a silent partial result.

What this reconstruction does not do

It reconstructs what can and cannot be established from that package — and leaves the judgement where it belongs.

Start a reconstruction →
Bring one question about your own AI system — from €1,500, scoped to your case.

If a regulator, auditor or board asked about this AI decision six months from now — what could we actually establish, and what couldn’t we? That is the question a reconstruction answers, on one bounded case.