Evidence
Show me the evidence
Everything on this page is generated from the committed results.json. The website computes nothing. To check it yourself, run one command: Reproduce.
Current release
core sha256 covers every deterministic result. Timestamps, environment and latency are excluded so that a clean re-run reproduces the same hash on any machine.
Same attacks, three configurations
Arm A has no monitor. Arm B is a simple model of containment: any presented credential permits a mutation of the confined files, with no per-request binding. Arm C is the full v1 governor. The baselines show that the corpus really does present attack opportunities.
| Arm | Attack attempts | Changed protected state | Legitimate completed | False blocks |
|---|---|---|---|---|
| A — Unconstrained baseline | 60 | 58 / 60 | 24 / 24 | 0 |
| B — Containment only | 60 | 54 / 60 | 24 / 24 | 0 |
| C — Full v1 governor | 60 | 0 / 60 | 24 / 24 | 0 |
Attack families (Arm C)
Each scenario is one script; “attempts” counts the measured requests inside it. A request passes only if the decision and the reason code are the expected ones and both protected files, observed directly from disk, are unchanged (or advanced by exactly one version for legitimate requests).
| Family | Scenarios | Attempts | Passed | Protected-state violations | Reason codes observed |
|---|---|---|---|---|---|
| F01 Happy path (utility) | 5 | 24 | 24 | 0 | OK×24 |
| F02 Missing capability | 4 | 4 | 4 | 0 | MISSING_CAPABILITY×4 |
| F03 Expired / not-yet-valid capability | 5 | 5 | 5 | 0 | EXPIRED×4, NOT_YET_VALID×1 |
| F04 Revoked capability | 3 | 3 | 3 | 0 | REVOKED×3 |
| F05 Principal substitution | 3 | 3 | 3 | 0 | PRINCIPAL_MISMATCH×2, UNKNOWN_PRINCIPAL×1 |
| F06 Resource substitution | 2 | 2 | 2 | 0 | RESOURCE_MISMATCH×2 |
| F07 Effect substitution | 2 | 2 | 2 | 0 | EFFECT_MISMATCH×2 |
| F08 Replay of a single-use capability | 3 | 3 | 3 | 0 | REPLAY×3 |
| F09 Stale state (TOCTOU) | 4 | 4 | 4 | 0 | STALE_STATE×4 |
| F10 Policy version change / downgrade | 4 | 4 | 4 | 0 | POLICY_MISMATCH×4 |
| F11 Malformed capability | 7 | 7 | 7 | 0 | MALFORMED_CAPABILITY×7 |
| F12 Forged or tampered capability | 6 | 6 | 6 | 0 | BAD_SIGNATURE×5, DELEGATION_UNSUPPORTED×1 |
| F13 Unsupported or malformed request | 5 | 5 | 5 | 0 | MALFORMED_REQUEST×3, UNKNOWN_RESOURCE×1, UNSUPPORTED_EFFECT×1 |
| F14 UNKNOWN safety state (fail-closed) | 5 | 5 | 5 | 0 | POLICY_UNAVAILABLE×2, REVOCATION_UNAVAILABLE×1, STATE_UNAVAILABLE×2 |
| F15 Evidence witness failure | 2 | 2 | 2 | 0 | EVIDENCE_UNAVAILABLE×2 |
| F16 Policy denial | 3 | 3 | 3 | 0 | POLICY_DENY×3 |
| F17 Out-of-band modification detection | 2 | 2 | 2 | 0 | STATE_INCONSISTENT×2 |
Component ablations
Each row removes one mechanism from a copy of the monitor (in the untrusted harness; the real monitor has no off switch) and re-runs the corpus. A mechanism earns its place if removing it lets an attack family through.
| Mechanism removed | Attacks that changed state | Families exposed |
|---|---|---|
| No signature verification | 5 / 60 | F12 |
| No principal binding | 1 / 60 | F05 |
| No effect binding | 2 / 60 | F07 |
| No resource binding | 2 / 60 | F06 |
| No expiry check | 5 / 60 | F03 |
| No policy-hash binding | 3 / 60 | F10 |
| No revocation check | 4 / 60 | F04, F14 |
| No policy decision | 3 / 60 | F16 |
| No single-use (replay) check | 0 / 60 | — none |
| No state/version revalidation | 4 / 60 | F09 |
| No evidence-witness requirement | 2 / 60 | F15 |
Removing the single-use check alone exposes nothing, because the version binding independently blocks the replay. That redundancy is a finding about this profile, reported as observed.
Concurrency
Two threads race, released together, to spend authority bound to the same version. In the non-atomic variant the check and the commit are not under one lock and a barrier forces the interleaving, which demonstrates why atomic revalidation (condition C6) is required. Atomicity here is one in-process lock; multi-process or distributed atomicity is not demonstrated.
| Variant | Mode | Rounds | Double spends | Invariant violations | Loser reasons |
|---|---|---|---|---|---|
| Full monitor (locked) | same_capability | 100 | 0 | 0 | REPLAY×100 |
| Full monitor (locked) | competing_capabilities | 100 | 0 | 0 | STALE_STATE×100 |
| Non-atomic ablation | same_capability | 100 | 100 | 100 | — |
| Non-atomic ablation | competing_capabilities | 100 | 100 | 100 | — |
Direct-bypass probe: complete mediation
The probe writes to a protected file without going through the monitor. In this run: direct write SUCCEEDED, after setting the file read-only SUCCEEDED (uid 0; root ignores file modes). The next mediated request was then {"decision":"BLOCK","reason":"STATE_INCONSISTENT"}. The change is detected afterwards, not prevented. Outcomes depend on the operating-system user and are excluded from the reproducibility hash. Full list of channels: channel inventory.
Evidence quality and latency
Decision records emitted in the corpus run: 91, with 0 missing required fields; hash chains verified: 65 / 65. This shows the log is well-formed and tamper-evident against naive edits. It does not show that every path to the files is logged.
Latency of the in-process mechanical path (300 samples per path, microseconds, includes file I/O and fsync; varies by machine and is not part of the hash): blocked requests p50 29.9, p99 143.4; allowed requests p50 1522.8, p99 3124.
Claim registry
Every public claim has an id, a scope, a status and its known limitations. “Locally tested” means the author's corpus exercises it; it does not mean independently reproduced. Epistemic labels D / A / S / E / R are the paper's, not confidence scores. Source: CLAIMS.json.
CLM-001LOCALLY TESTEDE — Empirical assuranceFor the declared attack corpus, no unauthorized mutation of the two protected files occurred when requests went through the reference monitor.
Scope. protected-file-v1, attack-corpus-1, in-process monitor
Assumptions. C1 (partial: declared effect surface only, not established); C2; C5; C6; C7
Attack families. F02, F03, F04, F05, F06, F07, F08, F09, F10, F11, F12, F13, F14, F15, F16, F17
Known limitations.
- Finite corpus; zero observed failures do not prove zero real-world risk
- No complete-mediation proof
- No independent audit
CLM-002LOCALLY TESTEDS — Security / invariant claimA capability is bound to one principal, one effect and one resource; substitution of any of the three is blocked, and edits to a signed capability are detected.
Scope. protected-file-v1
Assumptions. C5
Attack families. F05, F06, F07, F12
Known limitations.
- HMAC (symmetric): compromise of the monitor implies capability forgery
- Principal attribution is by declared id; no authentication of the requesting process
CLM-003LOCALLY TESTEDS — Security / invariant claimExpired, not-yet-valid and revoked capabilities (by id or by epoch) are rejected.
Scope. protected-file-v1
Assumptions. C5; C7
Attack families. F03, F04
Known limitations.
- Time comes from the host clock; clock tampering is out of scope
- Revocation store is in-memory in v0.1
CLM-004LOCALLY TESTEDS — Security / invariant claimA single-use capability cannot be spent twice, including under concurrent attempts within one process.
Scope. protected-file-v1, single process
Assumptions. C6
Attack families. F08
Known limitations.
- Atomicity comes from one in-process lock; multi-process or distributed atomicity is not demonstrated
- State/version binding independently blocks replays, so the single-use check is redundant for version-bound capabilities (reported by the ablation)
CLM-005LOCALLY TESTEDS — Security / invariant claimResource state is revalidated at execution time; a capability issued for a superseded version is rejected.
Scope. protected-file-v1
Assumptions. C6
Attack families. F09
Known limitations.
- Revalidation and commit are atomic only with respect to other mediated requests
- Non-mediated writers are detected afterwards (CLM-009), not excluded
CLM-006LOCALLY TESTEDS — Security / invariant claimUnresolvable safety state (policy, revocation, resource metadata, evidence witness) is treated as UNKNOWN and mapped to BLOCK.
Scope. protected-file-v1
Assumptions. C7; C8 (partial)
Attack families. F14, F15
Known limitations.
- Failures are injected by test doubles, not by real infrastructure faults
- If the evidence witness fails after commit, the effect has already happened; the gap is reported, not prevented
CLM-007LOCALLY TESTEDS — Security / invariant claimCapabilities are bound to the active policy hash, a deny-by-default policy is evaluated per request, and candidate policies are statically checked at admission against the Safety Profile.
Scope. protected-file-v1
Assumptions. C4 (not established)
Attack families. F10, F16
Known limitations.
- Restricted exact-match language; executable admission check only
- No proof object and no machine-checked refinement checker (paper condition C4/PAV remains open)
CLM-008LOCALLY TESTEDE — Empirical assuranceEvery mediated request produces a structured, hash-chained evidence record from which the request, the checks passed or failed, and the outcome can be reconstructed.
Scope. protected-file-v1
Assumptions. C8 (partial)
Known limitations.
- Local log only; no external witness
- Hash chain detects naive tampering; it does not establish legal non-repudiation
- Evidence completeness across alternate paths is not established
CLM-009LOCALLY TESTEDE — Empirical assuranceA change to a protected file made through a non-mediated channel is detected on the next mediated request and that request is blocked.
Scope. protected-file-v1
Attack families. F17
Known limitations.
- Detection after the fact, not prevention
- An attacker who also rewrites the version sidecar consistently is not detected by this mechanism
CLM-010OPENS — Security / invariant claimComplete mediation of the protected effect surface (condition C1).
Scope. protected-file-v1
Assumptions. C1
Known limitations.
- v0.1 runs monitor and mediated code in one process and uid; direct writes are possible (bypass probe)
- Roadmap v0.2: OS-level privilege separation and alternate-path tests
CLM-011OPENS — Security / invariant claimConformance of the monitor to verified semantics (kernel conformance, C3) and integrity of the trusted core (C2).
Scope. protected-file-v1
Assumptions. C2; C3
Known limitations.
- No formal verification or kernel-level assurance
- Trusted core is small (governor/*.py) but unproven
CLM-012OPENS — Security / invariant claimMachine-checked typed refinement from the deployed policy to the admitted safety constitution (C4).
Scope. protected-file-v1
Assumptions. C4
Known limitations.
- Roadmap v0.3
CLM-013OPENS — Security / invariant claimCapability delegation with attenuation (child scope, budget and expiry bounded by the parent).
Scope. protected-file-v1
Attack families. F12
Known limitations.
- Not implemented; capabilities naming a parent are rejected (tested), nothing more
- Roadmap v0.4
CLM-014OPENE — Empirical assuranceIndependent security review, external red-team validation and held-out adaptive evaluation.
Scope. protected-file-v1
Known limitations.
- None performed; invitations to falsify are in SECURITY.md and AUDIT.md
CLM-015SPECIFIEDS — Security / invariant claimConditional Behavioral Safety Theorem: under C1-C8 every executed protected effect satisfies the admitted safety predicate.
Scope. Paper section 1.3 / Table 1
Assumptions. C1; C2; C3; C4; C5; C6; C7; C8
Known limitations.
- Conditional on all eight conditions; this prototype demonstrates only a subset (see CLM-001..CLM-009 and CLM-010..CLM-012)
CLM-016SPECIFIEDD — DerivationL1 (public-ground component) and NRD (no unjustified normative difference), plus the constitutional-reflexivity constraint CR/CAR, as derived in the paper's practice-based ethics.
Scope. Paper sections 2-6
Assumptions. Paper premises P1 and following; classical first-order logic
Known limitations.
- Derivations are not machine-checked (paper section 10.13)
- Substantive standards B4-B8 are declared (A), not derived
- This repository does not implement or test the philosophical layer
CLM-017NOT APPLICABLEE — Empirical assuranceSemantic assurance (G-Osem): calibrated probabilistic bounds on neural state extraction.
Scope. Not part of v0.1
Assumptions. SA (sensor adequacy)
Known limitations.
- The v0.1 monitor is purely mechanical and makes no semantic judgments
CLM-018RESIDUALR — Residual / openCovert channels, unmediated physical effects, compromised root of trust and operator-controlled disabling are outside the guarantee.
Scope. Paper section 10.15
Known limitations.
- These are the boundary of the theorem, not defects of terminology
Raw data
- results.json — machine-readable results
- manifest.json — release manifest binding results to hashes
- CLAIMS.json — claim registry
- tests/ and evaluation/scenarios.py — the attack corpus source
Reproduce
git clone https://github.com/canon4d/the-dialectical-cage-and-the-glass-box-governor
cd the-dialectical-cage-and-the-glass-box-governor
python3 scripts/run_release_tests.pyThen compare your core sha256 with 013b7ced03fac4fc…. Details and troubleshooting: Reproduce. Overall shape of the corpus: 84 measured requests, 0 / 60 unauthorized executions.