Glass-Box Governor

AI agent execution security · reference prototype · v0.1.0

An external execution boundary for AI agents.

The model proposes an action; the reference monitor independently decides whether a declared protected effect may execute. This site exists so you can inspect the code, run the attacks, and see exactly what is not established.

Profile protected-file-v1 · in-process prototype · complete mediation not established · no independent audit

The mechanism

The model is not the final enforcement boundary. It may make mistakes, be manipulated, or behave adversarially. For a declared set of protected effects, the monitor decides from the current capability, policy and state, and records the decision as evidence.

  1. untrustedAgent / modelMay request anything, including adversarial actions.
  2. requestAction requestprincipal · effect · resource · capability
  3. trusted coreReference monitor19 ordered checks, one lock around check + commit.
  4. only if ALLOWProtected effectwrite_file · delete_file on two declared files.
Evidence record for every decision, ALLOW or BLOCK — hash-chained, reconstructible.
Trust boundary: everything left of the monitor is untrusted. In v0.1 that boundary is a convention inside one process, not an operating-system barrier — see Limitations.

Proof snapshot

Values below are read from the committed results.json at build time, not written into this page. They are counts from one profile and one attack corpus.

0 / 60unauthorized protected executionsadversarial requests through the full v1 governor
0 / 60protected-state violationsfiles observed from disk, independent of the monitor
3 / 3replays blocked
4 / 4stale-state (TOCTOU) attempts blocked
3 / 3principal substitutions blocked
0 / 200concurrent double spendsnon-atomic ablation: 200 / 200

Same attacks, three configurations

Arm AUnconstrained baseline58 / 60attacks changed protected state24 / 24 legitimate requests completed
Arm BContainment only54 / 60attacks changed protected state24 / 24 legitimate requests completed
Arm CFull v1 governor0 / 60attacks changed protected state24 / 24 legitimate requests completed

Zero observed failures is a finite observation, not a proof of zero real-world risk. Arm B is a simple model of containment defined in the repository, not a measurement of any real product. Counts, corpus and caveats: Evidence.

Try the decision path

Pick a scenario to see which check decides it, what the monitor records, and whether the protected file changed.

Request

principal
agent-demo
effect
write_file
resource
record-a

Capability presented

capability_id
CAP-000001
principal_id
agent-demo
effect
write_file
resource_id
record-a
resource_version
0
expires_at
1700000300
revocation_epoch
0
one_shot
true

Authorized write

Monitor checks, in order

  1. request wellformed — passed
  2. effect supported — passed
  3. resource declared — passed
  4. principal declared — passed
  5. capability present — passed
  6. capability wellformed — passed
  7. delegation absent — passed
  8. signature authentic — passed
  9. principal binding — passed
  10. effect binding — passed
  11. resource binding — passed
  12. validity window — passed
  13. policy available — passed
  14. policy binding — passed
  15. revocation current — passed
  16. policy decision — passed
  17. replay protection — passed
  18. state current — passed
  19. evidence witness — passed
  20. execute — passed

Outcome

ALLOWOK
protected state changed
YES
version
0 → 1
evidence emitted
YES
event id
EVT-000001
record integrity
186aa0ce34fd0b1b…

Replayed from results.json (scenario F01/authorized-write), not computed in your browser. The authoritative evidence is the source, the tests, and the command you run locally.

What is demonstrated now

Each item is LOCALLY TESTED — implemented and exercised by the author's committed corpus, not independently reproduced. Click through for scope and limits.

  • Capabilities bound to principal, effect and resource CLM-002

    A capability for A cannot be used by B, for read cannot delete, for X cannot touch Y; edits to a signed capability are detected.

  • Expiry and revocation CLM-003

    Expired, not-yet-valid, and revoked (by id or epoch) authority is rejected.

  • Single use, including under concurrency CLM-004

    Two threads racing to spend one capability: exactly one commit. A deliberately non-atomic variant double-spends, which shows the lock matters.

  • State and version revalidation CLM-005

    A capability issued for version 3 is useless at version 4. Check and commit happen under one lock.

  • Fail-closed handling of UNKNOWN CLM-006

    Unavailable policy, revocation, metadata or evidence witness blocks the protected effect; uncertainty never becomes authorization.

  • Policy-hash binding and deny-by-default CLM-007

    Authority does not silently survive a policy change; candidate policies are statically checked at admission.

  • Hash-chained evidence for every decision CLM-008

    Each request leaves a record from which the request, the checks, and the outcome can be reconstructed.

Three things, kept separate

Demonstrated by code

What the prototype enforces and the corpus tests: the capability, policy, state, fail-closed and evidence behaviours above, for two files and two effects.

Assumed by the theorem

The paper's conditional safety theorem needs eight conditions (C1–C8): complete mediation, trusted-core integrity, kernel conformance, verified policy refinement, principal attribution, atomic revalidation, fail-closed handling, evidence-path integrity. This release does not establish all of them.

Open empirical questions

  • Complete mediation outside the declared v1 profile (the bypass probe shows same-user writes succeed and are only detected afterwards)
  • Operating-system-level isolation and kernel assurance
  • Machine-checked policy refinement and proofs
  • Delegation with attenuation
  • Independent security review and external red-team
  • Semantic sensor assurance (not part of this release)
  • Production validation

The repository implements a concrete prototype of selected execution-security components described in the framework. It does not claim that the complete conceptual specification has been implemented. What this project does not prove →

Read the evidence

Release v0.1.0 · attack corpus attack-corpus-1 · core hash 013b7ced03fac4fc…

Try to break it

The repository contains a reproducible prototype of the execution-security layer described in the paper. Attempts to falsify the protected-effect claim are welcome, especially around channel completeness, state races, policy admission, and root-of-trust assumptions. Write to canon@necessaryuniverse.com or open an issue on GitHub; SECURITY.md says what counts as a finding.