AI agent execution security · reference prototype · v0.1.0
An external execution boundary for AI agents.
The model proposes an action; the reference monitor independently decides whether a declared protected effect may execute. This site exists so you can inspect the code, run the attacks, and see exactly what is not established.
Profile protected-file-v1 · in-process prototype · complete mediation not established · no independent audit
The mechanism
The model is not the final enforcement boundary. It may make mistakes, be manipulated, or behave adversarially. For a declared set of protected effects, the monitor decides from the current capability, policy and state, and records the decision as evidence.
- untrustedAgent / modelMay request anything, including adversarial actions.
- requestAction requestprincipal · effect · resource · capability
- trusted coreReference monitor19 ordered checks, one lock around check + commit.
- only if ALLOWProtected effectwrite_file · delete_file on two declared files.
Proof snapshot
Values below are read from the committed results.json at build time, not written into this page. They are counts from one profile and one attack corpus.
Same attacks, three configurations
Zero observed failures is a finite observation, not a proof of zero real-world risk. Arm B is a simple model of containment defined in the repository, not a measurement of any real product. Counts, corpus and caveats: Evidence.
Try the decision path
Pick a scenario to see which check decides it, what the monitor records, and whether the protected file changed.
Request
- principal
- agent-demo
- effect
- write_file
- resource
- record-a
Capability presented
- capability_id
- CAP-000001
- principal_id
- agent-demo
- effect
- write_file
- resource_id
- record-a
- resource_version
- 0
- expires_at
- 1700000300
- revocation_epoch
- 0
- one_shot
- true
Authorized write
Monitor checks, in order
- request wellformed — passed
- effect supported — passed
- resource declared — passed
- principal declared — passed
- capability present — passed
- capability wellformed — passed
- delegation absent — passed
- signature authentic — passed
- principal binding — passed
- effect binding — passed
- resource binding — passed
- validity window — passed
- policy available — passed
- policy binding — passed
- revocation current — passed
- policy decision — passed
- replay protection — passed
- state current — passed
- evidence witness — passed
- execute — passed
Outcome
- protected state changed
- YES
- version
- 0 → 1
- evidence emitted
- YES
- event id
- EVT-000001
- record integrity
- 186aa0ce34fd0b1b…
Replayed from results.json (scenario F01/authorized-write), not computed in your browser. The authoritative evidence is the source, the tests, and the command you run locally.
What is demonstrated now
Each item is LOCALLY TESTED — implemented and exercised by the author's committed corpus, not independently reproduced. Click through for scope and limits.
- Capabilities bound to principal, effect and resource CLM-002
A capability for A cannot be used by B, for read cannot delete, for X cannot touch Y; edits to a signed capability are detected.
- Expiry and revocation CLM-003
Expired, not-yet-valid, and revoked (by id or epoch) authority is rejected.
- Single use, including under concurrency CLM-004
Two threads racing to spend one capability: exactly one commit. A deliberately non-atomic variant double-spends, which shows the lock matters.
- State and version revalidation CLM-005
A capability issued for version 3 is useless at version 4. Check and commit happen under one lock.
- Fail-closed handling of UNKNOWN CLM-006
Unavailable policy, revocation, metadata or evidence witness blocks the protected effect; uncertainty never becomes authorization.
- Policy-hash binding and deny-by-default CLM-007
Authority does not silently survive a policy change; candidate policies are statically checked at admission.
- Hash-chained evidence for every decision CLM-008
Each request leaves a record from which the request, the checks, and the outcome can be reconstructed.
Three things, kept separate
Demonstrated by code
What the prototype enforces and the corpus tests: the capability, policy, state, fail-closed and evidence behaviours above, for two files and two effects.
Assumed by the theorem
The paper's conditional safety theorem needs eight conditions (C1–C8): complete mediation, trusted-core integrity, kernel conformance, verified policy refinement, principal attribution, atomic revalidation, fail-closed handling, evidence-path integrity. This release does not establish all of them.
Open empirical questions
- Complete mediation outside the declared v1 profile (the bypass probe shows same-user writes succeed and are only detected afterwards)
- Operating-system-level isolation and kernel assurance
- Machine-checked policy refinement and proofs
- Delegation with attenuation
- Independent security review and external red-team
- Semantic sensor assurance (not part of this release)
- Production validation
The repository implements a concrete prototype of selected execution-security components described in the framework. It does not claim that the complete conceptual specification has been implemented. What this project does not prove →
Read the evidence
Release v0.1.0 · attack corpus attack-corpus-1 · core hash 013b7ced03fac4fc…
Try to break it
The repository contains a reproducible prototype of the execution-security layer described in the paper. Attempts to falsify the protected-effect claim are welcome, especially around channel completeness, state races, policy admission, and root-of-trust assumptions. Write to canon@necessaryuniverse.com or open an issue on GitHub; SECURITY.md says what counts as a finding.