Glass-Box Governor

Paper · framework specification · independent research

The Dialectical Cage and the Glass-Box Governor

A Practice-Based Ethics, Safety Constitution, and Reference-Monitor Architecture for AI Agents

Paper scope versus prototype scope

The paper specifies a broader framework and a conditional execution theorem. The current repository demonstrates only the explicitly enumerated prototype profile and components. It does not claim that the complete conceptual specification has been implemented, and the paper itself lists machine-checked proofs, implementation, red-teaming and production evidence as release conditions.

Abstract

This paper specifies a layered framework for constraining the actions of AI agents without treating the agent’s own reasoning as the final safety boundary. The framework separates four questions that are often conflated. First, a practice-based ethics analyzes addressed justification, responsiveness, and answerability, deriving the public-ground component of L1 and the no-unjustified-difference result (NRD), while explicitly declaring the substantive standards that the selected constitution adds. Second, a constitutional layer records those substantive standards as a versioned, inspectable safety constitution rather than presenting them as consequences of Identity alone. Third, a safety-engineering layer translates that constitution into a machine-readable safety ontology, a deployment-specific safety profile, a typed policy representation, and proof obligations. Fourth, a reference-monitor layer mediates protected effects through deterministic policy evaluation, principal-bound capabilities, execution-time state validation, revocation, and tamper-evident evidence.

The central engineering claim is conditional rather than absolute. For a declared deployment profile, let protected effects have a defined state-transition safety predicate Safe(s, a, s′ ). If complete mediation, trusted-core integrity, kernel conformance, policy refinement, principal attribution, capability non-transferability, atomic execution-time revalidation, fail-closed handling of unresolved safety state, and evidence-path integrity all hold, then every executed protected effect satisfies the admitted safety invariant. The proof does not depend on the model being aligned, benevolent, rational, or normatively committed; it depends on the reference monitor and its stated assurance contract. The paper therefore distinguishes mechanical execution safety from semantic assurance: neural state extraction may be probabilistic, and a semantic monitor may report a measured error or coverage bound, but non-zero semantic uncertainty is never silently promoted into a categorical guarantee for a high-consequence effect.

The philosophical contribution is deliberately narrower than moral realism. Identity does not entail a unique substantive ethics. Constitutional reflexivity (CR/CAR) constrains self-authenticating constitutional authority within the specified reason-giving practice, but it does not determine all substantive values. The selected agency-complete and defensive constitution is presented as a declared constitution whose reasons, dependencies, rejection costs, and extensions are visible. The runtime machinery is constitution-parametric: another constitution can be instantiated provided its substantive safety requirements are represented in the same formal ontology and admitted through the same proof and governance pipeline.

The system is designed for enterprise adoption rather than rhetorical invulnerability. The paper defines deployment profiles, an explicit residual-risk budget, a SafetyBundle that binds the constitution, ontology, policy, kernel, sensor contract, model digest, compiler, and deployment profile, monotone safety-update rules, independent approval requirements, emergency-containment rules, and release gates. It also states what is not established: no universal channel-completeness theorem, no claim that all harmful speech is categorically prevented, no claim that natural-language policy compilation is automatically semantically correct, no empirical zero-failure result, and no claim that an AI system becomes a direct bearer of the philosophy merely because the philosophy is encoded in its governor. Implementation, machine-checked proofs, independent red-teaming, and production evidence are release conditions, not implied accomplishments.

Implementation status of this release

Release v0.1.0 implements selected execution-security components for one profile (Architecture). In the claim registry, SPECIFIED marks what the paper states and this repository does not test, LOCALLY TESTED what the author's corpus exercises, and OPEN what is not established. The philosophical derivations (L1, NRD, CR/CAR) are not implemented or machine-checked here.

The paper's epistemic labels

These are labels for the kind of claim, not confidence scores.

D
Derivation relative to explicit premises
A
Declared normative / deployment assumption
S
Security or invariant claim dependent on trusted-system conditions
E
Empirical assurance claim requiring measurement
R
Residual / open region

Verify the files you downloaded

shell
sha256sum the-dialectical-cage-and-the-glass-box-governor.pdf
# expected: 0fb34d207f97e01deaedbb8325f36b070b437a815441eae142af0bccfa35d2a2
sha256sum the-dialectical-cage-and-the-glass-box-governor.tex
# expected: bd4c2dedde78d830a533dea41074d3349e28941506a5c703305efa505df1e12d

Hashes come from manifest.json. The mapping from public claims to paper sections is in SOURCE_MAP.md.

Cite this paper

@misc{CanonGlassBox,
  author       = {Canon},
  title        = {{The Dialectical Cage and the Glass-Box Governor: A Practice-Based Ethics, Safety Constitution, and Reference-Monitor Architecture for AI Agents}},
  year         = {2026},
  publisher    = {Zenodo},
  doi          = {10.5281/zenodo.21278446},
  url          = {https://doi.org/10.5281/zenodo.21278446},
  note         = {Independent research; not peer-reviewed. Code: https://github.com/canon4d/the-dialectical-cage-and-the-glass-box-governor}
}

Citing software instead? The repository carries CITATION.cff.