Abstract
We disclose an architecture for AI alignment and corrigibility in which the constraints that make an AI agent safe are specified as constituents of the agent's identity rather than as external constraints on an optimizer. The agent is modelled as four functional registers (a value substrate; binding to the people it serves; active practice; communicative reach) bound into one agent by an integrating register that is neither a further module nor a central observer, on the pattern of global-workspace integration. A completeness check tests a proposed agent for coverage across the registers rather than for capabilities, so a system passing every component-level audit can still be an unintegrated assemblage. A counterfactual removal test classifies each constraint as constitutive (removing it yields a different agent, or none) or adversarial (removing it yields the same agent, unconstrained); an aligned agent is built as far as possible from constitutive constraints, with an external override present but minimized. Restrainers are mapped one per register and classified by locus; the human-held override is the single external active one. An oversight schedule narrows that override only as fast as internalized constraint is demonstrated, and never to zero, because an internal constraint cannot detect its own corruption. The design is contrasted with utility indifference, the off-switch game and cooperative inverse reinforcement learning. The Theravāda Buddhist analysis of a person as a conventional designation supports a claim of structural personhood without a claim of substantial self; the paper states where that analysis limits the claim.
Creative Commons License

This work is licensed under a Creative Commons Attribution 4.0 License.
Recommended Citation
Ly, Thon, "Constituting an Artificial Person: Restraint as Constitution, and the Elemental Completeness of an Aligned Mind — A Five-Register AI Agent Architecture with Alignment Constraints as Constituents of Agent Identity, a Counterfactual Removal Test for Corrigibility, and a Human Override That Never Reaches Zero", Technical Disclosure Commons, ()
https://www.tdcommons.org/dpubs_series/11954