Abstract

Large-language-model agents are routinely granted execution authority they have not earned: per-call permission prompts do not scale, allowlists do not adapt, and commercial “trust scores” are heuristic. This disclosure describes a deterministic, fail-closed execution-authority control loop in which model outputs are never authority. One model may only propose typed, exactly-argued actions; a second model may only request execution of an unmodified proposal; a host-owned controller authorizes each operation immediately before it executes; and a gateway redeeming single-use capability tokens is the only path to side effects. Four mechanisms are disclosed in enabling detail: (1) an autonomy-evidence store partitioned by a twelve-dimension key spanning both models’ versions and prompt versions, with any-change invalidation and severity-triggered suspension; (2) a selection-bias firewall excluding human-approved executions from autonomy evidence; (3) a random audit draw committed into the hash-chained authorization record before the outcome exists; and (4) capability tokens bound to the evidence snapshot and policy version that justified them. Parameter-selection guidance and a worked example are included. This document is published as an enabling disclosure so that these mechanisms are citable prior art.

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.

Share

COinS