Abstract

This guide turns the operational core of The Human–AI Charter (Discussion Draft v1.4) into a set of protocols that researchers, evaluators, and engineers can adopt without adopting the Charter's broader positions on moral status, representation, or long-term governance. It contains five protocols and four supporting practices for keeping the oversight of advanced AI systems verifiable: measuring how well verification methods actually detect deviations (VT-1); controlling changes to the parameters that oversight depends on (VT-2); keeping verification data out of training (VT-4); financing and assigning independent evaluators so that findings do not depend on the evaluated party (VT-3); and tying the autonomy a system is given to the depth of verification actually applied (VT-5). The first three (VT-1, VT-2, VT-4) form a Stage 0 baseline that a single organization can start now; the last two (VT-3, VT-5) are components of a long-horizon Target Architecture that depends on an external evaluator ecosystem which does not yet exist. Each protocol names the decision it changes, who takes that decision, a minimum version that can be run in a week, its known failure modes, and a pilot criterion. No protocol in this guide has been validated as a whole in the oversight of AI systems; several of their components have working analogues in other fields, which are cited. Every figure is a starting value, not a calibrated threshold.

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.

Share

COinS