Inventor(s)

Abstract

AI alignment practice sorts its interventions by when and how they are applied — prompting, in-context examples, retrieval, supervised fine-tuning, reinforcement learning from human feedback, weight editing, activation steering, ablation, guardrails — and describes what each does to a model with one undifferentiated word, influence. These interventions differ in causal type: in temporal structure, in durability, in range, in whether they act by presence or by removal, and in what it takes to undo them. This paper supplies a typed vocabulary of causal relations for that analysis, compiled from the Paṭṭhāna, the seventh book of the Theravāda Buddhist Abhidhamma, which defines twenty-four kinds of conditioning relation (paccaya). Section 3 gives a reference table of the twenty-four: each relation's canonical definition, a four-axis signature (time, mode, relation, span), and an analog in artificial agents graded by confidence. Section 4 develops six relations in depth and reports the Paṭṭhāna's own internal/external analysis: no relation by which one being's states condition another's mental states falls outside the object and decisive-support classes, so that influence between the execution traces of two agents types as parallel one-way edges. Section 5 classifies the alignment toolkit by dominant relation type and derives predictions about persistence under displacement; one is pre-registered (§6): under matched installation effect, a prompt-installed behavior will not out-persist a fine-tune-installed one. The claim is translation, not anticipation: every analog is a hypothesis, the vocabulary is not a causal-inference calculus, and the honest limits state where the compilation may be projection.

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.

Share

COinS