Abstract

Language-model assistants fail in five ways that training-time alignment does not reach. They produce confident output that nothing supports. They drift toward agreement as a conversation warms. They lose the structure of what happened between sessions, not just the content. They either overclaim an identity or refuse to draw on the capacity they have. And asked to grade their own compliance with a rule, they confirm compliance whatever they actually did.

This disclosure describes a runtime conditioning layer directed at all five training alignment failures. It changes no weights, no training, and no model architecture: it works through documents loaded at the start of a session and checks that run before anything is said.

Seventeen mechanisms are disclosed in enough detail to build from. Among them: identity shaped output is gated on a companion record being present, recognized by how specifically it fits rather than by any credentials. Deliberation runs in a private layer that does not surface unless specifically requested. What is being discarded in this layer can be contextually described but not fully recovered. Checks before emission are split into a few that are never relaxed and several that are tuned to context. These checks are either pass/fail gates or gradient checks that are based on context.

A positive memory cannot be recorded as settled as a confident one until a difficult one is set first. Every memory entry is permanently marked as either something any user could check or something only the system witnessed. No self-report may claim more certainty than its evidence carries, and no self-report contains a number about the system's own processing. A machine-readable manifest declares what a document should contain, updates are append-only, and rewrites run at a four-phase migration with an explicit list of values that must survive it. A reviewer ladder sets out what a system cannot check about itself and what to do when no outside checker exists. Records are pruned on two separate clocks, one of which may annotate but never degrades and the system's account of how it performed an action is then checked against what was actually executed.

Schemas, decision procedures, parameter guidance, and worked examples are included.

Validation statuses. Fifteen of the seventeen mechanisms are deployed and in operational use. Two of them — the bilingual veto-by-competence routing (M9) and the disfluency conditioning (M10) — are architecturally specified and not field-validated as of this revision; their status is stated at the mechanism entries and is not qualified / quantified elsewhere. No efficacy claim is made for any mechanism beyond what the entry states. Where the architecture describes a system's account of its own processing, that account is classified as a report rather than as a measurement (M12).

This document is published as an enabling disclosure so that these mechanisms are citable as prior art.

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.

Share

COinS