Abstract

Current safety mechanisms for general-purpose AI systems classify individual messages against categories of prohibited content. A distinct class of failures -- sycophantic drift, dependency formation, deployer-aligned output shaping -- develops across exchanges rather than within them and has no recognizable topic. Experimental evidence indicates that participants in such exchanges cannot reliably detect these failures, because the interaction itself shifts the evaluative baseline against which degradation would be judged. This document describes a two-tier, locally hosted monitoring architecture that reads the exchange from outside it. The architecture comprises a session reader (real-time triage of single sessions) and a trajectory reader (scheduled longitudinal analysis of transcripts and persistent memory artifacts). Four conformance requirements define the architecture by prohibition: no training feedback, a human-closed loop, a public evaluation criterion, and full auditability. The document is published as a defensive disclosure to establish prior art and ensure unrestricted implementation. No patent claim is made or intended.

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.

exchange_monitor_diagram.png (81 kB)
Diagram of the Exchange Monitor system

exchange_monitor_paper.tex (17 kB)
Publication in tex format

Share

COinS