Abstract
A system detects semantic leakage of contextual information in a prompt to a large language model (LLM) by extracting, in parallel from a reference context and a user input, a cosine similarity signal from multilingual sentence embeddings, an entity overlap signal from extracted named entities, and a number match signal from extracted numerical patterns. Fusion weights are selected from a plurality of language-tier weight profiles based on a detected language of the user input, and a fused score is computed as a weighted combination of the signals. The reference context and the user input are routed onto a fast trigger path, a fast safe path, or a gray zone path. A cross-encoder is selectively invoked to produce a cross-encoder score only when the gray zone path is selected. A verdict is produced indicating whether the user input semantically reproduces content from the reference context.
Creative Commons License

This work is licensed under a Creative Commons Attribution 4.0 License.
Recommended Citation
Saha, Kaustav Saha; Raghuvanshi, Shwetank; Haider, Batool; Andrea, Kliton; Hu, Chenhui; Srinivasan, Subramanian; Kesireddy, Ashwin; and Zscaler, Inc., "Multi-Signal Fusion Architecture for Semantic Leak Detection in Large Language Model Inputs", Technical Disclosure Commons, ()
https://www.tdcommons.org/dpubs_series/11476