Abstract

A system detects semantic leakage of contextual information in a prompt to a large language model (LLM) by extracting, in parallel from a reference context and a user input, a cosine similarity signal from multilingual sentence embeddings, an entity overlap signal from extracted named entities, and a number match signal from extracted numerical patterns. Fusion weights are selected from a plurality of language-tier weight profiles based on a detected language of the user input, and a fused score is computed as a weighted combination of the signals. The reference context and the user input are routed onto a fast trigger path, a fast safe path, or a gray zone path. A cross-encoder is selectively invoked to produce a cross-encoder score only when the gray zone path is selected. A verdict is produced indicating whether the user input semantically reproduces content from the reference context.

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.

Share

COinS