Inventor(s)

Abstract

Large language models can read a payer medical-policy document and produce structured coverage criteria that look right. Nothing in the model’s output distinguishes the criteria that are right from the criteria that are invented, and in clinical revenue-cycle use an invented criterion is acted on. We describe an architecture in which an LLM is treated as an untrusted proposer whose output becomes queryable only after a deterministic, separately built verifier has resolved every claim to a verbatim span in a hash-pinned rendering of the source document. Every served criterion carries that span, called a warrant, so any consumer can check any claim against the captured source bytes without trusting the extractor, the operator, or the model vendor. We define three evaluation metrics for this class of system (grounding rate, content recall, and exact-span rate) and report our own numbers, including the one that plateaued. The architecture is disclosed so that it cannot be enclosed; the measurements exist so that “the extraction is good” can be a claim with content.

keywords:

large language models, information extraction, verification, medical policy, coverage criteria, prior authorization, claims adjudication support, provenance, defensive publication

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.

Share

COinS