Abstract
A challenge in deploying generative models can be the occurrence of factual inaccuracies, particularly when attempting to maintain consistent reliability across diverse information categories. A disclosed technology describes a multi-agent pipeline that can quantify and calibrate the factual reliability of generated content. The system may operate as a black-box process where specialized agents can parse source material, generate categorized claims, and score their correctness. An evaluator agent can then apply category-specific conformal inference to filter claims that may not meet a defined statistical confidence level. This approach can provide a mechanism to manage the factuality of model outputs by providing statistically supported confidence estimations for generated claims, including those within less common, long-tail categories, without requiring internal model access.
Creative Commons License

This work is licensed under a Creative Commons Attribution 4.0 License.
Recommended Citation
Gao, Chenyin; Yu, Jiaqian; and Rao, Aniruddha, "Multi-Agent Pipeline for Category-Specific Factuality Calibration Using Conformal Inference", Technical Disclosure Commons, (August 27, 2026)
https://www.tdcommons.org/dpubs_series/11516