Abstract
Standard context caching for large language models can be inefficient in workflows with large, multi-part contexts, such as clinical documentation, where streaming data can invalidate static records and individual records may not meet caching thresholds. A computer-implemented system and method for optimizing context caching are described. The system may utilize a hierarchical structure that segments input context into static, semi-static, and dynamic layers based on update frequency. The system may also proactively aggregate historical records from multiple patients into a single cache prefix to meet service provider token minimums. This approach of hierarchical segmentation and predictive aggregation can reduce cache invalidations caused by streaming data, which may decrease redundant processing, lower latency, and reduce costs in some generative artificial intelligence applications.
Creative Commons License

This work is licensed under a Creative Commons Attribution 4.0 License.
Recommended Citation
Joshi, Sarita A.; Gustas, Leo J.; and Clark, Blane C., "Hierarchical Context Caching Using Data Segmentation and Predictive Aggregation", Technical Disclosure Commons, ()
https://www.tdcommons.org/dpubs_series/11384