Abstract
Systems and methods are described for caching and processing requests associated with a multi-turn agentic workflow. A proxy between a client and a large language model (LLM) API receives a request associated with a turn of the multi-turn agentic workflow and determines a cache key based on input information associated with the turn and an index of the turn. The proxy determines whether a cache entry corresponding to the cache key is valid based on a dependency of the cache entry on one or more external resources and, when valid, provides a cached response to the client. The proxy further determines, based on a sequence of tool calls associated with the multi-turn agentic workflow, a likelihood of a subsequent turn. In response to the likelihood satisfying a threshold, the proxy speculatively obtains a response for the subsequent turn from the LLM API before receiving a corresponding request from the client.
Creative Commons License

This work is licensed under a Creative Commons Attribution 4.0 License.
Recommended Citation
BASKARAN, RAJKUMAR; BASU, ISHANI; and DESAI, SUDHINDRA, "DEPENDENCY-AWARE SPECULATIVE PROXY CACHE FOR MULTI-TURN LLM AGENTIC WORKFLOWS", Technical Disclosure Commons, ()
https://www.tdcommons.org/dpubs_series/12041