Abstract
The present disclosure provides a method for key-value cache management in an autoregressive transformer inference system. The method includes maintaining, for each token of a plurality of tokens in a key-value cache, a rolling history of attention scores obtained over a plurality of decode steps. The method includes determining, for each token, a decay rate parameter by fitting an exponential decay model to the rolling history of attention scores. The method includes computing, for each token, an Attention Decay Residual Score based on the decay rate parameter and an elapsed interval associated with the token. The method includes classifying each token into one of a plurality of ordered decay classes based at least on the Attention Decay Residual Score, the plurality of ordered decay classes comprising a stable-anchor class, one or more intermediate decay classes, and a terminal-decay class. The method includes selecting, for each token according to the assigned decay class, a cache management action and applying the selected cache management action to a key-value entry associated with the token.
FIG. 1
Creative Commons License

This work is licensed under a Creative Commons Attribution 4.0 License.
Recommended Citation
SHIVHARE, SHUBHAM, "SYSTEM AND METHOD FOR PER-TOKEN ATTENTION DECAY RATE QUANTIFICATION, DECAY-CLASS-DRIVEN KV CACHE SEGMENT COMPRESSION AND SEMANTIC SUMMARY VECTOR GENERATION, AND MULTI-TIER CACHE RESIDENCY ARBITRATION IN AUTOREGRESSIVE TRANSFORMER INFERENCE SYSTEMS", Technical Disclosure Commons, (August 07, 2026)
https://www.tdcommons.org/dpubs_series/11299