Abstract
The present disclosure relates to the field of Artificial Intelligence (AI), in particular to a method and system for token-budget-constrained memory retrieval for AI agents. Agent Memory-as-a-Service (AMaaS) is an infrastructure system providing AI agents with a structured, externalized memory layer that dynamically compresses, prioritizes, and retrieves episodic, semantic, and procedural memory optimized for LLM context windows. Candidate memory items may be scored using a multi-dimensional relevance function based on task-type affinity, temporal recency, access frequency, and token cost, and an optimal memory subset may be selected by solving a bounded knapsack optimization that maximizes aggregate relevance score subject to a token budget constraint, prior to each agent inference call.
Creative Commons License

This work is licensed under a Creative Commons Attribution 4.0 License.
Recommended Citation
KVS, SANTOSH KUMAR and KATAKAM, ROOPESH, "DYNAMIC TOKEN-BUDGET-AWARE MEMORY COMPRESSION AND HIERARCHICAL RETRIEVAL SYSTEM FOR AI AGENTS", Technical Disclosure Commons, ()
https://www.tdcommons.org/dpubs_series/11082