Abstract

A system and method for managing intermediate computational states in a distributed artificial intelligence pipeline are disclosed. For each intermediate computational state generated during execution of the pipeline, a recomputation difficulty score is determined based on a plurality of factors comprising a computational cost of regenerating the intermediate computational state, a pipeline latency associated with regeneration, a depth of an upstream dependency chain, a semantic importance of the intermediate computational state, and a likelihood of subsequent reuse. The recomputation difficulty score is mapped to one of a plurality of difficulty classifications, and a storage placement action is selected based on the mapped difficulty classification. The storage placement action comprises retaining the intermediate computational state in a first memory tier, offloading the intermediate computational state to a second memory tier, storing the intermediate computational state in a remote storage tier, compressing and storing the intermediate computational state, or releasing the intermediate computational state for subsequent recomputation using a recorded upstream dependency path. In response to a memory-pressure condition, eviction candidates are ordered according to respective recomputation difficulty scores such that an intermediate computational state having a lower recomputation difficulty is released before an intermediate computational state having a higher recomputation difficulty.

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.

Share

COinS