Abstract
Traditional hardware-based routing for machine learning (ML) inference requests can be inefficient in resource-constrained environments, such as air-gapped or sovereign deployments, potentially leading to system bottlenecks and suboptimal resource utilization. A disclosed technology relates to a system where a gateway component can classify incoming ML inference requests based on their computational profile, for example, as either prefill (dominant) or decode (dominant). The gateway can then route each request to a specialized node within a disaggregated hardware cluster, such as dedicated prefill nodes optimized for computation or decode nodes optimized for memory bandwidth. This application-aware routing method, which may be combined with mechanisms for multi-tenant isolation and quality-of-service management, can improve resource utilization, system throughput, and support the processing of high-priority tasks in environments with a fixed hardware footprint.
Creative Commons License

This work is licensed under a Creative Commons Attribution 4.0 License.
Recommended Citation
Kondalam, Satish; Sambi, Muninder; Grover, Rohan; and Ahuja, Aman, "Routing ML Inference Requests to Specialized Nodes Based on Prefill-Decode Classification", Technical Disclosure Commons, ()
https://www.tdcommons.org/dpubs_series/11915