Abstract

Traditional hardware-based routing for machine learning (ML) inference requests can be inefficient in resource-constrained environments, such as air-gapped or sovereign deployments, potentially leading to system bottlenecks and suboptimal resource utilization. A disclosed technology relates to a system where a gateway component can classify incoming ML inference requests based on their computational profile, for example, as either prefill (dominant) or decode (dominant). The gateway can then route each request to a specialized node within a disaggregated hardware cluster, such as dedicated prefill nodes optimized for computation or decode nodes optimized for memory bandwidth. This application-aware routing method, which may be combined with mechanisms for multi-tenant isolation and quality-of-service management, can improve resource utilization, system throughput, and support the processing of high-priority tasks in environments with a fixed hardware footprint.

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.

Share

COinS