Abstract

Power consumption is a significant constraint in the effective deployment of machine learning computing hardware. The use of eager batching for ML serving can cause ML Accelerator (MLA) system process sub-optimal batches for a majority of ML workloads. This can result in very high duty cycles and provides limited opportunities for power saving. This disclosure describes impulse batching techniques for ML serving. Impulse batching exploits peaks and troughs in the diurnal cycle. Under impulse batching, queries that arrive within a configurable duration X are grouped into larger batch sizes, especially at lower query per second (QPS) values prior to being provided to the MLA system. The duration is configurable and can be selected based on service level objectives. With larger batch sizes, the throughput of the MLA system is higher, allowing the MLA system to stay in the idle state for longer durations. This reduction in duty cycle can provide substantial power savings.

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.

Share

COinS