Inventor(s)

Abstract

This paper describes a system architecture for managing high-density compute domains to address thermal and power constraints while facilitating continuous maintenance operations. The system utilizes an N plus x redundancy model for compute units and network switches, where supplemental units are managed as spares in varying low-power states. A central management controller orchestrates power steering and workload migration, aiming to adhere to designated power budgets while avoiding thermal hardware damage. This approach allows for ongoing service availability during hardware repairs and software updates without performance degradation.

Keywords: high-density compute domains, dynamic thermal management, workload migration, power steering, system redundancy, continuous maintenance, power-aware control loop.

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.

Share

COinS