Abstract
This paper describes a system architecture for managing high-density compute domains to address thermal and power constraints while facilitating continuous maintenance operations. The system utilizes an N plus x redundancy model for compute units and network switches, where supplemental units are managed as spares in varying low-power states. A central management controller orchestrates power steering and workload migration, aiming to adhere to designated power budgets while avoiding thermal hardware damage. This approach allows for ongoing service availability during hardware repairs and software updates without performance degradation.
Keywords: high-density compute domains, dynamic thermal management, workload migration, power steering, system redundancy, continuous maintenance, power-aware control loop.
Creative Commons License

This work is licensed under a Creative Commons Attribution 4.0 License.
Recommended Citation
N/A, "Domain N plus x Redundancy and Dynamic Thermal Management for High-Density Compute Domains", Technical Disclosure Commons, (August 14, 2026)
https://www.tdcommons.org/dpubs_series/11364