Inventor(s)

Abstract

A system integrates out-of-band physical infrastructure telemetry into software-layer workload orchestration for predictive scheduling. A multi-domain telemetry aggregator ingests and normalizes data from disparate sources, including host hardware via baseboard management controllers, facility infrastructure sensors, and network switches. A predictive reliability engine processes this aggregated data using a machine learning or heuristic model to generate a quantitative reliability score for each compute node. The engine asynchronously updates these scores in a high-speed, in-memory cache to ensure low-latency access for the scheduler. Upon receiving a new workload request, a scheduler plugin queries the cache to filter out nodes with scores below a defined threshold. The plugin then uses the reliability scores to weight and rank the remaining candidate nodes, binding the workload to the most reliable option. This approach proactively avoids at-risk hardware, reducing workload failures and extending component lifespan.

Keywords: multi-domain telemetry aggregator, predictive reliability engine, reliability score, workload placement, orchestration scheduler, in-memory cache

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.

Share

COinS