Abstract
A system integrates out-of-band physical infrastructure telemetry into software-layer workload orchestration for predictive scheduling. A multi-domain telemetry aggregator ingests and normalizes data from disparate sources, including host hardware via baseboard management controllers, facility infrastructure sensors, and network switches. A predictive reliability engine processes this aggregated data using a machine learning or heuristic model to generate a quantitative reliability score for each compute node. The engine asynchronously updates these scores in a high-speed, in-memory cache to ensure low-latency access for the scheduler. Upon receiving a new workload request, a scheduler plugin queries the cache to filter out nodes with scores below a defined threshold. The plugin then uses the reliability scores to weight and rank the remaining candidate nodes, binding the workload to the most reliable option. This approach proactively avoids at-risk hardware, reducing workload failures and extending component lifespan.
Keywords: multi-domain telemetry aggregator, predictive reliability engine, reliability score, workload placement, orchestration scheduler, in-memory cache
Creative Commons License

This work is licensed under a Creative Commons Attribution 4.0 License.
Recommended Citation
N/A, "Predictive Workload Scheduling Using Multi-Domain Telemetry and Reliability Scoring", Technical Disclosure Commons, (September 23, 2026)
https://www.tdcommons.org/dpubs_series/11839