Abstract

System performance trace data may be represented as a vector to allow the trace data to be efficiently processed by a machine learning model. First, the trace data is preprocessed to generate a hierarchical topology tree that maps software events to specific threads, processes, and processors. Next, an event-driven mechanism generates tokens when a system state change occurs, collapsing idle periods into a single duration attribute to achieve lossless compression and reduce overall sequence length. Each token is then converted into a concatenated input vector that aligns discrete software events and continuous hardware metrics onto a single temporal axis. Finally, a machine learning model processes these input vectors alongside topological encodings to generate a latent embedding, which can be utilized by modular task heads to automatically identify performance bottlenecks, detect anomalies, and/or perform semantic trace searches.

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.

Share

COinS