Abstract
System performance trace data may be represented as a vector to allow the trace data to be efficiently processed by a machine learning model. First, the trace data is preprocessed to generate a hierarchical topology tree that maps software events to specific threads, processes, and processors. Next, an event-driven mechanism generates tokens when a system state change occurs, collapsing idle periods into a single duration attribute to achieve lossless compression and reduce overall sequence length. Each token is then converted into a concatenated input vector that aligns discrete software events and continuous hardware metrics onto a single temporal axis. Finally, a machine learning model processes these input vectors alongside topological encodings to generate a latent embedding, which can be utilized by modular task heads to automatically identify performance bottlenecks, detect anomalies, and/or perform semantic trace searches.
Creative Commons License

This work is licensed under a Creative Commons Attribution 4.0 License.
Recommended Citation
Suggu, Sai Akhil; Adarsh, Anurag; Gupta, Somya; Ezeozue, Zimuzo; and Desai, Dwipal, "Representation of Processor Traces for Machine Learning Models", Technical Disclosure Commons, (September 16, 2026)
https://www.tdcommons.org/dpubs_series/11754