Abstract

This publication describes the Security Operations (SecOps) Decision Attribution and Reasoning Trace framework, an integrated ground-truth construction pipeline and evaluation harness for benchmarking autonomous security triage and investigation agents at enterprise scale. Evaluating security agents presents two fundamental challenges: first, historical human closure notes are inherently sparse, unstructured, and omit implicit investigative steps, making it difficult to systematically codify expert reasoning ground truth. Secondly, conventional verdict-based metrics fail to detect when an agent arrives at the correct conclusion through flawed logic, coincidental data lookups, or deficient reasoning paths. The SecOps Decision Attribution and Reasoning Trace framework aims to resolve both challenges. As a ground-truth pipeline, it utilizes the human resolution verdict as an anchor and leverages a language model to reverse-engineer sparse closure notes against alert telemetry, reconstructing latent decision-critical attributes, threat enrichments, and out-of-band context gaps while sanitizing any private or identifying information to ensure there are no biases. Concurrently, as an evaluation harness, this framework parses the agent’s final summary and execution footprint backwards to isolate the attributes the agent actually relied upon. A hierarchical scoring engine benchmarks the agent’s reasoning trace against the extracted human baseline, computing a primary attribute-level reasoning overlap score via multi-layer string, structural, and zero-temperature semantic matching, alongside multi-facet taxonomy evaluations for root-cause alignment, risk-action congruence, and tool-level signal-to-noise ratios. Evaluation failures and context gaps are materialized into machine-readable data objects that populate a domain knowledge base, which is exposed back to the autonomous agent suite as executable skills to enable continuous operational self-learning and improvement. . Keywords: security operations center, autonomous agent evaluation, reasoning trace, decision attributes, context gap, signal to noise ratio, unified data model, threat intelligence enrichment, root cause alignment, risk action congruence.

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.

Share

COinS