Chapter 7
A Graph-Based Approach for Mapping Kernel-Level Telemetry to MITRE ATT&CK
Why it matters
The cybersecurity knowledge graph in the thesis only earns its place if it can be fed by what actually happens on a machine, not just by reports written months later. This paper builds exactly that bridge: kernel events from a Linux host become a provenance graph, the graph is shrunk into a compact command graph, and a locally run language model maps it to MITRE ATT&CK techniques with a written rationale. That is the home-edge sovereign node's local AI plane doing security work without sending household telemetry to anyone's cloud, and it gives the cyber-KG a live, first-hand ingestion path (observed behavior → technique node) to sit beside the report-driven graphs we studied earlier (TRACE, BEACON).
What it proposes
Trace2ATT&CK, an end-to-end pipeline that starts from runtime evidence instead of threat-intel reports. It assumes an attack window has already been flagged (anomaly detection is out of scope) and focuses on explaining it: for each attacker session it returns a ranked top-5 list of ATT&CK techniques and sub-techniques, each with a confidence score and a natural-language justification. Three design goals drive it: context over isolated events, fully local execution (local models, embeddings and vector store), and explanations an analyst can check.
How it works
eBPF (via the Tracee tool) captures system calls with low overhead; a Rust preprocessor filters and chunks the stream. The pipeline finds the attacker's shell, walks its process tree, correlates each typed command to the kernel events it caused, and stores the result in Neo4j. It deliberately keeps only process and command nodes (files and sockets live inside command arguments) and then builds a weighted command graph where repeated commands carry a count, with a compressed variant for repetitive attacks. The graph is serialized as DOT text and classified three ways: plain prompting, prompting with the full ATT&CK ID-to-name table injected (to stop models pairing the right name with the wrong ID), and RAG over the ATT&CK knowledge base in a local Chroma store.
Evaluation used 347 Linux Atomic Red Team tests (about 62K preprocessed telemetry lines per scenario) across seven local open-weight models. RAG beat plain prompting in all 14 model/level combinations. The best local result was gpt-oss-120b with RAG at 66.57% technique-level HR@5 and 46.26% exact sub-technique HR@5. A security-specialized 8B model (foundation-sec-8b-reasoning) outscored llama-3.3-70b under base prompting (49.2% vs 35.26% technique HR@5). In the ablation, raw logs dropped to 37.88% technique HR@5 with empty outputs in over half the cases, while the full provenance graph scored higher when it fit but returned nothing in up to 49.76% of cases; the compact command graph never overflowed the 32K context. A frontier hosted model reached 90.91% technique HR@5, which the authors read as validation of the method, not as the deployment target.
Limits
The attack window is given, not detected. The trusted base includes the kernel and audit layer, so kernel-level attackers who tamper with telemetry, and hardware or side-channel attacks, are out of scope. Ground truth comes from scripted Atomic Red Team tests on Linux, not live intrusions, and about half the tests failed at execution; sub-technique accuracy drops sharply on those interrupted runs (for example 60.36% vs 31.07% HR@5 for the best RAG setup). Local models still trail the hosted one by roughly 24 points at technique level.
What we'd test
- Run the same eBPF → command-graph → local-RAG loop on a small home-edge box (one Linux node behind the access demark) and measure how many technique labels a modest local model gets right versus a home-sized hardware budget.
- Write each mapped technique back as an edge into our cyber-KG (host session → ATT&CK technique → affected asset or wallet hook), so the graph links observed attacks to what they threaten economically.
- Put the trust broker in front of it: telemetry stays in the on-prem plane, and only signed, technique-level summaries leave it. That makes "local inference without exposing household data" a measurable claim for the Trustworthy Agentic Operations paper.