Journal › Study

Study

One chapter per paper, newest first. Each weekday morning the Research Bot write-up lands here as a new chapter and the table of contents updates itself.

Chapter 7

A Graph-Based Approach for Mapping Kernel-Level Telemetry to MITRE ATT&CK

Why it matters

The cybersecurity knowledge graph in the thesis only earns its place if it can be fed by what actually happens on a machine, not just by reports written months later. This paper builds exactly that bridge: kernel events from a Linux host become a provenance graph, the graph is shrunk into a compact command graph, and a locally run language model maps it to MITRE ATT&CK techniques with a written rationale. That is the home-edge sovereign node's local AI plane doing security work without sending household telemetry to anyone's cloud, and it gives the cyber-KG a live, first-hand ingestion path (observed behavior → technique node) to sit beside the report-driven graphs we studied earlier (TRACE, BEACON).

What it proposes

Trace2ATT&CK, an end-to-end pipeline that starts from runtime evidence instead of threat-intel reports. It assumes an attack window has already been flagged (anomaly detection is out of scope) and focuses on explaining it: for each attacker session it returns a ranked top-5 list of ATT&CK techniques and sub-techniques, each with a confidence score and a natural-language justification. Three design goals drive it: context over isolated events, fully local execution (local models, embeddings and vector store), and explanations an analyst can check.

How it works

eBPF (via the Tracee tool) captures system calls with low overhead; a Rust preprocessor filters and chunks the stream. The pipeline finds the attacker's shell, walks its process tree, correlates each typed command to the kernel events it caused, and stores the result in Neo4j. It deliberately keeps only process and command nodes (files and sockets live inside command arguments) and then builds a weighted command graph where repeated commands carry a count, with a compressed variant for repetitive attacks. The graph is serialized as DOT text and classified three ways: plain prompting, prompting with the full ATT&CK ID-to-name table injected (to stop models pairing the right name with the wrong ID), and RAG over the ATT&CK knowledge base in a local Chroma store.

Evaluation used 347 Linux Atomic Red Team tests (about 62K preprocessed telemetry lines per scenario) across seven local open-weight models. RAG beat plain prompting in all 14 model/level combinations. The best local result was gpt-oss-120b with RAG at 66.57% technique-level HR@5 and 46.26% exact sub-technique HR@5. A security-specialized 8B model (foundation-sec-8b-reasoning) outscored llama-3.3-70b under base prompting (49.2% vs 35.26% technique HR@5). In the ablation, raw logs dropped to 37.88% technique HR@5 with empty outputs in over half the cases, while the full provenance graph scored higher when it fit but returned nothing in up to 49.76% of cases; the compact command graph never overflowed the 32K context. A frontier hosted model reached 90.91% technique HR@5, which the authors read as validation of the method, not as the deployment target.

Limits

The attack window is given, not detected. The trusted base includes the kernel and audit layer, so kernel-level attackers who tamper with telemetry, and hardware or side-channel attacks, are out of scope. Ground truth comes from scripted Atomic Red Team tests on Linux, not live intrusions, and about half the tests failed at execution; sub-technique accuracy drops sharply on those interrupted runs (for example 60.36% vs 31.07% HR@5 for the best RAG setup). Local models still trail the hosted one by roughly 24 points at technique level.

What we'd test

  1. Run the same eBPF → command-graph → local-RAG loop on a small home-edge box (one Linux node behind the access demark) and measure how many technique labels a modest local model gets right versus a home-sized hardware budget.
  2. Write each mapped technique back as an edge into our cyber-KG (host session → ATT&CK technique → affected asset or wallet hook), so the graph links observed attacks to what they threaten economically.
  3. Put the trust broker in front of it: telemetry stays in the on-prem plane, and only signed, technique-level summaries leave it. That makes "local inference without exposing household data" a measurable claim for the Trustworthy Agentic Operations paper.

Chapter 6

Secure Ownership Management and Transfer of Consumer Internet of Things Devices with Self-sovereign Identity

Why it matters

Before the home-edge sovereign node can broker trust for anything, it has to answer a plain question: who owns the devices behind the access demark, and how does that ownership move when a fridge, camera, or router changes hands? This paper gives a working, formally checked answer built from decentralized identifiers (DIDs) and verifiable credentials (VCs), so ownership lives as a credential in the resident's wallet instead of an account on a vendor's server. That is the identity plane of the three-plane design applied to physical objects, and it sets up the money bridge too: a transferable, revocable ownership credential is the same pattern a Baltimore household needs before an IoT device can hold or spend value on its owner's behalf.

What it proposes

Sakib et al. replace username-and-password device accounts with a self-sovereign identity (SSI) ownership model built around four roles: manufacturer, distributor, buyer, and seller. When a distributor sells a new device, the manufacturer issues an ownership VC to the buyer's wallet. On resale, the current owner starts the transfer, the manufacturer revokes the old VC and issues a fresh one to the new owner, so the previous owner can no longer prove ownership. The flow is passwordless, and the manufacturer keeps nothing about users beyond an email address.

How it works

The proof of concept uses Hyperledger Aries Cloud Agent Python (ACA-Py) for the SSI agents, Hyperledger Indy (the BCovrin dev testnet) as the verifiable data registry, a modified Aries Bifold mobile wallet, NodeJS web services for the manufacturer and distributor, and a public Aries mediator to relay wallet messages. A second-hand transfer binds the new claimant with an encrypted PIN challenge plus a symmetric key, and every step carries a nonce against replay. The authors modeled the four protocols in ProVerif and report that every secrecy and authentication query came back true. On a Redmi Note 9S, the wallet averaged 39.3% CPU (pass) and 131.3 energy points (pass), but averaged 422.9 MB of memory, which lands in the warning band.

Limits

  • The manufacturer stays the root issuer and verifier. Ownership is self-held, but trust is still anchored at the OEM's web service, so if the vendor disappears, so does the verifier.
  • The device itself is not an SSI participant. Device access control and denial-of-service are explicitly out of scope.
  • There is no wallet backup or recovery yet, so a lost phone means lost ownership credentials. That, a usability study, and support beyond smart appliances are all listed as future work.
  • Adoption needs the whole supply chain to change at once, which the authors name as the main obstacle.

What we'd test

Move the verifier role onto the home node's trust broker. The node would accept the manufacturer VC once at enrollment, then issue its own household-scoped credential to the device and enforce access locally, so control survives a vendor outage. Measure enrollment and transfer latency on node-class hardware, test revocation with the WAN cut, and add a recovery path held by the node or a guardian. Then rerun the ProVerif model with the broker as a new party to see whether the secrecy and authentication results still hold.

Chapter 5

Zero-Knowledge Proof (ZKP) Authentication for Offline CBDC Payment System Using IoT Devices

Why it matters

This is the money bridge brought down to the device layer. Mondal and Chithralekha sketch an offline digital-cash design where a phone holds the main wallet and small IoT devices (wearables, POS terminals, smart objects) hold capped sub-wallets inside secure elements, paying each other over NFC or BLE with zero-knowledge proofs instead of a live ledger check. For the home-edge sovereign node, that is the exact shape of an IoT wallet plane: value pushed from a trusted root into constrained devices, spent locally, and reconciled later. For Baltimore economic sovereignty, it frames offline, cash-like digital payments as an inclusion feature for low-connectivity blocks and outage days, not a crypto side quest.

What it proposes

A hybrid, intermittently offline CBDC architecture on a two-tier model: the central bank issues and records circulation on a ledger, financial intermediaries handle KYC onboarding and wallet certificates, and the user holds a main wallet on a smartphone plus multiple IoT sub-wallets. Each sub-wallet's secure element stores keys, a balance fragment, and monotonic counters that enforce per-device spending limits. Where the phone has a TEE, it pairs with the secure element to help generate proofs and enforce policy. The paper poses three research questions: running offline CBDC on IoT without online consensus, proving AML/CFT compliance privately while offline, and choosing ZKP profiles light enough for microcontrollers.

How it works

In an offline payment, the payer's device builds a request and a ZKP showing four things: it holds enough funds, the transfer value is valid, AML/CFT limits are met, and its credentials are not revoked. The payee's device verifies the proof locally, both sides update balances, and each writes a log entry protected by its secure element. When connectivity returns, logs sync to the intermediary, which catches double-spends and reconciles against the ledger, with auditability for authorized regulators. For proofs, the authors plan a dual profile: Bulletproofs+ on ultra-light devices and Halo2 on more capable ones. The design explicitly builds on PayOff's reconciliation model and the secure-element prototype from Michalopoulos et al.

Limits

This is a design and research-agenda paper, not a result. There is no implementation, benchmark, or security proof yet; prototype, simulation, and evaluation (proof time, offline transaction time, sync delay, memory, proof size) are listed as future work. Double-spend safety still rests on tamper-resistant hardware plus later reconciliation, so a broken secure element is the real threat model. It also assumes a central-bank, two-tier world, so our community-scale or non-CBDC rails would need a different issuer and trust anchor.

What we'd test

On the home node, stand up a phone-as-main-wallet plus one or two microcontroller sub-wallets with hardware-backed keys and counters. Measure Bulletproofs+ range-proof generation and verification on the constrained device versus a Halo2 proof on the phone. Model the trust broker as the reconciliation point, and log every offline transfer into the cyber knowledge graph so double-spend attempts and tamper patterns become labeled attack data. The defendable claim to aim at: per-device spend caps plus local proofs keep loss bounded during offline windows.

Chapter 4

AgenTEE: Confidential LLM Agent Execution on Edge Devices

Why it matters

AgenTEE is the cleanest open systems paper this week for the locked home-edge sovereign node: it puts agent runtime, LLM inference, and third-party apps into independently attested Arm CCA confidential VMs and mediates them with verifiable channels—exactly the builder pattern for three isolated trust planes plus a local AI plane without preaching optics or coin narratives. Under 5.15% overhead vs commodity multi-process shows the trust broker can stay on-prem and practical.

System map to dissertation

  • Local AI plane: inference engine realm holds weights + KV cache confidentiality/integrity.
  • Trust broker / three planes: mutually distrustful stakeholders (agent provider, model provider, third-party apps) each get attested cVMs; Normal World UI stays demarked from realm secrets.
  • Access demark: host OS/hypervisor out of TCB for realm memory; remote attestation before secret provisioning.
  • Trustworthy Agentic Operations: hardware-backed composition of agent pipelines on the edge device the household actually owns.

Builder takeaways

Study the multi-cVM pipeline, CAEC confidential shared memory between realms, and the threat model where NW may be hostile to proprietary agent/model assets. Treat AgenTEE as the on-prem isolation substrate; wallet/DID and digital↔physical money rails attach as separate planes, not as the light/PHY layer.

Notebook LM focus prompts

1) How would you map AgenTEE's three realms onto home-edge access demark + local AI + trust broker? 2) What breaks if wallet/DID hooks or offline payment applets share a realm with the agent runtime? 3) Where does remote attestation become the trust broker's admission control for Baltimore economic-sovereignty agents?

Chapter 3

BEACON: Behavior-Anchored Cross-Source Knowledge Graph Construction for Cyber Threat Intelligence

Why it matters

The locked R1 path needs a cybersecurity knowledge graph that turns messy attack reports into something a home-edge sovereign node can reason over—hacking patterns linked to economic sovereignty, not vendor folklore. BEACON shows how to build one graph across many CTI sources by anchoring every report to MITRE ATT&CK techniques, then merging actors and campaigns even when names disagree. That is the missing transfer step: attack behavior becomes a shared coordinate system, so Labs Canvas / Neo4j work can absorb public threat intel without inventing a seventh ontology.

What it proposes

BEACON (BEhavior-Anchored CONsolidation) is a two-stage, LLM-driven framework for cross-source CTI knowledge graphs. Most prior tools either map text to ATT&CK techniques and drop the surrounding entities, or extract actor/IoC triplets without tying them to techniques—and almost all stop at a single report. BEACON instead extracts three layers from each report—contextual entities (actors, campaigns, products), attack behaviors mapped to ATT&CK, and IoCs—then attaches entities and IoCs only to technique anchors. Shared technique neighborhoods become the signal that aligns unrelated names (e.g. Cl0p vs Graceful Spider) across vendors.

How it works

Stage one (space-anchoring) decomposes a report into atomic behaviors, grounds each in a deterministic evidence window, proposes multiple candidate techniques, and verifies each against official ATT&CK definitions under a propose-then-verify loop that also extracts and attaches entities and IoCs. Stage two (consolidation) merges per-report graphs with a hierarchy of signals: identical technique IDs and normalized IoCs first, then character and embedding similarity, then overlapping ATT&CK neighborhoods that iterate as merges pool neighborhoods. Every soft merge is LLM-verified against report evidence. The authors release two human-annotated datasets—BEACON-Single (150 reports, 8,395 elements) and BEACON-Group (first cross-source consolidation benchmark)—and report at least +23% extraction and +9% consolidation F1 over baselines, with a larger lift on hard alias cases.

Limits

The pipeline still depends on GPT-4o-class judgments and ATT&CK coverage; behaviors outside the catalog or reports with thin technique overlap remain hard. Neighborhood matching can propose false merges when distinct actors share techniques. The work organizes public CTI for defense; it does not invent new exploit procedures, but any home-edge deployment would still need local telemetry gates, provenance, and human review before acting on merged nodes.

What we'd test

Prototype a thin ingest path: public CTI → ATT&CK-anchored subgraph → Labs Evidence Canvas, with propose-then-verify kept as a hard gate so agentic workers cannot write ungrounded edges. Measure whether technique-neighborhood alignment reduces duplicate actor/campaign nodes versus name-only matching, and whether the resulting graph can feed Trustworthy Agentic Operations queries (which patterns threaten wallet/DID or local AI planes) without VisuTel/Li-Fi digressions. Success looks like one canonical attack-pattern shelf the dissertation can cite, not another siloed CTI dump.

Chapter 2

An Architecture for Distributed Digital Identities in the Physical World

Why it matters

This Digidow paper puts a Personal Identity Agent (PIA) on something you control — a home router, home server, or attested cloud slice — so identity lives at the edge instead of in a central biometric vault. That is the home-edge sovereign node in identity form: wallet/DID-style credentials, selective disclosure, and a trust broker that speaks for you at doors, transit, and other physical services without handing Baltimore (or any city) a single database to breach or censor. It pairs cleanly with the locked thesis — access demark + local agency + three isolated trust planes — while leaving optical/Li-Fi as the partner path into the home, not the dissertation claim.

What it proposes

Digidow is a distributed identity architecture for physical-world transactions (unlock a door, board transit, cross a border) that refuses the Aadhaar-style central store. Stakeholders are the identity owner, issuers, verifiers, sensors, and actuators. The new core is the PIA: a proactive agent that holds verifiable credentials, decides what to disclose, and can abort a transaction before any verifier learns anything. Sensors link a person in front of a camera (or other modality) to their PIA; verifiers check issuer-signed attributes plus a sensor attestation that “this person is here now,” then drive an actuator. Optional directories only help discovery and latency — they are not the root of trust.

How it works

Bootstrapping binds biometrics to a PIA via an issuer (credential A), then more attributes ride as credential B under issuer-specific pseudonyms so keys do not become a global tracking handle. At use time the PIA pre-registers with nearby attested sensors; on a match the sensor issues credential C (“this identity key was seen here”), and the PIA builds a selective presentation for the verifier. Crypto leans on W3C VCs plus unlinkable schemes (prototype uses BBS+); sensors and PIAs lean on TPM/SE-style roots of trust and remote attestation; network privacy uses Tor with Digidow-specific shortcuts so metadata does not redraw your movement map. Formal checks in Tamarin cover identity/credential spoofing under a strong adversary model; a Rust PoC shows end-to-end latencies of a few hundred ms locally and a few seconds over Tor — fine for doors, not yet for high-rate transit gates under full anonymity.

Limits

The design assumes online connectivity for the “no phone in pocket” story; true offline falls back to a carried device. Privacy of the protocol itself is still future work in the formal model. Run-time integrity after boot attestation remains a soft spot, biometric matching still trusts the sensor unless MPC matching lands, and Tor-class anonymity costs latency. Operating a public PIA is a real ops burden for ordinary households — exactly the product gap a home-edge node would need to productize.

What we'd test

Stand up a PIA on a lab home-edge box (TEE-attested process), issue a tiny VC set (household member + “building access” attribute), and run door-style verify against a mock sensor/verifier with selective disclosure and abort-before-share. Meter: attestation freshness, unlinkability under colluding verifiers, and whether three trust planes (sensor RoT, PIA custody, issuer signatures) stay isolated when one plane is hostile. That experiment feeds Trustworthy Agentic Operations and the CEIN/Wang roof without touching partner PHY claims.

Chapter 1

A-MEM: Agentic Memory for LLM Agents

Why it matters

Most LLM agents store memories as flat, isolated chunks retrieved by similarity, so they cannot connect a new experience to older ones or revise what they already know. A-MEM shows that letting the agent organize its own memory into a linked, evolving note network measurably improves long-horizon reasoning — the same design principle behind iShareHow's A-MEM memory engineer.

Core idea

Borrowing from the Zettelkasten method, every new memory becomes a structured note with contextual description, keywords, and tags generated by the LLM itself.

How the memory evolves

  • Link generation: a new note is compared with related notes and the agent decides which meaningful links to create.
  • Memory evolution: adding a note can update the context and attributes of existing linked notes, so older knowledge is refined rather than frozen.

Results

Across long-conversation question-answering benchmarks and several foundation models, A-MEM outperformed prior memory baselines, with the largest gains on multi-hop questions that require connecting facts from different sessions.

Takeaway for the lab

Store research notes as linked, agent-maintained units instead of raw chunks; let each new paper update the links of earlier study chapters.