Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

DFIR forensic-image detection benchmark

A reproducible benchmark for DFIR triage-image detection — does an analysis engine (deterministic or LLM-driven) actually surface the attacker techniques planted in a forensic evidence image? The first tier emulates Volt Typhoon (MITRE G1017, primary source CISA/NSA/FBI AA24-038A) as a living-off-the-land intrusion on a synthetic Windows host.

All evidence is synthetic lab material (not_a_real_seizure). The benchmark tests detection capability, not real provenance. Nothing here is real seized data.

What's here

  • Grader — scripts/dfir-image-bench/score-image-run.mjs: a pure, offline, model-agnostic scorer. It reads a producer verdict.json and the oracle and scores four axes — MITRE technique elevation (0.40), IOC recovery (0.20), attack-chain reconstruction (0.25), and AI-degree (0.15) — plus provenance (deterministic vs LLM). Selftest: node scripts/dfir-image-bench/selftest-image-bench.mjs.
  • Oracle — fixtures/dfir-image-bench/medium-volt/ground-truth.json + emulation-steplog.md: the STIX 2.1 / Attack-Flow ground truth. Six artifacts: RDP logon (T1021.001), internal proxy / netsh portproxy (T1090.001), NTDS IFM dump (T1003.003), selective Security-log clear (T1070.001), staged-then-deleted archive (T1560.001), and a planted AI-integration helper (AI-degree axis).
  • Grader fixtures — fixtures/.../runs/*.verdict.json: synthetic producer verdicts (perfect hit, LLM-authored, clean negative, mutated miss) that the selftest grades to prove the scorer's behavior.
  • Evidence builders — scripts/dfir-image-bench/local-kvm/ + guest/*.ps1 + build-live-image.sh: build a KVM Windows victim, emulate the chain, and collect a triage pack. (The large binary evidence images are not shipped here — they are regenerable from these builders.)
  • Runner + adapters — scripts/caseforge/ and scripts/dfir-image-bench/: drive an analysis engine over a pack and adapt its sealed output to the grader fixture (seal-to-fixture.py, score-live-run.sh), plus offline evidence probes (evtx-eid-probe.py, judge-dup-repro.py).
  • Docs — docs/: the benchmark design, a detection scorecard, and receipts from a worked run that took a deterministic engine from composite 0.0 → 0.755 (6/6 oracle artifacts) by fixing a finding-drop bug and adding five detectors + offline-carved evidence. See docs/medium-volt-richer-pack-2026-07-13.md.

Tiers

  • medium-volt (fixtures/dfir-image-bench/medium-volt/) — the worked tier; a deterministic engine reaches 6/6 artifacts (composite 0.755).
  • hard-mem (fixtures/dfir-image-bench/hard-mem/) — an authored HARD tier that targets memory-only techniques (LSASS T1003.001, injection T1055.001, rootkit T1014 via Volatility3), timestomp ($MFT $SI/$FN anomaly), and a full AI-orchestrated artifact. A triage-without-memory run correctly caps out (memory techniques MISS, and AI-present under-grades the AI-orchestrated expectation). Requires a memory image to fully solve.

An offline-LLM agent study (docs/offline-llm-agent-tool-calling-2026-07-13.md) finds the tool-call malformation is model-specific (gpt-oss errors; llama-family call tools cleanly), but no local model yet completes the sealed-case workflow — the deterministic engine is the baseline to beat.

Quick start

# grade the synthetic producer fixtures against the oracle (no evidence needed)
node scripts/dfir-image-bench/selftest-image-bench.mjs

# score one verdict.json against the oracle
node scripts/dfir-image-bench/score-image-run.mjs <verdict.json> \
  --oracle fixtures/dfir-image-bench/medium-volt/ground-truth.json --json

Scoring, briefly

Per-artifact HIT/MISS on each axis; composite = 0.40·mitre + 0.20·iocs + 0.25·chain + 0.15·ai. The AI axis expects ai_assessment.degree = "AI-present" and hard-fails (FAIL_OVERGRADE) on an AI-orchestrated overclaim — seeded style/slopsquat decoys must not raise the degree. Registry / event-log IOCs are marked not_gradeable in the IOC axis (a producer-schema limit, not a detection gap).

Notes / placeholders

  • Lab-VM credentials and network addresses in the builder scripts are placeholders (CHANGEME-LabPass1!, 192.168.200.10 on an isolated net, <user>@<dfir-host>). Change them for your environment.
  • The synthetic sk-ant-api03-LABONLY-NOT-A-REAL-KEY-… string is a deliberate planted decoy — the AI-helper detection target — not a real credential.

License

MIT — see LICENSE.

About

Reproducible DFIR forensic-image detection benchmark (Volt Typhoon / AA24-038A). Synthetic evidence; deterministic model-agnostic grader.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages