A reproducible benchmark for DFIR triage-image detection — does an analysis
engine (deterministic or LLM-driven) actually surface the attacker techniques
planted in a forensic evidence image? The first tier emulates Volt Typhoon
(MITRE G1017, primary source CISA/NSA/FBI AA24-038A) as a living-off-the-land
intrusion on a synthetic Windows host.
All evidence is synthetic lab material (
not_a_real_seizure). The benchmark tests detection capability, not real provenance. Nothing here is real seized data.
- Grader —
scripts/dfir-image-bench/score-image-run.mjs: a pure, offline, model-agnostic scorer. It reads a producerverdict.jsonand the oracle and scores four axes — MITRE technique elevation (0.40), IOC recovery (0.20), attack-chain reconstruction (0.25), and AI-degree (0.15) — plus provenance (deterministic vs LLM). Selftest:node scripts/dfir-image-bench/selftest-image-bench.mjs. - Oracle —
fixtures/dfir-image-bench/medium-volt/ground-truth.json+emulation-steplog.md: the STIX 2.1 / Attack-Flow ground truth. Six artifacts: RDP logon (T1021.001), internal proxy / netsh portproxy (T1090.001), NTDS IFM dump (T1003.003), selective Security-log clear (T1070.001), staged-then-deleted archive (T1560.001), and a planted AI-integration helper (AI-degree axis). - Grader fixtures —
fixtures/.../runs/*.verdict.json: synthetic producer verdicts (perfect hit, LLM-authored, clean negative, mutated miss) that the selftest grades to prove the scorer's behavior. - Evidence builders —
scripts/dfir-image-bench/local-kvm/+guest/*.ps1+build-live-image.sh: build a KVM Windows victim, emulate the chain, and collect a triage pack. (The large binary evidence images are not shipped here — they are regenerable from these builders.) - Runner + adapters —
scripts/caseforge/andscripts/dfir-image-bench/: drive an analysis engine over a pack and adapt its sealed output to the grader fixture (seal-to-fixture.py,score-live-run.sh), plus offline evidence probes (evtx-eid-probe.py,judge-dup-repro.py). - Docs —
docs/: the benchmark design, a detection scorecard, and receipts from a worked run that took a deterministic engine from composite 0.0 → 0.755 (6/6 oracle artifacts) by fixing a finding-drop bug and adding five detectors + offline-carved evidence. Seedocs/medium-volt-richer-pack-2026-07-13.md.
- medium-volt (
fixtures/dfir-image-bench/medium-volt/) — the worked tier; a deterministic engine reaches 6/6 artifacts (composite 0.755). - hard-mem (
fixtures/dfir-image-bench/hard-mem/) — an authored HARD tier that targets memory-only techniques (LSASS T1003.001, injection T1055.001, rootkit T1014 via Volatility3), timestomp ($MFT $SI/$FN anomaly), and a full AI-orchestrated artifact. A triage-without-memory run correctly caps out (memory techniques MISS, andAI-presentunder-grades theAI-orchestratedexpectation). Requires a memory image to fully solve.
An offline-LLM agent study (docs/offline-llm-agent-tool-calling-2026-07-13.md)
finds the tool-call malformation is model-specific (gpt-oss errors; llama-family call
tools cleanly), but no local model yet completes the sealed-case workflow — the
deterministic engine is the baseline to beat.
# grade the synthetic producer fixtures against the oracle (no evidence needed)
node scripts/dfir-image-bench/selftest-image-bench.mjs
# score one verdict.json against the oracle
node scripts/dfir-image-bench/score-image-run.mjs <verdict.json> \
--oracle fixtures/dfir-image-bench/medium-volt/ground-truth.json --jsonPer-artifact HIT/MISS on each axis; composite =
0.40·mitre + 0.20·iocs + 0.25·chain + 0.15·ai. The AI axis expects
ai_assessment.degree = "AI-present" and hard-fails (FAIL_OVERGRADE) on an
AI-orchestrated overclaim — seeded style/slopsquat decoys must not raise the
degree. Registry / event-log IOCs are marked not_gradeable in the IOC axis
(a producer-schema limit, not a detection gap).
- Lab-VM credentials and network addresses in the builder scripts are
placeholders (
CHANGEME-LabPass1!,192.168.200.10on an isolated net,<user>@<dfir-host>). Change them for your environment. - The synthetic
sk-ant-api03-LABONLY-NOT-A-REAL-KEY-…string is a deliberate planted decoy — the AI-helper detection target — not a real credential.
MIT — see LICENSE.