Compromise assessment with evidence an analyst can inspect.
An AI assisted system for examining Windows Event Logs and memory forensic artifacts. Candidate findings are mapped to MITRE ATT&CK and tied to specific evidence records for analyst review.
Public project summary · Local models · Analyst review
My contribution · Results · Evaluation scope · Public scope
Built by a ten person team during an eight week summer training engagement with Tamkeen Technologies as the industry client and mentor. The University of Prince Mugrin Entrepreneurship Centre hosted the training.
The team worked through two week sprints and four milestone reviews, followed by a final demonstration and handover. Tamkeen Technologies is not represented as an owner, certifier, or endorser of the software.
I served as Project Lead, with responsibilities in the Core AI, Architecture and Integration workstream.
My contribution combined delivery coordination with analysis track implementation, model selection and configuration, retrieval work, evaluation, and integration. The Core AI workstream involved shared ownership with colleagues. The final selected implementation and the reported benchmark results are team outcomes.
| Stage | Purpose |
|---|---|
| Examine artifacts | Read supported Windows Event Log and memory evidence |
| Propose findings | Relate supported observations to MITRE ATT&CK techniques |
| Preserve references | Connect findings to records an analyst can inspect |
| Support review | Keep approval and final judgment with the analyst |
Insufficient support is reported as a coverage or evidence gap. Models run locally. The assessment scope is identification; containment, eradication, and recovery are outside this project.
Reported team results from separate frozen controlled lab benchmarks. Each track used 18 samples. The Event Log and memory results are shown separately because their evidence and evaluation scope differ.
| Evidence domain | Micro precision | Micro recall | Micro F1 | Macro F1 |
|---|---|---|---|---|
| Windows Event Logs | 94.9% | 93.3% | 94.1% | 95.0% |
| Memory command line evidence | 92.5% | 92.5% | 92.5% | 93.8% |
| Check | Event Logs | Memory |
|---|---|---|
| Evidence grounding | 100% | 98.1% |
| Hallucination rate | 0% | 1.9% |
| Unresolvable record or observation IDs | 0 | 0 |
| Hallucinations among false positives | 0 | 0 |
Resolvable references and correct reasoning are separate checks. A citation can point to a real record while the interpretation is wrong. These results describe the evaluated runs, not a universal guarantee about system output.
The benchmarks used an AI only, one shot protocol over a purpose built corpus with team authored ground truth. The corpus includes attack samples, benign controls, and incomplete evidence controls.
Reported measures include micro and macro precision, recall and F1, evidence grounding, hallucinations, unresolvable citations, false positives per sample, and wall clock time per sample. No new benchmark was run for this public presentation.
| Limit | Implication |
|---|---|
| Team authored ground truth | The evaluation is not independent external validation |
| Limited controlled corpus | Results do not establish general production performance |
| Memory command line track | The benchmark covers one slice of memory forensic evidence |
| Serial local execution | Throughput remains suitable for research evaluation; operational throughput is not established |
| Partial parser coverage | Unsupported formats remain explicit coverage gaps |
This repository contains a public summary and illustrative cover. Application source, interface material, detailed system design, evaluation corpora, and results for individual samples remain private.
For a discussion of the work and evaluation methodology, contact me on LinkedIn.
A proposed finding is a hypothesis tied to evidence. A qualified analyst reviews the cited records and makes the final assessment. Output must not be presented as a certified forensic conclusion.
