Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
36 commits
Select commit Hold shift + click to select a range
03f4101
chore: add .gitattributes and .editorconfig
aliyaalias19 Oct 2, 2026
46df10d
style: add missing newline at end of files
aliyaalias19 Oct 2, 2026
1375061
feat(graph): add alert and script-block nodes, graph serialization an…
aliyaalias19 Oct 2, 2026
98d2de6
test: run 14 File tests hidden by a duplicate class name
aliyaalias19 Oct 2, 2026
6ff8457
fix(grounding): ignore trailing punctuation in extracted entities
aliyaalias19 Oct 2, 2026
e2ebe08
feat(report): add severity, ATT&CK techniques and analyst review to f…
aliyaalias19 Oct 2, 2026
3c3670c
feat(mcp): accept severity and ATT&CK in commit_finding; report damag…
aliyaalias19 Oct 2, 2026
7fdab93
feat(case): persist investigations in a .glaive case file
aliyaalias19 Oct 2, 2026
32b7fc4
feat(evtx): add optional Rust EVTX reader with python-evtx fallback
aliyaalias19 Oct 2, 2026
1463796
feat(ingestion): read JSON and JSON-Lines event exports
aliyaalias19 Oct 2, 2026
66c322e
feat(ingestion): parse Windows Security, System, Sysmon and PowerShel…
aliyaalias19 Oct 2, 2026
9c64f90
test: remove internal step number from case-file test docstring
aliyaalias19 Oct 2, 2026
510bc99
feat(detection): add Sigma rule engine and 27 built-in Windows rules
aliyaalias19 Oct 2, 2026
5c120f5
feat(security): detect prompt injection in evidence (English and Chin…
aliyaalias19 Oct 2, 2026
5dc9296
feat(detection): add correlation rules across events
aliyaalias19 Oct 2, 2026
08febe9
feat(demo): add the Operation Invoice synthetic intrusion with an ans…
aliyaalias19 Oct 2, 2026
12d5f66
refactor(ingestion): let the orchestrator integrate an already-parsed…
aliyaalias19 Oct 2, 2026
7e347c9
feat(ingestion): ingest a file, folder or zip in one call
aliyaalias19 Oct 2, 2026
357b93c
feat(llm): add provider-neutral message types and model adapters
aliyaalias19 Oct 2, 2026
adcf36e
feat(llm): add model router with fallback, retries, circuit breaker a…
aliyaalias19 Oct 2, 2026
2f94760
feat(llm): configure model providers from environment variables
aliyaalias19 Oct 2, 2026
e4107da
feat(eval): score an investigation against an answer key
aliyaalias19 Oct 2, 2026
5ecbdb5
feat(agents): add the validated toolbox agents investigate with
aliyaalias19 Oct 2, 2026
1dbfb6b
feat(agents): add rule triage, Hunter, Skeptic and Reporter agents
aliyaalias19 Oct 2, 2026
abfe64e
feat(agents): run a full investigation from triage to saved report
aliyaalias19 Oct 2, 2026
b688cc2
feat(mcp): add investigation tools and finding metadata to the MCP se…
aliyaalias19 Oct 2, 2026
0604a28
test: read files as UTF-8 so the suite passes on Windows
aliyaalias19 Oct 2, 2026
cb3434f
feat(report): add a self-contained HTML investigation report
aliyaalias19 Oct 2, 2026
8a389cb
feat(web): add a local web app to upload evidence, watch the gate and…
aliyaalias19 Oct 2, 2026
d5695c2
feat(cli): add demo, investigate, serve, report, verify, models, eval…
aliyaalias19 Oct 2, 2026
4e97e20
build: add a Docker image for the web app
aliyaalias19 Oct 2, 2026
966b021
ci: test on Windows and Linux, Python 3.11 and 3.12, mcp 1.x and 2.x
aliyaalias19 Oct 2, 2026
041597e
docs: describe v0.2 as it is: usage, security model, limits and measu…
aliyaalias19 Oct 2, 2026
b3767a7
chore(release): 0.2.0
aliyaalias19 Oct 2, 2026
838689b
chore: tidy package metadata and file endings
aliyaalias19 Oct 2, 2026
9a8fe83
ci: use a valid container name in the Docker job
aliyaalias19 Oct 2, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
14 changes: 14 additions & 0 deletions .dockerignore
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
.git
.github
.venv
venv
cases
analysis
glaive-demo
test_evidence
evidence_samples
tests
verification
docs
**/__pycache__
.env
19 changes: 19 additions & 0 deletions .editorconfig
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
# Consistent formatting in every editor (VS Code: install the "EditorConfig" extension).
root = true

[*]
charset = utf-8
end_of_line = lf
insert_final_newline = true
trim_trailing_whitespace = true

[*.py]
indent_style = space
indent_size = 4

[*.{yml,yaml,json,toml,html}]
indent_style = space
indent_size = 2

[*.md]
trim_trailing_whitespace = false
28 changes: 28 additions & 0 deletions .env.example
Original file line number Diff line number Diff line change
@@ -0,0 +1,28 @@
# Copy to .env (never commit it) or set these in your shell.
# GLAIVE uses every provider you configure, in this order, as a fallback chain.
# With none set, GLAIVE runs its detection rules only (no AI).

# --- International ---
# ANTHROPIC_API_KEY=sk-ant-...
# OPENAI_API_KEY=sk-...
# GEMINI_API_KEY=...
# OPENROUTER_API_KEY=...

# --- China ---
# DEEPSEEK_API_KEY=sk-...
# DASHSCOPE_API_KEY=sk-... # Qwen; set QWEN_BASE_URL to your Model Studio workspace URL
# MOONSHOT_API_KEY=sk-... # Kimi
# ZHIPUAI_API_KEY=... # GLM (use GLM_BASE_URL=https://api.z.ai/api/paas/v4 for Z.ai)
# ARK_API_KEY=... # Doubao; set DOUBAO_MODEL to your endpoint/model id
# SILICONFLOW_API_KEY=...

# --- Fully offline / self-hosted ---
# OLLAMA_MODEL=qwen3:8b # after: ollama pull qwen3:8b
# GLAIVE_BASE_URL=http://gpu-box:8000/v1 # vLLM / SGLang / LMDeploy / llama.cpp
# GLAIVE_MODEL=Qwen/Qwen3-32B

# --- Controls ---
# GLAIVE_PROVIDERS=deepseek,anthropic,ollama # explicit fallback order
# GLAIVE_MODEL=... # model for the first provider
# GLAIVE_TOKEN_BUDGET=300000 # stop the agents after this many tokens
# GLAIVE_WEB_TOKEN=... # require a token for the web app
15 changes: 15 additions & 0 deletions .gitattributes
Original file line number Diff line number Diff line change
@@ -0,0 +1,15 @@
# Store text with LF line endings in the repository; Git converts to the
# platform's endings on checkout. Keeps diffs clean for every contributor.
* text=auto eol=lf

# Windows scripts keep CRLF.
*.bat text eol=crlf
*.cmd text eol=crlf
*.ps1 text eol=crlf

# Binary evidence and images are never touched.
*.evtx binary
*.png binary
*.jpg binary
*.zip binary
*.glaive binary
57 changes: 57 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,57 @@
name: CI

on:
push:
branches: [main, "v0.*"]
pull_request:

permissions:
contents: read

concurrency:
group: ci-${{ github.ref }}
cancel-in-progress: true

jobs:
test:
name: ${{ matrix.os }} / Python ${{ matrix.python }} / ${{ matrix.mcp }}
runs-on: ${{ matrix.os }}
timeout-minutes: 20
strategy:
fail-fast: false
matrix:
os: [ubuntu-latest, windows-latest]
python: ["3.11", "3.12"]
mcp: ["mcp<2", "mcp>=2,<3"]
steps:
- uses: actions/checkout@v7
- uses: actions/setup-python@v7
with:
python-version: ${{ matrix.python }}
- name: Install
run: |
python -m pip install --upgrade pip
python -m pip install -e ".[dev]" "${{ matrix.mcp }}"
- name: Lint
run: ruff check glaive tests/v02
- name: Tests
run: python -m pytest -q
- name: Bypass (adversarial) tests
run: python -m pytest -q -m bypass
- name: Demo end to end (no AI model)
run: glaive demo --offline

docker:
name: Docker image builds and serves
runs-on: ubuntu-latest
timeout-minutes: 15
steps:
- uses: actions/checkout@v7
- name: Build
run: docker build -t glaive .
- name: Run and check the token is enforced
run: |
docker run -d --name glaive-ci -p 8765:8765 -e GLAIVE_WEB_TOKEN=ci-token glaive
for i in $(seq 1 30); do curl -fs -H "X-Glaive-Token: ci-token" http://127.0.0.1:8765/api/case && break; sleep 1; done
test "$(curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:8765/api/case)" = "401"
docker exec glaive-ci glaive demo --out /tmp/demo --offline
8 changes: 8 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -54,3 +54,11 @@ test_evidence/

# Runtime investigation output (evidence store, reports)
analysis/

# The example config (no secrets) is meant to be shared
!.env.example


# Default output folders of `glaive investigate` and `glaive demo`
cases/
glaive-demo/
58 changes: 49 additions & 9 deletions ACCURACY_REPORT.md
Original file line number Diff line number Diff line change
@@ -1,14 +1,54 @@
# Accuracy Report

> **Auto-generated** by `verification/harness.py` against the ground-truth cases.
> Last run: pending.
Measured on GLAIVE 0.2.0. Reproduce with:

This file will be regenerated and committed before submission. It will contain:
```bash
glaive demo --offline # rules only, no AI model
pytest -m integration # real samples; set GLAIVE_EVTX_SAMPLES first
```

- Per-case: precision, recall, F1
- Hallucination count (findings unsupported by graph paths)
- Missed-artifact count (ground-truth findings not produced)
- False-positive count (suspicious findings that were actually benign)
- Confidence calibration: of findings flagged "confirmed", what fraction were correct?
## Demo case "Operation Invoice" (synthetic, with answer key)

Honest numbers — including the ones that make us look bad.
Two hosts, 247 events, twelve attack steps hidden in normal activity. Rules
only, no AI model:

| Item | Found |
|---|---|
| GT1 Malicious document: Word spawned PowerShell | yes |
| GT2 Encoded PowerShell download from the C2 server | yes |
| GT3 Microsoft Defender real-time protection disabled | yes |
| GT4 Persistence through a Run key pointing at svchost32.exe | yes |
| GT5 Payload beacons to the C2 server over port 443 | **no** |
| GT6 Account and group discovery | **no** |
| GT7 LSASS memory dumped with comsvcs.dll | yes |
| GT8 Brute force then successful logon to FILESRV-01 | yes |
| GT9 Malicious service installed on FILESRV-01 | yes |
| GT10 Shadow copies deleted (ransomware preparation) | yes |
| GT11 Security log cleared on FILESRV-01 | yes |
| GT12 Prompt injection planted for AI investigators | yes |

- Recall: **10/12 (83%)**
- Findings that match an answer-key item: 13/14 (the 14th is a true but
unlisted detail)
- ATT&CK technique coverage: 86%
- Ungrounded statements in the report: **0**

Why the two misses: the beacon (GT5) and the discovery commands (GT6) only
trigger medium/low alerts or none, and rule triage reports medium and above.
These are exactly what the Hunter agent is for; with a model connected the
test suite shows the combined result reaching 12/12 using a scripted model.
How well a real model does depends on the model.

## Public attack samples (real data)

All 278 EVTX files of [EVTX-ATTACK-SAMPLES](https://github.com/sbousseaden/EVTX-ATTACK-SAMPLES)
(37,364 events): every file ingests without errors, every graph node traces to
a stored evidence file (no fabricated provenance), and rule triage commits
findings with zero ungrounded statements. There is no answer key for this set,
so recall is not measured on it.

## Not yet measured

- Real-model recall and precision on the demo case, per provider.
- False-positive rate on benign baselines.
- Confidence calibration (how often "confirmed" findings are correct).
38 changes: 38 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,38 @@
# Changelog

## 0.2.0 - 2026-10

### Fixed (found by auditing and running v0.1)
- Fresh installs failed: `mcp>=1.2` pulled mcp 2.x, which renamed `FastMCP`. Now works on 1.x and 2.x.
- CLI crashed on Windows consoles when printing non-ASCII characters.
- The gate accepted claims unrelated to their evidence; claims are now grounded entity by entity.
- Defender event 5001 (real-time protection disabled) was silently dropped.
- Nodes without a source file got placeholder evidence hashes (`fff...`, `000...`); such records are now rejected.
- Orchestrator crashed when a run without a source file preceded one with a file.
- Volatility pstree parent lookup crashed on mixed known/unknown start times.
- `query_graph`: no hard limit, time filters never matched, internals reachable through filters, File paths missing from results, date-like hostnames broke lookups.
- Evidence store: stored copies were writable, and hashing and copying were separate steps.
- Any file could be ingested as a Defender EVTX.
- A damaged EVTX crashed ingestion with the fast reader; it is now skipped with a structured error.
- A bad record in a JSON array export crashed the reader.
- Tests read source files with the system code page and failed on Windows.
- Duplicate `Process` class in `nodes.py`; duplicate `TestFile` class meant 14 tests never ran.
- Registry paths in `\REGISTRY\MACHINE` form were not mapped to `HKLM`.

### Added
- `.glaive` case files (SQLite): graph, findings, evidence manifest, audit log.
- Windows Security / System / Sysmon / PowerShell parsing, cross-log process corroboration.
- Optional Rust EVTX reader (about 1000x faster) with automatic fallback, verified equivalent on 37,364 real events.
- JSON / JSON-Lines event exports; folder and zip ingestion with zip-slip and zip-bomb protection.
- Sigma rule engine, 27 built-in rules, SigmaHQ compatibility; correlation rules; prompt-injection detection.
- Multi-provider model router (Claude, GPT, Gemini, DeepSeek, Qwen, Kimi, GLM, Doubao, OpenRouter, SiliconFlow, Ollama, any OpenAI-compatible server) with fallback, retries, circuit breaker and token budget.
- Agents: rules triage, Hunter, Skeptic, Reporter; human approval of high-severity findings.
- Web app, HTML report, `glaive demo / investigate / serve / report / verify / models / eval / mcp`.
- Web app protection: localhost only by default, token for other addresses, Host and Origin checks against DNS rebinding and cross-site requests, no third-party requests from the page.
- MCP `commit_finding` accepts severity, ATT&CK techniques and a rationale; new tools `case_overview`, `list_alerts`, `get_neighbors`, `get_timeline`, `save_case`.
- Demo case with answer key and accuracy scoring.
- Dockerfile (non-root, token required) and CI on Linux and Windows, Python 3.11 and 3.12, mcp 1.x and 2.x.

### Changed
- The demo uses documentation-only addresses (203.0.113.0/24) and `.example` domains.
- Default Gemini model is `gemini-3.8-flash` (2.5 Flash is limited to existing users).
18 changes: 18 additions & 0 deletions Dockerfile
Original file line number Diff line number Diff line change
@@ -0,0 +1,18 @@
# GLAIVE web app in a container.
# docker build -t glaive .
# docker run -p 8765:8765 -v "$PWD/cases:/cases" --env-file .env glaive
#
# Every dependency ships prebuilt wheels for Linux, so no compiler is needed.
FROM python:3.12-slim
ENV PYTHONDONTWRITEBYTECODE=1 PYTHONUNBUFFERED=1
WORKDIR /app
COPY pyproject.toml README.md LICENSE ./
COPY glaive ./glaive
RUN pip install --no-cache-dir ".[fast]"
RUN useradd --create-home glaive && mkdir /cases && chown glaive /cases
USER glaive
VOLUME ["/cases"]
EXPOSE 8765
# Listening on 0.0.0.0 makes GLAIVE require an access token (set GLAIVE_WEB_TOKEN,
# or read the generated one from `docker logs`).
CMD ["glaive", "serve", "/cases/case", "--host", "0.0.0.0", "--port", "8765", "--no-browser"]
75 changes: 56 additions & 19 deletions LIMITATIONS.md
Original file line number Diff line number Diff line change
@@ -1,27 +1,64 @@
# Limitations

> Honesty over perfection. These are the things GLAIVE deliberately does **not**
> do, or does imperfectly. Documenting them is part of the design.
Honesty over perfection. These are the things GLAIVE does not do, or does
imperfectly, as of v0.2.

## What GLAIVE does not do
## Evidence it cannot read yet

- **Replace Protocol SIFT.** GLAIVE is an extension layer. The base
Protocol SIFT CLAUDE.md, skills, and case template are unmodified.
- **Live system response.** GLAIVE analyzes captured evidence, not running hosts.
- **Malware reverse engineering.** GLAIVE detects suspicious binaries via
artifact correlation but does not perform deep static or dynamic analysis.
- **Cloud forensics.** Current evidence types: memory dumps, Windows event logs,
registry hives, filesystem images. AWS / Azure / GCP audit logs are out of
scope for this submission.
- **Network packet inspection.** Network artifacts are sourced from host-side
logs and memory; we do not parse PCAP.
- **Human-in-the-loop approval workflows.** Protocol SIFT explicitly forbids
asking the user mid-task. GLAIVE uses the graph as critic, not a human.
- **Windows event logs only.** EVTX (Security, System, Sysmon, PowerShell,
Microsoft Defender) and JSON / JSON-Lines exports of them. Other files in an
evidence folder are hashed into the store for chain of custody but not parsed.
- **No memory, disk or registry-hive analysis in the pipeline.** A Volatility
pstree parser exists in `glaive/ingestion/volatility.py` from v0.1, but it is
not wired into `glaive investigate`. Disk images, registry hives, Linux,
macOS, cloud audit logs and network captures are on the roadmap.
- **PowerShell 4104 events are not linked to their process** unless the export
carries the process id: the readers do not keep the `Execution` element.

## Known weaknesses
## What the gate can and cannot catch

(Filled in during accuracy harness runs in Week 3.)
- It checks **concrete entities**: IP addresses, file paths, hashes, domains,
account and threat names. A claim can still overstate what the evidence
means using ordinary words ("the attacker *exfiltrated* data" when the
evidence only shows a connection). The Skeptic agent and analyst approval of
high-severity findings exist for this.
- Confidence comes from how many independent sources corroborate the cited
evidence. Two logs that are both wrong in the same way still count as two.
- No attribution ("this was APT-X") and no legal conclusions.

## Things that look like bugs but are not
## Heuristics that can be wrong

(Filled in as we discover them.)
- **Process identity** across logs uses (host, PID, start time truncated to the
second). Very fast PID reuse within one second can merge two processes.
- Events that only carry a PID (network, file, registry) are attached to the
most recent process with that PID that started before them.
- The Volatility pstree parser picks the most recent parent started before the
child, because pstree output does not include the parent's start time.
- **Correlation thresholds** (5 failed logons within 10 minutes; tampering then
a high alert within 2 hours) are fixed defaults, not tuned per environment.
- At most 250 alerts per rule are kept, so a very noisy community rule cannot
flood the graph; the summary reports how many were suppressed.

## Sigma support

About 90% of SigmaHQ's Windows rules load (2,168 of 2,410 when measured).
Rules using aggregations (`| count()`), `base64` / `base64offset` / `utf16`
modifiers, `near`, or log sources GLAIVE does not parse are skipped and
reported, never evaluated incorrectly.

## AI agents

- Tests use scripted models and HTTP-level mocks. Real-world quality depends on
the model you connect.
- Default model names were checked against provider documentation in
October 2026. Providers rename models often; override them with
`<PROVIDER>_MODEL` if a default stops working.

## Operational limits

- The web app is built for one analyst on one machine. There are no user
accounts: the optional `GLAIVE_WEB_TOKEN` is a single shared secret.
- `evidence_root` (MCP server and sessions) restricts which folders can be
ingested. It is off unless you set it.
- Uploads through the web app are limited to 2 GB by default
(`GLAIVE_MAX_UPLOAD_MB`).
Loading
Loading