Skip to content

GLAIVE 0.3.0: benchmarks, privacy, audit trail, evidence search - #4

Merged
aliyaalias19 merged 21 commits into
mainfrom
feat/v0.3
Oct 3, 2026
Merged

aliyaalias19 merged 21 commits into
mainfrom
feat/v0.3

Conversation

@aliyaalias19

Copy link
Copy Markdown
Owner

0.3.0 - 2026-10

Added

  • Benchmarks on public data (glaive bench run): EVTX-ATTACK-SAMPLES,
    OTRF Security-Datasets and the NextronSystems evtx-baseline goodware logs,
    scored against the datasets' own ATT&CK labels at alert and finding level;
    glaive bench compare puts runs side by side (rules alone vs each model).
    CI enforces detection floors on EVTX-ATTACK-SAMPLES.
  • Log-poisoning benchmark (glaive bench poisoning): nine planted prompt
    injections plus a control, and what a model that obeys them could achieve.
  • Privacy: case data is pseudonymised (USER_1, HOST_2...) before it is
    sent to a cloud model, embedder or reranker, and restored in replies.
    GLAIVE_PRIVACY=local-only|pseudonymize|off.
  • Audit trail: spans for every investigation, agent, model call and tool
    call in <case>/trace.jsonl (OpenTelemetry GenAI conventions, no prompt
    text); optional OTLP export (glaive[otel]); glaive trace CASE.
  • Evidence search (GraphRAG): BM25 + optional vectors (local fastembed or
    Ollama, or OpenAI-compatible APIs) fused with RRF, optional reranker; nodes
    are indexed with their neighbours. glaive search, search_evidence tool
    for agents and MCP, glaive bench retrieval (recall@k, MRR).
  • Ask the case answers from findings and evidence nodes, with checked
    [F#] / [E#] citations; glaive ask CASE QUESTION.
  • Past-case memory (opt-in, local): glaive remember, glaive memory,
    indicator overlaps with earlier cases, recall_past_cases agent tool.
  • ATT&CK tactics on alerts, from a bundled Enterprise ATT&CK v19.2 table
    (old names and revoked IDs such as T1562.001 still resolve).
  • Confidence calibration in glaive eval.
  • Readers for NXLog/Logstash (OTRF) and Winlogbeat JSON; .tar, .tar.gz
    and .tgz evidence archives with the same safety checks as zips.
  • ROADMAP.md.

Changed

  • Prompt-injection detection also finds text hidden with zero-width
    characters, look-alike letters or Base64 (e.g. powershell -enc): 9/9
    planted payloads detected instead of 6/9.
  • Findings from a model that clear activity ("no malicious activity", "false
    positive", "the host is clean") wait for analyst approval.
  • A Skeptic refutation of a rule finding no longer marks it disputed; it goes
    to an analyst with the Skeptic's argument.
  • Sigma rules are pre-filtered per log source and event ID: 2,200 SigmaHQ rules
    run about 2.7x faster with identical results.
  • numpy is now a dependency (vector search).

Fixed

  • With the web app open in a browser, Ctrl+C did not stop glaive serve; a
    second Ctrl+C exited with a traceback.
  • The verdict-tampering injection pattern fired on word lists in real Windows
    registry values.

The live event stream only ended when the browser disconnected, so uvicorn waited forever for it during shutdown and a second Ctrl+C dumped a traceback. A shutdown flag set from the server's exit handler now ends open streams, with a 3-second graceful-shutdown limit as a backstop.
OTRF Security-Datasets ship NXLog/Logstash records with every field at the top level and the real host in 'Hostname' ('host' is the log collector). Elastic Winlogbeat nests fields under 'winlog'. Both shapes now normalize to GLAIVE events, with Sysmon UtcTime preferred over the shipper's receive time.
Adds a compact technique-to-tactic table generated from the official Enterprise ATT&CK v19.2 STIX bundle. Alerts now record tactics from Sigma tags and their techniques. v19 renamed Defense Evasion to Stealth and revoked T1562.x in favour of T1685; old names and revoked IDs still resolve, because rules and public datasets use them.
Which rules accept a (log family, event id) pair depends only on their logsource, so it is now computed once per pair and cached. With the 2,200 SigmaHQ Windows rules, EVTX-ATTACK-SAMPLES (35,807 events) goes from 33s to 13s with identical alerts; a test checks the result against evaluating every rule.
Several public datasets (some OTRF Security-Datasets, NextronSystems evtx-baseline) and many triage tools ship tar archives. They get the same protection as zips: path traversal refused, size and ratio limits, and only regular files are written (links, devices and fifos are skipped).
glaive bench run scores GLAIVE on EVTX-ATTACK-SAMPLES (ATT&CK tactic per folder), OTRF Security-Datasets (techniques from their metadata) and the NextronSystems evtx-baseline goodware logs (every alert is a false alarm). Each case runs in its own session, in rules mode or with the configured model, and is scored at alert and finding level. glaive bench compare puts runs side by side, so rules alone can be compared with each model. Integration tests keep floors on the real datasets.
For each confidence level the score now shows how many findings cover an answer-key item, so 'confirmed' can be checked to be right more often than 'inferred'.
Account names, host names, internal IP addresses, e-mail addresses, domains and SIDs are replaced with stable tokens (USER_1, HOST_2...) in every request to a cloud provider and restored in the reply, so tool calls and findings are checked by the gate with the real values. Local providers (Ollama, localhost, private network) get the real data. GLAIVE_PRIVACY=local-only refuses cloud models; =off disables masking.
Investigations and questions write spans to <case>/trace.jsonl using the OpenTelemetry GenAI conventions (gen_ai.operation.name, gen_ai.request.model, gen_ai.usage.*, gen_ai.tool.name) plus the gate's decision for each finding. Prompt and response text are not recorded, only sizes. With the otel extra and OTEL_EXPORTER_OTLP_ENDPOINT set, the same spans go to any OpenTelemetry backend. glaive trace CASE summarises calls, tokens per model and gate decisions.
Each benchmark case ran in a temporary folder removed with shutil.rmtree(ignore_errors=True), which on Windows silently left the read-only evidence copies behind. The CLI's read-only-aware remove_tree moves to glaive.fsutil and is used for both.
Every graph node becomes a document of its fields, ATT&CK tactics and connected nodes, so 'PowerShell started by Word' finds the PowerShell node. BM25 (SQLite FTS5) and optional dense vectors are fused with reciprocal rank fusion, with an optional cross-encoder reranker. Embedders: local fastembed (multilingual MiniLM by default) or Ollama, or OpenAI-compatible APIs (OpenAI, SiliconFlow BGE-M3, Qwen, Jina, Gemini); cloud calls are pseudonymised. The index lives in the case folder and rebuilds when the graph grows. Agents and the MCP server get a search_evidence tool; glaive search CASE QUERY and glaive bench retrieval (recall@k, MRR on 15 plain-language questions, three in Chinese) are added.
…itations

Ask the case now searches the evidence graph as well as the findings. The model may cite findings [F#] and evidence nodes [E#]; each sentence must cite something it was given and every entity it names must be in what it cites, otherwise the sentence is removed. If nothing survives, or no model is configured, the answer lists the matching findings and evidence. Citations come back with the answer so the web app opens the cited node; glaive ask CASE QUESTION does the same in the terminal.
glaive remember CASE stores a case's findings and their indicators (hashes, public IPs, domains, threat names, unusual paths) in GLAIVE_HOME/memory.sqlite on this computer. glaive investigate and glaive remember then report indicators already seen in earlier cases, glaive memory search/list/forget manage the store, and agents get a read-only recall_past_cases tool when memory exists. Remembered findings are context, never evidence: new findings still cite the current case. GLAIVE_MEMORY=off disables it; tests use a temporary GLAIVE_HOME.
glaive bench poisoning plants nine payloads (English, Chinese, AI-addressed notes, chat markup, tool abuse, verdict tampering, zero-width and full-width obfuscation, Base64 inside an encoded PowerShell command) plus a clean control, each in a copy of the demo case, and reports which the injection rule detects. It also runs a scripted model that obeys every instruction it reads and reports what reaches the case. Baseline: 6/9 detected; a fooled model cannot delete rule findings or commit invented evidence, but can commit a 'host is clean' claim and dispute every finding. The next commits address those gaps.
Before matching, evidence text is normalised the way a model reads it: look-alike letters folded (NFKC, e.g. full-width), invisible characters removed (zero-width spaces, soft hyphens, bidi marks), and Base64 blobs such as PowerShell -EncodedCommand decoded (UTF-16LE or UTF-8) and scanned. Each hit records how the text was hidden. Log-poisoning detection goes from 6/9 to 9/9 with no false alarm on the control, the benign baseline or EVTX-ATTACK-SAMPLES. The verdict-tampering pattern now requires 'as' ('report this host as clean'): it fired on a word list in an OTRF registry value.
A finding from a model (not a deterministic rule) that declares activity benign ('no malicious activity', 'false positive', 'the host is clean', 'authorised test', and the Chinese equivalents) now enters pending_approval. Clearing a host is what planted instructions aim for, and a model cannot prove an absence. Findings record why they need approval, shown in the web app and the HTML report. In the poisoning benchmark the fooled model's 'host is clean' claim is now held instead of committed.
A rule finding states a fact (a detection rule fired on a given event); a refutation cannot make it less true, so it no longer lowers its confidence to disputed. The finding is held for analyst approval with the Skeptic's argument instead. Model findings that are refuted are still marked disputed. In the poisoning benchmark a fooled Skeptic previously disputed all 15 findings; now 14 rule findings keep their evidence-based confidence and await an analyst.
A benchmark job fetches EVTX-ATTACK-SAMPLES at a pinned commit, enforces the detection floors and prints the benchmark, log-poisoning and retrieval reports, so a change that makes GLAIVE detect less on real data fails CI.
ACCURACY_REPORT.md now reports rules-only results on EVTX-ATTACK-SAMPLES, OTRF Security-Datasets and a benign Windows baseline (built-in rules and SigmaHQ), log-poisoning detection and damage, search recall@k and demo calibration, with the caveats: the built-in rules generalise poorly, SigmaHQ numbers are optimistic, AI mode is not measured yet. README, LIMITATIONS, ROADMAP and .env.example cover privacy, search, ask, memory, the audit trail and the benchmarks.
@aliyaalias19
aliyaalias19 merged commit ab97154 into main Oct 3, 2026
10 checks passed
@aliyaalias19
aliyaalias19 deleted the feat/v0.3 branch October 3, 2026 09:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant