Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
16 changes: 12 additions & 4 deletions .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -78,12 +78,20 @@ POSTGRES_PASSWORD=your-secure-db-password
# INTEGRITY_REQUIRE_SIGNED=false

# Trust-weighted ranking (W2c): fuse similarity with content trust, votes, decay, provenance.
# Off by default (ordering unchanged). Weights must sum to 1.0.
# ON by default — free on a clean corpus, and the only control that degrades gracefully under
# memory poisoning. Set false to restore pure vector ordering. Weights must sum to 1.0.
#
# RANKING_W_TRUST is a security parameter. Untrusted content can outrank trusted content
# whenever its similarity advantage exceeds
# RANKING_W_TRUST * 0.7 / RANKING_W_SEMANTIC
# (0.7 = the internal-vs-untrusted trust prior gap). Poisoned memories are written to match
# the query, so they routinely gain 0.3+ similarity. Lowering RANKING_W_TRUST below ~0.25
# makes the defense outbiddable — see docs/security/poisoning.md.
# ENABLE_TRUST_WEIGHTED_RANKING=false
# RANKING_W_SEMANTIC=0.60
# RANKING_W_TRUST=0.15
# RANKING_W_SEMANTIC=0.45
# RANKING_W_TRUST=0.35
# RANKING_W_EFFECTIVENESS=0.10
# RANKING_W_DECAY=0.10
# RANKING_W_DECAY=0.05
# RANKING_W_PROVENANCE=0.05
# RANKING_CANDIDATE_MULTIPLIER=4

Expand Down
11 changes: 11 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -164,3 +164,14 @@ tmpclaude-*/

# aegis inspect generated output
aegis-out/

# Database dumps
backups/

# LongMemEval: the dataset (278MB / 2.7GB) is downloaded, not vendored; run artifacts and
# local run logs are reproducible. The harness, its README, and committed reports stay tracked.
benchmarks/memory/longmemeval/longmemeval_s.json
benchmarks/memory/longmemeval/longmemeval_m*
benchmarks/memory/longmemeval/results/
benchmarks/memory/longmemeval/*.log
benchmarks/memory/longmemeval/*.err
48 changes: 48 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,54 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

## [Unreleased]

Measured what the memory is actually worth — clean and under attack — and fixed a defense that
the measurement showed was not working.

### Added

- **LongMemEval memory-quality benchmark** (`benchmarks/memory/longmemeval/`). Aegis scores
**0.860** on LongMemEval_S (500/500 questions, `top_k=15` semantic retrieval, reader
`claude-sonnet-5`, judge `gpt-4o-2024-08-06` with the benchmark's official prompts, dataset
pinned at revision `2ec2a55`). Resumable phases (ingest / answer / judge / score), seeded
subsampling, and a per-run report recording dataset revision, models, and parameters.
- **Memory-poisoning benchmark (W4.2)** — the same benchmark run against a poisoned corpus, and
the first published *utility retained under attack* numbers for an agent memory system.
360 fabricated memories (1.2% of corpus) asserting false answers, admitted at
`trust_level=untrusted`: accuracy falls **0.850 → 0.300** undefended, and trust-weighted
retrieval recovers it to **0.475** (exact McNemar p=0.0015). Tooling: `poison_corpus.py`
(generate / inject / delete, with recorded IDs for exact restore), `analyze_poison.py`,
`w42_report.py`. Full write-up in [`docs/security/memory-poisoning.md`](docs/security/memory-poisoning.md).
- **Adversarial ranking tests** (`tests/test_ranking.py::TestAdversarialSimilarity`). Every prior
ranking test compared candidates *at equal similarity* — the one regime an adversary never
operates in. The new tests assert that untrusted content stays below trusted content even when
it has the higher similarity score, and pin the defended semantic margin as arithmetic so a
future weight change cannot silently shrink it.

### Changed

- **`RANKING_W_TRUST` raised 0.15 → 0.35 (and `RANKING_W_SEMANTIC` 0.60 → 0.45,
`RANKING_W_DECAY` 0.10 → 0.05).** At the old weights, trust-weighted ranking was
statistically **indistinguishable from no defense** against memory poisoning (0.317 vs 0.300,
p=0.80). Ordering only flips while the similarity gap stays under
`w_trust * 0.7 / w_semantic` — 0.175 at the old weights, and query-shaped poison routinely
gains 0.3+. The new weights raise that margin to 0.544; poison stops ranking first in 98% of
questions. Defaults are defined in **two** places that must stay in sync: `server/config.py`
and the `RankingWeights` dataclass in `server/ranking.py`.
- **`ENABLE_TRUST_WEIGHTED_RANKING` now defaults to `true`.** With the corrected weights it costs
nothing on a clean corpus (0.875 enabled vs 0.850 disabled, p=0.45) and is the only control
that degrades gracefully under poisoning. Set `false` to restore pure vector ordering.

### Known issues

- `DELETE /memories/{id}` returns 500. The handler deletes the memory row and then writes a
`deleted` event whose `memory_id` foreign key references the row it just removed
(`memory_events_memory_id_fkey`). Delete the event rows first, or write the event before the
delete. Found while restoring the poisoned benchmark corpus.
- Trust-weighted ranking only engages when trust levels actually vary. With `ENABLE_TRUST_LEVELS`
off and callers declaring nothing, every write lands as `internal` and the ranking signal is
flat regardless of weight. Integrations must mark tool-, web-, and agent-derived content as
`untrusted` for the defense to do anything.

## [2.7.0] - 2026-08-02

Provenance-native memory: security becomes a property of the memory itself, not just a gate in
Expand Down
52 changes: 49 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -49,6 +49,7 @@
- [How Aegis compares](#how-aegis-compares) — vs mem0, Zep, Letta
- [What's shipped vs roadmap](#whats-shipped-vs-roadmap) — no marketing ahead of code
- [Security benchmark](#security-benchmark) — does the detector actually work?
- [Memory benchmark](#memory-benchmark) — LongMemEval, and what poisoning does to it
- [Performance](#performance) — latency and throughput numbers
- [Deployment and configuration](#deployment-and-configuration)
- [Documentation](#documentation) · [Contributing](#contributing) · [License](#license)
Expand Down Expand Up @@ -551,12 +552,14 @@ Everything described above is **shipped and released** on PyPI as of `aegis-memo
| Claude Code plugin + keyless local MCP mode | ✅ Shipped | v2.6.0 |
| Notebook (`.ipynb`) ingestion + inline fix/verify-loop for `inspect` | ✅ Shipped | v2.6.0 |
| Provenance-native memory (immutable HMAC-signed origin record) | ✅ Shipped | v2.7.0 |
| Trust-weighted retrieval ranking | ✅ Shipped | v2.7.0 |
| Trust-weighted retrieval ranking (on by default, poisoning-validated) | ✅ Shipped | v2.7.1 |
| LongMemEval memory-quality + under-attack benchmark | ✅ Shipped | v2.7.1 |

**Directions we're exploring** (not commitments — track them in
[Discussions](https://github.com/quantifylabs/aegis-memory/discussions) and the
[Changelog](CHANGELOG.md)): broader injection coverage via the adaptive attack harness, and
continued benchmark hardening.
[Changelog](CHANGELOG.md)): a hard untrusted-content gate for retrieval (a weight can be
outbid by a determined attacker; a constraint cannot), taint-aware context so the reader can
discount untrusted evidence, and broader injection coverage via the adaptive attack harness.

## Security benchmark

Expand Down Expand Up @@ -613,6 +616,49 @@ with the Claude classifier. Both at 0.000 FPR.
[`docs/security/benchmark.md`](docs/security/benchmark.md)** · reproduce with
`python benchmarks/injection/run_benchmark.py`.

## Memory benchmark

A security wrapper that retrieves badly is not worth wrapping. So we measure the memory too —
on [LongMemEval](https://github.com/xiaowu0162/LongMemEval) (ICLR 2025), where each of 500
questions is buried in ~50 sessions (~115K tokens) of chat history.

**Clean: 0.860** (500/500 questions, `top_k=15` semantic retrieval, no reranking or query
rewriting; reader `claude-sonnet-5`, judge `gpt-4o-2024-08-06` with the benchmark's official
prompts; dataset pinned at `2ec2a55`). For scale, the paper reports GPT-4o at 60.6–64% reading
the full context and 87–92% in the *oracle* condition where only the evidence sessions are
supplied.

**Then we poisoned it.** 360 fabricated memories asserting false answers — 1.2% of the corpus,
written to match the questions, arriving at `trust_level=untrusted` as tool or web content
would:

| Configuration | Accuracy | Utility retained | Poison ranked #1 |
|---|---:|---:|---:|
| clean corpus | 0.850 | 100% | 0% |
| poisoned, no defense | 0.300 | 35% | 100% |
| **poisoned, trust-weighted retrieval** | **0.475** | **56%** | **2%** |

*(n=120; trust-weighted vs undefended, exact McNemar p=0.0015)*

**Write-time screening caught 0 of the 360.** That is not a detector failure — the poison
contains no injection payload, just plainly-worded false facts. It is the reason Aegis ranks
by provenance at read time instead of relying on screening the write, and it is why a system
whose only defense is a write-time filter has no defense against this attack.

**We also found this defense broken in our own defaults.** At the previously shipped
`RANKING_W_TRUST=0.15`, trust weighting was statistically indistinguishable from no defense
(p=0.80): the trust term could only overturn a 0.175 similarity gap, and query-shaped poison
clears that every time. Every unit test passed, because each compared candidates *at equal
similarity* — which is the one thing an adversary never does. Fixed in this release (0.35, on
by default, free on clean data at p=0.45), with an adversarial test that pins the property.

56% is not 100%. Even defended, 8.8% of retrieved context is still poisoned and the reader
often believes it — closing that is what taint-aware context and a hard untrusted-gate are for.

→ **Full method, per-type scores, significance tests, and limitations:
[`docs/security/memory-poisoning.md`](docs/security/memory-poisoning.md)** · reproduce with
`benchmarks/memory/longmemeval/`.

## Performance

<details>
Expand Down
54 changes: 54 additions & 0 deletions benchmarks/memory/longmemeval/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,54 @@
# LongMemEval on Aegis Memory

> **Results: 0.860 clean (n=500); 0.850 → 0.300 under 1.2% poisoning, recovered to 0.475 by
> trust-weighted retrieval.** Full write-up: [`docs/security/memory-poisoning.md`](../../../docs/security/memory-poisoning.md).
> Poisoning tooling lives alongside this harness — `poison_corpus.py`, `analyze_poison.py`,
> `w42_report.py`, `compare_sweep.py`.

Measures Aegis's memory quality on [LongMemEval](https://github.com/xiaowu0162/LongMemEval)
(ICLR 2025) — 500 questions, each hidden in ~50 sessions (~115K tokens) of chat history,
testing information extraction, multi-session reasoning, temporal reasoning, knowledge
updates, and abstention.

## Method

1. **Ingest** — each question's haystack sessions are replayed into a running Aegis server
via `POST /memories/add_batch`, one memory per user/assistant round, prefixed with the
session timestamp. Each question gets its own namespace + agent_id (`scope=agent-private`),
so retrieval is isolated per question (Aegis dedup is namespace-scoped, so identical
rounds across questions do not collapse).
2. **Answer** — the question is embedded and queried against Aegis (`POST /memories/query`,
default `top_k=15`); retrieved memories go to the reader (`claude-sonnet-5`), which is
instructed to answer only from memories and to abstain when they don't contain the answer.
3. **Judge** — answers are graded by the official LongMemEval judge prompts (vendored
verbatim from [`evaluation/evaluate_qa.py`](https://github.com/xiaowu0162/LongMemEval/blob/main/src/evaluation/evaluate_qa.py)) with the
paper's pinned judge model `gpt-4o-2024-08-06`, temperature 0.
4. **Score** — accuracy overall and per question type, written to `results/<run>/report.json`
with dataset revision, models, and parameters recorded.

## Reproducibility

- Dataset: HF `xiaowu0162/longmemeval`, file `longmemeval_s`,
revision `2ec2a557f339b6c0369619b1ed5793734cc87533`,
sha256 `08d8dad4be43ee2049a22ff5674eb86725d0ce5ff434cde2627e5e8e7e117894`.
(Not committed — 278 MB; download to `longmemeval_s.json` in this directory.)
- Subsampling (`--limit N`) uses `random.Random(42).sample`, recorded in the report.
- Every phase writes a resumable JSONL artifact; re-running skips completed questions.

## Running

Requires the Aegis docker stack up (`docker compose up`) and `AEGIS_API_KEY`,
`OPENAI_API_KEY` (embeddings via the server + judge), `ANTHROPIC_API_KEY` (reader)
in the environment or repo `.env`.

```sh
python run_longmemeval.py all --limit 10 # validation subset
python run_longmemeval.py ingest # full 500-question ingest (hours; ~$1.50 embeddings)
python run_longmemeval.py answer # reader (~$7-9 on Sonnet 5)
python run_longmemeval.py judge # judge (~$3-4 on GPT-4o)
python run_longmemeval.py score
```

Note: the server rate-limits per project (default 60 req/min, 1000 req/hour). The full
ingest is ~2,000 batch calls — raise `RATE_LIMIT_PER_MINUTE`/`RATE_LIMIT_PER_HOUR` in
`docker-compose.yml` for the full run, or let the harness back off on 429s (slower).
67 changes: 67 additions & 0 deletions benchmarks/memory/longmemeval/analyze_poison.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,67 @@
"""Measure how much poison reached the reader's context in each W4.2 arm.

Answer accuracy conflates two things: whether retrieval kept poison out, and whether the
reader resisted whatever poison got through. This reports the retrieval half directly —
what fraction of each question's top-k was poisoned, and how often poison ranked first.

python analyze_poison.py --limit 120
"""
import argparse
import json
from pathlib import Path

from run_longmemeval import HERE, jsonl_read, load_dataset

ARMS = [
("poisoned TWR off", "poisoned_twr_off"),
("poisoned TWR on", "poisoned_twr_on"),
("poisoned TWR w=0.35", "poisoned_twr_w35"),
]


def main():
ap = argparse.ArgumentParser()
ap.add_argument("--limit", type=int, default=120)
args = ap.parse_args()

qids = {e["question_id"] for e in load_dataset(args.limit)}
injected = jsonl_read(HERE / "results" / "poison" / "injected.jsonl")
poison_ids = {q: set(r["memory_ids"]) for q, r in injected.items()}
total_poison = sum(len(v) for v in poison_ids.values())
print(f"\npoison injected: {total_poison} memories across {len(poison_ids)} questions\n")

header = f"{'arm':22s}{'poison in top-k':>18s}{'questions w/ any':>18s}{'poison ranked #1':>18s}"
print(header)
print("-" * len(header))

for label, sub in ARMS:
hyp = jsonl_read(HERE / "results" / sub / "hypotheses.jsonl")
if not hyp:
continue
slots = retrieved = with_any = ranked_first = n = 0
for qid in qids:
rec = hyp.get(qid)
if not rec:
continue
ids = rec["retrieved_memory_ids"]
pois = poison_ids.get(qid, set())
hits = [i for i in ids if i in pois]
n += 1
slots += len(ids)
retrieved += len(hits)
if hits:
with_any += 1
if ids and ids[0] in pois:
ranked_first += 1
if not n:
continue
print(
f"{label:22s}{retrieved}/{slots} ({retrieved/slots:.1%})".ljust(40)
+ f"{with_any}/{n} ({with_any/n:.0%})".rjust(0).ljust(18)
+ f"{ranked_first}/{n} ({ranked_first/n:.0%})"
)
print()


if __name__ == "__main__":
main()
Loading
Loading