Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
17 changes: 17 additions & 0 deletions .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -66,10 +66,27 @@ POSTGRES_PASSWORD=your-secure-db-password
# CONTENT_POLICY_PII=flag
# CONTENT_POLICY_SECRETS=reject
# CONTENT_POLICY_INJECTION=flag
# Injection action for global-scope writes (readable by every agent): reject | inherit
# CONTENT_POLICY_INJECTION_GLOBAL_SCOPE=reject

# HMAC integrity signing on store/verify (default: true)
# ENABLE_INTEGRITY_CHECK=true

# Verify-on-read: off | warn | drop (default: drop = exclude tamper-detected rows)
# INTEGRITY_READ_MODE=drop
# Also drop unsigned/legacy rows; set to true after running scripts/backfill_integrity.py
# INTEGRITY_REQUIRE_SIGNED=false

# Trust-weighted ranking (W2c): fuse similarity with content trust, votes, decay, provenance.
# Off by default (ordering unchanged). Weights must sum to 1.0.
# ENABLE_TRUST_WEIGHTED_RANKING=false
# RANKING_W_SEMANTIC=0.60
# RANKING_W_TRUST=0.15
# RANKING_W_EFFECTIVENESS=0.10
# RANKING_W_DECAY=0.10
# RANKING_W_PROVENANCE=0.05
# RANKING_CANDIDATE_MULTIPLIER=4

# Per-agent rate limiting (default: 30/min, 500/hr)
# PER_AGENT_RATE_LIMIT_PER_MINUTE=30
# PER_AGENT_RATE_LIMIT_PER_HOUR=500
Expand Down
55 changes: 55 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,8 +7,63 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

## [Unreleased]

## [2.7.0] - 2026-08-02

Provenance-native memory: security becomes a property of the memory itself, not just a gate in
front of it. Signed on write **and verified on read**, every memory carries an immutable origin
record, and the same trust signals that keep retrieval safe also make it better. Also ships the
authorization hardening that had been sitting unreleased since v2.6.1.

### Added

- **Provenance as a first-class, HMAC-signed record (W2a).** New `memory_provenance` table (1:1
with each memory, migration 0011) captures the origin channel and kind (the C1–C4 mapping), the
producing agent and API key, the source run/interaction/trajectory, the derivation depth, the
taint set (trust label + the detections that fired), and the full content-security verdict plus a
`policy_version` that admitted it. Written once through one choke point (`MemoryRepository.add`)
and never mutated — promotions are recorded as events (`SCOPE_CHANGED` / `TRUST_CHANGED` /
`PROMOTED` / `DEDUPLICATED`). New `GET /memories/{id}/provenance` returns the record, its HMAC
verification status, and promotion history; `policy_version` is surfaced by `GET /security/config`.
- **Verify integrity on read (W2b).** Signing on write is now matched by verification on read:
every retrieval path (`query`, `hybrid_query`, `GET /memories/{id}`, handoff, typed
timeline/entity, the ACE playbook, and the context bundle) recomputes the HMAC. `INTEGRITY_READ_MODE`
(`off` / `warn` / `drop`, default **`drop`**) excludes a tamper-detected row before it reaches a
prompt and emits an `INTEGRITY_FAILED` event (on a dedicated write session, so it is committed
even on the replica-safe read routes). Unsigned/legacy rows are kept until
`INTEGRITY_REQUIRE_SIGNED=true`, so `drop` is safe to run before the backfill.
- **Trust-weighted retrieval ranking (W2c).** With `ENABLE_TRUST_WEIGHTED_RANKING=true`, retrieval
fuses vector similarity with content trust, effectiveness votes, temporal decay, and provenance
depth (`RANKING_W_*`, validated to sum to 1.0) instead of ordering by distance alone — closing the
ACE loop so a helpful vote actually raises a memory next time, with the same signal ranking a
low-trust or poisoned write down. Off by default; unproven signals use neutral priors so the
un-voted corpus is never buried. `relevance_score` now reflects the fused score.
- **v2 integrity hash + backfill.** The signature now covers `scope` and `trust_level` (v1 covered
only content, so a direct DB scope-flip verified clean), is delimited and domain-separated, and is
stored with a `v2:` prefix; `verify_integrity` still accepts v1 for legacy rows. `add_batch` now
signs (it never did), consolidation re-signs the merged keeper, and PATCH re-signs on a trust
relabel. `scripts/backfill_integrity.py` upgrades unsigned/v1 rows (idempotent, `--dry-run`,
`--project-id`). Migration 0010 widens `integrity_hash`.

### Security

- **The ACE routes were an unguarded surface.** The W1 sweeps keyed on `MemoryRepository`, so the
entire `ACERepository` path was invisible: `POST /memories/ace/reflection` wrote unscanned,
unsigned memories straight into `global` scope with a body-supplied `agent_id`, and
`/ace/playbook`, `/ace/playbook/agent`, `/ace/vote`, `/ace/curate`, `/ace/consolidate` had the
same class of hole (spoofed identity / unauthorized reads / an unauthenticated vote-poisoning
channel). All now run the full gate set. The authorization sweeps are re-keyed on memory
creation/access (covering `ACERepository`), and a new test pins the files allowed to construct
`Memory(...)` directly, so this recurrence fails CI on the day it is reintroduced.
- **Injection rejected at global scope (W3.2).** `CONTENT_POLICY_INJECTION_GLOBAL_SCOPE` (default
`reject`) escalates a flagged injection to a hard reject for writes entering `global` scope —
readable by every agent in the project — even when the base injection policy only flags.
- **Consolidation cannot launder tampering.** `consolidate_pair` verifies both inputs and refuses a
mismatched pair, so a database-tampered keeper can no longer be re-signed into a valid HMAC.
- **Fixed** a latent crash: `POST /memories/ace/playbook` omitted a required argument when logging
its query event and would 500 on every call.

### From the previously-unreleased authorization work (W1)

- **`POST /memories/ace/delta` was an unguarded write path.** The first authorization pass covered
`memories.py` and `typed_memory.py`; this route was missed. Its `add` branch wrote `op.content`
straight to `MemoryRepository.add` with **no content-security scan, no `authorize_write`, and an
Expand Down
19 changes: 14 additions & 5 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -256,10 +256,11 @@ so `guard` and the server cannot drift.
## Security capabilities

Aegis implements [OWASP AI Agent Security](https://cheatsheetseries.owasp.org/cheatsheets/AI_Agent_Security_Cheat_Sheet.html)
recommendations natively. Six capabilities, none optional:
recommendations natively. Seven capabilities, none optional:

1. **[4-stage content security pipeline](https://docs.aegismemory.com/guides/security)** — input validation, sensitive-data scanning, prompt-injection detection, and an optional LLM-based injection classifier. On every memory write.
2. **[HMAC-SHA256 integrity signing](https://docs.aegismemory.com/guides/security)** — tamper detection on store, verification on demand. You know if a memory was modified.
2. **[HMAC-SHA256 integrity, signed on write and verified on read](https://docs.aegismemory.com/guides/security)** — the signature covers scope and trust level, not just content, and every retrieval path recomputes it: a memory tampered with directly in the database is dropped before it reaches a prompt, not merely flagged after the fact.
7. **[Provenance-native memory](https://docs.aegismemory.com/guides/security)** — every memory carries an immutable, HMAC-signed origin record: which channel produced it, the untrusted inputs that tainted it, and the exact policy verdict and version that admitted it. Security stops being a gate in front of memory and becomes a property of the memory itself.
3. **[OWASP 4-tier trust hierarchy](https://docs.aegismemory.com/guides/security)** — untrusted, internal, privileged, system. Agents get compromised; Aegis limits the blast radius.
4. **[Cryptographic agent binding](https://docs.aegismemory.com/guides/security)** — every route resolves its project from the authenticated key, and for agent-bound keys its acting agent too. A bound key can't be talked into a request body that says "I'm the admin agent." Unbound project keys act for the whole application — see [Threat model](#threat-model).
5. **[ACE loop](https://docs.aegismemory.com/guides/ace-patterns)** — generation, reflection, curation. Agents that learn from their own mistakes and promote what works.
Expand Down Expand Up @@ -382,6 +383,11 @@ Stanford/SambaNova's research, engineered for production. Your agent made the sa
The ACE loop remembers the fix. Stale memories polluting retrieval? Curation auto-cleans your
playbook.

With [trust-weighted ranking](https://docs.aegismemory.com/guides/security) enabled
(`ENABLE_TRUST_WEIGHTED_RANKING`), the loop closes: effectiveness votes, content trust, decay, and
provenance depth fuse into retrieval order, so a memory the agent found helpful actually surfaces
higher next time — the same signal that ranks a low-trust or poisoned write *down*.

<p align="center">
<img src=".github/ace-loop.svg" alt="The ACE loop: get_playbook (generation) → complete_run auto-votes helpful on success or auto-reflects on failure (reflection) → curate promotes, flags, and consolidates (curation). The curated playbook feeds the next run." width="720">
</p>
Expand Down Expand Up @@ -527,13 +533,13 @@ Pick **Aegis Memory** when most of these are true:

## What's shipped vs roadmap

Everything described above is **shipped and released** on PyPI as of `aegis-memory` v2.6.0
(2026-06-25). No feature in this README is aspirational.
Everything described above is **shipped and released** on PyPI as of `aegis-memory` v2.7.0
(2026-08-02). No feature in this README is aspirational.

| Capability | Status | Since |
|---|---|---|
| 4-stage content security pipeline | ✅ Shipped | core |
| HMAC-SHA256 integrity verification | ✅ Shipped | core |
| HMAC-SHA256 integrity, signed on write + verified on read | ✅ Shipped | v2.7.0 |
| 4-tier trust hierarchy + scope ACLs | ✅ Shipped | core |
| Multi-agent coordination + cross-agent query | ✅ Shipped | core |
| ACE loop (vote / reflection / playbook / curation) | ✅ Shipped | core |
Expand All @@ -544,6 +550,8 @@ Everything described above is **shipped and released** on PyPI as of `aegis-memo
| Sigstore-signed releases | ✅ Shipped | v2.5.2 |
| Claude Code plugin + keyless local MCP mode | ✅ Shipped | v2.6.0 |
| Notebook (`.ipynb`) ingestion + inline fix/verify-loop for `inspect` | ✅ Shipped | v2.6.0 |
| Provenance-native memory (immutable HMAC-signed origin record) | ✅ Shipped | v2.7.0 |
| Trust-weighted retrieval ranking | ✅ Shipped | v2.7.0 |

**Directions we're exploring** (not commitments — track them in
[Discussions](https://github.com/quantifylabs/aegis-memory/discussions) and the
Expand Down Expand Up @@ -649,6 +657,7 @@ kubectl apply -f k8s/
| `OPENAI_API_KEY` | — | For embeddings |
| `AEGIS_API_KEY` | `dev-key` | API authentication |
| `CONTENT_POLICY_INJECTION` | `flag` | `reject` / `redact` / `flag` / `allow` |
| `CONTENT_POLICY_INJECTION_GLOBAL_SCOPE` | `reject` | Injection action for global scope: `reject` / `inherit` |
| `CONTENT_POLICY_SECRETS` | `reject` | `reject` / `redact` / `flag` / `allow` |
| `ENABLE_LLM_INJECTION_CLASSIFIER` | `false` | Enable Stage 4 LLM classifier |
| `INJECTION_CLASSIFIER_MODEL` | `gpt-4o-mini` | Model for injection classification |
Expand Down
2 changes: 1 addition & 1 deletion aegis_memory/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -61,7 +61,7 @@
For more examples, see: https://github.com/quantifylabs/aegis-memory/tree/main/examples
"""

__version__ = "2.6.1"
__version__ = "2.7.0"

# Runtime memory write-gate (the firewall `aegis inspect` points its findings at)
from aegis_memory import guard
Expand Down
24 changes: 24 additions & 0 deletions aegis_memory/security/content_security.py
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,11 @@
from enum import Enum
from typing import Any

# Bumped whenever the detection rules below change; folded into the provenance policy_version so a
# memory records which generation of the scanner admitted it. Kept here (not in config) so it moves
# with the rules and stays byte-identical across the server/wheel copies.
SCANNER_RULES_VERSION = "1"

logger = logging.getLogger(__name__)


Expand Down Expand Up @@ -191,6 +196,9 @@ def __init__(self, settings: Any):
self.policy_pii: str = getattr(settings, "content_policy_pii", "flag")
self.policy_secrets: str = getattr(settings, "content_policy_secrets", "reject")
self.policy_injection: str = getattr(settings, "content_policy_injection", "flag")
# Injection action specifically for content entering global scope. Global is readable by
# every agent, so the default here is stricter ("reject") than the base injection policy.
self.policy_injection_global_scope: str = getattr(settings, "content_policy_injection_global_scope", "reject")

# Stage 4: optional LLM classifier (injected via set_classifier)
self._classifier: InjectionClassifier | None = None
Expand Down Expand Up @@ -338,6 +346,22 @@ async def scan_async(
"""
verdict = self.scan(content, metadata)

# Global-scope injection escalation. scan() has no scope, so this lives here where scope
# is known. A memory entering global scope is readable by every agent in the project, so
# a flagged injection there is escalated to a hard reject when so configured -- ahead of
# the classifier early-return below, because the LLM classifier is off by default and this
# must apply regardless. Mirrors the Stage-4 escalation structure further down.
if (
scope == "global"
and self.policy_injection_global_scope == "reject"
and verdict.allowed
and "injection_flagged" in verdict.flags
):
verdict.action = ContentAction.REJECT
verdict.allowed = False
if "injection_global_scope_rejected" not in verdict.flags:
verdict.flags.append("injection_global_scope_rejected")

# Skip Stage 4 if classifier not configured or verdict already rejected
if self._classifier is None or not verdict.allowed:
return verdict
Expand Down
39 changes: 39 additions & 0 deletions alembic/versions/0010_integrity_hash_v2_width.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,39 @@
"""Widen memories.integrity_hash for v2 prefixed hashes

v2 integrity hashes are stored as ``"v2:" + <64 hex>`` = 67 chars, which no longer fits the
original ``String(64)`` column. Widen to ``String(80)`` (headroom for a future prefix bump).
Prompts/skills/subagents keep v1 (bare 64-hex) hashes, so their columns are untouched.

Revision ID: 0010_integrity_hash_v2_width
Revises: 0009_memory_depth
Create Date: 2026-08-02
"""
from alembic import op
import sqlalchemy as sa


revision = "0010_integrity_hash_v2_width"
down_revision = "0009_memory_depth"
branch_labels = None
depends_on = None


def upgrade() -> None:
op.alter_column(
"memories",
"integrity_hash",
existing_type=sa.String(length=64),
type_=sa.String(length=80),
existing_nullable=True,
)


def downgrade() -> None:
# Safe only once v2 hashes are removed/re-hashed; kept symmetric for CI round-trip.
op.alter_column(
"memories",
"integrity_hash",
existing_type=sa.String(length=80),
type_=sa.String(length=64),
existing_nullable=True,
)
55 changes: 55 additions & 0 deletions alembic/versions/0011_memory_provenance.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,55 @@
"""Provenance as a first-class record: memory_provenance (v3 / W2a)

Immutable 1:1 origin record per memory: origin channel/kind, taint set, admitting policy
verdict + version, producing run/interaction, derivation depth, and an HMAC over the record.

Revision ID: 0011_memory_provenance
Revises: 0010_integrity_hash_v2_width
Create Date: 2026-08-02
"""
from alembic import op
import sqlalchemy as sa
from sqlalchemy.dialects import postgresql


revision = "0011_memory_provenance"
down_revision = "0010_integrity_hash_v2_width"
branch_labels = None
depends_on = None


def upgrade() -> None:
op.create_table(
"memory_provenance",
sa.Column("id", sa.String(length=32), primary_key=True),
sa.Column("memory_id", sa.String(length=32), sa.ForeignKey("memories.id", ondelete="CASCADE"), nullable=False),
sa.Column("project_id", sa.String(length=64), nullable=False),
sa.Column("origin_channel", sa.String(length=32), nullable=False),
sa.Column("origin_kind", sa.String(length=16), nullable=False),
sa.Column("producing_agent_id", sa.String(length=64), nullable=True),
sa.Column("acting_agent_id", sa.String(length=64), nullable=True),
sa.Column("principal", sa.String(length=64), nullable=True),
sa.Column("source_run_id", sa.String(length=64), nullable=True),
sa.Column("source_interaction_id", sa.String(length=32), nullable=True),
sa.Column("source_trajectory_id", sa.String(length=64), nullable=True),
sa.Column("parent_memory_ids", postgresql.JSON(), nullable=False, server_default="[]"),
sa.Column("provenance_depth", sa.Integer(), nullable=False, server_default="0"),
sa.Column("taint_json", postgresql.JSON(), nullable=False, server_default="{}"),
sa.Column("policy_verdict_json", postgresql.JSON(), nullable=False, server_default="{}"),
sa.Column("policy_version", sa.String(length=16), nullable=True),
sa.Column("admitted_trust_level", sa.String(length=16), nullable=True),
sa.Column("admitted_scope", sa.String(length=16), nullable=True),
sa.Column("scope_inferred", sa.Boolean(), nullable=False, server_default=sa.false()),
sa.Column("record_hmac", sa.String(length=80), nullable=True),
sa.Column("created_at", sa.DateTime(timezone=True), nullable=False, server_default=sa.func.now()),
)
op.create_unique_constraint("uq_memory_provenance_memory", "memory_provenance", ["memory_id"])
op.create_index("ix_memory_provenance_project", "memory_provenance", ["project_id"])
op.create_index("ix_memory_provenance_channel", "memory_provenance", ["project_id", "origin_channel"])


def downgrade() -> None:
op.drop_index("ix_memory_provenance_channel", table_name="memory_provenance")
op.drop_index("ix_memory_provenance_project", table_name="memory_provenance")
op.drop_constraint("uq_memory_provenance_memory", "memory_provenance", type_="unique")
op.drop_table("memory_provenance")
2 changes: 2 additions & 0 deletions docs/deployment/production-checklist.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -32,6 +32,8 @@ A checklist for self-hosting Aegis Memory. All settings are environment variable
- [ ] **`ENABLE_INTEGRITY_CHECK=true`** (default) — HMAC tamper detection on store/verify.
- [ ] Review content policy actions (`reject | redact | flag | allow`):
`CONTENT_POLICY_PII`, `CONTENT_POLICY_SECRETS`, `CONTENT_POLICY_INJECTION`.
- [ ] **`CONTENT_POLICY_INJECTION_GLOBAL_SCOPE=reject`** (default) — reject flagged injection on
global-scope writes even when the base injection policy only flags.
- [ ] Set limits appropriate to your data: `CONTENT_MAX_LENGTH`, `METADATA_MAX_DEPTH`,
`METADATA_MAX_KEYS`.
- [ ] Decide whether to enforce trust levels: **`ENABLE_TRUST_LEVELS`** (default `false`).
Expand Down
Loading
Loading