From 6e27aee00798371d29aa5ce7ba8c5865753db0bd Mon Sep 17 00:00:00 2001 From: Cipher208 <269750686+Cipher208@users.noreply.github.com> Date: Thu, 8 Oct 2026 01:47:38 +0200 Subject: [PATCH] chore(release): cut 1.11.0, and stop the release notes from naming people MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit [Unreleased] becomes [1.11.0] - 2026-10-08. It had accumulated 18 `###` headings against 5 canonical ones β€” four separate `### Added`, five `### Changed`, three `### Fixed` β€” because each was appended at the time of writing. Merged into Keep a Changelog order (Added, Changed, Removed, Fixed, Security) with all 57 items kept; verified item-for-item, not by eye. Three items had also been written without a newline before the next `- **`, so two entries rendered as one paragraph each. Fixed. Release plumbing, which is what actually leaked: - release.yml used `generate_release_notes: true`, composing the body from PR titles and author handles. That is where "@Cipher208 and Murat" came from in the dead v1.10.0 draft. The body now comes from the CHANGELOG section for the tag, and the step fails when that section is absent β€” a tag with no written notes should stop the release, not quietly generate a list of names. - release-drafter.yml carried `@$AUTHOR` in `change-template` and a `$CONTRIBUTORS` block. Both removed; the drafter remains a working aid. - The dead v1.10.0 draft (never tagged) is deleted. - bug_report.yml asked for `pip show mcp-ariel-memory`, which resolves nothing: the distribution is `a-memory`, and mcp-ariel-memory is one of three console scripts. Now `pip show a-memory`. Artifacts built and inspected before this commit: wheel 276 files with no tests, docs or logs; sdist 727 files carrying both, and no venv; the wheel is built from the sdist (so the sdist is self-sufficient); installed into a fresh venv it loads config.yaml, all 23 top-level keys, the entry point and all three console scripts. --- .github/ISSUE_TEMPLATE/bug_report.yml | 4 +- .github/release-drafter.yml | 10 +- .github/workflows/release.yml | 23 +++- CHANGELOG.md | 166 +++++++++++++------------- 4 files changed, 110 insertions(+), 93 deletions(-) diff --git a/.github/ISSUE_TEMPLATE/bug_report.yml b/.github/ISSUE_TEMPLATE/bug_report.yml index 218d8cc4..32cb3f3b 100644 --- a/.github/ISSUE_TEMPLATE/bug_report.yml +++ b/.github/ISSUE_TEMPLATE/bug_report.yml @@ -10,8 +10,8 @@ body: - type: input id: version attributes: - label: mcp-ariel-memory version - description: "Output of `pip show mcp-ariel-memory | grep Version`" + label: a-memory version + description: "Output of `pip show a-memory | grep Version`" placeholder: "1.0.0" validations: required: true diff --git a/.github/release-drafter.yml b/.github/release-drafter.yml index 0a0c9985..eddb224c 100644 --- a/.github/release-drafter.yml +++ b/.github/release-drafter.yml @@ -5,9 +5,8 @@ template: | $CHANGES - ## Contributors - - $CONTRIBUTORS + categories: - title: 'πŸš€ Features' @@ -36,7 +35,10 @@ categories: labels: - 'security' -change-template: '- $TITLE @$AUTHOR (#$NUMBER)' +# No `@$AUTHOR`: a commit author is a person, and this repository commits under +# a global git identity that is not meant to appear in release text. The +# `$CONTRIBUTORS` section was removed for the same reason. +change-template: '- $TITLE (#$NUMBER)' change-title-escapes: '\<*_&' version-resolver: major: diff --git a/.github/workflows/release.yml b/.github/workflows/release.yml index 8b681b67..e5e8e99a 100644 --- a/.github/workflows/release.yml +++ b/.github/workflows/release.yml @@ -27,11 +27,32 @@ jobs: pip install build python -m build + # The release body comes from CHANGELOG.md, never from generated notes. + # `generate_release_notes` built it out of PR titles and authors, which + # put a real committer identity into a public release ("@Cipher208 and + # Murat"). Notes written by hand in the CHANGELOG say what changed and + # why; a list of who touched it says neither. Failing here when the + # section is missing is deliberate β€” an empty body is better than a + # silently generated one, and the tag would be mis-built anyway. + - name: Take the release body from CHANGELOG.md + run: | + python - "${{ github.ref_name }}" > RELEASE_BODY.md <<'PY' + import re, sys + tag = sys.argv[1].lstrip("v") + text = open("CHANGELOG.md", encoding="utf-8").read() + m = re.search(rf"^## \[{re.escape(tag)}\][^\n]*\n(.*?)(?=^## \[|\Z)", + text, re.M | re.S) + if not m: + sys.exit(f"CHANGELOG.md has no '## [{tag}]' section β€” add one before tagging") + print(m.group(1).strip()) + PY + test -s RELEASE_BODY.md + - name: Create GitHub Release uses: softprops/action-gh-release@v2 with: files: dist/* - generate_release_notes: true + body_path: RELEASE_BODY.md - name: Publish to PyPI env: diff --git a/CHANGELOG.md b/CHANGELOG.md index 174169b8..6cd8c34c 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -3,17 +3,91 @@ All notable changes to mcp-ariel-memory are documented here. Format follows [Keep a Changelog](https://keepachangelog.com/). -## [Unreleased] +## [1.11.0] - 2026-10-08 -### Security -- **Incidental operator identifiers are gone from the published tree (2026-10-08).** A second pass over the same class the fixture scrub started: names that had leaked into prose rather than into data. Removed from `CHANGELOG.md` (two quoted utterances and the `agent_id` list), `AUDIT_REPORT_20260629.md` (the generator's name), `ARIEL_RULES.md`, `autohooks/examples/dsh.yaml`, one Alembic revision comment, `scripts/purge_foreign_rows.py`, and eleven test files where a name stood in for a role. `pyproject.toml`'s "handled by Lucy" comment became "handled by hand". The names used *as the anonymizer's dictionary* stay β€” `config.yaml` `rag.ru_personas`, `rag/synonyms.py`, `mcp_server/utils/privacy.py`, `tests/test_hooks/test_privacy_ru.py` and the `graph_miners` canon test: those are the privacy product, and deleting them would remove the feature that hides names. Six `docs/compose/` files that were force-added past `.gitignore` are untracked again; they remain on disk, and three tests that cite them in docstrings were never reading them. +### Added +- **`scripts/purge_foreign_rows.py` β€” remove already-distilled foreign memories, by provenance (2026-10-04).** Marking the raw journal `foreign_agent` cleans L0, but recall does not read L0: 69 of the contaminated rows had already been distilled, and they were sitting in `episodes`, `core_memory` and the `epi_nodes`/`epi_edges` graph. This removes them, with an archive. + + The first version selected by joining on ids and was **dangerously wrong**. `episodes.episode_id` and `epi_nodes.node_id` are independent sequences that overlap by value β€” 1402 of 1580 episode ids were simultaneously live node ids β€” and `epi_edges.source_id`/`target_id` reference **nodes**, not episodes. The id join produced "8359 linked rows", not one of which was foreign; deleting on it would have destroyed memories belonging to her. Same class of error as the rest of this list: the number looks plausible and means something else. + + Selection is now provenance-only: episodes via their `raw:` tag, `core_memory` via `metadata.source_raw_id`, graph nodes via `content` matching a foreign episode's summary β€” and a node is removed only when that text is **not** also produced by a surviving episode, so her own memory never leaves with someone else's. Every removal is copied into `archived_memories` first under reason `foreign_agent_contamination`, the same class compression itself archives with, so the operation is reversible by construction. `l0_journal` is never touched: the rows stay as `foreign_agent`, which keeps the hash chain a complete record of what arrived. Guards exit 2 if `MCP_MEMORY_DATA_DIR` and the database do not name the same base, and the report prints before/after counts plus `verify_chain`. + + Applied to her live base: episodes 1580 β†’ 1342, `core_memory` 436 β†’ 427, `epi_nodes` 1448 β†’ 1210, `epi_edges` 42246 β†’ 34759, journal unchanged at 3972. 247 rows archived. Verified independently that no foreign episode or fact remains, that her 1342 own episodes are intact, that the agent layer of `core_memory` is unchanged (33 β†’ 33), that recall still returns her own material, and that the chain has 0 violations. Worth recording why this could not wait for the automatic sweeper: of the 238 foreign episodes, **none** was older than 30 days so none purged today, 61 would have gone in three weeks, and **177 would have stayed forever** on weight β‰₯ 0.3. 7 of the 9 foreign facts had no `expires_at` at all. + +- **A journal-only import, and a second persona's history recovered with it (2026-10-04).** `scripts/l0_import_source.py` reads a source described by an autohooks agent config β€” same table, same `json_path`, same filter, same `dispatch_layer` as the live daemon β€” and writes rows to the raw journal via `capture()` only. Nothing is distilled. That is safe because `capture()` writes `status='received'` and the sole consumer of that status is `features.replay.replay()`, reachable only from `scripts/l0_cli.py replay` by hand; `lifecycle.l0_tiers` states outright that "the received status is NEVER archived or truncated", and the nightly graph miners select `raw_type='user-message'` while an import writes `raw_type='import'`. So a persona can have her history back without flooding L3/L4, and the decision to distil (and over what window) is deferred instead of made under pressure. The script re-checks the "did it distil anything" claim itself, printing `episodes` and `core_memory` before and after. It also **refuses to write** when the active `connection_manager.base_dir` is not the config's `data_dir` β€” not a hypothetical: a migration during this work landed in the shared base because the variable is `MCP_MEMORY_DATA_DIR`, not `MCP_DATA_DIR`, and only the printed `base_dir` caught it. + + Used on the third persona to restore her missing history: 3705 of her substantive messages, of which she had 941 rows total before. Result **3031 new L0 rows** (674 collapsed by content dedup β€” the source stores some messages twice), all 3705 accounted for, `episodes` 1580 β†’ 1580, `core_memory` 436 β†’ 436, **0 chain violations**. Agent-layer rows went 6 β†’ 2250: her own voice had been represented by *six* records. + +- **The intake door now counts what it refuses (2026-10-04).** A guard that runs *before* any write leaves no trace: `memory_dispatch_log` records saves, and every refusal in `auto_save_text` and every middleware block returns before its insert. Measured on a live base: **1419 dispatch rows and not one with `score = 0.0`** β€” the refusals were simply absent. That silence is what made the markdown bug (above) invisible; finding it required re-deriving the loss by hand from the source database, twice. + + New table `memory_dispatch_rejections` (migration `20261004_1300_door_rejections`), one row per refusal carrying `reason`, layer, `source_msg_id` and a 160-character preview. The refused text itself is **not** stored: keeping it would put back exactly what the guard declined to keep, and would make the refusal table a recall source. Writes are best-effort over a short-lived `sqlite3` connection with `timeout=1.0`, not `connection_manager` β€” this is observability beside the pipeline and must not join, or commit inside, the transaction `shared.l0.capture` holds around its hash-chain write. + + Both refusal sites record. The door does it in `_refuse()`, where returning and recording are one function so the verdict and its trace cannot diverge β€” the failure mode six bugs in this repository already share. The middleware pipeline does it where the fifth bug lived, and folds free-form block reasons onto stable tags (`"Rate limit exceeded (100/min)"` β†’ `rate_limit`), keeping unknown reasons as themselves rather than dropping them into a catch-all, since a new block site that cannot be seen is the original bug in miniature. The layer is computed *above* the guards so a refusal is attributed to the layer the save would have used. + + Reading it: `scripts/door_report.py` β€” counts by reason and layer, plus previews, because a count says something changed while the preview says whether the guard was right. `below_importance_threshold` is the designed filter, not a defect; `transcript`, `harness_limit` and `system_injection` are text judgements that can be wrong, and the script marks them as such. `verify_autohooks.py` grew to 26/26 with two rows that refuse a seeded input and assert both that the refusal was recorded *and* that it was not simultaneously logged as a save. 9 new tests, mutation-verified: removing the door's write fails 5, hard-coding the refusal layer fails 1, removing the middleware's write fails 1. + +- **A persona's discarded window was replayed and recovered (2026-10-04).** The owner approved it ("there isn't that much"), so the cursor was rewound to the first recoverable message and the daemon re-read the window with the fixed guards; the rate ceiling was raised to 20000/min for the run β€” the case the fifth fix was written for β€” and restored afterwards. A copy of the live base was taken first. + + Result: **227 of 227 distinct statements recovered, none lost**; 544 L0 rows; **0 chain violations** after 544 inserts. The 227 target texts were built by querying the SOURCE directly and then asking whether each `source_msg_id` was present in `l0_journal` β€” a check that shares no code with the filter it validates, which is the mistake the earlier "0 losses" verification made. The remaining rows of the 585 refused-by-the-old-filter lines are not losses: 197 were `duplicate_l0_block` (second copies of already-captured text, the source storing a message twice) and a few fell below the importance threshold. + + Live counter output right after the run β€” 228 refusals: `duplicate_l0_block` 197, `transcript` 24, `harness_limit` 6, `system_injection` 1. Not one of the 24 `transcript` refusals was the persona's speech, verified by the independent question "does the fixed filter accept this text?" (it did, zero times). + +- **MUSE Β§7.5 integration surface (2026-09-11).** `state_entered` / `state_exited` join KNOWN_EVENTS with agent-layer hook handlers: state lifecycle β†’ L3 episodes tagged `muse_state` (payload: state/intensity/trigger/duration/summary). Available through all three transports (memory_hook tool, POST /api/hooks/{event}, autohooks dispatch) β€” one dispatcher, one registration. Per muse-engine-spec v1.1: L4 facts like "state X was useful for task Y" remain the agent's deliberate think() call; the hook only captures the event. Pre-existing v0 precursors (`state_delta`, `emotion_trigger`, layer=user) coexist untouched. `evolve` now writes timestamped `evolution:` L4 keys β€” personality evolution is a timeline, the literal `agent_evolution` key overwrote every prior shift (auto_save anti-pattern). + +- **Stage 2 Plan C (2026-09-07) β€” tool-surface reorganization + meta-tool layer.** Slot map: `mcp_server/slots.py` maps every tool to one of 12 semantic slots (core/recall/context/episodes/sessions/graph/wiki/insight/write/review/admin/brief), registry cross-check test enforces full coverage and single-group membership. Tier rebuild: new `admin` tier (memory_api_key/backup/cleanup/data/lucidity_purge/saga/sync_replica β€” before C these were orphans visible only with `ARIEL_EXPOSE=all`); `memory_skill_promote` moved to `write`; tier `brief` dissolved into `review` (`daily_brief` joins review; legacy strings keep working β€” unknown tiers are ignored); orphans = 0 is now an executable invariant. `wake_up` primitive (E10): one-call session lift β€” memory_recap continuity pack + inject critical set on a shared token budget (recap ≀ half, inject the rest), markdown render with `cache:break`, read-only annotations. Tool count 65 β†’ 66. Presets: `ARIEL_EXPOSE=agent` (expands to primitives,context,insight,write,wiki,review β€” measured 59), `operator` (agent+admin β€” 66), `full` (all); preset expansion runs to a fixpoint so nested references work. Meta-tool layer behind `ARIEL_META=1` (default OFF β€” flat surface unchanged for eval and existing clients): six dispatchers (context/insight/write/wiki/review/admin) are generated programmatically from `EXTRA_TIERS` (single source of truth); each exposes a compact `(action, args)` wire schema (SDK-verified), `action='list'` returns the catalog with per-tool first-line description + Plan-B behavior hints + slot; `action=''` forwards with `user_id` scoping preserved. With all tiers on the visible surface collapses from 66 flat schemas to 13 (7 primitives + 6 dispatchers). Docs: `docs/tools/exposure.md` manifest (slots/tiers/presets/grammar) + reference.md surface counts corrected (the old "recommended combo β†’ 65 tools" line was wrong β€” it exposed 57). Live-agent configs are NOT migrated in C β€” flipping them to `agent` + `ARIEL_META=1` is a separate owner decision. +- **Stage 2 Plan A+B (2026-09-07).** Plan A β€” ariel URI scheme: `ariel:////` (fact/wiki/graph-node/episode/l0) built over existing keys, zero migrations; `shared/uris.py` (builders + `parse_uri` with strict validation and reserved `peer/` namespace + async `resolve_uri` where `user_id` is a required argument, so isolation holds β€” a foreign user's URI does not resolve); search items and `wiki_read` carry `uri`; `drill_down` accepts an ariel-URI in place of entry_id. Plan B β€” MCP behavior annotations: a single 65/65-tool map (`mcp_server/annotations.py`, registry cross-check test) with conservative MCP default (unknown tool = not read-only, destructive=True); ~30 read tools are read_only+idempotent, destructive ops explicitly flagged (forget/wiki_delete/cleanup/heal/data/api_key/saga); `server.py` passes annotations into `mcp.tool()` β€” verified on the wire via `list_tools`. Action-mixed tools (memory_history/proposals/backup) annotated by their worst action, honestly. +- **S20 eval run 3 β€” dense embeddings do NOT beat hash (2026-09-06).** LongMemEval-S n=50 with intfloat/multilingual-e5-small (384d, dim matches the MIB config): full_dense acc 0.600 / strict 0.340 / recall@5 0.031 / ndcg 0.558 / precision 0.252 vs full-hash 0.600 / 0.320 / 0.041 / 0.572 / 0.282 β€” the only win is strict +2pp; recall@5, ndcg and precision all DOWN, wall-time Γ—3.5 (CPU encoding of 2443 sessions). Negative control valid (shuffled 0.540 < 0.600). Hypothesis: e5-small without domain fine-tuning does not beat the lexical anchor of the RRF mix (FTS+tags+core exact matches) on this split β€” dense adds semantics where lexical already fired and noise where it didn't; the low recall of runs 1-2 looks sample-bound (50-question stride, evidence tails) rather than vector-bound. Practical outcomes: hash-fallback stays the prod embedding path; dense_per_kind arm and semantic-dedup/crosscheck/HDBSCAN flags stay OFF pending a larger corpus or dense-aware threshold retraining; e5 model cached in HF cache for future pilots. Infra: `scripts/run_lme_dense.py` β€” fail-fast runner (aborts if `_get_model()` returns None instead of silently degrading to hash; the `ARIEL_HASH_EMBEDDINGS=1` env is set inheritably for the whole session by the ariel-inject hook β€” the runner strips it), one run_eval per process (aiosqlite worker-thread race). +- **S19 miner tails + B6 post-eval (2026-09-06).** RU morphology via pymorphy3 (declared as a main dep): lemma normalization in `_canon_tokens` (topic_overlap, sessions binding, wire_new_node, embedding crosscheck) and in distiller `_canonical_key`, behind `rag.lemmatize` (default true). Guards: pymorphy score <0.3 (out-of-dictionary borrowings β€” Β«Π΄Π΅ΠΏΠ»ΠΎΠΉΒ»β†’Β«Π΄Π΅ΠΏΠ»Ρ‹ΠΉΒ») and common-prefix <4 reject the lemma, caller keeps the raw token; lemma shorter than 4 chars discarded; degradation is graceful when the lib is absent. Key-drift warning documented in config.yaml (old keys are not re-hashed). Test fixtures adapted β€” «сборка/сборку» is no longer an anti-case, that's the feature. B6 post-eval with live numbers from 3 deployed instances: false-merge NOT confirmed (0 duplicate person/org/entity groups) β€” UPDATE/MERGE skipped, `find_or_add_entity` exact-dedup is sufficient; bloat CONFIRMED β€” a multi-topic dump collected 109 co_mentions of 137 instance edges (one synonym class matched everything) β†’ degree cap `_CO_MENTIONS_TOPK=12` per node on the embedding miner precedent. Gate 1469/0, mypy clean. +- **S17 supplements 8-11 (start-FGH tail, 2026-09-06).** (8) `drill_down` survives cold archiving β€” when the L0 row has been moved to `l0_cold_archive` by tier_l0, the provenance reader falls back to the archive by `source_raw_id` and marks the result `archived: true` (was a dead end). (9) Confirming layer for the embedding miner β€” `graph.embedding_crosscheck` (default OFF): a `semantic_overlap` candidate pair is cross-checked against the keyword signal (shared canon-tokens or shared tags); agreement raises the edge weight to 0.6, no lexical trace drops the edge entirely (on hash vectors "vector-similar, lexically absent" is noise). (10) Junk-vector anomaly detector β€” degenerate MIB vectors (0 bits: token-less dump text) are flagged `anomaly:junk_vector` in epi_tags with an `anomalies: N` miner counter; anomaly tags are excluded from rich-embedding text so the flag never perturbs the node vector; density heuristics deliberately NOT added (no validation data β€” deterministic degeneracy only). (11) HDBSCAN clustering over the 384-bit MIB vectors (sklearn, precomputed Hamming) with a louvain-agreement ratio per cluster, reported in `graph_enrich` under `miners.embedding.clusters` behind `graph.embed_clusters` (default OFF, report-only, <12 cached vectors β†’ skipped; meaningful with dense embeddings, noise on the hash fallback). Gate 1463/0, mypy 227 clean. +- **S19.2 textcat pilot (2-class clause stability, default OFF).** Self-labeled from live core_memory across all 5 instances (310 rows: 48 stable / 262 ephemeral β€” the design doc's "training data already exists" was wrong in scale: L4 = 448 total, fact 81%, individual non-fact kinds in single digits, so the 14-class variant is untrainable). Stable = kinds with decay ≀ 0.005 (the same invariant/event line as `route_kind`); frozen ru_core_news_sm tok2vec + trainable textcat head, minority oversampling, threshold sweep. Integration: ONLY the keyword-map-missed FACT bucket β€” a confident 'stable' (β‰₯0.9) promotes the clause from L3 to L4 via `shared/textcat.route_promote_stable`; keyword matches and sub-threshold predictions untouched; breaker `textcat_model` 3/60s; `rag.textcat` default OFF. Verdict: NOT promoted to prod β€” 5-fold CV keyword P 0.636/R 0.352 vs keyword+textcat P 0.485/R 0.484 (+13pp recall bought with βˆ’15pp precision; L4 pollution costs more than a missed fact). Re-enable path: retrain at L4 non-fact β‰₯ 300, require 5-fold P β‰₯ 0.75, owner flips the flag. Infra: `shared/textcat_data.py` (collector), `scripts/train_textcat.py` (train + sweep + train_metrics.json), 11 tests. +- **S18-19 wave (15 items, 5 plans, 2026-09-06).** Retrieval rank: per-query ACT-R min-max multipliers `_minmax_actr` [1.0, 1.3] (top fact Γ—1.3, weakest Γ—1.0 β€” neutral floor, not mean); per-block `inject.max_chars` cap (400, default-ON) on all 9 `build_inject_blocks` build points; semantic dedup as 3rd dedup signal (`memory.semantic_dedup`, default OFF) β€” cosine > 0.92 same-kind LIMIT 50, checked BEFORE conflict-check so a paraphrase never births a conflict pair. Conflicts surface: `shared/channels.py` A4 channel granularity (`channel_of`/`canonical_source`, 9 known channels, unknown β†’ "other", prefix convention no migration); 3-option conflict contract via `memory_proposals action="conflict"` (supersede/retain through `ConflictResolver.resolve` β€” group closed, loser archived; annotate merges `{"annotated": ...}` into the EARLIER side's metadata via the deterministic `_canonical_key` link, existing metadata keys survive) β€” tool count stays 65; orphan-anchor GC in `graph_enrich` (`episode:%` fact/question anchors edgeless >7d deleted, runs BEFORE the dream phase β€” REM bridges isolated anchors first, episode tokens give Jaccard 1.0; report key `orphan_gc`); counter-signal pairs canonicalize distiller keys to the CURRENT name (`_canonical_key` unfolds superseded β†’ current before synonym canon) so a renaming falls into the same key and C4 builds the superseded chain itself. Temporal hygiene: B2 is_current view in `CoreMemory.search` β€” earlier-side hidden when the later version exists GLOBALLY (versioned `::vN` key OR same-key scope=later; private-later doesn't close), `include_superseded=True` returns hidden rows, every item carries `is_current`; `changed_since` delta mode in `get_intervals` (valid_from >= threshold, None = full chain). Infra: gap-registry `lifecycle/gap_registry.py` β€” L3 question episodes (cross-user SQL scan, search_by_tag is per-user) + repeated zero-result queries (β‰₯2, S17 journal) β†’ `memory_gaps` table, idempotent by `gap_hash` (check-then-write; rowcount/total_changes unreliable); nightly phase in graph_enrich + report key `gap_registry`; `shared/harness_breaker.py` β€” per-agent circuit for harness adapters, `harness_call` 5 consecutive fails β†’ 60s open, `HarnessUnavailableError` (adapters outside the repo opt in next visit); `_expand_graph(..., edge_exclude=None)` β€” S19 provenance filter: edges tagged `heuristic:` don't expand at graph-expand stage, None = status quo. Wiki pair: `miner_wiki_fact_links` backfills `metadata.wiki_ids` with the page NODE id (idempotent set, no-op save on rerun β€” ledger not bloated; existing metadata keys survive); `wiki_read` returns `related_facts` + `related_count` (reverse walk of wiki_fact_link edges, fact-node content == core_memory.value, private excluded; empty list is the norm). Full verdicts + plan deviations in the S20 design-doc journal. +- **S17 Stage-1 tail (7 lost draft keys + EMA gate, `f93f6c3`).** `deterministic_retrieval` flag (retrieval.deterministic, bypass inhibit/K in edm_rerank); ENGRAM procedural 14th MemoryKind (never_archive/decay 0, «сдСлай Ρ‚Π°ΠΊΒ»-markers, routes to agent-layer L4); AdaptiveRAG pre-gate (27 query-features, 7 groups, `pre_gate_flags` in route_query, `retrieval.pregate` default OFF); A2 `similar_to` advisory in distiller stats + auto_save_text answer; SHA-256 dedup of L0 blocks (content_hash, migration g23, repeat β†’ source rid, chain intact); zero-result miner (`recall_zero_results` journal + MINERS zero_results β†’ question nodes find_or_add, β‰₯2 repeats); counter-signal aliases (`rag.counter_signals`, Γ—0.7 pessimisation in route_query); EMA gate returned `adaptive_threshold.gate` to auto_save_text + config_hash replay. Tests test_s17_stage1_tail.py (24). +- **S19.1 ru-NER on the privacy gate (`2a91ff7`).** ru_core_news_sm 3.8.0 (model wheel via pip URL, not pyproject β€” same policy as en_core_web_sm); cyrillic input β†’ ru-NER (PER/ORG/LOC), latin β†’ en-NER, mixed β†’ both (en-NER skips pure-cyrillic text β€” guard filtered everything anyway); `rag.ru_ner` default on; lazy + circuit breaker `ru_ner_model` 3/60s (model missing β†’ None WITHOUT tripping); empiric garbage guard: single-token PER = noise («Кисонька» β€” contract test caught it), real persons are 2-3 tokens, single-token ORG/LOC legit (Baltschug), numeric spans dropped, ≀4 tokens, β‰₯2 chars. Tests test_privacy_ru_ner.py (6, skipif without model). +- **S20 eval runs (MINI + LongMemEval-S).** Run 1 (MINI n=10, proxy-judge): full arm WINS (acc 1.0, ndcg 0.73, recall 1.0, 1008 tok) vs rrf/dense_per_kind/gated; dense_per_kind dead (0.20/0.00/0.00 β€” ENGRAM schema unreproducible on multi-layer memory without a dense model); reacq clamp in harness. Run 2 (LongMemEval-S n=50 stride, 2443 sessions, hash embeddings): full wins the second dataset (acc 0.600 vs rrf 0.500, strict 0.320 vs 0.200, recall@5 4.1% vs 1.0%, precision 0.282 vs 0.142); **pre-gate verdict AGAINST prod** β€” gated ≑ rrf byte-for-byte (gate matrix = full fan-out for English question queries; the MINI βˆ’8.5% savings were Russian short/enumerative queries) and full_pregate ≑ full (pre-gate trims the fan-out pool but EDM-rerank returns the same top hits; zero construction_tokens savings on S); `retrieval.pregate` stays default OFF. Honest judge caveat: proxy-judge is weakly discriminative on English (full_shuffled 0.540 vs real 0.600; MINI was 1.0β†’0.3) β€” strict judge is the discriminative metric. Infra: eval adapter reads the official xiaowu0162/longmemeval S-file from a local cache (HF card is broken β€” extensionless files), evidence = `answer_session_ids`, `_stride_sample` for category coverage; one run_eval per process + `os._exit(0)` (aiosqlite worker-thread race kills the next run in-process); crash-safe runner `scripts/run_lme_s.py`. +- **Phase E β€” hardening & closure (18 items, 3 waves).** Durability: E1 atomic L1 persistence β€” `ReflexBuffer._save` writes tempβ†’fsyncβ†’`os.replace`, `MemoryLayer` now wires `persist_path` (`/l1__.json`, hostile ids hashed; L1 finally survives restarts), `restore()` persists immediately. E2 circuit breaker β€” restored from commit `538c61b` (purged as unwired dead code in `94a7c52`), modernized, and WIRED: `shared/circuit_breaker.py` (`CircuitBreaker` + `CircuitBreakerRegistry`), `_embedding_breaker` gates the embedding model path β€” 3 consecutive `model.encode` failures open the circuit for 30s, then hash-fallback vectors keep recall serving (cached under `hash-fallback/` so they never masquerade as model embeddings); created through the registry factory so `memory_diagnose` sees it and `memory_heal reset_breakers` works. E7 least privilege β€” `sync_external` rejects external dirs resolving inside the data dir (symlink aliases included), backup `restore()` runs `safe_resolve` on the name before any fs access, error-outcome counting fixed (`result["errors"]`). Operations: E3 `memory_diagnose` (DB quick_check, alembic head, L1-file JSON validity, pending-proposal backlog, breaker states) + `memory_heal` (remigrate / reset_breakers / purge_invalid_l1) β€” insight/write tiers, tools 60β†’62. E5 integrity score β€” `AuditTrail.log_verify` records `{verified, dropped}` per recall; `memory_report_card` returns `card["integrity"]` = `{score: 100*verified/(verified+dropped), verified, dropped}` (`None` = no data). E9 `` + stable prompt prefix β€” inject blocks reorder post-budget (stable `{rehydrate, important}` β†’ marker β†’ dynamic), the marker renders as a bare line; provider prompt caches can now hit the stable prefix. Retrieval: E10 faceted tag queries β€” `memory_query` gains `tags` (episodes-only): `dimension:value` entries, same-dimension OR, dimensions AND via correlated `json_each` EXISTS; legacy `tag` unchanged. E15 memory-kind weights β€” ACT-R core scoring multiplies by a kind table (`instruction`/`rule`/`commitment` Γ—1.1), config `retrieval.kind_weights` merges over defaults; `core.search` returns `memory_kind`; access-count columns deliberately NOT added (frequency stays derived from `recall_useful` β€” no write amplification on the hot path). E11 disclosure triggers β€” `memory_disclose` (63): recall-side rules Β«when text contains X, surface YΒ» (`trigger_keywords` OR, case-insensitive), surfacing as a `triggered` block (0.95) in `recall_protocol` and the session-start inject; feature-private `disclosure_rules` table. Wiring: E16 `memory_pressure` emitter in the autohooks daemon (L1 ring >40, +10 hysteresis) and `context_threshold` in the Hermes plugin (`sync_turn` cumulative bytes, env `ARIEL_CONTEXT_THRESHOLD_BYTES`, re-arms per crossing) + Hermes `on_session_end` β†’ `session_ended` (no-LLM gist β†’ L3 `session_summary`); E14 `on_turn_end` event (same auto_save pipeline, sync result) as a surface for future harness emitters; E17a causal-link producer on `memory_graph_add` (`action="causal"` β†’ idempotent action/outcome nodes + strength edge); E17b `wiki_write` staging kind with a `propose` producer action on `memory_proposals` (apply β†’ `wiki:{path}`, revert deletes the page; WikiManager instance-scoped to the caller's data dir); E17c `revert_transition` β€” undo one consolidation promotion by its `memory_transitions` id (removes the promoted L4 fact, source episode stays, audit-logged). Validation: E18 DREAM markers anchored to message start (`^\s*DREAM:`) β€” mid-text mentions in documents no longer create junk `dream_skill` episodes (6 found in live dirs); E13 post-compaction semantic audit β€” payload-INDEPENDENT (real harnesses send no summary): mean cosine between the pre-compaction episode window and the L4 set rehydrate actually re-injects, logged to `audit_log` (`action='semantic_audit'`), returned in the handler result. Tool count 60 β†’ 63. E2E (9 full pipelines) + chaos/hypothesis suites (state-machine invariants, hostile ids/titles/payloads, crash-truncated L1 files, threading stress) β€” chaos found and fixed 4 real bugs (`_load` schema-drift escape, unbounded persist/waypoint filenames, duplicate facet params). + +- **Phase D E2E + chaos suite** β€” `tests/test_mcp/test_phase_d_e2e.py`: MCP-level flows with a real AppContext across the new tools (session lifecycle β†’ recap; remember β†’ query β†’ quality feedback β†’ fact blame; typed save β†’ query; rules β†’ hook gate β†’ tagged episode; scratchpad β†’ recap β†’ promote). `tests/test_features/test_phase_d_chaos.py`: 24h window boundaries, zero/negative budgets, limit clamping, unicode/injection-safe filters, scratchpad LRU cap, malformed rules.yaml/schemas degradation, 50k-line log bounding, concurrent remember+query races. `tests/test_mcp/test_server_tiers.py`: exact tier sets, live-agent combo, admin surfaces stay hidden. +- **Query DSL + typed schemas + rules engine (Phase D D1.7 + D1.8 + D1.9)** β€” `memory_query` (54): whitelisted filters β†’ parameterized SQL over core_memory/episodes (importance band, key/content LIKE, created_at window, episodes-only tag) β€” power-user analytics with no injection surface. `memory_save_typed` (55): validate structured fields against built-in schemas (`decision`/`error_pattern`/`relationship`) or custom `/schemas/*.yaml`, store as an L4 fact with `metadata.typed` (key `:`). `memory_load_rules` (56): declarative `/rules.yaml` rules (`when_content_contains` β†’ `importance_boost` sum cap 0.3 + episode tags) applied in the `auto_save_text` write gate β€” user-configurable memory behavior without code; mtime-cached, missing file = unchanged behavior. Tool count 53 β†’ 56. +- **Provenance tracking (Phase D D1.6)** β€” `memory_fact_blame` (53): who wrote this fact, when, and why β€” the evidence trail for debugging hallucinations. Rides the EXISTING `core_memory.source` column (no migration): canonical provenance values `user_explicit` (now the default for `MemoryManager.remember`), `staging_promotion`/`episode_promotion` (B1.4 consolidation products), legacy `manual`; `emotion_trigger`/`wiki_import`/`graph_inference` reserved until such core_memory write paths exist. Blame assembles importance_audit history (chunk_id = entry_id) + audit_log events (target_id = entry_id). +- **Tool-output compression + recall verification (Phase D D1.4 + D1.5)** β€” `memory_compress` (52): `log` mode keeps error/warn/fail lines (+ header, collapses consecutive duplicates, caps lines); `code` mode skeletonizes Python via ast (signatures kept, bodies dropped; non-Python falls back to log); `auto` tries code first. D1.5 verification rides the recall protocol: semantic hits with zero meaningful-token overlap against the query are dropped as retrieval noise (graph-expand hits exempt β€” their relevance is structural, not lexical). Ceiling documented: exact-zero filter only, overlap scoring = v2. Tool count 51 β†’ 52. +- **Session continuity + steering hints (Phase D D1.2 + D1.3)** β€” `memory_recap` (50): the /new recovery pack β€” last closed session (summary, topics, state deltas) β†’ pending work (scratchpad, diff gaps, staged proposals) β†’ markers β†’ day digest, budget-capped to ~2K tokens of recovery instead of re-reading raw history; also available as `autohooks recap`. `memory_steering` (51): deterministic route table (8 routes) + RU/EN keyword intent match steering agents to the best tool per intent (max 3 hints; empty query = full table) β€” advisory, the harness decides. Tool count 49 β†’ 51. +- **Counterfactual memory + skill evolution (Phase D D1.20 + D2.4)** β€” `memory_counterfactual` (49): "what could have been" notes (`anchor`/`premise`/`projection`) saved and listed by anchor. Skills now evolve from sessions: `wiki_read` audit-logs `skill_read` rows; the nightly hook's 6th phase (`skill_reinforce`) boosts read skills' importance +0.05 (cap 1.0, self-limiting via file mtime β€” only reads newer than the page's last write count); promotion MERGES into an existing skill page when the episode's title matches (full title or topic prefix before ":") instead of forking duplicates. Tool count 48 β†’ 49. +- **Agent scratchpad + memory quality loop (Phase D D1.15 + D1.19)** β€” `memory_scratchpad` (47): L2.5 working memory between session and episodes β€” the agent writes hypotheses/plans/drafts (cap 20, oldest evicted), entries re-inject at session start as a `scratchpad` block, and `promote` moves agent-judged-useful entries into L3/L4 (the agent is the distiller β€” D2.2 ceiling). `memory_quality` (48): was_useful β†’ score feedback loop β€” `feedback` writes a `recall_useful` audit row (feeds ACT-R frequency, D1.17) with importance Β±0.05 (cap 1.0 / floor 0.05, importance_audit-logged as `agent_feedback`); `report` aggregates per-entry useful counts with current importance. +- **Smart context budget + reflection system (Phase D D1.10 + D1.16)** β€” `memory_get_smart_context` (45): weighted token distribution across memory sources β€” important 30% / relevant 30% / recent 15% / day 15% / ops 10%, each source gets a floor first, leftover redistributes with a 2x-floor ceiling (a fat source can no longer starve the rest, unlike the sequential inject builder). `memory_reflect` (46) + `reflections` table (migration `20260830_1000_d116`): deterministic meta-memories β€” windowed episode counts + recurring topics (no LLM), `action="write"` stores a reflection row, `action="list"` reads them back; the nightly hook's 5th phase writes the daily reflection automatically. +- **Store pipeline + shared skill SSOT (Phase D D2.2 + D2.3)** β€” `memory_skill_promote` tool (44) and `features/skill_pipeline.py`: promote distilled episodes into `skill` wiki pages verbatim with provenance footers, idempotent via a new `skill_promoted` episode tag (`EpisodicMemory.get_by_id`/`add_tag` added); the nightly hook gains a 4th phase (`skill_promotion`) auto-promoting fresh `dream_skill` episodes. Ceiling (documented): no LLM in ariel β€” dream_skill episodes are agent-distilled at write time (C1.12); raw auto_save chatter needs harness-side distillation (v2). D2.3: `scripts/sync_skills.py` β€” a git-versioned skill SSOT (`~/skills-ssot/skill/*.md`) synced into every live agent's wiki via `WikiManager.sync_external` (copy + sha256 dedup, per-agent subprocesses so `connection_manager` picks up each `MCP_MEMORY_DATA_DIR`; aiosqlite `close_all()` before exit per the C1.10 hang lesson). `scripts/sync_skills.py --bootstrap` creates the skeleton; the sync is cron-ready. +- **Skill = Memory (Phase D D2.1)** β€” skills are wiki pages of the new first-class `skill` type: plain Markdown (SKILL.md convention), agent-read, no embeddings for retrieval β€” progressive disclosure through the wiki surface (`wiki_list(wiki_type="skill")` β†’ `wiki_search` β†’ `wiki_read`). New `wiki_read` tool closes the missing read leg (full page content by path; `wiki_list`/`wiki_search` now expose `path`/`snippet` for chaining). Wiki lint gains the `skill_too_large` rule (concept cap: 4KB per skill page). Tool count 41 β†’ 43. The C1.12 `dream_skill` L3 episodes are the mineable promotion trace (store pipeline = D2.2). +- **/recall protocol β€” multi-axis proportional recall (Phase D D1.1)** β€” `features/recall.py::recall_protocol` fuses five axes in concept order: `markers` (dream-marker facts at 0.95 + `dream_skill` episodes β€” outrank everything) β†’ `session` (last-24h L1 chatter + latest session summary) β†’ `semantic` (MultiSourceRAG hybrid hits) β†’ `expand` (B1.6 1-hop graph neighbors) β†’ `day` (last-24h auto-save digest). Proportional: empty query = zero-state (markers + day only, ~3 lines); dedupe across axes (markers' parts pre-registered, first axis wins); budget-capped. New `autohooks recall --query … [--budget N] [--format md|json]` subcommand and the `memory_recall_protocol` MCP tool (operator tier, tool count 41 β†’ 42). The Hermes native plugin's `queue_prefetch`/`prefetch` now drive this protocol per turn β€” every Hermes turn gets multi-axis recall of the memory store. +- **Compaction-aware memory rehydrate (Phase D D3.5)** β€” ariel now learns when an agent's context was compacted: every `post_context_compression` dispatch logs a drift row to a new `compaction_events` table (migration `20260829_1900_d35_rehydrate`), and the session-start inject gains a `rehydrate` block (important L4 facts, score 0.9) when a compaction happened within `rehydrate.window_hours` (default 6h; `rehydrate.enabled` toggles). `autohooks dispatch` accepts `--payload` (JSON object merged into the event context) and `autohooks inject` accepts `--blocks` (kind filter) so harnesses can pull exactly the rehydrate block. MiMoCode plugin: the compaction hook appends the critical set to the summarizer prompt (salvage), the `session.compacted` event logs the drift and arms a one-shot `[ariel rehydrate]` system block; Hermes gateway emits a `compaction` boundary event (new `~/.hermes/hooks/ariel-compaction/` hook) and the rotated session's inject carries the block. +- **Phase C closeout (C1.12 + C1.13 + C1.14)** β€” `DREAM: memory:/fact:/skill:` markers detected in the auto-save pipeline and routed through staging at importance 0.95 (`skill:` also writes an L3 `dream_skill` episode for D2 to mine). `memory_proposals(action="revert")` undoes applied proposals with exact provenance: core_write (forget by key) and archive (`restore_entries` from `archived_memories` back into core β€” the backup already exists by design); consolidation revert deferred (ids not pinned). `memory_report_card(period_hours=24)` β€” operator digest: proposal counts by status + last decisions, auto-save tier sums, open gaps, dream-marker count. Both tools hidden by default; new `ARIEL_EXPOSE=primitives,review` tier opts an instance in. Tool count 40 β†’ 41. **Phase C auto-hooks keystone is COMPLETE.** +- **Staged mutation β€” proposal β†’ review β†’ apply (Phase C C1.11)** β€” risk-tier true staging: L4-destined auto-saves (score β‰₯ 0.8), consolidation staging promotions (hook + lifespan sweep), and the forgetting ritual's archive sweep now create `mutation_proposals` instead of writing immediately (L3/graph/decays stay direct; propose failure falls back to direct write so bookkeeping never loses memory). Pending proposals surface in the session-start inject as a `proposals` block; the decision is one `memory_proposals` tool call (list / decide approve|reject), same surface for agent and operator. 7-day lazy expiry (no cron). Apply executes the exact write the direct path would have (pinned ids/items); every decision audit-logged. `staging.enabled=false` restores direct behavior. Tool count 39 β†’ 40. Migration `20260829_1600_c111`. +- **`memory_watch` operator tool + `post_session_diff` auto-handler (Phase C C1.10)** β€” one server-side event fires after each session end, computes "dispatched vs persisted" gaps from a new `memory_dispatch_log` table, materializes them as L3 `diff_gap` episodes, and the next `session_started` inject surfaces them as a new `gap` block kind. `memory_watch(action=list|add|disable|delete)` is the operator CRUD over `watch_rules` (introspection of the rules ariel already applies via `auto_save_text`; no new behavior). Migration `20260829_1300_d1c110` adds both tables and seeds one default rule. Tool count 38 β†’ 39. Gaps now reach the agent without any explicit call. +- **Universal autohooks runtime (Phase C C1.9)** β€” `python -m autohooks` (package `autohooks/`): per-agent daemon tails each agent's SQLite conversation store and pushes `new_message` events into the hook dispatcher in-process (no HTTP β€” works for stdio agents); `inject` subcommand returns the session-start critical set for harness embedding. Declarative per-agent YAML configs in `autohooks/examples/` (hermes, mimocode, cowagent β€” all live-verified). First-run baseline prevents history replay; cursor persists per batch (at-least-once). Isolation inherited via `MCP_MEMORY_DATA_DIR`; one runtime, one config = new agent. +- **Auto-hooks foundation (Phase C)** β€” external event dispatcher (`hooks/external.py::dispatch_event`) with two transports: `POST /api/hooks/{event}` (HTTP harnesses) and the new `memory_hook` MCP tool (stdio agents; primitive tier, tool count 37β†’38). Seven events with day-one handlers: `session_started` (returns the critical inject block), `session_ended` (session summary β†’ L3), `new_message`/`auto_save_candidate` (`evaluate_importance` heuristic β€” no LLM; score β‰₯ 0.5 β†’ L3 episodic + graph node, β‰₯ 0.8 β†’ also L4 core), `post_context_compression` (rehydrate candidates via retrieval), `context_threshold`/`memory_pressure` (thin advice; decisions stay harness-side). New `POST /api/context-inject` returns the budget-capped critical set (ACT-R top-5 relevant when a query is given + recent L1 (24h) + important facts β‰₯ `inject.important_min`). Graph is now threaded through the hook registry (`fire(..., graph=)`, C1.1). Per-agent isolation inherited (own process + `MCP_MEMORY_DATA_DIR`; D1.13 API-key binding applies to both surfaces). Harness-side daemon (C1.9) is the next step. ### Changed +- **The release body is the CHANGELOG now, not generated notes (2026-10-08).** `release.yml` published releases with `generate_release_notes: true`, which composes the body out of PR titles and author handles β€” the draft it left behind ended in a real committer identity, "@Cipher208 and Murat". A note that lists who touched a change says neither what changed nor why. The workflow now extracts the `## []` section from `CHANGELOG.md` into the release body and **fails loudly** when that section is missing, so a tag with no written notes cannot silently become a generated list. The same author handles came out of `release-drafter.yml` (`@$AUTHOR` in `change-template`, plus a `$CONTRIBUTORS` block): the drafter stays as a working aid, but the published text comes from the file a human wrote. The dead `v1.10.0` draft β€” never tagged, body generated β€” was deleted. + +- **The bug-report form asked for a version from a package that does not exist (2026-10-08).** `pip show mcp-ariel-memory` cannot resolve: the distribution on PyPI is **`a-memory`**, and `mcp-ariel-memory` is the name of a console script (one of three). Anyone following the form got "Package(s) not found". Now `pip show a-memory`. + - **The wheel now ships the package instead of the repository (2026-10-08).** `[tool.hatch.build.targets.wheel]` was `packages = ["."]`, which made the repo root the package root: the 1.9.0 wheel carried `tests/` (115 files), `docs/`, `.github/`, `uv.lock`, the `Dockerfile` and the raw pytest logs β€” **~445 files against 276 now**. `only-include` names the runtime trees explicitly. The obvious risk was dropping something the code reads at import, and it was real: `config.py` loads `config.yaml` from its own directory, so the first attempt installed cleanly and then failed on a missing default config. Caught by installing into a fresh venv rather than trusting the in-repo test run, where the file is always present; `config.yaml` is now included and the install loads all 23 top-level keys. `test_output.txt` / `full_test_output.txt` joined `.gitignore` β€” they were never tracked yet still shipped, because hatch walked the filesystem. +- **`source:` is now optional in the autohooks agent config (2026-10-03).** The block describes a chat database the `daemon` TAILS; every other command (`inject`, `recall`, `recap`, `context`, `dispatch`) is push-driven and never reads it. The check nevertheless sat above the command dispatch, so a push-only platform β€” one that fires its own lifecycle events and has no chat log to poll β€” was refused at `load_config` with `source: must be a mapping` for a driver that command would never touch. Found wiring the DSH house: the ariel-memory server and its MCP tools worked there, while every hook-side call died on the config. Now absent means "this platform pushes events": the daemon fails fast with an actionable message (checked before `build_app_context()`, so the refusal costs milliseconds instead of a full ariel load) and everything else works. An absent block is a choice, a malformed one is a mistake β€” a present `source:` is still validated as strictly as before (unknown keys, missing keys and non-sqlite drivers stay hard errors), because only the first case should be silent. `SqliteSource.from_config` keeps its own guard for direct callers. 5 new tests, including one that pins the boundary by asserting a push command reaches ariel instead of being turned away. + +- **Procedural memory minimal core (Phase D D2.5)** β€” `memory_procedure` (60): HOW-knowledge with execution stats, author-trimmed from Β§13's 8 tables/12 tools/8 hooks down to ONE table + ONE multi-action tool (Stage-1 directive). `save` registers a named step list (branch-regex names, duplicate β†’ error), `use` records an execution outcome β€” counters update and an optional `learned` note APPENDS to notes (the trimmed Β§13.2 learn hook); `success_rate` is computed on read, never stored (no drift); `list` is payload-free, `get` returns full steps, `delete` removes. No execute-engine: procedures are a declarative cheat-sheet β€” the agent reads get() and acts. Discoverability rides the D1.3 steering route (updated: memory_procedure for repeatable procedures, wiki stays for rich skill docs). Lazy `procedural_memory` table (no migration). Write tier. Tool count 59 β†’ 60. + +- **Memory stash (Phase D D1.12)** β€” `memory_stash` (59): git-stash for the working context β€” `save` captures the L1 reflex buffer (role/content/tokens; timestamps refreshed on pop β€” resumed chatter becomes "recent" again) + scratchpad entries under a name in a lazily-created `memory_stash` table and clears both; `pop` restores the exact working set and consumes the row (refuses while the current scratchpad is non-empty β€” stash first, no silent data loss); `list` (payload-free) and `drop` manage the set. L4 facts and sessions are deliberately NOT stashed β€” identity is D1.11 branches / D1.14 snapshots territory; after per-user layer keying the live use case is one agent switching between contexts (project A ↔ project B). Write tier. Tool count 58 β†’ 59. + +- **Memory versioning: snapshots + rollback (Phase D D1.14)** β€” `memory_history` (57) gains four actions: `snapshot_create/snapshot_list/snapshot_restore` (named point-in-time captures of the (layer, user) fact set; restore makes L4 exactly equal the payload β€” upserts through save, deletes for keys absent from the snapshot, every step ledger-traced as `snapshot_restore:`, re-run idempotent) and `rollback` (git revert of ONE ledger mutation: reinstate its pre-state from the new full-row JSON β€” insert β†’ delete the fact, update/delete β†’ restore exact value/importance/kind/expires/source/metadata; legacy pre-d114 rows fall back to value+importance). The A2.2 ledger is amended in migration `20260901_1300_d114` with `old_row_json`/`new_row_json` (full before/after rows β€” typed schemas and provenance survive rollbacks; `list` stays slim, `get` returns full rows) plus the `core_memory_snapshots` table. Tool becomes dual-tier (insight read legs + write tier). Ceilings: entry_id/created_at not restored (blame chain ends at the revert β€” the ledger keeps history); restore not transactional (idempotent re-run); no auto-snapshots; ledger never pruned (scars forever). Tool count 58 β†’ 58 (no new tools). + +- **Memory versioning ledger + branches (A2.2 + Phase D D1.11)** β€” `memory_history` (57): the `core_memory_history` mutation ledger (migration `20260901_1000_a22`) β€” every L4 insert/update/delete recorded with old/new values, importance, a deterministic `commit_hash` and `triggered_by` provenance (rides the D1.6 `source` contract; `list`/`get` read surface; write failures degrade to a warning so memory writes never fail on history). `memory_branch` (58): A/B persona staging for L4 facts β€” a branch is a `@` namespace in the existing `layer` column (zero schema change, invisible to RAG/inject/stats until merged): `create` full-clones the base layer's facts, `write` lands experiment facts branch-locally, `diff` reports added/changed/unchanged, `merge` cherry-picks keys (default: all differing) back into base with `source='branch_merge'` + `triggered_by='branch_merge:'`, `delete`/`list` manage the set. Ceilings (documented): no deletion propagation, no conflict detection, no checkout (inject always follows main β€” switch = merge). D1.14 (snapshot/rollback) reuses both. Tool count 56 β†’ 58. + +- **Coherent tool tiers for ARIEL_EXPOSE** β€” three new opt-in tier groups alongside wiki/brief/review: `context` (recall protocol, recap, smart context, context inject, steering, compression β€” build/recover context), `insight` (query DSL, fact blame, quality, reflect, stats, raw search, episodes/sessions/graph reads β€” read-side analytics), `write` (remember, typed saves, rules, scratchpad, counterfactuals, episode/graph writes, session lifecycle β€” shape memory). Admin surfaces (backup/api_key/cleanup/watch/saga/data/sync/cleanup_purge, memory_forget) deliberately stay out of tiers β€” expose only via `ARIEL_EXPOSE=all`. Recommended live-agent value: `ARIEL_EXPOSE=primitives,context,insight,write,wiki,brief,review`. Default remains `primitives` (6 tools). + ### Removed - **Five files left the tree without leaving the disk (2026-10-08).** `AUDIT_REPORT_20260629.md` audits 18 of what are now 731 files; `ROADMAP.md` opens by declaring itself legacy and frozen against reality; `.ai-memory.toml` and `.codegraph/` are local tool state; `.compose/context/gates.md` is a working note. Each was checked for live references first β€” none had any, beyond `.compose` in docstrings that turned out to mean the operator's `~/.compose`, not this one. They stay locally, named in `.gitignore`. `.codegraph/` is ignored from the root file now instead of by its own nested `.codegraph/.gitignore`, because git honours a nested ignore file and build backends do not β€” the same asymmetry that put a 5.6 GB venv into a source distribution in the previous commit. +- **auto_save staging branch (`hooks/external.py`, 2026-09-11).** The scoreβ‰₯0.8 branch staged raw chat text under the literal core key `"auto_save"` for manual review β€” 52 same-key proposals flooded `mutation_proposals` in a week (288 total; review tier had also been broken, see Fixed), no code reads the `"auto_save"` core key, and approving any one would overwrite the slot with the next. L4 routing stays with the distiller (canonical keys, near-dup, conflict detection) β€” verified live: `routes.l4_saved=1`, zero `mem.remember` calls for the same input that previously staged. Staging remains for deliberate mutations only: dream markers, consolidation promotions, agent-side `propose`, conflict resolution. Backlog hygiene: the 52 live duplicates were rejected in bulk via the real `staging.decide` path (audit-logged); the 1 consolidation proposal stays pending for review. + ### Fixed - **`is_encrypted_blob` answered a security question by sniffing one byte, and was wrong 1.68% of the time (2026-10-08).** The envelope is `nonce(24) || ciphertext`, so its first byte is the first byte of a random nonce β€” which is `{`, `[`, a space or a newline in **4 cases out of 256**. The old test (`head not in (b"{", b"[", b" ", b"\n")`) therefore called real, readable ciphertext "plain JSON"; measured over 4096 draws: **69 misclassifications, 1.68%**. Detection now attempts decryption: a blob is encrypted when it decrypts under the given key, and nothing else counts. Truncated blobs, filler bytes, plain JSON, tampered ciphertext and blobs encrypted under a foreign key all correctly report `False`. @@ -79,100 +153,20 @@ Format follows [Keep a Changelog](https://keepachangelog.com/). **Correction (2026-10-04, later the same day).** The figures "591 of 734 (80%)" in the entries above were measured on a replay window (`id > 61299`) and counted source ROWS, not distinct texts. Re-measured over the whole source with a stated predicate β€” `role IN ('user','assistant') AND LENGTH(TRIM(content)) > 0` β€” the persona's assistant side is 790 rows / 417 distinct texts, of which the guard refused **630 rows / 266 distinct texts (79.7%)**. The source stores a message twice (`compacted=1`, identical `timestamp`, differing `id` and `message_uid`), so "591 messages" overstated the loss 2.3Γ—. The conclusion is unchanged β€” the guard ate ~80% of her voice β€” but the reproducible number is 266 distinct statements, and the predicate is now written down. -### Added -- **`scripts/purge_foreign_rows.py` β€” remove already-distilled foreign memories, by provenance (2026-10-04).** Marking the raw journal `foreign_agent` cleans L0, but recall does not read L0: 69 of the contaminated rows had already been distilled, and they were sitting in `episodes`, `core_memory` and the `epi_nodes`/`epi_edges` graph. This removes them, with an archive. - - The first version selected by joining on ids and was **dangerously wrong**. `episodes.episode_id` and `epi_nodes.node_id` are independent sequences that overlap by value β€” 1402 of 1580 episode ids were simultaneously live node ids β€” and `epi_edges.source_id`/`target_id` reference **nodes**, not episodes. The id join produced "8359 linked rows", not one of which was foreign; deleting on it would have destroyed memories belonging to her. Same class of error as the rest of this list: the number looks plausible and means something else. - - Selection is now provenance-only: episodes via their `raw:` tag, `core_memory` via `metadata.source_raw_id`, graph nodes via `content` matching a foreign episode's summary β€” and a node is removed only when that text is **not** also produced by a surviving episode, so her own memory never leaves with someone else's. Every removal is copied into `archived_memories` first under reason `foreign_agent_contamination`, the same class compression itself archives with, so the operation is reversible by construction. `l0_journal` is never touched: the rows stay as `foreign_agent`, which keeps the hash chain a complete record of what arrived. Guards exit 2 if `MCP_MEMORY_DATA_DIR` and the database do not name the same base, and the report prints before/after counts plus `verify_chain`. - - Applied to her live base: episodes 1580 β†’ 1342, `core_memory` 436 β†’ 427, `epi_nodes` 1448 β†’ 1210, `epi_edges` 42246 β†’ 34759, journal unchanged at 3972. 247 rows archived. Verified independently that no foreign episode or fact remains, that her 1342 own episodes are intact, that the agent layer of `core_memory` is unchanged (33 β†’ 33), that recall still returns her own material, and that the chain has 0 violations. Worth recording why this could not wait for the automatic sweeper: of the 238 foreign episodes, **none** was older than 30 days so none purged today, 61 would have gone in three weeks, and **177 would have stayed forever** on weight β‰₯ 0.3. 7 of the 9 foreign facts had no `expires_at` at all. - -- **A journal-only import, and a second persona's history recovered with it (2026-10-04).** `scripts/l0_import_source.py` reads a source described by an autohooks agent config β€” same table, same `json_path`, same filter, same `dispatch_layer` as the live daemon β€” and writes rows to the raw journal via `capture()` only. Nothing is distilled. That is safe because `capture()` writes `status='received'` and the sole consumer of that status is `features.replay.replay()`, reachable only from `scripts/l0_cli.py replay` by hand; `lifecycle.l0_tiers` states outright that "the received status is NEVER archived or truncated", and the nightly graph miners select `raw_type='user-message'` while an import writes `raw_type='import'`. So a persona can have her history back without flooding L3/L4, and the decision to distil (and over what window) is deferred instead of made under pressure. The script re-checks the "did it distil anything" claim itself, printing `episodes` and `core_memory` before and after. It also **refuses to write** when the active `connection_manager.base_dir` is not the config's `data_dir` β€” not a hypothetical: a migration during this work landed in the shared base because the variable is `MCP_MEMORY_DATA_DIR`, not `MCP_DATA_DIR`, and only the printed `base_dir` caught it. - - Used on the third persona to restore her missing history: 3705 of her substantive messages, of which she had 941 rows total before. Result **3031 new L0 rows** (674 collapsed by content dedup β€” the source stores some messages twice), all 3705 accounted for, `episodes` 1580 β†’ 1580, `core_memory` 436 β†’ 436, **0 chain violations**. Agent-layer rows went 6 β†’ 2250: her own voice had been represented by *six* records. - -- **The intake door now counts what it refuses (2026-10-04).** A guard that runs *before* any write leaves no trace: `memory_dispatch_log` records saves, and every refusal in `auto_save_text` and every middleware block returns before its insert. Measured on a live base: **1419 dispatch rows and not one with `score = 0.0`** β€” the refusals were simply absent. That silence is what made the markdown bug (above) invisible; finding it required re-deriving the loss by hand from the source database, twice. - - New table `memory_dispatch_rejections` (migration `20261004_1300_door_rejections`), one row per refusal carrying `reason`, layer, `source_msg_id` and a 160-character preview. The refused text itself is **not** stored: keeping it would put back exactly what the guard declined to keep, and would make the refusal table a recall source. Writes are best-effort over a short-lived `sqlite3` connection with `timeout=1.0`, not `connection_manager` β€” this is observability beside the pipeline and must not join, or commit inside, the transaction `shared.l0.capture` holds around its hash-chain write. - - Both refusal sites record. The door does it in `_refuse()`, where returning and recording are one function so the verdict and its trace cannot diverge β€” the failure mode six bugs in this repository already share. The middleware pipeline does it where the fifth bug lived, and folds free-form block reasons onto stable tags (`"Rate limit exceeded (100/min)"` β†’ `rate_limit`), keeping unknown reasons as themselves rather than dropping them into a catch-all, since a new block site that cannot be seen is the original bug in miniature. The layer is computed *above* the guards so a refusal is attributed to the layer the save would have used. - - Reading it: `scripts/door_report.py` β€” counts by reason and layer, plus previews, because a count says something changed while the preview says whether the guard was right. `below_importance_threshold` is the designed filter, not a defect; `transcript`, `harness_limit` and `system_injection` are text judgements that can be wrong, and the script marks them as such. `verify_autohooks.py` grew to 26/26 with two rows that refuse a seeded input and assert both that the refusal was recorded *and* that it was not simultaneously logged as a save. 9 new tests, mutation-verified: removing the door's write fails 5, hard-coding the refusal layer fails 1, removing the middleware's write fails 1. - -- **A persona's discarded window was replayed and recovered (2026-10-04).** The owner approved it ("there isn't that much"), so the cursor was rewound to the first recoverable message and the daemon re-read the window with the fixed guards; the rate ceiling was raised to 20000/min for the run β€” the case the fifth fix was written for β€” and restored afterwards. A copy of the live base was taken first. - - Result: **227 of 227 distinct statements recovered, none lost**; 544 L0 rows; **0 chain violations** after 544 inserts. The 227 target texts were built by querying the SOURCE directly and then asking whether each `source_msg_id` was present in `l0_journal` β€” a check that shares no code with the filter it validates, which is the mistake the earlier "0 losses" verification made. The remaining rows of the 585 refused-by-the-old-filter lines are not losses: 197 were `duplicate_l0_block` (second copies of already-captured text, the source storing a message twice) and a few fell below the importance threshold. - - Live counter output right after the run β€” 228 refusals: `duplicate_l0_block` 197, `transcript` 24, `harness_limit` 6, `system_injection` 1. Not one of the 24 `transcript` refusals was the persona's speech, verified by the independent question "does the fixed filter accept this text?" (it did, zero times). - -### Changed -- **`source:` is now optional in the autohooks agent config (2026-10-03).** The block describes a chat database the `daemon` TAILS; every other command (`inject`, `recall`, `recap`, `context`, `dispatch`) is push-driven and never reads it. The check nevertheless sat above the command dispatch, so a push-only platform β€” one that fires its own lifecycle events and has no chat log to poll β€” was refused at `load_config` with `source: must be a mapping` for a driver that command would never touch. Found wiring the DSH house: the ariel-memory server and its MCP tools worked there, while every hook-side call died on the config. Now absent means "this platform pushes events": the daemon fails fast with an actionable message (checked before `build_app_context()`, so the refusal costs milliseconds instead of a full ariel load) and everything else works. An absent block is a choice, a malformed one is a mistake β€” a present `source:` is still validated as strictly as before (unknown keys, missing keys and non-sqlite drivers stay hard errors), because only the first case should be silent. `SqliteSource.from_config` keeps its own guard for direct callers. 5 new tests, including one that pins the boundary by asserting a push command reaches ariel instead of being turned away. - -### Added -- **MUSE Β§7.5 integration surface (2026-09-11).** `state_entered` / `state_exited` join KNOWN_EVENTS with agent-layer hook handlers: state lifecycle β†’ L3 episodes tagged `muse_state` (payload: state/intensity/trigger/duration/summary). Available through all three transports (memory_hook tool, POST /api/hooks/{event}, autohooks dispatch) β€” one dispatcher, one registration. Per muse-engine-spec v1.1: L4 facts like "state X was useful for task Y" remain the agent's deliberate think() call; the hook only captures the event. Pre-existing v0 precursors (`state_delta`, `emotion_trigger`, layer=user) coexist untouched. `evolve` now writes timestamped `evolution:` L4 keys β€” personality evolution is a timeline, the literal `agent_evolution` key overwrote every prior shift (auto_save anti-pattern). - -### Removed -- **auto_save staging branch (`hooks/external.py`, 2026-09-11).** The scoreβ‰₯0.8 branch staged raw chat text under the literal core key `"auto_save"` for manual review β€” 52 same-key proposals flooded `mutation_proposals` in a week (288 total; review tier had also been broken, see Fixed), no code reads the `"auto_save"` core key, and approving any one would overwrite the slot with the next. L4 routing stays with the distiller (canonical keys, near-dup, conflict detection) β€” verified live: `routes.l4_saved=1`, zero `mem.remember` calls for the same input that previously staged. Staging remains for deliberate mutations only: dream markers, consolidation promotions, agent-side `propose`, conflict resolution. Backlog hygiene: the 52 live duplicates were rejected in bulk via the real `staging.decide` path (audit-logged); the 1 consolidation proposal stays pending for review. - -### Fixed - **S2-exhaustive wiki route crashed on live WikiEntry models (2026-09-12).** `rag/dual_route.py::s2_exhaustive` consumed `wiki.list_all()` rows as dicts (`r.get(...)`), but `wiki/manager.list_all` returns `WikiEntry` pydantic models β€” the live `memory_search` hybrid path died with "'WikiEntry' object has no attribute 'get'" on the first real query. The existing `FakeWiki` test returned dicts and masked the live path (env-parity gap, same family as the ARIEL_HASH_EMBEDDINGS bypass). Fix: boundary normalization (`model_dump()` for non-dict rows) + regression test `test_s2_route_list_all_wiki_models` (red with the exact prod error, then green). P7.1 pattern sweep: the only other manager-level `list_all` consumers (ops.py inject builders, wiki.py list) already use attribute access β€” clean. - **proposal dedup-guard + surface-aware review hints (2026-09-11 follow-up).** `staging.propose` absorbs a pending same-identity proposal (source/kind/user/layer + payload key|title) instead of adding a row β€” latest payload wins, TTL refreshed; distinct identities (different key/user/kind) stay granular. `staging.decision_hint()` renders the review-call hint for the ACTIVE surface β€” the inject proposals header and the steering route table taught the flat `memory_proposals(action=…)` shape even under `ARIEL_META=1`, where that flat tool does not exist for the agent (same failure family as the blind-args errors). 4 new tests. - **review tier meta-tool dispatch was unusable (2026-09-11 incident).** Three defects: (1) dispatcher `ctx: Any` β€” the SDK injects Context only into Context-annotated params (`find_context_parameter` β†’ `get_type_hints`), so ctx stayed `None` and every member calling `_get_ctx` raised "Context is required but was None"; now Context-typed with a runtime import (TYPE_CHECKING-only names break annotation resolution). (2) `memory_proposals` called `_get_ctx` unconditionally, killing `action="list"` which needs no app context; strict only in the mutate branches (decide/revert), mirroring `memory_report_card`. (3) the meta catalog didn't surface member parameter names β€” the free-form `args` dict made agents guess (`unexpected keyword argument context` was a blind guess); catalog now carries `params` (ctx excluded). 4 new tests; the flat (non-meta) surface is unchanged. -### Added -- **Stage 2 Plan C (2026-09-07) β€” tool-surface reorganization + meta-tool layer.** Slot map: `mcp_server/slots.py` maps every tool to one of 12 semantic slots (core/recall/context/episodes/sessions/graph/wiki/insight/write/review/admin/brief), registry cross-check test enforces full coverage and single-group membership. Tier rebuild: new `admin` tier (memory_api_key/backup/cleanup/data/lucidity_purge/saga/sync_replica β€” before C these were orphans visible only with `ARIEL_EXPOSE=all`); `memory_skill_promote` moved to `write`; tier `brief` dissolved into `review` (`daily_brief` joins review; legacy strings keep working β€” unknown tiers are ignored); orphans = 0 is now an executable invariant. `wake_up` primitive (E10): one-call session lift β€” memory_recap continuity pack + inject critical set on a shared token budget (recap ≀ half, inject the rest), markdown render with `cache:break`, read-only annotations. Tool count 65 β†’ 66. Presets: `ARIEL_EXPOSE=agent` (expands to primitives,context,insight,write,wiki,review β€” measured 59), `operator` (agent+admin β€” 66), `full` (all); preset expansion runs to a fixpoint so nested references work. Meta-tool layer behind `ARIEL_META=1` (default OFF β€” flat surface unchanged for eval and existing clients): six dispatchers (context/insight/write/wiki/review/admin) are generated programmatically from `EXTRA_TIERS` (single source of truth); each exposes a compact `(action, args)` wire schema (SDK-verified), `action='list'` returns the catalog with per-tool first-line description + Plan-B behavior hints + slot; `action=''` forwards with `user_id` scoping preserved. With all tiers on the visible surface collapses from 66 flat schemas to 13 (7 primitives + 6 dispatchers). Docs: `docs/tools/exposure.md` manifest (slots/tiers/presets/grammar) + reference.md surface counts corrected (the old "recommended combo β†’ 65 tools" line was wrong β€” it exposed 57). Live-agent configs are NOT migrated in C β€” flipping them to `agent` + `ARIEL_META=1` is a separate owner decision. -- **Stage 2 Plan A+B (2026-09-07).** Plan A β€” ariel URI scheme: `ariel:////` (fact/wiki/graph-node/episode/l0) built over existing keys, zero migrations; `shared/uris.py` (builders + `parse_uri` with strict validation and reserved `peer/` namespace + async `resolve_uri` where `user_id` is a required argument, so isolation holds β€” a foreign user's URI does not resolve); search items and `wiki_read` carry `uri`; `drill_down` accepts an ariel-URI in place of entry_id. Plan B β€” MCP behavior annotations: a single 65/65-tool map (`mcp_server/annotations.py`, registry cross-check test) with conservative MCP default (unknown tool = not read-only, destructive=True); ~30 read tools are read_only+idempotent, destructive ops explicitly flagged (forget/wiki_delete/cleanup/heal/data/api_key/saga); `server.py` passes annotations into `mcp.tool()` β€” verified on the wire via `list_tools`. Action-mixed tools (memory_history/proposals/backup) annotated by their worst action, honestly.- **S20 eval run 3 β€” dense embeddings do NOT beat hash (2026-09-06).** LongMemEval-S n=50 with intfloat/multilingual-e5-small (384d, dim matches the MIB config): full_dense acc 0.600 / strict 0.340 / recall@5 0.031 / ndcg 0.558 / precision 0.252 vs full-hash 0.600 / 0.320 / 0.041 / 0.572 / 0.282 β€” the only win is strict +2pp; recall@5, ndcg and precision all DOWN, wall-time Γ—3.5 (CPU encoding of 2443 sessions). Negative control valid (shuffled 0.540 < 0.600). Hypothesis: e5-small without domain fine-tuning does not beat the lexical anchor of the RRF mix (FTS+tags+core exact matches) on this split β€” dense adds semantics where lexical already fired and noise where it didn't; the low recall of runs 1-2 looks sample-bound (50-question stride, evidence tails) rather than vector-bound. Practical outcomes: hash-fallback stays the prod embedding path; dense_per_kind arm and semantic-dedup/crosscheck/HDBSCAN flags stay OFF pending a larger corpus or dense-aware threshold retraining; e5 model cached in HF cache for future pilots. Infra: `scripts/run_lme_dense.py` β€” fail-fast runner (aborts if `_get_model()` returns None instead of silently degrading to hash; the `ARIEL_HASH_EMBEDDINGS=1` env is set inheritably for the whole session by the ariel-inject hook β€” the runner strips it), one run_eval per process (aiosqlite worker-thread race).- **S19 miner tails + B6 post-eval (2026-09-06).** RU morphology via pymorphy3 (declared as a main dep): lemma normalization in `_canon_tokens` (topic_overlap, sessions binding, wire_new_node, embedding crosscheck) and in distiller `_canonical_key`, behind `rag.lemmatize` (default true). Guards: pymorphy score <0.3 (out-of-dictionary borrowings β€” Β«Π΄Π΅ΠΏΠ»ΠΎΠΉΒ»β†’Β«Π΄Π΅ΠΏΠ»Ρ‹ΠΉΒ») and common-prefix <4 reject the lemma, caller keeps the raw token; lemma shorter than 4 chars discarded; degradation is graceful when the lib is absent. Key-drift warning documented in config.yaml (old keys are not re-hashed). Test fixtures adapted β€” «сборка/сборку» is no longer an anti-case, that's the feature. B6 post-eval with live numbers from 3 deployed instances: false-merge NOT confirmed (0 duplicate person/org/entity groups) β€” UPDATE/MERGE skipped, `find_or_add_entity` exact-dedup is sufficient; bloat CONFIRMED β€” a multi-topic dump collected 109 co_mentions of 137 instance edges (one synonym class matched everything) β†’ degree cap `_CO_MENTIONS_TOPK=12` per node on the embedding miner precedent. Gate 1469/0, mypy clean.- **S17 supplements 8-11 (start-FGH tail, 2026-09-06).** (8) `drill_down` survives cold archiving β€” when the L0 row has been moved to `l0_cold_archive` by tier_l0, the provenance reader falls back to the archive by `source_raw_id` and marks the result `archived: true` (was a dead end). (9) Confirming layer for the embedding miner β€” `graph.embedding_crosscheck` (default OFF): a `semantic_overlap` candidate pair is cross-checked against the keyword signal (shared canon-tokens or shared tags); agreement raises the edge weight to 0.6, no lexical trace drops the edge entirely (on hash vectors "vector-similar, lexically absent" is noise). (10) Junk-vector anomaly detector β€” degenerate MIB vectors (0 bits: token-less dump text) are flagged `anomaly:junk_vector` in epi_tags with an `anomalies: N` miner counter; anomaly tags are excluded from rich-embedding text so the flag never perturbs the node vector; density heuristics deliberately NOT added (no validation data β€” deterministic degeneracy only). (11) HDBSCAN clustering over the 384-bit MIB vectors (sklearn, precomputed Hamming) with a louvain-agreement ratio per cluster, reported in `graph_enrich` under `miners.embedding.clusters` behind `graph.embed_clusters` (default OFF, report-only, <12 cached vectors β†’ skipped; meaningful with dense embeddings, noise on the hash fallback). Gate 1463/0, mypy 227 clean. -- **S19.2 textcat pilot (2-class clause stability, default OFF).** Self-labeled from live core_memory across all 5 instances (310 rows: 48 stable / 262 ephemeral β€” the design doc's "training data already exists" was wrong in scale: L4 = 448 total, fact 81%, individual non-fact kinds in single digits, so the 14-class variant is untrainable). Stable = kinds with decay ≀ 0.005 (the same invariant/event line as `route_kind`); frozen ru_core_news_sm tok2vec + trainable textcat head, minority oversampling, threshold sweep. Integration: ONLY the keyword-map-missed FACT bucket β€” a confident 'stable' (β‰₯0.9) promotes the clause from L3 to L4 via `shared/textcat.route_promote_stable`; keyword matches and sub-threshold predictions untouched; breaker `textcat_model` 3/60s; `rag.textcat` default OFF. Verdict: NOT promoted to prod β€” 5-fold CV keyword P 0.636/R 0.352 vs keyword+textcat P 0.485/R 0.484 (+13pp recall bought with βˆ’15pp precision; L4 pollution costs more than a missed fact). Re-enable path: retrain at L4 non-fact β‰₯ 300, require 5-fold P β‰₯ 0.75, owner flips the flag. Infra: `shared/textcat_data.py` (collector), `scripts/train_textcat.py` (train + sweep + train_metrics.json), 11 tests. -- **S18-19 wave (15 items, 5 plans, 2026-09-06).** Retrieval rank: per-query ACT-R min-max multipliers `_minmax_actr` [1.0, 1.3] (top fact Γ—1.3, weakest Γ—1.0 β€” neutral floor, not mean); per-block `inject.max_chars` cap (400, default-ON) on all 9 `build_inject_blocks` build points; semantic dedup as 3rd dedup signal (`memory.semantic_dedup`, default OFF) β€” cosine > 0.92 same-kind LIMIT 50, checked BEFORE conflict-check so a paraphrase never births a conflict pair. Conflicts surface: `shared/channels.py` A4 channel granularity (`channel_of`/`canonical_source`, 9 known channels, unknown β†’ "other", prefix convention no migration); 3-option conflict contract via `memory_proposals action="conflict"` (supersede/retain through `ConflictResolver.resolve` β€” group closed, loser archived; annotate merges `{"annotated": ...}` into the EARLIER side's metadata via the deterministic `_canonical_key` link, existing metadata keys survive) β€” tool count stays 65; orphan-anchor GC in `graph_enrich` (`episode:%` fact/question anchors edgeless >7d deleted, runs BEFORE the dream phase β€” REM bridges isolated anchors first, episode tokens give Jaccard 1.0; report key `orphan_gc`); counter-signal pairs canonicalize distiller keys to the CURRENT name (`_canonical_key` unfolds superseded β†’ current before synonym canon) so a renaming falls into the same key and C4 builds the superseded chain itself. Temporal hygiene: B2 is_current view in `CoreMemory.search` β€” earlier-side hidden when the later version exists GLOBALLY (versioned `::vN` key OR same-key scope=later; private-later doesn't close), `include_superseded=True` returns hidden rows, every item carries `is_current`; `changed_since` delta mode in `get_intervals` (valid_from >= threshold, None = full chain). Infra: gap-registry `lifecycle/gap_registry.py` β€” L3 question episodes (cross-user SQL scan, search_by_tag is per-user) + repeated zero-result queries (β‰₯2, S17 journal) β†’ `memory_gaps` table, idempotent by `gap_hash` (check-then-write; rowcount/total_changes unreliable); nightly phase in graph_enrich + report key `gap_registry`; `shared/harness_breaker.py` β€” per-agent circuit for harness adapters, `harness_call` 5 consecutive fails β†’ 60s open, `HarnessUnavailableError` (adapters outside the repo opt in next visit); `_expand_graph(..., edge_exclude=None)` β€” S19 provenance filter: edges tagged `heuristic:` don't expand at graph-expand stage, None = status quo. Wiki pair: `miner_wiki_fact_links` backfills `metadata.wiki_ids` with the page NODE id (idempotent set, no-op save on rerun β€” ledger not bloated; existing metadata keys survive); `wiki_read` returns `related_facts` + `related_count` (reverse walk of wiki_fact_link edges, fact-node content == core_memory.value, private excluded; empty list is the norm). Full verdicts + plan deviations in the S20 design-doc journal. -- **S17 Stage-1 tail (7 lost draft keys + EMA gate, `f93f6c3`).** `deterministic_retrieval` flag (retrieval.deterministic, bypass inhibit/K in edm_rerank); ENGRAM procedural 14th MemoryKind (never_archive/decay 0, «сдСлай Ρ‚Π°ΠΊΒ»-markers, routes to agent-layer L4); AdaptiveRAG pre-gate (27 query-features, 7 groups, `pre_gate_flags` in route_query, `retrieval.pregate` default OFF); A2 `similar_to` advisory in distiller stats + auto_save_text answer; SHA-256 dedup of L0 blocks (content_hash, migration g23, repeat β†’ source rid, chain intact); zero-result miner (`recall_zero_results` journal + MINERS zero_results β†’ question nodes find_or_add, β‰₯2 repeats); counter-signal aliases (`rag.counter_signals`, Γ—0.7 pessimisation in route_query); EMA gate returned `adaptive_threshold.gate` to auto_save_text + config_hash replay. Tests test_s17_stage1_tail.py (24). -- **S19.1 ru-NER on the privacy gate (`2a91ff7`).** ru_core_news_sm 3.8.0 (model wheel via pip URL, not pyproject β€” same policy as en_core_web_sm); cyrillic input β†’ ru-NER (PER/ORG/LOC), latin β†’ en-NER, mixed β†’ both (en-NER skips pure-cyrillic text β€” guard filtered everything anyway); `rag.ru_ner` default on; lazy + circuit breaker `ru_ner_model` 3/60s (model missing β†’ None WITHOUT tripping); empiric garbage guard: single-token PER = noise («Кисонька» β€” contract test caught it), real persons are 2-3 tokens, single-token ORG/LOC legit (Baltschug), numeric spans dropped, ≀4 tokens, β‰₯2 chars. Tests test_privacy_ru_ner.py (6, skipif without model). -- **S20 eval runs (MINI + LongMemEval-S).** Run 1 (MINI n=10, proxy-judge): full arm WINS (acc 1.0, ndcg 0.73, recall 1.0, 1008 tok) vs rrf/dense_per_kind/gated; dense_per_kind dead (0.20/0.00/0.00 β€” ENGRAM schema unreproducible on multi-layer memory without a dense model); reacq clamp in harness. Run 2 (LongMemEval-S n=50 stride, 2443 sessions, hash embeddings): full wins the second dataset (acc 0.600 vs rrf 0.500, strict 0.320 vs 0.200, recall@5 4.1% vs 1.0%, precision 0.282 vs 0.142); **pre-gate verdict AGAINST prod** β€” gated ≑ rrf byte-for-byte (gate matrix = full fan-out for English question queries; the MINI βˆ’8.5% savings were Russian short/enumerative queries) and full_pregate ≑ full (pre-gate trims the fan-out pool but EDM-rerank returns the same top hits; zero construction_tokens savings on S); `retrieval.pregate` stays default OFF. Honest judge caveat: proxy-judge is weakly discriminative on English (full_shuffled 0.540 vs real 0.600; MINI was 1.0β†’0.3) β€” strict judge is the discriminative metric. Infra: eval adapter reads the official xiaowu0162/longmemeval S-file from a local cache (HF card is broken β€” extensionless files), evidence = `answer_session_ids`, `_stride_sample` for category coverage; one run_eval per process + `os._exit(0)` (aiosqlite worker-thread race kills the next run in-process); crash-safe runner `scripts/run_lme_s.py`. -- **Phase E β€” hardening & closure (18 items, 3 waves).** Durability: E1 atomic L1 persistence β€” `ReflexBuffer._save` writes tempβ†’fsyncβ†’`os.replace`, `MemoryLayer` now wires `persist_path` (`/l1__.json`, hostile ids hashed; L1 finally survives restarts), `restore()` persists immediately. E2 circuit breaker β€” restored from commit `538c61b` (purged as unwired dead code in `94a7c52`), modernized, and WIRED: `shared/circuit_breaker.py` (`CircuitBreaker` + `CircuitBreakerRegistry`), `_embedding_breaker` gates the embedding model path β€” 3 consecutive `model.encode` failures open the circuit for 30s, then hash-fallback vectors keep recall serving (cached under `hash-fallback/` so they never masquerade as model embeddings); created through the registry factory so `memory_diagnose` sees it and `memory_heal reset_breakers` works. E7 least privilege β€” `sync_external` rejects external dirs resolving inside the data dir (symlink aliases included), backup `restore()` runs `safe_resolve` on the name before any fs access, error-outcome counting fixed (`result["errors"]`). Operations: E3 `memory_diagnose` (DB quick_check, alembic head, L1-file JSON validity, pending-proposal backlog, breaker states) + `memory_heal` (remigrate / reset_breakers / purge_invalid_l1) β€” insight/write tiers, tools 60β†’62. E5 integrity score β€” `AuditTrail.log_verify` records `{verified, dropped}` per recall; `memory_report_card` returns `card["integrity"]` = `{score: 100*verified/(verified+dropped), verified, dropped}` (`None` = no data). E9 `` + stable prompt prefix β€” inject blocks reorder post-budget (stable `{rehydrate, important}` β†’ marker β†’ dynamic), the marker renders as a bare line; provider prompt caches can now hit the stable prefix. Retrieval: E10 faceted tag queries β€” `memory_query` gains `tags` (episodes-only): `dimension:value` entries, same-dimension OR, dimensions AND via correlated `json_each` EXISTS; legacy `tag` unchanged. E15 memory-kind weights β€” ACT-R core scoring multiplies by a kind table (`instruction`/`rule`/`commitment` Γ—1.1), config `retrieval.kind_weights` merges over defaults; `core.search` returns `memory_kind`; access-count columns deliberately NOT added (frequency stays derived from `recall_useful` β€” no write amplification on the hot path). E11 disclosure triggers β€” `memory_disclose` (63): recall-side rules Β«when text contains X, surface YΒ» (`trigger_keywords` OR, case-insensitive), surfacing as a `triggered` block (0.95) in `recall_protocol` and the session-start inject; feature-private `disclosure_rules` table. Wiring: E16 `memory_pressure` emitter in the autohooks daemon (L1 ring >40, +10 hysteresis) and `context_threshold` in the Hermes plugin (`sync_turn` cumulative bytes, env `ARIEL_CONTEXT_THRESHOLD_BYTES`, re-arms per crossing) + Hermes `on_session_end` β†’ `session_ended` (no-LLM gist β†’ L3 `session_summary`); E14 `on_turn_end` event (same auto_save pipeline, sync result) as a surface for future harness emitters; E17a causal-link producer on `memory_graph_add` (`action="causal"` β†’ idempotent action/outcome nodes + strength edge); E17b `wiki_write` staging kind with a `propose` producer action on `memory_proposals` (apply β†’ `wiki:{path}`, revert deletes the page; WikiManager instance-scoped to the caller's data dir); E17c `revert_transition` β€” undo one consolidation promotion by its `memory_transitions` id (removes the promoted L4 fact, source episode stays, audit-logged). Validation: E18 DREAM markers anchored to message start (`^\s*DREAM:`) β€” mid-text mentions in documents no longer create junk `dream_skill` episodes (6 found in live dirs); E13 post-compaction semantic audit β€” payload-INDEPENDENT (real harnesses send no summary): mean cosine between the pre-compaction episode window and the L4 set rehydrate actually re-injects, logged to `audit_log` (`action='semantic_audit'`), returned in the handler result. Tool count 60 β†’ 63. E2E (9 full pipelines) + chaos/hypothesis suites (state-machine invariants, hostile ids/titles/payloads, crash-truncated L1 files, threading stress) β€” chaos found and fixed 4 real bugs (`_load` schema-drift escape, unbounded persist/waypoint filenames, duplicate facet params). - -### Changed -- **Procedural memory minimal core (Phase D D2.5)** β€” `memory_procedure` (60): HOW-knowledge with execution stats, author-trimmed from Β§13's 8 tables/12 tools/8 hooks down to ONE table + ONE multi-action tool (Stage-1 directive). `save` registers a named step list (branch-regex names, duplicate β†’ error), `use` records an execution outcome β€” counters update and an optional `learned` note APPENDS to notes (the trimmed Β§13.2 learn hook); `success_rate` is computed on read, never stored (no drift); `list` is payload-free, `get` returns full steps, `delete` removes. No execute-engine: procedures are a declarative cheat-sheet β€” the agent reads get() and acts. Discoverability rides the D1.3 steering route (updated: memory_procedure for repeatable procedures, wiki stays for rich skill docs). Lazy `procedural_memory` table (no migration). Write tier. Tool count 59 β†’ 60. - -### Changed -- **Memory stash (Phase D D1.12)** β€” `memory_stash` (59): git-stash for the working context β€” `save` captures the L1 reflex buffer (role/content/tokens; timestamps refreshed on pop β€” resumed chatter becomes "recent" again) + scratchpad entries under a name in a lazily-created `memory_stash` table and clears both; `pop` restores the exact working set and consumes the row (refuses while the current scratchpad is non-empty β€” stash first, no silent data loss); `list` (payload-free) and `drop` manage the set. L4 facts and sessions are deliberately NOT stashed β€” identity is D1.11 branches / D1.14 snapshots territory; after per-user layer keying the live use case is one agent switching between contexts (project A ↔ project B). Write tier. Tool count 58 β†’ 59. - -### Changed -- **Memory versioning: snapshots + rollback (Phase D D1.14)** β€” `memory_history` (57) gains four actions: `snapshot_create/snapshot_list/snapshot_restore` (named point-in-time captures of the (layer, user) fact set; restore makes L4 exactly equal the payload β€” upserts through save, deletes for keys absent from the snapshot, every step ledger-traced as `snapshot_restore:`, re-run idempotent) and `rollback` (git revert of ONE ledger mutation: reinstate its pre-state from the new full-row JSON β€” insert β†’ delete the fact, update/delete β†’ restore exact value/importance/kind/expires/source/metadata; legacy pre-d114 rows fall back to value+importance). The A2.2 ledger is amended in migration `20260901_1300_d114` with `old_row_json`/`new_row_json` (full before/after rows β€” typed schemas and provenance survive rollbacks; `list` stays slim, `get` returns full rows) plus the `core_memory_snapshots` table. Tool becomes dual-tier (insight read legs + write tier). Ceilings: entry_id/created_at not restored (blame chain ends at the revert β€” the ledger keeps history); restore not transactional (idempotent re-run); no auto-snapshots; ledger never pruned (scars forever). Tool count 58 β†’ 58 (no new tools). - -### Changed -- **Memory versioning ledger + branches (A2.2 + Phase D D1.11)** β€” `memory_history` (57): the `core_memory_history` mutation ledger (migration `20260901_1000_a22`) β€” every L4 insert/update/delete recorded with old/new values, importance, a deterministic `commit_hash` and `triggered_by` provenance (rides the D1.6 `source` contract; `list`/`get` read surface; write failures degrade to a warning so memory writes never fail on history). `memory_branch` (58): A/B persona staging for L4 facts β€” a branch is a `@` namespace in the existing `layer` column (zero schema change, invisible to RAG/inject/stats until merged): `create` full-clones the base layer's facts, `write` lands experiment facts branch-locally, `diff` reports added/changed/unchanged, `merge` cherry-picks keys (default: all differing) back into base with `source='branch_merge'` + `triggered_by='branch_merge:'`, `delete`/`list` manage the set. Ceilings (documented): no deletion propagation, no conflict detection, no checkout (inject always follows main β€” switch = merge). D1.14 (snapshot/rollback) reuses both. Tool count 56 β†’ 58. - -### Changed -- **Coherent tool tiers for ARIEL_EXPOSE** β€” three new opt-in tier groups alongside wiki/brief/review: `context` (recall protocol, recap, smart context, context inject, steering, compression β€” build/recover context), `insight` (query DSL, fact blame, quality, reflect, stats, raw search, episodes/sessions/graph reads β€” read-side analytics), `write` (remember, typed saves, rules, scratchpad, counterfactuals, episode/graph writes, session lifecycle β€” shape memory). Admin surfaces (backup/api_key/cleanup/watch/saga/data/sync/cleanup_purge, memory_forget) deliberately stay out of tiers β€” expose only via `ARIEL_EXPOSE=all`. Recommended live-agent value: `ARIEL_EXPOSE=primitives,context,insight,write,wiki,brief,review`. Default remains `primitives` (6 tools). - -### Fixed - `memory_load_rules`: a rules.yaml with a valid-but-wrong shape (e.g. a bare YAML list) crashed `load_rules` with `AttributeError` instead of degrading to an empty ruleset (caught by the new chaos suite). - `memory_query`: the `metadata` column was missing from the core_memory SELECT. - Test-infra: per-test deterministic importance gate (the adaptive-threshold EMA singleton drifted with execution order β€” the long-standing dir-isolated `test_mcp` order-dependent failures) and per-test hook-registry snapshot/restore (`AppContext()` used in tests registers real UserHooks/AgentHooks into the global registry and never unregisters). Full suite is now order-independent. - **Phase E audit round (2026-09-02)**: E13 semantic audit was dead in live β€” all three harnesses dispatch `post_context_compression` without a `query` payload, so the summary-vs-episodes design never fired; the audit now compares the pre-compaction window against the L4 set (payload-independent) and returns its result in the handler dict (was discarded). E17b `wiki_write` had no producer (nothing staged that kind) β€” `memory_proposals action="propose"` added, and the apply/revert path scoped `WikiManager` to the caller's data dir (was writing to `_DEFAULT_DIR`). E2 the live embedding breaker bypassed `breaker_registry` (direct construction) β€” the diagnostics breaker check was vacuously "all closed" and `reset_breakers` a no-op; the breaker is now created via the registry factory. `ReflexBuffer._load` escaped schema-drift rows (TypeError outside its except tuple) β€” entries now load field-tolerantly; persist-file names and wiki `safe_title`s are length-capped (OSError "File name too long" on hostile inputs); facet IN-params deduplicate. -### Added -- **Phase D E2E + chaos suite** β€” `tests/test_mcp/test_phase_d_e2e.py`: MCP-level flows with a real AppContext across the new tools (session lifecycle β†’ recap; remember β†’ query β†’ quality feedback β†’ fact blame; typed save β†’ query; rules β†’ hook gate β†’ tagged episode; scratchpad β†’ recap β†’ promote). `tests/test_features/test_phase_d_chaos.py`: 24h window boundaries, zero/negative budgets, limit clamping, unicode/injection-safe filters, scratchpad LRU cap, malformed rules.yaml/schemas degradation, 50k-line log bounding, concurrent remember+query races. `tests/test_mcp/test_server_tiers.py`: exact tier sets, live-agent combo, admin surfaces stay hidden. -- **Query DSL + typed schemas + rules engine (Phase D D1.7 + D1.8 + D1.9)** β€” `memory_query` (54): whitelisted filters β†’ parameterized SQL over core_memory/episodes (importance band, key/content LIKE, created_at window, episodes-only tag) β€” power-user analytics with no injection surface. `memory_save_typed` (55): validate structured fields against built-in schemas (`decision`/`error_pattern`/`relationship`) or custom `/schemas/*.yaml`, store as an L4 fact with `metadata.typed` (key `:`). `memory_load_rules` (56): declarative `/rules.yaml` rules (`when_content_contains` β†’ `importance_boost` sum cap 0.3 + episode tags) applied in the `auto_save_text` write gate β€” user-configurable memory behavior without code; mtime-cached, missing file = unchanged behavior. Tool count 53 β†’ 56. -- **Provenance tracking (Phase D D1.6)** β€” `memory_fact_blame` (53): who wrote this fact, when, and why β€” the evidence trail for debugging hallucinations. Rides the EXISTING `core_memory.source` column (no migration): canonical provenance values `user_explicit` (now the default for `MemoryManager.remember`), `staging_promotion`/`episode_promotion` (B1.4 consolidation products), legacy `manual`; `emotion_trigger`/`wiki_import`/`graph_inference` reserved until such core_memory write paths exist. Blame assembles importance_audit history (chunk_id = entry_id) + audit_log events (target_id = entry_id). -- **Tool-output compression + recall verification (Phase D D1.4 + D1.5)** β€” `memory_compress` (52): `log` mode keeps error/warn/fail lines (+ header, collapses consecutive duplicates, caps lines); `code` mode skeletonizes Python via ast (signatures kept, bodies dropped; non-Python falls back to log); `auto` tries code first. D1.5 verification rides the recall protocol: semantic hits with zero meaningful-token overlap against the query are dropped as retrieval noise (graph-expand hits exempt β€” their relevance is structural, not lexical). Ceiling documented: exact-zero filter only, overlap scoring = v2. Tool count 51 β†’ 52. -- **Session continuity + steering hints (Phase D D1.2 + D1.3)** β€” `memory_recap` (50): the /new recovery pack β€” last closed session (summary, topics, state deltas) β†’ pending work (scratchpad, diff gaps, staged proposals) β†’ markers β†’ day digest, budget-capped to ~2K tokens of recovery instead of re-reading raw history; also available as `autohooks recap`. `memory_steering` (51): deterministic route table (8 routes) + RU/EN keyword intent match steering agents to the best tool per intent (max 3 hints; empty query = full table) β€” advisory, the harness decides. Tool count 49 β†’ 51. -- **Counterfactual memory + skill evolution (Phase D D1.20 + D2.4)** β€” `memory_counterfactual` (49): "what could have been" notes (`anchor`/`premise`/`projection`) saved and listed by anchor. Skills now evolve from sessions: `wiki_read` audit-logs `skill_read` rows; the nightly hook's 6th phase (`skill_reinforce`) boosts read skills' importance +0.05 (cap 1.0, self-limiting via file mtime β€” only reads newer than the page's last write count); promotion MERGES into an existing skill page when the episode's title matches (full title or topic prefix before ":") instead of forking duplicates. Tool count 48 β†’ 49. -- **Agent scratchpad + memory quality loop (Phase D D1.15 + D1.19)** β€” `memory_scratchpad` (47): L2.5 working memory between session and episodes β€” the agent writes hypotheses/plans/drafts (cap 20, oldest evicted), entries re-inject at session start as a `scratchpad` block, and `promote` moves agent-judged-useful entries into L3/L4 (the agent is the distiller β€” D2.2 ceiling). `memory_quality` (48): was_useful β†’ score feedback loop β€” `feedback` writes a `recall_useful` audit row (feeds ACT-R frequency, D1.17) with importance Β±0.05 (cap 1.0 / floor 0.05, importance_audit-logged as `agent_feedback`); `report` aggregates per-entry useful counts with current importance. -- **Smart context budget + reflection system (Phase D D1.10 + D1.16)** β€” `memory_get_smart_context` (45): weighted token distribution across memory sources β€” important 30% / relevant 30% / recent 15% / day 15% / ops 10%, each source gets a floor first, leftover redistributes with a 2x-floor ceiling (a fat source can no longer starve the rest, unlike the sequential inject builder). `memory_reflect` (46) + `reflections` table (migration `20260830_1000_d116`): deterministic meta-memories β€” windowed episode counts + recurring topics (no LLM), `action="write"` stores a reflection row, `action="list"` reads them back; the nightly hook's 5th phase writes the daily reflection automatically. -- **Store pipeline + shared skill SSOT (Phase D D2.2 + D2.3)** β€” `memory_skill_promote` tool (44) and `features/skill_pipeline.py`: promote distilled episodes into `skill` wiki pages verbatim with provenance footers, idempotent via a new `skill_promoted` episode tag (`EpisodicMemory.get_by_id`/`add_tag` added); the nightly hook gains a 4th phase (`skill_promotion`) auto-promoting fresh `dream_skill` episodes. Ceiling (documented): no LLM in ariel β€” dream_skill episodes are agent-distilled at write time (C1.12); raw auto_save chatter needs harness-side distillation (v2). D2.3: `scripts/sync_skills.py` β€” a git-versioned skill SSOT (`~/skills-ssot/skill/*.md`) synced into every live agent's wiki via `WikiManager.sync_external` (copy + sha256 dedup, per-agent subprocesses so `connection_manager` picks up each `MCP_MEMORY_DATA_DIR`; aiosqlite `close_all()` before exit per the C1.10 hang lesson). `scripts/sync_skills.py --bootstrap` creates the skeleton; the sync is cron-ready. -- **Skill = Memory (Phase D D2.1)** β€” skills are wiki pages of the new first-class `skill` type: plain Markdown (SKILL.md convention), agent-read, no embeddings for retrieval β€” progressive disclosure through the wiki surface (`wiki_list(wiki_type="skill")` β†’ `wiki_search` β†’ `wiki_read`). New `wiki_read` tool closes the missing read leg (full page content by path; `wiki_list`/`wiki_search` now expose `path`/`snippet` for chaining). Wiki lint gains the `skill_too_large` rule (concept cap: 4KB per skill page). Tool count 41 β†’ 43. The C1.12 `dream_skill` L3 episodes are the mineable promotion trace (store pipeline = D2.2). -- **/recall protocol β€” multi-axis proportional recall (Phase D D1.1)** β€” `features/recall.py::recall_protocol` fuses five axes in concept order: `markers` (dream-marker facts at 0.95 + `dream_skill` episodes β€” outrank everything) β†’ `session` (last-24h L1 chatter + latest session summary) β†’ `semantic` (MultiSourceRAG hybrid hits) β†’ `expand` (B1.6 1-hop graph neighbors) β†’ `day` (last-24h auto-save digest). Proportional: empty query = zero-state (markers + day only, ~3 lines); dedupe across axes (markers' parts pre-registered, first axis wins); budget-capped. New `autohooks recall --query … [--budget N] [--format md|json]` subcommand and the `memory_recall_protocol` MCP tool (operator tier, tool count 41 β†’ 42). The Hermes native plugin's `queue_prefetch`/`prefetch` now drive this protocol per turn β€” every Hermes turn gets multi-axis recall of the memory store. -- **Compaction-aware memory rehydrate (Phase D D3.5)** β€” ariel now learns when an agent's context was compacted: every `post_context_compression` dispatch logs a drift row to a new `compaction_events` table (migration `20260829_1900_d35_rehydrate`), and the session-start inject gains a `rehydrate` block (important L4 facts, score 0.9) when a compaction happened within `rehydrate.window_hours` (default 6h; `rehydrate.enabled` toggles). `autohooks dispatch` accepts `--payload` (JSON object merged into the event context) and `autohooks inject` accepts `--blocks` (kind filter) so harnesses can pull exactly the rehydrate block. MiMoCode plugin: the compaction hook appends the critical set to the summarizer prompt (salvage), the `session.compacted` event logs the drift and arms a one-shot `[ariel rehydrate]` system block; Hermes gateway emits a `compaction` boundary event (new `~/.hermes/hooks/ariel-compaction/` hook) and the rotated session's inject carries the block. -- **Phase C closeout (C1.12 + C1.13 + C1.14)** β€” `DREAM: memory:/fact:/skill:` markers detected in the auto-save pipeline and routed through staging at importance 0.95 (`skill:` also writes an L3 `dream_skill` episode for D2 to mine). `memory_proposals(action="revert")` undoes applied proposals with exact provenance: core_write (forget by key) and archive (`restore_entries` from `archived_memories` back into core β€” the backup already exists by design); consolidation revert deferred (ids not pinned). `memory_report_card(period_hours=24)` β€” operator digest: proposal counts by status + last decisions, auto-save tier sums, open gaps, dream-marker count. Both tools hidden by default; new `ARIEL_EXPOSE=primitives,review` tier opts an instance in. Tool count 40 β†’ 41. **Phase C auto-hooks keystone is COMPLETE.** -- **Staged mutation β€” proposal β†’ review β†’ apply (Phase C C1.11)** β€” risk-tier true staging: L4-destined auto-saves (score β‰₯ 0.8), consolidation staging promotions (hook + lifespan sweep), and the forgetting ritual's archive sweep now create `mutation_proposals` instead of writing immediately (L3/graph/decays stay direct; propose failure falls back to direct write so bookkeeping never loses memory). Pending proposals surface in the session-start inject as a `proposals` block; the decision is one `memory_proposals` tool call (list / decide approve|reject), same surface for agent and operator. 7-day lazy expiry (no cron). Apply executes the exact write the direct path would have (pinned ids/items); every decision audit-logged. `staging.enabled=false` restores direct behavior. Tool count 39 β†’ 40. Migration `20260829_1600_c111`. -- **`memory_watch` operator tool + `post_session_diff` auto-handler (Phase C C1.10)** β€” one server-side event fires after each session end, computes "dispatched vs persisted" gaps from a new `memory_dispatch_log` table, materializes them as L3 `diff_gap` episodes, and the next `session_started` inject surfaces them as a new `gap` block kind. `memory_watch(action=list|add|disable|delete)` is the operator CRUD over `watch_rules` (introspection of the rules ariel already applies via `auto_save_text`; no new behavior). Migration `20260829_1300_d1c110` adds both tables and seeds one default rule. Tool count 38 β†’ 39. Gaps now reach the agent without any explicit call. -- **Universal autohooks runtime (Phase C C1.9)** β€” `python -m autohooks` (package `autohooks/`): per-agent daemon tails each agent's SQLite conversation store and pushes `new_message` events into the hook dispatcher in-process (no HTTP β€” works for stdio agents); `inject` subcommand returns the session-start critical set for harness embedding. Declarative per-agent YAML configs in `autohooks/examples/` (hermes, mimocode, cowagent β€” all live-verified). First-run baseline prevents history replay; cursor persists per batch (at-least-once). Isolation inherited via `MCP_MEMORY_DATA_DIR`; one runtime, one config = new agent. -- **Auto-hooks foundation (Phase C)** β€” external event dispatcher (`hooks/external.py::dispatch_event`) with two transports: `POST /api/hooks/{event}` (HTTP harnesses) and the new `memory_hook` MCP tool (stdio agents; primitive tier, tool count 37β†’38). Seven events with day-one handlers: `session_started` (returns the critical inject block), `session_ended` (session summary β†’ L3), `new_message`/`auto_save_candidate` (`evaluate_importance` heuristic β€” no LLM; score β‰₯ 0.5 β†’ L3 episodic + graph node, β‰₯ 0.8 β†’ also L4 core), `post_context_compression` (rehydrate candidates via retrieval), `context_threshold`/`memory_pressure` (thin advice; decisions stay harness-side). New `POST /api/context-inject` returns the budget-capped critical set (ACT-R top-5 relevant when a query is given + recent L1 (24h) + important facts β‰₯ `inject.important_min`). Graph is now threaded through the hook registry (`fire(..., graph=)`, C1.1). Per-agent isolation inherited (own process + `MCP_MEMORY_DATA_DIR`; D1.13 API-key binding applies to both surfaces). Harness-side daemon (C1.9) is the next step. - -### Fixed - **MiMoCode plugin exported non-existent hook keys** β€” `~/.config/mimocode/hooks/ariel-inject.ts` used `session.start`/`session.end`, which are not keys of the fork's `Hooks` interface (verified: zero occurrences in the repo), so the C1.9 session-start inject and C1.10 `post_session_diff` never fired for MiMoCode. Rewired to the real keys: per-session one-shot inject via `experimental.chat.system.transform`, debounced diff via `session.post` (plus the D3.5 compaction hooks). +### Security +- **Incidental operator identifiers are gone from the published tree (2026-10-08).** A second pass over the same class the fixture scrub started: names that had leaked into prose rather than into data. Removed from `CHANGELOG.md` (two quoted utterances and the `agent_id` list), `AUDIT_REPORT_20260629.md` (the generator's name), `ARIEL_RULES.md`, `autohooks/examples/dsh.yaml`, one Alembic revision comment, `scripts/purge_foreign_rows.py`, and eleven test files where a name stood in for a role. `pyproject.toml`'s "handled by Lucy" comment became "handled by hand". The names used *as the anonymizer's dictionary* stay β€” `config.yaml` `rag.ru_personas`, `rag/synonyms.py`, `mcp_server/utils/privacy.py`, `tests/test_hooks/test_privacy_ru.py` and the `graph_miners` canon test: those are the privacy product, and deleting them would remove the feature that hides names. Six `docs/compose/` files that were force-added past `.gitignore` are untracked again; they remain on disk, and three tests that cite them in docstrings were never reading them. + ## [1.9.0] - 2026-08-29 ### Added