A token-efficient persistent knowledge base for teams using LLMs on large internal corpora — architecture docs, ADRs, RFCs, runbooks, playbooks, and code.
The core problem: internal knowledge is scattered across Notion, Confluence, GitHub, and Slack. The naive fix is dumping files into a prompt. That breaks immediately — a single RFC uses 3,000–8,000 tokens, the context window fills after a few documents, and you're back to copy-pasting manually.
This system solves it differently. Content is compressed at write time into small, typed summaries. At query time, only the relevant pieces load — progressively, starting from the smallest. A direct lookup costs ~120 tokens. A cross-document synthesis costs ~600–2,000. The full corpus never loads.
This is not vanilla RAG (chunk everything → embed → retrieve top-k → stuff the prompt). A vector index, if you use one, is only the recall substrate — one step, subordinate to a memory-management layer that is the actual point of the system:
| Layer | Role | RAG equivalent |
|---|---|---|
| Semantic/keyword index (e.g. qmd) | Generate candidate chunks | The whole of RAG |
_memory-scores.json decay |
Re-rank by importance × recency × frequency (a forgetting curve), not raw similarity | — |
| Rules → abstracts → summaries → source | Progressive levels-of-detail loading | — |
| Query cache | Skip retrieval entirely for near-duplicate questions | — |
consolidate / prune-memories |
Merge, link, and forget over time | — |
So it's closer to a cache-and-memory hierarchy (or a cognitive-architecture memory model) than to retrieve-then-stuff RAG. Vanilla RAG is just the first row; everything below it is what makes token cost scale with task complexity instead of corpus size. The retrieval backend is pluggable — grep works under ~200 chunks; a vector DB takes over above that (see Retrieval backend below).
Six documents covering the full system:
| Doc | What it covers |
|---|---|
01-architecture.md |
System overview, components, file layout, session lifecycle |
02-chunking.md |
How to chunk different content types (RFCs, ADRs, runbooks, code) and the three-layer summary model |
03-context-engineering.md |
Static vs fluid context tracks, the _headline.md / _state.md / _scratchpad.md pattern |
04-retrieval.md |
How retrieval works: rules → abstracts → summaries → full text |
05-implementation.md |
Step-by-step setup guide with file templates |
06-maintenance.md |
Keeping the knowledge base healthy over time |
07-hosting.md |
Sharing across a team |
consolidate.md — The "dreaming cycle": score chunks from LOG.md, merge overlapping summaries, find new cross-chunk connections. Run when the repo feels fragmented or after a burst of ingest activity.
prune-memories.md — Score all chunks using exponential decay, categorise by health (healthy / fading / decayed), and propose archival. Never auto-prunes — always waits for approval.
session-wrapup/SKILL.md — End-of-session wrapup: self-critique, grounding check, state update, scratchpad fold, plan check, model grade, strategy distillation. Two modes: lightweight (7 steps) and full (11 steps).
save-plan/SKILL.md — Capture a completed task as a reusable plan template in _plans/. Extracts load order, steps, trigger keywords, and guardrails. Run after any repeatable task.
Every chunk of knowledge exists at three sizes:
| Tier | Size | When it loads |
|---|---|---|
| Rule | 20–50 tok | Always — tells Claude when to load the chunk |
| Abstract | ~100 tok | On a query match — bullet facts, enough for a direct lookup |
| Summary | ~300 tok | When the abstract isn't enough — full narrative |
| Source | Full file | Only on explicit request |
Rules for all chunks stay in context permanently (~20 rules per 1k tokens). The rest load progressively based on query relevance. This means token cost scales with task complexity, not corpus size.
Two context tracks with different freshness semantics:
- Static track — loaded once per session. Repo index (
_tree.md), glossary, active project summaries. Treat as immutable mid-session. - Fluid track — injected each turn via a
UserPromptSubmithook or manual read. Current task headline, state, scratchpad top layer.
The patterns work with any agent harness. Skills use the Agent Skills standard format and are compatible with Claude Code, pi, and any harness that supports SKILL.md discovery.
The retrieval layer (docs/04-retrieval.md) is deliberately pluggable — the architecture only requires something that turns a query into ranked candidate chunk IDs. Two tiers:
- Grep / INDEX.md — the zero-dependency default. Fine under ~200 chunks.
- Vector/hybrid index — once the corpus outgrows grep. The reference implementation is qmd ("Quick Markdown Search"): a local hybrid engine over your markdown, with BM25 (
qmd search), vector similarity (qmd vsearch), and an auto-expanding, reranked hybrid mode (qmd query) supportinglex:/vec:/hyde:typed lines. It exposes collections,qmd get/multi-getfor single-doc fetch, and an MCP server (qmd mcp) so agents can call it as a tool.
Crucially, qmd only does the recall step. Its ranked output is then re-ranked by _memory-scores.json and fed into progressive tier-loading — the memory hierarchy described above. Swap qmd for any other index (pgvector, Chroma, Turbopuffer, plain BM25) without touching the rest of the system.