Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

knowledge-base-ops

A token-efficient persistent knowledge base for teams using LLMs on large internal corpora — architecture docs, ADRs, RFCs, runbooks, playbooks, and code.

The core problem: internal knowledge is scattered across Notion, Confluence, GitHub, and Slack. The naive fix is dumping files into a prompt. That breaks immediately — a single RFC uses 3,000–8,000 tokens, the context window fills after a few documents, and you're back to copy-pasting manually.

This system solves it differently. Content is compressed at write time into small, typed summaries. At query time, only the relevant pieces load — progressively, starting from the smallest. A direct lookup costs ~120 tokens. A cross-document synthesis costs ~600–2,000. The full corpus never loads.


How this relates to RAG

This is not vanilla RAG (chunk everything → embed → retrieve top-k → stuff the prompt). A vector index, if you use one, is only the recall substrate — one step, subordinate to a memory-management layer that is the actual point of the system:

Layer Role RAG equivalent
Semantic/keyword index (e.g. qmd) Generate candidate chunks The whole of RAG
_memory-scores.json decay Re-rank by importance × recency × frequency (a forgetting curve), not raw similarity —
Rules → abstracts → summaries → source Progressive levels-of-detail loading —
Query cache Skip retrieval entirely for near-duplicate questions —
consolidate / prune-memories Merge, link, and forget over time —

So it's closer to a cache-and-memory hierarchy (or a cognitive-architecture memory model) than to retrieve-then-stuff RAG. Vanilla RAG is just the first row; everything below it is what makes token cost scale with task complexity instead of corpus size. The retrieval backend is pluggable — grep works under ~200 chunks; a vector DB takes over above that (see Retrieval backend below).


What's here

docs/

Six documents covering the full system:

Doc What it covers
01-architecture.md System overview, components, file layout, session lifecycle
02-chunking.md How to chunk different content types (RFCs, ADRs, runbooks, code) and the three-layer summary model
03-context-engineering.md Static vs fluid context tracks, the _headline.md / _state.md / _scratchpad.md pattern
04-retrieval.md How retrieval works: rules → abstracts → summaries → full text
05-implementation.md Step-by-step setup guide with file templates
06-maintenance.md Keeping the knowledge base healthy over time
07-hosting.md Sharing across a team

commands/

consolidate.md — The "dreaming cycle": score chunks from LOG.md, merge overlapping summaries, find new cross-chunk connections. Run when the repo feels fragmented or after a burst of ingest activity.

prune-memories.md — Score all chunks using exponential decay, categorise by health (healthy / fading / decayed), and propose archival. Never auto-prunes — always waits for approval.

skills/

session-wrapup/SKILL.md — End-of-session wrapup: self-critique, grounding check, state update, scratchpad fold, plan check, model grade, strategy distillation. Two modes: lightweight (7 steps) and full (11 steps).

save-plan/SKILL.md — Capture a completed task as a reusable plan template in _plans/. Extracts load order, steps, trigger keywords, and guardrails. Run after any repeatable task.


The core idea

Every chunk of knowledge exists at three sizes:

Tier Size When it loads
Rule 20–50 tok Always — tells Claude when to load the chunk
Abstract ~100 tok On a query match — bullet facts, enough for a direct lookup
Summary ~300 tok When the abstract isn't enough — full narrative
Source Full file Only on explicit request

Rules for all chunks stay in context permanently (~20 rules per 1k tokens). The rest load progressively based on query relevance. This means token cost scales with task complexity, not corpus size.

Two context tracks with different freshness semantics:

  • Static track — loaded once per session. Repo index (_tree.md), glossary, active project summaries. Treat as immutable mid-session.
  • Fluid track — injected each turn via a UserPromptSubmit hook or manual read. Current task headline, state, scratchpad top layer.

Compatibility

The patterns work with any agent harness. Skills use the Agent Skills standard format and are compatible with Claude Code, pi, and any harness that supports SKILL.md discovery.


Retrieval backend

The retrieval layer (docs/04-retrieval.md) is deliberately pluggable — the architecture only requires something that turns a query into ranked candidate chunk IDs. Two tiers:

  • Grep / INDEX.md — the zero-dependency default. Fine under ~200 chunks.
  • Vector/hybrid index — once the corpus outgrows grep. The reference implementation is qmd ("Quick Markdown Search"): a local hybrid engine over your markdown, with BM25 (qmd search), vector similarity (qmd vsearch), and an auto-expanding, reranked hybrid mode (qmd query) supporting lex: / vec: / hyde: typed lines. It exposes collections, qmd get/multi-get for single-doc fetch, and an MCP server (qmd mcp) so agents can call it as a tool.

Crucially, qmd only does the recall step. Its ranked output is then re-ranked by _memory-scores.json and fed into progressive tier-loading — the memory hierarchy described above. Swap qmd for any other index (pgvector, Chroma, Turbopuffer, plain BM25) without touching the rest of the system.

About

A token-efficient persistent knowledge base for teams running LLMs over large internal corpora (RFCs, ADRs, runbooks). Compress at write time, load progressively at query time.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors