Skip to content

Latest commit

 

History

History
180 lines (155 loc) · 10.2 KB

File metadata and controls

180 lines (155 loc) · 10.2 KB

Handover

Everything a new contributor needs to pick Lorekeeper up. Read this first, then README.md for the reader-facing description.

What this is, and why it exists

Lorekeeper is a campaign codex that a person and their AI agent keep together in one browser tab. You tell the agent what happened at the table, it records entries; weeks later you ask a question in plain language and it recalls the answer by meaning rather than by keyword. Everything, including the database and the embedding model, runs in the page.

The problem it solves is agent amnesia with a twist that matters: the memory is visible and editable by the human who owns it. Most agent memory is either a server-side black box or a blob of JSON in local storage. Here the codex is a real database with real semantics, the reader can read and correct any entry, and the agent and the reader arbitrate their edits through a conditional write.

It exists as an entry to the OpenAI WebMCP Challenge (submission deadline September 3, 1pm PT, judged on usefulness, originality, execution, thoughtful use of WebMCP, and the quality of the human and agent experience). The other reason it exists: it is the first public demonstration of the ExtendDB WebAssembly target, so what the app shows about the engine matters as much as what it shows about the app.

Architecture in one pass

  • web/vendor/extenddb/ is ExtendDB, a DynamoDB-compatible engine, compiled to WebAssembly over SQLite. It exposes one entry point, dispatch(target, body), which takes DynamoDB JSON exactly as the managed service would, plus export_db() and import_db(bytes) for whole-database snapshots. See docs/BUILDING-THE-ENGINE.md to rebuild it, and patches/ for the persistence patch that build carries.
  • web/vendor/transformers/ plus web/vendor/models/ is a quantized MiniLM model, 384 dimensions, loaded with remote models disabled. Embeddings never leave the page. web/embed.mjs wraps it.
  • web/tools.mjs is the whole data layer: five tools (record_lore, recall_lore, list_lore, revise_lore, forget_lore) over PutItem, vector search, Query, UpdateItem with a condition, and DeleteItem. It also exports TOOL_DEFINITIONS, the single description of the tool surface. Descriptors registered through the shim carry annotations (read-only on the read tools, destructive on forget_lore), and the write tools' descriptions carry one sentence of untrusted-content guidance for the model. record_lore accepts an optional epoch-milliseconds ts that only the seeder uses; it is kept out of the public schema, and it is how the sample campaign gets entry dates spread over weeks instead of one second.
  • Three things drive those same five tools:
    1. WebMCP, for agent-native browsers. web/webmcp-shim.mjs normalizes the API namespace and registration shape, and passes annotations through.
    2. The in-page scribe, web/agent.mjs, which runs a tool-calling loop against any OpenAI-compatible endpoint, one plain JSON completion request per model round (there is no streaming path). Its schemas are derived from TOOL_DEFINITIONS, so the two machine consumers cannot drift apart, and when the browser provides a WebMCP surface its calls route through the registered tools rather than a private path.
    3. The reader, through the composer, inline card editing, the search box, and the one-click suggested queries beside it.
  • web/app.mjs wires all of it to the page and logs every engine call to the Wire view. Cards show relevance as a percentage (raw distances stay in the wire view). Codex refreshes coalesce to one per completed scribe turn, while human actions refresh immediately. On a first visit it seeds the sample campaign automatically and shows a notice that clears it; export and import in the codex header go through the engine's export_db / import_db.
  • web/seed.mjs holds the starter campaign, each entry with a daysAgo offset for the backdated timestamps, and the SUGGESTED_QUERIES the chips render.
  • Persistence is an IndexedDB-backed store with relaxed durability, plus a snapshot flush when the page is hidden and after an import, so a codex survives reload and a fast tab close.

Running it

npm install
npm run serve            # http://127.0.0.1:4173
npx playwright test      # the whole suite should pass

The suite (tests/lorekeeper.spec.mjs plus tests/webmcp-surface.spec.mjs) drives the real engine, the real model, and the real WebMCP surface in Chromium with --enable-features=WebMCP. It intercepts the model endpoint, so no test ever makes a network call. The test server binds port 4173 by default; set LOREKEEPER_TEST_PORT when something already holds that port.

To exercise the in-page scribe you need an OpenAI-compatible endpoint. Any will do. If you have AWS credentials but no OpenAI key, there is a small local proxy outside this repo that presents Bedrock as an OpenAI chat-completions endpoint; ask Lee for bedrock_openai_proxy.py. Point the scribe's setup at it and use any non-empty key.

Ground rules for changes here

  • Never an em dash or en dash in any file, including code comments and commit messages. Use a comma, a colon, parentheses, or two sentences.
  • The privacy claim is scoped deliberately: "your codex never leaves the page". It is not "nothing leaves the page", because prompts, replies, and the tool calls and results inside an agent conversation go to the configured model endpoint. Keep that distinction wherever you touch the copy.
  • Tests use real clicks. A synthetic MouseEvent skips the checks that catch a control a reader cannot actually reach, which is a bug this app already had once. Only use one to test a re-entry guard, and say why in a comment.
  • Two files carry a frozen contract that other files code against: the DOM ids in web/index.html and the module surface of web/agent.mjs. Renaming either breaks the wiring silently. Grep before you rename.

Known issues

Closed in the latest pass, so you do not chase them: the per-mutation refresh storm is gone (refreshes now coalesce to one per completed scribe turn), the streaming path in web/agent.mjs is deleted rather than dormant (the module sends one plain JSON completion per model round), and the page no longer lands empty (a first visit auto-seeds the sample campaign with backdated timestamps, plus the suggested queries and export/import described above), multi-campaign shipped (a campaign picker in the header, one partition per campaign, see the data model below), revise_lore can retag an entry in the same write (absent keeps the tags, an empty array clears them, an array replaces them; the inline editor's threads field drives it), and a Web Lock now guards the single-writer engine against a second tab.

Still open:

  1. Multi-tab is guarded, not solved. The storage VFS is single-writer, and two tabs writing the same IndexedDB store could corrupt the codex. A Web Lock (lorekeeper-engine) now guards it: the first tab owns the engine for the page's lifetime, and any later tab skips the persistent store, shows #multitab-notice, and offers a Take over button that waits for the lock and reloads. A browser without navigator.locks boots as before the guard. A worker-owned engine that every tab shares is still the real fix.

The data model

The table is codex, keyed by pk (HASH) and sk (RANGE). One campaign is one partition: every entry in a campaign's codex shares the partition value codex#<campaign>, and the default campaign uses codex#default, the same value every pre-campaign codex already used, so old codices open unchanged. The campaign picker in the header switches the active campaign; the list of campaigns and the open one live in a single item of the app_meta table, so a reload reopens where the reader left off. An entry's id IS its sort key: a 13-digit zero-padded creation timestamp joined to a unique suffix. Padding makes lexicographic order match chronological order, so list_lore issues a Query with ScanIndexForward: false and gets the codex newest first with the limit applied to the newest entries, all server side. Nothing sorts client side, and nothing scans.

Two consequences worth knowing. The sort key holds the creation timestamp and never changes, while the ts attribute is the version the conditional write in revise_lore compares, so revising an entry leaves it in place in reading order rather than jumping it to the top. And the table was previously named memories with a single id key, so a browser that ran an older build still holds that table in IndexedDB; it is orphaned rather than migrated, which is deliberate at this stage.

Semantic recall is campaign-scoped too. The vector index declares a SearchSchema with pk as its HASH element, so every SearchVectors request carries SearchConditionExpression: "#pk = :pk" and the engine confines the search to the active campaign's partition. A codex persisted by a build before multi-campaign has a vector index with no SearchSchema; ensureTable detects that once via DescribeTable, and recall_lore then over-fetches and applies the partition scope client side instead (capped by the service's TopK limit of 100, which is acceptable for a legacy codex).

One timing consequence: each tool captures the active campaign's partition when the call starts, before the in-page embed runs, so a campaign switch made while an embed is in flight cannot re-target that write or read.

Where the submission stands

Complete: the app, the tool surface with annotations, persistence, export and import, the auto-seeded sample campaign with suggested queries, tests, the engine reproducibility docs, and the Devpost assets in docs/ (devpost-story.md, demo-script.md with a 2:48 shot list, SUBMISSION.md with the live checklist). Hosting exists: https://lorekeeper.extenddb.org, a CloudFront distribution behind the extenddb.org subdomain.

Outstanding, and all of it needs Lee: confirming the distribution serves, the Chrome origin-trial token for that origin, verification in the ChatGPT in-app browser, the demo video, flipping this repository public with an incognito check, and the Devpost submit itself. Note the freeze rule: after submission, no changes to the repository or the site until winners are announced.