From 2ef783ec0ab704aa5d9f43598a64a2f44a570486 Mon Sep 17 00:00:00 2001 From: Kyle Tse Date: Wed, 30 Sep 2026 00:05:12 +0000 Subject: [PATCH] docs: revamp AGENTS.md, README and docs to owner standards (RUL-78) --- AGENTS.md | 42 ++++++++++++++++++++++++++---------------- PROJECT.md | 2 +- docs/benchmarks.md | 2 +- docs/capabilities.md | 1 - 4 files changed, 28 insertions(+), 19 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index da308a3..a50a6fd 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -1,19 +1,29 @@ # lockdocs: agent notes -Read [PROJECT.md](PROJECT.md) for the layout and the release flow. +lockdocs answers an agent's library questions from the exact versions in the +project's lockfile, locally and offline, over a three-tool MCP surface. Work +here is judged by the version-sensitive benchmark: the answer must be right for +the pinned version and cheap in tokens and time. Layout and release flow: +[PROJECT.md](PROJECT.md); destination: [docs/vision.md](docs/vision.md). -- Run the narrowest check first: `cargo test -p lockdocs-core`, then - `cargo test --workspace`. The benchmark (`bench/`) installs real packages; - run it in CI (`bench.yml`), not on a shared machine. -- Keep the MCP surface at three tools: `resolve`, `docs`, `api`. -- Bump `index::FORMAT` whenever extraction output or `Entry` changes, so old - caches are rebuilt; bump `upstream::FORMAT` when `fetch` downloads more, so - `lockdocs fetch` refreshes old copies. -- Ranking changes are judged on the whole benchmark in CI, never on one - question; questions marked `held-out` are not used for tuning. -- A new lockfile format: add a pure parser to `lockfile.rs` with a unit test, - and register it in `LOCKFILES` or `parse`. -- One version everywhere: `bun scripts/set-version.ts` (CI runs - `scripts/check-version.ts`). -- Format with `rustfmt --edition 2021 ` (config in `rustfmt.toml`). -- Never commit secrets, tokens or `.env` files. +## Hard lines + +- The MCP surface stays at three tools (`resolve`, `docs`, `api`): agents pay + for every tool description on every call, so new abilities go into these. +- Bump `index::FORMAT` when extraction output or `Entry` changes, and + `upstream::FORMAT` when `fetch` downloads more: otherwise old caches serve + stale answers. +- Questions marked `held-out` in `bench/questions.json` are not used for tuning: + they are the check that ranking changes generalize. +- No secrets, tokens or `.env` files in the repository. + +## How a result is judged + +- `cargo test -p lockdocs-core`, then `cargo test --workspace`; format with + `rustfmt --edition 2021` (see `rustfmt.toml`). +- Ranking changes are judged on the whole benchmark (`bench.yml`, which installs + real packages, so run it in CI), never on one question. +- CI also runs `scripts/check-version.ts`, `scripts/check-capabilities.ts` and + `scripts/check-tagline.ts`. One version everywhere: `bun scripts/set-version.ts`. +- A new lockfile format is a pure parser in `lockfile.rs` with a unit test, + registered in `LOCKFILES` or `parse`. diff --git a/PROJECT.md b/PROJECT.md index 8c10456..7d60d73 100644 --- a/PROJECT.md +++ b/PROJECT.md @@ -4,7 +4,7 @@ Exact-version library docs for AI agents, from the project's lockfile, offline. It ships as a Rust MCP server and CLI, runs locally, needs no API key, and is MIT licensed. -- Lifecycle: `active`, published as `@sylphx/lockdocs` (npm) and +- Published as `@sylphx/lockdocs` (npm) and `io.github.SylphxAI/lockdocs` (MCP Registry) - Docs: https://sylphxai.github.io/lockdocs/ diff --git a/docs/benchmarks.md b/docs/benchmarks.md index 3e8bebf..b20b151 100644 --- a/docs/benchmarks.md +++ b/docs/benchmarks.md @@ -3,7 +3,7 @@ ## Method - **Questions:** 105 questions whose correct answer depends on the version, in [`bench/questions.json`](https://github.com/SylphxAI/lockdocs/blob/main/bench/questions.json), over zod 3/4, Next.js 14/15/16, React Router 6/7, pydantic 1/2, axum 0.7/0.8, tokio, Tailwind CSS 3/4, ESLint 8/9, Prisma 5/6, React 18/19, Vite 5/6, Express 4/5, SQLAlchemy 1.4/2.0, Django 4.2/5.1 and FastAPI 0.88/0.115. Twin questions are worded identically for both versions; each records its source of truth. -- **Held-out questions:** 18 questions marked `held-out` are not used for tuning: they were written before the ranking changes they measure and have their own column. Held-out questions that are later used to diagnose a miss join the main set (their `history` field says so), and new ones replace them. The 0.3 round did this once: 17 questions first written as held-out (lockdocs 11/17, Context7 12/17 on their first run) are now in the main set. +- **Held-out questions:** 18 questions marked `held-out` are not used for tuning: they were written before the ranking changes they measure and have their own column. Held-out questions that are later used to diagnose a miss join the main set (their `history` field says so), and new ones replace them. - **Projects:** one project per version in [`bench/projects`](https://github.com/SylphxAI/lockdocs/tree/main/bench/projects), installed at those pins by [`bench/setup.sh`](https://github.com/SylphxAI/lockdocs/blob/main/bench/setup.sh). - **Grading:** an answer passes when it contains at least one string from every `expect` group (the version-correct API, in code or prose form) and none of the `reject` strings (the other version's API). Case-insensitive; the same grader for every tool. - **lockdocs**, default settings (1,200-token budget), in three configurations: keyword only (`LOCKDOCS_EMBED=0`) on package files; hybrid on package files (the offline default once the model is downloaded); hybrid after `lockdocs fetch` added upstream docs at each version's tag. Index and fetch times are reported separately. diff --git a/docs/capabilities.md b/docs/capabilities.md index d671d51..a235a49 100644 --- a/docs/capabilities.md +++ b/docs/capabilities.md @@ -18,4 +18,3 @@ is in [vision.md](vision.md). | LD-CLI | Command line with the same queries, plus `index`, `fetch` and `setup` | supported | crates/lockdocs/src/main.rs, crates/lockdocs/src/setup.rs | LD-SEARCH | | LD-NPM | npm launcher and native binaries for five platforms | supported | packages/lockdocs, packages/npm | LD-CLI | | LD-BENCH | Version-sensitive benchmark against Context7, run on GitHub-hosted runners | supported | bench/run.py, bench/questions.json, .github/workflows/bench.yml | LD-CLI | -| LD-SEMANTIC-RERANK | Rerank the top results with a stronger local model for questions worded differently from the docs | planned | | LD-SEARCH |