Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
42 changes: 26 additions & 16 deletions AGENTS.md
Original file line number Diff line number Diff line change
@@ -1,19 +1,29 @@
# lockdocs: agent notes

Read [PROJECT.md](PROJECT.md) for the layout and the release flow.
lockdocs answers an agent's library questions from the exact versions in the
project's lockfile, locally and offline, over a three-tool MCP surface. Work
here is judged by the version-sensitive benchmark: the answer must be right for
the pinned version and cheap in tokens and time. Layout and release flow:
[PROJECT.md](PROJECT.md); destination: [docs/vision.md](docs/vision.md).

- Run the narrowest check first: `cargo test -p lockdocs-core`, then
`cargo test --workspace`. The benchmark (`bench/`) installs real packages;
run it in CI (`bench.yml`), not on a shared machine.
- Keep the MCP surface at three tools: `resolve`, `docs`, `api`.
- Bump `index::FORMAT` whenever extraction output or `Entry` changes, so old
caches are rebuilt; bump `upstream::FORMAT` when `fetch` downloads more, so
`lockdocs fetch` refreshes old copies.
- Ranking changes are judged on the whole benchmark in CI, never on one
question; questions marked `held-out` are not used for tuning.
- A new lockfile format: add a pure parser to `lockfile.rs` with a unit test,
and register it in `LOCKFILES` or `parse`.
- One version everywhere: `bun scripts/set-version.ts` (CI runs
`scripts/check-version.ts`).
- Format with `rustfmt --edition 2021 <files>` (config in `rustfmt.toml`).
- Never commit secrets, tokens or `.env` files.
## Hard lines

- The MCP surface stays at three tools (`resolve`, `docs`, `api`): agents pay
for every tool description on every call, so new abilities go into these.
- Bump `index::FORMAT` when extraction output or `Entry` changes, and
`upstream::FORMAT` when `fetch` downloads more: otherwise old caches serve
stale answers.
- Questions marked `held-out` in `bench/questions.json` are not used for tuning:
they are the check that ranking changes generalize.
- No secrets, tokens or `.env` files in the repository.

## How a result is judged

- `cargo test -p lockdocs-core`, then `cargo test --workspace`; format with
`rustfmt --edition 2021` (see `rustfmt.toml`).
- Ranking changes are judged on the whole benchmark (`bench.yml`, which installs
real packages, so run it in CI), never on one question.
- CI also runs `scripts/check-version.ts`, `scripts/check-capabilities.ts` and
`scripts/check-tagline.ts`. One version everywhere: `bun scripts/set-version.ts`.
- A new lockfile format is a pure parser in `lockfile.rs` with a unit test,
registered in `LOCKFILES` or `parse`.
2 changes: 1 addition & 1 deletion PROJECT.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ Exact-version library docs for AI agents, from the project's lockfile, offline.
It ships as a Rust MCP server and CLI, runs locally, needs no API key, and is
MIT licensed.

- Lifecycle: `active`, published as `@sylphx/lockdocs` (npm) and
- Published as `@sylphx/lockdocs` (npm) and
`io.github.SylphxAI/lockdocs` (MCP Registry)
- Docs: https://sylphxai.github.io/lockdocs/

Expand Down
2 changes: 1 addition & 1 deletion docs/benchmarks.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@
## Method

- **Questions:** 105 questions whose correct answer depends on the version, in [`bench/questions.json`](https://github.com/SylphxAI/lockdocs/blob/main/bench/questions.json), over zod 3/4, Next.js 14/15/16, React Router 6/7, pydantic 1/2, axum 0.7/0.8, tokio, Tailwind CSS 3/4, ESLint 8/9, Prisma 5/6, React 18/19, Vite 5/6, Express 4/5, SQLAlchemy 1.4/2.0, Django 4.2/5.1 and FastAPI 0.88/0.115. Twin questions are worded identically for both versions; each records its source of truth.
- **Held-out questions:** 18 questions marked `held-out` are not used for tuning: they were written before the ranking changes they measure and have their own column. Held-out questions that are later used to diagnose a miss join the main set (their `history` field says so), and new ones replace them. The 0.3 round did this once: 17 questions first written as held-out (lockdocs 11/17, Context7 12/17 on their first run) are now in the main set.
- **Held-out questions:** 18 questions marked `held-out` are not used for tuning: they were written before the ranking changes they measure and have their own column. Held-out questions that are later used to diagnose a miss join the main set (their `history` field says so), and new ones replace them.
- **Projects:** one project per version in [`bench/projects`](https://github.com/SylphxAI/lockdocs/tree/main/bench/projects), installed at those pins by [`bench/setup.sh`](https://github.com/SylphxAI/lockdocs/blob/main/bench/setup.sh).
- **Grading:** an answer passes when it contains at least one string from every `expect` group (the version-correct API, in code or prose form) and none of the `reject` strings (the other version's API). Case-insensitive; the same grader for every tool.
- **lockdocs**, default settings (1,200-token budget), in three configurations: keyword only (`LOCKDOCS_EMBED=0`) on package files; hybrid on package files (the offline default once the model is downloaded); hybrid after `lockdocs fetch` added upstream docs at each version's tag. Index and fetch times are reported separately.
Expand Down
1 change: 0 additions & 1 deletion docs/capabilities.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,4 +18,3 @@ is in [vision.md](vision.md).
| LD-CLI | Command line with the same queries, plus `index`, `fetch` and `setup` | supported | crates/lockdocs/src/main.rs, crates/lockdocs/src/setup.rs | LD-SEARCH |
| LD-NPM | npm launcher and native binaries for five platforms | supported | packages/lockdocs, packages/npm | LD-CLI |
| LD-BENCH | Version-sensitive benchmark against Context7, run on GitHub-hosted runners | supported | bench/run.py, bench/questions.json, .github/workflows/bench.yml | LD-CLI |
| LD-SEMANTIC-RERANK | Rerank the top results with a stronger local model for questions worded differently from the docs | planned | | LD-SEARCH |
Loading