Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 8 additions & 1 deletion .github/workflows/bench.yml
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,10 @@ on:
description: Also query Context7 (anonymous tier)
type: boolean
default: true
variants:
description: 'Extra configurations, space-separated name:KEY=VALUE,KEY=VALUE (e.g. head0:LOCKDOCS_HEAD_WEIGHT=0)'
type: string
default: ''
pull_request:
paths: ['bench/**', '.github/workflows/bench.yml']

Expand Down Expand Up @@ -44,10 +48,13 @@ jobs:
- name: Run (package files, then `lockdocs fetch` + upstream docs)
env:
C7: ${{ (github.event_name != 'workflow_dispatch' || inputs.context7) && '--context7' || '' }}
VARIANTS: ${{ inputs.variants }}
GITHUB_TOKEN: ${{ github.token }}
run: |
set -euo pipefail
python3 bench/run.py target/release/lockdocs bench/projects bench.json --fetch $C7 --context7-cache docs/public/bench-results.json | tee -a "$GITHUB_STEP_SUMMARY"
extra=()
for v in $VARIANTS; do extra+=(--variant "$v"); done
python3 bench/run.py target/release/lockdocs bench/projects bench.json --fetch $C7 --context7-cache docs/public/bench-results.json "${extra[@]}" | tee -a "$GITHUB_STEP_SUMMARY"
- uses: actions/upload-artifact@v7
with:
name: bench
Expand Down
3 changes: 2 additions & 1 deletion .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -50,12 +50,13 @@ jobs:
components: clippy, rustfmt
- uses: Swatinem/rust-cache@6323deb102c322ba6fcbdcafc7e3dddab59af2b6 # v2
- run: bun install --frozen-lockfile
- name: Manifests agree on one version and one tagline
- name: Manifests agree on one version and one tagline; capability paths exist
env:
GH_TOKEN: ${{ github.token }}
run: |
bun scripts/check-version.ts
bun scripts/check-tagline.ts "$(gh api "repos/$GITHUB_REPOSITORY" --jq .description)"
bun scripts/check-capabilities.ts
- name: Format and lints
run: cargo fmt --all --check && cargo clippy --workspace --locked -- -D warnings
- name: Build release binary
Expand Down
5 changes: 4 additions & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,10 @@ Read [PROJECT.md](PROJECT.md) for the layout and the release flow.
run it in CI (`bench.yml`), not on a shared machine.
- Keep the MCP surface at three tools: `resolve`, `docs`, `api`.
- Bump `index::FORMAT` whenever extraction output or `Entry` changes, so old
caches are rebuilt.
caches are rebuilt; bump `upstream::FORMAT` when `fetch` downloads more, so
`lockdocs fetch` refreshes old copies.
- Ranking changes are judged on the whole benchmark in CI, never on one
question; questions marked `held-out` are not used for tuning.
- A new lockfile format: add a pure parser to `lockfile.rs` with a unit test,
and register it in `LOCKFILES` or `parse`.
- One version everywhere: `bun scripts/set-version.ts` (CI runs
Expand Down
9 changes: 9 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,14 @@
# Changelog

## 0.3.0

- **Docs sites follow your major.** For packages whose docs live in a separate website repository, `lockdocs fetch` takes the docs for the pinned major: the default branch when your major is the latest, a `vN` / `N.x` branch when the site keeps one (Tailwind CSS v3, Prisma v6), or the last commit before the next major was released. Pages about a later major are skipped. React keeps the latest-only rule (react.dev documents APIs before they ship). tokio's website (tutorial and topics) is added, and docs sites now work for crates and PyPI packages too.
- **More of the docs are read.** Django's docs (reStructuredText in `.txt` files) were downloaded but not indexed; they are now. Docs pages written as React components (Tailwind's installation guides) are indexed. HTML headings in MDX split sections, and `export const title` names the page.
- **Ranking.** A second BM25 over just the heading (or name) and first sentence keeps a long body from burying what an entry says it is. Question words that name a documented top-level API of the package ("run code *after* the response" in Next.js, which exports `after`) count as identifiers. Generic headings (Parameters, Returns, Examples) take their topic from the heading above; MDX heading ids (`{/*usage*/}`) are dropped; capitalized words in headings stay whole (TypeScript no longer matches "type"). Upgrade guides to an older major than yours and pages titled "(Deprecated)" rank lower unless the question is about changes. Code-only sections are embedded with their code, not their title alone. The stemmer pairs -ation/-ate and -ability/-able, and the synonym table adds parameter/param, JavaScript/JS, TypeScript/TS and database/DB.
- **Answers.** The top result quotes the first paragraph of its page and parent section, with the short list or code block that follows, when they add something (React's "In React 19, forwardRef is no longer necessary" at the top of the page; the `@custom-variant` setup a Tailwind subsection builds on).
- **Fetch.** GitHub API redirects (renamed repositories) keep the token, docs-site errors appear in the fetch note, and `lockdocs fetch` refreshes copies made by older versions.
- **Benchmark:** 105 questions (was 70). 17 questions written as held-out were used to diagnose misses after their first run and joined the main set; 18 new held-out questions, not used for tuning, have their own column. Context7 answers are reused only when the question and its grading are unchanged, and a Context7 library that answers HTTP 404 falls back to the next search result. The pydantic 1 graders reject the v2 idiom `model_config = ConfigDict` instead of `ConfigDict`, which pydantic 1.10 also ships. `bench.yml` takes a `variants` input for weight sweeps.

## 0.2.1

- **MCP server** now runs on [mcp-kit](https://github.com/SylphxAI/mcp-kit), which uses rmcp, the official Rust MCP SDK, instead of lockdocs' own JSON-RPC loop. Tools and answers are unchanged. The server now also handles protocol negotiation across every spec version, cancellation, progress and pagination.
Expand Down
4 changes: 2 additions & 2 deletions Cargo.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

2 changes: 1 addition & 1 deletion Cargo.toml
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@ members = ["crates/lockdocs-core", "crates/lockdocs"]
resolver = "2"

[workspace.package]
version = "0.2.1"
version = "0.3.0"
edition = "2021"
license = "MIT"
repository = "https://github.com/SylphxAI/lockdocs"
Expand Down
18 changes: 9 additions & 9 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -121,18 +121,18 @@ Out of the box lockdocs reads only your disk (plus the one-time embedding model
## Benchmarks

<!-- bench:start -->
70 questions whose correct answer depends on the version, over 15 libraries (zod, Next.js, React Router, pydantic, axum, tokio, Tailwind CSS, ESLint, Prisma, React, Vite, Express, SQLAlchemy, Django, FastAPI), each asked in a real project with that version installed. An answer passes when it contains the version-correct API and none of the other version's. Same questions and grader against Context7's anonymous API, on a GitHub-hosted runner ([run](https://github.com/SylphxAI/lockdocs/actions/runs/36125903106)):
105 questions whose correct answer depends on the version, over 15 libraries (zod, Next.js, React Router, pydantic, axum, tokio, Tailwind CSS, ESLint, Prisma, React, Vite, Express, SQLAlchemy, Django, FastAPI), each asked in a real project with that version installed. An answer passes when it contains the version-correct API and none of the other version's. Same questions and grader against Context7's anonymous API, on a GitHub-hosted runner ([run](https://github.com/SylphxAI/lockdocs/actions/runs/36204404336)):

| | correct | older majors | newer majors | tokio | median tokens | median latency |
|---|---|---|---|---|---|---|
| lockdocs + `lockdocs fetch` | 55/70 | 24/33 | 29/34 | 2/3 | 875 | 87 ms |
| lockdocs, package files only | 46/70 | 22/33 | 22/34 | 2/3 | 915 | 49 ms |
| Context7 (anonymous) | 49/70 | 12/33 | 34/34 | 3/3 | 908 | 2,011 ms |
| | correct | older majors | newer majors | tokio | held-out | median tokens | median latency |
|---|---|---|---|---|---|---|---|
| lockdocs + `lockdocs fetch` | 96/105 | 38/45 | 53/55 | 5/5 | 16/18 | 866 | 97 ms |
| lockdocs, package files only | 60/105 | 27/45 | 29/55 | 4/5 | 6/18 | 903 | 52 ms |
| Context7 (anonymous) | 77/105 | 19/45 | 53/55 | 5/5 | 15/18 | 908 | 2,583 ms |
<!-- bench:end -->

- **Where versions matter most, lockdocs wins by 2x.** On older majors Context7 often answers with the newest API (all five pydantic 1 questions got pydantic 2 answers).
- **Context7 is ahead on the newest majors (34/34 vs 29/34) and on tokio (3/3 vs 2/3)**, and we are working to close that. Its index covers docs websites that no package or tag ships (Prisma's docs now describe a later major), and lockdocs has a few ranking misses. The benchmark page lists every question and answer.
- **~23x faster, no quota.** lockdocs latency is a fresh CLI process per question; `lockdocs fetch` is a one-time 0.6-5 s per project (median 2.4 s).
- **Where versions matter most, lockdocs wins by 2x** (38/45 vs 19/45 on older majors). Context7 often answers with the newest API (pydantic 1 questions get pydantic 2 answers).
- **Tied on the newest majors (53/55 each) and on tokio (5/5 each); ahead on held-out questions (16/18 vs 15/18)**, which were written before the ranking changes they measure and not used for tuning.
- **~27x faster, fewer tokens, no quota.** lockdocs latency is a fresh CLI process per question; `lockdocs fetch` is a one-time download per project.

Method, questions, per-question results and scripts: [benchmark page](https://sylphxai.github.io/lockdocs/benchmarks) and [`bench/`](bench/).

Expand Down
Loading
Loading