Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
14 changes: 7 additions & 7 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -70,8 +70,8 @@ lockdocs takes the version question off the table:

- **Exact version, zero config.** It reads `package-lock.json`, `pnpm-lock.yaml`, `yarn.lock`, `bun.lock`, `Cargo.lock`, `uv.lock`, `poetry.lock`, `Pipfile.lock`, `requirements*.txt` and `go.mod`. No library IDs, no "use v14" in the prompt.
- **Docs that ship with the code.** READMEs, changelogs and `docs/` folders, plus the API reference in the package itself: `.d.ts` declarations with JSDoc, Python docstrings and stubs, rustdoc comments, Go doc comments. If it is installed, it is documented, including your private and internal packages.
- **Offline and unlimited.** Everything is read from `node_modules`, your virtualenv, `~/.cargo/registry` and the Go module cache. Offline after a one-time model download (129 MB, kept as 32 MB); keyword-only mode (`LOCKDOCS_EMBED=0`) needs no network at all. No account, no rate limit, and nothing about your dependencies leaves your machine.
- **Upstream docs at the exact tag, when you want them.** Packages like Next.js, Django and FastAPI ship no docs. `lockdocs fetch` pulls their docs folders from GitHub at the git tag of your pinned version, once, then stays offline.
- **Offline and unlimited after caching.** Package files are read from `node_modules`, your virtualenv, `~/.cargo/registry` and the Go module cache. The optional model downloads once (129 MB, kept as 32 MB). Use `--offline` (or `LOCKDOCS_OFFLINE=1`) to prohibit all downloads; `LOCKDOCS_EMBED=0` disables only the model. No account or hosted query quota. First-use upstream requests reveal the public repository and version being fetched, not your question or project files.
- **Upstream docs on first use.** Packages like Next.js, Django and FastAPI ship no docs. Queries automatically add public GitHub docs at the immutable commit resolved from your pinned release tag, anonymously and with bounded downloads. `--no-fetch` or `LOCKDOCS_FETCH=0` opts out; cached docs still work offline. Failures are reported alongside local answers, never replaced with latest-version docs.
- **Meaning, not just words.** Hybrid retrieval: BM25 fused with a small local embedding model (downloaded once, 32 MB on disk), plus API redirects from deprecation notes ("use `model_validate` instead").
- **Small, cited answers.** Packed into a token budget (1,200 by default), every section cited as `package@version path:line`.

Expand Down Expand Up @@ -111,13 +111,13 @@ The same call in a pydantic 1 project answers that pydantic 1.10.18 has no `mode

Legacy copies bundled inside a package (`zod/v3` inside zod 4, `pydantic/v1` inside pydantic 2) rank below the current API.

### Upstream docs and missing packages (opt-in)
### Upstream docs on first use

`lockdocs fetch` adds, once, each direct dependency's upstream docs: it finds the GitHub repository in the package's own metadata and the git tag of your pinned version, and downloads only the docs folders at that tag (Markdown, MDX, reStructuredText, docs examples). Answers then cite `next@15.1.0 upstream:docs/01-app/.../cookies.mdx:12`. See [Upstream docs and fetching](https://sylphxai.github.io/lockdocs/guide/fetch).
The first `docs` or `api` query adds the selected dependencies' public release-tag docs once: it finds the GitHub repository in the package's own metadata and the git tag of your pinned version, resolves that tag to an immutable commit, and downloads only its docs folders (Markdown, MDX, reStructuredText, docs examples). It does not use ambient GitHub credentials or download major-version website docs by default. Previously opted-in docs-site caches are preserved and their provenance stays visible. `lockdocs fetch` remains available to prewarm docs and explicitly add major-version docs sites. Answers then cite `next@15.1.0 upstream:docs/01-app/.../cookies.mdx:12`. See [Upstream docs and fetching](https://sylphxai.github.io/lockdocs/guide/fetch).

### Not installed? Fetch the exact version (opt-in)

Out of the box lockdocs reads only your disk (plus the one-time embedding model download). If a pinned package is not installed (a fresh clone, CI, a lockfile you are reviewing), it says so and tells you how to install it. Pass `--fetch` (or set `LOCKDOCS_FETCH=1`, or `lockdocs setup --fetch`) to let it download exactly that version from the registry (npm tarball, PyPI wheel or sdist, crates.io `.crate`, Go module proxy zip) into its cache. Fetched answers say `fetched from registry.npmjs.org`. You can also ask for a version you do not use: `lockdocs npm:zod@4.1.5 "strict object" --fetch`.
Registry package downloads remain opt-in; default upstream enrichment uses installed package metadata (plus the one-time embedding model download). If a pinned package is not installed (a fresh clone, CI, a lockfile you are reviewing), it says so and tells you how to install it. Pass `--fetch` (or set `LOCKDOCS_FETCH=1`, or `lockdocs setup --fetch`) to let it download exactly that version from the registry (npm tarball, PyPI wheel or sdist, crates.io `.crate`, Go module proxy zip) into its cache. Fetched answers say `fetched from registry.npmjs.org`. You can also ask for a version you do not use: `lockdocs npm:zod@4.1.5 "strict object" --fetch`.

## Benchmarks

Expand Down Expand Up @@ -176,14 +176,14 @@ lockdocs cache [clean] Show or delete the cache
lockdocs setup Configure MCP clients (--client a,b --dry-run --remove --fetch)
lockdocs mcp MCP server on stdio

Options: -C/--root <dir>, --pkg <package>, --tokens <n>, --fetch, --offline, --json
Options: -C/--root <dir>, --pkg <package>, --tokens <n>, --fetch, --no-fetch, --offline, --json
```

Prebuilt binaries for macOS (arm64, x64), Linux glibc (x64, arm64) and Windows x64 ship through npm; each [GitHub release](https://github.com/SylphxAI/lockdocs/releases) has them too. From source: `cargo install --git https://github.com/SylphxAI/lockdocs lockdocs`.

## Privacy

lockdocs reads files on your machine and answers over stdio. Network use: the embedding model once from huggingface.co (pinned revision, SHA-256 checked; `LOCKDOCS_EMBED=0` or `--offline` skips it), and, only when you run `lockdocs fetch` or enable fetching, public registries and GitHub for the exact package versions requested. Nothing about your project is sent. The cache lives in your OS cache directory (`LOCKDOCS_CACHE` overrides it).
lockdocs reads files on your machine and answers over stdio. Network use: the embedding model once from huggingface.co (pinned revision, SHA-256 checked; `LOCKDOCS_EMBED=0` or `--offline` skips it), and anonymous GitHub requests for public docs at the resolved release commit on first query. `--fetch` / `lockdocs fetch` also access public registries and major-version docs-site repositories; only these explicit fetches may use `GITHUB_TOKEN` / `GH_TOKEN`. Upstream requests disclose the repository, version and file paths, not your question, lockfile or project files. `--no-fetch` / `LOCKDOCS_FETCH=0` disables query package/docs downloads; `--offline` / `LOCKDOCS_OFFLINE=1` prohibits all downloads. The cache lives in your OS cache directory (`LOCKDOCS_CACHE` overrides it).

## Also from Sylphx

Expand Down
44 changes: 36 additions & 8 deletions bench/run.py
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,7 @@
context for the question) on the anonymous tier, trying the next search result
when a library answers HTTP 404; 429s are recorded, not retried.
"""
import json, os, subprocess, sys, time, urllib.parse, urllib.request
import json, os, subprocess, sys, tempfile, time, urllib.parse, urllib.request

HERE = os.path.dirname(os.path.abspath(__file__))

Expand Down Expand Up @@ -135,15 +135,21 @@ def run_context7(q, version):


VARIANTS = [
("keyword", "lockdocs, keyword only (BM25), package files", {"LOCKDOCS_EMBED": "0", "LOCKDOCS_NO_UPSTREAM": "1"}),
("hybrid", "lockdocs, hybrid (BM25 + embeddings), package files", {"LOCKDOCS_NO_UPSTREAM": "1"}),
("fetched", "lockdocs, hybrid + upstream docs (after `lockdocs fetch`)", {}),
("keyword", "lockdocs, keyword only (BM25), package files", {"LOCKDOCS_EMBED": "0", "LOCKDOCS_NO_UPSTREAM": "1", "LOCKDOCS_FETCH": "0"}),
("hybrid", "lockdocs, hybrid (BM25 + embeddings), package files", {"LOCKDOCS_NO_UPSTREAM": "1", "LOCKDOCS_FETCH": "0"}),
("first-use-default", "lockdocs, real first-use defaults (empty isolated cache, anonymous release-tag fetch)",
{"LOCKDOCS_FETCH": None, "LOCKDOCS_NO_UPSTREAM": None, "LOCKDOCS_EMBED": None, "LOCKDOCS_OFFLINE": None, "GITHUB_TOKEN": None, "GH_TOKEN": None}),
("fetched", "lockdocs, hybrid + upstream docs (after `lockdocs fetch`)", {"LOCKDOCS_FETCH": "1"}),
]


def run_lockdocs_env(binary, proj, q, env):
old = {k: os.environ.get(k) for k in env}
os.environ.update(env)
for k, v in env.items():
if v is None:
os.environ.pop(k, None)
else:
os.environ[k] = v
try:
return run_lockdocs(binary, proj, q)
finally:
Expand Down Expand Up @@ -191,12 +197,17 @@ def main():
for p in projs:
proj = os.path.join(projects, p)
t = time.perf_counter()
r = subprocess.run([binary, "index", "-C", proj, "--json"], capture_output=True, text=True, env={**os.environ, "LOCKDOCS_NO_UPSTREAM": "1"})
r = subprocess.run([binary, "index", "-C", proj, "--json"], capture_output=True, text=True, env={**os.environ, "LOCKDOCS_NO_UPSTREAM": "1", "LOCKDOCS_FETCH": "0"})
index[p] = {"ms": round((time.perf_counter() - t) * 1000), "report": json.loads(r.stdout) if r.returncode == 0 else r.stderr}
variants = [v for v in VARIANTS if do_fetch or v[0] in ("keyword", "hybrid")]
variants = [v for v in VARIANTS if do_fetch or v[0] in ("keyword", "hybrid", "first-use-default")]
results = {}
fetch = {}
# A real query run before prefetch, with its own initially empty supported cache.
# It never borrows upstream files from the explicit-fetch score.
first_use_cache = tempfile.TemporaryDirectory(prefix="lockdocs-first-use-")
for name, _, env in variants:
if name == "first-use-default":
env = {**env, "LOCKDOCS_CACHE": first_use_cache.name}
if name == "fetched" and not fetch:
for p in projs:
t = time.perf_counter()
Expand All @@ -209,6 +220,7 @@ def main():
first = text.splitlines()[0] if text else ""
version = first.split(" · ")[0].rsplit("@", 1)[-1] if "@" in first else ""
results[(name, q["id"])] = ({"pass": ok and code == 0, "missing": missing, "rejected": rej, "tokens": count(text), "ms": round(ms, 1)}, version)
first_use_cache.cleanup()
main_variant = "fetched" if do_fetch else variants[-1][0]
rows = []
for q in qs:
Expand All @@ -232,6 +244,15 @@ def main():
"context7_calls": dict(c7_state, reused=reused) if with_c7 else None, "runner": {"os": os.uname().sysname, "machine": os.uname().machine}}
json.dump(res, open(out, "w"), indent=1)
print(markdown(res, with_c7))
check_floors(summary, len(qs))


def check_floors(summary, total):
if total != 105:
return
for variant, floor in [("fetched", 96), ("hybrid", 60)]:
if variant in summary and summary[variant]["passed"] < floor:
raise RuntimeError(f"{variant} regressed: {summary[variant]['passed']}/105, required >= {floor}/105")


def agg(rs, total):
Expand All @@ -247,6 +268,13 @@ def result_of(r, key):
return r.get("context7", {}) if key == "context7" else r["variants"].get(key, {})


def fetch_file_count(package):
"""Accept the original flat fetch report and the additive manifest layout."""
if "files" in package:
return package["files"] or 0
return (package.get("upstream") or {}).get("files", 0)


def markdown(res, with_c7):
s = res["summary"]
cols = [(n, label) for n, label in res.get("variants", [["lockdocs", "lockdocs"]])]
Expand Down Expand Up @@ -290,7 +318,7 @@ def markdown(res, with_c7):
for p, v in res["fetch"].items():
rep = v["report"]
if isinstance(rep, dict):
pk = ", ".join(f"{x['package']} {x.get('files', 0)} files" for x in rep.get("packages", []) if x.get("files"))
pk = ", ".join(f"{x['package']} {fetch_file_count(x)} files" for x in rep.get("packages", []) if fetch_file_count(x))
out.append(f"- {p}: {v['ms']} ms ({pk or 'no upstream docs'})")
else:
out.append(f"- {p}: {v['ms']} ms (error)")
Expand Down
90 changes: 90 additions & 0 deletions bench/test_run.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,90 @@
#!/usr/bin/env python3
"""Network-free tests of benchmark isolation and regression gates."""
import contextlib
import importlib.util
import io
import json
import os
from pathlib import Path
import sys
import tempfile
import unittest
from unittest.mock import patch

spec = importlib.util.spec_from_file_location("bench_runner", Path(__file__).with_name("run.py"))
runner = importlib.util.module_from_spec(spec)
spec.loader.exec_module(runner)


class BenchmarkTests(unittest.TestCase):
def test_floors_apply_only_to_full_suite(self):
runner.check_floors({"fetched": {"passed": 96}, "hybrid": {"passed": 60}}, 105)
runner.check_floors({"fetched": {"passed": 0}}, 1)
for variant, score in [("fetched", 95), ("hybrid", 59)]:
with self.assertRaises(RuntimeError):
runner.check_floors({variant: {"passed": score}}, 105)

def test_fetch_formatter_preserves_flat_and_nested_file_counts(self):
for package in [
{"package": "axum@0.7.9", "files": 20},
{"package": "axum@0.7.9", "upstream": {"files": 20}},
{"package": "axum@0.7.9", "files": 20, "upstream": {"files": 20}},
]:
self.assertEqual(runner.fetch_file_count(package), 20)
result = {"tokenizer": "test", "runner": {"os": "Test", "machine": "test"},
"rows": [], "summary": {}, "variants": [], "index": {},
"fetch": {"axum07": {"ms": 1, "report": {"packages": [package]}}}}
formatted = runner.markdown(result, False)
self.assertIn("axum@0.7.9 20 files", formatted)
self.assertNotIn("no upstream docs", formatted)
self.assertEqual(runner.fetch_file_count({"files": 0, "upstream": {"files": 20}}), 0)

def test_first_use_has_own_empty_cache_and_no_opt_in(self):
with tempfile.TemporaryDirectory(prefix="lockdocs-bench-test-") as tmp:
root = Path(tmp)
(root / "questions.json").write_text(json.dumps({"questions": [{
"id": "fixture", "project": "project", "package": "fixture",
"question": "documented API", "why": "harness test", "expect": [["right_api"]],
}]}))
(root / "projects" / "project").mkdir(parents=True)
default_cache = root / "regular-cache"
default_cache.mkdir()
(default_cache / "prefetched").write_text("already present")
log = root / "calls.jsonl"
binary = root / "lockdocs"
binary.write_text('''#!/usr/bin/env python3
import json, os, pathlib, sys
cache = pathlib.Path(os.environ["LOCKDOCS_CACHE"])
cmd = sys.argv[1]
with open(os.environ["TEST_LOG"], "a") as log:
log.write(json.dumps({"cmd": cmd, "cache": str(cache), "empty": not any(cache.iterdir()),
"fetch": os.environ.get("LOCKDOCS_FETCH"), "token": os.environ.get("GITHUB_TOKEN"),
"no_upstream": os.environ.get("LOCKDOCS_NO_UPSTREAM")}) + "\\n")
if cmd in ("index", "fetch"):
print("{}")
else:
(cache / "queried").write_text("cached")
print("fixture@1.0.0 · npm · source\\nright_api")
''')
binary.chmod(0o755)
out = root / "results.json"
with patch.object(runner, "HERE", str(root)), patch.object(sys, "argv", [
"run.py", str(binary), str(root / "projects"), str(out), "--fetch",
]), patch.dict(os.environ, {
"LOCKDOCS_CACHE": str(default_cache), "TEST_LOG": str(log),
"LOCKDOCS_FETCH": "1", "GITHUB_TOKEN": "test-placeholder",
}), contextlib.redirect_stdout(io.StringIO()), contextlib.redirect_stderr(io.StringIO()):
runner.main()
calls = [json.loads(line) for line in log.read_text().splitlines()]
cold = next(c for c in calls if c["cmd"] == "docs" and c["cache"] != str(default_cache))
self.assertTrue(cold["empty"])
self.assertIsNone(cold["fetch"])
self.assertIsNone(cold["token"])
self.assertIsNone(cold["no_upstream"])
self.assertLess(calls.index(cold), next(i for i, c in enumerate(calls) if c["cmd"] == "fetch"))
self.assertEqual(json.loads(out.read_text())["summary"]["first-use-default"]["passed"], 1)
self.assertTrue((default_cache / "prefetched").exists())


if __name__ == "__main__":
unittest.main()
22 changes: 19 additions & 3 deletions crates/lockdocs-core/src/fetch.rs
Original file line number Diff line number Diff line change
Expand Up @@ -41,6 +41,9 @@ pub fn fetched_dir(dep: &Dep) -> PathBuf {

/// A previously fetched copy, if any (never touches the network).
pub fn cached(dep: &Dep) -> Option<Source> {
if is_git(dep) {
return None;
}
let dir = fetched_dir(dep);
if dir.join(".lockdocs-complete").is_file() {
return Some(source_for(dep, &dir));
Expand Down Expand Up @@ -78,13 +81,26 @@ fn source_for(dep: &Dep, dir: &Path) -> Source {
}

/// Download and unpack `dep` if it is not cached yet.
pub fn is_git(dep: &Dep) -> bool {
dep.from.ends_with("(git)")
}

/// Shared guard: a git checkout is not interchangeable with a registry release tag.
pub fn require_registry_origin(dep: &Dep) -> Result<()> {
if is_git(dep) {
bail!(
"{} is a git dependency; use its resolved checkout files, not registry/release-tag docs",
dep.id()
);
}
Ok(())
}

pub fn fetch(dep: &Dep) -> Result<Source> {
require_registry_origin(dep)?;
if let Some(s) = cached(dep) {
return Ok(s);
}
if dep.from.ends_with("(git)") {
bail!("{} is a git dependency; run `cargo fetch` to check it out", dep.id());
}
let dir = fetched_dir(dep);
let _ = std::fs::remove_dir_all(&dir);
std::fs::create_dir_all(&dir)?;
Expand Down
Loading
Loading