From c95db5e5986d92462c0d4f69d7c4289363395b51 Mon Sep 17 00:00:00 2001 From: Sylphx Builder Date: Thu, 1 Oct 2026 00:50:05 +0000 Subject: [PATCH 1/5] feat: fetch immutable release docs on first query by default --- README.md | 14 +- bench/run.py | 35 ++++- bench/test_run.py | 75 ++++++++++ crates/lockdocs-core/src/index.rs | 17 ++- crates/lockdocs-core/src/query.rs | 145 +++++++++++++++----- crates/lockdocs-core/src/upstream.rs | 159 ++++++++++++++++++---- crates/lockdocs-core/tests/cargo_git.rs | 2 +- crates/lockdocs-core/tests/integration.rs | 17 ++- crates/lockdocs-core/tests/pnp.rs | 2 +- crates/lockdocs/src/main.rs | 29 +++- crates/lockdocs/src/mcp.rs | 2 +- crates/lockdocs/src/tools.rs | 20 ++- docs/benchmarks.md | 4 +- docs/capabilities.md | 2 +- docs/guide/fetch.md | 20 ++- docs/reference/cli.md | 5 +- docs/reference/tools.md | 9 +- docs/vision.md | 4 +- 18 files changed, 462 insertions(+), 99 deletions(-) create mode 100644 bench/test_run.py diff --git a/README.md b/README.md index c28cc49..e6b7e5d 100644 --- a/README.md +++ b/README.md @@ -70,8 +70,8 @@ lockdocs takes the version question off the table: - **Exact version, zero config.** It reads `package-lock.json`, `pnpm-lock.yaml`, `yarn.lock`, `bun.lock`, `Cargo.lock`, `uv.lock`, `poetry.lock`, `Pipfile.lock`, `requirements*.txt` and `go.mod`. No library IDs, no "use v14" in the prompt. - **Docs that ship with the code.** READMEs, changelogs and `docs/` folders, plus the API reference in the package itself: `.d.ts` declarations with JSDoc, Python docstrings and stubs, rustdoc comments, Go doc comments. If it is installed, it is documented, including your private and internal packages. -- **Offline and unlimited.** Everything is read from `node_modules`, your virtualenv, `~/.cargo/registry` and the Go module cache. Offline after a one-time model download (129 MB, kept as 32 MB); keyword-only mode (`LOCKDOCS_EMBED=0`) needs no network at all. No account, no rate limit, and nothing about your dependencies leaves your machine. -- **Upstream docs at the exact tag, when you want them.** Packages like Next.js, Django and FastAPI ship no docs. `lockdocs fetch` pulls their docs folders from GitHub at the git tag of your pinned version, once, then stays offline. +- **Offline and unlimited after caching.** Package files are read from `node_modules`, your virtualenv, `~/.cargo/registry` and the Go module cache. Offline after a one-time model download (129 MB, kept as 32 MB); Use `--offline` (or `LOCKDOCS_OFFLINE=1`) to prohibit all downloads; `LOCKDOCS_EMBED=0` disables only the model. No account or hosted query quota. First-use upstream requests reveal the public repository and version being fetched, not your question or project files. +- **Upstream docs on first use.** Packages like Next.js, Django and FastAPI ship no docs. Queries automatically add public GitHub docs at the immutable commit resolved from your pinned release tag, anonymously and with bounded downloads. `--no-fetch` or `LOCKDOCS_FETCH=0` opts out; cached docs still work offline. Failures are reported alongside local answers, never replaced with latest-version docs. - **Meaning, not just words.** Hybrid retrieval: BM25 fused with a small local embedding model (downloaded once, 32 MB on disk), plus API redirects from deprecation notes ("use `model_validate` instead"). - **Small, cited answers.** Packed into a token budget (1,200 by default), every section cited as `package@version path:line`. @@ -111,13 +111,13 @@ The same call in a pydantic 1 project answers that pydantic 1.10.18 has no `mode Legacy copies bundled inside a package (`zod/v3` inside zod 4, `pydantic/v1` inside pydantic 2) rank below the current API. -### Upstream docs and missing packages (opt-in) +### Upstream docs on first use -`lockdocs fetch` adds, once, each direct dependency's upstream docs: it finds the GitHub repository in the package's own metadata and the git tag of your pinned version, and downloads only the docs folders at that tag (Markdown, MDX, reStructuredText, docs examples). Answers then cite `next@15.1.0 upstream:docs/01-app/.../cookies.mdx:12`. See [Upstream docs and fetching](https://sylphxai.github.io/lockdocs/guide/fetch). +The first `docs` or `api` query adds the selected dependencies' public release-tag docs once: it finds the GitHub repository in the package's own metadata and the git tag of your pinned version, resolves that tag to an immutable commit, and downloads only its docs folders (Markdown, MDX, reStructuredText, docs examples). It does not use ambient GitHub credentials or substitute major-version website docs. `lockdocs fetch` remains available to prewarm docs and explicitly add major-version docs sites. Answers then cite `next@15.1.0 upstream:docs/01-app/.../cookies.mdx:12`. See [Upstream docs and fetching](https://sylphxai.github.io/lockdocs/guide/fetch). ### Not installed? Fetch the exact version (opt-in) -Out of the box lockdocs reads only your disk (plus the one-time embedding model download). If a pinned package is not installed (a fresh clone, CI, a lockfile you are reviewing), it says so and tells you how to install it. Pass `--fetch` (or set `LOCKDOCS_FETCH=1`, or `lockdocs setup --fetch`) to let it download exactly that version from the registry (npm tarball, PyPI wheel or sdist, crates.io `.crate`, Go module proxy zip) into its cache. Fetched answers say `fetched from registry.npmjs.org`. You can also ask for a version you do not use: `lockdocs npm:zod@4.1.5 "strict object" --fetch`. +Registry package downloads remain opt-in; default upstream enrichment uses installed package metadata (plus the one-time embedding model download). If a pinned package is not installed (a fresh clone, CI, a lockfile you are reviewing), it says so and tells you how to install it. Pass `--fetch` (or set `LOCKDOCS_FETCH=1`, or `lockdocs setup --fetch`) to let it download exactly that version from the registry (npm tarball, PyPI wheel or sdist, crates.io `.crate`, Go module proxy zip) into its cache. Fetched answers say `fetched from registry.npmjs.org`. You can also ask for a version you do not use: `lockdocs npm:zod@4.1.5 "strict object" --fetch`. ## Benchmarks @@ -176,14 +176,14 @@ lockdocs cache [clean] Show or delete the cache lockdocs setup Configure MCP clients (--client a,b --dry-run --remove --fetch) lockdocs mcp MCP server on stdio -Options: -C/--root , --pkg , --tokens , --fetch, --offline, --json +Options: -C/--root , --pkg , --tokens , --fetch, --no-fetch, --offline, --json ``` Prebuilt binaries for macOS (arm64, x64), Linux glibc (x64, arm64) and Windows x64 ship through npm; each [GitHub release](https://github.com/SylphxAI/lockdocs/releases) has them too. From source: `cargo install --git https://github.com/SylphxAI/lockdocs lockdocs`. ## Privacy -lockdocs reads files on your machine and answers over stdio. Network use: the embedding model once from huggingface.co (pinned revision, SHA-256 checked; `LOCKDOCS_EMBED=0` or `--offline` skips it), and, only when you run `lockdocs fetch` or enable fetching, public registries and GitHub for the exact package versions requested. Nothing about your project is sent. The cache lives in your OS cache directory (`LOCKDOCS_CACHE` overrides it). +lockdocs reads files on your machine and answers over stdio. Network use: the embedding model once from huggingface.co (pinned revision, SHA-256 checked; `LOCKDOCS_EMBED=0` or `--offline` skips it), and anonymous GitHub requests for public docs at the resolved release commit on first query. `--fetch` / `lockdocs fetch` also access public registries and major-version docs-site repositories; only these explicit fetches may use `GITHUB_TOKEN` / `GH_TOKEN`. Upstream requests disclose the repository, version and file paths, not your question, lockfile or project files. `--no-fetch` / `LOCKDOCS_FETCH=0` disables query package/docs downloads; `--offline` / `LOCKDOCS_OFFLINE=1` prohibits all downloads. The cache lives in your OS cache directory (`LOCKDOCS_CACHE` overrides it). ## Also from Sylphx diff --git a/bench/run.py b/bench/run.py index 1a9cfad..a6357f5 100755 --- a/bench/run.py +++ b/bench/run.py @@ -19,7 +19,7 @@ context for the question) on the anonymous tier, trying the next search result when a library answers HTTP 404; 429s are recorded, not retried. """ -import json, os, subprocess, sys, time, urllib.parse, urllib.request +import json, os, subprocess, sys, tempfile, time, urllib.parse, urllib.request HERE = os.path.dirname(os.path.abspath(__file__)) @@ -135,15 +135,21 @@ def run_context7(q, version): VARIANTS = [ - ("keyword", "lockdocs, keyword only (BM25), package files", {"LOCKDOCS_EMBED": "0", "LOCKDOCS_NO_UPSTREAM": "1"}), - ("hybrid", "lockdocs, hybrid (BM25 + embeddings), package files", {"LOCKDOCS_NO_UPSTREAM": "1"}), - ("fetched", "lockdocs, hybrid + upstream docs (after `lockdocs fetch`)", {}), + ("keyword", "lockdocs, keyword only (BM25), package files", {"LOCKDOCS_EMBED": "0", "LOCKDOCS_NO_UPSTREAM": "1", "LOCKDOCS_FETCH": "0"}), + ("hybrid", "lockdocs, hybrid (BM25 + embeddings), package files", {"LOCKDOCS_NO_UPSTREAM": "1", "LOCKDOCS_FETCH": "0"}), + ("first-use-default", "lockdocs, real first-use defaults (empty isolated cache, anonymous release-tag fetch)", + {"LOCKDOCS_FETCH": None, "LOCKDOCS_NO_UPSTREAM": None, "LOCKDOCS_EMBED": None, "LOCKDOCS_OFFLINE": None, "GITHUB_TOKEN": None, "GH_TOKEN": None}), + ("fetched", "lockdocs, hybrid + upstream docs (after `lockdocs fetch`)", {"LOCKDOCS_FETCH": "1"}), ] def run_lockdocs_env(binary, proj, q, env): old = {k: os.environ.get(k) for k in env} - os.environ.update(env) + for k, v in env.items(): + if v is None: + os.environ.pop(k, None) + else: + os.environ[k] = v try: return run_lockdocs(binary, proj, q) finally: @@ -191,12 +197,17 @@ def main(): for p in projs: proj = os.path.join(projects, p) t = time.perf_counter() - r = subprocess.run([binary, "index", "-C", proj, "--json"], capture_output=True, text=True, env={**os.environ, "LOCKDOCS_NO_UPSTREAM": "1"}) + r = subprocess.run([binary, "index", "-C", proj, "--json"], capture_output=True, text=True, env={**os.environ, "LOCKDOCS_NO_UPSTREAM": "1", "LOCKDOCS_FETCH": "0"}) index[p] = {"ms": round((time.perf_counter() - t) * 1000), "report": json.loads(r.stdout) if r.returncode == 0 else r.stderr} - variants = [v for v in VARIANTS if do_fetch or v[0] in ("keyword", "hybrid")] + variants = [v for v in VARIANTS if do_fetch or v[0] in ("keyword", "hybrid", "first-use-default")] results = {} fetch = {} + # A real query run before prefetch, with its own initially empty supported cache. + # It never borrows upstream files from the explicit-fetch score. + first_use_cache = tempfile.TemporaryDirectory(prefix="lockdocs-first-use-") for name, _, env in variants: + if name == "first-use-default": + env = {**env, "LOCKDOCS_CACHE": first_use_cache.name} if name == "fetched" and not fetch: for p in projs: t = time.perf_counter() @@ -209,6 +220,7 @@ def main(): first = text.splitlines()[0] if text else "" version = first.split(" · ")[0].rsplit("@", 1)[-1] if "@" in first else "" results[(name, q["id"])] = ({"pass": ok and code == 0, "missing": missing, "rejected": rej, "tokens": count(text), "ms": round(ms, 1)}, version) + first_use_cache.cleanup() main_variant = "fetched" if do_fetch else variants[-1][0] rows = [] for q in qs: @@ -232,6 +244,15 @@ def main(): "context7_calls": dict(c7_state, reused=reused) if with_c7 else None, "runner": {"os": os.uname().sysname, "machine": os.uname().machine}} json.dump(res, open(out, "w"), indent=1) print(markdown(res, with_c7)) + check_floors(summary, len(qs)) + + +def check_floors(summary, total): + if total != 105: + return + for variant, floor in [("fetched", 96), ("hybrid", 60)]: + if variant in summary and summary[variant]["passed"] < floor: + raise RuntimeError(f"{variant} regressed: {summary[variant]['passed']}/105, required >= {floor}/105") def agg(rs, total): diff --git a/bench/test_run.py b/bench/test_run.py new file mode 100644 index 0000000..7838236 --- /dev/null +++ b/bench/test_run.py @@ -0,0 +1,75 @@ +#!/usr/bin/env python3 +"""Network-free tests of benchmark isolation and regression gates.""" +import contextlib +import importlib.util +import io +import json +import os +from pathlib import Path +import sys +import tempfile +import unittest +from unittest.mock import patch + +spec = importlib.util.spec_from_file_location("bench_runner", Path(__file__).with_name("run.py")) +runner = importlib.util.module_from_spec(spec) +spec.loader.exec_module(runner) + + +class BenchmarkTests(unittest.TestCase): + def test_floors_apply_only_to_full_suite(self): + runner.check_floors({"fetched": {"passed": 96}, "hybrid": {"passed": 60}}, 105) + runner.check_floors({"fetched": {"passed": 0}}, 1) + for variant, score in [("fetched", 95), ("hybrid", 59)]: + with self.assertRaises(RuntimeError): + runner.check_floors({variant: {"passed": score}}, 105) + + def test_first_use_has_own_empty_cache_and_no_opt_in(self): + with tempfile.TemporaryDirectory(prefix="lockdocs-bench-test-") as tmp: + root = Path(tmp) + (root / "questions.json").write_text(json.dumps({"questions": [{ + "id": "fixture", "project": "project", "package": "fixture", + "question": "documented API", "why": "harness test", "expect": [["right_api"]], + }]})) + (root / "projects" / "project").mkdir(parents=True) + default_cache = root / "regular-cache" + default_cache.mkdir() + (default_cache / "prefetched").write_text("already present") + log = root / "calls.jsonl" + binary = root / "lockdocs" + binary.write_text('''#!/usr/bin/env python3 +import json, os, pathlib, sys +cache = pathlib.Path(os.environ["LOCKDOCS_CACHE"]) +cmd = sys.argv[1] +with open(os.environ["TEST_LOG"], "a") as log: + log.write(json.dumps({"cmd": cmd, "cache": str(cache), "empty": not any(cache.iterdir()), + "fetch": os.environ.get("LOCKDOCS_FETCH"), "token": os.environ.get("GITHUB_TOKEN"), + "no_upstream": os.environ.get("LOCKDOCS_NO_UPSTREAM")}) + "\\n") +if cmd in ("index", "fetch"): + print("{}") +else: + (cache / "queried").write_text("cached") + print("fixture@1.0.0 · npm · source\\nright_api") +''') + binary.chmod(0o755) + out = root / "results.json" + with patch.object(runner, "HERE", str(root)), patch.object(sys, "argv", [ + "run.py", str(binary), str(root / "projects"), str(out), "--fetch", + ]), patch.dict(os.environ, { + "LOCKDOCS_CACHE": str(default_cache), "TEST_LOG": str(log), + "LOCKDOCS_FETCH": "1", "GITHUB_TOKEN": "test-placeholder", + }), contextlib.redirect_stdout(io.StringIO()), contextlib.redirect_stderr(io.StringIO()): + runner.main() + calls = [json.loads(line) for line in log.read_text().splitlines()] + cold = next(c for c in calls if c["cmd"] == "docs" and c["cache"] != str(default_cache)) + self.assertTrue(cold["empty"]) + self.assertIsNone(cold["fetch"]) + self.assertIsNone(cold["token"]) + self.assertIsNone(cold["no_upstream"]) + self.assertLess(calls.index(cold), next(i for i, c in enumerate(calls) if c["cmd"] == "fetch")) + self.assertEqual(json.loads(out.read_text())["summary"]["first-use-default"]["passed"], 1) + self.assertTrue((default_cache / "prefetched").exists()) + + +if __name__ == "__main__": + unittest.main() diff --git a/crates/lockdocs-core/src/index.rs b/crates/lockdocs-core/src/index.rs index dcf926e..e837d43 100644 --- a/crates/lockdocs-core/src/index.rs +++ b/crates/lockdocs-core/src/index.rs @@ -154,7 +154,9 @@ pub fn embed_text(e: &Entry) -> String { } fn cache_path(dep: &Dep, src: &Source, types: Option<&Path>, up: Option<&(PathBuf, Manifest)>, embed: &str) -> PathBuf { - let up_key = up.map(|(_, m)| format!("{}@{:?}:{}:{:?}", m.repo, m.tag, m.files, m.site)).unwrap_or_default(); + let up_key = up + .map(|(_, m)| format!("{}@{:?}:{}:{:?}:{:?}", m.repo, m.tag, m.files, m.site, m.commit)) + .unwrap_or_default(); let key = cache::hash(&[ embed, &up_key, @@ -210,9 +212,16 @@ pub fn build(dep: &Dep, src: &Source, root: &Path, up: Option<&(PathBuf, Manifes head, head_lens, build_ms: t.elapsed().as_millis() as u64, - upstream: up - .filter(|(_, m)| m.files > 0) - .map(|(_, m)| format!("{}@{} ({} files)", m.repo, m.tag.as_deref().unwrap_or("?"), m.files)), + upstream: up.filter(|(_, m)| m.files > 0).map(|(_, m)| { + format!( + "{}@{}{}{} ({} files)", + m.repo, + m.tag.as_deref().unwrap_or("?"), + m.commit.as_ref().map(|c| format!(" commit {c}")).unwrap_or_default(), + m.site.as_ref().map(|s| format!(" + docs site {s}")).unwrap_or_default(), + m.files + ) + }), embed: if model.is_some() { embed::MODEL_ID.to_string() } else { String::new() }, vecs, } diff --git a/crates/lockdocs-core/src/query.rs b/crates/lockdocs-core/src/query.rs index 9269b7a..c2ed1c2 100644 --- a/crates/lockdocs-core/src/query.rs +++ b/crates/lockdocs-core/src/query.rs @@ -24,19 +24,30 @@ const HEAD_WEIGHT: f32 = 0.25; /// Cross-dependency searches index at most this many direct dependencies. const MAX_PACKAGES: usize = 80; -#[derive(Debug, Clone)] +#[derive(Debug, Clone, PartialEq, Eq)] pub struct Options { /// Allow downloading exact versions that are not installed. pub fetch: bool, + /// Anonymous release-tag docs on first use; independent of registry downloads. + pub upstream: bool, } impl Default for Options { fn default() -> Self { - let fetch = std::env::var("LOCKDOCS_FETCH").is_ok_and(|v| matches!(v.as_str(), "1" | "true" | "yes")); - Options { fetch } + let mut fetch = std::env::var("LOCKDOCS_FETCH").is_ok_and(|v| matches!(v.as_str(), "1" | "true" | "yes")); + let mut upstream = automatic_upstream(std::env::var("LOCKDOCS_FETCH").ok().as_deref(), std::env::var("LOCKDOCS_NO_UPSTREAM").is_ok()); + if std::env::var("LOCKDOCS_OFFLINE").is_ok_and(|v| v != "0") { + fetch = false; + upstream = false; + } + Options { fetch, upstream } } } +fn automatic_upstream(fetch: Option<&str>, disabled: bool) -> bool { + !disabled && !fetch.is_some_and(|v| matches!(v, "0" | "false" | "no")) +} + pub struct Answer { pub text: String, pub json: Value, @@ -46,6 +57,7 @@ pub struct Engine { pub project: Project, opts: Options, indexes: Mutex>>, + upstream_notes: Mutex>, } /// An indexed package, the dependency it answers for, and a drift note. @@ -57,6 +69,18 @@ enum Resolved { Missing(Dep, String), } +fn provenance(ready: &[ReadyPkg]) -> Vec { + ready + .iter() + .map(|(idx, dep, note)| { + json!({ + "package": idx.id(), "requested_version": dep.version, "source": idx.source, + "registry_fetched": idx.fetched, "upstream": idx.upstream, "note": note, + }) + }) + .collect() +} + fn lang(eco: Eco) -> &'static str { match eco { Eco::Npm => "ts", @@ -125,6 +149,7 @@ impl Engine { project: Project::load(root), opts, indexes: Mutex::new(HashMap::new()), + upstream_notes: Mutex::new(HashMap::new()), } } @@ -159,21 +184,13 @@ impl Engine { } } if let Some(s) = local { - let note = format!( - "{} pins {}@{}, but {} has {}; showing the installed {}. Reinstall to sync{}.", + return Err(format!( + "{} pins {}, but {} has {}; refusing to substitute the installed version. Reinstall to sync, or enable exact registry fetching (--fetch).", dep.from, - dep.name, - dep.version, + dep.id(), s.label, - s.version, - s.version, - if self.opts.fetch { - "" - } else { - ", or enable fetching (--fetch / LOCKDOCS_FETCH=1) to read the pinned version" - } - ); - return Ok((s, Some(note))); + s.version + )); } let how = match dep.eco { Eco::Npm => "run your package manager's install", @@ -190,15 +207,28 @@ impl Engine { fn index(&self, dep: &Dep) -> Resolved { match self.source(dep) { - Ok((src, note)) => { + Ok((src, mut note)) => { let key = format!("{}:{}@{}:{}", dep.eco, dep.name, src.version, src.dir.display()); if let Some(i) = self.indexes.lock().unwrap().get(&key) { // Rebuild once the embedding model has arrived. if !i.embed.is_empty() || embed::get().is_none() { + if let Some(n) = self.upstream_notes.lock().unwrap().get(&dep.id()) { + note = Some(match note { + Some(old) => format!("{old} {n}"), + None => n.clone(), + }); + } return Resolved::Ready(i.clone(), dep.clone(), note); } } - let up = self.upstream(dep, &src); + let (up, up_note) = self.upstream(dep, &src); + if let Some(n) = up_note { + self.upstream_notes.lock().unwrap().insert(dep.id(), n.clone()); + note = Some(match note { + Some(old) => format!("{old} {n}"), + None => n, + }); + } let idx = Arc::new(index::load_or_build(dep, &src, &self.project.root, up.as_ref())); self.indexes.lock().unwrap().insert(key, idx.clone()); Resolved::Ready(idx, dep.clone(), note) @@ -209,19 +239,42 @@ impl Engine { /// Upstream docs for this exact version: a cached copy, or (when fetching /// is enabled) a one-time download. - fn upstream(&self, dep: &Dep, src: &Source) -> Option<(PathBuf, upstream::Manifest)> { + fn upstream(&self, dep: &Dep, src: &Source) -> (Option<(PathBuf, upstream::Manifest)>, Option) { if src.version != dep.version || std::env::var("LOCKDOCS_NO_UPSTREAM").is_ok() { - return None; + return (None, None); + } + let enabled = self.opts.fetch || self.opts.upstream; + if dep.from.ends_with("(git)") { + return ( + None, + enabled.then(|| "upstream docs skipped: git dependency's resolved commit is not a release tag; using its checkout files".into()), + ); } let cached = upstream::cached(dep); - if let Some(c) = cached.as_ref().filter(|(_, m)| !(self.opts.fetch && upstream::stale(m))) { - return Some(c.clone()); + if let Some(c) = cached + .as_ref() + .filter(|(_, m)| !enabled || (!upstream::stale(m) && (self.opts.fetch || (m.site.is_none() && (m.commit.is_some() || m.files == 0))))) + { + return (Some(c.clone()), c.1.note.clone()); } - if self.opts.fetch && upstream::repo_of(dep, src).is_some() { - let _ = upstream::fetch(dep, src); - return upstream::cached(dep); + if enabled { + if upstream::repo_of(dep, src).is_none() { + return ( + None, + Some("upstream docs unavailable: no GitHub repository in package metadata; using package files".into()), + ); + } + let result = if self.opts.fetch { + upstream::fetch(dep, src) + } else { + upstream::fetch_automatic(dep, src) + }; + return match result { + Ok(m) => (Some((upstream::dir(dep), m.clone())), m.note.clone()), + Err(e) => (None, Some(format!("upstream docs fetch failed: {e:#}; using package files"))), + }; } - None + (None, None) } /// `lockdocs fetch`: download what makes answers complete, once: missing @@ -453,7 +506,7 @@ impl Engine { if let Some(n) = note { text.push_str(&format!(" note: {n}\n")); } - items.push(json!({"package": idx.id(), "ecosystem": idx.eco.as_str(), "symbols": symbols, "sections": idx.entries.len() - symbols, "source": idx.source, "build_ms": idx.build_ms})); + items.push(json!({"package": idx.id(), "ecosystem": idx.eco.as_str(), "symbols": symbols, "sections": idx.entries.len() - symbols, "source": idx.source, "upstream": idx.upstream, "note": note, "build_ms": idx.build_ms})); } for (d, e) in &missing { text.push_str(&format!(" {:<40} missing: {e}\n", d.id())); @@ -625,7 +678,7 @@ impl Engine { if !used_pkgs.contains(&idx.id()) { used_pkgs.push(idx.id()); } - hits_json.push(json!({"package": idx.id(), "kind": e.kind.as_str(), "path": e.path, "file": e.file, "line": e.line, "score": score})); + hits_json.push(json!({"provenance": provenance(&ready), "package": idx.id(), "kind": e.kind.as_str(), "path": e.path, "file": e.file, "line": e.line, "score": score})); } if out.blocks == 0 { out.push_raw(&format!( @@ -647,7 +700,7 @@ impl Engine { } } Ok(Answer { - json: json!({"query": q, "packages": ready.iter().map(|r| r.0.id()).collect::>(), "hits": hits_json, "tokens": est_tokens(&out.text)}), + json: json!({"query": q, "packages": ready.iter().map(|r| r.0.id()).collect::>(), "hits": hits_json, "provenance": provenance(&ready), "tokens": est_tokens(&out.text)}), text: out.text, }) } @@ -656,6 +709,9 @@ impl Engine { let mut out = Pack::new(tokens); for (idx, dep, note) in ready { out.push_raw(&format!("{} · {} · {} · pinned in {}\n", idx.id(), idx.eco, idx.source, dep.from)); + if let Some(u) = &idx.upstream { + out.push_raw(&format!("Upstream: {u}\n")); + } if let Some(n) = note { out.push_raw(&format!("Note: {n}\n")); } @@ -689,7 +745,7 @@ impl Engine { } } Answer { - json: json!({"packages": ready.iter().map(|r| r.0.id()).collect::>(), "tokens": est_tokens(&out.text)}), + json: json!({"packages": ready.iter().map(|r| r.0.id()).collect::>(), "provenance": provenance(ready), "tokens": est_tokens(&out.text)}), text: out.text, } } @@ -770,6 +826,9 @@ impl Engine { format!(" · pinned in {}", dep.from) } )); + if let Some(u) = &idx.upstream { + out.push_raw(&format!("Upstream: {u}\n")); + } if let Some(n) = note { out.push_raw(&format!("Note: {n}\n")); } @@ -839,7 +898,7 @@ impl Engine { } Ok(Answer { json: json!({ - "package": idx.id(), "kind": best.kind.as_str(), "path": best.path, "signature": best.sig, "doc": best.doc, + "provenance": provenance(&ready), "package": idx.id(), "kind": best.kind.as_str(), "path": best.path, "signature": best.sig, "doc": best.doc, "file": best.file, "line": best.line, "members": members_json, "tokens": est_tokens(&out.text), }), text: out.text, @@ -1472,7 +1531,7 @@ impl Workspace { let stamp = lock_stamp(&root); let mut m = self.engines.lock().unwrap(); if let Some((s, e)) = m.get(&root) { - if *s == stamp { + if *s == stamp && e.opts == *opts { return e.clone(); } } @@ -1491,6 +1550,28 @@ pub fn dep_key(d: &Dep) -> String { mod tests { use super::*; + #[test] + fn upstream_is_default_with_explicit_opt_out() { + assert!(automatic_upstream(None, false)); + assert!(automatic_upstream(Some("1"), false)); + for value in ["0", "false", "no"] { + assert!(!automatic_upstream(Some(value), false)); + } + assert!(!automatic_upstream(None, true)); + } + + #[test] + fn workspace_does_not_reuse_network_policy() { + let ws = Workspace::default(); + let root = std::env::temp_dir(); + let online = Options { fetch: false, upstream: true }; + let offline = Options { fetch: false, upstream: false }; + let a = ws.engine(&root, &online); + let b = ws.engine(&root, &offline); + assert!(!Arc::ptr_eq(&a, &b)); + assert!(!b.opts.upstream); + } + #[test] fn old_upgrade_guides_and_deprecated_pages() { let e = |name: &str, file: &str| Entry { diff --git a/crates/lockdocs-core/src/upstream.rs b/crates/lockdocs-core/src/upstream.rs index 5e116c0..56a7b73 100644 --- a/crates/lockdocs-core/src/upstream.rs +++ b/crates/lockdocs-core/src/upstream.rs @@ -3,7 +3,8 @@ //! This module finds the repository from the package's own metadata, finds //! the tag for the pinned version, and downloads only the docs folders //! (Markdown, MDX, reStructuredText) from GitHub into the cache, once. -//! Network use is opt-in: `lockdocs fetch`, `--fetch` or `LOCKDOCS_FETCH=1`. +//! First-use queries fetch public docs at an immutable release commit. Explicit +//! `fetch` also supports major-version docs sites and optional GitHub credentials. use crate::locate::Source; use crate::{cache, Dep, Eco}; @@ -28,7 +29,7 @@ pub struct Repo { /// Bump when what `fetch` downloads changes, so `lockdocs fetch` refreshes /// older copies (a stale copy is still used until then). -pub const FORMAT: u32 = 2; +pub const FORMAT: u32 = 3; #[derive(Debug, Clone, Serialize, Deserialize)] pub struct Manifest { @@ -37,6 +38,9 @@ pub struct Manifest { pub format: u32, pub repo: String, pub tag: Option, + /// Immutable package-repository commit resolved from the release tag. + #[serde(default)] + pub commit: Option, pub files: usize, pub bytes: u64, /// Why nothing was downloaded, when files == 0. @@ -123,7 +127,7 @@ const DOCS_SITES: &[DocsSite] = &[ ]; /// The latest stable major of a package, from its registry. -fn latest_major(agent: &ureq::Agent, dep: &Dep) -> Option { +fn latest_major(agent: &Client, dep: &Dep) -> Option { let name = dep.name.strip_prefix("@types/").unwrap_or(&dep.name); let (url, pointer) = match dep.eco { Eco::Npm => (format!("https://registry.npmjs.org/{}/latest", name.replace('/', "%2F")), "/version"), @@ -140,7 +144,7 @@ fn latest_major(agent: &ureq::Agent, dep: &Dep) -> Option { /// When the package's next major was released: the commit date of its /// `{major+1}.0.0` tag in the package repository. -fn next_major_date(agent: &ureq::Agent, dep: &Dep, repo: &Repo, major: u64) -> Result> { +fn next_major_date(agent: &Client, dep: &Dep, repo: &Repo, major: u64) -> Result> { let next = Dep { version: format!("{}.0.0", major + 1), ..dep.clone() @@ -165,7 +169,7 @@ struct SiteFiles { } /// Which files of the docs site describe `dep`'s major, if any. -fn site_files(agent: &ureq::Agent, dep: &Dep, pkg_repo: Option<&Repo>) -> Result> { +fn site_files(agent: &Client, dep: &Dep, pkg_repo: Option<&Repo>) -> Result> { let Some(site) = DOCS_SITES.iter().find(|s| s.eco == dep.eco && s.names.contains(&dep.name.as_str())) else { return Ok(None); }; @@ -276,7 +280,7 @@ fn page_file(p: &str) -> bool { && !p.contains("[") } -fn download(agent: &ureq::Agent, repo: &Repo, reference: &str, files: &[(String, u64)], out: &Path) -> Result> { +fn download(agent: &Client, repo: &Repo, reference: &str, files: &[(String, u64)], out: &Path) -> Result> { let pool = rayon::ThreadPoolBuilder::new().num_threads(16).build()?; let results: Vec> = pool.install(|| { files @@ -289,12 +293,16 @@ fn download(agent: &ureq::Agent, repo: &Repo, reference: &str, files: &[(String, enc(reference), enc_path(p) ); + agent.check_deadline()?; let mut res = agent.get(&url).call()?; if res.status().as_u16() != 200 { bail!("HTTP {} for {p}", res.status()); } let mut buf = Vec::new(); res.body_mut().as_reader().take(MAX_FILE + 1).read_to_end(&mut buf)?; + if buf.len() as u64 > MAX_FILE { + bail!("upstream file too large: {p}"); + } let Some(rel) = crate::fetch::safe_rel(Path::new(p), false) else { bail!("unsafe path {p}") }; @@ -325,7 +333,8 @@ pub fn parse_github(url: &str) -> Option<(String, String)> { let mut it = rest.split(['/', '#', '?']).filter(|s| !s.is_empty()); let owner = it.next()?.to_string(); let name = it.next()?.trim_end_matches(".git").to_string(); - if owner.is_empty() || name.is_empty() { + let safe = |s: &str| !s.is_empty() && s != "." && s != ".." && s.bytes().all(|b| b.is_ascii_alphanumeric() || b"-_.".contains(&b)); + if !safe(&owner) || !safe(&name) { return None; } Some((owner, name)) @@ -419,23 +428,52 @@ fn token() -> Option { .find_map(|k| std::env::var(k).ok().filter(|v| !v.is_empty())) } -fn agent() -> ureq::Agent { - ureq::Agent::config_builder() - .timeout_global(Some(std::time::Duration::from_secs(60))) +struct Client { + agent: ureq::Agent, + authenticated: bool, + deadline: Option, +} + +impl std::ops::Deref for Client { + type Target = ureq::Agent; + fn deref(&self) -> &Self::Target { + &self.agent + } +} + +impl Client { + fn check_deadline(&self) -> Result<()> { + if self.deadline.is_some_and(|d| std::time::Instant::now() >= d) { + bail!("automatic upstream fetch time budget exhausted"); + } + Ok(()) + } +} + +fn agent(automatic: bool) -> Client { + let agent = ureq::Agent::config_builder() + .timeout_global(Some(std::time::Duration::from_secs(if automatic { 10 } else { 60 }))) + .max_redirects(if automatic { 0 } else { 10 }) .http_status_as_error(false) // Renamed repositories answer with a redirect; keep the token for it. .redirect_auth_headers(ureq::config::RedirectAuthHeaders::SameHost) .user_agent(concat!("lockdocs/", env!("CARGO_PKG_VERSION"), " (+https://github.com/SylphxAI/lockdocs)")) .build() - .into() + .into(); + Client { + agent, + authenticated: !automatic, + deadline: automatic.then(|| std::time::Instant::now() + std::time::Duration::from_secs(45)), + } } /// GitHub API GET; Ok(None) on 404. -fn api(agent: &ureq::Agent, path: &str) -> Result> { +fn api(agent: &Client, path: &str) -> Result> { + agent.check_deadline()?; let mut req = agent .get(&format!("https://api.github.com{path}")) .header("Accept", "application/vnd.github+json"); - if let Some(t) = token() { + if let Some(t) = agent.authenticated.then(token).flatten() { req = req.header("Authorization", &format!("Bearer {t}")); } let mut res = req.call()?; @@ -446,9 +484,9 @@ fn api(agent: &ureq::Agent, path: &str) -> Result> { let mut body = String::new(); res.body_mut().as_reader().take(64 << 20).read_to_string(&mut body)?; if status == 403 || status == 429 { - bail!("GitHub API rate limit (HTTP {status}); set GITHUB_TOKEN to raise it"); + bail!("GitHub API rate limit (HTTP {status}); automatic fetch is anonymous; explicit `lockdocs fetch` can use GITHUB_TOKEN"); } - if status >= 400 { + if status >= 300 { bail!("GitHub API HTTP {status} for {path}"); } Ok(Some(serde_json::from_str(&body)?)) @@ -498,7 +536,7 @@ struct Item { size: u64, } -fn tree(agent: &ureq::Agent, repo: &Repo, sha_or_ref: &str, recursive: bool) -> Result>> { +fn tree(agent: &Client, repo: &Repo, sha_or_ref: &str, recursive: bool) -> Result>> { let q = if recursive { "?recursive=1" } else { "" }; let Some(v) = api(agent, &format!("/repos/{}/{}/git/trees/{}{q}", repo.owner, repo.name, enc(sha_or_ref)))? else { return Ok(None); @@ -551,16 +589,65 @@ fn is_lang(s: &str) -> bool { /// Download the docs folders of `dep`'s repository at the pinned version's tag. pub fn fetch(dep: &Dep, src: &Source) -> Result { + fetch_with(dep, src, false) +} + +/// Anonymous, bounded first-use enrichment; never substitutes a major docs site. +pub fn fetch_automatic(dep: &Dep, src: &Source) -> Result { + fetch_with(dep, src, true) +} + +struct FetchGuard { + lock: PathBuf, + staging: PathBuf, +} +impl Drop for FetchGuard { + fn drop(&mut self) { + let _ = std::fs::remove_dir_all(&self.staging); + let _ = std::fs::remove_file(&self.lock); + } +} + +fn publish(out: &Path, target: &Path, m: &Manifest) -> Result<()> { + std::fs::write(out.join(".lockdocs-upstream.json"), serde_json::to_string(m)?)?; + if target.exists() { + std::fs::remove_dir_all(target)?; + } + std::fs::rename(out, target)?; + Ok(()) +} + +fn fetch_with(dep: &Dep, src: &Source, automatic: bool) -> Result { let repo = repo_of(dep, src).context("no GitHub repository in the package metadata")?; - let out = dir(dep); - let _ = std::fs::remove_dir_all(&out); + let target = dir(dep); + let parent = target.parent().context("upstream cache parent")?; + std::fs::create_dir_all(parent)?; + let stem = target.file_name().context("upstream cache name")?.to_string_lossy(); + let lock = target.with_file_name(format!("{stem}.fetch-lock")); + std::fs::OpenOptions::new() + .write(true) + .create_new(true) + .open(&lock) + .context("upstream fetch already in progress (or an interrupted fetch left its lock; use cache clean)")?; + let out = target.with_file_name(format!("{stem}.staging-{}", std::process::id())); + let _guard = FetchGuard { lock, staging: out.clone() }; std::fs::create_dir_all(&out)?; - let agent = agent(); + let agent = agent(automatic); let label = format!("github.com/{}/{}", repo.owner, repo.name); let mut tag = None; let mut root = None; + let mut commit = None; for t in tag_candidates(dep) { - if let Some(items) = tree(&agent, &repo, &t, false)? { + let Some(release) = api(&agent, &format!("/repos/{}/{}/commits/{}", repo.owner, repo.name, enc(&t)))? else { + continue; + }; + let sha = release + .get("sha") + .and_then(Value::as_str) + .filter(|s| s.len() == 40 && s.bytes().all(|b| b.is_ascii_hexdigit())) + .context("GitHub release did not resolve to an immutable commit")?; + if let Some(items) = tree(&agent, &repo, sha, false)? { + commit = Some(sha.to_string()); tag = Some(t); root = Some(items); break; @@ -571,13 +658,14 @@ pub fn fetch(dep: &Dep, src: &Source) -> Result { format: FORMAT, repo: label, tag: None, + commit: None, files: 0, bytes: 0, note: Some(format!("no git tag found for {}", dep.version)), site: None, pages: Vec::new(), }; - match site_files(&agent, dep, Some(&repo)) { + match if automatic { Ok(None) } else { site_files(&agent, dep, Some(&repo)) } { Ok(Some(site)) => { let (n, bytes, pages) = download_site(&agent, &site, &out)?; if n > 0 { @@ -591,7 +679,7 @@ pub fn fetch(dep: &Dep, src: &Source) -> Result { Ok(None) => {} Err(e) => m.note = Some(format!("{}; docs site skipped: {e:#}", m.note.unwrap_or_default())), } - std::fs::write(out.join(".lockdocs-upstream.json"), serde_json::to_string(&m)?)?; + publish(&out, &target, &m)?; return Ok(m); }; // Find docs roots: top-level docs dirs, and docs dirs one or two levels @@ -668,11 +756,14 @@ pub fn fetch(dep: &Dep, src: &Source) -> Result { files.dedup(); let mut total = 0u64; files.retain(|(_, s)| { + if *s > MAX_FILE || total.saturating_add(*s) > MAX_BYTES { + return false; + } total += s; - total <= MAX_BYTES + true }); files.truncate(MAX_FILES); - let ok = download(&agent, &repo, &tag, &files, &out)?; + let ok = download(&agent, &repo, commit.as_deref().context("missing release commit")?, &files, &out)?; let failed = files.len() - ok.len(); // A separate docs-site repository, when the package keeps its docs there. let mut site = None; @@ -680,7 +771,7 @@ pub fn fetch(dep: &Dep, src: &Source) -> Result { let mut site_bytes = 0u64; let mut pages = Vec::new(); let mut site_err = None; - match site_files(&agent, dep, Some(&repo)) { + match if automatic { Ok(None) } else { site_files(&agent, dep, Some(&repo)) } { Ok(Some(s)) => { (site_n, site_bytes, pages) = download_site(&agent, &s, &out)?; if site_n > 0 { @@ -694,6 +785,7 @@ pub fn fetch(dep: &Dep, src: &Source) -> Result { format: FORMAT, repo: label, tag: Some(tag), + commit, files: ok.len() + site_n, bytes: ok.iter().sum::() + site_bytes, note: if ok.is_empty() && site_n == 0 { @@ -706,12 +798,12 @@ pub fn fetch(dep: &Dep, src: &Source) -> Result { site, pages, }; - std::fs::write(out.join(".lockdocs-upstream.json"), serde_json::to_string(&m)?)?; + publish(&out, &target, &m)?; Ok(m) } /// Download a docs site's files into `//`: (files, bytes, page paths). -fn download_site(agent: &ureq::Agent, site: &SiteFiles, out: &Path) -> Result<(usize, u64, Vec)> { +fn download_site(agent: &Client, site: &SiteFiles, out: &Path) -> Result<(usize, u64, Vec)> { let dst = out.join(&site.repo.name); let docs = download(agent, &site.repo, &site.reference, &site.files, &dst)?; let got = download(agent, &site.repo, &site.reference, &site.pages, &dst)?; @@ -731,6 +823,16 @@ fn download_site(agent: &ureq::Agent, site: &SiteFiles, out: &Path) -> Result<(u #[cfg(test)] mod tests { use super::*; + #[test] + fn automatic_client_is_anonymous_and_bounded() { + let mut client = agent(true); + assert!(!client.authenticated); + assert!(client.deadline.is_some()); + client.deadline = Some(std::time::Instant::now()); + assert!(client.check_deadline().is_err()); + assert!(agent(false).authenticated); + } + #[test] fn github_urls() { let g = |s: &str| parse_github(s).map(|(o, n)| format!("{o}/{n}")); @@ -740,6 +842,9 @@ mod tests { assert_eq!(g("github:tokio-rs/axum").as_deref(), Some("tokio-rs/axum")); assert_eq!(g("expressjs/express").as_deref(), Some("expressjs/express")); assert_eq!(g("https://gitlab.com/x/y"), None); + assert_eq!(g("github:../repo"), None); + assert_eq!(g("github:owner/.."), None); + assert_eq!(g("github:owner/repo%2f.."), None); let d = Dep { eco: Eco::PyPI, name: "sqlalchemy".into(), diff --git a/crates/lockdocs-core/tests/cargo_git.rs b/crates/lockdocs-core/tests/cargo_git.rs index 526fae0..6f2dd28 100644 --- a/crates/lockdocs-core/tests/cargo_git.rs +++ b/crates/lockdocs-core/tests/cargo_git.rs @@ -11,7 +11,7 @@ fn finds_git_dependencies_in_cargo_checkouts() { std::env::set_var("LOCKDOCS_CACHE", &cache); std::env::set_var("CARGO_HOME", tests.join("fixture-git-home")); std::env::set_var("LOCKDOCS_NO_SYSTEM_PYTHON", "1"); - let e = Engine::new(&tests.join("fixture-git"), Options { fetch: false }); + let e = Engine::new(&tests.join("fixture-git"), Options { fetch: false, upstream: false }); let r = e.resolve(Some("tinygit")); assert!(r.text.contains("tinygit@0.2.0") && r.text.contains("docs ready"), "{}", r.text); let a = e.api("tinygit::clone_into", None, 800).unwrap(); diff --git a/crates/lockdocs-core/tests/integration.rs b/crates/lockdocs-core/tests/integration.rs index 32aa130..742a395 100644 --- a/crates/lockdocs-core/tests/integration.rs +++ b/crates/lockdocs-core/tests/integration.rs @@ -13,7 +13,7 @@ fn engine() -> Engine { std::env::set_var("LOCKDOCS_NO_SYSTEM_PYTHON", "1"); }); let root = PathBuf::from(env!("CARGO_MANIFEST_DIR")).join("tests").join("fixture"); - Engine::new(&root, Options { fetch: false }) + Engine::new(&root, Options { fetch: false, upstream: false }) } #[test] @@ -85,3 +85,18 @@ fn budget_is_respected() { a.text ); } + +#[test] +fn first_use_reports_unavailable_upstream_and_reuses_note() { + let e = engine(); + let root = e.root().to_path_buf(); + let e = Engine::new(&root, Options { fetch: false, upstream: true }); + for _ in 0..2 { + let a = e.docs(Some("tiny-schema"), "reject unknown keys", 1500).unwrap(); + assert!(a.text.contains("upstream docs unavailable"), "{}", a.text); + assert_eq!(a.json["provenance"][0]["requested_version"], "1.2.0"); + assert!(a.json["provenance"][0]["note"].as_str().unwrap().contains("no GitHub repository")); + } + let offline = engine().docs(Some("tiny-schema"), "reject unknown keys", 1500).unwrap(); + assert!(offline.json["provenance"][0]["note"].is_null()); +} diff --git a/crates/lockdocs-core/tests/pnp.rs b/crates/lockdocs-core/tests/pnp.rs index 259351d..8260088 100644 --- a/crates/lockdocs-core/tests/pnp.rs +++ b/crates/lockdocs-core/tests/pnp.rs @@ -10,7 +10,7 @@ fn reads_packages_from_the_yarn_zip_cache() { std::env::set_var("YARN_CACHE_FOLDER", cache.join("no-global-yarn-cache")); std::env::set_var("LOCKDOCS_NO_SYSTEM_PYTHON", "1"); let root = PathBuf::from(env!("CARGO_MANIFEST_DIR")).join("tests").join("fixture-pnp"); - let e = Engine::new(&root, Options { fetch: false }); + let e = Engine::new(&root, Options { fetch: false, upstream: false }); let r = e.resolve(Some("tiny-pnp")); assert!(r.text.contains("tiny-pnp@1.0.0") && r.text.contains("docs ready"), "{}", r.text); assert!(r.text.contains(".yarn/cache/tiny-pnp-npm-1.0.0-"), "{}", r.text); diff --git a/crates/lockdocs/src/main.rs b/crates/lockdocs/src/main.rs index 07b13e6..467b5b6 100644 --- a/crates/lockdocs/src/main.rs +++ b/crates/lockdocs/src/main.rs @@ -33,9 +33,10 @@ Options: -C, --root Project directory (default: current directory) --pkg Package for docs/api: zod, npm:zod, pydantic@2.9.2, tokio --tokens Answer budget in tokens (default 1200) - --fetch Allow downloads during queries: missing packages and upstream docs - (cached; also LOCKDOCS_FETCH=1). Without it, only `lockdocs fetch` - and the one-time model download use the network + --fetch Also download missing exact packages and major-version docs sites + (cached; also LOCKDOCS_FETCH=1) + --no-fetch Disable query downloads (also LOCKDOCS_FETCH=0) + Public release-tag docs download on first use by default --offline Never download (not even the embedding model) --json Machine-readable output @@ -100,8 +101,12 @@ impl Args { if self.on("fetch") { o.fetch = true; } - if self.on("offline") { + if self.on("fetch") { + o.upstream = true; + } + if self.on("no-fetch") || self.flag("fetch") == Some("false") || self.on("offline") || std::env::var("LOCKDOCS_OFFLINE").is_ok_and(|v| v != "0") { o.fetch = false; + o.upstream = false; } o } @@ -193,6 +198,9 @@ fn run() -> Result<()> { } "cache" => cache(&args), "fetch" | "pull" => { + if args.on("offline") || std::env::var("LOCKDOCS_OFFLINE").is_ok_and(|v| v != "0") { + bail!("fetch cannot download with --offline; use docs/api to read cached files"); + } let p = if args.positional.is_empty() { None } else { Some(args.positional.join(",")) }; print(engine().fetch_all(p.as_deref().or(args.pkg())).map_err(anyhow::Error::msg)?, json) } @@ -252,3 +260,16 @@ fn cache(args: &Args) -> Result<()> { ); Ok(()) } + +#[cfg(test)] +mod tests { + use super::*; + #[test] + fn offline_and_opt_out_override_fetch() { + for flag in ["--offline", "--no-fetch", "--fetch=false"] { + let args = Args::parse(vec!["--fetch".into(), flag.into()]); + let opts = args.opts(); + assert!(!opts.fetch && !opts.upstream); + } + } +} diff --git a/crates/lockdocs/src/mcp.rs b/crates/lockdocs/src/mcp.rs index 8a315d3..ba660ff 100644 --- a/crates/lockdocs/src/mcp.rs +++ b/crates/lockdocs/src/mcp.rs @@ -7,7 +7,7 @@ use mcp_kit::server::{run_stdio, App, Call, Info}; use serde_json::Value; use std::path::PathBuf; -const INSTRUCTIONS: &str = "lockdocs answers from the exact dependency versions pinned in this project's lockfiles, using the installed packages' own docs and type declarations, offline. Call `resolve` to see versions, `docs` with a question (and optionally `package`) before using an API you are unsure of in this version, and `api` for the exact signature of a symbol like `z.object` or `tokio::spawn`. Every section cites package@version path:line."; +const INSTRUCTIONS: &str = "lockdocs answers from the exact dependency versions pinned in this project's lockfiles, using the installed packages' own docs and type declarations, with anonymous release-tag docs fetched on first use unless offline. Cached docs work offline. Call `resolve` to see versions, `docs` with a question (and optionally `package`) before using an API you are unsure of in this version, and `api` for the exact signature of a symbol like `z.object` or `tokio::spawn`. Every section cites package@version path:line."; struct Lockdocs { ws: Workspace, diff --git a/crates/lockdocs/src/tools.rs b/crates/lockdocs/src/tools.rs index 064180c..087294e 100644 --- a/crates/lockdocs/src/tools.rs +++ b/crates/lockdocs/src/tools.rs @@ -8,6 +8,7 @@ pub fn definitions() -> Value { let root = json!({"type": "string", "description": "Absolute path of the project (defaults to the client's workspace root or the server's working directory)."}); let tokens = json!({"type": "integer", "description": "Answer budget in tokens (default 1200).", "minimum": 200, "maximum": 20000}); + let offline = json!({"type": "boolean", "description": "Use only installed and cached files for this call (no package or upstream downloads)."}); json!([ { "name": "resolve", @@ -22,14 +23,15 @@ pub fn definitions() -> Value { { "name": "docs", "title": "Search version-exact docs", - "description": "Answer a question from the docs of the exact installed version of a dependency: README, changelog, docs folder, and API reference (signatures + doc comments from .d.ts, Python, Rust and Go sources). Returns the most relevant sections within a token budget (hybrid keyword + embedding search), each cited as `package@version path:line`. Use it before using a library API you are not sure about for this version. Omit `package` to search all direct dependencies.", + "description": "Answer a question from the docs of the exact installed version of a dependency: README, changelog, docs folder (public release-tag docs fetched on first use unless offline), and API reference (signatures + doc comments from .d.ts, Python, Rust and Go sources). Returns the most relevant sections within a token budget (hybrid keyword + embedding search), each cited as `package@version path:line`. Use it before using a library API you are not sure about for this version. Omit `package` to search all direct dependencies.", "inputSchema": {"type": "object", "properties": { "query": {"type": "string", "description": "What you need, in words or identifiers, e.g. \"strict object unknown keys\" or \"cookies async\"."}, "package": {"type": "string", "description": "Package name, optionally ecosystem- or version-qualified: zod, npm:zod, pydantic, tokio, github.com/gin-gonic/gin. Comma-separate several."}, "tokens": tokens, - "root": root + "root": root, + "offline": offline }, "required": ["query"]}, - "annotations": {"readOnlyHint": true, "openWorldHint": false} + "annotations": {"readOnlyHint": true, "openWorldHint": true} }, { "name": "api", @@ -39,9 +41,10 @@ pub fn definitions() -> Value { "symbol": {"type": "string", "description": "Symbol path, optionally prefixed by the package name."}, "package": {"type": "string", "description": "Package to look in when the symbol has no package prefix."}, "tokens": tokens, - "root": root + "root": root, + "offline": offline }, "required": ["symbol"]}, - "annotations": {"readOnlyHint": true, "openWorldHint": false} + "annotations": {"readOnlyHint": true, "openWorldHint": true} } ]) } @@ -49,7 +52,12 @@ pub fn definitions() -> Value { pub fn call(ws: &Workspace, opts: &Options, name: &str, args: &Value, root: &Path) -> Result { let s = |k: &str| args.get(k).and_then(|v| v.as_str()).map(str::trim).filter(|v| !v.is_empty()); let tokens = args.get("tokens").and_then(|v| v.as_u64()).map_or(DEFAULT_TOKENS, |t| t as usize); - let engine = ws.engine(root, opts); + let mut opts = opts.clone(); + if args.get("offline").and_then(Value::as_bool) == Some(true) { + opts.fetch = false; + opts.upstream = false; + } + let engine = ws.engine(root, &opts); match name { "resolve" => Ok(engine.resolve(s("filter"))), "docs" => { diff --git a/docs/benchmarks.md b/docs/benchmarks.md index b20b151..1eb8ee7 100644 --- a/docs/benchmarks.md +++ b/docs/benchmarks.md @@ -6,7 +6,7 @@ - **Held-out questions:** 18 questions marked `held-out` are not used for tuning: they were written before the ranking changes they measure and have their own column. Held-out questions that are later used to diagnose a miss join the main set (their `history` field says so), and new ones replace them. - **Projects:** one project per version in [`bench/projects`](https://github.com/SylphxAI/lockdocs/tree/main/bench/projects), installed at those pins by [`bench/setup.sh`](https://github.com/SylphxAI/lockdocs/blob/main/bench/setup.sh). - **Grading:** an answer passes when it contains at least one string from every `expect` group (the version-correct API, in code or prose form) and none of the `reject` strings (the other version's API). Case-insensitive; the same grader for every tool. -- **lockdocs**, default settings (1,200-token budget), in three configurations: keyword only (`LOCKDOCS_EMBED=0`) on package files; hybrid on package files (the offline default once the model is downloaded); hybrid after `lockdocs fetch` added upstream docs at each version's tag. Index and fetch times are reported separately. +- **lockdocs**, default settings (1,200-token budget), in four configurations: keyword only (`LOCKDOCS_EMBED=0`) on package files; hybrid on package files (`LOCKDOCS_FETCH=0`, upstream disabled); real first-use defaults in an initially empty isolated `LOCKDOCS_CACHE` with fetch-policy and credential variables removed, before explicit prefetch; hybrid after explicit `lockdocs fetch` added upstream docs and major-version sites. The first-use row includes cold-download latency and anonymous rate-limit failures; it is never copied from the prefetched score. Full runs enforce at least 96/105 prefetched and 60/105 package-only hybrid. Index and fetch times are reported separately. - **Context7:** the anonymous API as its MCP server uses it: search the library, pick the top result and its listed version with the same major (exact when listed), then fetch context for the question. When that library answers HTTP 404, the next search result is tried, as an agent would. Rate-limit responses are recorded, not retried. - **Tokens:** tiktoken `o200k_base`. **Latency:** wall time per call from the same GitHub-hosted runner (for lockdocs: a fresh CLI process per question, including loading the model). - **Runner:** [`bench/run.py`](https://github.com/SylphxAI/lockdocs/blob/main/bench/run.py) via the [`bench` workflow](https://github.com/SylphxAI/lockdocs/actions/workflows/bench.yml). Reproduce: `bash bench/setup.sh && python3 bench/run.py target/release/lockdocs bench/projects out.json --fetch --context7`. @@ -15,6 +15,8 @@ +The historical results below predate automatic first-use enrichment; they are not a measurement of the new default. New first-use scores are emitted in the benchmark job summary and artifact. + From [this run](https://github.com/SylphxAI/lockdocs/actions/runs/36204404336). Tokenizer: tiktoken o200k_base. Runner: Linux x86_64. 105 questions. diff --git a/docs/capabilities.md b/docs/capabilities.md index d1036de..e842895 100644 --- a/docs/capabilities.md +++ b/docs/capabilities.md @@ -9,7 +9,7 @@ is in [vision.md](vision.md). | LD-LOCKFILE | Read exact versions from npm, pnpm, Yarn, Bun, Cargo, uv, Poetry, PDM, Pipenv, pip requirements and Go lockfiles | supported | crates/lockdocs-core/src/lockfile.rs, crates/lockdocs-core/src/project.rs | | | LD-LOCATE | Find each package's files on disk (node_modules, Yarn Plug'n'Play, virtualenvs, Cargo registry and git checkouts, Go module cache) | supported | crates/lockdocs-core/src/locate.rs | LD-LOCKFILE | | LD-REGISTRY-FETCH | Download a pinned package that is not installed, from its registry (opt-in) | supported | crates/lockdocs-core/src/fetch.rs | LD-LOCKFILE | -| LD-UPSTREAM | Download the docs folders of a package's GitHub repository at the pinned version's git tag (opt-in) | supported | crates/lockdocs-core/src/upstream.rs | LD-LOCATE | +| LD-UPSTREAM | Download the docs folders of a package's GitHub repository at the immutable commit resolved from the pinned release tag (automatic on first query; opt-out/offline supported) | supported | crates/lockdocs-core/src/upstream.rs | LD-LOCATE | | LD-DOCS-SITE | Add an official docs-site repository for the pinned major: the default branch for the latest major, a `vN` branch or the last commit before the next major for older ones (React, Express, Tailwind CSS, Prisma, tokio) | partial | crates/lockdocs-core/src/upstream.rs | LD-UPSTREAM | | LD-EXTRACT | Split Markdown, MDX, reStructuredText and component-based docs pages into sections; extract TypeScript, Python, Rust and Go symbols with signatures and doc comments | supported | crates/lockdocs-core/src/extract.rs, crates/lockdocs-core/src/markdown.rs | LD-LOCATE | | LD-INDEX | Cache one index per package, version and source location | supported | crates/lockdocs-core/src/index.rs, crates/lockdocs-core/src/cache.rs | LD-EXTRACT | diff --git a/docs/guide/fetch.md b/docs/guide/fetch.md index d095ffe..aa5680e 100644 --- a/docs/guide/fetch.md +++ b/docs/guide/fetch.md @@ -1,6 +1,20 @@ # Fetching and upstream docs -lockdocs answers from what is on your machine. Two optional downloads make answers better; both happen once and are cached. +lockdocs starts with installed package files and automatically adds public release-tag docs on first `docs` or `api` use. Downloads are cached under the supported OS cache root (`LOCKDOCS_CACHE` overrides it). + +## First-use defaults and controls + +The query's selected packages are enriched, not every dependency in the project. The existing upstream fetcher reads the repository from package metadata, resolves a candidate tag for exactly the requested version to a full immutable commit, then reads only that commit's docs. Automatic fetches are anonymous: ambient `GITHUB_TOKEN` and `GH_TOKEN` are not read. They do not use major-version docs sites or fall back to a default branch/latest version. + +- `--no-fetch` or `LOCKDOCS_FETCH=0`: no query package/upstream downloads; installed and cached docs still work. The model has its own policy. +- `--offline` or `LOCKDOCS_OFFLINE=1`: no downloads, including the embedding model. `fetch --offline` is rejected. +- MCP `docs` / `api` with `offline: true`: no package/upstream downloads for that call. Use server `--offline` to disable the background model download too. +- `LOCKDOCS_NO_UPSTREAM=1`: package-only answers, even if upstream docs are cached. +- `--fetch` or `LOCKDOCS_FETCH=1`: explicitly enable missing registry packages and major-version docs-site enrichment too. + +Automatic upstream work has a 45-second scheduling budget per package and a 10-second timeout per request (an in-flight request may finish after the scheduling deadline), no redirects, at most 2,500 files / 40 MB of advertised content and a 2 MB per-file limit, with 16 downloads in flight. Failed or missing enrichment is reported in text and structured `provenance`; answers retain the requested version's package files. Successful caches include repository, release tag, immutable commit, file counts and partial-download notes. Later queries reuse them without network. A failed fetch preserves the previous complete cache; concurrent fetches fail without waiting. An interrupted process can leave a fetch lock; `lockdocs cache clean` removes it. + +Network requests reveal the public repository/version/file being fetched, not the question or project contents. Anonymous GitHub limits can stop a cold multi-package session; failures are reported, not retried within an indexed MCP session. ## The embedding model (automatic, once) @@ -29,9 +43,9 @@ Some projects keep their docs in a separate website repository: React (react.dev - an older major: the site's `vN` or `N.x` branch when it has one (Tailwind CSS `v3`, Prisma `v6`), otherwise the last commit before the next major was released (the date of its `N+1.0.0` tag). React is excluded from the second rule because react.dev documents APIs before they ship; - pages about a later major (`v4-beta.mdx` on the v3 branch) are skipped. -The header says which: `github.com/prisma/docs@v6 (branch v6 for major 6)`. After upgrading lockdocs, run `lockdocs fetch` again to refresh copies made by an older version. Set `GITHUB_TOKEN` to raise GitHub's API limit (60 requests an hour without it; a package takes 2 to 20). +The header says which: `github.com/prisma/docs@v6 (branch v6 for major 6)`. After upgrading lockdocs, run `lockdocs fetch` again to refresh copies made by an older version. For explicit `fetch` / `--fetch` only, set `GITHUB_TOKEN` (or `GH_TOKEN`) to raise GitHub's API limit (60 requests an hour without it; a package takes 2 to 20). -To have it happen automatically on first query instead, enable fetching: `--fetch`, `LOCKDOCS_FETCH=1`, or register the MCP server with `lockdocs setup --fetch`. +Release-tag enrichment already happens on first query. Separate website repositories above describe a major, not an exact patch version, and remain explicit-fetch-only; their provenance is included in the answer header. ## Packages that are not installed diff --git a/docs/reference/cli.md b/docs/reference/cli.md index b2510eb..6822fb2 100644 --- a/docs/reference/cli.md +++ b/docs/reference/cli.md @@ -18,8 +18,11 @@ lockdocs version | `-C`, `--root ` | Project directory (default: current) | | `--pkg ` | Package for `docs` / `api` | | `--tokens ` | Answer budget (default 1200) | -| `--fetch` | Allow downloads during queries (missing packages, upstream docs) | +| `--fetch` | Also allow exact registry packages and major-version docs sites during queries | +| `--no-fetch` | Disable query package/docs downloads; cached docs remain usable | | `--offline` | Never download, not even the embedding model | | `--json` | Machine-readable output | Exit code 1 with a message on stderr when a package is not a dependency or not installed. + +Public release-tag docs are fetched anonymously on first `docs` / `api` use by default. `LOCKDOCS_FETCH=0` opts out; `LOCKDOCS_OFFLINE=1` also disables the embedding model download. diff --git a/docs/reference/tools.md b/docs/reference/tools.md index 7ad6195..66bf47a 100644 --- a/docs/reference/tools.md +++ b/docs/reference/tools.md @@ -1,6 +1,9 @@ # MCP tools -All tools are read-only and take an optional `root` (absolute project path). +All tools take an optional `root` (absolute project path). `resolve` stays local. +`docs` and `api` may populate the cache with anonymous public release-tag docs +on first use; `offline: true` restricts that call to installed/cached files. +Start the server with `--offline` to prevent all downloads, including the model. ## resolve @@ -20,6 +23,7 @@ Answer a question from the installed version's docs and API reference. |---|---|---| | `query` | string, required | Words or identifiers | | `package` | string | `zod`, `npm:zod`, `pydantic@2.9.2`, `tokio`, `github.com/gin-gonic/gin`; comma-separate several. Omit to search direct dependencies (a package named in the query is picked automatically) | +| `offline` | boolean | No package/upstream network access for this call | | `tokens` | integer | Budget, default 1200 (200-20000) | ## api @@ -30,8 +34,11 @@ Exact signature and doc comment of one symbol. |---|---|---| | `symbol` | string, required | `zod.z.object`, `z.object`, `tokio::spawn`, `axum::Router::route`, `pydantic.BaseModel.model_dump`, `gin.Context.JSON` | | `package` | string | When the symbol has no package prefix | +| `offline` | boolean | No package/upstream network access for this call | | `tokens` | integer | Budget, default 1200 | Returns the best match (following re-exports), other declarations in the same file (overloads), members for classes, interfaces, structs and traits, and other matches. If no symbol has that name, it falls back to `docs`. Add `"format": "json"` to any call for structured output. + +Structured docs/api answers include `provenance`: the requested version, actual package source, registry-fetch flag, upstream repository/tag/commit label, and any fetch failure note. Text answers include the same source and failure notes. Failed enrichment falls back only to the requested package version, never latest. diff --git a/docs/vision.md b/docs/vision.md index c02626b..53950de 100644 --- a/docs/vision.md +++ b/docs/vision.md @@ -23,7 +23,9 @@ newest one. That is the one job. - Local first: it runs on the developer's machine, needs no account or API key, and works offline after the one-time downloads. -- Network use is opt-in and limited to public sources: package registries, +- First-use release-tag docs are anonymous and automatic, with an explicit + opt-out and an offline mode. Registry downloads and major-version docs sites + remain opt-in. Network use is limited to public sources: package registries, GitHub (docs folders at a tag, and official docs-site repositories), and the embedding model on Hugging Face. - Three MCP tools only: `resolve`, `docs`, `api`. New abilities go into those From 9cd0e47aff0973620ae94756eb5c1ca64959498b Mon Sep 17 00:00:00 2001 From: Sylphx Builder Date: Thu, 1 Oct 2026 00:52:36 +0000 Subject: [PATCH 2/5] fix: expose complete pinned-fetch provenance and registry errors --- README.md | 2 +- crates/lockdocs-core/src/index.rs | 5 ++++- crates/lockdocs-core/src/lib.rs | 3 ++- crates/lockdocs-core/src/query.rs | 8 ++------ docs/guide/fetch.md | 4 ++-- docs/reference/tools.md | 2 +- 6 files changed, 12 insertions(+), 12 deletions(-) diff --git a/README.md b/README.md index e6b7e5d..45486a8 100644 --- a/README.md +++ b/README.md @@ -70,7 +70,7 @@ lockdocs takes the version question off the table: - **Exact version, zero config.** It reads `package-lock.json`, `pnpm-lock.yaml`, `yarn.lock`, `bun.lock`, `Cargo.lock`, `uv.lock`, `poetry.lock`, `Pipfile.lock`, `requirements*.txt` and `go.mod`. No library IDs, no "use v14" in the prompt. - **Docs that ship with the code.** READMEs, changelogs and `docs/` folders, plus the API reference in the package itself: `.d.ts` declarations with JSDoc, Python docstrings and stubs, rustdoc comments, Go doc comments. If it is installed, it is documented, including your private and internal packages. -- **Offline and unlimited after caching.** Package files are read from `node_modules`, your virtualenv, `~/.cargo/registry` and the Go module cache. Offline after a one-time model download (129 MB, kept as 32 MB); Use `--offline` (or `LOCKDOCS_OFFLINE=1`) to prohibit all downloads; `LOCKDOCS_EMBED=0` disables only the model. No account or hosted query quota. First-use upstream requests reveal the public repository and version being fetched, not your question or project files. +- **Offline and unlimited after caching.** Package files are read from `node_modules`, your virtualenv, `~/.cargo/registry` and the Go module cache. The optional model downloads once (129 MB, kept as 32 MB). Use `--offline` (or `LOCKDOCS_OFFLINE=1`) to prohibit all downloads; `LOCKDOCS_EMBED=0` disables only the model. No account or hosted query quota. First-use upstream requests reveal the public repository and version being fetched, not your question or project files. - **Upstream docs on first use.** Packages like Next.js, Django and FastAPI ship no docs. Queries automatically add public GitHub docs at the immutable commit resolved from your pinned release tag, anonymously and with bounded downloads. `--no-fetch` or `LOCKDOCS_FETCH=0` opts out; cached docs still work offline. Failures are reported alongside local answers, never replaced with latest-version docs. - **Meaning, not just words.** Hybrid retrieval: BM25 fused with a small local embedding model (downloaded once, 32 MB on disk), plus API redirects from deprecation notes ("use `model_validate` instead"). - **Small, cited answers.** Packed into a token budget (1,200 by default), every section cited as `package@version path:line`. diff --git a/crates/lockdocs-core/src/index.rs b/crates/lockdocs-core/src/index.rs index e837d43..654bd2a 100644 --- a/crates/lockdocs-core/src/index.rs +++ b/crates/lockdocs-core/src/index.rs @@ -12,7 +12,7 @@ use std::path::{Path, PathBuf}; use std::time::Instant; /// Bump when extraction or the on-disk format changes. -pub const FORMAT: u32 = 9; +pub const FORMAT: u32 = 10; #[derive(Serialize, Deserialize)] pub struct PackageIndex { @@ -33,6 +33,8 @@ pub struct PackageIndex { pub build_ms: u64, /// `github.com/o/r@tag (N files)` when upstream docs are included. pub upstream: Option, + /// Fetch manifest used by this index, including immutable commit and failure notes. + pub upstream_provenance: Option, /// Embedding model id, or empty for keyword-only. pub embed: String, /// One embedding per entry (empty without a model). @@ -222,6 +224,7 @@ pub fn build(dep: &Dep, src: &Source, root: &Path, up: Option<&(PathBuf, Manifes m.files ) }), + upstream_provenance: up.map(|(_, m)| m.clone()), embed: if model.is_some() { embed::MODEL_ID.to_string() } else { String::new() }, vecs, } diff --git a/crates/lockdocs-core/src/lib.rs b/crates/lockdocs-core/src/lib.rs index 70e3883..a3f3e44 100644 --- a/crates/lockdocs-core/src/lib.rs +++ b/crates/lockdocs-core/src/lib.rs @@ -1,6 +1,7 @@ //! lockdocs core: read lockfiles, find each dependency's installed sources, //! extract API reference and prose, and answer queries with BM25 under a -//! token budget. Everything is local unless fetching is explicitly enabled. +//! token budget. Public release-tag docs enrich first-use queries by default; +//! opt-out and offline controls keep installed/cached answers local. pub mod bm25; pub mod cache; diff --git a/crates/lockdocs-core/src/query.rs b/crates/lockdocs-core/src/query.rs index c2ed1c2..c352dc3 100644 --- a/crates/lockdocs-core/src/query.rs +++ b/crates/lockdocs-core/src/query.rs @@ -75,7 +75,7 @@ fn provenance(ready: &[ReadyPkg]) -> Vec { .map(|(idx, dep, note)| { json!({ "package": idx.id(), "requested_version": dep.version, "source": idx.source, - "registry_fetched": idx.fetched, "upstream": idx.upstream, "note": note, + "registry_fetched": idx.fetched, "upstream": idx.upstream_provenance, "upstream_label": idx.upstream, "note": note, }) }) .collect() @@ -176,11 +176,7 @@ impl Engine { if self.opts.fetch { match fetch::fetch(dep) { Ok(s) => return Ok((s, None)), - Err(e) => { - if local.is_none() { - return Err(format!("{} is pinned but not installed, and fetching failed: {e:#}", dep.id())); - } - } + Err(e) => return Err(format!("{}: exact registry fetching failed: {e:#}; no other version substituted", dep.id())), } } if let Some(s) = local { diff --git a/docs/guide/fetch.md b/docs/guide/fetch.md index aa5680e..91b2d3e 100644 --- a/docs/guide/fetch.md +++ b/docs/guide/fetch.md @@ -4,7 +4,7 @@ lockdocs starts with installed package files and automatically adds public relea ## First-use defaults and controls -The query's selected packages are enriched, not every dependency in the project. The existing upstream fetcher reads the repository from package metadata, resolves a candidate tag for exactly the requested version to a full immutable commit, then reads only that commit's docs. Automatic fetches are anonymous: ambient `GITHUB_TOKEN` and `GH_TOKEN` are not read. They do not use major-version docs sites or fall back to a default branch/latest version. +The query's selected packages are enriched, not every dependency in the project. The existing upstream fetcher reads the repository from package metadata, resolves a candidate tag for exactly the requested version to a full immutable commit, then reads only that commit's docs. Automatic fetches are anonymous: ambient `GITHUB_TOKEN` and `GH_TOKEN` are not read. Git dependencies use their checkout files rather than guessing a release tag for their resolved commit. Automatic fetches do not use major-version docs sites or fall back to a default branch/latest version. - `--no-fetch` or `LOCKDOCS_FETCH=0`: no query package/upstream downloads; installed and cached docs still work. The model has its own policy. - `--offline` or `LOCKDOCS_OFFLINE=1`: no downloads, including the embedding model. `fetch --offline` is rejected. @@ -12,7 +12,7 @@ The query's selected packages are enriched, not every dependency in the project. - `LOCKDOCS_NO_UPSTREAM=1`: package-only answers, even if upstream docs are cached. - `--fetch` or `LOCKDOCS_FETCH=1`: explicitly enable missing registry packages and major-version docs-site enrichment too. -Automatic upstream work has a 45-second scheduling budget per package and a 10-second timeout per request (an in-flight request may finish after the scheduling deadline), no redirects, at most 2,500 files / 40 MB of advertised content and a 2 MB per-file limit, with 16 downloads in flight. Failed or missing enrichment is reported in text and structured `provenance`; answers retain the requested version's package files. Successful caches include repository, release tag, immutable commit, file counts and partial-download notes. Later queries reuse them without network. A failed fetch preserves the previous complete cache; concurrent fetches fail without waiting. An interrupted process can leave a fetch lock; `lockdocs cache clean` removes it. +Automatic upstream work has a 45-second scheduling budget per package and a 10-second timeout per request (an in-flight request may finish after the scheduling deadline), no redirects, at most 2,500 files / 40 MB of advertised content and a 2 MB per-file limit, with 16 downloads in flight. Failed or missing enrichment is reported in text and structured `provenance`; answers retain the requested version's package files. Successful caches include repository, release tag, immutable commit, file counts and partial-download notes. Later queries reuse them without network. A failed download preserves the previous complete cache; concurrent fetches fail without waiting. An interrupted process can leave a fetch lock; `lockdocs cache clean` removes it. Network requests reveal the public repository/version/file being fetched, not the question or project contents. Anonymous GitHub limits can stop a cold multi-package session; failures are reported, not retried within an indexed MCP session. diff --git a/docs/reference/tools.md b/docs/reference/tools.md index 66bf47a..bf7326b 100644 --- a/docs/reference/tools.md +++ b/docs/reference/tools.md @@ -41,4 +41,4 @@ Returns the best match (following re-exports), other declarations in the same fi Add `"format": "json"` to any call for structured output. -Structured docs/api answers include `provenance`: the requested version, actual package source, registry-fetch flag, upstream repository/tag/commit label, and any fetch failure note. Text answers include the same source and failure notes. Failed enrichment falls back only to the requested package version, never latest. +Structured docs/api answers include `provenance`: the requested version, actual package source, registry-fetch flag, the upstream manifest (repository, tag, immutable commit, counts and notes), its display label, and any fetch failure note. Text answers include the same source and failure notes. Failed enrichment falls back only to the requested package version, never latest. From 443f0ae5f63d84821d0dc3c3f14b8e3951fd587e Mon Sep 17 00:00:00 2001 From: Sylphx Builder Date: Thu, 1 Oct 2026 00:53:49 +0000 Subject: [PATCH 3/5] test: cover pinned upstream cache readback online and offline --- crates/lockdocs-core/tests/integration.rs | 37 +++++++++++++++++++++++ 1 file changed, 37 insertions(+) diff --git a/crates/lockdocs-core/tests/integration.rs b/crates/lockdocs-core/tests/integration.rs index 742a395..ff35e27 100644 --- a/crates/lockdocs-core/tests/integration.rs +++ b/crates/lockdocs-core/tests/integration.rs @@ -100,3 +100,40 @@ fn first_use_reports_unavailable_upstream_and_reuses_note() { let offline = engine().docs(Some("tiny-schema"), "reject unknown keys", 1500).unwrap(); assert!(offline.json["provenance"][0]["note"].is_null()); } + +#[test] +fn pinned_upstream_cache_is_reused_online_and_offline_with_provenance() { + let _ = engine(); // initialize the isolated test cache + let root = std::env::temp_dir().join(format!("lockdocs-cached-provenance-{}", std::process::id())); + let dep = lockdocs_core::Dep { + eco: lockdocs_core::Eco::Npm, + name: "cached-provenance-fixture".into(), + version: "1.2.3".into(), + direct: true, + from: "package-lock.json".into(), + }; + let src = root.join("node_modules").join(&dep.name); + std::fs::create_dir_all(&src).unwrap(); + // No repository metadata: any mistaken refresh would fail instead of contacting GitHub. + std::fs::write(src.join("package.json"), r#"{"name":"cached-provenance-fixture","version":"1.2.3"}"#).unwrap(); + let cache = lockdocs_core::upstream::dir(&dep); + std::fs::create_dir_all(cache.join("docs")).unwrap(); + std::fs::write(cache.join("docs/guide.md"), "# Pinned guide\n\nUse release_only_api to reject unknown keys.\n").unwrap(); + let manifest = serde_json::json!({ + "format": lockdocs_core::upstream::FORMAT, + "repo": "github.com/example/pinned", "tag": "v1.2.3", + "commit": "0123456789abcdef0123456789abcdef01234567", + "files": 1, "bytes": 73, "note": null, "site": null, "pages": [], + }); + std::fs::write(cache.join(".lockdocs-upstream.json"), manifest.to_string()).unwrap(); + for upstream in [true, false] { + let mut e = Engine::new(&root, Options { fetch: false, upstream }); + e.project.deps = vec![dep.clone()]; + let a = e.docs(Some(&dep.name), "reject unknown keys", 1500).unwrap(); + assert!(a.text.contains("release_only_api") && a.text.contains("upstream:docs/guide.md"), "{}", a.text); + assert_eq!(a.json["provenance"][0]["upstream"]["commit"], manifest["commit"]); + assert_eq!(a.json["provenance"][0]["requested_version"], "1.2.3"); + } + let _ = std::fs::remove_dir_all(&root); + let _ = std::fs::remove_dir_all(&cache); +} From b7cddf407e8fa9320999b123e4dd7e909018d82f Mon Sep 17 00:00:00 2001 From: Sylphx Builder Date: Thu, 1 Oct 2026 01:18:00 +0000 Subject: [PATCH 4/5] fix: enforce exact tag provenance and preserve complete enrichment caches --- README.md | 2 +- crates/lockdocs-core/src/fetch.rs | 22 +- crates/lockdocs-core/src/index.rs | 9 +- crates/lockdocs-core/src/lockfile.rs | 79 ++++++- crates/lockdocs-core/src/query.rs | 98 +++++--- crates/lockdocs-core/src/upstream.rs | 260 ++++++++++++++++++++-- crates/lockdocs-core/tests/cargo_git.rs | 14 +- crates/lockdocs-core/tests/integration.rs | 23 +- docs/guide/fetch.md | 6 +- 9 files changed, 446 insertions(+), 67 deletions(-) diff --git a/README.md b/README.md index 45486a8..85708a4 100644 --- a/README.md +++ b/README.md @@ -113,7 +113,7 @@ Legacy copies bundled inside a package (`zod/v3` inside zod 4, `pydantic/v1` ins ### Upstream docs on first use -The first `docs` or `api` query adds the selected dependencies' public release-tag docs once: it finds the GitHub repository in the package's own metadata and the git tag of your pinned version, resolves that tag to an immutable commit, and downloads only its docs folders (Markdown, MDX, reStructuredText, docs examples). It does not use ambient GitHub credentials or substitute major-version website docs. `lockdocs fetch` remains available to prewarm docs and explicitly add major-version docs sites. Answers then cite `next@15.1.0 upstream:docs/01-app/.../cookies.mdx:12`. See [Upstream docs and fetching](https://sylphxai.github.io/lockdocs/guide/fetch). +The first `docs` or `api` query adds the selected dependencies' public release-tag docs once: it finds the GitHub repository in the package's own metadata and the git tag of your pinned version, resolves that tag to an immutable commit, and downloads only its docs folders (Markdown, MDX, reStructuredText, docs examples). It does not use ambient GitHub credentials or download major-version website docs by default. Previously opted-in docs-site caches are preserved and their provenance stays visible. `lockdocs fetch` remains available to prewarm docs and explicitly add major-version docs sites. Answers then cite `next@15.1.0 upstream:docs/01-app/.../cookies.mdx:12`. See [Upstream docs and fetching](https://sylphxai.github.io/lockdocs/guide/fetch). ### Not installed? Fetch the exact version (opt-in) diff --git a/crates/lockdocs-core/src/fetch.rs b/crates/lockdocs-core/src/fetch.rs index db63ef3..75d6aa0 100644 --- a/crates/lockdocs-core/src/fetch.rs +++ b/crates/lockdocs-core/src/fetch.rs @@ -41,6 +41,9 @@ pub fn fetched_dir(dep: &Dep) -> PathBuf { /// A previously fetched copy, if any (never touches the network). pub fn cached(dep: &Dep) -> Option { + if is_git(dep) { + return None; + } let dir = fetched_dir(dep); if dir.join(".lockdocs-complete").is_file() { return Some(source_for(dep, &dir)); @@ -78,13 +81,26 @@ fn source_for(dep: &Dep, dir: &Path) -> Source { } /// Download and unpack `dep` if it is not cached yet. +pub fn is_git(dep: &Dep) -> bool { + dep.from.ends_with("(git)") +} + +/// Shared guard: a git checkout is not interchangeable with a registry release tag. +pub fn require_registry_origin(dep: &Dep) -> Result<()> { + if is_git(dep) { + bail!( + "{} is a git dependency; use its resolved checkout files, not registry/release-tag docs", + dep.id() + ); + } + Ok(()) +} + pub fn fetch(dep: &Dep) -> Result { + require_registry_origin(dep)?; if let Some(s) = cached(dep) { return Ok(s); } - if dep.from.ends_with("(git)") { - bail!("{} is a git dependency; run `cargo fetch` to check it out", dep.id()); - } let dir = fetched_dir(dep); let _ = std::fs::remove_dir_all(&dir); std::fs::create_dir_all(&dir)?; diff --git a/crates/lockdocs-core/src/index.rs b/crates/lockdocs-core/src/index.rs index 654bd2a..3609c29 100644 --- a/crates/lockdocs-core/src/index.rs +++ b/crates/lockdocs-core/src/index.rs @@ -12,7 +12,7 @@ use std::path::{Path, PathBuf}; use std::time::Instant; /// Bump when extraction or the on-disk format changes. -pub const FORMAT: u32 = 10; +pub const FORMAT: u32 = 11; #[derive(Serialize, Deserialize)] pub struct PackageIndex { @@ -157,7 +157,12 @@ pub fn embed_text(e: &Entry) -> String { fn cache_path(dep: &Dep, src: &Source, types: Option<&Path>, up: Option<&(PathBuf, Manifest)>, embed: &str) -> PathBuf { let up_key = up - .map(|(_, m)| format!("{}@{:?}:{}:{:?}:{:?}", m.repo, m.tag, m.files, m.site, m.commit)) + .map(|(_, m)| { + format!( + "{}@{:?}:{}:{:?}:{:?}:{:?}:{}", + m.repo, m.tag, m.files, m.site, m.commit, m.note, m.docs_sites_checked + ) + }) .unwrap_or_default(); let key = cache::hash(&[ embed, diff --git a/crates/lockdocs-core/src/lockfile.rs b/crates/lockdocs-core/src/lockfile.rs index d9a4946..0ddf716 100644 --- a/crates/lockdocs-core/src/lockfile.rs +++ b/crates/lockdocs-core/src/lockfile.rs @@ -21,6 +21,16 @@ pub const LOCKFILES: &[(&str, Eco)] = &[ ("go.mod", Eco::Go), ]; +fn git_url(s: &str) -> bool { + s.contains("git+") + || s.contains("git://") + || s.contains("git@") + || s.contains("github:") + || s.contains(".git#") + || s.contains(".git?") + || s.ends_with(".git") +} + fn dep(eco: Eco, name: &str, version: &str, direct: bool, from: &str) -> Dep { Dep { eco, @@ -95,13 +105,23 @@ pub fn npm_lock(text: &str, from: &str) -> anyhow::Result> { } let Some(ver) = p.get("version").and_then(|v| v.as_str()) else { continue }; let name = p.get("name").and_then(|n| n.as_str()).unwrap_or(name); - out.push(dep(Eco::Npm, name, ver, direct.contains(name), from)); + let origin = if p.get("resolved").and_then(Value::as_str).is_some_and(git_url) || git_url(ver) { + format!("{from} (git)") + } else { + from.to_string() + }; + out.push(dep(Eco::Npm, name, ver, direct.contains(name), &origin)); } } else if let Some(deps) = v.get("dependencies").and_then(|d| d.as_object()) { // lockfileVersion 1 for (name, p) in deps { if let Some(ver) = p.get("version").and_then(|v| v.as_str()) { - out.push(dep(Eco::Npm, name, ver, true, from)); + let origin = if p.get("resolved").and_then(Value::as_str).is_some_and(git_url) || git_url(ver) { + format!("{from} (git)") + } else { + from.to_string() + }; + out.push(dep(Eco::Npm, name, ver, true, &origin)); } } // v1 has no reliable direct marker: every top-level entry counts. @@ -260,6 +280,7 @@ pub fn yarn_lock(text: &str, direct: &HashSet) -> Vec { let mut out = Vec::new(); let mut seen = HashSet::new(); let mut names: Vec = Vec::new(); + let mut git_origin = false; for line in text.lines() { if line.is_empty() || line.starts_with('#') { continue; @@ -270,6 +291,7 @@ pub fn yarn_lock(text: &str, direct: &HashSet) -> Vec { continue; } let header = line.trim_end_matches(':'); + git_origin = git_url(header); for spec in header.split(", ") { let spec = spec.trim().trim_matches('"'); if let Some((name, _)) = split_at_version(spec) { @@ -286,7 +308,13 @@ pub fn yarn_lock(text: &str, direct: &HashSet) -> Vec { let ver = ver.trim().trim_matches('"'); for n in &names { if seen.insert(format!("{n}@{ver}")) { - out.push(dep(Eco::Npm, n, ver, direct.contains(n), "yarn.lock")); + out.push(dep( + Eco::Npm, + n, + ver, + direct.contains(n), + if git_origin { "yarn.lock (git)" } else { "yarn.lock" }, + )); } } names.clear(); @@ -417,7 +445,12 @@ pub fn uv_lock(text: &str) -> anyhow::Result> { if local.contains(name) { continue; } - out.push(dep(Eco::PyPI, name, ver, direct.contains(&norm_name(Eco::PyPI, name)), "uv.lock")); + let from = if p.get("source").is_some_and(|s| s.get("git").is_some()) { + "uv.lock (git)" + } else { + "uv.lock" + }; + out.push(dep(Eco::PyPI, name, ver, direct.contains(&norm_name(Eco::PyPI, name)), from)); } Ok(out) } @@ -430,7 +463,15 @@ pub fn poetry_lock(text: &str, from: &str, direct: &HashSet) -> anyhow:: continue; }; let is_direct = direct.is_empty() || direct.contains(&norm_name(Eco::PyPI, name)); - out.push(dep(Eco::PyPI, name, ver, is_direct, from)); + let origin = if p + .get("source") + .is_some_and(|s| s.get("git").is_some() || s.get("type").and_then(toml::Value::as_str) == Some("git")) + { + format!("{from} (git)") + } else { + from.to_string() + }; + out.push(dep(Eco::PyPI, name, ver, is_direct, &origin)); } Ok(out) } @@ -441,7 +482,13 @@ pub fn pipfile_lock(text: &str) -> anyhow::Result> { for sect in ["default", "develop"] { for (name, p) in v.get(sect).and_then(|s| s.as_object()).into_iter().flatten() { if let Some(ver) = p.get("version").and_then(|v| v.as_str()) { - out.push(dep(Eco::PyPI, name, ver.trim_start_matches("=="), true, "Pipfile.lock")); + out.push(dep( + Eco::PyPI, + name, + ver.trim_start_matches("=="), + true, + if p.get("git").is_some() { "Pipfile.lock (git)" } else { "Pipfile.lock" }, + )); } } } @@ -537,6 +584,26 @@ fn parse_replace(line: &str, out: &mut BTreeMap) { mod tests { use super::*; + #[test] + fn python_and_npm_git_origins_are_not_registry_releases() { + let uv = r#"[[package]] +name = "git-lib" +version = "1.2.3" +source = { git = "https://github.com/o/r?rev=abc#01234567" } +"#; + assert_eq!(uv_lock(uv).unwrap()[0].from, "uv.lock (git)"); + let poetry = r#"[[package]] +name = "git-lib" +version = "1.2.3" +[package.source] +type = "git" +url = "https://github.com/o/r" +"#; + assert_eq!(poetry_lock(poetry, "poetry.lock", &HashSet::new()).unwrap()[0].from, "poetry.lock (git)"); + let npm = r#"{"packages":{"node_modules/git-lib":{"version":"1.2.3","resolved":"git+https://github.com/o/r.git#abc"}}}"#; + assert_eq!(npm_lock(npm, "package-lock.json").unwrap()[0].from, "package-lock.json (git)"); + } + #[test] fn npm_v3() { let t = r#"{"lockfileVersion":3,"packages":{"":{"dependencies":{"zod":"^3.23.0"},"devDependencies":{"@types/node":"^20"}}, diff --git a/crates/lockdocs-core/src/query.rs b/crates/lockdocs-core/src/query.rs index c352dc3..ca88e89 100644 --- a/crates/lockdocs-core/src/query.rs +++ b/crates/lockdocs-core/src/query.rs @@ -81,6 +81,50 @@ fn provenance(ready: &[ReadyPkg]) -> Vec { .collect() } +/// Keep network fallbacks visible on ordinary broad MCP answers without listing +/// every dependency's long header or crowding out the actual answer. +fn broad_header(ready: &[ReadyPkg], missing: &[(Dep, String)], tokens: usize) -> String { + let fetched = ready.iter().filter(|r| r.0.fetched).count(); + let upstream = ready.iter().filter(|r| r.0.upstream.is_some()).count(); + let notes = ready.iter().filter(|r| r.2.is_some()).count(); + let mut header = format!("Searched {} direct dependencies: {} local package sources, {fetched} registry sources, {upstream} upstream sources; {notes} fallback notes, {} unavailable.\n", ready.len(), ready.len() - fetched, missing.len()); + let max_chars = (tokens / 3).clamp(100, 400) * 3; + for (idx, _, note) in ready { + if let Some(note) = note { + let row = format!("Fallback {}: {}\n", idx.id(), truncate(note, 160)); + if header.chars().count() + row.chars().count() > max_chars { + break; + } + header.push_str(&row); + } + } + for (dep, error) in missing { + let row = format!("Skipped {}: {}\n", dep.id(), truncate(error, 140)); + if header.chars().count() + row.chars().count() > max_chars { + break; + } + header.push_str(&row); + } + for (idx, _, _) in ready { + if let Some(m) = &idx.upstream_provenance { + let row = format!( + "Upstream {}: {}@{} commit {}{}\n", + idx.id(), + m.repo, + m.tag.as_deref().unwrap_or("none"), + m.commit.as_deref().map(|s| &s[..s.len().min(7)]).unwrap_or("none"), + if m.site.is_some() { " + opted-in docs site" } else { "" } + ); + if header.chars().count() + row.chars().count() > max_chars { + break; + } + header.push_str(&row); + } + } + header.push_str("Focus `package` or request JSON for full provenance.\n"); + header +} + fn lang(eco: Eco) -> &'static str { match eco { Eco::Npm => "ts", @@ -224,6 +268,8 @@ impl Engine { Some(old) => format!("{old} {n}"), None => n, }); + } else { + self.upstream_notes.lock().unwrap().remove(&dep.id()); } let idx = Arc::new(index::load_or_build(dep, &src, &self.project.root, up.as_ref())); self.indexes.lock().unwrap().insert(key, idx.clone()); @@ -240,26 +286,17 @@ impl Engine { return (None, None); } let enabled = self.opts.fetch || self.opts.upstream; - if dep.from.ends_with("(git)") { + if fetch::is_git(dep) { return ( None, enabled.then(|| "upstream docs skipped: git dependency's resolved commit is not a release tag; using its checkout files".into()), ); } let cached = upstream::cached(dep); - if let Some(c) = cached - .as_ref() - .filter(|(_, m)| !enabled || (!upstream::stale(m) && (self.opts.fetch || (m.site.is_none() && (m.commit.is_some() || m.files == 0))))) - { + if let Some(c) = cached.as_ref().filter(|(_, m)| !enabled || !upstream::needs_refresh(m, self.opts.fetch)) { return (Some(c.clone()), c.1.note.clone()); } if enabled { - if upstream::repo_of(dep, src).is_none() { - return ( - None, - Some("upstream docs unavailable: no GitHub repository in package metadata; using package files".into()), - ); - } let result = if self.opts.fetch { upstream::fetch(dep, src) } else { @@ -301,12 +338,20 @@ impl Engine { let rows: Vec<(Dep, String, Value)> = deps .par_iter() .map(|d| { + if fetch::is_git(d) { + let src = locate::locate(d, &self.project.root); + return ( + d.clone(), + "upstream skipped: git dependency; use resolved checkout files".into(), + json!({"package": d.id(), "status": "git-checkout", "source": src.map(|s| s.label)}), + ); + } let src = match locate::locate(d, &self.project.root) .filter(|s| s.version == d.version) .or_else(|| fetch::cached(d)) { Some(s) => Some(s), - None if !d.from.ends_with("(git)") => fetch::fetch(d).ok(), + None if !fetch::is_git(d) => fetch::fetch(d).ok(), None => None, }; let Some(src) = src else { @@ -316,18 +361,15 @@ impl Engine { json!({"package": d.id(), "status": "missing"}), ); }; - let status = match upstream::cached(d).filter(|(_, m)| !upstream::stale(m)) { - Some((_, m)) => m, - None => match upstream::fetch(d, &src) { - Ok(m) => m, - Err(e) => { - return ( - d.clone(), - format!("upstream docs: {e:#}"), - json!({"package": d.id(), "status": "no-upstream", "error": format!("{e:#}")}), - ); - } - }, + let status = match upstream::fetch(d, &src) { + Ok(m) => m, + Err(e) => { + return ( + d.clone(), + format!("upstream docs fetch failed: {e:#}"), + json!({"package": d.id(), "status": "fetch-failed", "error": format!("{e:#}")}), + ) + } }; let line = match (&status.tag, status.files) { (_, n) if n > 0 => format!( @@ -340,8 +382,8 @@ impl Engine { }; ( d.clone(), - line, - json!({"package": d.id(), "repo": status.repo, "tag": status.tag, "files": status.files, "note": status.note}), + if let Some(note) = &status.note { format!("{line}; {note}") } else { line }, + json!({"package": d.id(), "upstream": status}), ) }) .collect(); @@ -633,7 +675,7 @@ impl Engine { } } if broad { - header = format!("Searched {} direct dependencies. Pass `package` to focus.\n", ready.len()); + header = broad_header(&ready, &missing, tokens); } for (d, e) in &missing { if !broad { @@ -674,7 +716,7 @@ impl Engine { if !used_pkgs.contains(&idx.id()) { used_pkgs.push(idx.id()); } - hits_json.push(json!({"provenance": provenance(&ready), "package": idx.id(), "kind": e.kind.as_str(), "path": e.path, "file": e.file, "line": e.line, "score": score})); + hits_json.push(json!({"package": idx.id(), "kind": e.kind.as_str(), "path": e.path, "file": e.file, "line": e.line, "score": score})); } if out.blocks == 0 { out.push_raw(&format!( diff --git a/crates/lockdocs-core/src/upstream.rs b/crates/lockdocs-core/src/upstream.rs index 56a7b73..cc70533 100644 --- a/crates/lockdocs-core/src/upstream.rs +++ b/crates/lockdocs-core/src/upstream.rs @@ -29,7 +29,7 @@ pub struct Repo { /// Bump when what `fetch` downloads changes, so `lockdocs fetch` refreshes /// older copies (a stale copy is still used until then). -pub const FORMAT: u32 = 3; +pub const FORMAT: u32 = 4; #[derive(Debug, Clone, Serialize, Deserialize)] pub struct Manifest { @@ -41,6 +41,9 @@ pub struct Manifest { /// Immutable package-repository commit resolved from the release tag. #[serde(default)] pub commit: Option, + /// Explicit enrichment completed a docs-site lookup (including a genuine empty result). + #[serde(default)] + pub docs_sites_checked: bool, pub files: usize, pub bytes: u64, /// Why nothing was downloaded, when files == 0. @@ -150,7 +153,10 @@ fn next_major_date(agent: &Client, dep: &Dep, repo: &Repo, major: u64) -> Result ..dep.clone() }; for t in tag_candidates(&next) { - if let Some(v) = api(agent, &format!("/repos/{}/{}/commits/{}", repo.owner, repo.name, enc(&t)))? { + if let Some(sha) = release_commit(repo, &t, |path| api(agent, path))? { + let Some(v) = api(agent, &format!("/repos/{}/{}/commits/{sha}", repo.owner, repo.name))? else { + continue; + }; if let Some(d) = v.pointer("/commit/committer/date").and_then(|d| d.as_str()) { return Ok(Some(d.to_string())); } @@ -315,7 +321,16 @@ fn download(agent: &Client, repo: &Repo, reference: &str, files: &[(String, u64) }) .collect() }); - Ok(results.into_iter().filter_map(|r| r.ok()).collect()) + download_results(results) +} + +fn download_results(results: Vec>) -> Result> { + let failed = results.iter().filter(|r| r.is_err()).count(); + if failed > 0 { + let first = results.iter().find_map(|r| r.as_ref().err()).expect("failed download"); + bail!("{failed}/{} upstream files failed to download: {first:#}", results.len()); + } + results.into_iter().collect() } /// Parse a GitHub URL or shorthand into owner/name. @@ -415,6 +430,19 @@ pub fn stale(m: &Manifest) -> bool { m.format < FORMAT } +/// A prior explicit enrichment is also usable by automatic and offline queries. +pub fn needs_refresh(m: &Manifest, explicit: bool) -> bool { + if explicit { + stale(m) || !m.docs_sites_checked + } else { + !compatible(m) + } +} + +fn compatible(m: &Manifest) -> bool { + m.format >= 3 && (m.commit.is_some() || m.site.is_some() || m.files == 0) +} + /// A previous fetch (with or without files). Never touches the network. pub fn cached(dep: &Dep) -> Option<(PathBuf, Manifest)> { let d = dir(dep); @@ -492,6 +520,41 @@ fn api(agent: &Client, path: &str) -> Result> { Ok(Some(serde_json::from_str(&body)?)) } +/// Resolve an exact tag ref, then peel at most eight annotated tag objects. +/// The generic GET is only to make branch rejection and peeling network-free tests. +fn release_commit(repo: &Repo, tag: &str, mut get: impl FnMut(&str) -> Result>) -> Result> { + let path = format!("/repos/{}/{}/git/ref/tags/{}", repo.owner, repo.name, enc(tag)); + let Some(reference) = get(&path)? else { return Ok(None) }; + if reference.get("ref").and_then(Value::as_str) != Some(format!("refs/tags/{tag}").as_str()) { + bail!("GitHub returned a non-exact tag ref for {tag}"); + } + let mut object = reference.get("object").cloned().context("tag ref has no object")?; + for depth in 0..=8 { + let sha = object + .get("sha") + .and_then(Value::as_str) + .filter(|s| s.len() == 40 && s.bytes().all(|b| b.is_ascii_hexdigit())) + .context("tag object has no immutable SHA")? + .to_string(); + match object.get("type").and_then(Value::as_str) { + Some("commit") => return Ok(Some(sha)), + Some("tag") => { + if depth == 8 { + bail!("release tag exceeds annotated-tag peel limit"); + } + let path = format!("/repos/{}/{}/git/tags/{sha}", repo.owner, repo.name); + object = get(&path)? + .context("annotated tag object unavailable")? + .get("object") + .cloned() + .context("annotated tag has no object")?; + } + _ => bail!("release tag does not point to a commit"), + } + } + bail!("release tag exceeds annotated-tag peel limit") +} + fn enc(s: &str) -> String { let mut out = String::new(); for b in s.bytes() { @@ -610,14 +673,52 @@ impl Drop for FetchGuard { fn publish(out: &Path, target: &Path, m: &Manifest) -> Result<()> { std::fs::write(out.join(".lockdocs-upstream.json"), serde_json::to_string(m)?)?; - if target.exists() { - std::fs::remove_dir_all(target)?; + let backup = target.with_file_name(format!( + "{}.previous-{}", + target.file_name().context("cache name")?.to_string_lossy(), + std::process::id() + )); + let had_prior = target.exists(); + if had_prior { + if backup.exists() { + bail!("previous cache backup already exists; refusing to overwrite it"); + } + std::fs::rename(target, &backup)?; + } + if let Err(e) = std::fs::rename(out, target) { + if had_prior { + let _ = std::fs::rename(&backup, target); + } + return Err(e.into()); + } + if had_prior { + let _ = std::fs::remove_dir_all(backup); } - std::fs::rename(out, target)?; Ok(()) } fn fetch_with(dep: &Dep, src: &Source, automatic: bool) -> Result { + crate::fetch::require_registry_origin(dep)?; + enrich_cached(cached(dep), !automatic, || fetch_attempt(dep, src, automatic)) +} + +fn enrich_cached(prior: Option<(PathBuf, Manifest)>, explicit: bool, attempt: impl FnOnce() -> Result) -> Result { + if let Some((_, m)) = prior.as_ref().filter(|(_, m)| !needs_refresh(m, explicit)) { + return Ok(m.clone()); + } + match attempt() { + Ok(m) => Ok(m), + Err(e) => match prior.filter(|(_, m)| compatible(m)) { + Some((_, mut m)) => { + m.note = Some(format!("upstream refresh failed: {e:#}; previous complete cache retained")); + Ok(m) + } + None => Err(e), + }, + } +} + +fn fetch_attempt(dep: &Dep, src: &Source, automatic: bool) -> Result { let repo = repo_of(dep, src).context("no GitHub repository in the package metadata")?; let target = dir(dep); let parent = target.parent().context("upstream cache parent")?; @@ -638,15 +739,10 @@ fn fetch_with(dep: &Dep, src: &Source, automatic: bool) -> Result { let mut root = None; let mut commit = None; for t in tag_candidates(dep) { - let Some(release) = api(&agent, &format!("/repos/{}/{}/commits/{}", repo.owner, repo.name, enc(&t)))? else { + let Some(sha) = release_commit(&repo, &t, |path| api(&agent, path))? else { continue; }; - let sha = release - .get("sha") - .and_then(Value::as_str) - .filter(|s| s.len() == 40 && s.bytes().all(|b| b.is_ascii_hexdigit())) - .context("GitHub release did not resolve to an immutable commit")?; - if let Some(items) = tree(&agent, &repo, sha, false)? { + if let Some(items) = tree(&agent, &repo, &sha, false)? { commit = Some(sha.to_string()); tag = Some(t); root = Some(items); @@ -659,6 +755,7 @@ fn fetch_with(dep: &Dep, src: &Source, automatic: bool) -> Result { repo: label, tag: None, commit: None, + docs_sites_checked: !automatic, files: 0, bytes: 0, note: Some(format!("no git tag found for {}", dep.version)), @@ -677,7 +774,7 @@ fn fetch_with(dep: &Dep, src: &Source, automatic: bool) -> Result { } } Ok(None) => {} - Err(e) => m.note = Some(format!("{}; docs site skipped: {e:#}", m.note.unwrap_or_default())), + Err(e) => return Err(e.context("docs-site selection failed")), } publish(&out, &target, &m)?; return Ok(m); @@ -764,13 +861,11 @@ fn fetch_with(dep: &Dep, src: &Source, automatic: bool) -> Result { }); files.truncate(MAX_FILES); let ok = download(&agent, &repo, commit.as_deref().context("missing release commit")?, &files, &out)?; - let failed = files.len() - ok.len(); // A separate docs-site repository, when the package keeps its docs there. let mut site = None; let mut site_n = 0usize; let mut site_bytes = 0u64; let mut pages = Vec::new(); - let mut site_err = None; match if automatic { Ok(None) } else { site_files(&agent, dep, Some(&repo)) } { Ok(Some(s)) => { (site_n, site_bytes, pages) = download_site(&agent, &s, &out)?; @@ -779,21 +874,20 @@ fn fetch_with(dep: &Dep, src: &Source, automatic: bool) -> Result { } } Ok(None) => {} - Err(e) => site_err = Some(format!("docs site skipped: {e:#}")), + Err(e) => return Err(e.context("docs-site selection failed")), } let m = Manifest { format: FORMAT, repo: label, tag: Some(tag), commit, + docs_sites_checked: !automatic, files: ok.len() + site_n, bytes: ok.iter().sum::() + site_bytes, note: if ok.is_empty() && site_n == 0 { Some("no docs folder at that tag".into()) - } else if failed > 0 { - Some(format!("{failed} files failed to download")) } else { - site_err + None }, site, pages, @@ -823,6 +917,132 @@ fn download_site(agent: &Client, site: &SiteFiles, out: &Path) -> Result<(usize, #[cfg(test)] mod tests { use super::*; + fn manifest() -> Manifest { + Manifest { + format: FORMAT, + repo: "github.com/o/r".into(), + tag: Some("v1.2.3".into()), + commit: Some("0123456789abcdef0123456789abcdef01234567".into()), + docs_sites_checked: false, + files: 1, + bytes: 10, + note: None, + site: None, + pages: Vec::new(), + } + } + + #[test] + fn release_ref_rejects_branch_only_and_peels_annotated_tags() { + let repo = Repo { + owner: "o".into(), + name: "r".into(), + subdir: None, + }; + let sha = "0123456789abcdef0123456789abcdef01234567"; + let mut paths = Vec::new(); + let none = release_commit(&repo, "v1.2.3", |p| { + paths.push(p.to_string()); + Ok(None) + }) + .unwrap(); + assert!(none.is_none()); + assert_eq!(paths, vec!["/repos/o/r/git/ref/tags/v1.2.3"]); + let branch = release_commit(&repo, "v1.2.3", |_| { + Ok(Some(serde_json::json!({ + "ref": "refs/heads/v1.2.3", "object": {"type":"commit", "sha":sha} + }))) + }); + assert!(branch.is_err()); + let mut calls = 0; + let commit = release_commit(&repo, "v1.2.3", |p| { + calls += 1; + if p.contains("/git/ref/tags/") { + Ok(Some(serde_json::json!({"ref":"refs/tags/v1.2.3", "object":{"type":"tag", "sha":sha}}))) + } else { + Ok(Some(serde_json::json!({"object":{"type":"commit", "sha":sha}}))) + } + }) + .unwrap(); + assert_eq!(commit.as_deref(), Some(sha)); + assert_eq!(calls, 2); + let mut calls = 0; + let cycle = release_commit(&repo, "v1.2.3", |_| { + calls += 1; + Ok(Some(serde_json::json!({"ref":"refs/tags/v1.2.3", "object":{"type":"tag", "sha":sha}}))) + }); + assert!(cycle.is_err()); + assert_eq!(calls, 9); // one exact-ref lookup plus at most eight tag-object lookups + } + + #[test] + fn enrichment_policy_upgrades_once_and_preserves_opted_in_sites() { + let automatic = manifest(); + assert!(needs_refresh(&automatic, true)); + assert!(!needs_refresh(&automatic, false)); + let mut explicit = automatic.clone(); + explicit.docs_sites_checked = true; + explicit.site = Some("github.com/o/site@resolved (major-version docs)".into()); + explicit.files = 2; + let upgraded = enrich_cached(Some((PathBuf::new(), automatic)), true, || Ok(explicit.clone())).unwrap(); + assert_eq!(upgraded.files, 2); + for policy in [false, true] { + let reused = enrich_cached(Some((PathBuf::new(), upgraded.clone())), policy, || panic!("needless network refresh")).unwrap(); + assert_eq!(reused.site, explicit.site); + } + } + + #[test] + fn failed_downloads_are_not_empty_docs_or_published_partial_success() { + assert!(download_results(vec![]).unwrap().is_empty()); + assert!(download_results(vec![Err(anyhow::anyhow!("HTTP 503"))]) + .unwrap_err() + .to_string() + .contains("1/1")); + assert!(download_results(vec![Ok(10), Err(anyhow::anyhow!("HTTP 503"))]) + .unwrap_err() + .to_string() + .contains("1/2")); + let old = manifest(); + let preserved = enrich_cached(Some((PathBuf::new(), old.clone())), true, || { + download_results(vec![Ok(10), Err(anyhow::anyhow!("HTTP 503"))]).map(|_| unreachable!()) + }) + .unwrap(); + assert_eq!(preserved.commit, old.commit); + assert_eq!(preserved.files, old.files); + assert!(!preserved.docs_sites_checked); + assert!(preserved.note.unwrap().contains("previous complete cache retained")); + assert!(enrich_cached(None, true, || Err(anyhow::anyhow!("download failed"))).is_err()); + } + + #[test] + fn git_origins_are_rejected_before_repository_or_tag_requests() { + let source = Source { + dir: PathBuf::new(), + files: None, + metadata: None, + version: "1.2.3".into(), + label: "checkout".into(), + fetched: false, + }; + for (eco, from) in [(Eco::Cargo, "Cargo.lock (git)"), (Eco::PyPI, "uv.lock (git)")] { + let dep = Dep { + eco, + name: "git-fixture".into(), + version: "1.2.3".into(), + direct: true, + from: from.into(), + }; + for result in [ + fetch(&dep, &source), + fetch_automatic(&dep, &source), + crate::fetch::fetch(&dep).map(|_| manifest()), + ] { + assert!(result.unwrap_err().to_string().contains("git dependency")); + } + } + } + #[test] fn automatic_client_is_anonymous_and_bounded() { let mut client = agent(true); diff --git a/crates/lockdocs-core/tests/cargo_git.rs b/crates/lockdocs-core/tests/cargo_git.rs index 6f2dd28..e5fc9bc 100644 --- a/crates/lockdocs-core/tests/cargo_git.rs +++ b/crates/lockdocs-core/tests/cargo_git.rs @@ -11,11 +11,23 @@ fn finds_git_dependencies_in_cargo_checkouts() { std::env::set_var("LOCKDOCS_CACHE", &cache); std::env::set_var("CARGO_HOME", tests.join("fixture-git-home")); std::env::set_var("LOCKDOCS_NO_SYSTEM_PYTHON", "1"); - let e = Engine::new(&tests.join("fixture-git"), Options { fetch: false, upstream: false }); + std::env::set_var("LOCKDOCS_EMBED", "0"); + let mut e = Engine::new(&tests.join("fixture-git"), Options { fetch: true, upstream: true }); let r = e.resolve(Some("tinygit")); assert!(r.text.contains("tinygit@0.2.0") && r.text.contains("docs ready"), "{}", r.text); let a = e.api("tinygit::clone_into", None, 800).unwrap(); assert!(a.text.contains("Clones a repository into a directory"), "{}", a.text); assert!(a.text.contains("pinned in Cargo.lock (git)"), "{}", a.text); + assert!(a.text.contains("git dependency's resolved commit"), "{}", a.text); + let fetched = e.fetch_all(Some("tinygit")).unwrap(); + assert_eq!(fetched.json["packages"][0]["status"], "git-checkout"); + let uv = r#"[[package]] +name = "uv-git-fixture" +version = "1.2.3" +source = { git = "https://github.com/example/git-fixture?rev=abc#01234567" } +"#; + e.project.deps = lockdocs_core::lockfile::uv_lock(uv).unwrap(); + let fetched = e.fetch_all(Some("uv-git-fixture")).unwrap(); + assert_eq!(fetched.json["packages"][0]["status"], "git-checkout"); let _ = std::fs::remove_dir_all(&cache); } diff --git a/crates/lockdocs-core/tests/integration.rs b/crates/lockdocs-core/tests/integration.rs index ff35e27..4bd76d2 100644 --- a/crates/lockdocs-core/tests/integration.rs +++ b/crates/lockdocs-core/tests/integration.rs @@ -93,10 +93,20 @@ fn first_use_reports_unavailable_upstream_and_reuses_note() { let e = Engine::new(&root, Options { fetch: false, upstream: true }); for _ in 0..2 { let a = e.docs(Some("tiny-schema"), "reject unknown keys", 1500).unwrap(); - assert!(a.text.contains("upstream docs unavailable"), "{}", a.text); + assert!(a.text.contains("upstream docs fetch failed"), "{}", a.text); assert_eq!(a.json["provenance"][0]["requested_version"], "1.2.0"); assert!(a.json["provenance"][0]["note"].as_str().unwrap().contains("no GitHub repository")); } + let broad = e.docs(None, "reject unknown keys", 1200).unwrap(); + assert!( + broad.text.contains("fallback notes") && broad.text.contains("Fallback tiny-schema@1.2.0"), + "{}", + broad.text + ); + assert!(lockdocs_core::est_tokens(&broad.text) <= 1230, "{}", broad.text); + for hit in broad.json["hits"].as_array().unwrap() { + assert!(hit.get("provenance").is_none()); + } let offline = engine().docs(Some("tiny-schema"), "reject unknown keys", 1500).unwrap(); assert!(offline.json["provenance"][0]["note"].is_null()); } @@ -123,15 +133,20 @@ fn pinned_upstream_cache_is_reused_online_and_offline_with_provenance() { "format": lockdocs_core::upstream::FORMAT, "repo": "github.com/example/pinned", "tag": "v1.2.3", "commit": "0123456789abcdef0123456789abcdef01234567", - "files": 1, "bytes": 73, "note": null, "site": null, "pages": [], + "files": 1, "bytes": 73, "note": null, "site": "github.com/example/site@fixed (major docs)", "pages": [], "docs_sites_checked": true, }); std::fs::write(cache.join(".lockdocs-upstream.json"), manifest.to_string()).unwrap(); - for upstream in [true, false] { - let mut e = Engine::new(&root, Options { fetch: false, upstream }); + for opts in [ + Options { fetch: true, upstream: true }, + Options { fetch: false, upstream: true }, + Options { fetch: false, upstream: false }, + ] { + let mut e = Engine::new(&root, opts); e.project.deps = vec![dep.clone()]; let a = e.docs(Some(&dep.name), "reject unknown keys", 1500).unwrap(); assert!(a.text.contains("release_only_api") && a.text.contains("upstream:docs/guide.md"), "{}", a.text); assert_eq!(a.json["provenance"][0]["upstream"]["commit"], manifest["commit"]); + assert_eq!(a.json["provenance"][0]["upstream"]["site"], manifest["site"]); assert_eq!(a.json["provenance"][0]["requested_version"], "1.2.3"); } let _ = std::fs::remove_dir_all(&root); diff --git a/docs/guide/fetch.md b/docs/guide/fetch.md index 91b2d3e..0a50ec1 100644 --- a/docs/guide/fetch.md +++ b/docs/guide/fetch.md @@ -4,7 +4,7 @@ lockdocs starts with installed package files and automatically adds public relea ## First-use defaults and controls -The query's selected packages are enriched, not every dependency in the project. The existing upstream fetcher reads the repository from package metadata, resolves a candidate tag for exactly the requested version to a full immutable commit, then reads only that commit's docs. Automatic fetches are anonymous: ambient `GITHUB_TOKEN` and `GH_TOKEN` are not read. Git dependencies use their checkout files rather than guessing a release tag for their resolved commit. Automatic fetches do not use major-version docs sites or fall back to a default branch/latest version. +The query's selected packages are enriched, not every dependency in the project. The existing upstream fetcher reads the repository from package metadata, resolves an exact `refs/tags/` ref for the requested version, peeling at most eight annotated tags to a full immutable commit (a same-named branch is rejected), then reads only that commit's docs. Automatic fetches are anonymous: ambient `GITHUB_TOKEN` and `GH_TOKEN` are not read. Git dependencies use their checkout files rather than guessing a release tag for their resolved commit. Automatic fetches do not download major-version docs sites or fall back to a default branch/latest version. - `--no-fetch` or `LOCKDOCS_FETCH=0`: no query package/upstream downloads; installed and cached docs still work. The model has its own policy. - `--offline` or `LOCKDOCS_OFFLINE=1`: no downloads, including the embedding model. `fetch --offline` is rejected. @@ -12,7 +12,9 @@ The query's selected packages are enriched, not every dependency in the project. - `LOCKDOCS_NO_UPSTREAM=1`: package-only answers, even if upstream docs are cached. - `--fetch` or `LOCKDOCS_FETCH=1`: explicitly enable missing registry packages and major-version docs-site enrichment too. -Automatic upstream work has a 45-second scheduling budget per package and a 10-second timeout per request (an in-flight request may finish after the scheduling deadline), no redirects, at most 2,500 files / 40 MB of advertised content and a 2 MB per-file limit, with 16 downloads in flight. Failed or missing enrichment is reported in text and structured `provenance`; answers retain the requested version's package files. Successful caches include repository, release tag, immutable commit, file counts and partial-download notes. Later queries reuse them without network. A failed download preserves the previous complete cache; concurrent fetches fail without waiting. An interrupted process can leave a fetch lock; `lockdocs cache clean` removes it. +Automatic upstream work has a 45-second scheduling budget per package and a 10-second timeout per request (an in-flight request may finish after the scheduling deadline), no redirects, at most 2,500 files / 40 MB of advertised content and a 2 MB per-file limit, with 16 downloads in flight. Failed or missing enrichment is reported in text and structured `provenance`; answers retain the requested version's package files. Successful caches include repository, release tag, immutable commit, file counts and whether explicit docs-site enrichment completed. Later queries reuse them without network. Any failed file or docs-site lookup aborts publication: a compatible previous complete cache is retained and returned with a transient failure note. Without a compatible prior cache, the failure is reported and no permanent negative cache is written. Concurrent fetches fail without waiting. An interrupted process can leave a fetch lock; `lockdocs cache clean` removes it. + +A release-only automatic cache is upgraded when explicit enrichment is requested. Once docs sites have been explicitly cached, default and offline queries preserve and reuse those files and their provenance without destructive refresh. Truly empty selections are cached separately from failed downloads. Broad text answers include a bounded fallback/source summary; JSON carries complete provenance once at the top level. Network requests reveal the public repository/version/file being fetched, not the question or project contents. Anonymous GitHub limits can stop a cold multi-package session; failures are reported, not retried within an indexed MCP session. From b05aecda72fcfabce4453f5ec52bae65fb61668b Mon Sep 17 00:00:00 2001 From: Sylphx Builder Date: Thu, 1 Oct 2026 02:04:08 +0000 Subject: [PATCH 5/5] fix: retry failed site lookups and revalidate legacy cache provenance --- bench/run.py | 9 ++- bench/test_run.py | 15 ++++ crates/lockdocs-core/src/index.rs | 4 +- crates/lockdocs-core/src/query.rs | 41 +++++++++- crates/lockdocs-core/src/upstream.rs | 91 +++++++++++++++++++++-- crates/lockdocs-core/tests/integration.rs | 19 +++++ docs/guide/fetch.md | 2 +- 7 files changed, 167 insertions(+), 14 deletions(-) diff --git a/bench/run.py b/bench/run.py index a6357f5..22bc82a 100755 --- a/bench/run.py +++ b/bench/run.py @@ -268,6 +268,13 @@ def result_of(r, key): return r.get("context7", {}) if key == "context7" else r["variants"].get(key, {}) +def fetch_file_count(package): + """Accept the original flat fetch report and the additive manifest layout.""" + if "files" in package: + return package["files"] or 0 + return (package.get("upstream") or {}).get("files", 0) + + def markdown(res, with_c7): s = res["summary"] cols = [(n, label) for n, label in res.get("variants", [["lockdocs", "lockdocs"]])] @@ -311,7 +318,7 @@ def markdown(res, with_c7): for p, v in res["fetch"].items(): rep = v["report"] if isinstance(rep, dict): - pk = ", ".join(f"{x['package']} {x.get('files', 0)} files" for x in rep.get("packages", []) if x.get("files")) + pk = ", ".join(f"{x['package']} {fetch_file_count(x)} files" for x in rep.get("packages", []) if fetch_file_count(x)) out.append(f"- {p}: {v['ms']} ms ({pk or 'no upstream docs'})") else: out.append(f"- {p}: {v['ms']} ms (error)") diff --git a/bench/test_run.py b/bench/test_run.py index 7838236..2acd819 100644 --- a/bench/test_run.py +++ b/bench/test_run.py @@ -24,6 +24,21 @@ def test_floors_apply_only_to_full_suite(self): with self.assertRaises(RuntimeError): runner.check_floors({variant: {"passed": score}}, 105) + def test_fetch_formatter_preserves_flat_and_nested_file_counts(self): + for package in [ + {"package": "axum@0.7.9", "files": 20}, + {"package": "axum@0.7.9", "upstream": {"files": 20}}, + {"package": "axum@0.7.9", "files": 20, "upstream": {"files": 20}}, + ]: + self.assertEqual(runner.fetch_file_count(package), 20) + result = {"tokenizer": "test", "runner": {"os": "Test", "machine": "test"}, + "rows": [], "summary": {}, "variants": [], "index": {}, + "fetch": {"axum07": {"ms": 1, "report": {"packages": [package]}}}} + formatted = runner.markdown(result, False) + self.assertIn("axum@0.7.9 20 files", formatted) + self.assertNotIn("no upstream docs", formatted) + self.assertEqual(runner.fetch_file_count({"files": 0, "upstream": {"files": 20}}), 0) + def test_first_use_has_own_empty_cache_and_no_opt_in(self): with tempfile.TemporaryDirectory(prefix="lockdocs-bench-test-") as tmp: root = Path(tmp) diff --git a/crates/lockdocs-core/src/index.rs b/crates/lockdocs-core/src/index.rs index 3609c29..ff8fa10 100644 --- a/crates/lockdocs-core/src/index.rs +++ b/crates/lockdocs-core/src/index.rs @@ -159,8 +159,8 @@ fn cache_path(dep: &Dep, src: &Source, types: Option<&Path>, up: Option<&(PathBu let up_key = up .map(|(_, m)| { format!( - "{}@{:?}:{}:{:?}:{:?}:{:?}:{}", - m.repo, m.tag, m.files, m.site, m.commit, m.note, m.docs_sites_checked + "{}@{:?}:{}:{:?}:{:?}:{:?}:{}:{}", + m.repo, m.tag, m.files, m.site, m.commit, m.note, m.docs_sites_checked, m.format ) }) .unwrap_or_default(); diff --git a/crates/lockdocs-core/src/query.rs b/crates/lockdocs-core/src/query.rs index ca88e89..99fab8c 100644 --- a/crates/lockdocs-core/src/query.rs +++ b/crates/lockdocs-core/src/query.rs @@ -69,6 +69,12 @@ enum Resolved { Missing(Dep, String), } +/// Keep the original flat fields for CLI consumers; richer provenance is additive. +fn fetch_report(dep: &Dep, status: &upstream::Manifest) -> Value { + json!({"package": dep.id(), "repo": status.repo, "tag": status.tag, + "files": status.files, "note": status.note, "upstream": status}) +} + fn provenance(ready: &[ReadyPkg]) -> Vec { ready .iter() @@ -294,7 +300,16 @@ impl Engine { } let cached = upstream::cached(dep); if let Some(c) = cached.as_ref().filter(|(_, m)| !enabled || !upstream::needs_refresh(m, self.opts.fetch)) { - return (Some(c.clone()), c.1.note.clone()); + let note = if upstream::needs_refresh(&c.1, false) { + Some(format!( + "offline fallback: legacy upstream cache format {} has not been exact-tag revalidated; its old commit candidate may have been a branch{}", + c.1.format, + c.1.note.as_ref().map(|n| format!("; {n}")).unwrap_or_default() + )) + } else { + c.1.note.clone() + }; + return (Some(c.clone()), note); } if enabled { let result = if self.opts.fetch { @@ -383,7 +398,7 @@ impl Engine { ( d.clone(), if let Some(note) = &status.note { format!("{line}; {note}") } else { line }, - json!({"package": d.id(), "upstream": status}), + fetch_report(d, &status), ) }) .collect(); @@ -1588,6 +1603,28 @@ pub fn dep_key(d: &Dep) -> String { mod tests { use super::*; + #[test] + fn fetch_report_preserves_flat_counts_and_adds_full_provenance() { + let dep = Dep { + eco: Eco::Cargo, + name: "axum".into(), + version: "0.7.9".into(), + direct: true, + from: "Cargo.lock".into(), + }; + let status: upstream::Manifest = serde_json::from_value(json!({ + "format": upstream::FORMAT, "repo": "github.com/tokio-rs/axum", "tag": "axum-v0.7.9", + "commit": "0123456789abcdef0123456789abcdef01234567", "files": 20, "bytes": 100, "note": null, + })) + .unwrap(); + let report = fetch_report(&dep, &status); + assert_eq!(report["files"], 20); + assert_eq!(report["files"], report["upstream"]["files"]); + assert_eq!(report["tag"], report["upstream"]["tag"]); + assert_eq!(report["repo"], report["upstream"]["repo"]); + assert!(report.get("note").is_some()); + } + #[test] fn upstream_is_default_with_explicit_opt_out() { assert!(automatic_upstream(None, false)); diff --git a/crates/lockdocs-core/src/upstream.rs b/crates/lockdocs-core/src/upstream.rs index cc70533..04a7195 100644 --- a/crates/lockdocs-core/src/upstream.rs +++ b/crates/lockdocs-core/src/upstream.rs @@ -130,19 +130,47 @@ const DOCS_SITES: &[DocsSite] = &[ ]; /// The latest stable major of a package, from its registry. -fn latest_major(agent: &Client, dep: &Dep) -> Option { +fn latest_major(agent: &Client, dep: &Dep) -> Result> { let name = dep.name.strip_prefix("@types/").unwrap_or(&dep.name); let (url, pointer) = match dep.eco { Eco::Npm => (format!("https://registry.npmjs.org/{}/latest", name.replace('/', "%2F")), "/version"), Eco::Cargo => (format!("https://crates.io/api/v1/crates/{name}"), "/crate/max_stable_version"), Eco::PyPI => (format!("https://pypi.org/pypi/{name}/json"), "/info/version"), - Eco::Go => return None, + Eco::Go => return Ok(None), }; - let mut res = agent.get(&url).call().ok()?; + agent.check_deadline()?; + let mut res = agent.get(&url).call().with_context(|| format!("latest-major lookup: {url}"))?; + let status = res.status().as_u16(); let mut body = String::new(); - res.body_mut().as_reader().take(16 << 20).read_to_string(&mut body).ok()?; - let v: Value = serde_json::from_str(&body).ok()?; - v.pointer(pointer)?.as_str()?.split('.').next()?.parse().ok() + res.body_mut() + .as_reader() + .take((16 << 20) + 1) + .read_to_string(&mut body) + .context("latest-major response read failed")?; + if body.len() > 16 << 20 { + bail!("latest-major response exceeds 16 MB"); + } + latest_major_response(status, &body, pointer) +} + +/// Only a genuine missing registry record is an empty lookup. Network, HTTP, +/// malformed JSON and invalid version responses must leave enrichment retryable. +fn latest_major_response(status: u16, body: &str, pointer: &str) -> Result> { + if status == 404 { + return Ok(None); + } + if status != 200 { + bail!("latest-major registry lookup failed: HTTP {status}"); + } + let v: Value = serde_json::from_str(body).context("invalid latest-major registry JSON")?; + let version = v.pointer(pointer).and_then(Value::as_str).context("latest-major response missing version")?; + let major = version + .split('.') + .next() + .context("latest-major response has empty version")? + .parse() + .context("latest-major response has invalid version")?; + Ok(Some(major)) } /// When the package's next major was released: the commit date of its @@ -180,7 +208,7 @@ fn site_files(agent: &Client, dep: &Dep, pkg_repo: Option<&Repo>) -> Result bool { } fn compatible(m: &Manifest) -> bool { - m.format >= 3 && (m.commit.is_some() || m.site.is_some() || m.files == 0) + m.format >= FORMAT && (m.commit.is_some() || m.site.is_some() || m.files == 0) } /// A previous fetch (with or without files). Never touches the network. @@ -975,6 +1003,53 @@ mod tests { assert_eq!(calls, 9); // one exact-ref lookup plus at most eight tag-object lookups } + #[test] + fn latest_major_failures_remain_retryable_and_explicit_retry_adds_site() { + assert_eq!(latest_major_response(404, "not found", "/version").unwrap(), None); + assert_eq!(latest_major_response(200, r#"{"version":"2.3.4"}"#, "/version").unwrap(), Some(2)); + for (status, body) in [ + (503, "unavailable"), + (429, "limited"), + (200, "not json"), + (200, "{}"), + (200, r#"{"version":"invalid"}"#), + ] { + assert!(latest_major_response(status, body, "/version").is_err()); + } + let prior = manifest(); + let failed = enrich_cached(Some((PathBuf::new(), prior.clone())), true, || { + latest_major_response(503, "unavailable", "/version")?; + unreachable!("failed latest lookup cannot mark docs-site enrichment complete") + }) + .unwrap(); + assert_eq!(failed.files, prior.files); + assert!(!failed.docs_sites_checked); + assert!(failed.note.as_deref().unwrap().contains("HTTP 503")); + let retried = enrich_cached(Some((PathBuf::new(), failed)), true, || { + assert_eq!(latest_major_response(200, r#"{"version":"2.3.4"}"#, "/version")?, Some(2)); + let mut enriched = prior; + enriched.docs_sites_checked = true; + enriched.site = Some("github.com/o/site@fixed (major 2)".into()); + enriched.files += 1; + Ok(enriched) + }) + .unwrap(); + assert!(retried.docs_sites_checked && retried.site.is_some()); + assert_eq!(retried.files, 2); + assert!(retried.note.is_none()); + } + + #[test] + fn legacy_candidate_commit_cache_cannot_be_trusted_online() { + let mut legacy = manifest(); + legacy.format = 3; + legacy.docs_sites_checked = true; + assert!(needs_refresh(&legacy, false)); + assert!(needs_refresh(&legacy, true)); + let rejected = enrich_cached(Some((PathBuf::new(), legacy)), false, || Err(anyhow::anyhow!("exact tag revalidation failed"))); + assert!(rejected.is_err()); + } + #[test] fn enrichment_policy_upgrades_once_and_preserves_opted_in_sites() { let automatic = manifest(); diff --git a/crates/lockdocs-core/tests/integration.rs b/crates/lockdocs-core/tests/integration.rs index 4bd76d2..c484440 100644 --- a/crates/lockdocs-core/tests/integration.rs +++ b/crates/lockdocs-core/tests/integration.rs @@ -149,6 +149,25 @@ fn pinned_upstream_cache_is_reused_online_and_offline_with_provenance() { assert_eq!(a.json["provenance"][0]["upstream"]["site"], manifest["site"]); assert_eq!(a.json["provenance"][0]["requested_version"], "1.2.3"); } + // A format-3 commit candidate was resolved through /commits/{candidate}, + // which could have been a branch. Online use cannot silently trust it. + let mut legacy = manifest.clone(); + legacy["format"] = serde_json::json!(3); + std::fs::write(cache.join(".lockdocs-upstream.json"), legacy.to_string()).unwrap(); + let mut online = Engine::new(&root, Options { fetch: false, upstream: true }); + online.project.deps = vec![dep.clone()]; + let rejected = online.docs(Some(&dep.name), "reject unknown keys", 1500).unwrap(); + assert!(rejected.json["provenance"][0]["upstream"].is_null()); + assert!(rejected.text.contains("upstream docs fetch failed")); + assert!(!rejected.text.contains("release_only_api")); + // Failed revalidation does not destroy the disk cache. Explicitly offline + // callers may read it, with its unverified provenance clearly disclosed. + let mut offline = Engine::new(&root, Options { fetch: false, upstream: false }); + offline.project.deps = vec![dep]; + let fallback = offline.docs(Some("cached-provenance-fixture"), "reject unknown keys", 1500).unwrap(); + assert!(fallback.text.contains("release_only_api") && fallback.text.contains("offline fallback")); + assert!(fallback.text.contains("may have been a branch")); + assert_eq!(fallback.json["provenance"][0]["upstream"]["format"], 3); let _ = std::fs::remove_dir_all(&root); let _ = std::fs::remove_dir_all(&cache); } diff --git a/docs/guide/fetch.md b/docs/guide/fetch.md index 0a50ec1..49161ba 100644 --- a/docs/guide/fetch.md +++ b/docs/guide/fetch.md @@ -14,7 +14,7 @@ The query's selected packages are enriched, not every dependency in the project. Automatic upstream work has a 45-second scheduling budget per package and a 10-second timeout per request (an in-flight request may finish after the scheduling deadline), no redirects, at most 2,500 files / 40 MB of advertised content and a 2 MB per-file limit, with 16 downloads in flight. Failed or missing enrichment is reported in text and structured `provenance`; answers retain the requested version's package files. Successful caches include repository, release tag, immutable commit, file counts and whether explicit docs-site enrichment completed. Later queries reuse them without network. Any failed file or docs-site lookup aborts publication: a compatible previous complete cache is retained and returned with a transient failure note. Without a compatible prior cache, the failure is reported and no permanent negative cache is written. Concurrent fetches fail without waiting. An interrupted process can leave a fetch lock; `lockdocs cache clean` removes it. -A release-only automatic cache is upgraded when explicit enrichment is requested. Once docs sites have been explicitly cached, default and offline queries preserve and reuse those files and their provenance without destructive refresh. Truly empty selections are cached separately from failed downloads. Broad text answers include a bounded fallback/source summary; JSON carries complete provenance once at the top level. +A release-only automatic cache is upgraded when explicit enrichment is requested. Once docs sites have been explicitly cached, default and offline queries preserve and reuse those files and their provenance without destructive refresh. Transient registry HTTP, network, body-read or parsing failures leave docs-site enrichment incomplete and retryable; only a successful lookup (including a genuine missing registry record) can mark it checked. Truly empty selections are cached separately from failed downloads. Legacy format-3 caches are revalidated online because their commit candidates could have named branches. Offline use can retain those files, with an explicit unverified-cache fallback note. Broad text answers include a bounded fallback/source summary; JSON carries complete provenance once at the top level. Network requests reveal the public repository/version/file being fetched, not the question or project contents. Anonymous GitHub limits can stop a cold multi-package session; failures are reported, not retried within an indexed MCP session.