Skip to content

Latest commit

 

History

History
61 lines (45 loc) · 18.3 KB

File metadata and controls

61 lines (45 loc) · 18.3 KB

Analysis engine

analyseRepository(handle, options) in packages/core/src/engine/ is the one entry point the CLI (#39) and the GitHub App check reporter (#32) call. It runs pipeline steps 2-8 from architecture.md over a RepositoryHandle and returns one AnalysisResult.

import { analyseRepository } from "@ghostdeps/core";

const result = await analyseRepository(handle, {
  adapters: [javascriptAdapter, pythonAdapter],
  network: { mode: "offline" }, // default
  detectionThreshold: 0.5, // default, DEFAULT_DETECTION_THRESHOLD
  adapterTimeoutMs: 60_000, // default, per adapter stage (async waits only, see below)
  usageConcurrency: 8, // default, concurrent findUsage calls per adapter
  recommend: policy, // optional; omit for facts only
});

Behaviour

  • Detection gate. Every adapter's detect() runs in parallel. Only adapters at or above the threshold go further. Scores map to Confidence as high (>= 0.8), medium (>= 0.5), low.
  • Capabilities. buildDependencyGraph and findUsage run only when the adapter declares the capability and implements the method. Graph and usage stages run in parallel per adapter.
  • API version. Adapters whose apiVersion does not match core's adapterApiVersion (major.minor while pre-1.0) are skipped with an info finding.
  • Error isolation. A stage that throws, or whose async work passes adapterTimeoutMs, becomes an info finding (confidence low, with limitations) for that adapter. Other adapters and earlier facts are kept. Adapter error text is flattened and truncated before it enters a finding.
  • Timeout limits (important). The engine has two isolation tiers. In-process (analyseRepository), the timeout only bounds async waits: synchronous CPU work inside an adapter (a big JSON.parse, a YAML parse, a pathological source file) cannot be preempted, so the run waits for it; a fired timeout aborts AdapterContext.signal and stops scheduling that adapter's remaining usage calls, but cannot kill work already running. The worker-thread tier (analyseRepositoryIsolated, #90) runs each adapter in its own worker_threads Worker loaded by module specifier, with a heap ceiling (resourceLimits, default 512 MiB old generation) and a watchdog that terminate()s the worker when a stage outlives adapterTimeoutMs (plus a small grace so a healthy worker reports its own async timeout first). A busy-loop adapter is killed at the budget and a heap-hungry adapter dies on its own ceiling; both become the same info findings as in-process failures, and untrusted parsing never touches the main thread's heap. Worker stdout/stderr is captured, never inherited by the parent (a stray adapter log line must not corrupt ghostdeps scan --json), and drained to the optional debugLog sink. A worker that posts its outcome resolves at exit so captured output has drained before callers see the result. What a worker may post back is count-capped main-side (OUTCOME_CAPS: 10,000 dependencies, 50,000 usages, 100,000 graph nodes-plus-closure-entries, 100 evidence entries per finding, 1,000 findings); the structured clone has already landed in the main heap when the caps run, so they bound what flows downstream, not peak memory. Overflow becomes an info finding with limitations - capped usages also mark the outcome's usage analysis incomplete so policy never reads a truncated usage list as "analysed, none found", and oversized graphs are dropped whole rather than truncated into inconsistency. Workers start in waves of maxParallelAdapters (default 4), so worst-case adapter heap is maxParallelAdapters x adapterHeapMb (2 GiB with defaults), not N x the ceiling. Adapters must still check signal between files and keep their own input ceilings (security-model rule 3). Adapter module specifiers are trusted configuration: they go to import() in the worker, so they must come from CLI flags or project config and never from repository content - a manifest or lockfile must never be able to choose what code the engine loads. The rule is enforced: a file URL or absolute path that resolves inside the analysed repository root (symlinks resolved) is rejected with an info finding before a worker starts, and relative specifiers are rejected outright (#150).
  • Bounded concurrency. findUsage runs for at most usageConcurrency dependencies at once per adapter.
  • Recommendation policy. Core-owned and injected via recommend. It receives dependencies, usages, graphs, and usageAnalysedEcosystems, the ecosystems where usage analysis actually completed. Policy must not call a dependency unused outside that set. A policy failure becomes an info finding; facts survive. Findings that are not plain JSON data (Map, Date, class instances, non-finite numbers) are dropped with an info finding, because renderJsonReport rejects them and one bad finding must not abort the run. As a second guard, the canonical sort (normaliseAnalysisResult) never throws, even on such values, so ordering can't abort a run on its own.
  • Peer-host attribution (#395). The JS/TS lockfile graph keeps an additive direct-host peer-edge index alongside its transitive closure. For npm and Bun, an unverified-no-imports note names the direct package whose own lockfile entry declares the peer, rather than the first package whose wider closure reaches it. A declared but unresolved peer stays a low-confidence manual-review note, not an unused verdict. The lockfile alone does not establish host usage or peer-range compatibility; other lockfile formats and transitive-only paths still receive the conservative closure note. No registry lookup or repository code execution is involved.
  • Determinism. Projects, dependencies, usages, findings, detections and surfaces are sorted by stable keys, so adapter order and completion timing never change the output. The engine uses the same canonical ordering as the JSON reporter (normaliseAnalysisResult), so there is one sorter to keep correct.
  • Surface. direct counts unique direct dependency names per ecosystem; transitive counts unique node names across that ecosystem's graphs (0 without a graph). graphs says how far that count can be trusted (#114): "none" means no usable graph was built (no lockfile, an unparseable one, or a failed graph stage), so the 0 means unknown; "partial" means some graph was built but a detected project has no graph, or a graph is marked incomplete, so transitive is a lower bound; "complete" means every detected project has a complete graph. Renderers must annotate or withhold totals that are not "complete". The field is additive in schema version 1: a result without it must be read as unknown.
  • Project tree (#55). projectTree lists every detected project, from every adapter, with a deterministic id ("<ecosystem>:<path>") and a parent: the nearest enclosing project of any ecosystem, meaning the deepest project whose path is a strict ancestor directory. Ties at that path go to the lowest ecosystem name. Projects at the same path never parent each other, and top-level projects have no parent. Adapters still discover their own projects (ADR 0002). Core only relates them, so a Python service inside a JS workspace root reads as one tree. The field is additive in schema version 1.
  • Unified graph (#55). graph merges every project's lockfile graph into one package index. Each node is keyed "<ecosystem>:<name>@<version>", so npm debug and PyPI debug never merge. A node lists its dependency names, dev (true only when every project graph containing it marks it dev), the project ids whose graph contains it (projects), and the project ids that declare it directly (directIn, matched by name). graph.ecosystems gives each ecosystem's completeness (same meaning as surface.graphs) with total and emitted node counts. Only the emitted graph is capped (MAX_EMITTED_GRAPH_NODES, 5,000; directly declared packages are kept first). The graph is built after the policy has run on the full in-memory graphs, so a capped graph never changes a verdict. When the cap bites, graph.truncated is true and the per-ecosystem nodes/emitted counts show how much was left out. No finding or run note is added, so a clean repo with a large graph still gets a clean check. Edges (dependencies) are package names, not node ids, and after truncation an edge can name a package that isn't in the emitted nodes. The field is additive in schema version 1.
  • Transitive impact (#59). impact has one entry per direct dependency: ecosystem, project (path), name, graph (that project's #114 completeness), transitive and exclusive. transitive counts the unique packages in the dependency's lockfile closure, not counting itself. exclusive counts closure packages that no other direct dependency of the project reaches and that aren't declared directly, roughly what removing it would drop. Counts come from lockfile graphs only (ADR 0004). With no graph or no closure entry, both are null (unknown, never 0). On a partial graph, transitive is a lower bound and exclusive is null, since a partial graph can overstate exclusivity. For the same reason, exclusive is null for the whole project when a direct dependency that is a graph node has no closure entry (for example, PyYAML declared and pyyaml locked; nodes are matched with case and -_. folded). A direct dependency that isn't in the graph at all, such as a never-locked Python build backend, doesn't block exclusive. There's one row per name per project, even when the name is declared in several sections. Impact is a fact, never a verdict: it adds no finding and never changes severity, confidence or a check conclusion. The walk has a work budget (MAX_IMPACT_CLOSURE_ENTRIES, 2,000,000 closure entries, applied per project in a stable order). A project past it gets null counts with limited: true, plus one impact-limited run note that is non-capping (findingGroup "note", so the check keeps its conclusion). It's computed for every dependency in PR mode too, and presenters choose what to show. Each entry can also carry footprint (slice B): an approximate install size from the caller's cached registry metadata (AnalyseOptions.metadata), with approximate: true, the provider's basis and coverage (sized package names out of all counted). bytes is a lower bound (#288): each name, the dependency or a closure member, counts at its smallest locked version, and only when every locked version of it is sized. With no provider, a provider failure or timeout (10 s per ecosystem), more than 50,000 package versions in an ecosystem, or no sizes, footprint is absent. That's not a note and not incomplete. See transitive-impact.md. The field is additive in schema version 1.
  • Declaration lines (#198). A dependency may carry declaredLine, the 1-based line in declaredIn where it is declared. Adapters report it; the JavaScript/TypeScript adapter finds each key's line in the dependency sections of the raw package.json text. Right after dependency listing, the engine keeps it only if that manifest line contains the dependency name as a whole token; any other value is dropped, never guessed. Declaration-anchored findings (unused, removed-last-usage, unverified-no-imports, should-be-dev) put it on their declaration evidence (file plus line), so the app can annotate the declaration itself. When any of them only knows the file, one declaration-line-unavailable info note says how many (cap-and-note, #154).
  • Unparsed manifests (#269). When a detected adapter reports detection evidence of kind manifest-malformed, the engine maps it into scan completeness: one manifest-malformed info note per (ecosystem, manifest), with up to 5 evidence lines (file and line), at most 100 notes plus one overflow note. The notes are unmarked, so findingGroup is "incomplete". They cap verdicts like any scan-completeness note (#154) and make the run neutral (#197), so an unparseable manifest never reads as "Findings: none". The neutral conclusion is run-wide: one broken manifest anywhere, including a vendored or fixture one, makes the whole run neutral. This is uniform for every ecosystem, with no adapter contract change. Should-cap gaps don't use the #205 notes channel.
  • Adapter notes (#205). After its graph and usage stages the engine calls an adapter's optional notes(). A note naming a dependency the adapter listed becomes an adapter-capability info finding with that dependency and awareness: true. A note with no dependency becomes an adapter-note run note with adapterNote: true (findingGroup "note"). Both are added after the policy and the marker strip, so they never feed or cap a verdict and never touch reference analysis. Malformed notes and unknown dependencies are dropped. Statements are cleaned to one line and cut to 300 characters, duplicates are merged, and the run keeps 100 notes plus one overflow note. A failing, slow or worker-killing notes() keeps everything analysed before it and adds one incomplete adapter-failure note; the worker posts its outcome just before the notes stage so a hang or heap death there cannot drop it. The worker tier bounds the raw array before it reaches assembly.
  • Capability catalogue and cross-ecosystem overlap (#55). CAPABILITY_CATALOGUE (in capabilities/) lists packages that do the same job, grouped into clusters such as http-client or date-time, per ecosystem. The shape is versioned (CAPABILITY_CATALOGUE_VERSION, now 2). Adding clusters or members doesn't change the version; adding, removing or changing the meaning of a field does. Version 2 (#211) added crossEcosystem: a cluster with crossEcosystem: false (today only test-runner, since pytest next to vitest or jest is normal) is left out of the overlap view but still counts for #58. Python names match per PEP 503. #58 builds same-ecosystem duplicate detection on this data. When one cluster is declared in two or more ecosystems (say axios in a JS app and requests in a Python service), each declared package gets one info finding, rule cross-ecosystem-capability-overlap, naming the packages in the other ecosystems. These are for awareness only (awareness: true, #234) and never a verdict. The engine strips awareness from every adapter and policy finding, so only this core rule can set it. They are added after the policy runs, and each one carries its dependency, so they are never run notes. In a pull request, only packages the PR added are reported.

Local directories

analyseDirectory(path, options) scans a checkout with the repository scanner (repository-scanner.md) and runs analyseRepository over the resulting FsRepositoryHandle. Scan options go in options.scan.

Scan limits that could hide the project's own files become info findings, so a partial scan is never presented as a complete analysis:

  • a truncated scan (max-files, max-directories, max-total-bytes)
  • files or directories skipped as too large, too deep, over-long, unsafely named or unreadable, with counts from skippedCounts and up to 5 example paths

analyseDirectory passes these to analyseRepository as AnalyseOptions.scanCompleteness. Core then appends them, treats the scan as incomplete (scanIncomplete), caps unused and potentially-unnecessary findings at medium confidence with a limitation saying the scan was incomplete, and never counts an ecosystem as reference-analysed, so a skipped file can never produce a confident false "not needed". The GitHub App passes the same fields from its tarball scan (scanIncomplete alone caps without adding the per-path notes). Either way, core adds one run-level info finding saying that every verdict in the result, including should-be-dev and type-only, was computed from a partial scan. Callers never post-process the result (ADR 0004, #154).

Skips that are by design (excluded vendor/generated directories, generated files, symlinks, special files) are not reported. scanCompletenessFindings(scan) is exported for callers that scan on their own.

Pull-request mode (#128)

AnalyseOptions.pullRequestChanges (and the same field on analyseRepositoryIsolated) takes the DependencyChange[] built by the GitHub App from the PR diff: parseUnifiedDiff, listDirectDependencies over base/head manifests, then extractDependencyChanges (#31, wired in #115). It is versioned with @ghostdeps/core, not the adapter contract, so adapterApiVersion is unchanged.

AnalyseOptions.pullRequestSourceChanges (#101) is a separate field for the removed and added source lines, taken from extractDependencyChanges(...).sourceLineChanges (every non-dependency, non-binary file the PR touched, deleted files included). It only applies when pullRequestChanges is set, so full scans never pass PR data to adapters. The engine bounds it first (PR_SOURCE_CHANGE_LIMITS in limits.ts: 5,000 lines per side per file, 2,000 characters per line, 8 Mi characters in total). It also drops unsafe paths and malformed entries, and reports anything dropped in one pr-source-changes-capped info finding. The bounded payload reaches adapters as AdapterContext.pullRequestSourceChanges, including through worker threads. Adapters report matches on removed lines as Usage records with removedInPr: true. To scan removed lines in context, an adapter can rebuild each changed file's base version with reconstructBase(headLines, change) from @ghostdeps/core (#259). It returns undefined when the diff and head don't fit together, and the adapter then records nothing for that file. Both fields are additive, so there is no adapterApiVersion bump (same precedent as Usage.via, #130).

  • Absent: full scan. Policy gets mode: "full" and gives verdicts over every dependency.
  • Present: PR scan. Policy gets mode: "pull-request" plus the changes. It scopes findings to the dependencies those changes touch: unchanged dependencies produce no findings, and findings about removed dependencies are allowed only when the evidence covers them. An empty array means a source-only PR (#101).
  • Coverage. An added or changed dependency in an ecosystem nobody analysed, or one the adapter did not list from that manifest, becomes an info finding. It is never silently dropped.

runRecommendationPolicyContractTests(name, policy) in the contract-test kit checks any policy against policyContractFixtures in both modes: evidence on every verdict, no unused where usage was not analysed, and PR-mode findings limited to touched dependencies.