Which suite proves what, and which failure each tier exists to catch.
The commands themselves — setup, what to run before pushing, how a pull request lands — are CONTRIBUTING.md's and are not repeated here. This page is about why the suites are split the way they are, and what a new test has to do to be worth adding.
Every test needs a case that goes red in the silent direction.
A test that only pins the message text is half a test. It fails when the wording changes and passes when the analyzer stops finding the import at all — which is the failure that matters. Concretely, for a rule test: one case where the violation must be reported, and one near-miss where it must not be, so a rule that stopped matching anything cannot stay green by finding nothing.
That is the same reason the Semgrep rules each carry a ruleid: case as well as
an ok: case. A pattern that stopped matching passes every scan by matching
nothing.
pnpm test # node --test over scripts/ only
moon run ...:lint ...:test ...:typecheck # each package's own targetpnpm test is node --test over scripts/*.test.mjs — the gate scripts, and
nothing else. Those scripts are what make a green build mean something, and a
broken gate reports nothing rather than reporting a failure. They run first in CI
for exactly that reason.
moon run ...:test is each package's own target — Vitest in both packages.
For archkeep that includes the differential against a real
@nx/enforce-module-boundaries; for archkeep-vscode it is every decision the
extension makes, driven as pure functions with no editor running.
Running one and assuming the other passed is the most common local mistake. Run the whole list before pushing; a shorter local run just moves the red to the pull request.
Pure functions over injected data. No filesystem, no mocking library.
That is a consequence rather than a preference: the gate scripts and the
analyzers take their facts as arguments. parseCiTargets gets text, evaluate
gets records, an analyzer gets { sourceFile, text, workspace }. A function that
reads a file and decides something has to be split before it can be tested, and
the split is the improvement.
The one deliberate exception is readMoonProjects, which touches the outside
world and is untested on purpose — a test that stubbed the answer would pin the
stub.
Two rules specific to this codebase:
- An analyzer test that would pass against a hard-coded name→project map is
not a test. Resolution is driven over an in-memory workspace whose
readFilebacks a real fixture tree, sots.resolveModuleNameruns for real.typescript.test.mjsrepoints atsconfig.base.jsonalias without changing the specifier and requires the answer to move with it. - Positions are computed from the fixture, never written as literals.
src/lsp/diagnostics.test.mjsfinds the expected 0-based position by searching the fixture text, so changing the fixture moves both sides and changing the conversion moves only one. A diagnostic one line off is worse than no diagnostic: it sends every reader to the wrong import, confidently.
Property-based cases use fast-check, with the seed pinned on CI —
vitest.property-seed.mjs says why.
Most tests here are examples: this input produces this verdict. Two rarer shapes earn their place when an example per case would enumerate forever:
- Metamorphic / invariant — a relation between inputs: transform the
input in a way that is not an architecture fact, require the same answer.
src/rules/invariants.test.mjsholds the rule engine's (site order and repetition change nothing but the documented output order);src/analysis/metamorphic.test.mjsholds the analyzers' (comments, blank lines and local renames are not dependencies). These catch ordering leaks, comment-mask drift and duplicate handling without a fixture per spelling. - Property-based — generated inputs over a stated total relation (totality, determinism, round-trip). Fast-check where a generator earns its keep; seeded sweeps where a dependency would not.
A relation needs the same guard an example needs, plus one more: prove the
comparison can see a change at all. Every family above carries a negative
control (one real added crossing must move the records) and refuses an empty
baseline before comparing — [] "equal" to [] is the vacuous pass this
repository exists to prevent, and it satisfies every equality silently.
Whether an order is semantic is decided where it is pinned, not assumed.
Import-site order survives evaluateWithSuppressions deliberately (documented
there) and becomes canonical only at sortViolations, which the CLI-level
tests pin; graph snapshots sort projects, edges and tags so no provider's
discovery order can leak; but a project's tags array is semantic — reordering
it changes the snapshot identity history keys transitions on — and so is
array element order in the intent fingerprint. Before asserting
order-invariance on any sequence, check which kind it is.
Several commands answer overlapping questions over the same facts, and most of
them already share one implementation (trajectory reads history's own
snapshots and classifier; report calls health/waivers/fitness
verbatim; delta and change re-judge through the same engine check uses).
Each command's own suite would stay green if a future edit quietly grew a
second, slightly different implementation — two commands answering differently
about the same tree while nothing fails. The consistency suites exist for that
defect: src/commands/cross-command-history.integration.test.mjs drives
history, trajectory, classifyTransition and the raw computeDiff over
ONE synthetic snapshot set with known transitions and requires the counts to
reconcile exactly; src/commands/cross-command-gates.integration.test.mjs
requires diff's edge-constraint impact and delta's full-engine
classification to agree on introduced tag violations (including pinning the
documented narrowing between them as a subset relation), requires every number
in report to be the number its section's command returns over the same
context, law and clock, and requires plan mode's impact to equal impact's.
These prove composition, not equivalence: no claim is made that any two
commands are interchangeable, only that the relationships stated above hold —
and where they deliberately do not (diff's narrower rule impact), the
divergence itself is pinned as a fact rather than left to drift silently.
Real seams, over throwaway fixture trees. Each one exists for a failure the unit tier structurally cannot see.
| file | the failure it is for |
|---|---|
graph/create-dependencies.integration.test.mjs |
Nx changing CreateDependenciesContext. It reaches through index.mjs deliberately — an entry that stopped re-exporting createDependencies would drop every polyglot edge while a test pointed at the implementation stayed green |
config.integration.test.mjs |
A malformed row in this repository's own boundary config, and the proof that the reader takes its filename from a caller — one case asks for a name that does not exist and requires the error to name it back |
config-spelling.integration.test.mjs |
A dialect that "loads" a policy SMALLER than what was written — one fewer constraint row, one option silently defaulted instead of read — or a marker/route/filename spelling that silently diverges in verdict from another spelling of the identical law. Covers path source, filename, workspace-root marker, all four config dialects, and where each dialect's stated-vs-defaulted options and suppressions actually come from; every axis carries a red twin the differential must catch, never only a case where the spellings agree |
cli.integration.test.mjs |
Both halves: the real binary spawned, for the exit-code and usage contract; and check()/runCli() in-process over a fixture Go workspace with only Nx and git injected, so the exact file:line:column a developer acts on is pinned rather than assumed |
lsp.integration.test.mjs |
The server answering over real stdio |
report/sarif.integration.test.mjs |
An upload GitHub silently rejects. It builds one result per messageId from the real message table and asserts the fields a rejection turns on — a ruleId that resolves in the catalogue, a non-empty message, a repository-relative uri with a 1-based startLine/startColumn. A SARIF test that only checks the file parses is not a test |
lsp/editor-config.integration.test.mjs |
A language reaching the analyzer registry and not the Nx integration's manifest — checked by the CLI, never by an editor, which reads exactly like a clean tree |
rules/upstream.integration.test.mjs |
A copied message or option drifting from the installed @nx/eslint-plugin |
analysis/vue.integration.test.mjs |
The line a reader finally sees, from the real analyzer pair, against positions computed from the fixture |
conformance/rule-sdks.integration.test.mjs |
Four rule-authoring SDKs drifting into four laws while all four suites stay green — every SDK's fixture copies held byte-identical, every committed .wasm loaded through the engine's real host at its recorded digest, and one verdict document required from all of them per fixture, which must also be the recorded one. Three of the four SDK suites replay their reference rule as native code, so without this the artifact a consumer runs was proven only by its digest |
conformance/corpus.integration.test.mjs |
A rule that stops firing in Go, Rust or Python, where no upstream engine can disagree — hand-labeled architecture workspaces driven through the whole check path, with each near-miss measured against a forbid-everything policy so its silence cannot be an engine that never looked |
The Vue analyzer is pinned from both sides, and it is the model for anything
with a coordinate conversion in it: vue.test.mjs mocks the TypeScript analyzer
to pin the text handed over — the whole file with everything outside the script
block blanked, so no arithmetic can be wrong — and the integration test drives the
real pair.
scripts/coverage-real-trees.mjs asks the question the ESLint differential
structurally cannot. .go, .rs, .py, .java, .kt and .cs never reach
the upstream rule, so upstream's silence about them is inability rather than a
verdict, and there is nothing to disagree with. This lane needs no oracle: it
clones real public repositories at pinned shas — two each of Go, Rust and
Python, one JVM (.java and .kt together, one package index), one C# —
runs the real analyzers over every tracked source file, and holds three
counts EXACTLY in both directions: files read, import records produced,
failures reported.
Exactness is legitimate because the sha is pinned: the tree cannot move under the harness, so any movement is the harness changing. Fewer records is an analyzer that went quiet on a syntax it used to read — the silent direction, and the reason the lane exists. Fewer failures is a breach too: a failure that stopped being reported is a file that now looks read.
It runs weekly beside the differential, in the same workflow and with the same
posture — not a required check, because it depends on third-party repositories
being reachable, and a red run is a regression rather than a flake. Its tree
table records what each count means, including the gap its first run found and
the fix that followed: the Rust analyzer could not read a top-level
use { a::b, c::d }; brace group, which ripgrep writes 29 times, and reading it
as the list of paths it is turned 29 failures into 72 records.
src/conformance/
is the only thing in the repository that puts this engine's verdict beside real
ESLint's on the same code. Its catalogue sizes — fixture workspaces, probes,
projects, near-misses — live in that directory's README, held to the catalogue
by stated-counts.integration.test.mjs; a copy of the numbers here would drift
with nothing holding it.
Nothing is stubbed on either side. The eight option values and the fifteen
message ids come off the installed rule's own defaultOptions and
meta.messages, so a case states only what it overrides and an Nx upgrade moves
both engines' input at once.
Two mechanics to know before changing anything there, both documented at length
in that directory's README: one workspace root per process (@nx/devkit
resolves it once, on first load, and never again), and same site, not same
column (ESLint reports the whole statement, this engine reports the specifier,
and a pair matches when this engine's position falls inside ESLint's range).
The real-tree half of that comparison lives outside the package:
scripts/differential-real-trees.mjs
drives the same two engines over public Nx workspaces pinned at fixed commits,
from .github/workflows/differential.yml rather than from any test target —
the conformance README's condition 3 states what it measures and why a red run
there is a regression, and its pure halves (ledger matching, the empty-verdict
claim) are what scripts/differential-real-trees.test.mjs pins under
pnpm test.
The differential can only speak where ESLint can read the file, which is
JavaScript, TypeScript and Vue. On .go, .rs, .py, .java, .kt and
.cs — the languages upstream cannot read at all — upstream is silent by
inability, so no comparison there can
catch a rule that stopped firing: every spelling of the import would go quiet
together and still agree.
corpus.mjs and corpus.integration.test.mjs are that half. Architecture
styles built in Go, Rust and Python, each probe carrying the findings a
person decided in advance, run end to end through cli.mjs's check over a
native (archkeep.json) workspace. The mechanism that keeps a near-miss from
being an engine that never looked — a second, forbid-everything policy every
case tree carries — and the three numbers each probe states are documented in
that directory's README rather than here, along with the corpus's own sizes,
which stated-counts.integration.test.mjs holds to the catalogue.
Three cheaper files sit beside those two and check the project against its own
declarations rather than against ESLint: boundary.test.mjs holds the shipped
tool to what it may depend on, stated-counts.integration.test.mjs holds the
README's catalogue sizes to the catalogue, and
plugin-catalogue.integration.test.mjs holds the Claude Code plugin manifests
to each other.
The same script carries a second, native-provider leg over the same pinned
trees — never a second clone, never a second script. It derives a
archkeep.json-equivalent model mechanically from the graph JSON the upstream
leg already fetched (deriveNativeModel), runs nativeProvider.discover/
buildGraph over it, and compares the node set, edge set, and rule verdicts
against that same tree's real Nx-graph-based run — reusing LEDGER and
classifyDifferences with its own "native-extra"/"native-missing"
direction pair rather than a parallel mechanism
(packages/archkeep/src/providers/native/README.md's "What proves this
provider against a tree it was not tested on" has the fuller account,
including what the first real run against code-pushup found and why a
populated ledger there is the expected outcome, not a regression). The tool
run on this tree, described just below, has a native-provider twin too:
.github/workflows/ci.yml's "Check this repository's own module boundaries
(native provider)" step runs the checker against a throwaway copy of this
same tracked tree with nx.json swapped for a tracked
.github/native-selfcheck/archkeep.json, and asserts that copy reads the
byte-identical module-boundaries.config.mjs the Nx-based step just above it
judged — so a disagreement between the two is a provider defect, never two
copies of the law drifting apart.
Three things in CI prove something no unit test can, and they are worth knowing about because a change can break them without breaking a single test:
check-packages— asserts every directory underpackages/is a project the graph can see, declaring at least one CI target. Without it,moon runexits 0 on three different states and only one of them is good: nothing there, something there with no matching target (skipped in silence), or a directory with sources and nomoon.yml(invisible tomoon projects).check-docs-links— fails on any doc reference that cannot resolve: a markdown link whose target file is gone, a#anchorthat names no heading, adocs/…citation pointing at nothing. Prettier formats a broken link and marks it correct; this gate is where "does the link land" is answered.- The tool run on this tree — the final
checkstep, overmodule-boundaries.config.mjsat the root. Everything above it proves the code correct against fixtures it built itself; this is the only step where the enforcer meets real source under a tag vocabulary nothing insrc/knows. verify-package.mjs— packs the real tarball, installs it into a throwaway workspace, and drives what a consumer actually buys. Every other gate runs where the tool's dependencies already exist, which is why none of them could see the state this package was really in once: zero declared dependencies, green suite, and aimport ts from "typescript"that would have thrown on a consumer's disk.
V8 provider, 80% floor on lines, functions, branches and statements. 80 is a floor, not a target — it exists so a change that deletes coverage fails instead of passing quietly. Both suites sit well above it.
The four numbers live in coverage.config.json at the root, and each package's
vitest.config.mjs reads them. They started as literals in the one package that
had tests, with a note saying they would hoist the day a second package earned a
test target — because four numbers copied into two configs are not a hardcode,
they are an unsynced config. archkeep-vscode was that second package. The file
is read rather than imported: JSON has no import, and a relative import from
inside a project up to a root-level file is a violation this repository's own
checker reports.
Three files are deliberately absent from the coverage include, all for one
reason. cli.mjs and lsp.mjs call process.exit and are driven end-to-end
as spawned subprocesses; extension.mjs only runs inside a VS Code extension
host. In-process V8 coverage sees none of the three. That is the exclusion's
reason, not a convenience — and it is why each of them holds wiring only:
everything with a decision in it lives under src/, where coverage can see it.
A missed violation — an import that crosses a boundary and produced no output — is the worst class of bug this project has, and it earns a permanent regression fixture rather than just a fix. Which tier depends on what was silent:
| what was silent | where the fixture goes |
|---|---|
| an import the analyzer did not see | that analyzer's unit tests, plus a line in its limits header if the shape stays unreadable |
| a rule that should have fired | src/conformance/ — cases.mjs in a language ESLint reads, so its verdict is the check; corpus.mjs in Go, Rust or Python, where the label is |
| an edge that never reached the graph | graph/create-dependencies.integration.test.mjs |
| a file the editor never received | lsp/editor-config.integration.test.mjs |
docs/reference/evidence.md owns the verdict vocabulary and its invariants;
this table audits where each incomplete-evidence state is held to them. The
invariant every row answers to is the same: insufficient evidence never reads
as clean, unchanged, zero, or pass. Each row names the modules whose tests pin
the behavior — a state missing from this table, or a row whose tests have
vanished, is exactly the silent hole the matrix exists to prevent.
| evidence failure | what must surface | pinned at |
|---|---|---|
| baseline missing or invalid | refusal naming the file (exit 3) — never an empty diff | each reader's own tests; src/exit-matrix.integration.test.mjs drives every verb against shared fixture worlds |
| provenance one-sided or cross-repository | a disclosure note — or, where the command claims identity (change), a refusal |
diff.test.mjs, delta.test.mjs, change.test.mjs, history.test.mjs |
| a side captured from a dirty tree | a disclosure note from every command over that metadata | same files; the comparator's own booleans pinned in history.test.mjs |
| incomplete graph coverage | refusal (no-verdict, exit 3) from every judging verb; a named index gap in the editor |
analyzer/refusal suites beside each command; lsp/workspace-index.test.mjs |
| file the analyzer could not read | a whole-file failure record (line/column null) — never silence | each analyzer's "misbehaves" tests, Error and non-Error throws alike |
| boundary law malformed | exit 3 naming the law — a broken table must not read as "no constraints" | config.test.mjs and the broken world of the exit matrix |
| manifest unreadable or unparseable | the skip disclosed where coverage is claimed; at CLI level, a not-analyzed file (exit 3) | the Go/Rust/Python analyzer and resolver suites |
| a documented parse limit fires | a spurious record naming text the file really contains — never a missed project | the analyzer's limits header plus the fixture that pins it (python.test.mjs) |
| decisionRef unresolved | unknown — no-verdict, exit 3, never pass |
adr.test.mjs, governance/adr-registry.test.mjs, provenance-command.test.mjs |
| historical observation missing | incomparable, derived numbers null — never folded into unchanged |
history.test.mjs, trajectory.test.mjs |
| external readiness evidence absent | unmeasured, naming what would measure it — neither met nor not met |
scripts/check-readiness.test.mjs |
A new incomplete-evidence state gets a row here in the same change that makes it behave — the table is the audit trail of what "unknown" means across the surface, and prose without a test behind it does not belong in it.