Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
60 changes: 60 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -54,6 +54,66 @@ All notable changes to AgentFlow are documented in this file.
failing at import. The stale-shim hazard is now written into the test's own
design rules, since the next rename will ripple the same way.

* **The gate is measurable on this machine now, not only on Sundays.**
`python scripts/mutation_local.py --module <target>` runs one module of the
same gate locally in about two minutes. It exists because `mutmut run` calls
`sys.exit(1)` at import time on native Windows, so between weekly runs nobody
here could see a score at all — which is part of why nine consecutive red
Sundays went unnoticed. The driver never invokes the `mutmut` CLI: it reads
its targets from `scripts/mutation_report.MODULE_TARGETS`, builds the
workspace with `prepare_workspace`, generates mutants with mutmut's own
`mutate_file_contents`, and runs each one as a plain `pytest` subprocess
selected through `MUTANT_UNDER_TEST`. The gate's definition of a target is
still declared in exactly one place.
* **It reproduces CI rather than approximating it.** Measured against run
34266462154 on `serving/semantic_layer/query/sql_builder.py`: 141 mutants,
the same population, down to the surviving mutant names. At `10e20a5` the
module scores 90.8% (128 killed, 13 survived) — the 88.7% CI last reported
plus the three mutants `2cda8da` and `10e20a5` killed since. Only pytest exit
0 (survived) and 1 (killed) count as verdicts; anything else is reported as a
harness failure and fails the run, where `mutation_report.py` counts exit 3
as a kill. That is the one deliberate divergence, and `CONTRIBUTING.md` says
so rather than claiming exact parity.
* **A score you can trust to be about your own tree.** Three failure modes are
closed by construction: the workspace is stamped with the root, the module,
its source, the materialized package tree, the target's tests and
`pyproject.toml`, and is rebuilt whenever any of those move, so a second run
never reports the first one's sources; a mutant that comes back without a
verdict is retried once serially before it is called a harness failure, so
the number does not drift with machine load; and a `--workspace` that is a
checkout — this repository, anything inside it, or any directory holding a
`.git` — is refused instead of emptied. The mutated module never leaves the
temp workspace: the working tree is clean after a run.

* **`sql_builder.py` is off the threshold line, and its residue is honest.**
It cleared 90% by a single mutant (90.8%, 128 killed of 141), which is not a
margin worth keeping: the next covered line added to the module would have
put the gate back in the red for reasons unrelated to the change. Nine of the
thirteen survivors could never have been killed. Eight mutated a
`typing.cast` type argument — a cast returns its second argument untouched
and never evaluates the first — so both casts are plain annotations now and
the mutants stop existing; the ninth turned `rows = []` into `rows = None` in
a branch whose next statement is `bool(rows)`, and carries a
`# pragma: no mutate` with the reason above it. The remaining four were the
`dialect="duckdb"` argument, and two of them are now dead: DuckDB list
indexing is 1-based where sqlglot's default dialect is not, so
`list_value(1, 2)[1]` read without the dialect comes back out of the scoper
as `[2]` — the tenant scoper would have changed which element the query asked
for while it added a WHERE clause. The module measures 98.4% (124 killed of
126) with `scripts/mutation_local.py` on py3.13.
* **The two mutants still alive are named in the test file, not suppressed.**
`_scope_sql__mutmut_41` and `_43` drop the dialect from the parse of the
relation `_qualify_table` generated itself, and that string has one fixed
shape which — parsed with the dialect or without it — renders identically
under the `sql(dialect="duckdb")` `_scope_sql` applies on the way out, so no
input reaches them with a difference to observe. Not the same as neutral: the
*default-dialect render* of that shape rewrites `EXCLUDE` to `EXCEPT`, which
is why the mutants on the render itself stay killable and dead.
They are not equivalent — a `_qualify_table` that ever emitted
DuckDB-specific syntax would make them killable — so they get a written
record of what was tried and came out identical rather than a pragma that
would outlive its reason.

### Terraform — an exact core pin took the provider update channel down with it

* **`required_version = "= 1.15.4"` broke Dependabot's terraform ecosystem the
Expand Down
37 changes: 37 additions & 0 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -148,6 +148,43 @@ replaceable runtime artifacts, not reviewed evidence or production acceptance.
Promote a reviewed snapshot only under a new date-stamped identity with
provenance.

`python scripts/mutation_local.py --module <target>` measures one module of
that same gate on this machine, so a mutation score is available before you
push instead of only after the weekly workflow. `--list-modules` prints the
targets; the score, the mutant population and the surviving mutant names match
the CI run for the same commit — up to pytest's internal-error exits, which
`scripts/mutation_report.py` counts as kills (exit 3) while this driver refuses
to score them at all — because the driver reads its targets from
`scripts/mutation_report.MODULE_TARGETS`, builds its workspace with
`prepare_workspace`, and generates mutants with mutmut's own engine
(`mutmut.mutation.file_mutation`). It needs `mutmut` installed (it is in the
`dev` extra) but never invokes the `mutmut` CLI: mutants are executed as plain
`pytest` subprocesses selected through `MUTANT_UNDER_TEST`, which is what makes
this work on native Windows, where `mutmut run` exits at import time. Expect
minutes, not seconds — one pytest process per mutant, `--jobs` in parallel.
The driver exits 1 when the module is below its threshold or when any mutant
got no verdict (only pytest exit 0 = survived and 1 = killed are verdicts;
anything else is a harness failure, never a kill). A mutant that comes back
without a verdict — usually a timeout from running `--jobs` of them at once —
is retried once serially with a longer timeout before it is reported that way,
so the score does not move with the machine's load. Its workspace and JSON
report live under the OS temp directory, outside the repository; the workspace
is reused across runs only when it is stamped with the same root, module,
module source, materialized top-level package tree, target tests and
`pyproject.toml` (which `prepare_workspace` always renders into the workspace as
a real file, carrying pytest addopts, filterwarnings and `[tool.mutmut]`), and
is rebuilt otherwise, so a second run never reports the first one's copy of
those sources. The stamp does not cover the trees `prepare_workspace` normally
symlinks — `src/`, `sdk/`, `config/`, `scripts/` and the rest of `tests/`
beyond the target's own test files — which it copies instead where the OS
refuses symlinks; on such a machine, pass a fresh `--workspace` after editing
them. A `--workspace` is emptied on rebuild,
so one that is a checkout — this repository, anything inside the tree the
sources come from, or any directory holding a `.git` — is refused instead, and
only a directory carrying the driver's own marker is ever cleared. `--root`
points it at another checkout; `--only` re-runs named mutants, written either
bare or exactly as the report prints them (`<module>.<mutant>`).

`python scripts/evaluate_trivy_policy.py` writes ignored Trivy policy
summaries under `.artifacts/trivy/`. Relative `--report`, `--waivers`, and
`--output` paths resolve from the project root, not the caller CWD, and every
Expand Down
Loading