Skip to content

test(mutation): green the mutation gate, make it measurable locally, and take sql_builder off its threshold line - #255

Merged
brownjuly2003-code merged 6 commits into
mainfrom
test/sql-builder-foreign-tenant-mutants
Sep 10, 2026
Merged

brownjuly2003-code merged 6 commits into
mainfrom
test/sql-builder-foreign-tenant-mutants

Conversation

@brownjuly2003-code

Copy link
Copy Markdown
Owner

What this is

Six commits that take the weekly Mutation Testing gate from "red for nine consecutive Sundays, and unverifiable until the next Sunday" to green, measurable on a laptop in about two minutes, and honest about what it still leaves alive.

3820a2f (already on main) fixed the two causes of the redness. This branch is the follow-through: it pins the survivors the gate exposed, gives the gate a local driver, and takes the one module that was clearing its threshold by a single mutant off that line.

The commits

Commit What it does
6bed38f Pins the tenant guard the gate had found unpinned — the mutants that survived were in the code that decides whether a table holds another tenant's rows.
2cda8da Pins the retry rule that sat exactly on its threshold, so an off-by-one in the backoff comparison is now caught.
10e20a5 Makes the tenant guard's tests state what they refused over, not just that they refused.
02b7506 Stops a Docker-only assertion in the 4h soak verifier from failing the whole local suite on a machine with no Docker.
25769d7 Adds scripts/mutation_local.py: the mutation gate, runnable per module against a private workspace, so a survivor can be reproduced now instead of on Sunday.
49fe3f9 Takes sql_builder.py off the threshold line: 90.8% (128/141) → 98.4% (124/126).

The last one in more detail

sql_builder.py cleared the 90% threshold by one mutant. Of its thirteen survivors, nine could never have been killed:

  • Eight mutated the type argument of a typing.cast. A cast returns its second argument untouched and never evaluates the first, so no test can observe the difference. Both casts are plain annotations now — identical to mypy, and the mutants stop being generated. The cast calls' own six killable mutants went with them, which is why the denominator drops to 126.
  • One turned rows = [] into rows = None in a branch whose next statement is bool(rows). Equivalent, and marked # pragma: no mutate in place with the reason written above it.

Two of the four dialect="duckdb" survivors are now dead, killed by a behaviour statement rather than an assertion about sqlglot's internals: DuckDB indexes lists from 1 where sqlglot's default dialect does not, so list_value(1, 2)[1] parsed without the dialect leaves the scoper as [2] — the tenant scoper would have changed which element the query asked for while it added a WHERE clause.

_scope_sql__mutmut_41 and _43 stay alive on purpose, named in the test file with what was tried against them. They drop the dialect from the parse of the relation _qualify_table generates itself, and that one fixed shape parses to an AST that renders identically under the sql(dialect="duckdb") applied on the way out. That is not the same as dialect-neutral: the default-dialect render of that shape rewrites EXCLUDE to EXCEPT — which is exactly why the render-side dialect mutants are dead and must stay pinned. An honest survivor with its evidence beats a pragma that would outlive its reason.

Verification

  • .venv/Scripts/python.exe scripts/mutation_local.py --module serving/semantic_layer/query/sql_builder.py → 126 generated, 124 killed, 2 survived, 98.4% (threshold 90%), py3.13 / sqlglot 30.12.0. The run's stamp covers the bytes that ship.
  • Full local gate on the final tree: ruff format, ruff check, mypy, contract tests and the whole unit suite — green.
  • The first real end-to-end verification of the scheduled gate is the Sunday 2026-09-13 04:00 UTC run.

🤖 Generated with Claude Code

https://claude.ai/code/session_018GqSmXR1bp1xPTZdsdjgTE

JuliaEdom and others added 6 commits September 8, 2026 15:00
The first mutation run after 3820a2f repaired the import shim (run
34265359911, 2026-09-08) scored sql_builder.py at 80.3% against a 90%
threshold -- killed 106, survived 26. That is not a regression from the
repair; it is the first honest measurement in nine weeks. The shim had been
broken since 1096e2e, so the module scored n/a and the gate's complaint was
about the harness, not the tests.

14 of the 26 survivors were in `_holds_foreign_tenant_rows`, which had no
tests. It is the fail-closed probe that decides whether a request carrying no
tenant context may read a table at all (audit p2_1 #5): if the table holds a
row belonging to anyone but DEFAULT_TENANT, the read is refused. The only
path any test reached was the `_backend is None` early return the host
doubles fell into, so the probe SQL, the per-table cache and the refusal
branch in `_qualify_table` were all unexercised -- a guard against
cross-tenant reads with nothing holding it in place.

Thirteen tests, through a `_Backend` double that records the SQL rather than
only replaying a verdict. The probe text is the check: a mutant that widens
`<>` to `=`, drops `LIMIT 1`, or asks about a tenant other than the default
still returns a truthy row, so a test reading only the boolean would call all
three correct. Also pinned: the store's own "cannot read that" means an
unmaterialized table and stays permissive, while an unexpected failure is not
laundered into permission; the cache answers without probing, is written
once, and is keyed per table; and `_qualify_table` refuses the unscoped read
of a multi-tenant table, allows it on a single-tenant one, and does not probe
at all when a tenant is in context.

The docstring's "96.0%, the 7 survivors are equivalent mutants" paragraph is
replaced. It was measured on a py3.10 harness before the shim broke, and it
read as reassurance about a mutant population that no longer existed.

54 passed (was 41), and 54 passed again inside a rebuilt mutmut workspace --
top-level `serving`, no `src` -- which is the only shape the shims run in.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015HsJh4f7Lqv3XmkCPSA1uF
Run 34266462154 scored `sdk/agentflow/retry.py` at 75.0% against a 75%
threshold: 15 killed, 5 survived, and all five survivors were in
`is_retryable_method`. A module sitting exactly on its floor is not passing --
it fails the gate the moment anything shifts, and it says nothing about whether
the rule underneath is held in place.

`is_retryable_method` decides whether a request is replayed after a 429 or a
5xx. The only branch any test reached was the verb allowlist, and even that was
partial: `OPTIONS` was in the set and in nobody's assertion. The whole
POST-with-an-Idempotency-Key branch -- the one that decides whether a *write*
is retried -- had no test at all. Mutants that dropped the case-folding, that
matched the header name against a pair's value instead of its key, or that
extended the rescue to other non-idempotent verbs therefore all survived.

Eleven tests, one per decision the function makes: the missing `OPTIONS` verb,
case normalisation of the verb itself, the key found in a mapping and in a
header sequence, matched case-insensitively, found among unrelated headers,
rejected on a merely similar header name (`Idempotency`), on no headers and on
an empty mapping, matched on the name and not on the value, and ignored
entirely for PATCH.

The tests widen coverage into a branch that was previously invisible to the
gate, so mutmut now generates 30 mutants for this module where it generated 20.
All 30 are killed: 75.0% -> 100%. Measured by driving mutmut's own mutation
engine (`mutate_file_contents` plus the `MUTANT_UNDER_TEST` trampoline) under
pytest, which is the only way to run it on this machine; the same driver
reproduces run 34266462154's sql_builder figures exactly, mutant name for
mutant name.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvU8wMbmkbJ6JhapXomgd2
…over

Three of the sixteen mutants that survived run 34266462154 lived in assertions
that were satisfied by the wrong thing:

- `_qualify_table`'s refusal of an unscoped read was proved by the exception
  alone. A mutant that probes a different table still finds a row and still
  raises, so `pytest.raises` passed while the guard asked the wrong question.
  The test now pins the SQL the backend actually saw.
- The recursive-CTE refusal matched only the first half of its message. The
  table it refused over is the half an operator needs when they read the 503,
  and it is also what stops a mutant that reports `['ORDERS']` from looking
  correct. The `match=` now covers the rendered name.
- Nothing exercised a recursive CTE that shadows *nothing*. A guard that
  refused every `WITH RECURSIVE` and one that refused only the dangerous ones
  were indistinguishable; the new test separates them.

sql_builder.py: 88.7% -> 90.8% (128 of 141 killed), measured locally against
the same mutant population CI generates. That clears the 90% threshold by one
mutant, which is not where this module should stay -- the thirteen remaining
survivors are a separate piece of work.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CvU8wMbmkbJ6JhapXomgd2
… suite

`test_merged_soak_compose_overrides_api_healthcheck_for_background_consumers`
merges the three compose files by shelling out to `docker compose config`.
There was no availability guard, so on a machine without the Docker CLI the
call raises `FileNotFoundError: [WinError 2]` and the test fails rather than
skipping — one red in 3953 tests, enough to fail every local full-suite run
and hide anything that goes red after it.

The repository already declares `requires_docker` ("marks tests that require
local Docker") for exactly this, but the unit lane selects tests by path and
never deselects by marker, so the marker alone skips nothing. The test now
carries both: the marker for the vocabulary, and a `skipif` on
`shutil.which("docker")` that actually takes effect. CI installs Docker, so the
assertion still runs where the compose contract is worth checking.

Verified on this machine: the file goes from `1 failed, 7 passed` to `7 passed,
1 skipped`, with the skip reason naming the missing CLI.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FHfk1kcS9om9nHXW32tyfo
… week

`.github/workflows/mutation.yml` runs on Sundays and on dispatch, and it was
the only way anyone here could see a mutation score: `mutmut run` calls
`sys.exit(1)` at import time on native Windows, and this machine has no WSL.
The gate was red for nine consecutive Sundays partly because nobody could see a
number between pushes.

`scripts/mutation_local.py --module <target>` measures one module of that same
gate in about two minutes, and never invokes the `mutmut` CLI. It reads its
targets from `scripts/mutation_report.MODULE_TARGETS` and builds its workspace
with `prepare_workspace`, so the gate's definition of a target still lives in
one place; it generates mutants with mutmut's own `mutate_file_contents`, so
the population is the engine's, not an imitation; and it runs each mutant as a
plain `pytest` subprocess selected through `MUTANT_UNDER_TEST`, which is what
sidesteps the Windows guard.

Four mechanics carry it. Coverage is measured first, because
`mutate_only_covered_lines` makes the population coverage-dependent and getting
it wrong drifts the mutant numbering away from CI's. A generated
`sitecustomize.py` pre-registers a stub `mutmut.__main__` for the child
processes, without which every mutant dies during collection and scores a
silent, meaningless 100%. The symlinked package is materialized before the
module is mutated, so a mutated source never lands in the working tree. And
each mutant gets a private `--basetemp`, namespaced per invocation, because
pytest wipes and recreates that directory at startup.

The number is only worth having if it is about the tree in front of you:
the workspace is stamped with the root, the module, its source, the
materialized package, the target's tests and `pyproject.toml`, and rebuilt
whenever any of those move; a mutant that returns without a verdict is retried
once serially before it is called a harness failure, so the score does not
track machine load; and a `--workspace` that is a checkout — this repository,
anything inside it, or any directory holding a `.git` — is refused rather than
emptied.

Verified against CI, not just under pytest: at 10e20a5,
`serving/semantic_layer/query/sql_builder.py` generates 141 mutants, the same
population as CI run 34266462154 down to the surviving names, and scores 90.8%
(128 killed, 13 survived, none without a verdict) in 124s. The 88.7% CI last
reported plus the three mutants killed since by 2cda8da and 10e20a5. The
working tree was clean afterwards.

Only pytest exit 0 (survived) and 1 (killed) count as verdicts; anything else
is a harness failure and fails the run, where `mutation_report.py` counts exit
3 as a kill. That is the one deliberate divergence from CI, and CONTRIBUTING
says so rather than claiming exact parity.

43 unit tests cover the driver's own logic with subprocess and mutmut stubbed —
no real mutation run in the suite, which would be minutes long.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FHfk1kcS9om9nHXW32tyfo
…t stays alive

The module cleared the 90% mutation threshold by a single mutant (128 killed of
141 = 90.8%), which is not a margin worth keeping: the next covered line added
to it would have put the gate back in the red for reasons unrelated to the
change. Thirteen mutants survived; nine of them could never have been killed.

Eight mutated the type argument of a `typing.cast` — a cast returns its second
argument untouched and never evaluates the first, so no test can observe the
difference. Both casts are plain annotations now, which says the same thing to
mypy and leaves nothing to mutate; the cast calls' own six killable mutants went
with them, so the denominator shrank to 126. The ninth turned `rows = []` into
`rows = None` in a branch whose next statement is `bool(rows)` — equivalent, and
marked `# pragma: no mutate` in place, with the reason above it. Nothing was
silenced that was not first shown to be equivalent, and the threshold in
scripts/mutation_report.py is untouched.

The remaining four were the `dialect="duckdb"` argument, and two of them are now
dead: DuckDB indexes lists from 1 where sqlglot's default dialect does not, so
`list_value(1, 2)[1]` parsed without the dialect leaves the scoper as `[2]` —
the tenant scoper would have changed which element the query asked for while it
added a WHERE clause. The new test states that as a property of the query, not
of sqlglot.

`_scope_sql__mutmut_41` and `_43` stay alive and are named in the test file with
what was tried against them, rather than pragma'd: they drop the dialect from
the parse of the relation `_qualify_table` generates itself, whose one fixed
shape parses to an AST that renders identically under the `sql(dialect="duckdb")`
applied on the way out. That is not the same as dialect-neutral — the
default-dialect *render* of that shape rewrites `EXCLUDE` to `EXCEPT`, which is
exactly why the render-side mutants are dead and must stay pinned.

Measured with scripts/mutation_local.py (py3.13, sqlglot 30.12.0) against the
bytes that ship: 126 generated, 124 killed, 2 survived, 98.4%.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018GqSmXR1bp1xPTZdsdjgTE
@github-actions

github-actions Bot commented Sep 9, 2026

Copy link
Copy Markdown

DORA Metrics

  • Window: last 30 days
  • Branch: main
  • Deployment frequency: 19 total / 4.43 per week
  • Lead time for changes: avg 0.59h / median 0.0h
  • Change failure rate: 73.68% (14/19)
  • MTTR: 79.52h across 11 incident(s)

@brownjuly2003-code
brownjuly2003-code merged commit 127be28 into main Sep 10, 2026
32 checks passed
@brownjuly2003-code
brownjuly2003-code deleted the test/sql-builder-foreign-tenant-mutants branch September 10, 2026 06:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants