test(mutation): green the mutation gate, make it measurable locally, and take sql_builder off its threshold line - #255
Merged
brownjuly2003-code merged 6 commits intoSep 10, 2026
Conversation
The first mutation run after 3820a2f repaired the import shim (run 34265359911, 2026-09-08) scored sql_builder.py at 80.3% against a 90% threshold -- killed 106, survived 26. That is not a regression from the repair; it is the first honest measurement in nine weeks. The shim had been broken since 1096e2e, so the module scored n/a and the gate's complaint was about the harness, not the tests. 14 of the 26 survivors were in `_holds_foreign_tenant_rows`, which had no tests. It is the fail-closed probe that decides whether a request carrying no tenant context may read a table at all (audit p2_1 #5): if the table holds a row belonging to anyone but DEFAULT_TENANT, the read is refused. The only path any test reached was the `_backend is None` early return the host doubles fell into, so the probe SQL, the per-table cache and the refusal branch in `_qualify_table` were all unexercised -- a guard against cross-tenant reads with nothing holding it in place. Thirteen tests, through a `_Backend` double that records the SQL rather than only replaying a verdict. The probe text is the check: a mutant that widens `<>` to `=`, drops `LIMIT 1`, or asks about a tenant other than the default still returns a truthy row, so a test reading only the boolean would call all three correct. Also pinned: the store's own "cannot read that" means an unmaterialized table and stays permissive, while an unexpected failure is not laundered into permission; the cache answers without probing, is written once, and is keyed per table; and `_qualify_table` refuses the unscoped read of a multi-tenant table, allows it on a single-tenant one, and does not probe at all when a tenant is in context. The docstring's "96.0%, the 7 survivors are equivalent mutants" paragraph is replaced. It was measured on a py3.10 harness before the shim broke, and it read as reassurance about a mutant population that no longer existed. 54 passed (was 41), and 54 passed again inside a rebuilt mutmut workspace -- top-level `serving`, no `src` -- which is the only shape the shims run in. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015HsJh4f7Lqv3XmkCPSA1uF
Run 34266462154 scored `sdk/agentflow/retry.py` at 75.0% against a 75% threshold: 15 killed, 5 survived, and all five survivors were in `is_retryable_method`. A module sitting exactly on its floor is not passing -- it fails the gate the moment anything shifts, and it says nothing about whether the rule underneath is held in place. `is_retryable_method` decides whether a request is replayed after a 429 or a 5xx. The only branch any test reached was the verb allowlist, and even that was partial: `OPTIONS` was in the set and in nobody's assertion. The whole POST-with-an-Idempotency-Key branch -- the one that decides whether a *write* is retried -- had no test at all. Mutants that dropped the case-folding, that matched the header name against a pair's value instead of its key, or that extended the rescue to other non-idempotent verbs therefore all survived. Eleven tests, one per decision the function makes: the missing `OPTIONS` verb, case normalisation of the verb itself, the key found in a mapping and in a header sequence, matched case-insensitively, found among unrelated headers, rejected on a merely similar header name (`Idempotency`), on no headers and on an empty mapping, matched on the name and not on the value, and ignored entirely for PATCH. The tests widen coverage into a branch that was previously invisible to the gate, so mutmut now generates 30 mutants for this module where it generated 20. All 30 are killed: 75.0% -> 100%. Measured by driving mutmut's own mutation engine (`mutate_file_contents` plus the `MUTANT_UNDER_TEST` trampoline) under pytest, which is the only way to run it on this machine; the same driver reproduces run 34266462154's sql_builder figures exactly, mutant name for mutant name. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CvU8wMbmkbJ6JhapXomgd2
…over Three of the sixteen mutants that survived run 34266462154 lived in assertions that were satisfied by the wrong thing: - `_qualify_table`'s refusal of an unscoped read was proved by the exception alone. A mutant that probes a different table still finds a row and still raises, so `pytest.raises` passed while the guard asked the wrong question. The test now pins the SQL the backend actually saw. - The recursive-CTE refusal matched only the first half of its message. The table it refused over is the half an operator needs when they read the 503, and it is also what stops a mutant that reports `['ORDERS']` from looking correct. The `match=` now covers the rendered name. - Nothing exercised a recursive CTE that shadows *nothing*. A guard that refused every `WITH RECURSIVE` and one that refused only the dangerous ones were indistinguishable; the new test separates them. sql_builder.py: 88.7% -> 90.8% (128 of 141 killed), measured locally against the same mutant population CI generates. That clears the 90% threshold by one mutant, which is not where this module should stay -- the thirteen remaining survivors are a separate piece of work. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CvU8wMbmkbJ6JhapXomgd2
… suite
`test_merged_soak_compose_overrides_api_healthcheck_for_background_consumers`
merges the three compose files by shelling out to `docker compose config`.
There was no availability guard, so on a machine without the Docker CLI the
call raises `FileNotFoundError: [WinError 2]` and the test fails rather than
skipping — one red in 3953 tests, enough to fail every local full-suite run
and hide anything that goes red after it.
The repository already declares `requires_docker` ("marks tests that require
local Docker") for exactly this, but the unit lane selects tests by path and
never deselects by marker, so the marker alone skips nothing. The test now
carries both: the marker for the vocabulary, and a `skipif` on
`shutil.which("docker")` that actually takes effect. CI installs Docker, so the
assertion still runs where the compose contract is worth checking.
Verified on this machine: the file goes from `1 failed, 7 passed` to `7 passed,
1 skipped`, with the skip reason naming the missing CLI.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FHfk1kcS9om9nHXW32tyfo
… week `.github/workflows/mutation.yml` runs on Sundays and on dispatch, and it was the only way anyone here could see a mutation score: `mutmut run` calls `sys.exit(1)` at import time on native Windows, and this machine has no WSL. The gate was red for nine consecutive Sundays partly because nobody could see a number between pushes. `scripts/mutation_local.py --module <target>` measures one module of that same gate in about two minutes, and never invokes the `mutmut` CLI. It reads its targets from `scripts/mutation_report.MODULE_TARGETS` and builds its workspace with `prepare_workspace`, so the gate's definition of a target still lives in one place; it generates mutants with mutmut's own `mutate_file_contents`, so the population is the engine's, not an imitation; and it runs each mutant as a plain `pytest` subprocess selected through `MUTANT_UNDER_TEST`, which is what sidesteps the Windows guard. Four mechanics carry it. Coverage is measured first, because `mutate_only_covered_lines` makes the population coverage-dependent and getting it wrong drifts the mutant numbering away from CI's. A generated `sitecustomize.py` pre-registers a stub `mutmut.__main__` for the child processes, without which every mutant dies during collection and scores a silent, meaningless 100%. The symlinked package is materialized before the module is mutated, so a mutated source never lands in the working tree. And each mutant gets a private `--basetemp`, namespaced per invocation, because pytest wipes and recreates that directory at startup. The number is only worth having if it is about the tree in front of you: the workspace is stamped with the root, the module, its source, the materialized package, the target's tests and `pyproject.toml`, and rebuilt whenever any of those move; a mutant that returns without a verdict is retried once serially before it is called a harness failure, so the score does not track machine load; and a `--workspace` that is a checkout — this repository, anything inside it, or any directory holding a `.git` — is refused rather than emptied. Verified against CI, not just under pytest: at 10e20a5, `serving/semantic_layer/query/sql_builder.py` generates 141 mutants, the same population as CI run 34266462154 down to the surviving names, and scores 90.8% (128 killed, 13 survived, none without a verdict) in 124s. The 88.7% CI last reported plus the three mutants killed since by 2cda8da and 10e20a5. The working tree was clean afterwards. Only pytest exit 0 (survived) and 1 (killed) count as verdicts; anything else is a harness failure and fails the run, where `mutation_report.py` counts exit 3 as a kill. That is the one deliberate divergence from CI, and CONTRIBUTING says so rather than claiming exact parity. 43 unit tests cover the driver's own logic with subprocess and mutmut stubbed — no real mutation run in the suite, which would be minutes long. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FHfk1kcS9om9nHXW32tyfo
…t stays alive The module cleared the 90% mutation threshold by a single mutant (128 killed of 141 = 90.8%), which is not a margin worth keeping: the next covered line added to it would have put the gate back in the red for reasons unrelated to the change. Thirteen mutants survived; nine of them could never have been killed. Eight mutated the type argument of a `typing.cast` — a cast returns its second argument untouched and never evaluates the first, so no test can observe the difference. Both casts are plain annotations now, which says the same thing to mypy and leaves nothing to mutate; the cast calls' own six killable mutants went with them, so the denominator shrank to 126. The ninth turned `rows = []` into `rows = None` in a branch whose next statement is `bool(rows)` — equivalent, and marked `# pragma: no mutate` in place, with the reason above it. Nothing was silenced that was not first shown to be equivalent, and the threshold in scripts/mutation_report.py is untouched. The remaining four were the `dialect="duckdb"` argument, and two of them are now dead: DuckDB indexes lists from 1 where sqlglot's default dialect does not, so `list_value(1, 2)[1]` parsed without the dialect leaves the scoper as `[2]` — the tenant scoper would have changed which element the query asked for while it added a WHERE clause. The new test states that as a property of the query, not of sqlglot. `_scope_sql__mutmut_41` and `_43` stay alive and are named in the test file with what was tried against them, rather than pragma'd: they drop the dialect from the parse of the relation `_qualify_table` generates itself, whose one fixed shape parses to an AST that renders identically under the `sql(dialect="duckdb")` applied on the way out. That is not the same as dialect-neutral — the default-dialect *render* of that shape rewrites `EXCLUDE` to `EXCEPT`, which is exactly why the render-side mutants are dead and must stay pinned. Measured with scripts/mutation_local.py (py3.13, sqlglot 30.12.0) against the bytes that ship: 126 generated, 124 killed, 2 survived, 98.4%. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018GqSmXR1bp1xPTZdsdjgTE
DORA Metrics
|
brownjuly2003-code
deleted the
test/sql-builder-foreign-tenant-mutants
branch
September 10, 2026 06:35
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this is
Six commits that take the weekly Mutation Testing gate from "red for nine consecutive Sundays, and unverifiable until the next Sunday" to green, measurable on a laptop in about two minutes, and honest about what it still leaves alive.
3820a2f(already onmain) fixed the two causes of the redness. This branch is the follow-through: it pins the survivors the gate exposed, gives the gate a local driver, and takes the one module that was clearing its threshold by a single mutant off that line.The commits
6bed38f2cda8da10e20a502b750625769d7scripts/mutation_local.py: the mutation gate, runnable per module against a private workspace, so a survivor can be reproduced now instead of on Sunday.49fe3f9sql_builder.pyoff the threshold line: 90.8% (128/141) → 98.4% (124/126).The last one in more detail
sql_builder.pycleared the 90% threshold by one mutant. Of its thirteen survivors, nine could never have been killed:typing.cast. A cast returns its second argument untouched and never evaluates the first, so no test can observe the difference. Both casts are plain annotations now — identical to mypy, and the mutants stop being generated. The cast calls' own six killable mutants went with them, which is why the denominator drops to 126.rows = []intorows = Nonein a branch whose next statement isbool(rows). Equivalent, and marked# pragma: no mutatein place with the reason written above it.Two of the four
dialect="duckdb"survivors are now dead, killed by a behaviour statement rather than an assertion about sqlglot's internals: DuckDB indexes lists from 1 where sqlglot's default dialect does not, solist_value(1, 2)[1]parsed without the dialect leaves the scoper as[2]— the tenant scoper would have changed which element the query asked for while it added a WHERE clause._scope_sql__mutmut_41and_43stay alive on purpose, named in the test file with what was tried against them. They drop the dialect from the parse of the relation_qualify_tablegenerates itself, and that one fixed shape parses to an AST that renders identically under thesql(dialect="duckdb")applied on the way out. That is not the same as dialect-neutral: the default-dialect render of that shape rewritesEXCLUDEtoEXCEPT— which is exactly why the render-side dialect mutants are dead and must stay pinned. An honest survivor with its evidence beats a pragma that would outlive its reason.Verification
.venv/Scripts/python.exe scripts/mutation_local.py --module serving/semantic_layer/query/sql_builder.py→ 126 generated, 124 killed, 2 survived, 98.4% (threshold 90%), py3.13 / sqlglot 30.12.0. The run's stamp covers the bytes that ship.🤖 Generated with Claude Code
https://claude.ai/code/session_018GqSmXR1bp1xPTZdsdjgTE