Skip to content

Close the remaining Bölüm C items, and say what the parity gate proves - #18

Merged
emrefbulut merged 8 commits into
mainfrom
fix/gate-semantics-and-provenance
Sep 14, 2026
Merged

emrefbulut merged 8 commits into
mainfrom
fix/gate-semantics-and-provenance

Conversation

@emrefbulut

Copy link
Copy Markdown
Owner

Seven commits, each its own change. Closes the remaining Bölüm C items and the two clarifications asked for on top of them.

1. What PARITY_GATE_PASSED actually asserts (4c6ff03)

The gate was documented in two lines, neither of which said what it checks. A reader could not tell "ran with the same settings" from "produced the same numbers". It is the stronger one, and that is now written down in SPEC §5.10.1 and methodology §8.

compare_cell compares test_accuracy, train_accuracy, train_windows, test_windows by exact equality, row by row, matched by (strategy, split seed, train seed) rather than by position — on top of the run count and the seed-pair set. Demonstrated against the real rows of leakage_real_stride_runs.json (stride 1024, 30 runs):

difference introduced result
one test_accuracy moved by 1e-12 caught
one train_accuracy changed caught
one train_windows off by one caught
grid cut to the first seed pair, every value exact caught (run count)
right count, shifted seeds caught (seed-pair set)
different environment, same numbers not caught
different stride recorded, same numbers not caught

Both exclusions are now documented rather than left to be discovered. stride/noise_sigma/snr_db select the cell rather than being measured by it — a wrong stride changes the window counts, and those are compared.

2. A library upgrade that did not move the numbers (same commit)

The environment exclusion is load-bearing, not missing: because the gate ignores it, a re-run on a different stack is a real difference and equality across it is a result.

artifacts/leakage_real_stride_runs.json was produced 2026-08-11. Re-measured 2026-09-13 under sigmf 1.13.0, torch 2.13.0+cpu, numpy 2.5.1, scipy 1.18.0 — three cells, 30 runs each — bit-identical on all four compared fields.

On the provenance, stated as an inference rather than a measurement: the artifact carries environment: null, so nothing in it records which sigmf produced it. The version tripwire in tests/test_io.py records each release as first encountered and did not record 1.12.0 until 2026-08-19, eight days later. That puts the original run on 1.11.1. methodology says so in exactly those terms — the gap is precisely what prompted environment stamping.

Written at the width it supports: one table, three cells, one direction of upgrade, on CPU.

3. --force now refuses the two categories that are not judgements (520af35)

Categories 2–5 are inferences about what a pattern means; a user who knows their recordings can be right where the inference is wrong. Categories 1 (reader cannot open the files) and 6 (build would refuse the split) say no measurement can be constructed. Those stay REFUSED, exit 1, and the report prints forced refused. … rather than merely omitting the FORCED header.

4. The report says whether it trained (46aa665)

started no. This version of the command stops before training was printed unconditionally, and on a machine with torch the measurement followed four lines later. It is now derived: yes. Both arms train after this block…, or no. with the actual obstacle. Module docstring, default reason and WorkEstimate docstring corrected; JSON gains trains and no_train_reason. CHANGELOG's "No training CLI yet." and "This version does not train." rewritten.

One mutation initially passed here — on this machine every run trains, so yes was right by accident. The torch-free branch is now forced open by patching _torch_available.

5. measurement_schema in the sweep payload (b2f5a3a) — closes #3

The sweep path emitted no schema and hand-rolled _run_row's nine fields. A sweep payload written today was indistinguishable from a pre-schema one.

6–7. CHANGELOG, ROADMAP, category numbering (7bfc03d, 8bb6f8f, ae0d750)

  • Three shipped changes had no CHANGELOG entry — including the seed-pair restoration, whose commit touched seven files and not this one.
  • The ROADMAP grouping item was unchecked while its own body said --group-by collection is shipped. Split into the half that shipped and the half that has not started.
  • methodology §6 never used the word "category" while the command cites it by number. It now carries a cross-reference table — and states that category 5 and §6.5 are opposite things: a refusal, and the one dataset that was not eliminated.

Verification

ruff check .              All checks passed!
ruff format --check .     82 files already formatted
pytest -q -rs             385 passed, 1 skipped

Every behaviour change mutation-verified; the mutations and which test caught each are in the individual commit messages.

Still open, not in this PR

The parity gate's LoRaIQ table is running now. Three of four tables have passed; "Faz 5 doğrulandı" is not claimed until the fourth does.

🤖 Generated with Claude Code

emrefbulut and others added 8 commits September 14, 2026 00:34
The gate was documented in two lines, neither of which said what it checks.
"The published tables and the parity gate stay on CPU" is a policy, not a
contract, and a reader had no way to tell whether a pass meant "ran with the
same settings" or "produced the same numbers". Those are different claims and
only one of them is true of this gate.

It is the stronger one. compare_cell compares test_accuracy, train_accuracy,
train_windows and test_windows by exact equality, row by row, matched by
(strategy, split seed, train seed) rather than by position -- on top of the
run count and the seed-pair set. Demonstrated against the real rows of
leakage_real_stride_runs.json: one test_accuracy moved by 1e-12 fails the
cell.

What it does not compare is now written down too, because an undocumented
exclusion reads as coverage. It ignores the environment block, and it ignores
stride / noise_sigma / snr_db, which select the cell rather than being
measured by it -- a wrong stride changes the window counts, and those are
compared.

The environment exclusion is load-bearing rather than missing. Because the
gate does not compare environments, a re-run on a different library stack is
a real difference between the two runs, and equality across it is a result.
methodology 8 now records one: leakage_real_stride_runs.json was produced on
2026-08-11 under sigmf 1.11.1 -- inferred from the tripwire not recording
1.12.0 until eight days later, and labelled as an inference because the
artifact carries environment: null -- and three cells of it came back
bit-identical when re-measured under sigmf 1.13.0 on a torch, numpy and scipy
that also moved. Stated at the width it supports: one table, three cells, one
direction of upgrade, on CPU.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
--force overrode every refuse category, including the two where there is
nothing to overrule.

Categories 2, 3, 4 and 5 are inferences: this tool deciding what a shared
timestamp, a 43-second gap, a separable axis or an audit LEAK probably means.
Someone who knows their own recordings can be right where the inference is
wrong, and that is what the flag is for -- the overridden category stays in
the header so the block cannot be pasted as a clean run.

Categories 1 and 6 are not inferences. Category 1 means the reader could not
open the files; there is nothing to measure and forcing produces a crash
rather than a questionable number. Category 6 means build would refuse the
split, so the paired measurement cannot be constructed at all. Accepting
--force on either is accepting a request that cannot be fulfilled, and then
failing somewhere further in with a message about something else -- which is
exactly the shape of failure the refuse path exists to prevent.

They now stay REFUSED with exit code 1, and the report says so rather than
merely omitting the FORCED header: `forced  refused. category 1 ... cannot be
overridden`. The JSON carries force_refused alongside forced and forced_past.
Silence would have been indistinguishable from the flag not being passed.

Mutations, each red on the test that should catch it: emptying
STRUCTURAL_CATEGORIES lets the structural refusals be forced again; filling it
with every category breaks the four that must stay forcible; dropping the
report line makes the refusal silent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The WOULD MEASURE path printed

    started       no. This version of the command stops before training

unconditionally, and on a machine with torch the measurement it denied
followed four lines later in the same block. A refuse path whose report
cannot be trusted about its own behaviour has no standing to be trusted
about the recordings.

The line is now derived rather than fixed. The CLI decides before the block
is rendered whether training will follow -- torch present, and a folder of
recordings rather than a built dataset -- and passes the obstacle to decide()
when there is one. `started` then reads `yes. Both arms train after this
block and the result follows below`, or `no.` followed by the actual reason,
or `no. Refused above; nothing was built and nothing was trained`.

The reason is passed in rather than inferred inside decide(), because the
decision genuinely does not know whether its caller intends to train.
Guessing is what produced the wrong sentence.

Three other places asserted the same stale thing and are corrected: the
module docstring, the default WOULD MEASURE reason, and the WorkEstimate
docstring. The JSON gains `trains` and `no_train_reason` so a machine reader
does not have to parse the sentence.

CHANGELOG: "No training CLI yet." and "This version does not train." were
both false and are rewritten, the second one also picking up the --force
restriction from the previous commit.

Mutations: hardcoding the old sentence back fails two tests; making the CLI
stop passing the reason fails the torch-free test. That second mutation
initially passed, because on this machine every run trains and "yes" is right
by accident -- so the branch is now forced open by patching _torch_available
rather than left unreachable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… row writer

measure-leakage --format json had two payload builders and only one was
updated when measurement_schema was introduced. The single-cell path emitted
it and serialised rows through _run_row; the sweep path emitted neither, and
repeated _run_row's nine fields by hand.

runs_from_payload treats a missing measurement_schema as "unversioned, assume
compatible". That is right for payloads written before the field existed, and
it meant a sweep payload written today was indistinguishable from one of
those -- so sweep payloads would be exactly the ones that cannot be told
apart when schema 2 arrives.

The hand-rolled copy is the same shape of duplication that dropped the seed
lists from one of three near-identical measurement calls. A field added to
_run_row would not have reached the sweep payload.

The sweep payload also gains `forced`, which the single-cell payload already
carried. A reader should not have to know which mode produced a file to know
whether --force was used.

Closes #3. Mutations: removing the schema key fails two tests; hand-rolling
the rows again fails the third.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three shipped changes had no entry. The most consequential one had none at
all: the commit that restored the 15 seed pairs touched seven files and not
this one, so the single change that fixed the sample-size regression -- and
added the guard that stops a published grid shrinking again -- was visible
only as a passing mention inside the parity-gate entry.

Added:

- The seed-pair restoration, the --split-seeds / --train-seeds flags,
  guard_artifact_rows, and the check_environment hole that made the guard
  ineffective on exactly the files it was meant to protect.
- measurement_schema in the JSON payload, with the reason it exists: the same
  one read_manifest already has.
- docs/release-notes/v0.5.0.md being labelled an unpublished draft, which is
  what makes __version__, CITATION.cff, the CHANGELOG and the notes agree.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The item was unchecked and its own body said "`--group-by collection` is now
shipped, so the first half of that is done". A checkbox that disagrees with
the paragraph under it teaches a reader to stop trusting the checkboxes.

It was one item covering two things: read the SigMF field, and decide whether
the field needs an extension. The first shipped -- `collection` is the third
scheme beside `path:` and `csv:`, with a test, and a recording that declares
no collection stays its own unit. The second has not started, and is the one
with a bar to clear: use it on a public dataset and record with evidence what
`core:collection` could not express.

They are now two items, ticked according to what is true of each. The
reasoning note and the sigmf-python#159 / SigMF#233 precedent stay attached to
the half they inform, which is the proposal rather than the scheme.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
measure-leakage prints `category 4  ceiling  (methodology 6.4)` and a reader
follows that into §6. §6 did not contain the word "category" anywhere. The
numbering matched for the four eliminated datasets by construction and nothing
held it there, and past four it does not match at all -- which mattered more
than the missing word:

  category 5 is a refusal, a structural leak audit named
  §6.5 is LoRaIQ, the one dataset in that section that was NOT eliminated

Same digit, opposite meaning, and nothing warned a reader following a citation.

§6 now opens with a cross-reference table covering all six categories, marks
the two that have no case here and what they cite instead, states the 5 / 6.5
collision explicitly, and notes which categories --force can override.

A test pins all three documents together: for every member of the Category
enum, SPEC and methodology must each carry a table row with that number and
that name, and the collision warning must still be there. Mutations:
renumbering a category in the code alone fails it, deleting one methodology
row fails it, softening the warning to "Note." fails it.

The assertions compare with whitespace removed. These documents are
hard-wrapped and "There is deliberately no category for it" already straddles
two lines -- an exact-substring assertion passed only until the paragraph was
re-wrapped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
test_loraiq_pattern_is_not_refused_when_grouped asserted "stops before
training" in its torch-free branch. That branch is unreachable on a machine
with torch, so the full local suite passed after the output changed and CI's
test (3.11) and test (3.12) jobs are what caught it.

That is precisely the failure this PR describes one commit earlier -- an
assertion whose coverage depends silently on the environment -- and I walked
into it while fixing it. Worth saying rather than quietly amending.

The branch now asserts what the report actually prints, the torch branch
additionally pins `started  yes.`, and the old sentence is asserted absent on
both paths so neither can drift back.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@emrefbulut
emrefbulut merged commit 22e5423 into main Sep 14, 2026
5 checks passed
emrefbulut added a commit that referenced this pull request Sep 14, 2026
The stacked PRs #18, #19 and #20 were merged at the same moment, so only #18
reached main -- #19 merged into fix/gate-semantics-and-provenance and #20 into
docs/refresh, their own bases. Everything from those two is therefore sitting
on this branch and main is still on 0.4.0. That was my mistake in stacking
three deep instead of retargeting; this merge is the fix.

One conflict, in CHANGELOG.md. main carried the entry describing
docs/release-notes/v0.5.0.md as an unpublished draft, which commit 6491f27 on
this branch had already rewritten into "the four places that carry a version
agree again" -- the draft label was the state this release ends. Kept this
branch's side, which also carries three entries main does not have.

Checked rather than assumed, because a merge across branches that both edited
the test files could silently drop one side: #17's CUDA comparability tests
and #16's forced-measurement width test are both present afterwards, and the
suite goes from 385 to 388 passing, which is main's three additions arriving
rather than anything being lost.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

measure-leakage --sweep JSON omits measurement_schema and duplicates the row serialiser

1 participant