Close the remaining Bölüm C items, and say what the parity gate proves - #18
Merged
Merged
Conversation
The gate was documented in two lines, neither of which said what it checks. "The published tables and the parity gate stay on CPU" is a policy, not a contract, and a reader had no way to tell whether a pass meant "ran with the same settings" or "produced the same numbers". Those are different claims and only one of them is true of this gate. It is the stronger one. compare_cell compares test_accuracy, train_accuracy, train_windows and test_windows by exact equality, row by row, matched by (strategy, split seed, train seed) rather than by position -- on top of the run count and the seed-pair set. Demonstrated against the real rows of leakage_real_stride_runs.json: one test_accuracy moved by 1e-12 fails the cell. What it does not compare is now written down too, because an undocumented exclusion reads as coverage. It ignores the environment block, and it ignores stride / noise_sigma / snr_db, which select the cell rather than being measured by it -- a wrong stride changes the window counts, and those are compared. The environment exclusion is load-bearing rather than missing. Because the gate does not compare environments, a re-run on a different library stack is a real difference between the two runs, and equality across it is a result. methodology 8 now records one: leakage_real_stride_runs.json was produced on 2026-08-11 under sigmf 1.11.1 -- inferred from the tripwire not recording 1.12.0 until eight days later, and labelled as an inference because the artifact carries environment: null -- and three cells of it came back bit-identical when re-measured under sigmf 1.13.0 on a torch, numpy and scipy that also moved. Stated at the width it supports: one table, three cells, one direction of upgrade, on CPU. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
--force overrode every refuse category, including the two where there is nothing to overrule. Categories 2, 3, 4 and 5 are inferences: this tool deciding what a shared timestamp, a 43-second gap, a separable axis or an audit LEAK probably means. Someone who knows their own recordings can be right where the inference is wrong, and that is what the flag is for -- the overridden category stays in the header so the block cannot be pasted as a clean run. Categories 1 and 6 are not inferences. Category 1 means the reader could not open the files; there is nothing to measure and forcing produces a crash rather than a questionable number. Category 6 means build would refuse the split, so the paired measurement cannot be constructed at all. Accepting --force on either is accepting a request that cannot be fulfilled, and then failing somewhere further in with a message about something else -- which is exactly the shape of failure the refuse path exists to prevent. They now stay REFUSED with exit code 1, and the report says so rather than merely omitting the FORCED header: `forced refused. category 1 ... cannot be overridden`. The JSON carries force_refused alongside forced and forced_past. Silence would have been indistinguishable from the flag not being passed. Mutations, each red on the test that should catch it: emptying STRUCTURAL_CATEGORIES lets the structural refusals be forced again; filling it with every category breaks the four that must stay forcible; dropping the report line makes the refusal silent. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The WOULD MEASURE path printed
started no. This version of the command stops before training
unconditionally, and on a machine with torch the measurement it denied
followed four lines later in the same block. A refuse path whose report
cannot be trusted about its own behaviour has no standing to be trusted
about the recordings.
The line is now derived rather than fixed. The CLI decides before the block
is rendered whether training will follow -- torch present, and a folder of
recordings rather than a built dataset -- and passes the obstacle to decide()
when there is one. `started` then reads `yes. Both arms train after this
block and the result follows below`, or `no.` followed by the actual reason,
or `no. Refused above; nothing was built and nothing was trained`.
The reason is passed in rather than inferred inside decide(), because the
decision genuinely does not know whether its caller intends to train.
Guessing is what produced the wrong sentence.
Three other places asserted the same stale thing and are corrected: the
module docstring, the default WOULD MEASURE reason, and the WorkEstimate
docstring. The JSON gains `trains` and `no_train_reason` so a machine reader
does not have to parse the sentence.
CHANGELOG: "No training CLI yet." and "This version does not train." were
both false and are rewritten, the second one also picking up the --force
restriction from the previous commit.
Mutations: hardcoding the old sentence back fails two tests; making the CLI
stop passing the reason fails the torch-free test. That second mutation
initially passed, because on this machine every run trains and "yes" is right
by accident -- so the branch is now forced open by patching _torch_available
rather than left unreachable.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… row writer measure-leakage --format json had two payload builders and only one was updated when measurement_schema was introduced. The single-cell path emitted it and serialised rows through _run_row; the sweep path emitted neither, and repeated _run_row's nine fields by hand. runs_from_payload treats a missing measurement_schema as "unversioned, assume compatible". That is right for payloads written before the field existed, and it meant a sweep payload written today was indistinguishable from one of those -- so sweep payloads would be exactly the ones that cannot be told apart when schema 2 arrives. The hand-rolled copy is the same shape of duplication that dropped the seed lists from one of three near-identical measurement calls. A field added to _run_row would not have reached the sweep payload. The sweep payload also gains `forced`, which the single-cell payload already carried. A reader should not have to know which mode produced a file to know whether --force was used. Closes #3. Mutations: removing the schema key fails two tests; hand-rolling the rows again fails the third. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three shipped changes had no entry. The most consequential one had none at all: the commit that restored the 15 seed pairs touched seven files and not this one, so the single change that fixed the sample-size regression -- and added the guard that stops a published grid shrinking again -- was visible only as a passing mention inside the parity-gate entry. Added: - The seed-pair restoration, the --split-seeds / --train-seeds flags, guard_artifact_rows, and the check_environment hole that made the guard ineffective on exactly the files it was meant to protect. - measurement_schema in the JSON payload, with the reason it exists: the same one read_manifest already has. - docs/release-notes/v0.5.0.md being labelled an unpublished draft, which is what makes __version__, CITATION.cff, the CHANGELOG and the notes agree. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The item was unchecked and its own body said "`--group-by collection` is now shipped, so the first half of that is done". A checkbox that disagrees with the paragraph under it teaches a reader to stop trusting the checkboxes. It was one item covering two things: read the SigMF field, and decide whether the field needs an extension. The first shipped -- `collection` is the third scheme beside `path:` and `csv:`, with a test, and a recording that declares no collection stays its own unit. The second has not started, and is the one with a bar to clear: use it on a public dataset and record with evidence what `core:collection` could not express. They are now two items, ticked according to what is true of each. The reasoning note and the sigmf-python#159 / SigMF#233 precedent stay attached to the half they inform, which is the proposal rather than the scheme. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
measure-leakage prints `category 4 ceiling (methodology 6.4)` and a reader follows that into §6. §6 did not contain the word "category" anywhere. The numbering matched for the four eliminated datasets by construction and nothing held it there, and past four it does not match at all -- which mattered more than the missing word: category 5 is a refusal, a structural leak audit named §6.5 is LoRaIQ, the one dataset in that section that was NOT eliminated Same digit, opposite meaning, and nothing warned a reader following a citation. §6 now opens with a cross-reference table covering all six categories, marks the two that have no case here and what they cite instead, states the 5 / 6.5 collision explicitly, and notes which categories --force can override. A test pins all three documents together: for every member of the Category enum, SPEC and methodology must each carry a table row with that number and that name, and the collision warning must still be there. Mutations: renumbering a category in the code alone fails it, deleting one methodology row fails it, softening the warning to "Note." fails it. The assertions compare with whitespace removed. These documents are hard-wrapped and "There is deliberately no category for it" already straddles two lines -- an exact-substring assertion passed only until the paragraph was re-wrapped. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
test_loraiq_pattern_is_not_refused_when_grouped asserted "stops before training" in its torch-free branch. That branch is unreachable on a machine with torch, so the full local suite passed after the output changed and CI's test (3.11) and test (3.12) jobs are what caught it. That is precisely the failure this PR describes one commit earlier -- an assertion whose coverage depends silently on the environment -- and I walked into it while fixing it. Worth saying rather than quietly amending. The branch now asserts what the report actually prints, the torch branch additionally pins `started yes.`, and the old sentence is asserted absent on both paths so neither can drift back. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
emrefbulut
added a commit
that referenced
this pull request
Sep 14, 2026
The stacked PRs #18, #19 and #20 were merged at the same moment, so only #18 reached main -- #19 merged into fix/gate-semantics-and-provenance and #20 into docs/refresh, their own bases. Everything from those two is therefore sitting on this branch and main is still on 0.4.0. That was my mistake in stacking three deep instead of retargeting; this merge is the fix. One conflict, in CHANGELOG.md. main carried the entry describing docs/release-notes/v0.5.0.md as an unpublished draft, which commit 6491f27 on this branch had already rewritten into "the four places that carry a version agree again" -- the draft label was the state this release ends. Kept this branch's side, which also carries three entries main does not have. Checked rather than assumed, because a merge across branches that both edited the test files could silently drop one side: #17's CUDA comparability tests and #16's forced-measurement width test are both present afterwards, and the suite goes from 385 to 388 passing, which is main's three additions arriving rather than anything being lost. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Seven commits, each its own change. Closes the remaining Bölüm C items and the two clarifications asked for on top of them.
1. What
PARITY_GATE_PASSEDactually asserts (4c6ff03)The gate was documented in two lines, neither of which said what it checks. A reader could not tell "ran with the same settings" from "produced the same numbers". It is the stronger one, and that is now written down in SPEC §5.10.1 and methodology §8.
compare_cellcomparestest_accuracy,train_accuracy,train_windows,test_windowsby exact equality, row by row, matched by(strategy, split seed, train seed)rather than by position — on top of the run count and the seed-pair set. Demonstrated against the real rows ofleakage_real_stride_runs.json(stride 1024, 30 runs):test_accuracymoved by 1e-12train_accuracychangedtrain_windowsoff by oneenvironment, same numbersstriderecorded, same numbersBoth exclusions are now documented rather than left to be discovered.
stride/noise_sigma/snr_dbselect the cell rather than being measured by it — a wrong stride changes the window counts, and those are compared.2. A library upgrade that did not move the numbers (same commit)
The
environmentexclusion is load-bearing, not missing: because the gate ignores it, a re-run on a different stack is a real difference and equality across it is a result.artifacts/leakage_real_stride_runs.jsonwas produced 2026-08-11. Re-measured 2026-09-13 under sigmf 1.13.0, torch 2.13.0+cpu, numpy 2.5.1, scipy 1.18.0 — three cells, 30 runs each — bit-identical on all four compared fields.On the provenance, stated as an inference rather than a measurement: the artifact carries
environment: null, so nothing in it records which sigmf produced it. The version tripwire intests/test_io.pyrecords each release as first encountered and did not record1.12.0until 2026-08-19, eight days later. That puts the original run on 1.11.1. methodology says so in exactly those terms — the gap is precisely what prompted environment stamping.Written at the width it supports: one table, three cells, one direction of upgrade, on CPU.
3.
--forcenow refuses the two categories that are not judgements (520af35)Categories 2–5 are inferences about what a pattern means; a user who knows their recordings can be right where the inference is wrong. Categories 1 (reader cannot open the files) and 6 (
buildwould refuse the split) say no measurement can be constructed. Those stayREFUSED, exit 1, and the report printsforced refused. …rather than merely omitting the FORCED header.4. The report says whether it trained (
46aa665)started no. This version of the command stops before trainingwas printed unconditionally, and on a machine with torch the measurement followed four lines later. It is now derived:yes. Both arms train after this block…, orno.with the actual obstacle. Module docstring, default reason andWorkEstimatedocstring corrected; JSON gainstrainsandno_train_reason. CHANGELOG's "No training CLI yet." and "This version does not train." rewritten.One mutation initially passed here — on this machine every run trains, so
yeswas right by accident. The torch-free branch is now forced open by patching_torch_available.5.
measurement_schemain the sweep payload (b2f5a3a) — closes #3The sweep path emitted no schema and hand-rolled
_run_row's nine fields. A sweep payload written today was indistinguishable from a pre-schema one.6–7. CHANGELOG, ROADMAP, category numbering (
7bfc03d,8bb6f8f,ae0d750)--group-by collectionis shipped. Split into the half that shipped and the half that has not started.category 5and§6.5are opposite things: a refusal, and the one dataset that was not eliminated.Verification
Every behaviour change mutation-verified; the mutations and which test caught each are in the individual commit messages.
Still open, not in this PR
The parity gate's LoRaIQ table is running now. Three of four tables have passed; "Faz 5 doğrulandı" is not claimed until the fourth does.
🤖 Generated with Claude Code