Skip to content
Merged
44 changes: 40 additions & 4 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,12 @@ can tell whether the format it is looking at is one it understands.
parity gate, and `artifacts/*.json` stay on CPU. `TrainingResult.environment`
already recorded the device; the measurement path now stamps the device that
was actually requested.
- **`measurement_schema` in the `measure-leakage --format json` payload**, and
`runs_from_payload` as the single reader of it. A reader that guesses at a
shape it does not recognise produces a plausible wrong answer, which is why
`read_manifest` already refuses a `manifest_schema` newer than it
understands; measurement payloads now get the same treatment. Both the
single-cell and the stride-sweep payloads declare it.
- `iqforge measure-leakage` now accepts `--balance-by`, so the command path can
run the same nuisance-balancing setup that the published synthetic measurement
tables used.
Expand All @@ -40,21 +46,51 @@ can tell whether the format it is looking at is one it understands.
still a placeholder (`examples/` does not false-positive).
- `iqforge.measurement` is the paired leakage-measurement core: one `BuildSpec`,
recording-level build, window-level re-deal, paired training, paired
statistics. No training CLI yet. The three experiment scripts now call it;
statistics. Reached from the CLI by `measure-leakage`, which trains the
paired cell after its refuse-path classification. The three experiment
scripts now call it;
dataset-specific `prepare` stays in `scripts/`. The LoRaIQ bit-exact cell
(stride 1024 / split 42 / train 0) is the acceptance gate and is skipped in
CI when the recordings are not present; published tables are reproduced from
the recorded run files.
- `iqforge measure-leakage` is the refuse path: it runs `audit`, classifies the
result into six categories (methodology §6.1–§6.4 plus remaining leaks and
unsplittable sets), estimates the work a paired cell would do, and stops.
This version does not train. `--force` overrides a refusal and puts the
overridden category in the header (`FORCED PAST audit VERDICT 'ceiling'`).
unsplittable sets), estimates the work a paired cell would do, and — when
nothing fired — trains it. The report's `started` line says which of those
happened rather than asserting one: `yes` when the measurement follows, `no`
with the reason when it does not (no torch, or a built dataset rather than a
folder). `--force` overrides a refusal and puts the overridden category in
the header (`FORCED PAST audit VERDICT 'ceiling'`); it does not apply to
categories 1 and 6, which say no measurement can be built rather than
inferring what the recordings mean.
LoRaIQ-like simultaneous receptions are not refused when `--group-by` holds
them together.

### Changed

- **The published grids are measured at 15 seed pairs again, and cannot
silently shrink.** The Phase 5 migration hardcoded `[42]` and `[0]` into the
command's measurement path, cutting every grid from 15 seed pairs to 1. The
reduced grid reproduces the first pair exactly, so nothing that compared
values noticed; it was found by reading the code, not by reading a result. A
table built that way reports a standard error of zero and calls it a
measurement.
`measure-leakage` now takes `--split-seeds` and `--train-seeds`, defaulting to
the five split seeds and three training seeds every published table used.
They are flags rather than constants so a cheaper run is a visible choice,
and the count is printed with the result: a measurement whose sample size is
not on the page cannot be read. The three experiment scripts pass the same
lists, and `guard_artifact_rows` refuses, before anything is trained, to
overwrite a file under `artifacts/` with fewer runs than it already holds.
`check_environment` was part of the same failure: it returned quietly when a
checkpoint recorded no environment at all, which is the state every published
grid is in, so the guard had never protected one. It now refuses that case
instead of waving it through.
- **`docs/release-notes/v0.5.0.md` says it is an unpublished draft.** The file
read as a shipped release while `__version__`, `CITATION.cff` and the newest
released CHANGELOG section all said `0.4.0` and no `v0.5.0` tag existed. It
now names that state at the top and points at `[Unreleased]`, so the four
places that carry a version agree about which one is real.
- **The experiment scripts and their tests no longer carry a hardcoded path.**
`scripts/leakage_real.py`, `scripts/leakage_loraiq.py`, `tests/test_preflight.py`
and `tests/test_measurement.py` all fell back to an absolute path inside one
Expand Down
16 changes: 11 additions & 5 deletions ROADMAP.md
Original file line number Diff line number Diff line change
Expand Up @@ -72,7 +72,13 @@ After Now is done — still reliability-first:
burst are slices of one continuous capture. See
[docs/methodology.md](docs/methodology.md) §6.

- [ ] **`--group-by` by SigMF field.** Deliberately not in the first release.
- [x] **`--group-by collection`: group by the SigMF field.** Shipped as the
third scheme alongside `path:` and `csv:`. It reads `core:collection`
from the Global Object; a recording that declares none stays its own
unit. The note below is the reasoning that produced it, and the item
after it is the half that is still open.

Deliberately not in the first release, for a reason worth keeping.

The obvious third scheme would read a metadata key, the way
`--balance-by` does. It was left out because it would have solved none of
Expand Down Expand Up @@ -113,10 +119,10 @@ After Now is done — still reliability-first:
An extension proposal is therefore about a *qualifier* on existing
grouping, not a new grouping mechanism.

**`--group-by collection` is now shipped**, so the first half of that is
done. What is not done is the part that decides whether a proposal is
worth writing: using it on a public dataset and recording what it could
not express.
- [ ] **Measure what `core:collection` cannot express, then decide whether to
propose an extension.** The scheme is shipped; this is the part that
decides whether a proposal is worth writing — using it on a public
dataset and recording, with evidence, what it could not say.

**The precedent says do not propose before that.** SigMF accepted an ML
extension and then took it back. `rfml.sigmf-ext.md` was contributed by
Expand Down
39 changes: 39 additions & 0 deletions SPEC.md
Original file line number Diff line number Diff line change
Expand Up @@ -350,6 +350,45 @@ Overlapping air time that `--group-by` already holds together is not category 5:

`--force` does not hide the category. The header becomes `iqforge leakage measurement -- FORCED PAST audit VERDICT 'ceiling'` (or `category N 'name'`), ASCII `--` not a typographic dash, so a pasted block cannot be mistaken for a clean run.

**`--force` applies to categories 2, 3, 4 and 5 only.** Those are inferences: this tool deciding what a timestamp, a gap, a separable axis or an audit finding probably means, and a user who knows their own recordings can be right where the inference is wrong. Categories **1** (the reader cannot open the files) and **6** (`build` would refuse the split) are not inferences — they say no measurement can be constructed. `--force` on either is refused, the decision stays `REFUSED`, the exit code stays 1, and the report says `forced refused. ...` with the reason. A flag that accepts a request it cannot fulfil and fails somewhere further in is worse than one that says no at the point of asking.

### 5.10.1 The parity gate, and what `PARITY_GATE_PASSED` asserts

`scripts/parity_gate.py` re-measures selected cells of the published tables in
`artifacts/` and compares them against what is on disk. It is a deliberate,
hours-long run, not a test, and it is not part of any command's contract.

It compares three things per cell, and a pass needs all three:

| | compared | why |
|---|---|---|
| run count | number of rows for the cell | a grid cut from 15 seed pairs to 1 reproduces the first pair exactly and is a different measurement |
| seed pairs | the set of `(strategy, split seed, train seed)` | fifteen runs from the wrong fifteen seeds is the same count and a different grid |
| results | `test_accuracy`, `train_accuracy`, `train_windows`, `test_windows`, **exact equality, per row**, matched by the key above rather than by position | this is the numerical comparison; a difference of 1e-12 in one accuracy fails the cell |

So `PARITY_GATE_PASSED` asserts **the numbers, not merely the configuration**:
every compared row is bit-identical to the recorded one, and there are exactly
as many of them, from exactly the same seeds.

What it deliberately does **not** compare, so that the claim is not read wider
than it is:

- **`environment`** — device, torch / numpy / scipy / sigmf versions. The
published artifacts predate environment stamping and carry `null`, so there is
nothing to compare against. This is a feature of the check rather than a gap:
it is what lets the gate demonstrate that a result survives a library upgrade
(methodology §8).
- **`stride`, `noise_sigma`, `snr_db`** — these select the cell rather than
being measured by it. A wrong stride changes the window counts, which *are*
compared.

The seed lists are not passed on the command line. The published grids were
measured at the command's defaults, so the defaults are part of what is under
test; passing them would make the run-count check a tautology.

A partial run is not a pass. `--tables` exists because an hours-long run gets
interrupted, and the verdict line names the tables a run actually covered.

**The command is read-only.** It consumes a folder of recordings (or a built dataset) and writes a report. It does not write modified recordings. Every other user-facing command in §4 is already read-only with respect to the user's captures; measurement is not an exception.

**There is no `--sweep snr`.** Adding noise to a user's recordings requires writing altered copies, and doing it correctly is dataset-specific. On DASH7 the carrier is on air 6.8% of the time and about 26 dB of processing gain sits between a wideband SNR figure and the SNR the task sees (methodology §6.4). The pilot that motivated this tool produced a silently useless grid by getting those wrong. An opt-in flag does not fix that — it would be the one place the command touches the user's data, and the one place it can fail silently.
Expand Down
71 changes: 71 additions & 0 deletions docs/methodology.md
Original file line number Diff line number Diff line change
Expand Up @@ -409,6 +409,37 @@ per class — enough that a recording-level split has something to split — a
format the reader can be trusted on, and a task that is neither trivial nor
impossible. Format turned out to be the easy one.

### How these cases map to the refuse categories

`iqforge measure-leakage` refuses a dataset by **category number**, and cites
this section: `category 4 ceiling (methodology 6.4)`. The numbering was taken
from the cases below so the command could point at a paragraph. It matches for
the four eliminated datasets and **does not extend past them**, which is worth
stating here rather than leaving a reader to discover it:

| command category | name | case here | what it cites |
|---|---|---|---|
| 1 | unreadable format | §6.1 AirID | this section |
| 2 | shared timestamp | §6.2 Vega-C | this section |
| 3 | physical independence | §6.3 DASH7 `ds_indoor` | this section |
| 4 | ceiling | §6.4 DASH7 `ds_indoor_cabled` | this section |
| 5 | structural leak | — | the `audit` LEAK finding that fired |
| 6 | cannot split | — | SPEC §5.6 |
| — | *(not refused)* | §6.5 LoRaIQ | — |

**`category 5` and `§6.5` are not the same thing, and they point in opposite
directions.** Category 5 is a refusal: an `audit` LEAK that `--group-by` does
not already hold together. §6.5 is LoRaIQ — the dataset that passed, the one
case in this section that was *not* eliminated. There is deliberately no
category for it. Category 6 likewise has no case here; it is a split `build`
would refuse, and it cites SPEC §5.6.

Categories 1 and 6 cannot be overridden with `--force`; 2 through 5 can. The
line is whether the category is an inference about what the recordings mean —
those are judgements a user may know better than the tool — or a statement
that no measurement can be constructed. SPEC §5.10 carries the full table and
the trigger for each.

### 6.1 Case 1 — AirID

**AirID** (GENESYS Lab, 4 UAV transmitters with deliberately distinct IQ
Expand Down Expand Up @@ -945,6 +976,46 @@ and inspects it; any warning from `build` aborts the run. This exists because th
first version discarded that output and consequently measured a confounded split
for an entire grid. A warning that no one reads is equivalent to no warning.

**Re-measuring the published tables, and saying exactly what that proves.**
`scripts/parity_gate.py` re-runs selected cells of the tables in `artifacts/`
through the shipped command and compares them against the recorded runs. It
compares the run count, the set of `(strategy, split seed, train seed)` pairs,
and — row by row, matched by that key rather than by position — `test_accuracy`,
`train_accuracy`, `train_windows` and `test_windows`, by exact equality.

`PARITY_GATE_PASSED` therefore claims the **numbers**, not merely that the run
used the same configuration. Demonstrated against the real rows of
`artifacts/leakage_real_stride_runs.json` (stride 1024, 30 runs): changing one
`test_accuracy` by 1e-12 fails the cell, as does changing one `train_windows` by
one, cutting the grid to its first seed pair while keeping every value exact, or
keeping the count and shifting the seeds.

It does **not** compare the `environment` block, and that is deliberate rather
than an oversight — see the following note, which depends on it.

**A library upgrade that did not move the numbers.**
`artifacts/leakage_real_stride_runs.json` was produced on 2026-08-11. The
version tripwire in `tests/test_io.py` records each sigmf release as it is first
encountered, and it did not record `1.12.0` until 2026-08-19 — eight days later,
when that release interrupted release preparation. The run therefore used
**sigmf 1.11.1**. That is an inference from the project's own record of which
versions it had seen, not a measurement: the artifact itself carries
`environment: null`, which is precisely the gap that prompted environment
stamping (§7).

Re-measured on 2026-09-13 under **sigmf 1.13.0** — with `torch 2.13.0+cpu`,
`numpy 2.5.1` and `scipy 1.18.0`, none of which match the original stack either
— three cells of that table (stride 1024, 768, 512; 30 runs each) came back
**bit-identical** on all four compared fields. Since the gate does not compare
environments, the upgrade is a genuine difference between the two runs and the
equality is the result: the sigmf 1.11.1 → 1.13.0 transition, including the
`SigMFFile` deep-copy change that `sigmf-python#160` introduced, did not move
this measurement.

This is narrower than "library versions do not matter". It is one table, three
cells, one direction of upgrade, on CPU. It is evidence that the reader change
did not reach the numbers, not that no numeric-stack change could.

**Measuring rather than reasoning.** Where a claim could be checked by running
something, it was — including claims that turned out to be wrong. The initial
diagnosis of a "uniform spectrogram bug" on the cellular recording was incorrect:
Expand Down
Loading
Loading