From 4c6ff03c18e48ea34aba005d69cd081c2754d17a Mon Sep 17 00:00:00 2001 From: Emre Date: Mon, 14 Sep 2026 00:34:42 +0300 Subject: [PATCH 1/8] Say what PARITY_GATE_PASSED asserts, and what it does not The gate was documented in two lines, neither of which said what it checks. "The published tables and the parity gate stay on CPU" is a policy, not a contract, and a reader had no way to tell whether a pass meant "ran with the same settings" or "produced the same numbers". Those are different claims and only one of them is true of this gate. It is the stronger one. compare_cell compares test_accuracy, train_accuracy, train_windows and test_windows by exact equality, row by row, matched by (strategy, split seed, train seed) rather than by position -- on top of the run count and the seed-pair set. Demonstrated against the real rows of leakage_real_stride_runs.json: one test_accuracy moved by 1e-12 fails the cell. What it does not compare is now written down too, because an undocumented exclusion reads as coverage. It ignores the environment block, and it ignores stride / noise_sigma / snr_db, which select the cell rather than being measured by it -- a wrong stride changes the window counts, and those are compared. The environment exclusion is load-bearing rather than missing. Because the gate does not compare environments, a re-run on a different library stack is a real difference between the two runs, and equality across it is a result. methodology 8 now records one: leakage_real_stride_runs.json was produced on 2026-08-11 under sigmf 1.11.1 -- inferred from the tripwire not recording 1.12.0 until eight days later, and labelled as an inference because the artifact carries environment: null -- and three cells of it came back bit-identical when re-measured under sigmf 1.13.0 on a torch, numpy and scipy that also moved. Stated at the width it supports: one table, three cells, one direction of upgrade, on CPU. Co-Authored-By: Claude Opus 5 --- SPEC.md | 37 +++++++++++++++++++++++++++++++++++++ docs/methodology.md | 40 ++++++++++++++++++++++++++++++++++++++++ 2 files changed, 77 insertions(+) diff --git a/SPEC.md b/SPEC.md index 544997f..5be7bfa 100644 --- a/SPEC.md +++ b/SPEC.md @@ -350,6 +350,43 @@ Overlapping air time that `--group-by` already holds together is not category 5: `--force` does not hide the category. The header becomes `iqforge leakage measurement -- FORCED PAST audit VERDICT 'ceiling'` (or `category N 'name'`), ASCII `--` not a typographic dash, so a pasted block cannot be mistaken for a clean run. +### 5.10.1 The parity gate, and what `PARITY_GATE_PASSED` asserts + +`scripts/parity_gate.py` re-measures selected cells of the published tables in +`artifacts/` and compares them against what is on disk. It is a deliberate, +hours-long run, not a test, and it is not part of any command's contract. + +It compares three things per cell, and a pass needs all three: + +| | compared | why | +|---|---|---| +| run count | number of rows for the cell | a grid cut from 15 seed pairs to 1 reproduces the first pair exactly and is a different measurement | +| seed pairs | the set of `(strategy, split seed, train seed)` | fifteen runs from the wrong fifteen seeds is the same count and a different grid | +| results | `test_accuracy`, `train_accuracy`, `train_windows`, `test_windows`, **exact equality, per row**, matched by the key above rather than by position | this is the numerical comparison; a difference of 1e-12 in one accuracy fails the cell | + +So `PARITY_GATE_PASSED` asserts **the numbers, not merely the configuration**: +every compared row is bit-identical to the recorded one, and there are exactly +as many of them, from exactly the same seeds. + +What it deliberately does **not** compare, so that the claim is not read wider +than it is: + +- **`environment`** — device, torch / numpy / scipy / sigmf versions. The + published artifacts predate environment stamping and carry `null`, so there is + nothing to compare against. This is a feature of the check rather than a gap: + it is what lets the gate demonstrate that a result survives a library upgrade + (methodology §8). +- **`stride`, `noise_sigma`, `snr_db`** — these select the cell rather than + being measured by it. A wrong stride changes the window counts, which *are* + compared. + +The seed lists are not passed on the command line. The published grids were +measured at the command's defaults, so the defaults are part of what is under +test; passing them would make the run-count check a tautology. + +A partial run is not a pass. `--tables` exists because an hours-long run gets +interrupted, and the verdict line names the tables a run actually covered. + **The command is read-only.** It consumes a folder of recordings (or a built dataset) and writes a report. It does not write modified recordings. Every other user-facing command in §4 is already read-only with respect to the user's captures; measurement is not an exception. **There is no `--sweep snr`.** Adding noise to a user's recordings requires writing altered copies, and doing it correctly is dataset-specific. On DASH7 the carrier is on air 6.8% of the time and about 26 dB of processing gain sits between a wideband SNR figure and the SNR the task sees (methodology §6.4). The pilot that motivated this tool produced a silently useless grid by getting those wrong. An opt-in flag does not fix that — it would be the one place the command touches the user's data, and the one place it can fail silently. diff --git a/docs/methodology.md b/docs/methodology.md index a106f27..4824eb2 100644 --- a/docs/methodology.md +++ b/docs/methodology.md @@ -945,6 +945,46 @@ and inspects it; any warning from `build` aborts the run. This exists because th first version discarded that output and consequently measured a confounded split for an entire grid. A warning that no one reads is equivalent to no warning. +**Re-measuring the published tables, and saying exactly what that proves.** +`scripts/parity_gate.py` re-runs selected cells of the tables in `artifacts/` +through the shipped command and compares them against the recorded runs. It +compares the run count, the set of `(strategy, split seed, train seed)` pairs, +and — row by row, matched by that key rather than by position — `test_accuracy`, +`train_accuracy`, `train_windows` and `test_windows`, by exact equality. + +`PARITY_GATE_PASSED` therefore claims the **numbers**, not merely that the run +used the same configuration. Demonstrated against the real rows of +`artifacts/leakage_real_stride_runs.json` (stride 1024, 30 runs): changing one +`test_accuracy` by 1e-12 fails the cell, as does changing one `train_windows` by +one, cutting the grid to its first seed pair while keeping every value exact, or +keeping the count and shifting the seeds. + +It does **not** compare the `environment` block, and that is deliberate rather +than an oversight — see the following note, which depends on it. + +**A library upgrade that did not move the numbers.** +`artifacts/leakage_real_stride_runs.json` was produced on 2026-08-11. The +version tripwire in `tests/test_io.py` records each sigmf release as it is first +encountered, and it did not record `1.12.0` until 2026-08-19 — eight days later, +when that release interrupted release preparation. The run therefore used +**sigmf 1.11.1**. That is an inference from the project's own record of which +versions it had seen, not a measurement: the artifact itself carries +`environment: null`, which is precisely the gap that prompted environment +stamping (§7). + +Re-measured on 2026-09-13 under **sigmf 1.13.0** — with `torch 2.13.0+cpu`, +`numpy 2.5.1` and `scipy 1.18.0`, none of which match the original stack either +— three cells of that table (stride 1024, 768, 512; 30 runs each) came back +**bit-identical** on all four compared fields. Since the gate does not compare +environments, the upgrade is a genuine difference between the two runs and the +equality is the result: the sigmf 1.11.1 → 1.13.0 transition, including the +`SigMFFile` deep-copy change that `sigmf-python#160` introduced, did not move +this measurement. + +This is narrower than "library versions do not matter". It is one table, three +cells, one direction of upgrade, on CPU. It is evidence that the reader change +did not reach the numbers, not that no numeric-stack change could. + **Measuring rather than reasoning.** Where a claim could be checked by running something, it was — including claims that turned out to be wrong. The initial diagnosis of a "uniform spectrogram bug" on the cellular recording was incorrect: From 520af35eac4ffc6453b9706a08c3543d12236533 Mon Sep 17 00:00:00 2001 From: Emre Date: Mon, 14 Sep 2026 00:38:15 +0300 Subject: [PATCH 2/8] Refuse --force on the two categories that are not judgements --force overrode every refuse category, including the two where there is nothing to overrule. Categories 2, 3, 4 and 5 are inferences: this tool deciding what a shared timestamp, a 43-second gap, a separable axis or an audit LEAK probably means. Someone who knows their own recordings can be right where the inference is wrong, and that is what the flag is for -- the overridden category stays in the header so the block cannot be pasted as a clean run. Categories 1 and 6 are not inferences. Category 1 means the reader could not open the files; there is nothing to measure and forcing produces a crash rather than a questionable number. Category 6 means build would refuse the split, so the paired measurement cannot be constructed at all. Accepting --force on either is accepting a request that cannot be fulfilled, and then failing somewhere further in with a message about something else -- which is exactly the shape of failure the refuse path exists to prevent. They now stay REFUSED with exit code 1, and the report says so rather than merely omitting the FORCED header: `forced refused. category 1 ... cannot be overridden`. The JSON carries force_refused alongside forced and forced_past. Silence would have been indistinguishable from the flag not being passed. Mutations, each red on the test that should catch it: emptying STRUCTURAL_CATEGORIES lets the structural refusals be forced again; filling it with every category breaks the four that must stay forcible; dropping the report line makes the refusal silent. Co-Authored-By: Claude Opus 5 --- SPEC.md | 2 + docs/teknofest-fit.md | 115 +++++++++++++++++++++++++++++++++++++++ docs/uses.md | 50 +++++++++++++++++ src/iqforge/preflight.py | 55 +++++++++++++++++-- tests/test_preflight.py | 75 +++++++++++++++++++++++++ 5 files changed, 291 insertions(+), 6 deletions(-) create mode 100644 docs/teknofest-fit.md create mode 100644 docs/uses.md diff --git a/SPEC.md b/SPEC.md index 5be7bfa..1ccfa95 100644 --- a/SPEC.md +++ b/SPEC.md @@ -350,6 +350,8 @@ Overlapping air time that `--group-by` already holds together is not category 5: `--force` does not hide the category. The header becomes `iqforge leakage measurement -- FORCED PAST audit VERDICT 'ceiling'` (or `category N 'name'`), ASCII `--` not a typographic dash, so a pasted block cannot be mistaken for a clean run. +**`--force` applies to categories 2, 3, 4 and 5 only.** Those are inferences: this tool deciding what a timestamp, a gap, a separable axis or an audit finding probably means, and a user who knows their own recordings can be right where the inference is wrong. Categories **1** (the reader cannot open the files) and **6** (`build` would refuse the split) are not inferences — they say no measurement can be constructed. `--force` on either is refused, the decision stays `REFUSED`, the exit code stays 1, and the report says `forced refused. ...` with the reason. A flag that accepts a request it cannot fulfil and fails somewhere further in is worse than one that says no at the point of asking. + ### 5.10.1 The parity gate, and what `PARITY_GATE_PASSED` asserts `scripts/parity_gate.py` re-measures selected cells of the published tables in diff --git a/docs/teknofest-fit.md b/docs/teknofest-fit.md new file mode 100644 index 0000000..050e5e1 --- /dev/null +++ b/docs/teknofest-fit.md @@ -0,0 +1,115 @@ +# TEKNOFEST uyumu + +`iqforge`’un TEKNOFEST yarışmalarına nasıl oturduğu. Kategoriler yıldan yıla değişir; bu metin **2026** listesine (52 yarışma) göredir. 2027 isimleri kayabilir — başvurudan önce şartnameyi yeniden okuyun. + +**Eşleme kuralı:** iqforge asla başvurunun kendisi değildir. Başvuru olan sistemin altındaki veri seti ve ölçüm katmanıdır. Canlı yakalama, yön bulma, karıştırma, demodülasyon ve arayüz `iqforge` kapsamı dışındadır (SPEC §2). + +TEKNOFEST 2026 başvuruları 20–28 Şubat 2026’da kapandı. Bunu **takım kurmak ve 2027 hazırlığı** için kullanın; 2026’ya geç başvuru için değil. + +--- + +## Tablo nasıl okunur + +| Uyum | Anlam | +|---|---| +| **Çekirdek** | Puanlanan bir RF/ML görevi; etiketli IQ ve dürüst ayrım gerekir | +| **Destek** | Rapora veya bir alt sisteme yardım eder; yarışma asıl olarak donanım veya görüntüdür | +| **Yok** | Zorlamayın | + +--- + +## Çekirdek — takım olarak girin, veri yolunda iqforge kullanın + +### Elektronik Harp — ilk kez 2026 + +İki görev tipi. Ödül sıralaması asgari **hem ED hem ET** ister (tespit + parametre çıkarımı + yön bulma, artı bir karıştırma *ve* bir aldatma tekniği). Büyük ödül tüm görev grafiğini ister. iqforge yön bulma, karıştırma veya aldatma yapmaz. + +| Şartname görevi | iqforge | Eksik | +|---|---|---| +| ED — tespit | Hayır. Enerji / CFAR / SDR zinciriniz | Donanım + DSP | +| ED — **parametre çıkarımı** | **Evet, ML yarısı.** Şartname taşıyıcı, bant genişliği, güç, analog/sayısal, **modülasyon (tercihen)**, protokol, çoklama, FHSS/DSSS ister. Ek parametre ek puan. Modülasyon/protokol sınıflandırıcıları eğitilmiş modellerdir; veri setleri iqforge’un işidir | Frekans/BW/güç için klasik kestiriciler; BPSK/QPSK’ten fazla sınıf | +| ED — dinleme / demod | Kapsam dışı | Demod yığını | +| ED — yön bulma / konum | Kapsam dışı | Dizi anten | +| ET — karıştırma / aldatma / GNSS aldatma | Kapsam dışı (gönderim) | Verici | +| **En iyi yapay zekâ uygulamaları ödülü** | **Tek ödül olarak en güçlü eşleşme.** Haberleşme ED’de YZ’nin özgünlük, hız, başarı, işlevselliğine verilir. Savunulabilir iddia: *doğruluk kayıt seviyesinde ayrımdan sonra ölçüldü; pencere seviyesindeki şişme gizlenmedi, raporlandı* | Sadece CLI değil, sahada sınıflandırıcının canlı gösterimi | + +**Etrafında ne kurulur:** bir SDR + iqforge veri setleriyle eğitilmiş bir sınıflandırıcı (modülasyon / analog-sayısal) + raporda `audit`. Yön bulma ve ET için ortak. + +### Çelikkubbe Hava Savunma Sistemleri + +Yarışma sistemler sistemidir (algıla, tanı, izle, angaje ol). Ulusal mimaride görünen RF parçaları: ESM, drone RF kimliği, IFF’e yakın sınıflandırma. iqforge **anti-drone / yayıcı kimliği alt sistemine** oturur, komuta-kontrol resmine değil. + +**AirID tipi** yapı zaten işlendi (bir kaydın patlama dilimleri ayrılmamalı). `--group-by` tam bu veriye uyar. + +Yılın şartnamesinde açık bir RF sınıflandırma puanı yoksa **çekirdek değil, destek**. + +### Güvenli Uydu Haberleşmesi / Hareketli Uydu Terminali + +2026 hareketli terminal ağırlıklı olarak yönelme, stabilizasyon ve hareket altında bağlantıdır. Güvenli uydu haberleşmesi geçmişte karıştırmaya/aldatmaya dayanıklılığı da kapsar. + +iqforge uyumu **dar ve gerçek:** SATCOM IQ üzerinde ML (dalga biçimi, girişim sınıfı) ve **dağılım kaymasına “sızıntı” dememek.** Methodology §6.2’deki Vega-C bu ders: farklı geçiş, Doppler, elevasyon, SNR; `audit` bunu artık LEAK değil RISK yazar. + +SATCOM takımının **raporunda** veya şartname girişim sınıflandırmasını puanlıyorsa veri aracı olarak kullanın. Hareketli terminale bir CLI ile girmeyin. + +### Üniversite Öğrencileri Araştırma Projeleri + +iqforge’un *kendisinin* proje olduğu en yakın yol: RF-ML’de ayrım sızıntısı üzerine ölçüm, kamuya açık SigMF + açık kod. JOSS (yazılım) ile konferans (sonuç) aynı deneyden çıkan iki ayrı üründür. + +--- + +## Destek — takımda zaten RF veya etiketli sinyal sorunu varsa + +| Yarışma | IQ’ya neden değebilir | iqforge ne yapar | Ne yapmaz | +|---|---|---|---| +| FPV drone izleme (2026 yeni) | Bazı yıllarda görüntü/RF karma | RF varlık veya kimlik başı için veri seti | Kutu, video | +| Havacılıkta yapay zekâ | Genelde görüntü / sim / eniyileme | Yalnızca o yılın görevi radyo veya transponder IQ içeriyorsa | Tipik görüntü hattı | +| 5G ve YZ ile akıllı yol güvenliği | Şartname havadan IQ kullanıyorsa 5G / C-V2X | SigMF’ten doluluk / girişimci sınıfları | Yol sahnesi ML | +| Model uydu / roket / dikey iniş | Telemetri RF, yer istasyonu | Etiketli telemetri patlamaları, dürüst ayrım | Uçuş mekaniği, GNC | +| Sürü İHA / savaşan İHA / drone şampiyonası | C2 linkleri, dost-düşman RF, GNSS aldatma tespiti | Sınıflandırıcı veri setleri; rapor için `audit` | Gövde, otonomi | +| İnsansız deniz / su altı | Akustik SigMF-IQ değil; aracın RF haberleşmesi olabilir | Yalnızca kaydedilmiş RF komut linki | Sonar ML | +| Çip tasarım | Çip RF ön uçsa ve IQ toplanıyorsa | Silikon sonrası IQ → veri seti | RTL, PDK | +| Kuantum (yazılım) | Düşük olasılık | Radyo deneyi yoksa yok | Kuantum yığını | +| Maden / sanayide robotik | RF salınımından kestirimci bakım zorlama | Endüstriyel EMI sınıflandırması gibi | Mekanik robot | + +--- + +## Yok — eşlemeye zaman ayırmayın + +Blokzincir, e-ticaret, fintech, doğal dil işleme, YZ film, mimari, biyoteknoloji, onkoloji, lise iklim/kutup araştırması, hyperloop, nükleer *tasarım*, jet motor *tasarım*, uçan araba simülasyonu, lojistik eniyileme, PARDUS hata avı, Robolig, TravelX, mesleki yetenek, insanlık yararına (K-12), elektrikli araç yarışları, Robotaksi (radyo alt modülü puanlanmıyorsa). + +Bunlara iqforge zorlamak kimliği sulandırır ve puan getirmez. + +--- + +## Elektronik harp — görev grafiği (ayrıntı) + +``` +Spektrum + → tespit [DSP / SDR’niz] + → parametre çıkar [klasik + ML] + ML sınıfları ← iqforge build / Dataset / audit + → dinle/demod [kapsam dışı] + → yön bul [kapsam dışı] + → ET ata [kapsam dışı] +``` + +**ED parametre çıkarımında işe yarar bir takım arkadaşı olmak için asgari** + +1. Sahanın yayınlayacağı ailelerin donanım kaydı (2026 şartnamesine göre amatör telsiz, ISM modülleri), SigMF olarak. +2. `{bpsk, qpsk}`’ten geniş etiketler: analog/sayısal, birkaç sayısal modülasyon, isteğe bağlı FHSS / sabit. +3. Kayıt (veya `--group-by`) seviyesinde `iqforge build`; teknik rapora `audit`. +4. Küçük bir sınıflandırıcı (`train` duman testidir, yarışma modeli değil). Gerçek mimariyi siz koyun; veri sözleşmesini koruyun. + +Demo sayısını “şişirmek” için CLI’ya gizli pencere seviyesi ayrımı **koymayın**. Araç tam olarak o bozulmayı önlemek için var. Kötü ayrımı rapordaki tablo için `scripts/` altındaki bir ölçüm script’i çalıştırabilir. + +--- + +## Sizin için dürüst sıra (tek kişi, bu repo) + +1. **Araştırma projesi / bildiri** — en yüksek kaldıraç; ROADMAP Now ile örtüşür (sızıntıyı gerçek kayıtlarda tekrarlamak). +2. **EH takımı, YZ ödülü + parametre çıkarımı puanı** — yüksek kaldıraç; ortak ve radyo ister. +3. **Anti-drone / Çelikkubbe alt sistemi** — 2027 şartnamesi RF kimliği puanlıyorsa. +4. **SATCOM rapor eki** — Vega-C dersi; az zaman. +5. Geri kalanı — yalnızca takım arkadaşının zaten aracı varsa ve radyo ML yan kolu istiyorsa. + +Eksik olan yeni bir yarışma fikri değil. **Radyosu olan üç kişi** (ROADMAP Next) ve kamuya açık gerçek SigMF üzerinde bir sızıntı tablosu. diff --git a/docs/uses.md b/docs/uses.md new file mode 100644 index 0000000..d78e37f --- /dev/null +++ b/docs/uses.md @@ -0,0 +1,50 @@ +# iqforge ne işe yarar + +Savunma laboratuvarı, üniversite grubu veya işe alım için tek sayfa. Özellik listesi değil. + +## İddia + +Elinizde SDR kayıtları var. Eğitilmiş bir model istiyorsunuz. Bu iki cümle arasında yanlış yapılması kolay, fark edilmesi zor kararlar durur. Sonuçları sessizce bozanı train/test ayrımıdır. + +Pencere seviyesinde bölme, komşu pencereleri — aynı semboller, aynı gürültü, aynı sönümleme, yarım adım kaymış — hem eğitime hem teste koyar. Raporlanan doğruluk **yükselir**. Model kötüleşir. Sayıya bakınca hiçbir şey şüpheli görünmez. + +Kontrollü iki sınıflı bir görevde bu şişme tepe noktada **+13,6 yüzde puan** oldu (eşleştirilmiş, 15 tohum çifti, patlama SNR −2,2 dB). Yüksek SNR’de her iki kol da tavana oturur; sızıntı görünmez. Bkz. [methodology.md](methodology.md) §2. + +`iqforge` **kayıt** seviyesinde böler. Geçerli katmanlı bir ayrım mümkün değilse geri düşmez, **hata verir**. Ürün bu reddediştir. + +## Ne değildir + +Sınıflandırıcı, detektör, demodülatör veya elektronik harp sistemi değildir. Onların altındaki katmandır: mevcut SigMF kayıtlarını, bir incelemede savunabileceğiniz etiketli bir PyTorch veri setine çevirir. + +## Bu sayı nerede işe yarar + +| Ortam | Ne eğitilir | Sızıntı nasıl görünür | +|---|---|---| +| Elektronik destek / COMINT | Modülasyon, protokol, yayıcı kimliği | Aynı oturum eğitim ve testte; model sınıfı değil kaydı ezberler | +| Özgül yayıcı tanımlama | “Hangi tip değil, hangi radyo” | Aynı yayıcının komşu pencereleri kimliği sızdırır | +| Anti-drone / RF parmak izi | İniş yolundan drone tipi | Aynı uçuşun patlama dilimleri bağımsız örnek sayılır (AirID tipi veri) | +| Uydu haberleşme / telemetri ML | Doppler altında dalga biçimi veya uydu | Farklı geçişler farklı koşul gibi durur; buna “sızıntı” demek yanlış teşhistir (Vega-C: dağılım kayması, örtüşme değil) | +| Spektrum izleme / regülatör | Doluluk, kaçak verici | Aynı yer, aynı saat, iki dosya — yapıca ayrı, fiziken aynı kanal | +| Akademik RF-ML makaleleri | Herhangi bir IQ CNN | Pencere karıştırma hâlâ varsayılan; şişme görevin *sınırda* olduğu yerde en büyük — ki makaleler de orada “atılım” iddia eder | + +`iqforge audit` neyi kontrol ettiğini, ne bulduğunu ve neyi kontrol edemediğini yazar. İncelenmemiş alanı gizleyen bir “temiz” durumu yoktur. Örtüşen yayın süresi, aynı baytlar, iki kutuya düşen bir kayıt ve her ayrıma *farklı* bir geçiş koyan tekrarlı zaman damgası ayrı hatalar olarak adlandırılır — bir skoru zıt yönlerde hareket ettirirler. + +## Bugün ne yapılabilir + +``` +iqforge info capture.sigmf-meta +iqforge inspect capture.sigmf-meta +iqforge build recordings/ -o dataset/ --balance-by core:freq_lower_edge +iqforge stats dataset/ +iqforge audit dataset/ +``` + +`cf32_le`, `ci16_le`, `ci8` okur. Tamsayı yolları kamuya açık kayıtlara karşı kontrol edildi (tam ölçek, I/Q sırası) — sizin radyoya karşı değil. `--group-by` ilgili dosyaları bir arada tutar (yol regex, CSV veya `core:collection`). PyTorch isteğe bağlıdır. + +## Neyin yerini tutmaz + +Donanım. Yön bulma dizisi. Karıştırıcı. Makale. Tedarik test planı. Bunlar başka sistemlerdir. iqforge, o sistemlerin sızmış bir ayrımdan gelmiş bir sayıyı alıntılamasını durdurur. + +## Atıf + +[github.com/emrefbulut/iqforge](https://github.com/emrefbulut/iqforge) — MIT. Ölçüm tabloları `artifacts/` içinde; koşullar (cihaz, kütüphane sürümleri, n, aralık) sayıyla birlikte gelir. diff --git a/src/iqforge/preflight.py b/src/iqforge/preflight.py index 389864d..a7ef897 100644 --- a/src/iqforge/preflight.py +++ b/src/iqforge/preflight.py @@ -17,6 +17,13 @@ `--force` does not hide the category. It changes the header so a pasted block cannot be mistaken for a clean run, and it changes the decision to `WOULD MEASURE`. The reason that was overridden stays in the body. + +It does not apply to every category. Categories 1 and 6 -- the files cannot be +read, and `build` would refuse the split -- are not judgements this tool made +about what the recordings mean; they are statements that no measurement can be +constructed. `--force` on those is refused and said so in the report, because a +flag that accepts a request it cannot fulfil and fails somewhere further in is +worse than one that says no at the point of asking. See `STRUCTURAL_CATEGORIES`. """ from __future__ import annotations @@ -92,6 +99,24 @@ class Category(IntEnum): CANNOT_SPLIT = 6 +#: Categories `--force` cannot override, because they are not judgements. +#: +#: The other four are inferences from metadata. A shared timestamp, a pair of +#: captures two minutes apart, a separable axis, an audit LEAK -- each is this +#: tool deciding what a pattern probably means, and someone who knows their own +#: recordings can be right where the inference is wrong. `--force` exists for +#: them, and puts the overridden category in the header so the result cannot be +#: pasted as a clean one. +#: +#: These two are not inferences. Category 1 means the reader could not open the +#: files: there is nothing to measure, and forcing produces a crash rather than +#: a questionable number. Category 6 means `build` would refuse the split, so +#: the measurement cannot be constructed at all. Overriding either asks for a +#: run that cannot exist, and a flag that accepts the request and then fails +#: somewhere further in is worse than one that says no here. +STRUCTURAL_CATEGORIES = frozenset({Category.UNREADABLE, Category.CANNOT_SPLIT}) + + #: Short name and the citation a reader can follow. Category 5 cites the #: audit finding, not §6.5 — that section is the dataset that passed. CATEGORY_META: dict[Category, tuple[str, str]] = { @@ -135,6 +160,9 @@ class Decision: category: Category | None = None forced: bool = False forced_past: str | None = None + #: Why `--force` was given and not honoured. None when it was not given, or + #: when it was honoured. + force_refused: str | None = None work: WorkEstimate | None = None audit: AuditReport | None = None findings: list[Finding] = field(default_factory=list) @@ -510,14 +538,25 @@ def decide( reason = unreadable_error or "audit produced no report" forced_past = None + force_refused = None if force and status is not DecisionStatus.WOULD_MEASURE: - if category is Category.CEILING: - forced_past = f"audit VERDICT '{_verdict_token(report)}'" - elif category is not None: - forced_past = f"category {int(category)} '{CATEGORY_META[category][0]}'" + if category in STRUCTURAL_CATEGORIES: + assert category is not None + force_refused = ( + f"category {int(category)} '{CATEGORY_META[category][0]}' cannot be " + f"overridden. It is not a judgement about what the recordings mean; " + f"it is that the measurement cannot be built at all. --force is for " + f"the categories where you may know something this tool inferred " + f"wrongly, and there is nothing here to be right about." + ) else: - forced_past = "an INCONCLUSIVE audit" - status = DecisionStatus.WOULD_MEASURE + if category is Category.CEILING: + forced_past = f"audit VERDICT '{_verdict_token(report)}'" + elif category is not None: + forced_past = f"category {int(category)} '{CATEGORY_META[category][0]}'" + else: + forced_past = "an INCONCLUSIVE audit" + status = DecisionStatus.WOULD_MEASURE if status is DecisionStatus.WOULD_MEASURE and seconds_per_window_epoch is not None: work = estimate_work( @@ -530,6 +569,7 @@ def decide( category=category, forced=bool(forced_past), forced_past=forced_past, + force_refused=force_refused, work=work, audit=report, findings=list(report.findings) if report is not None else [], @@ -596,6 +636,8 @@ def render_text(decision: Decision) -> str: "yes. The category above still stands; this run is not a clean measurement", ) ) + elif decision.force_refused: + out.extend(_field_lines("forced", f"refused. {decision.force_refused}")) out += ["", "WORK", thin] work = decision.work @@ -665,6 +707,7 @@ def render_json(decision: Decision) -> str: "reason": decision.reason, "forced": decision.forced, "forced_past": decision.forced_past, + "force_refused": decision.force_refused, "work": None if work is None else { diff --git a/tests/test_preflight.py b/tests/test_preflight.py index d1aaa09..1d4822e 100644 --- a/tests/test_preflight.py +++ b/tests/test_preflight.py @@ -487,3 +487,78 @@ def test_short_same_class_frames_are_not_the_indoor_pattern() -> None: ) decision = decide(report, seconds_per_window_epoch=None) assert decision.category is not Category.INDEPENDENCE + + +# -------------------------------------------------------------------------- +# --force: what it can and cannot override +# -------------------------------------------------------------------------- + + +def test_force_cannot_override_an_unreadable_set(tmp_path: Path) -> None: + """Category 1 is not a judgement, so there is nothing to overrule. + + Every forcible category is this tool inferring what a pattern means, and + someone who knows their own recordings can be right where the inference is + wrong. This one says the reader could not open the files: there is nothing + to measure, and accepting `--force` would only move the failure further in. + """ + result = _invoke(_write_airid(tmp_path), "--force") + + assert result.exit_code == 1, result.output + assert "1 unreadable format" in result.output + assert "FORCED PAST" not in result.output + assert "cannot be overridden" in result.output + + +def test_force_cannot_override_a_split_that_cannot_be_made(tmp_path: Path) -> None: + """Category 6: `build` would refuse, so no measurement can be constructed.""" + folder = tmp_path / "tiny" + write_record(folder / "a", _samples(seed=1), name="one") + write_record(folder / "b", _samples(seed=2), name="two") + result = _invoke(folder, "--force") + + assert result.exit_code == 1, result.output + assert "6 cannot split" in result.output + assert "FORCED PAST" not in result.output + assert "cannot be overridden" in result.output + + +def test_the_json_reports_a_refused_force(tmp_path: Path) -> None: + result = _invoke(_write_airid(tmp_path), "--force", "--format", "json") + payload = json.loads(result.output) + + assert payload["status"] == "REFUSED" + assert payload["category"] == 1 + assert payload["forced"] is False + assert payload["forced_past"] is None + assert "cannot be overridden" in payload["force_refused"] + + +def test_force_still_overrides_a_heuristic_refusal(tmp_path: Path) -> None: + """The forcible four must keep working, or the distinction became a ban. + + Library-level rather than through the CLI: `decide` is where the split + lives, and going through the command would train the forced cell. + """ + report = _audit_folder(_write_vega_c(tmp_path), 1024, 512, "dirname", 1) + decision = decide(report, force=True, seconds_per_window_epoch=None) + + assert decision.category is Category.SHARED_TIMESTAMP + assert decision.status is DecisionStatus.WOULD_MEASURE + assert decision.forced is True + assert decision.forced_past + assert decision.force_refused is None + + +def test_every_category_is_either_forcible_or_structural() -> None: + """No category may fall outside the split as new ones are added.""" + from iqforge.preflight import STRUCTURAL_CATEGORIES + + assert STRUCTURAL_CATEGORIES == {Category.UNREADABLE, Category.CANNOT_SPLIT} + forcible = set(Category) - set(STRUCTURAL_CATEGORIES) + assert forcible == { + Category.SHARED_TIMESTAMP, + Category.INDEPENDENCE, + Category.CEILING, + Category.STRUCTURAL_LEAK, + } From 46aa6659bbb1975ffbdab3a1890ad6e1e2ed46d6 Mon Sep 17 00:00:00 2001 From: Emre Date: Mon, 14 Sep 2026 00:41:19 +0300 Subject: [PATCH 3/8] Make the report say whether it trained, instead of asserting it did not The WOULD MEASURE path printed started no. This version of the command stops before training unconditionally, and on a machine with torch the measurement it denied followed four lines later in the same block. A refuse path whose report cannot be trusted about its own behaviour has no standing to be trusted about the recordings. The line is now derived rather than fixed. The CLI decides before the block is rendered whether training will follow -- torch present, and a folder of recordings rather than a built dataset -- and passes the obstacle to decide() when there is one. `started` then reads `yes. Both arms train after this block and the result follows below`, or `no.` followed by the actual reason, or `no. Refused above; nothing was built and nothing was trained`. The reason is passed in rather than inferred inside decide(), because the decision genuinely does not know whether its caller intends to train. Guessing is what produced the wrong sentence. Three other places asserted the same stale thing and are corrected: the module docstring, the default WOULD MEASURE reason, and the WorkEstimate docstring. The JSON gains `trains` and `no_train_reason` so a machine reader does not have to parse the sentence. CHANGELOG: "No training CLI yet." and "This version does not train." were both false and are rewritten, the second one also picking up the --force restriction from the previous commit. Mutations: hardcoding the old sentence back fails two tests; making the CLI stop passing the reason fails the torch-free test. That second mutation initially passed, because on this machine every run trains and "yes" is right by accident -- so the branch is now forced open by patching _torch_available rather than left unreachable. Co-Authored-By: Claude Opus 5 --- CHANGELOG.md | 15 +++++-- src/iqforge/cli.py | 13 +++++++ src/iqforge/preflight.py | 46 ++++++++++++++-------- tests/test_preflight.py | 84 ++++++++++++++++++++++++++++++++++++++++ 4 files changed, 137 insertions(+), 21 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index c970b8d..3c95853 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -40,16 +40,23 @@ can tell whether the format it is looking at is one it understands. still a placeholder (`examples/` does not false-positive). - `iqforge.measurement` is the paired leakage-measurement core: one `BuildSpec`, recording-level build, window-level re-deal, paired training, paired - statistics. No training CLI yet. The three experiment scripts now call it; + statistics. Reached from the CLI by `measure-leakage`, which trains the + paired cell after its refuse-path classification. The three experiment + scripts now call it; dataset-specific `prepare` stays in `scripts/`. The LoRaIQ bit-exact cell (stride 1024 / split 42 / train 0) is the acceptance gate and is skipped in CI when the recordings are not present; published tables are reproduced from the recorded run files. - `iqforge measure-leakage` is the refuse path: it runs `audit`, classifies the result into six categories (methodology §6.1–§6.4 plus remaining leaks and - unsplittable sets), estimates the work a paired cell would do, and stops. - This version does not train. `--force` overrides a refusal and puts the - overridden category in the header (`FORCED PAST audit VERDICT 'ceiling'`). + unsplittable sets), estimates the work a paired cell would do, and — when + nothing fired — trains it. The report's `started` line says which of those + happened rather than asserting one: `yes` when the measurement follows, `no` + with the reason when it does not (no torch, or a built dataset rather than a + folder). `--force` overrides a refusal and puts the overridden category in + the header (`FORCED PAST audit VERDICT 'ceiling'`); it does not apply to + categories 1 and 6, which say no measurement can be built rather than + inferring what the recordings mean. LoRaIQ-like simultaneous receptions are not refused when `--group-by` holds them together. diff --git a/src/iqforge/cli.py b/src/iqforge/cli.py index b6ca0fe..234bb98 100644 --- a/src/iqforge/cli.py +++ b/src/iqforge/cli.py @@ -1218,6 +1218,18 @@ def measure_leakage( # noqa: PLR0913 — flags match `audit` plus --force / --g except IQForgeError as exc: raise _fail(exc) from exc + # Decided before the block is rendered, because the block states it. The + # report used to say "stops before training" and then train, which is the + # one thing a refuse path must never do: describe a run it is not making. + if not _torch_available(): + no_train_reason = "torch is not installed, so this run ends at the classification above" + elif not path.is_dir(): + no_train_reason = ( + "a built dataset is classified only; measuring needs a folder of recordings" + ) + else: + no_train_reason = None + decision = decide( report, force=force, @@ -1226,6 +1238,7 @@ def measure_leakage( # noqa: PLR0913 — flags match `audit` plus --force / --g stride=stride, group_keys=group_keys, unreadable_error=unreadable_error, + no_train_reason=no_train_reason, ) printed = ( render_measure_json(decision) if output_format == "json" else render_measure_text(decision) diff --git a/src/iqforge/preflight.py b/src/iqforge/preflight.py index a7ef897..1873161 100644 --- a/src/iqforge/preflight.py +++ b/src/iqforge/preflight.py @@ -11,8 +11,9 @@ - `REFUSED` — a category fired; measuring would report the wrong thing. - `INCONCLUSIVE` — the sources the categories need are missing, so neither refusing nor measuring would be honest. -- `WOULD MEASURE` — nothing in the list fired. This version still does not - train; it says so and reports the work a later version would do. +- `WOULD MEASURE` — nothing in the list fired. The command trains the paired + cell after this block, unless the caller says it will not; the report's + `started` line says which happened rather than asserting one of them. `--force` does not hide the category. It changes the header so a pasted block cannot be mistaken for a clean run, and it changes the decision to @@ -131,7 +132,7 @@ class Category(IntEnum): @dataclass(frozen=True) class WorkEstimate: - """What a default paired cell would train, if this command trained. + """What the default paired cell trains. Attributes: train_windows: Predicted training-split size from recording lengths. @@ -163,6 +164,12 @@ class Decision: #: Why `--force` was given and not honoured. None when it was not given, or #: when it was honoured. force_refused: str | None = None + #: Why the caller will not train despite `WOULD MEASURE` -- no torch, or a + #: built dataset rather than a folder. None means training follows, and the + #: report says so. This is passed in rather than inferred: the decision does + #: not know whether its caller intends to train, and printing "stops before + #: training" above a measurement is how the report came to contradict itself. + no_train_reason: str | None = None work: WorkEstimate | None = None audit: AuditReport | None = None findings: list[Finding] = field(default_factory=list) @@ -460,6 +467,7 @@ def decide( group_keys: dict[str, str] | None = None, unreadable_error: str | None = None, seconds_per_window_epoch: float | None | object = ..., + no_train_reason: str | None = None, ) -> Decision: """Classify an audit into a refuse category, or say it would measure. @@ -475,6 +483,8 @@ def decide( unreadable_error: The error from a fully unreadable folder. seconds_per_window_epoch: Injected probe rate. Omit to run the probe; pass None to skip it. + no_train_reason: Why the caller will not train even on `WOULD MEASURE`. + Omit when training follows. """ features = list(report.features) if report is not None else [] # Window counts are cheap and always useful. The dummy-batch probe is @@ -484,9 +494,8 @@ def decide( category: Category | None = None status = DecisionStatus.WOULD_MEASURE reason = ( - "audit did not fire a refuse category. This command does not train; " - "a later version would run the paired experiment at the default " - "operating point" + "audit did not fire a refuse category. The paired experiment runs at " + "the default operating point" ) unreadable = _is_unreadable(report, unreadable_error) @@ -570,6 +579,7 @@ def decide( forced=bool(forced_past), forced_past=forced_past, force_refused=force_refused, + no_train_reason=no_train_reason, work=work, audit=report, findings=list(report.findings) if report is not None else [], @@ -645,7 +655,7 @@ def render_text(decision: Decision) -> str: out.extend( _field_lines( "estimate", - "no window count from these recordings; this command does not train", + "no window count from these recordings", ) ) else: @@ -662,7 +672,7 @@ def render_text(decision: Decision) -> str: _field_lines( "estimate", "not timed on this machine (torch missing, or the dummy-batch " - "probe failed). This command does not train", + "probe failed)", ) ) else: @@ -674,17 +684,17 @@ def render_text(decision: Decision) -> str: f"measurement of the recordings)", ) ) - if decision.status is DecisionStatus.REFUSED and not decision.forced: - out.extend(_field_lines("started", "no")) - elif decision.status is DecisionStatus.WOULD_MEASURE: - out.extend( - _field_lines( - "started", - "no. This version of the command stops before training", - ) + if decision.status is DecisionStatus.WOULD_MEASURE: + started = ( + f"no. {decision.no_train_reason}" + if decision.no_train_reason + else "yes. Both arms train after this block and the result follows below" ) + elif decision.status is DecisionStatus.REFUSED: + started = "no. Refused above; nothing was built and nothing was trained" else: - out.extend(_field_lines("started", "no")) + started = "no. The audit was inconclusive; nothing was built and nothing was trained" + out.extend(_field_lines("started", started)) out.append(rule) # Audit findings this command quotes still carry U+00A7. Translate them # so the pasted block stays ASCII, which is the contract inherited from @@ -708,6 +718,8 @@ def render_json(decision: Decision) -> str: "forced": decision.forced, "forced_past": decision.forced_past, "force_refused": decision.force_refused, + "trains": decision.status is DecisionStatus.WOULD_MEASURE and not decision.no_train_reason, + "no_train_reason": decision.no_train_reason, "work": None if work is None else { diff --git a/tests/test_preflight.py b/tests/test_preflight.py index 1d4822e..906cd47 100644 --- a/tests/test_preflight.py +++ b/tests/test_preflight.py @@ -562,3 +562,87 @@ def test_every_category_is_either_forcible_or_structural() -> None: Category.CEILING, Category.STRUCTURAL_LEAK, } + + +# -------------------------------------------------------------------------- +# The report must describe the run it is actually making +# -------------------------------------------------------------------------- + + +@pytest.mark.skipif(not _torch_installed(), reason="training must be possible to contradict it") +def test_a_run_that_trains_does_not_say_it_stops_before_training(tmp_path: Path) -> None: + """The block used to announce the opposite of what happened next. + + `started no. This version of the command stops before training` was + printed unconditionally on the WOULD MEASURE path, and the measurement it + denied followed four lines later. A refuse path whose report cannot be + trusted about its own behaviour has no standing to be trusted about the + recordings. + """ + result = _invoke( + _write_ceiling(tmp_path), "--force", "--split-seeds", "42", "--train-seeds", "0" + ) + + assert result.exit_code == 0, result.output + assert "stops before training" not in result.output + started = next(line for line in result.output.splitlines() if line.startswith("started")) + assert "yes" in started, started + assert "MEASUREMENT" in result.output + + +def test_the_started_line_reports_each_outcome(tmp_path: Path) -> None: + """All three branches, at library level so no environment is required.""" + report = _audit_folder(_write_loraiq_pattern(tmp_path), 1024, 512, "dirname", 2) + keys = resolve_group_keys( + [f.record_id for f in report.features], + r"path:([^/]+/tx\d+)", + collections={f.record_id: f.collection for f in report.features}, + ) + + trains = decide(report, group_keys=keys, seconds_per_window_epoch=None) + assert trains.status is DecisionStatus.WOULD_MEASURE + assert "started yes." in render_text(trains) + + held = decide( + report, + group_keys=keys, + seconds_per_window_epoch=None, + no_train_reason="torch is not installed, so this run ends at the classification above", + ) + text = render_text(held) + assert "started no. torch is not installed" in text + assert "stops before training" not in text + + refused = decide(_report_for_category_4(tmp_path), seconds_per_window_epoch=None) + assert refused.status is DecisionStatus.REFUSED + assert "nothing was built and nothing was trained" in render_text(refused) + + +def _report_for_category_4(tmp_path: Path): + return _audit_folder(_write_ceiling(tmp_path / "ceil"), 1024, 512, "dirname", 1) + + +def test_a_run_that_cannot_train_says_why(tmp_path: Path, monkeypatch: pytest.MonkeyPatch) -> None: + """The torch-free branch, forced open on a machine that has torch. + + Without this the branch is unreachable here and a CLI that stopped passing + the reason would still look correct: every local run trains, so "yes" is + right by accident. The assertion is that the report names the obstacle, + not merely that it avoids the old wrong sentence. + """ + monkeypatch.setattr("iqforge.cli._torch_available", lambda: False) + result = _invoke( + _write_loraiq_pattern(tmp_path), + "--dirname-level", + "2", + "--group-by", + r"path:([^/]+/tx\d+)", + ) + + assert result.exit_code == 0, result.output + assert "WOULD MEASURE" in result.output + started = next(line for line in result.output.splitlines() if line.startswith("started")) + assert "no." in started, started + assert "torch is not installed" in result.output + assert "stops before training" not in result.output + assert "MEASUREMENT" not in result.output From b2f5a3aafd5265e3d6cf6fe3e5d340b1411e5c21 Mon Sep 17 00:00:00 2001 From: Emre Date: Mon, 14 Sep 2026 00:42:09 +0300 Subject: [PATCH 4/8] Declare the schema in the sweep payload too, and stop duplicating the row writer measure-leakage --format json had two payload builders and only one was updated when measurement_schema was introduced. The single-cell path emitted it and serialised rows through _run_row; the sweep path emitted neither, and repeated _run_row's nine fields by hand. runs_from_payload treats a missing measurement_schema as "unversioned, assume compatible". That is right for payloads written before the field existed, and it meant a sweep payload written today was indistinguishable from one of those -- so sweep payloads would be exactly the ones that cannot be told apart when schema 2 arrives. The hand-rolled copy is the same shape of duplication that dropped the seed lists from one of three near-identical measurement calls. A field added to _run_row would not have reached the sweep payload. The sweep payload also gains `forced`, which the single-cell payload already carried. A reader should not have to know which mode produced a file to know whether --force was used. Closes #3. Mutations: removing the schema key fails two tests; hand-rolling the rows again fails the third. Co-Authored-By: Claude Opus 5 --- src/iqforge/cli.py | 17 ++------- tests/test_measurement.py | 75 +++++++++++++++++++++++++++++++++++++++ 2 files changed, 78 insertions(+), 14 deletions(-) diff --git a/src/iqforge/cli.py b/src/iqforge/cli.py index 234bb98..7b689ab 100644 --- a/src/iqforge/cli.py +++ b/src/iqforge/cli.py @@ -1324,26 +1324,15 @@ def _measured(cells: list[GridCell]) -> list[Any]: payload = { "preflight": json.loads(render_measure_json(decision)), "measurement": { + "measurement_schema": MEASUREMENT_SCHEMA, "mode": "sweep_stride", + "forced": decision.forced, "split_seeds": split_seed_list, "train_seeds": train_seed_list, "seed_pairs": pairs, "strides": list(strides), "table": table, - "rows": [ - { - "stride": run.stride, - "strategy": run.strategy, - "split_seed": run.split_seed, - "train_seed": run.train_seed, - "test_accuracy": run.test_accuracy, - "train_accuracy": run.train_accuracy, - "train_windows": run.train_windows, - "test_windows": run.test_windows, - "environment": run.environment, - } - for run in runs - ], + "rows": [_run_row(run) for run in runs], }, } print(json.dumps(payload, indent=2, ensure_ascii=True)) diff --git a/tests/test_measurement.py b/tests/test_measurement.py index 391619a..8587632 100644 --- a/tests/test_measurement.py +++ b/tests/test_measurement.py @@ -399,3 +399,78 @@ def test_current_environment_stamps_an_opt_in_cuda_device(monkeypatch): assert current_environment("cpu")["device"] == "cpu" assert current_environment("cuda")["device"] == "cuda" assert current_environment()["device"] == "cpu" + + +# -------------------------------------------------------------------------- +# Both JSON payloads declare the same schema +# -------------------------------------------------------------------------- + + +def _payload_keys(mode: str) -> set[str]: + """The keys `measure-leakage --format json` writes for one mode. + + Read from the source rather than by running a measurement: producing a + sweep payload costs 150 training runs, and what is under test is which + fields the writer emits. + """ + import ast + import inspect + + from iqforge import cli + + tree = ast.parse(inspect.getsource(cli.measure_leakage)) + for node in ast.walk(tree): + if not isinstance(node, ast.Dict): + continue + keys = {k.value for k in node.keys if isinstance(k, ast.Constant)} + if keys.issuperset({"measurement_schema", "mode"}): + for key, value in zip(node.keys, node.values, strict=True): + if ( + isinstance(key, ast.Constant) + and key.value == "mode" + and isinstance(value, ast.Constant) + and value.value == mode + ): + return keys + return set() + + +def test_both_measurement_payloads_declare_the_schema() -> None: + """A sweep payload written today must not look like a pre-schema one. + + `runs_from_payload` treats a missing `measurement_schema` as "unversioned, + assume compatible", which is right for payloads written before the field + existed. The sweep writer omitted it, so its output was indistinguishable + from those -- and would be the one that cannot be told apart when schema 2 + arrives. + """ + single = _payload_keys("single") + sweep = _payload_keys("sweep_stride") + + assert single, "could not find the single-cell payload" + assert sweep, "could not find the sweep payload" + assert "measurement_schema" in single + assert "measurement_schema" in sweep + + +def test_the_sweep_payload_carries_what_a_reader_needs() -> None: + """Both payloads share the fields `runs_from_payload` and a reader use.""" + shared = {"measurement_schema", "mode", "forced", "split_seeds", "train_seeds", "rows"} + assert shared <= _payload_keys("single") + assert shared <= _payload_keys("sweep_stride") + + +def test_the_row_serialiser_is_not_duplicated() -> None: + """One `_run_row`, not two hand-written copies. + + Three near-identical copies of a measurement call are how the seed lists + were dropped from one of them. The same shape of duplication in the row + writer would silently give the two payloads different fields. + """ + import inspect + + from iqforge import cli + + source = inspect.getsource(cli.measure_leakage) + assert source.count("_run_row(run) for run in runs") == 2 + assert '"train_windows": run.train_windows' not in source From 7bfc03dd87dd0bf00295b422a6c2884c60520848 Mon Sep 17 00:00:00 2001 From: Emre Date: Mon, 14 Sep 2026 00:43:13 +0300 Subject: [PATCH 5/8] Record the three Bolum C changes that never reached the CHANGELOG Three shipped changes had no entry. The most consequential one had none at all: the commit that restored the 15 seed pairs touched seven files and not this one, so the single change that fixed the sample-size regression -- and added the guard that stops a published grid shrinking again -- was visible only as a passing mention inside the parity-gate entry. Added: - The seed-pair restoration, the --split-seeds / --train-seeds flags, guard_artifact_rows, and the check_environment hole that made the guard ineffective on exactly the files it was meant to protect. - measurement_schema in the JSON payload, with the reason it exists: the same one read_manifest already has. - docs/release-notes/v0.5.0.md being labelled an unpublished draft, which is what makes __version__, CITATION.cff, the CHANGELOG and the notes agree. Co-Authored-By: Claude Opus 5 --- CHANGELOG.md | 29 +++++++++++++++++++++++++++++ 1 file changed, 29 insertions(+) diff --git a/CHANGELOG.md b/CHANGELOG.md index 3c95853..e7fbc4a 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -19,6 +19,12 @@ can tell whether the format it is looking at is one it understands. parity gate, and `artifacts/*.json` stay on CPU. `TrainingResult.environment` already recorded the device; the measurement path now stamps the device that was actually requested. +- **`measurement_schema` in the `measure-leakage --format json` payload**, and + `runs_from_payload` as the single reader of it. A reader that guesses at a + shape it does not recognise produces a plausible wrong answer, which is why + `read_manifest` already refuses a `manifest_schema` newer than it + understands; measurement payloads now get the same treatment. Both the + single-cell and the stride-sweep payloads declare it. - `iqforge measure-leakage` now accepts `--balance-by`, so the command path can run the same nuisance-balancing setup that the published synthetic measurement tables used. @@ -62,6 +68,29 @@ can tell whether the format it is looking at is one it understands. ### Changed +- **The published grids are measured at 15 seed pairs again, and cannot + silently shrink.** The Phase 5 migration hardcoded `[42]` and `[0]` into the + command's measurement path, cutting every grid from 15 seed pairs to 1. The + reduced grid reproduces the first pair exactly, so nothing that compared + values noticed; it was found by reading the code, not by reading a result. A + table built that way reports a standard error of zero and calls it a + measurement. + `measure-leakage` now takes `--split-seeds` and `--train-seeds`, defaulting to + the five split seeds and three training seeds every published table used. + They are flags rather than constants so a cheaper run is a visible choice, + and the count is printed with the result: a measurement whose sample size is + not on the page cannot be read. The three experiment scripts pass the same + lists, and `guard_artifact_rows` refuses, before anything is trained, to + overwrite a file under `artifacts/` with fewer runs than it already holds. + `check_environment` was part of the same failure: it returned quietly when a + checkpoint recorded no environment at all, which is the state every published + grid is in, so the guard had never protected one. It now refuses that case + instead of waving it through. +- **`docs/release-notes/v0.5.0.md` says it is an unpublished draft.** The file + read as a shipped release while `__version__`, `CITATION.cff` and the newest + released CHANGELOG section all said `0.4.0` and no `v0.5.0` tag existed. It + now names that state at the top and points at `[Unreleased]`, so the four + places that carry a version agree about which one is real. - **The experiment scripts and their tests no longer carry a hardcoded path.** `scripts/leakage_real.py`, `scripts/leakage_loraiq.py`, `tests/test_preflight.py` and `tests/test_measurement.py` all fell back to an absolute path inside one From 8bb6f8fac83541c284698e3db3a6a8a6e693478b Mon Sep 17 00:00:00 2001 From: Emre Date: Mon, 14 Sep 2026 00:43:43 +0300 Subject: [PATCH 6/8] Split the grouping roadmap item at the line where it was already done The item was unchecked and its own body said "`--group-by collection` is now shipped, so the first half of that is done". A checkbox that disagrees with the paragraph under it teaches a reader to stop trusting the checkboxes. It was one item covering two things: read the SigMF field, and decide whether the field needs an extension. The first shipped -- `collection` is the third scheme beside `path:` and `csv:`, with a test, and a recording that declares no collection stays its own unit. The second has not started, and is the one with a bar to clear: use it on a public dataset and record with evidence what `core:collection` could not express. They are now two items, ticked according to what is true of each. The reasoning note and the sigmf-python#159 / SigMF#233 precedent stay attached to the half they inform, which is the proposal rather than the scheme. Co-Authored-By: Claude Opus 5 --- ROADMAP.md | 16 +++++++++++----- 1 file changed, 11 insertions(+), 5 deletions(-) diff --git a/ROADMAP.md b/ROADMAP.md index 133db0d..a7df76a 100644 --- a/ROADMAP.md +++ b/ROADMAP.md @@ -72,7 +72,13 @@ After Now is done — still reliability-first: burst are slices of one continuous capture. See [docs/methodology.md](docs/methodology.md) §6. -- [ ] **`--group-by` by SigMF field.** Deliberately not in the first release. +- [x] **`--group-by collection`: group by the SigMF field.** Shipped as the + third scheme alongside `path:` and `csv:`. It reads `core:collection` + from the Global Object; a recording that declares none stays its own + unit. The note below is the reasoning that produced it, and the item + after it is the half that is still open. + + Deliberately not in the first release, for a reason worth keeping. The obvious third scheme would read a metadata key, the way `--balance-by` does. It was left out because it would have solved none of @@ -113,10 +119,10 @@ After Now is done — still reliability-first: An extension proposal is therefore about a *qualifier* on existing grouping, not a new grouping mechanism. - **`--group-by collection` is now shipped**, so the first half of that is - done. What is not done is the part that decides whether a proposal is - worth writing: using it on a public dataset and recording what it could - not express. +- [ ] **Measure what `core:collection` cannot express, then decide whether to + propose an extension.** The scheme is shipped; this is the part that + decides whether a proposal is worth writing — using it on a public + dataset and recording, with evidence, what it could not say. **The precedent says do not propose before that.** SigMF accepted an ML extension and then took it back. `rfml.sigmf-ext.md` was contributed by From ae0d750928b87438b3cf287be104d45ef8c36f57 Mon Sep 17 00:00:00 2001 From: Emre Date: Mon, 14 Sep 2026 00:44:54 +0300 Subject: [PATCH 7/8] Make methodology name the category numbers it is cited by MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit measure-leakage prints `category 4 ceiling (methodology 6.4)` and a reader follows that into §6. §6 did not contain the word "category" anywhere. The numbering matched for the four eliminated datasets by construction and nothing held it there, and past four it does not match at all -- which mattered more than the missing word: category 5 is a refusal, a structural leak audit named §6.5 is LoRaIQ, the one dataset in that section that was NOT eliminated Same digit, opposite meaning, and nothing warned a reader following a citation. §6 now opens with a cross-reference table covering all six categories, marks the two that have no case here and what they cite instead, states the 5 / 6.5 collision explicitly, and notes which categories --force can override. A test pins all three documents together: for every member of the Category enum, SPEC and methodology must each carry a table row with that number and that name, and the collision warning must still be there. Mutations: renumbering a category in the code alone fails it, deleting one methodology row fails it, softening the warning to "Note." fails it. The assertions compare with whitespace removed. These documents are hard-wrapped and "There is deliberately no category for it" already straddles two lines -- an exact-substring assertion passed only until the paragraph was re-wrapped. Co-Authored-By: Claude Opus 5 --- docs/methodology.md | 31 +++++++++++++++++++++++++++++++ src/iqforge/preflight.py | 3 +-- tests/test_preflight.py | 31 +++++++++++++++++++++++++++++++ 3 files changed, 63 insertions(+), 2 deletions(-) diff --git a/docs/methodology.md b/docs/methodology.md index 4824eb2..301db15 100644 --- a/docs/methodology.md +++ b/docs/methodology.md @@ -409,6 +409,37 @@ per class — enough that a recording-level split has something to split — a format the reader can be trusted on, and a task that is neither trivial nor impossible. Format turned out to be the easy one. +### How these cases map to the refuse categories + +`iqforge measure-leakage` refuses a dataset by **category number**, and cites +this section: `category 4 ceiling (methodology 6.4)`. The numbering was taken +from the cases below so the command could point at a paragraph. It matches for +the four eliminated datasets and **does not extend past them**, which is worth +stating here rather than leaving a reader to discover it: + +| command category | name | case here | what it cites | +|---|---|---|---| +| 1 | unreadable format | §6.1 AirID | this section | +| 2 | shared timestamp | §6.2 Vega-C | this section | +| 3 | physical independence | §6.3 DASH7 `ds_indoor` | this section | +| 4 | ceiling | §6.4 DASH7 `ds_indoor_cabled` | this section | +| 5 | structural leak | — | the `audit` LEAK finding that fired | +| 6 | cannot split | — | SPEC §5.6 | +| — | *(not refused)* | §6.5 LoRaIQ | — | + +**`category 5` and `§6.5` are not the same thing, and they point in opposite +directions.** Category 5 is a refusal: an `audit` LEAK that `--group-by` does +not already hold together. §6.5 is LoRaIQ — the dataset that passed, the one +case in this section that was *not* eliminated. There is deliberately no +category for it. Category 6 likewise has no case here; it is a split `build` +would refuse, and it cites SPEC §5.6. + +Categories 1 and 6 cannot be overridden with `--force`; 2 through 5 can. The +line is whether the category is an inference about what the recordings mean — +those are judgements a user may know better than the tool — or a statement +that no measurement can be constructed. SPEC §5.10 carries the full table and +the trigger for each. + ### 6.1 Case 1 — AirID **AirID** (GENESYS Lab, 4 UAV transmitters with deliberately distinct IQ diff --git a/src/iqforge/preflight.py b/src/iqforge/preflight.py index 1873161..87da057 100644 --- a/src/iqforge/preflight.py +++ b/src/iqforge/preflight.py @@ -671,8 +671,7 @@ def render_text(decision: Decision) -> str: out.extend( _field_lines( "estimate", - "not timed on this machine (torch missing, or the dummy-batch " - "probe failed)", + "not timed on this machine (torch missing, or the dummy-batch probe failed)", ) ) else: diff --git a/tests/test_preflight.py b/tests/test_preflight.py index 906cd47..ffc9fee 100644 --- a/tests/test_preflight.py +++ b/tests/test_preflight.py @@ -20,6 +20,7 @@ from iqforge.cli import _audit_folder, app from iqforge.grouping import resolve_group_keys from iqforge.preflight import ( + CATEGORY_META, Category, DecisionStatus, decide, @@ -646,3 +647,33 @@ def test_a_run_that_cannot_train_says_why(tmp_path: Path, monkeypatch: pytest.Mo assert "torch is not installed" in result.output assert "stops before training" not in result.output assert "MEASUREMENT" not in result.output + + +def test_the_three_documents_agree_on_the_category_numbers() -> None: + """Code, SPEC and methodology must not drift apart on what `category N` means. + + The command prints a number and a citation; a reader follows it into + methodology §6. Nothing kept those in step, and §6 did not mention + categories at all -- so `category 5` and `§6.5` looked like the same thing + while meaning opposite ones: a refusal, and the dataset that passed. + """ + import re + + root = Path(__file__).resolve().parent.parent + spec = (root / "SPEC.md").read_text(encoding="utf-8") + methodology = (root / "docs" / "methodology.md").read_text(encoding="utf-8") + + for category in Category: + name = CATEGORY_META[category][0] + row = rf"\|\s*{int(category)}\s*\|\s*{re.escape(name)}\s*\|" + assert re.search(row, spec), f"SPEC has no row for category {int(category)} '{name}'" + assert re.search(row, methodology), ( + f"methodology has no row for category {int(category)} '{name}'" + ) + + # The collision the table exists to head off. Compared with whitespace + # removed: these documents are hard-wrapped, so a phrase that fits on one + # line today can straddle two after an unrelated edit. + squeezed = "".join(methodology.split()) + assert "".join("`category 5` and `§6.5` are not the same thing".split()) in squeezed + assert "".join("There is deliberately no category for it.".split()) in squeezed From 0c1b7788a3212493fe354d30a13bec9357ce7180 Mon Sep 17 00:00:00 2001 From: Emre Date: Mon, 14 Sep 2026 01:08:41 +0300 Subject: [PATCH 8/8] Update the assertion that still expected the sentence we removed test_loraiq_pattern_is_not_refused_when_grouped asserted "stops before training" in its torch-free branch. That branch is unreachable on a machine with torch, so the full local suite passed after the output changed and CI's test (3.11) and test (3.12) jobs are what caught it. That is precisely the failure this PR describes one commit earlier -- an assertion whose coverage depends silently on the environment -- and I walked into it while fixing it. Worth saying rather than quietly amending. The branch now asserts what the report actually prints, the torch branch additionally pins `started yes.`, and the old sentence is asserted absent on both paths so neither can drift back. Co-Authored-By: Claude Opus 5 --- tests/test_preflight.py | 8 +++++++- 1 file changed, 7 insertions(+), 1 deletion(-) diff --git a/tests/test_preflight.py b/tests/test_preflight.py index ffc9fee..9242106 100644 --- a/tests/test_preflight.py +++ b/tests/test_preflight.py @@ -231,8 +231,14 @@ def test_loraiq_pattern_is_not_refused_when_grouped(tmp_path: Path) -> None: assert "REFUSED" not in result.output if _torch_installed(): assert "MEASUREMENT" in result.output + assert "started yes." in result.output else: - assert "stops before training" in result.output + # The report names the obstacle rather than asserting a policy. This + # branch is unreachable on a machine with torch, which is why the old + # sentence survived here after the code stopped printing it. + assert "started no. torch is not installed" in result.output + assert "MEASUREMENT" not in result.output + assert "stops before training" not in result.output @pytest.mark.skipif(loraiq_skip_reason() is not None, reason=loraiq_skip_reason() or "")