diff --git a/SPEC-ISSUES.md b/SPEC-ISSUES.md index 342d4a8..aa59426 100644 --- a/SPEC-ISSUES.md +++ b/SPEC-ISSUES.md @@ -87,6 +87,18 @@ to Milestone 2, when primitive-frequency extraction exists; committing the value identification dimension it has named but not yet quantified — reserving weight under an explicit `unmeasured` status is the mechanism proposed here. +**Resolved for the reconstruction (Phase 9, SI-026).** `primitive_frequency_mix` is now +measured and committed: `iso-002.yaml` at `grammar_version 1.1.0` gives it +`status: measured`, an `expected` instance-share vector and a scalar `tolerance`, and the +locus weights rebalance (ink 0.75 → 0.60, mix reserved-0.25 → committed-0.40). The +`grammar_version` bump (not a schema change) is exactly the mechanism this entry predicted. +Two caveats stand. (a) The committed vector describes the **reconstruction's** composition +model (the generator's uniform-type dialect, measured back out), **not** Studio.Build's +original proportions — measuring those would need the real assets (SI-018); the sheet says so. +(b) The reservation-then-commit round trip worked as designed, but the spec still owes the +general rule (how any audit records a named-but-unquantified identification dimension, and +what committing it later requires); this is one worked instance, not that rule. + ## SI-009 · §3.7 · decided-here — Canonical-peak rule only bites on numeric expected values Spec §3.7 lists canonical peaks as numbers (0°/45°/90°, ratios 2.0/1.5/φ, duty 1/2·1/3). @@ -432,6 +444,31 @@ measured in a shared feature space. **Two consequences the spec should price in: distance is single-dimensional and therefore that same-family collision is unguarded — because a one-feature signature has no minimum distance to defend once that feature is shared. +**Dangerous half substantially closed (Phase 9, SI-026).** Consequence 2's warning — "two grid +dialects that share an ink set are indistinguishable" — is now measured against, not just noted. +`primitive_frequency_mix` is committed as a second identification feature (SI-008 resolved), so a +same-ink impostor must also match the composition. Rebuilt as constrained generator compositions in +002's **exact** master inks (`generator.grid.render_with_truth(..., types=[...])`), **no same-ink +single-primitive impostor reaches identified at any fragment fraction** (0.2/0.35/0.5/1.0; 0 of 624 +impostor fragments): **all-circles** peaks at aggregate 0.632 (frac 0.2) and sits at 0.554 from frac +0.5 up; **all-stripes** at 0.554 everywhere; the **nearest** family — all-staircase / all-filled, +whose depth-2-cap remnants (SI-019) read as filled cells — peaks at **0.684** (frac 0.35–0.5), still +below the 0.70 line. That family is the honest bound of an L1-on-mix feature and its margin (0.016) +is the thinnest; it stays below identified only because the classifier's recall fixes (half-cell-bar +stripe rule, stadium extent band, disc-IoU circle gate — SI-026) spread the committed expected vector +across all five bins, pushing the filled+staircase two-bin L1 floor to ~1.0, above the committed +tolerance 0.95. Genuine held-out surfaces (seeds outside the derivation corpus) identify at **0.92** +full-surface (0.89 on the report corpus), 0.76 at frac 0.5, 0.63 at frac 0.35. The general cost is +inherent, not incidental — a fragment too small to measure the mix makes a genuine surface and a +same-ink impostor **identical** (same inks, second carrier unobservable) — but it is priced honestly, +not punished: border-clipped components are excluded from the count and the scorer's n-dependent +tolerance (SI-026) treats small-n L1 as sampling noise, so genuine frac-0.2 fragments hold a mean +aggregate of 0.64 (solid candidate; identified rate 0.20 vs ~0.91 for the ink-only v1.0 sheet — that +delta is exactly the price of the closed gap). **Spec consequence (unchanged, sharpened):** §5's +grid-vs-grid minimum distance is now two-dimensional but still bounded by primitive +measurement-confusability; the spec should say whether a family may lean on a mix feature whose +discrimination is uneven across the alphabet, and consequence 1 (the control corpus) remains open. + ## SI-023 · §11/§10 · decided-here — The field battery runs the pipeline UNMODIFIED on raw photos, and puts its human-readable label in the page margin Spec §11 makes L3 depend on recognition "under the hostile-conditions test battery (print, camera, @@ -473,3 +510,76 @@ relationship/ordering colour signature (Milestone 2, §5) can only be validated from captures rich enough to estimate the correction, so the field-battery protocol must state a minimum per-feature framing (here: ≥3 inks in frame for the WB-diagonal check) rather than treating "photograph the surface under varied conditions" as sufficient for every question the battery is meant to answer. +## SI-026 · §3.7/§5 · decided-here — Measuring and committing the primitive-frequency mix as a second identification carrier + +Audit 002 §6 point 2 names "the primitive mix and composition statistics" as one of the two carriers +of ISO's identification load, and SI-008 reserved weight for it under `status: unmeasured` pending a +measurer. Phase 9 builds that measurer and commits the value, to close the dangerous half of SI-022. +Several sub-decisions, all v0 choices the spec does not fix: + +1. **Per-ink footprint separation.** The obvious measurer — connected components over one all-ink mask + — fails: at the audit's densities (0.35–0.55) overlapping and adjacent primitives fuse into a few + giant blobs and only ~3 % of instances stay separable, giving a mix estimate too noisy to carry + identification. Instead each ink's footprint is reconstructed from the committed overprint + arithmetic (a pixel carries ink X where it is X or `multiply(X, Y)` for another extracted ink Y), + so components merge only with **same-ink** neighbours (rare in a 4–6 ink subset). A vertical + morphological close regroups a stripe block's separated bars before classification (which runs on + the un-closed mask so the M/2 rhythm survives). + +2. **Interior-only counting (edge exclusion).** A component touching the image border may be a + primitive clipped by the fragment boundary; its shape statistics are corrupted, so it is excluded + and `n` counts **interior** classified instances only (the exclusion count is reported in the + measurement detail). The cost — whole edge-cell primitives on a full surface are excluded too, + shrinking n (typical full-surface interior n ≈ 16 on the derivation grids) — is priced consistently: + the committed expected vector is derived interior-only, and the scorer widens the tolerance at + small n (point 4) instead of punishing sampling noise as disagreement. + +3. **Honest unclassified share, zero cross-confusion, full recall.** A component that does not match a + single primitive's module footprint (a merged multi-primitive blob, or an off-alphabet shape) is + counted **unclassified**, never forced into a bin; the mix is the share over classified instances + and the unclassified share is reported. On isolated primitives the confusion matrix is **diagonal**: + recall 1.0 AND precision 1.0 for all five types — three targeted rules close what were systematic + recall gaps: (a) a **half-cell-bar rule** catches 1-cell-tall stripe blocks (one solid M/2-tall bar, + no periodicity to vote on — no other alphabet form is half-a-cell tall); (b) the stadium extent band + extends to 0.955 (a 4-cell stadium fills 0.946 of its bbox and was silently missed at 0.94); (c) a + **disc-IoU gate** (≥ 0.90 against the bbox-inscribed disc) keeps depth-2-cap-clipped cell remnants + (SI-019) out of the circle bin — the extent band alone was spoofable, and that leakage had flattered + the all-staircase impostor's L1. End-to-end (real compositions, merging + edge exclusion included) + the mean absolute per-type share error is ~0.11, with the residual bias (filled_cell + over-represented by survival and cap remnants) consistent and **baked into the committed expected + vector**, which genuine surfaces match. + +4. **Expected, tolerance and n-scaling from data.** The `expected` vector `[0.30, 0.10, 0.24, 0.20, + 0.16]` (order: filled_cell, inscribed_circle, stripe_bar, staircase_diagonal, rounded_cap) is the + **mean** interior-only measured mix over a genuine corpus (seeds × densities 0.35/0.45/0.55 × + modules 40/48/64, varying ink subsets; 108 surfaces). Agreement is clipped-linear (SI-001) on + `L1(measured, expected) / tolerance_eff` with the **n-dependent effective tolerance** + + tolerance_eff = tolerance × max(1, sqrt(N_REF_PRIMITIVES / n)), N_REF_PRIMITIVES = 16 + + (a share vector from n instances of a 5-bin multinomial has ~sqrt(1/n) L1 sampling deviation, so a + small fragment's larger L1 is noise, not disagreement; at n ≥ n_ref the committed value applies + unchanged, so full-measurement impostor rejection is never loosened). Confidence is n-scaled + (SI-002, `sample_unit: primitives_observed`) with **k = 2**, deliberately harder than the scalar + features' k = 1: the widened tolerance raises small-n agreement, so the saturation must discount + harder or a lucky small-n impostor fragment could cross the identified line — with k = 2 the + worst-case (L1 exactly at the impostor floor, all six inks visible) stays below 0.70 at every n. + The committed `tolerance = 0.95` sits in the window bounded **below** by genuine coverage (max + n-normalised genuine L1 `L1 × min(1, sqrt(n/n_ref))` over derivation + held-out corpora is 0.83 → + ~15 % margin) and **above** by impostor rejection (the nearest impostor family's measured L1 floor + is ~1.0 — the filled+staircase two-bin floor `2 × (1 − 0.30 − 0.20)` — so its full-measurement mix + agreement is 0). The claim's working carries n, tolerance_eff and the L1. + +5. **Weight split from data.** Ink 0.60 / mix 0.40 (from 0.75 / reserved-0.25). The split is set so a + same-ink impostor's best case — ink agreement ~1, mix agreement ~0 — scores `0.60 × saturation(inks)` + ≈ 0.55, below the 0.70 identified line by ~0.15, while genuine surfaces, whose mix also agrees, clear + it. A higher ink weight thins the impostor margin; a lower one costs genuine identification. The + `max`-combinator, the 0.70 threshold and the k constants are all prior v0 choices (SI-013, SI-002). + +**Spec consequence:** the meta-grammar should say (a) whether a composition/primitive-mix feature is +enrollable at all (SI-018 asked this from the generator side; here it is answered "yes, as a +reconstruction-scoped measured feature"); (b) how its expected/tolerance are to be derived and +published so two recognisers agree; and (c) that its discrimination may be **uneven across the +alphabet** (strong where compositions differ in primitive *vocabulary*, weak where they reconcentrate +the same primitives) — so a family leaning on it must publish where its mix signature is load-bearing, +exactly as §3.7 already asks each grammar to declare *where* its signature lives. diff --git a/generator/grid.py b/generator/grid.py index eaa8d6d..2dd30e2 100644 --- a/generator/grid.py +++ b/generator/grid.py @@ -453,6 +453,7 @@ def render_with_truth( seed: int, ink_subset=None, density: float = 0.45, + types=None, ): """Render a grid composition and its GroundTruth; return (surface, truth). @@ -465,6 +466,12 @@ def render_with_truth( subset of the master set (audit s3). density composition density (audit s3 dialect freedom, SI-018): the number of placed primitive instances is round(density*cols*rows). + types optional restriction of the primitive pool to a subset of + ``PRIMITIVE_TYPES`` (composition freedom, SI-018). Default (None) + draws uniformly from all five, the audit-neutral dialect. A + single-type restriction builds a same-ink/different-composition + impostor for the SI-022/SI-026 distinctiveness test; the genome + (grid, alphabet, rhythm, overprint, ground) is untouched. Surface is (H, W, 3) BGR uint8 on a white ground, built by multiplying flat ink layers (exact overprint, depth capped at 2 -- module docstring). @@ -498,10 +505,17 @@ def render_with_truth( label_first = np.full((h, w), -1, dtype=np.int64) label_second = np.full((h, w), -1, dtype=np.int64) + if types is None: + type_pool = PRIMITIVE_TYPES + else: + type_pool = tuple(t for t in PRIMITIVE_TYPES if t in set(types)) + if not type_pool: + raise ValueError("types must name at least one of PRIMITIVE_TYPES") + n_primitives = max(1, round(density * cols * rows)) primitives = [] for _ in range(n_primitives): - type_ = PRIMITIVE_TYPES[int(rng.integers(len(PRIMITIVE_TYPES)))] + type_ = type_pool[int(rng.integers(len(type_pool)))] ink_k = int(rng.integers(len(ink_hexes))) mask, cells, params = _place_primitive(rng, rows, cols, module_px, type_) diff --git a/grammars/iso-002.yaml b/grammars/iso-002.yaml index ab1b940..444f902 100644 --- a/grammars/iso-002.yaml +++ b/grammars/iso-002.yaml @@ -16,7 +16,18 @@ spec_version: "0.1" sheet: id: iso-002 name: ISO (Studio.Build) -- audit reconstruction - grammar_version: "0.0.0" # reconstruction, not an owner-committed version + # 1.1.0: commits primitive_frequency_mix as a measured identification carrier + # (was reserved-unmeasured at weight 0.25, SI-008), and rebalances the locus + # weights (ink 0.75->0.60, mix ->0.40) so a same-ink/different-composition + # impostor no longer identifies -- closing the dangerous half of SI-022 + # (SI-026). IMPORTANT (SI-018): the committed mix vector describes THIS + # audit reconstruction's composition model (the generator's uniform-type + # dialect, measured back out), NOT Studio.Build's original proportions. The + # audit s6.2 names the primitive mix as an identification carrier but publishes + # no measured proportions; measuring the real identity's mix would need + # Studio.Build's actual assets. This sheet commits what it can measure and + # says so. Still an audit reconstruction, never an enrolment (spec 6.2). + grammar_version: "1.1.0" # reconstruction, not an owner-committed version status: audit-reconstruction issuer: null # not enrolled; the owner alone may issue identity_endpoint: null @@ -98,8 +109,12 @@ variation_model: - ground signature_locus: - # audit s6: after discounting canonical structure, the ink set carries the load; - # the primitive/composition mix carries the rest but is unmeasured here. + # audit s6: after discounting canonical structure, identification is carried by + # TWO measured features -- the ink set (point 1, the address) and the primitive + # mix (point 2, the proportions the alphabet is used in). At v1.1.0 both are + # measured; the weights are split so neither alone reaches the identified + # threshold, so a same-ink impostor with a different composition cannot pass on + # inks alone (SI-022 / SI-026). features: - id: ink_set measure: ink_set_match @@ -116,18 +131,48 @@ signature_locus: tolerance: delta_e: 10 role: identification - weight: 0.75 # audit s6 point 1: the inks are the address + # 0.60 (was 0.75). Lowered so ink agreement alone (a same-ink impostor's + # best case) sits below the 0.70 identified threshold; the mix must also + # agree to identify (SI-026). Still the larger share: inks are the address. + weight: 0.60 sample_unit: inks_observed status: measured - id: primitive_frequency_mix measure: primitive_frequency_mix - expected: null # audit s6 point 2 names it, publishes no proportions - tolerance: null + # The INTERIOR instance-share vector over the five alphabet primitives, in + # the order [filled_cell, inscribed_circle, stripe_bar, staircase_diagonal, + # rounded_cap]. Derived (SI-026) as the MEAN measured mix over a genuine + # corpus (seeds x densities 0.35/0.45/0.55 x modules 40/48/64, varying ink + # subsets; 108 surfaces), with border-touching components EXCLUDED from the + # measurement (a clipped shape biases the estimate) -- the same rule the + # measurer applies at recognition time. This is the RECONSTRUCTION's + # composition model (SI-018), not the original identity's. Residual biases + # (filled_cell over-represented by survival -- compact 1x1 cells merge less + # than multi-cell primitives -- and their remnants under the depth-2 cap, + # SI-019) are consistent and baked into the vector genuine surfaces match. + expected: [0.30, 0.10, 0.24, 0.20, 0.16] + # Clipped-linear agreement on L1(measured, expected) / tolerance_eff (SI-001 + # convention), where tolerance_eff = tolerance x max(1, sqrt(n_ref / n)) + # with n_ref = 16 (the derivation corpus's typical full-surface interior + # instance count; recogniser/score.py N_REF_PRIMITIVES): a share vector from + # n instances has ~sqrt(1/n) L1 sampling noise, so small fragments widen the + # tolerance instead of being punished as disagreement (the harder + # primitives_observed saturation k=2 keeps small-n estimates from + # dominating), while at n >= n_ref the committed value applies unchanged -- + # full-measurement impostor rejection is never loosened. The committed 0.95 + # is chosen from data (SI-026): bounded BELOW by genuine coverage (max + # n-normalised genuine L1 over derivation + held-out corpora is 0.83, so + # 0.95 covers every genuine surface with ~15% margin) and ABOVE by impostor + # rejection. The nearest same-ink impostor family (all-staircase / + # all-filled, whose depth-2-cap remnants read as filled cells) has a + # measured full-surface L1 floor of 1.0 -- above tolerance, so its mix + # agreement is 0 and its aggregate rests on the ink weight alone (~0.55). + tolerance: 0.95 role: identification - weight: 0.25 # weight reserved; see SPEC-ISSUES SI-008 + weight: 0.40 # committed (was reserved-unmeasured, SI-008) sample_unit: primitives_observed - status: unmeasured # recogniser skips and renormalises at runtime + status: measured - id: grid_module measure: grid_module_detect diff --git a/recogniser/measure_grid.py b/recogniser/measure_grid.py index 9693027..98877cc 100644 --- a/recogniser/measure_grid.py +++ b/recogniser/measure_grid.py @@ -53,13 +53,44 @@ report the duty; if no striped region is found we report it unobserved (n = 0), never a failure. - * PRIMITIVE FREQUENCY MIX (audit s6 point 2; sheet ``status: unmeasured``, - SI-008). A rough connected-component classifier bins components into - {circle_like, rectangle_like, stripe_like, diagonal_like, other}. The sheet - declares this feature unmeasured, so the scorer SKIPS it regardless (SI-008); - we expose the measured mix in the working as "measured but uncommitted", - with an honest caveat that overlapping primitives merge into one component so - the classifier is coarse. We deliberately do not over-engineer it. + * PRIMITIVE FREQUENCY MIX (audit s6 point 2; committed as an identification + carrier at grammar_version 1.1.0, SI-008/SI-026). We classify inked regions + into the five alphabet primitives (filled_cell, inscribed_circle, stripe_bar, + staircase_diagonal, rounded_cap) by shape statistics in MODULE units and + report the share of primitive INSTANCES by type (matching + ``GroundTruth.primitive_counts`` semantics) plus n = instances observed. + + Separation is the hard part: overlapping primitives merge into one connected + component. Three moves keep the estimate honest (SI-026): + 1. PER-INK LAYERING. Instead of one all-ink mask (where every adjacent or + overlapping primitive fuses into a few giant blobs -- only ~3% of + instances stay separable), we reconstruct each ink's footprint from the + overprint arithmetic: a pixel "carries ink X" when it is exactly X (a + depth-1 region) OR exactly multiply(X, Y) for some other extracted ink Y + (a depth-2 overlap). Components then merge only with SAME-ink neighbours + (rare: each instance draws a random ink from a 4-6 subset). + 2. INTERIOR-ONLY COUNTING. A component touching the image border may be a + primitive clipped by the fragment boundary; its shape statistics are + corrupted, so it is excluded (count reported) and n counts INTERIOR + classified instances only. The scorer's n-dependent tolerance (score.py, + SI-026) absorbs the sampling noise of the smaller n. + 3. HONEST UNCLASSIFIED SHARE. A component that does not match a single + primitive's module footprint (a merged multi-primitive blob, or a shape + off the alphabet) is counted UNCLASSIFIED, never forced into a bin. The + mix is the share over CLASSIFIED instances; the unclassified share is + reported so the caller sees how much was absorbed. Tolerance on the sheet + (SI-026) covers the residual bias this leaves. + + The classifier never cross-confuses (measured recall 1.0 AND precision 1.0 + per type on isolated primitives -- single-bar stripe blocks are caught by the + half-cell-bar rule, 4-cell stadiums by the widened extent band, and the + disc-IoU gate keeps clipped-cell remnants out of the circle bin): on real + compositions a merged or cap-clipped (SI-019) component becomes unclassified + or a filled-cell remnant, never a wrong multi-cell label. The residual bias + (filled_cell over-represented by survival and remnants) is consistent and + baked into the committed expected vector, which genuine surfaces match and a + same-ink impostor with a different composition does not (its measured L1 + floor sits at ~1.0, above the committed tolerance). * INK-SET matching is NOT done here. The two-path (absolute + relationship / white-balance-gain) match lives in ``recogniser/score.py`` because it needs @@ -142,6 +173,63 @@ # within this fraction of the module of the expected M/2 bar/gap. STRIPE_RUN_TOL_FRAC = 0.30 +# --- primitive-mix classifier constants (SI-026) ---------------------------- + +# The five alphabet primitives, in the generator's PRIMITIVE_TYPES order. The +# committed mix vector (sheet ``expected``) and the measured value are lists in +# THIS order, so the scorer's L1 is index-aligned. +PRIMITIVE_MIX_ORDER = ( + "filled_cell", + "inscribed_circle", + "stripe_bar", + "staircase_diagonal", + "rounded_cap", +) + +# A component whose area is below this fraction of a module cell is noise (a stray +# overprint sliver, a degraded edge) and is skipped before classification. +MIX_MIN_AREA_FRAC = 0.10 + +# A single primitive spans at most a few modules (staircase n<=5, stripe 3x3). +# A component wider/taller than this in modules is a merged multi-primitive blob +# and is counted unclassified, never binned. +MIX_MAX_FOOTPRINT_MODULES = 5 + +# A non-stripe component's bounding box must sit within this many modules of an +# integer cell footprint to be a clean single primitive (else unclassified). +MIX_FOOTPRINT_TOL_MODULES = 0.33 + +# Extent (filled area / bbox area) bands per primitive, in module-normalised +# shape space (all measured, not tuned to fit): a solid cell fills ~1.0, an +# inscribed circle ~pi/4, a stripe block ~0.5, a staircase ~1/n, a stadium +# ~0.85-0.93. Bands are set wide enough to absorb rasterisation yet never overlap +# between types (confusion matrix: precision 1.0, no cross-type errors). +MIX_STRIPE_EXTENT = (0.30, 0.72) +MIX_STAIRCASE_EXTENT_MAX = 0.62 +MIX_CIRCLE_EXTENT = (0.64, 0.88) +# Stadium upper bound 0.955: an n-cell stadium fills (n-1+pi/4)/n of its bbox = +# 0.893 (n=2), 0.928 (n=3), 0.946 (n=4) -- the earlier 0.94 bound silently missed +# every 4-cell stadium (SI-026 recall fix). +MIX_STADIUM_EXTENT = (0.70, 0.955) +MIX_FILLED_EXTENT_MIN = 0.85 +MIX_QUARTER_EXTENT_MAX = 0.985 + +# Single-bar stripe blocks: a 1-cell-tall stripe block renders as ONE solid bar, +# height M/2 -- no periodicity to vote on, but the alphabet contains no other +# half-cell-tall solid form, so a solid bar whose bbox height is ~M/2 and width a +# whole number of modules IS a stripe bar (SI-026 recall fix; previously ~1/3 of +# stripe placements were systematically missed). +MIX_HALF_BAR_HEIGHT = (0.35, 0.65) # bbox height in modules +MIX_HALF_BAR_EXTENT_MIN = 0.90 # solid bar + +# An inscribed-circle candidate must also BE a disc: IoU between the component +# and the ideal disc inscribed in its bounding box. Exact (1.0) on clean renders; +# the margin absorbs mild degradation. This gate exists because the extent band +# alone is spoofable -- a cell partially clipped by the depth-2 overprint cap +# (SI-019) can land in the circle extent band without being disc-shaped, and that +# leakage flattered a same-ink all-staircase impostor's mix (SI-026 / SI-022). +MIX_CIRCLE_DISC_IOU = 0.90 + # --- multiply (matches generator.grid exactly) ------------------------------ @@ -457,54 +545,236 @@ def _run_lengths(mask_1d): # --- primitive frequency mix (SI-008: measured but uncommitted) ------------- -def measure_primitive_mix(surface: np.ndarray, module_px): - """Rough connected-component classification of the primitive mix. +def _classify_primitive(comp: np.ndarray, module_px: int): + """Classify one connected component mask into a primitive type, or ``None``. - Returns ``(mix, n_components, detail)``. Bins each ink component into - {circle_like, rectangle_like, stripe_like, diagonal_like, other} by simple - shape statistics (extent, aspect, circularity, and internal stripe periodicity). - This is deliberately coarse: overlapping primitives merge into one component, - so the counts are indicative, not exact -- honesty note carried in the detail. - The sheet declares ``primitive_frequency_mix`` unmeasured (SI-008), so the - scorer skips this regardless; it is exposed only as working evidence. + ``comp`` is a bool bounding-box mask of a single ink component; ``module_px`` + the measured module M. Returns one of ``PRIMITIVE_MIX_ORDER`` or ``None`` when + the component is not a clean single primitive (a merged blob or off-alphabet + shape -> unclassified). Decisions use only module-normalised shape statistics + (footprint in cells, extent, internal stripe periodicity, a corner test) and + are ordered so the distinctive stripe periodicity is tested before the + footprint gate a half-integer stripe bbox would fail (SI-026). + """ + bh, bw = comp.shape + area = int(comp.sum()) + if area == 0: + return None + bw_m = bw / module_px + bh_m = bh / module_px + nw = int(round(bw_m)) + nh = int(round(bh_m)) + if nw < 1 or max(nw, nh) > MIX_MAX_FOOTPRINT_MODULES: + return None + extent = area / float(bw * bh) + + # (1) Stripe FIRST. Regular M/2 ink runs separated by ~M/2 gaps (period ~M) in + # several columns -- a distinctive periodicity, tested before the footprint + # gate because a multi-bar block's bbox height is the half-integer (h-0.5)M. + if MIX_STRIPE_EXTENT[0] <= extent <= MIX_STRIPE_EXTENT[1] and _is_striped(comp, module_px): + return "stripe_bar" + + # (1b) Single-bar stripe block: a 1-cell-tall block is ONE solid bar (height + # M/2, width a whole number of modules) with no periodicity to vote on. No + # other alphabet form is half-a-cell tall, so the shape alone identifies it. + if MIX_HALF_BAR_HEIGHT[0] <= bh_m <= MIX_HALF_BAR_HEIGHT[1] \ + and abs(bw_m - nw) <= MIX_FOOTPRINT_TOL_MODULES \ + and extent >= MIX_HALF_BAR_EXTENT_MIN: + return "stripe_bar" + + # Remaining types are single cells on the grid: bbox must be near-integer. + if nh < 1 or abs(bw_m - nw) > MIX_FOOTPRINT_TOL_MODULES \ + or abs(bh_m - nh) > MIX_FOOTPRINT_TOL_MODULES: + return None + + # (2) Staircase: square-ish multi-cell footprint left sparsely filled (~1/n). + if nw >= 2 and nh >= 2 and abs(nw - nh) <= 1 and extent <= MIX_STAIRCASE_EXTENT_MAX: + return "staircase_diagonal" + + # (3) Inscribed circle: square 1x1 or 2x2 footprint, extent ~pi/4, AND the + # component actually IS the inscribed disc (IoU gate, MIX_CIRCLE_DISC_IOU): + # the extent band alone is spoofable by depth-2-clipped cell remnants. + if nw == nh and nw in (1, 2) and MIX_CIRCLE_EXTENT[0] <= extent <= MIX_CIRCLE_EXTENT[1]: + if _disc_iou(comp) >= MIX_CIRCLE_DISC_IOU: + return "inscribed_circle" + return None # circle-sized but not disc-shaped: unclassified, never binned + + # (4) Rounded cap -- stadium: elongated 1xn / nx1 footprint, high extent with + # rounded (area-shaving) ends. + if ((nw >= 2 and nh == 1) or (nh >= 2 and nw == 1)) \ + and MIX_STADIUM_EXTENT[0] <= extent <= MIX_STADIUM_EXTENT[1]: + return "rounded_cap" + + # (5) 1x1 near-solid: filled_cell, unless exactly one corner is shaved -> the + # rounded_cap quarter-round. + if nw == 1 and nh == 1 and extent >= MIX_FILLED_EXTENT_MIN: + q = max(2, module_px // 4) + corners = [comp[:q, :q], comp[:q, -q:], comp[-q:, :q], comp[-q:, -q:]] + empty = sum(1 for c in corners if float(c.mean()) < 0.55) + if empty == 1 and extent < MIX_QUARTER_EXTENT_MAX: + return "rounded_cap" + return "filled_cell" + return None + + +def _is_striped(comp: np.ndarray, module_px: int) -> bool: + """True when a component's columns carry the M/2 bar / M/2 gap stripe rhythm. + + Samples several columns; a column votes when its ink runs are all ~M/2 and its + interior gaps ~M/2 (audit s2 rhythm). Requires a majority of sampled columns + to agree, so a lone bar (a 1-cell-tall block) or an incidental gap never reads + as a stripe block. """ - ink = (surface != 255).any(axis=2).astype(np.uint8) - n_labels, labels, stats, _ = cv2.connectedComponentsWithStats(ink, connectivity=8) - bins = {"circle_like": 0, "rectangle_like": 0, "stripe_like": 0, - "diagonal_like": 0, "other": 0} - min_area = max(16, (module_px * module_px) // 8) if module_px else 16 - n_components = 0 - for lab in range(1, n_labels): - area = int(stats[lab, cv2.CC_STAT_AREA]) - if area < min_area: + bh, bw = comp.shape + tol = 0.35 * module_px + half = module_px / 2.0 + votes = 0 + checked = 0 + for frac in (0.2, 0.35, 0.5, 0.65, 0.8): + cx = int(bw * frac) + if cx >= bw: continue - n_components += 1 - bw = int(stats[lab, cv2.CC_STAT_WIDTH]) - bh = int(stats[lab, cv2.CC_STAT_HEIGHT]) - extent = area / float(bw * bh) if bw * bh else 0.0 - aspect = bw / float(bh) if bh else 0.0 - comp = (labels[stats[lab, cv2.CC_STAT_TOP]:stats[lab, cv2.CC_STAT_TOP] + bh, - stats[lab, cv2.CC_STAT_LEFT]:stats[lab, cv2.CC_STAT_LEFT] + bw] - == lab) - # Internal stripe periodicity: a stripe block's centre column alternates. - mid = comp[:, bw // 2] if bw else np.array([], dtype=bool) - runs = [n for on, n in _run_lengths(mid) if on] - striped = (module_px and len(runs) >= 2 - and all(abs(n - module_px / 2.0) <= 0.4 * module_px for n in runs)) - if striped: - bins["stripe_like"] += 1 - elif extent >= 0.95 and 0.8 <= aspect <= 1.25: - bins["rectangle_like"] += 1 - elif 0.70 <= extent <= 0.85 and 0.8 <= aspect <= 1.25: - bins["circle_like"] += 1 # inscribed circle fills ~pi/4 of its box - elif extent <= 0.6: - bins["diagonal_like"] += 1 # staircase leaves the box sparsely filled - else: - bins["other"] += 1 - detail = {"caveat": "overlapping primitives merge into one component; counts " - "are coarse and this feature is uncommitted (SI-008)", - "min_component_area": min_area} - return bins, n_components, detail + checked += 1 + runs = _run_lengths(comp[:, cx]) + interior = runs[1:-1] if len(runs) >= 3 else [] + on = [n for o, n in runs if o] + gap = [n for o, n in interior if not o] + if len(on) >= 2 and all(abs(n - half) <= tol for n in on) \ + and (not gap or all(abs(n - half) <= tol for n in gap)): + votes += 1 + return checked >= 2 and votes >= max(2, checked - 1) + + +def _disc_iou(comp: np.ndarray) -> float: + """IoU between a component and the ideal disc inscribed in its bounding box. + + The generator's circle mask is a pixel-centre distance test against the + inscribed disc, so a true inscribed circle scores ~1.0; a clipped cell + remnant that merely lands in the circle EXTENT band scores far lower. + Deterministic, pure numpy. + """ + bh, bw = comp.shape + if bh == 0 or bw == 0: + return 0.0 + ys = np.arange(bh, dtype=np.float64) + 0.5 + xs = np.arange(bw, dtype=np.float64) + 0.5 + X, Y = np.meshgrid(xs, ys) + r = min(bw, bh) / 2.0 + disc = (X - bw / 2.0) ** 2 + (Y - bh / 2.0) ** 2 <= r ** 2 + inter = float(np.logical_and(comp, disc).sum()) + union = float(np.logical_or(comp, disc).sum()) + return inter / union if union else 0.0 + + +def _ink_present_mask(surface_i64: np.ndarray, ink_bgr, other_inks) -> np.ndarray: + """Reconstruct the footprint of one ink from the overprint arithmetic (SI-026). + + A pixel carries ink ``ink_bgr`` when it is exactly that colour (a depth-1 + region) or exactly ``multiply(ink_bgr, other)`` for some other extracted ink + (a depth-2 overlap). Reconstructing the overlap regions keeps a primitive's + footprint whole, so components merge only with SAME-ink neighbours. + """ + x = np.asarray(ink_bgr, dtype=np.int64) + present = np.all(surface_i64 == x, axis=2) + for y in other_inks: + if np.array_equal(x, y): + continue + present |= np.all(surface_i64 == _multiply(x, y), axis=2) + return present.astype(np.uint8) + + +def measure_primitive_mix(surface: np.ndarray, module_px, ink_set): + """Classify the primitive-frequency mix (audit s6 point 2; SI-026). + + ``ink_set`` is the extracted ink set ``[(bgr, fraction), ...]`` from + ``separate_products``. Returns ``(mix, n_classified, detail)`` where ``mix`` is + a list of five instance shares in ``PRIMITIVE_MIX_ORDER`` (summing to 1 over + classified instances, or all-zero when none classified) and ``n_classified`` + the number of instances observed (the ``primitives_observed`` sample unit). + + Method (SI-026): per-ink footprint reconstruction (``_ink_present_mask``), a + vertical morphological close to regroup a stripe block's separated bars into + one component, connected components, then ``_classify_primitive`` on the + ORIGINAL (un-closed) per-ink mask so the stripe rhythm survives. Components + that are not a clean single primitive are counted unclassified. The detail + carries the raw counts, the unclassified count/share and the ink layers used. + """ + counts = {t: 0 for t in PRIMITIVE_MIX_ORDER} + if not module_px or module_px < 4 or not ink_set: + return [0.0] * len(PRIMITIVE_MIX_ORDER), 0, { + "reason": "no module/ink set to anchor the primitive mix", + "counts": counts, "n_unclassified": 0} + + surface_i64 = surface.astype(np.int64) + ink_bgrs = [np.asarray(bgr, dtype=np.int64) for bgr, _ in ink_set] + min_area = max(16, int(MIX_MIN_AREA_FRAC * module_px * module_px)) + # Close vertically enough to bridge the M/2 stripe gap but not a full-cell gap. + kh = int(round(0.6 * module_px)) | 1 + kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (1, kh)) + + H, W = surface.shape[:2] + n_classified = 0 + n_unclassified = 0 + n_excluded_edge = 0 + for x in ink_bgrs: + present = _ink_present_mask(surface_i64, x, ink_bgrs) + grouped = cv2.morphologyEx(present, cv2.MORPH_CLOSE, kernel) + n_labels, labels, stats, _ = cv2.connectedComponentsWithStats(grouped, connectivity=8) + present_bool = present.astype(bool) + for lab in range(1, n_labels): + if int(stats[lab, cv2.CC_STAT_AREA]) < min_area: + continue + top = int(stats[lab, cv2.CC_STAT_TOP]) + left = int(stats[lab, cv2.CC_STAT_LEFT]) + bh = int(stats[lab, cv2.CC_STAT_HEIGHT]) + bw = int(stats[lab, cv2.CC_STAT_WIDTH]) + # EDGE EXCLUSION (SI-026): a component touching the image border may be + # a primitive clipped by the fragment boundary -- its shape statistics + # are corrupted, so counting it biases the mix (standard practice in + # counting problems: count interior events only). n therefore counts + # INTERIOR classified instances; the exclusion count is reported. The + # cost -- whole edge-cell primitives on a full surface are excluded + # too, shrinking n -- is priced in consistently: the committed expected + # vector is derived from interior-only measurements, and the scorer's + # n-dependent tolerance (score.py) absorbs the sampling noise of a + # smaller n instead of punishing it as disagreement. + if top == 0 or left == 0 or top + bh >= H or left + bw >= W: + n_excluded_edge += 1 + continue + # Classify on the ORIGINAL mask (restricted to this grouped component) + # so the stripe rhythm the close would fill in is preserved. + comp = (present_bool[top:top + bh, left:left + bw] + & (labels[top:top + bh, left:left + bw] == lab)) + if int(comp.sum()) < min_area: + continue + t = _classify_primitive(comp, module_px) + if t is None: + n_unclassified += 1 + else: + counts[t] += 1 + n_classified += 1 + + if n_classified > 0: + mix = [counts[t] / float(n_classified) for t in PRIMITIVE_MIX_ORDER] + else: + mix = [0.0] * len(PRIMITIVE_MIX_ORDER) + total_components = n_classified + n_unclassified + detail = { + "order": list(PRIMITIVE_MIX_ORDER), + "counts": counts, + "n_classified": n_classified, + "n_unclassified": n_unclassified, + "n_excluded_edge": n_excluded_edge, + "unclassified_share": (round(n_unclassified / total_components, 4) + if total_components else 0.0), + "n_ink_layers": len(ink_bgrs), + "note": "share of INTERIOR classified primitive INSTANCES by type; " + "border-touching components excluded (clipped shapes bias the " + "mix), merged/off-alphabet components counted unclassified " + "(SI-026). This describes the reconstruction's composition " + "model, not Studio.Build's original (SI-018).", + } + return mix, n_classified, detail # --- top-level grid-surface measurement ------------------------------------- @@ -581,11 +851,17 @@ def measure_grid_surface(image: np.ndarray) -> dict: working["stripe"] = stripe_detail stripe = Measurement("stripe_duty", duty, n_stripe, stripe_detail) - # 5. Primitive mix (measured but uncommitted -- SI-008). - mix, n_comp, mix_detail = measure_primitive_mix(image, module_px) - working["primitive_mix"] = {"mix": mix, "n_components": n_comp, **mix_detail} - primitive_mix = Measurement("primitive_frequency_mix", mix, n_comp, - {"note": "measured but uncommitted (SI-008)", **mix_detail}) + # 5. Primitive mix (committed identification carrier -- SI-026, closing the + # same-ink half of SI-022). Classified from the extracted ink set via + # per-ink footprint reconstruction; None value when nothing classifiable. + mix, n_comp, mix_detail = measure_primitive_mix(image, module_px, ink_set) + working["primitive_mix"] = {"mix": mix, "n_classified": n_comp, **mix_detail} + primitive_mix = Measurement( + "primitive_frequency_mix", + mix if n_comp > 0 else None, + n_comp, + mix_detail, + ) # 6. Staircase angle: not separately measured in the v0 grid family (weight-0 # verification; robustly recovering 45deg from a cell staircase inside a diff --git a/recogniser/score.py b/recogniser/score.py index 7e5db46..bc39140 100644 --- a/recogniser/score.py +++ b/recogniser/score.py @@ -32,6 +32,21 @@ "delta_e" tolerance). Per-ink agreement is the same clipped-linear curve on delta-E / delta_e; the feature agreement is the mean over observed inks. + 3b. Primitive-frequency mix (measure ``primitive_frequency_mix``, SI-026): the + measured and expected values are 5-bin instance-share vectors; agreement is + the clipped-linear curve on their L1 distance over an n-DEPENDENT effective + tolerance, + + tolerance_eff = tolerance * max(1, sqrt(N_REF_PRIMITIVES / n)) + + because a share vector estimated from n interior instances of a 5-bin + multinomial has L1 sampling deviation ~ sqrt(1/n): a small fragment's larger + L1 is sampling noise, not disagreement (the n/(n+k) saturation below already + handles "we saw little, so confidence is low"). At n >= N_REF_PRIMITIVES the + committed tolerance applies unchanged, so full-measurement impostor + rejection is never loosened. The claim's working carries n, tolerance_eff + and the L1. + 4. Sample-size scaling (SI-002): a feature's confidence discounts its agreement by how many samples of its declared ``sample_unit`` were seen, @@ -127,6 +142,20 @@ # arithmetic (audit s5 "c1*c2 != c3") and flags the claim's verification. OVERPRINT_RESIDUAL_TOL = 12.0 +# Reference sample size for the primitive-frequency-mix tolerance (SI-026): the +# typical INTERIOR classified-instance count of a full surface in the derivation +# corpus that produced the committed expected vector (grammars/iso-002.yaml, +# seeds x densities 0.35/0.45/0.55 x modules 40/48/64 on 14x10 grids, border +# components excluded). The mix feature's effective tolerance is +# +# tolerance_eff = committed_tolerance * max(1, sqrt(N_REF_PRIMITIVES / n)) +# +# -- a share vector from n instances of a 5-bin multinomial has L1 sampling +# deviation ~ sqrt(1/n), so a small-n fragment's larger L1 is sampling noise, not +# disagreement; at n >= N_REF the committed tolerance applies unchanged, so +# full-surface impostor rejection is never loosened by this rule. +N_REF_PRIMITIVES = 16 + # Sample-size saturation constants per declared sample_unit (SI-002). Small # integers: two boundaries already give real cascade evidence, a couple of inks # complete the pair, but periods are cheap so they saturate slower. @@ -139,7 +168,12 @@ "striped_regions": 1.0, "diagonals": 1.0, "overlaps": 1.0, - "primitives_observed": 1.0, + # 2.0 (not 1.0): the mix is a 5-BIN SHARE VECTOR, not a scalar -- a handful of + # instances pins it far less than a handful of periods pins a period, and the + # n-dependent tolerance (SI-026) simultaneously widens at small n, so the + # saturation must discount harder to keep a noisy-but-lucky small-n mix from + # pushing a same-ink impostor fragment over the identified line. + "primitives_observed": 2.0, } DEFAULT_SATURATION_K = 1.0 @@ -630,6 +664,52 @@ def score_sheet(sheet: dict, measured: dict) -> dict: n = m.n entry["measured"] = ink["measured"] entry["detail"] = ink["detail"] + elif measure_name == "primitive_frequency_mix": + # Primitive-mix identification carrier (SI-026, closing the same-ink + # half of SI-022). Agreement is clipped-linear on the L1 distance + # between the measured and expected instance-share vectors, over the + # declared scalar tolerance (SI-001 convention). n = classified + # instances scales confidence (SI-002): a fragment with a handful of + # primitives yields a noisy mix, so its small n discounts it rather + # than letting the estimate dominate. n = 0 -> unobserved (skipped and + # renormalised away, never scored 0). + m = measurements.get("primitive_frequency_mix") + expected_vec = feature.get("expected") + tol = feature.get("tolerance") + if m is None or m.n <= 0 or m.value is None: + entry["note"] = "no classifiable primitives observed in this fragment" + per_feature.append(entry) + continue + if not isinstance(expected_vec, (list, tuple)) or not isinstance(tol, (int, float)): + entry["note"] = ("primitive_frequency_mix expects a vector expected " + "and scalar tolerance; skipped") + per_feature.append(entry) + continue + measured_vec = list(m.value) + l1 = float(np.abs(np.asarray(measured_vec, dtype=float) + - np.asarray(expected_vec, dtype=float)).sum()) + n = m.n + # n-DEPENDENT EFFECTIVE TOLERANCE (SI-026). The mix is a share vector + # estimated from n interior instances of a 5-bin multinomial: an + # HONEST measurement's expected L1 sampling deviation scales + # ~ sqrt(1/n), so a small-n fragment's larger L1 is sampling noise, + # not disagreement, and must not be punished as such (the existing + # n/(n+k) saturation already handles "we saw little -> low + # confidence"). The committed tolerance is calibrated at the typical + # full-surface interior instance count N_REF_PRIMITIVES (from the + # derivation corpus); below it the tolerance widens by sqrt(n_ref/n), + # at or above it tolerance_eff == the committed value, so full-surface + # impostor rejection is never loosened. + tol_eff = float(tol) * max(1.0, np.sqrt(N_REF_PRIMITIVES / n)) + agree = agreement_linear(l1, 0.0, tol_eff) + entry["measured"] = [round(float(v), 4) for v in measured_vec] + entry["detail"] = {"l1_distance": round(l1, 4), + "tolerance_committed": float(tol), + "tolerance_eff": round(tol_eff, 4), + "n_ref": N_REF_PRIMITIVES, + "n": int(n), + "order": getattr(m, "detail", {}).get("order"), + **getattr(m, "detail", {})} elif measure_name == "overprint_multiply_consistency": # Verification (audit s5): agreement falls with the worst product # residual; a large residual (broken multiply arithmetic) flags the diff --git a/tests/test_primitive_mix.py b/tests/test_primitive_mix.py new file mode 100644 index 0000000..411cbb1 --- /dev/null +++ b/tests/test_primitive_mix.py @@ -0,0 +1,383 @@ +"""Tests for the committed primitive-frequency mix (Phase 9; SI-026, closing the +dangerous same-ink half of SI-022). + +At grammar_version 1.1.0 iso-002 carries identification on TWO measured features: +the ink set (weight 0.60) and the primitive-frequency mix (weight 0.40, was +reserved-unmeasured, SI-008). This file validates, by measurement against +generator ground truth (README rule): + + 1. CLASSIFIER ACCURACY. Per-type recall on isolated primitives, and the key + safety property -- ZERO cross-type confusion (a miss becomes unclassified, + never a wrong label). Plus an end-to-end share-error floor on full + compositions (includes the honest merged-blob + interior-only bias). + 2. GENUINE surfaces (held-out seeds NOT used to derive the expected vector) + still identify. + 3. SAME-INK IMPOSTORS. all-circles and all-stripes compositions rendered with + 002's exact inks do NOT reach identified at any fragment fraction >= 0.2, + with a stated margin (SI-022 dangerous half) -- and the NEAREST impostor + family (all-staircase / all-filled, whose depth-2-cap remnants read as + filled cells, SI-019) stays below identified at every fraction too. + 4. SMALL-FRAGMENT HONESTY. A fragment with only a few classified primitives + yields a noisy mix estimate; the n-dependent tolerance (sampling noise is + not disagreement) plus n-scaling (SI-002, k=2) must let the ink evidence + carry such a fragment rather than mix noise dominating either way. + 5. The sheet validates at v1.1.0 with the mix committed. +""" + +import numpy as np +import pytest +from pathlib import Path + +from sheets import load_sheet +from generator import grid, fragments +from generator.grid import (filled_cell_mask, circle_mask, stripe_block_mask, + staircase_mask, stadium_mask, quarter_round_mask, + _hex_to_bgr) +from recogniser import measure_grid as mg +from recogniser import score +from recogniser.claim import recognise + +REPO_ROOT = Path(__file__).resolve().parent.parent +GRAMMARS = REPO_ROOT / "grammars" +ISO_SHEET = GRAMMARS / "iso-002.yaml" +TYPES = mg.PRIMITIVE_MIX_ORDER + + +@pytest.fixture(scope="module") +def sheet(): + return load_sheet(ISO_SHEET) + + +def _result(claim, sheet_id): + for r in claim["results"]: + if r["sheet_id"] == sheet_id: + return r + raise KeyError(sheet_id) + + +def _feature(result, fid): + for f in result["per_feature"]: + if f["id"] == fid: + return f + raise KeyError(fid) + + +def _isolated_mask(type_, rng, rows, cols, M): + """Build one isolated primitive mask of a known type (mirrors the generator's + seeded placement, but for a single instance whose true type we know).""" + if type_ == "filled_cell": + r = int(rng.integers(0, rows)); c = int(rng.integers(0, cols)) + return filled_cell_mask(rows, cols, M, r, c) + if type_ == "inscribed_circle": + scale = 2 if rng.random() < 0.35 else 1 + r = int(rng.integers(0, rows - scale + 1)); c = int(rng.integers(0, cols - scale + 1)) + return circle_mask(rows, cols, M, r, c, scale) + if type_ == "stripe_bar": + w = int(rng.integers(1, 4)); h = int(rng.integers(1, 4)) + r = int(rng.integers(0, rows - h + 1)); c = int(rng.integers(0, cols - w + 1)) + return stripe_block_mask(rows, cols, M, r, c, w, h) + if type_ == "staircase_diagonal": + n = int(rng.integers(2, 6)); direction = 1 if rng.random() < 0.5 else -1 + c = int(rng.integers(0, max(1, cols - n + 1))) + if direction == 1: + r = int(rng.integers(0, max(1, rows - n + 1))) + else: + r = int(rng.integers(n - 1, rows)) if rows >= n else rows - 1 + return staircase_mask(rows, cols, M, r, c, n, direction) + # rounded_cap + vertical = rng.random() < 0.5 + axis = rows if vertical else cols + n = int(rng.integers(1, min(4, axis) + 1)) + if n == 1: + corner = int(rng.integers(0, 4)); r = int(rng.integers(0, rows)); c = int(rng.integers(0, cols)) + return quarter_round_mask(rows, cols, M, r, c, corner) + if vertical: + c = int(rng.integers(0, cols)); r = int(rng.integers(0, rows - n + 1)) + return stadium_mask(rows, cols, M, r, c, n, vertical=True) + r = int(rng.integers(0, rows)); c = int(rng.integers(0, cols - n + 1)) + return stadium_mask(rows, cols, M, r, c, n, vertical=False) + + +def _bbox_comp(mask): + ys, xs = np.nonzero(mask) + return mask[ys.min():ys.max() + 1, xs.min():xs.max() + 1] + + +# ========================================================================= +# 1. Classifier accuracy floor + zero cross-type confusion (SI-026) +# ========================================================================= + + +def test_classifier_recall_and_no_cross_confusion(): + """On isolated primitives the classifier's INTRINSIC accuracy: per-type recall + above a stated (data-derived) floor, and -- the load-bearing safety property + -- ZERO cross-type confusion. A miss falls to unclassified, never to a wrong + primitive label, so the classifier cannot manufacture a plausible-but-false + mix (SI-026).""" + M = 48 + rows = cols = 8 + rng = np.random.default_rng(0) + N = 120 + confusion = {t: {u: 0 for u in list(TYPES) + [None]} for t in TYPES} + for _ in range(N): + for t in TYPES: + comp = _bbox_comp(_isolated_mask(t, rng, rows, cols, M)) + pred = mg._classify_primitive(comp, M) + confusion[t][pred] += 1 + + recall = {t: confusion[t][t] / N for t in TYPES} + # Off-diagonal (a true type predicted as a DIFFERENT type) must be exactly 0. + for t in TYPES: + for u in TYPES: + if u != t: + assert confusion[t][u] == 0, f"cross-confusion {t}->{u}: {confusion[t][u]}" + + # Per-type recall floors, below the measured values (all five types measure + # 1.0: single-bar stripe blocks are caught by the half-cell-bar rule, + # 4-cell stadiums by the widened extent band). + for t in TYPES: + assert recall[t] >= 0.95, f"{t} recall {recall[t]}" + assert np.mean(list(recall.values())) >= 0.98 + + +def test_end_to_end_share_error_floor(sheet): + """On full compositions the measured instance-share vector tracks the ground- + truth mix within a stated mean-absolute-share-error floor. This includes the + honest merged-blob bias (overlapping primitives -> unclassified), so the floor + is looser than the isolated-classifier accuracy -- and stated as such.""" + errs = [] + classified_fracs = [] + for M in (40, 48, 64): + for d in (0.35, 0.45, 0.55): + for s in range(6): + surface, gt = grid.render_with_truth( + sheet, cols=14, rows=10, module_px=M, seed=s, density=d) + meas = mg.measure_grid_surface(surface)["measurements"]["primitive_frequency_mix"] + if meas.n == 0: + continue + gtc = gt.primitive_counts + total = sum(gtc.values()) + true_vec = np.array([gtc[t] / total for t in TYPES]) + errs.append(np.abs(np.array(meas.value) - true_vec)) + classified_fracs.append(meas.n / total) + errs = np.array(errs) + # Measured ~0.105 (interior-only); floor 0.15 with margin. Every type < 0.20. + assert errs.mean() < 0.15 + assert (errs.mean(axis=0) < 0.20).all() + # A meaningful separable INTERIOR sample is recovered (per-ink separation + + # edge exclusion, SI-026; measured ~0.25 of placed instances). + assert np.mean(classified_fracs) > 0.18 + + +# ========================================================================= +# 2. Genuine held-out surfaces still identify (SI-026 requirement b) +# ========================================================================= + +# Seeds 100+ are NOT in the derivation corpus (seeds 0-11) the expected vector +# came from -- a genuine hold-out. +HELDOUT_SEEDS = [100, 101, 102, 103, 104, 105, 106, 107] + + +def test_genuine_heldout_surfaces_identify(sheet): + """Held-out genuine full surfaces (seeds never used to derive the expected + vector) still reach identified -- both id-features observed, coverage 1.0.""" + verdicts = [] + for s in HELDOUT_SEEDS: + for d in (0.35, 0.45, 0.55): + surface = grid.render(sheet, cols=16, rows=12, module_px=48, seed=s, density=d) + r = _result(recognise(surface, str(GRAMMARS)), "iso-002") + assert r["aggregate_confidence"] >= score.CANDIDATE_THRESHOLD + verdicts.append(r["verdict"]) + identified_rate = np.mean([v == "identified" for v in verdicts]) + # Full genuine surfaces identify at >= 0.80 (director acceptance; measured + # 0.92 on this corpus). + assert identified_rate >= 0.80 + + +def test_genuine_small_fragments_stay_solid_candidates(sheet): + """Genuine frac-0.2 fragments: the mix is barely measurable there (few + interior instances), so identification honestly weakens -- but the mean + aggregate must stay a solid candidate (>= 0.45; measured ~0.64). The + n-dependent tolerance keeps mix sampling noise from being punished as + disagreement, and the ink evidence carries the fragment.""" + aggs = [] + for s in (100, 102, 104): + for d in (0.35, 0.45, 0.55): + surface = grid.render(sheet, cols=16, rows=12, module_px=48, + seed=s, density=d) + rng = np.random.default_rng(s * 7 + int(d * 100)) + for _ in range(3): + frag = fragments.sample_fragment(surface, frac=0.2, rng=rng)[0] + r = _result(recognise(frag, str(GRAMMARS)), "iso-002") + aggs.append(r["aggregate_confidence"]) + assert float(np.mean(aggs)) >= 0.45 + + +# ========================================================================= +# 3. Same-ink impostors do NOT identify at any fragment >= 0.2 (SI-022) +# ========================================================================= + +IDENT = score.IDENTIFIED_THRESHOLD # 0.70 +IMPOSTOR_MARGIN = 0.04 # stated margin below the identified line + + +@pytest.mark.parametrize("impostor", ["inscribed_circle", "stripe_bar"]) +@pytest.mark.parametrize("frac", [0.2, 0.5, 1.0]) +def test_same_ink_impostor_never_identifies(sheet, impostor, frac): + """A same-ink / different-composition impostor -- an all-circles or all-stripes + field rendered with 002's EXACT master inks -- must not reach identified at any + fragment fraction >= 0.2, with a stated margin. Its inks match (agreement ~1) + but its mix is a one-hot vector far from the committed expected mix, so mix + agreement is ~0 and the ink weight alone (0.60) cannot cross 0.70 (SI-022 / + SI-026).""" + for d in (0.35, 0.45, 0.55): + surface = grid.render(sheet, cols=16, rows=12, module_px=48, seed=0, + density=d, types=[impostor]) + rng = np.random.default_rng(int(d * 100) + 7) + for _ in range(4): + frag = surface if frac >= 1.0 else fragments.sample_fragment( + surface, frac=frac, rng=rng)[0] + r = _result(recognise(frag, str(GRAMMARS)), "iso-002") + assert r["verdict"] != "identified" + assert r["aggregate_confidence"] <= IDENT - IMPOSTOR_MARGIN, ( + f"{impostor} frac {frac} d {d}: agg {r['aggregate_confidence']}") + if frac >= 1.0: + break + + +@pytest.mark.parametrize("impostor", ["staircase_diagonal", "filled_cell"]) +@pytest.mark.parametrize("frac", [0.35, 0.5, 1.0]) +def test_nearest_impostor_family_stays_below_identified(sheet, impostor, frac): + """The NEAREST same-ink impostor family: all-staircase and all-filled, whose + depth-2-cap remnants (SI-019) read as filled cells, concentrating measured + mass on the filled/staircase bins. Their measured L1 floor (~1.0) sits above + the committed tolerance because the classifier's recall fixes (half-cell-bar + stripe rule, stadium band, disc-IoU gate) spread the expected vector across + all five bins -- so even this family must stay below identified at every + fraction. Margin here is thinner than the circles/stripes one and asserted + as strictly-below-threshold (empirical worst 0.684).""" + for s in range(3): + for d in (0.35, 0.45, 0.55): + surface = grid.render(sheet, cols=16, rows=12, module_px=48, seed=s, + density=d, types=[impostor]) + rng = np.random.default_rng(s * 13 + int(d * 100) + 3) + reps = 1 if frac >= 1.0 else 3 + for _ in range(reps): + frag = surface if frac >= 1.0 else fragments.sample_fragment( + surface, frac=frac, rng=rng)[0] + r = _result(recognise(frag, str(GRAMMARS)), "iso-002") + assert r["verdict"] != "identified" + assert r["aggregate_confidence"] < IDENT, ( + f"all-{impostor} frac {frac} s {s} d {d}: " + f"agg {r['aggregate_confidence']}") + + +def test_impostor_mix_disagrees_genuine_agrees(sheet): + """The discriminator, stated directly: on a full surface the all-circles + impostor's mix agreement is ~0 while a genuine surface's is well above it.""" + imp = grid.render(sheet, cols=16, rows=12, module_px=48, seed=0, density=0.45, + types=["inscribed_circle"]) + gen = grid.render(sheet, cols=16, rows=12, module_px=48, seed=101, density=0.45) + r_imp = _result(recognise(imp, str(GRAMMARS)), "iso-002") + r_gen = _result(recognise(gen, str(GRAMMARS)), "iso-002") + imp_mix = _feature(r_imp, "primitive_frequency_mix") + gen_mix = _feature(r_gen, "primitive_frequency_mix") + # Both see the exact inks, so ink agreement is high for both... + assert _feature(r_imp, "ink_set")["agreement"] >= 0.9 + # ...but the mix separates them. + assert imp_mix["agreement"] == 0.0 + assert gen_mix["agreement"] >= 0.3 + assert r_gen["aggregate_confidence"] > r_imp["aggregate_confidence"] + 0.1 + + +# ========================================================================= +# 4. Small-fragment honesty: n-scaling stops a noisy mix dominating (SI-002) +# ========================================================================= + + +def test_small_mix_saturation_discounts(): + """A 3-primitive mix estimate is noisy; the sample-size saturation (SI-002, + k=2 for primitives_observed) must discount it so it cannot dominate the + score. At n=3 the mix confidence is agreement x 3/(3+2) = 0.6 x agreement -- + strictly below the agreement, and below a large-n mix's near-1 saturation.""" + k = score.SATURATION_K["primitives_observed"] + assert k == 2.0 + assert score.saturation(3, "primitives_observed") == pytest.approx(3 / (3 + k)) + assert score.saturation(3, "primitives_observed") < 0.8 + # Monotone: more primitives -> more confidence in the mix estimate. + assert (score.saturation(3, "primitives_observed") + < score.saturation(25, "primitives_observed")) + + +def test_n_dependent_tolerance_widens_then_pins(): + """The mix tolerance is n-dependent (SI-026): tolerance_eff = committed x + max(1, sqrt(n_ref/n)). Small n widens (sampling noise is not disagreement); + n >= n_ref pins to the committed value so full-measurement impostor rejection + is never loosened.""" + n_ref = score.N_REF_PRIMITIVES + tol = 0.95 + for n in (1, 3, 6): + eff = tol * max(1.0, np.sqrt(n_ref / n)) + assert eff > tol + for n in (n_ref, n_ref + 10): + assert tol * max(1.0, np.sqrt(n_ref / n)) == pytest.approx(tol) + + +def test_small_fragment_mix_does_not_dominate(sheet): + """On a small genuine fragment carrying only a few classified primitives, the + ink evidence -- not the noisy mix -- must carry the score. We assert the mix's + weighted confidence contribution is below the ink's, so a noisy few-primitive + mix estimate cannot swing the verdict on its own (SI-002 / SI-026), and that + the claim's working carries n, tolerance_eff and the L1 (director requirement) + with the widened tolerance applied at small n.""" + surface = grid.render(sheet, cols=16, rows=12, module_px=48, seed=102, density=0.35) + rng = np.random.default_rng(11) + found = False + for _ in range(40): + frag = fragments.sample_fragment(surface, frac=0.12, rng=rng)[0] + r = _result(recognise(frag, str(GRAMMARS)), "iso-002") + mix = _feature(r, "primitive_frequency_mix") + ink = _feature(r, "ink_set") + if not (mix["observed"] and ink["observed"] and 1 <= mix["n"] <= 6): + continue + found = True + # The working shows the n-dependent tolerance handling (SI-026). + det = mix["detail"] + assert det["n"] == mix["n"] + assert "l1_distance" in det + assert det["tolerance_eff"] > det["tolerance_committed"] # widened (n small) + assert det["n_ref"] == score.N_REF_PRIMITIVES + # Small n -> saturation strictly discounts the mix confidence below its + # agreement (confidence = agreement x saturation, SI-002). + assert mix["saturation"] < 0.8 + if mix["agreement"] > 0: + assert mix["confidence"] < mix["agreement"] + assert mix["confidence"] == pytest.approx( + mix["agreement"] * mix["saturation"], abs=1e-3) + # Ink (weight 0.60, saturated) contributes more than the small-n mix + # (weight 0.40, discounted): the fragment leans on the ink, not mix noise. + ink_contrib = 0.60 * (ink["confidence"] or 0.0) + mix_contrib = 0.40 * (mix["confidence"] or 0.0) + assert ink_contrib > mix_contrib + assert found, "no small-mix fragment sampled; adjust frac/seed" + + +# ========================================================================= +# 5. Sheet validates at v1.1.0 with the mix committed +# ========================================================================= + + +def test_sheet_committed_mix_v110(sheet): + assert sheet["sheet"]["grammar_version"] == "1.1.0" + feats = {f["id"]: f for f in sheet["signature_locus"]["features"]} + mix = feats["primitive_frequency_mix"] + assert mix["status"] == "measured" + assert mix["role"] == "identification" + assert mix["weight"] == pytest.approx(0.40) + assert feats["ink_set"]["weight"] == pytest.approx(0.60) + assert isinstance(mix["expected"], list) and len(mix["expected"]) == 5 + assert isinstance(mix["tolerance"], (int, float)) + # Identification weights still sum to 1.0 (the load-bearing invariant). + idw = sum(f["weight"] for f in sheet["signature_locus"]["features"] + if f["role"] == "identification") + assert idw == pytest.approx(1.0) diff --git a/tests/test_recogniser.py b/tests/test_recogniser.py index 7c7919e..046a080 100644 --- a/tests/test_recogniser.py +++ b/tests/test_recogniser.py @@ -276,33 +276,34 @@ def test_iso002_does_not_crash_and_scores_below_001(surface0): def test_iso002_grid_family_measures_gracefully_on_band_image(surface0): - """Phase 6 changed this: the grid measurer family now EXISTS, so a 001 band - image is also measured by it (claim.STRUCTURE_MEASURERS) and iso-002 is scored - against real grid measurements -- honestly, and still well below 001. - - The pre-Phase-6 version of this test asserted 'no measurer for grid_module'; - that documented the absence this phase fills. The invariant that MUST hold is - honest cross-discrimination: iso-002's nine-ink address is not met by a band - surface's two greens, and primitive_frequency_mix stays reserved (SI-008).""" + """Phase 6 gave the grid family a measurer, so a 001 band image is measured by + it and iso-002 is scored against real grid measurements -- honestly, and still + well below 001. Phase 9 (SI-026) committed primitive_frequency_mix as measured, + so it now runs here too; the invariant that MUST hold is unchanged honest + cross-discrimination: neither the nine-ink address nor the ISO primitive mix is + met by a band surface, so both id-features agree ~0 and iso-002 does not + identify.""" r002 = _result(recognise(surface0, str(GRAMMARS)), "iso-002") # The grid family ran end to end: the grid module (weight-0 normalisation) and # the ink-set extractor both produced real measurements -- never a crash. assert _feature(r002, "grid_module")["observed"] is True assert _feature(r002, "ink_set")["observed"] is True - # ...but two greens are not the nine ISO inks: colour agreement stays low and - # the relationship path is disabled (< 3 inks, SI-020), so no leniency leaks. + # ...but a band surface's greens are not the nine ISO inks: colour agreement + # stays low and the relationship path is disabled (< 3 inks, SI-020). assert _feature(r002, "ink_set")["agreement"] < 0.34 - # primitive_frequency_mix is reserved + skipped regardless (SI-008). + # primitive_frequency_mix is now measured (SI-026), but a band composition is + # nothing like the ISO primitive mix, so its agreement is ~0 -- it cannot + # rescue the cross-family score. prim = _feature(r002, "primitive_frequency_mix") - assert prim["observed"] is False - assert "SI-008" in prim.get("note", "") + if prim["observed"]: + assert prim["agreement"] == 0.0 - # No two-ink overprints on a two-green band surface: overprint unobserved, - # which is NOT a verification failure. + # No two-ink overprints on a band surface: overprint unobserved (NOT a + # verification failure), and iso-002 does not identify a band image. assert _feature(r002, "overprint_consistency")["observed"] is False - assert "iso-002" not in [r002["sheet_id"]] or r002["verdict"] != "identified" + assert r002["verdict"] != "identified" # ========================================================================= diff --git a/tests/test_recogniser_grid.py b/tests/test_recogniser_grid.py index 709f9a5..113fe5a 100644 --- a/tests/test_recogniser_grid.py +++ b/tests/test_recogniser_grid.py @@ -296,19 +296,19 @@ def test_relationship_path_disabled_below_min_inks(sheet): def test_grid_render_top_is_iso002(render): - """A clean 002 render is the top verdict against iso-002 (candidate-or-better), - with coverage reflecting the unmeasured primitive mix (SI-008): only the ink - set carries measurable identification weight, so coverage is 1.0 and the - reserved mix weight is renormalised away, not counted.""" + """A clean 002 render is the top verdict against iso-002 (candidate-or-better). + At grammar_version 1.1.0 the primitive mix is a MEASURED identification carrier + (SI-026), so on a full render both id-features (ink set + mix) are observed and + coverage is 1.0.""" surface, _ = render claim = recognise(surface, str(GRAMMARS)) top = claim["results"][0] assert top["sheet_id"] == "iso-002" assert top["aggregate_confidence"] >= score.CANDIDATE_THRESHOLD assert top["verdict"] in ("candidate", "identified") - # SI-008: the primitive mix is reserved + skipped, so it is not in coverage. + # SI-026: the primitive mix is now measured and observed on a full render. prim = _feature(top, "primitive_frequency_mix") - assert prim["observed"] is False and "SI-008" in prim.get("note", "") + assert prim["observed"] is True and prim["n"] > 0 assert top["coverage"] == pytest.approx(1.0, abs=1e-6)