Equivalence harness: explain HF-26 and HF-24 by reconstruction instead of hand-measured waivers - #819
Merged
Merged
Conversation
…six hand-measured waivers
HF-26 was signed off by six `max_abs_pct` waiver rows scoped to
`interconnect: western, prong: 2`. That does not scale and was already
incomplete: the western smoke run's `p11 | solar` row (189.4 -> 239.9 MW) was
never measured, so it was never waived, and on the USA leg the same difference
reaches +28,620 % on rows where master is ~0 MW, where a percentage bound is
not a usable statement at all.
Replace them with two `reconstruction:` waivers, one per metric, whose bound is
COMPUTED per row from the artifacts:
fleet_mw = MW of existing c-plants whose nearest master nodal bus is in z
profiled_mw = the part at substations master's profile_{c}.nc covers
dropped_mw = fleet_mw - profiled_mw <- the prediction
gate G1: profiled_mw ~= the table's master column
gate G2: fleet_mw ~= the table's develop column
residual = (develop - master) - dropped_mw
Both gates predict a side SEPARATELY, so "explained" is earned rather than
restated: a wrong plant population fails G2, a wrong coverage set fails G1, and
a zone relocation (the HF-27 shape) fails both with equal and opposite
residuals. A reconstruction that cannot run sets `error`, every lookup returns
None, and the rows stay UNEXPLAINED -- failing open would be worse than an
unexplained row.
New `tests/equivalence/reconstructions.py`: a pure, file-free core
(reconstruct_drop / gate_tolerance / build_rows) behind loaders that open only
master's `elec_base_network.nc`, `bus2sub.csv`, `powerplants.csv`, the region
geojsons and the `bus` COORDINATE of the profile files -- never the 6.5 GB
`elec_base_network_l_pp.pkl` or the 622 MB USA profile payload. The nearest-bus
match is a BallTree over raw degrees, master's own metric, degree distortion
included. `filter_fleet` mirrors `add_electricity.load_powerplants` rather than
importing master-benchmark's copy (a sys.path collision), with a parity test
against develop's own `load_powerplants` as the drift alarm.
tables.py (additive): `reconstruction` joins WAIVER_BOUND_KEYS, checked before
the numeric bounds; `reconstructions=None` means the waiver does NOT hold;
`tables/hf26_reconstruction.csv`, `reconstructions.json` and a
**Reconstructions** section in comparison.md. plots.py: two figures with CSV
twins -- per-zone bars with the dropped MW hatched onto master, and a zone
choropleth with a residual panel floored at the gate tolerance so an exact
reconstruction does not render as a red map.
Replaces the six western waivers (capacity_existing_by_carrier onwind/solar,
p_nom_existing_by_zone_carrier p10|p9 x onwind|solar). They are deleted, not
kept: keeping both would let a row be explained by the weaker percentage bound
exactly when the reconstruction disagrees with it.
Verified on the western 2127 artifacts: every residual 0.0 MW (max 4.5e-13),
both gates hold on all 10 rows, and the run goes from 1 UNEXPLAINED row to 0
(72 equivalent / 9 explained). 352 fast tests pass.
Also clears two PRE-EXISTING pre-commit failures in the two files this commit
already touches, so the hooks pass on them: `add-trailing-comma` + `ruff format`
on two `paired_bar(...)` calls in plots.py, and RUF059 (`develop` unpacked and
never used) in test_plots.py. Both fail identically on 3e490b2; neither is
caused by this change.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… s no-op Three verifier findings on 56895c8, plus a latent bug the third one exposed. 1. The residual panel of hf26_dropped_mw_map.png could not be read as a failure indicator, for two independent reasons. The residual was SUMMED over carriers, so a +X onwind / -X solar zone cancelled to exactly zero and drew as perfect. And the colour scale was floored at the largest GLOBAL gate tolerance, while the gate is per row (0.5 % of master, 1 MW floor) -- so on the USA leg `p129 | solar` (+2.0 MW against its own 1.0 MW gate) and `p123 | solar` (+9.4 against 4.9) both rendered at under 10 % of a +/-105 MW ramp: failures drawn as "inside tolerance". The residual now gets its own panel PER CARRIER, coloured by `residual_mw / gate_tol_mw` on a diverging scale fixed to +/-3, so +/-1 is the band edge and a gate failure is visibly outside it; failing zones are hatched and black-edged, because colour alone cannot carry a pass/fail distinction on a continuous ramp. The colourbar names the unit. The CSV twin is now long-format: zone, carrier, dropped_mw, master_mw, develop_mw, residual_mw, gate_tol_mw, residual_over_gate, ok. Normalising by the row's own gate keeps the property the global floor was reaching for -- an exact reconstruction is white -- but earns it, rather than widening the scale until the failures disappear too. 2. `gate_tolerance` was documented as "the row's own tolerance" and was not: it scaled by `max(|master|, |develop|)` while `tables._row_verdict` scales by `|master|` (the denominator of delta_pct). Since develop exceeds master throughout the HF-26 regime, the gate was systematically LOOSER than the verdict it guards -- USA `p126 | solar` 6.4 MW against a 1.1 MW row tolerance, western `p10 | onwind` 23.3 against 10.5. Now `|master|`, with the atol floor still carrying the master-is-zero rows. A new test asserts the PROPERTY rather than the formula: a delta just inside the gate is `equivalent` in the table and just outside it is not. This moves one verdict, as it should: USA `p126 | solar` carried a 3.0 MW G2 error that the old, too-wide gate waved through. It is now UNEXPLAINED. 3. `investment_year` opened master's assembled network with `pypsa.Network` purely to read `investment_periods`, ~100 s per USA run, when `plots.load_artifacts` has already loaded that same file as `art.assembled_master`. It now takes it off `art` and falls back to the file for the CLI and the unit tests. Doing so uncovered why nobody had noticed the cost: the expression was `list(getattr(n, "investment_periods", []) or [])`, and on a real Network that attribute is a pandas Index, so `Index or []` raises "The truth value of a Index is ambiguous". A bare `except` swallowed it and the config fallback answered every time. The file was opened and then never read. Extracted as `_first_investment_period`, which length-checks instead. Re-verified read-only, no rebuild: western 2127 -- 7 HF-26 rows explained, 0 UNEXPLAINED, 0 gate failures, max |residual| 0.0 MW, with gates now tighter (p10 | onwind 10.5 MW, was 23.3); whole-USA 2127 -- G1 exact on all 270 keys, 6 G2 failures (was 5; the sixth is p126 | solar from finding 2), sum |residual| 228.8 MW unchanged, 138 explained / 22 UNEXPLAINED over the two existing-capacity metrics, against 160 UNEXPLAINED before the reconstruction existed. The regenerated map hatches exactly those 6 zones. 357 fast tests pass. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Spec: section 6.1 of
experiments/runs/eq-usa-80abb3bc125f/diagnosis/p_nom_max_zone_rows.md.
Master's `build_renewable_profiles` filters on
`ds["profile"].mean("time") > min_p_max_pu`. A substation whose summed GODEEEP
land availability is zero gets a NaN CF out of
`weighted_bus_aggregation`, `NaN > 0` is False, and the whole bus leaves the
file -- TAKING ITS NREL CAPS `p_nom_max` WITH IT, which was never zero. Develop
sums the caps per s{simpl} cluster BEFORE that filter
(`remap_caps_to_cluster`), so the MW survives on a cluster whose CF is defined.
Same shape as HF-26 -- master drops, develop keeps, delta >= 0 -- and the same
two gates, so "explained" is still earned rather than asserted:
caps_mw = caps of every substation whose busmap cluster is common (-> develop, G2)
covered_mw = the part of that at substations master's profile kept (-> master, G1)
dropped_mw = caps_mw - covered_mw (-> the delta)
Three small inputs: the caps artifact (1.4/1.6 MB), busmap_s{simpl}, and the
`bus` COORDINATE of master's profile file. No network, no CF array.
`nrel_caps_path` mirrors `build_electricity.smk::nrel_exclusion_artifact` rather
than hard-coding a name, and RAISES when `renewable_land_access` is unset: that
is the atlite path, which has no caps file, so the mechanism does not exist
there and must not hold. Nine tests cover the arithmetic, the two gates, the id
formats, the busmap and common-cluster exclusions, the path naming and the
table integration.
Replaces the four `interconnect: western` DL-18/HF-24 table waivers
(`p_nom_max_by_zone_onwind` p8/p9/p10, `p_nom_max_by_zone_solar` p8, each
`max_abs_pct: 5`). The three HF-24 CELL waivers are untouched -- they are
compare.py findings, not table rows. The hand bounds did not scale: the USA leg
has 46 onwind + 4 solar rows over tolerance and `p128 | onwind` moves
126 -> 174 MW (+38 %), which no percentage written for western would cover.
Generalised alongside it, because there are two providers now:
- `_make_row` takes a `key` override, so HF-24's per-tech metric is keyed by the
bare zone while HF-26 keeps "p10 | onwind";
- `rows_frame` keeps every row that NAMES A ZONE rather than only HF-26's
metric, so one frame, one CSV and one pair of figures serve both;
- `tables/hf26_reconstruction.csv` -> `tables/reconstructions.csv`, which is
what it now holds;
- `Reconstruction.totals()` sums the ZONE rows. The markdown summary used
`national()`, which HF-24 has none of, so its line read "0 MW" beside 50
explained rows. The residual column also prints three decimals now: the whole
point of it is to distinguish 0.000 from 0.118 from 100;
- the two figures take a per-provider headline and `plots.export_all` draws a
pair for each.
And one figure defect the 134-zone metric exposed: the capped remainder was
drawn as a pooled "other (109 zones)" BAR, which at 7.9 million MW against
individual zones of 2e5 collapsed every bar the figure exists to show, and gave
that row a "gate tolerance" band that was the sum of 109 unrelated gates. The
remainder is now stated in the panel title instead -- reconciliation kept, axis
range given back.
Acceptance, read-only, no rebuild:
usa-p2-3h-20260915-2127 -- all 46 onwind + 4 solar rows explained, 0
UNEXPLAINED (was 50), 0 of 268 zone rows failing a gate, max |residual|
0.000 MW onwind and 0.118 MW solar, max G1 error 2.9e-11 MW. National
master 9,686,706 / develop 9,700,392 / dropped 13,686 MW onwind and
76,074,993 / 76,078,586 / 3,593 MW solar, matching section 4 of the spec.
`p128 | onwind` survives the master-scaled gate on the atol floor, as
predicted. The common cluster set was derived from the two profile files'
bus coordinates (292/292/292 onwind, 299/299/299 solar, 0 unmapped);
western 2127 -- 0 UNEXPLAINED, all 8 (zone, tech) rows passing both gates,
max |residual| 0.000 MW onwind / 0.033 MW solar, and its two over-tolerance
rows now explained by the provider instead of the deleted percentage bounds.
380 fast tests pass.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… quantity being explained
Two defects in the shared plotting code, both surfaced by drawing HF-24 through
functions written for HF-26.
1. The axis and legend were hard-coded to HF-26's quantity. HF-24's panels were
labelled `existing capacity [MW]` with a grey series called
`reconstructed fleet`, but HF-24 reconstructs installable POTENTIAL
(`p_nom_max`) and has nothing to do with the existing fleet -- a reader
validating from the PNG would take 9.7 million MW of onwind as installed
plant. The labels now belong to the provider: a `ReconLabels` record per
`Reconstruction.name` carrying the headline, the quantity axis label, the
fleet-series legend and the dropped-series legend, with a NEUTRAL default so
a future provider cannot silently inherit HF-26's wording.
Three more places said "fleet" about a quantity that may not be one:
- the `comparison.md` summary header (`national fleet` / `national master`
-> `predicts develop` / `predicts master` / `dropped`, and it now reports
what the reconstruction PREDICTS for each side rather than echoing the
table's own columns);
- the figure suptitle, same change;
- the gate note in `reconstructions._make_row`, which lands in the
comparison table's hot-fix column: "reconstructed fleet X MW vs develop Y"
is now "reconstruction predicts X MW for develop, table says Y".
The functions are renamed `hf26_dropped_*_figure` ->
`reconstruction_*_figure` for the same reason: they draw whichever provider
they are handed.
2. The left panel could not show what its title claims. "master + dropped
should meet develop" is unverifiable by eye when the drop is a small
fraction of the total: HF-24 moves 13,686 MW against 9,700,392 MW (0.14 %),
so master, develop and the reconstruction render as three identical bars and
the hatched segment is sub-pixel. The verification was really being done by
the residual panel alone.
There are now three panels per carrier: totals (reconciliation), dropped
(the quantity actually being explained, on its own axis), residual (the
check). HF-26's stacked view is unchanged and still reads well at 14-112 %;
HF-24's dropped panel goes from invisible to a legible 0-1,536 MW range.
Only the figure changes -- the CSV twin already carried every column.
Also, since the map panels of a row share one scale: one colourbar per row
instead of one per panel, built from a bare ScalarMappable (the geo and point
branches produce different artists, only the norm is common). That gives the
maps back a fifth of their width each. The map figure moves to
`layout="constrained"`, because a colourbar spanning a row of axes is not
something `tight_layout` can place -- it warned and overlapped.
No number moves; no re-acceptance run needed. 383 fast tests pass, including
new ones pinning the per-provider labels, the neutral default, the dedicated
dropped panel and the provider-neutral summary headers. All four PNGs
regenerated and checked: the USA HF-26 residual map still hatches exactly its 6
failing solar zones, the HF-24 maps still hatch none, and `fleet_mw` equals
`develop_mw` exactly on the HF-24 totals panel (verified against the CSV after
the bars looked unequal in a downsampled render -- they are not).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
G1 and G2 decompose develop - master into the caps at substations master's
profile does not carry, but they cannot fail for the wrong reason: master's
profile p_nom_max equals the caps file on every substation it keeps. A
substation dropped by min_p_nom_max, or by a low CF under a non-zero
min_p_max_pu, would pass with residual 0 and be signed off as HF-24.
Gate G3 reads avail_{tech}_{access}.nc and master's cell->bus mapping cache,
sums land availability per substation (as weighted_bus_aggregation does), and
splits dropped_mw into dropped_mw_zero_avail (HF-24's mechanism, including
substations with no intersecting cell) and dropped_mw_other. ReconRow gains
mechanism_mw / unattributed_mw; a row whose unattributed MW exceeds its
tolerance fails with the MW named. HF-26 rows carry NaN, so the gate never
fires there.
Also:
- NREL artifact names resolve from the layered config the Snakefile loads
(config.{slurm,common,plotting,api,sector,default}.yaml under the scenario),
not the harness config alone; caps and avail share one suffix rule.
- "could not run" errors say which reason (no metric frames vs no common
cluster set) instead of one shared, sometimes false, message.
- the shipped HF-24 cell-waiver test checks the bounds survive, not only that
three entries exist.
- figure tests: failing zones outlined on the zone panel; dedicated dropped
panel.
Written 2026-09-16/17 and left uncommitted on Sherlock; committed 2026-09-30.
Tests run after develop is merged in.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…t never needed Two fast tests failed on 3941e1d: the WIP had tightened them to assert WHICH "could not run" reason is reported, but the code read files before checking the in-memory preconditions, so a run with no profiles or no common cluster set reported a missing busmap or FileNotFoundError on master's cell->bus cache instead. - art.profiles is checked before busmap_s{simpl} is read; - master's cell->bus cache is read only once a tech has both a metric frame and a common cluster set. Every path still fails closed; only the message changes. Fast tier 519 passed / 10 skipped (the master-benchmark worktree checks); ruff 0.15.0 clean. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this does
Replaces the hand-measured HF-26 and HF-24 waivers in the equivalence harness with reconstructions: for each affected table row the harness now computes, from the run's own artifacts, what master and develop should show if the documented mechanism is the whole story, and only then marks the row explained.
Test-harness only: every change is under
tests/equivalence/. No workflow code, config or model result changes.Two providers (
tests/equivalence/reconstructions.py)hf26_existing_renewable_drop: master drops existing renewable plants at substations its profile file does not cover; develop keeps them. Predicts per zone and carrier:fleet_mw(develop),profiled_mw(master),dropped_mw(the delta). Replaces sixmax_abs_pctrows scoped to western, which did not scale to USA (rows where master is ~0 MW reach +28,620 %).hf24_nrel_caps_drop: master drops the NREL capsp_nom_maxof substations whose summed GODEEEP land availability is zero; develop sums caps per cluster before that filter. Replaces four westernmax_abs_pct: 5rows; the three HF-24 cell waivers (compare.py findings) are untouched and now have their bounds pinned by a test.Gates
Each row is explained only if:
(develop - master) - dropped_mwis within tolerance;avail_{tech}_{access}.nc(or no intersecting cell). G1/G2 alone cannot fail for the wrong reason on HF-24, because master's profilep_nom_maxequals the caps file wherever master keeps a substation. G3 catches a substation dropped bymin_p_nom_maxor a low CF being signed off as HF-24.A reconstruction that cannot run sets
errorand its rows stay UNEXPLAINED, so it fails closed. The error names the reason: missing profiles, missing metric frames, or missing common cluster set.Reporting (
tables.py,plots.py)comparison.mdgains a reconstruction summary.tables/hf26_reconstruction.csvandreconstructions.jsonare written per run.Status
56895c80) had an adversarial review; its findings are fixed in9137d2ab. I found no review record for the HF-24 provider and figure commits. The latest work (HF-24 mechanism gate G3, plus two ordering fixes for the "could not run" messages found by the test run today) is unit-tested only; it has not been run against a benchmark's artifacts.usa-p2-3h-20260915-2127), before G3 existed:Tests
pytest -m fast tests: 519 passed, 0 failed, 10 skipped (job 46103754)pytest -m fast tests/equivalence: 399 passed, 0 failed, 10 skippedtest_master_benchmark_branch.py, which needs a local.worktrees/master-benchmarkcheckout.ruff checkandruff format --checkat the CI pin (0.15.0) are clean.🤖 Generated with Claude Code