diff --git a/.claude/skills/docs-writing/SKILL.md b/.claude/skills/docs-writing/SKILL.md index 2b13c850..cd0e83b2 100644 --- a/.claude/skills/docs-writing/SKILL.md +++ b/.claude/skills/docs-writing/SKILL.md @@ -178,7 +178,7 @@ not ## 6. Vocabulary - **Gloss house vocabulary at first use** — _spec_, _program_, _declaration_, - _dimension_, _coordinate_, _frame_, _lookup_, _absence_, _macro_, _named + _dimension_, _coordinate_, _frame_, _relation_, _absence_, _macro_, _named expression_, _reported expression_, _escape_. One clause with a concrete instance: "one point of it, one generator in one snapshot, is a coordinate". - **Gloss every acronym and domain term at first use**, in parentheses, six diff --git a/README.md b/README.md index 8fc64850..4aea960d 100644 --- a/README.md +++ b/README.md @@ -21,7 +21,7 @@ with no data and no solver.** A math-spec file declares four things: the axes the model runs over, such as `snapshot` and `generator`; the data it expects, such as `load` and `cost`; the decisions the solver makes, such as `dispatch`; and the rules those decisions obey, such -as `sum(dispatch, over=generator) == load`. The file [below](#example) is a complete +as `sum(dispatch, consume=generator) == load`. The file [below](#example) is a complete model. math-spec reads that file, checks everything that can be checked without data, @@ -211,7 +211,7 @@ through are a dependency rather than one engine's internals. The keys themselves which are YAML math, a block per component, `dims:` and a `where:` string, come from [Calliope](https://github.com/calliope-project/calliope). [linopy](https://github.com/PyPSA/linopy) supplies the vocabulary that -`sum(over=)` and the dimension rules are named against. Issue numbers in these +`sum(consume=)` and the dimension rules are named against. Issue numbers in these pages point at lpspec, where the arguments happened. ## Status diff --git a/docs/about/limits.md b/docs/about/limits.md index 00b7b042..776cddf3 100644 --- a/docs/about/limits.md +++ b/docs/about/limits.md @@ -41,7 +41,7 @@ macro can write `over=d` and let the caller supply `d`. It could not do that if the dimension were the keyword itself. **An operator may read the whole table. It pays one full pass over the data.** -`sum(p, over=g)` reads one row per generator. `shift(p, over=t, offset=1)` reads +`sum(p, over=g)` reads one row per generator. `shift(p, along=t, offset=1)` reads one row, the one before it. `x * y * a` reads the rows of `a` that pair an `x` with a `y`. Each reads a bounded number of rows per output row, so an engine builds the model one chunk of rows at a time. @@ -55,14 +55,14 @@ the list of snapshots, not at the data. **An operator that calls itself is refused.** Nothing bounds how far it expands, so no number of passes over the data is enough. -| The operator | Allowed? | -| ---------------------------------------------------- | ----------------------------------------------- | -| filters rows on a column they already carry | yes | -| joins each row against a parameter or a lookup table | yes | -| reads a fixed number of neighbouring rows | yes | -| reads only the coordinate labels | yes | -| reads every row | yes, at one full pass before any chunk builds | -| calls itself | no, and the message names what to write instead | +| The operator | Allowed? | +| ------------------------------------------------ | ----------------------------------------------- | +| filters rows on a column they already carry | yes | +| joins each row against a parameter or a relation | yes | +| reads a fixed number of neighbouring rows | yes | +| reads only the coordinate labels | yes | +| reads every row | yes, at one full pass before any chunk builds | +| calls itself | no, and the message names what to write instead | **Degree is not a third test.** `p * q` at one coordinate is a join of a table with itself, so the objective and the constraints take it. Two things limit the @@ -138,7 +138,7 @@ sentence tells them apart: A cycle basis is the first kind. It needs the network's topology, which only the data has, so `cycle_incidence` arrives as a parameter. A minimum up time is the second kind. `min_up_time` is a column the model already binds, and the window -"the last `min_up_time` hours" follows from it, so `sum_back(within=min_up_time)` +"the last `min_up_time` hours" follows from it, so `sum_back(window=min_up_time)` reads the width off the column and you ship no window mask ([#849](https://github.com/fluxopt/lpspec/issues/849)). diff --git a/docs/examples/commitment.md b/docs/examples/commitment.md index 5906e2a4..23dc4d08 100644 --- a/docs/examples/commitment.md +++ b/docs/examples/commitment.md @@ -60,7 +60,7 @@ expressions: boundary: when: "committable and position(snapshot) == 0" expression: status_initial - otherwise: shift(status, over=snapshot, offset=1) + otherwise: shift(status, along=snapshot, offset=1) constraints: power_balance: @@ -80,7 +80,7 @@ constraints: `ramp_limit`, a unit starting up to `start_up_limit`. dims: [snapshot, generator] expression: >- - p - shift(p, over=snapshot, offset=1, edge=0) + p - shift(p, along=snapshot, offset=1, edge=0) <= ramp_limit * previous_status + start_up_limit * (1 - previous_status) objective: diff --git a/docs/examples/operators.md b/docs/examples/operators.md index 29b5abe5..45ad80c8 100644 --- a/docs/examples/operators.md +++ b/docs/examples/operators.md @@ -73,14 +73,14 @@ objective: { sense: minimize, expression: sum(p) } $`\sum_{g \in \mathcal{G}} p_{t,g} \le \mathrm{limit}_{t} \qquad \forall\, t \in \mathcal{T}`$ -### `sum(array, by=lookup)` +### `sum(array, by=relation)` `examples/operators/sum_by.yaml` ```yaml description: >- - The membership reduction — `sum(array, by=lookup)` lands the result on the - dimension the lookup maps into, which is what makes topology data rather than + The membership reduction — `sum(array, by=relation)` lands the result on the + column the relation is walked to, which is what makes topology data rather than structure. dimensions: @@ -88,8 +88,8 @@ dimensions: generator: { dtype: str } bus: { dtype: str } -lookups: - gen_bus: { over: generator, into: bus } +relations: + gen_bus: { columns: [generator, bus], key: generator } parameters: limit: { dims: [snapshot, bus] } @@ -109,14 +109,14 @@ objective: { sense: minimize, expression: sum(p) } $`\sum_{g \in \mathcal{G} \,:\, \mathrm{gen\_bus}(g) = b} p_{t,g} \le \mathrm{limit}_{t,b} \qquad \forall\, t \in \mathcal{T},\ b \in \mathcal{B}`$ -### `sum(array, by=[lookup, …])` +### `sum(array, by=[relation, …])` -`examples/operators/sum_by_lookups.yaml` +`examples/operators/sum_by_relations.yaml` ```yaml description: >- - Grouping through several maps at once — `sum(array, by=[lookup, …])` lands - the result on every dimension the lookups map into, which is one grouping + Grouping through several maps at once — `sum(array, by=[relation, …])` lands + the result on every dimension the relations map into, which is one grouping rather than a composition of two: the generator dimension is consumed once. dimensions: @@ -125,9 +125,9 @@ dimensions: bus: { dtype: str } technology: { dtype: str } -lookups: - gen_bus: { over: generator, into: bus } - gen_tech: { over: generator, into: technology } +relations: + gen_bus: { columns: [generator, bus], key: generator } + gen_tech: { columns: [generator, technology], key: generator } parameters: limit: { dims: [snapshot, bus, technology] } @@ -147,21 +147,21 @@ objective: { sense: minimize, expression: sum(p) } $`\sum_{g \in \mathcal{G} \,:\, \mathrm{gen\_bus}(g) = b \wedge \mathrm{gen\_tech}(g) = e} p_{t,g} \le \mathrm{limit}_{t,b,e} \qquad \forall\, t \in \mathcal{T},\ b \in \mathcal{B},\ e \in \mathcal{E}`$ -### `at(array, by=lookup)` +### `at(array, by=relation)` `examples/operators/at.yaml` ```yaml description: >- - The adjoint of the membership reduction — `at(array, by=lookup)` reads one + The adjoint of the membership reduction — `at(array, by=relation)` reads one coarse value once per fine label pointing at it. dimensions: snapshot: { dtype: int } period: { dtype: int } -lookups: - period_of: { over: snapshot, into: period } +relations: + period_of: { columns: [snapshot, period], key: snapshot } parameters: cap: { dims: [period] } @@ -181,7 +181,7 @@ objective: { sense: minimize, expression: sum(p) } $`p_{t} \le \mathrm{cap}_{\mathrm{period\_of}(t)} \qquad \forall\, t \in \mathcal{T}`$ -### `shift(array, over=dim, offset=n)` +### `shift(array, along=dim, offset=n)` `examples/operators/shift.yaml` @@ -201,14 +201,14 @@ variables: constraints: no_faster_than_before: dims: [snapshot] - expression: p <= shift(p, over=snapshot, offset=1) + expression: p <= shift(p, along=snapshot, offset=1) objective: { sense: minimize, expression: sum(p) } ``` $`p_{t} \le p_{t - 1} \qquad \forall\, t \in \mathcal{T}`$ -### `shift(array, over=dim, offset=n, edge='wrap')` +### `shift(array, along=dim, offset=n, edge='wrap')` `examples/operators/shift_wrap.yaml` @@ -228,14 +228,14 @@ variables: constraints: no_faster_than_before: dims: [snapshot] - expression: p <= shift(p, over=snapshot, offset=1, edge='wrap') + expression: p <= shift(p, along=snapshot, offset=1, edge='wrap') objective: { sense: minimize, expression: sum(p) } ``` $`p_{t} \le p_{t \ominus 1} \qquad \forall\, t \in \mathcal{T}`$ -### `shift(array, over=dim, offset=n, edge=v)` +### `shift(array, along=dim, offset=n, edge=v)` `examples/operators/shift_edge.yaml` @@ -255,14 +255,14 @@ variables: constraints: no_faster_than_before: dims: [snapshot] - expression: p <= shift(p, over=snapshot, offset=1, edge=0) + expression: p <= shift(p, along=snapshot, offset=1, edge=0) objective: { sense: minimize, expression: sum(p) } ``` $`p_{t} \le p_{t \boxminus_{0} 1} \qquad \forall\, t \in \mathcal{T}`$ -### `shift(array, over=dim, offset=p, edge=…)` +### `shift(array, along=dim, offset=p, edge=…)` `examples/operators/shift_by_parameter.yaml` @@ -288,14 +288,14 @@ variables: constraints: arrives_after_its_lead: dims: [technology, month] - expression: shift(order, over=month, offset=lead, edge=0) >= demand + expression: shift(order, along=month, offset=lead, edge=0) >= demand objective: { sense: minimize, expression: sum(order) } ``` $`\mathit{order}_{t,m \boxminus_{0} \mathrm{lead}} \ge \mathrm{demand}_{t,m} \qquad \forall\, t \in \mathcal{T},\ m \in \mathcal{M}`$ -### `shift(array, over=dim, offset=n, by=lookup)` +### `shift(array, along=dim, offset=n, by=relation)` `examples/operators/shift_partitioned.yaml` @@ -308,8 +308,8 @@ dimensions: snapshot: { dtype: int } season: { dtype: str } -lookups: - season_of: { over: snapshot, into: season } +relations: + season_of: { columns: [snapshot, season], key: snapshot } variables: p: @@ -319,14 +319,14 @@ variables: constraints: no_faster_than_before_in_season: dims: [snapshot] - expression: p <= shift(p, over=snapshot, offset=1, edge='wrap', by=season_of) + expression: p <= shift(p, along=snapshot, offset=1, edge='wrap', by=season_of) objective: { sense: minimize, expression: sum(p) } ``` $`p_{t} \le p_{t \ominus^{\mathrm{season\_of}(t)} 1} \qquad \forall\, t \in \mathcal{T}`$ -### `sum_back(array, over=dim, within=n)` +### `sum_back(array, along=dim, window=n)` `examples/operators/sum_back.yaml` @@ -353,14 +353,14 @@ variables: constraints: stays_up_its_own_time: dims: [unit, hour] - expression: sum_back(started, over=hour, within=3) <= on + expression: sum_back(started, along=hour, window=3) <= on objective: { sense: minimize, expression: sum(on) } ``` $`\sum_{h' \in \mathcal{H} \,:\, 0 \le h - h' < 3} \mathit{started}_{u,h'} \le \mathit{on}_{u,h} \qquad \forall\, u \in \mathcal{U},\ h \in \mathcal{H}`$ -### `sum_back(array, over=dim, within=p)` +### `sum_back(array, along=dim, window=p)` `examples/operators/sum_back_by_parameter.yaml` @@ -387,14 +387,14 @@ variables: constraints: stays_up_its_own_time: dims: [unit, hour] - expression: sum_back(started, over=hour, within=min_up) <= on + expression: sum_back(started, along=hour, window=min_up) <= on objective: { sense: minimize, expression: sum(on) } ``` $`\sum_{h' \in \mathcal{H} \,:\, 0 \le h - h' < \mathrm{min\_up}} \mathit{started}_{u,h'} \le \mathit{on}_{u,h} \qquad \forall\, u \in \mathcal{U},\ h \in \mathcal{H}`$ -### `sum_back(array, over=dim, within=p, edge='wrap')` +### `sum_back(array, along=dim, window=p, edge='wrap')` `examples/operators/sum_back_wrap.yaml` @@ -421,14 +421,14 @@ variables: constraints: stays_up_its_own_time: dims: [unit, hour] - expression: sum_back(started, over=hour, within=min_up, edge='wrap') <= on + expression: sum_back(started, along=hour, window=min_up, edge='wrap') <= on objective: { sense: minimize, expression: sum(on) } ``` $`\sum_{h' \in \mathcal{H} \,:\, 0 \le h \ominus h' < \mathrm{min\_up}} \mathit{started}_{u,h'} \le \mathit{on}_{u,h} \qquad \forall\, u \in \mathcal{U},\ h \in \mathcal{H}`$ -### `sum_back(array, over=dim, within=n, by=lookup)` +### `sum_back(array, along=dim, window=n, by=relation)` `examples/operators/sum_back_partitioned.yaml` @@ -443,8 +443,8 @@ dimensions: hour: { dtype: int } day: { dtype: str } -lookups: - day_of: { over: hour, into: day } +relations: + day_of: { columns: [hour, day], key: hour } variables: started: @@ -457,7 +457,7 @@ variables: constraints: stays_up_inside_its_day: dims: [unit, hour] - expression: sum_back(started, over=hour, within=3, by=day_of) <= on + expression: sum_back(started, along=hour, window=3, by=day_of) <= on objective: { sense: minimize, expression: sum(on) } ``` diff --git a/docs/examples/pypsa.md b/docs/examples/pypsa.md index 2c4a6d93..f71824cc 100644 --- a/docs/examples/pypsa.md +++ b/docs/examples/pypsa.md @@ -5,41 +5,34 @@ SPDX-License-Identifier: CC-BY-4.0 # PyPSA in one file -This file states the model that a plain `n.optimize()` builds. It grows one -**rung** at a time, where a rung is one `n.optimize()` keyword stated in full, -towards [milestone 1](https://github.com/energy-models/math-spec/milestone/1). - -The index below lists every row that PyPSA `1.3.0` emits from -`pypsa/optimization/`. Once a row is stated here, the index links it to its block -in the file. The blocks are generated, so a row that stops loading or changes its +The model a plain `n.optimize()` builds, stated as one file and grown a rung +at a time towards +[milestone 1](https://github.com/energy-models/math-spec/milestone/1). The +index below lists every row PyPSA emits (PyPSA `1.3.0`, +`pypsa/optimization/`) and links each to its block in the file once it is +there. The blocks are generated, so a row that stops loading or changes its math fails CI. -Three rules shape the file: - -- Bounds are the explicit rows that PyPSA writes, so their duals are row duals. -- Regimes are data columns and `where:` masks, never variants of the file. -- Names are PyPSA's own, in the form `Component_attribute`. The symbol table - `examples/symbols/pypsa.yaml` makes the math read as math. +Three rules shape the file. Bounds are the explicit rows PyPSA writes, so +their duals are row duals. Regimes are data columns and `where:` masks, never +file variants. Names are PyPSA's, `Component_attribute`, with a symbol table +(`examples/symbols/pypsa.yaml`) making the math read as math. ## Index -A row is **done**, and gets a link, once the file states it as the one block that -PyPSA builds, on this branch as it stands. A fix still on its way stays not-done, -and the note carries its PR or issue. Three words say how far a row has to go: - -- **split**: the same feasible region and the same optimum, stated differently, - as several `where:` blocks or with a bookkeeping difference the note names. -- **open**: not stated yet. -- **out**: never stated, deliberately. PyPSA emits it only under the keyword, - scope or version the note names. +A row is **done** and links once the file states it as the one block PyPSA +builds — on this branch, as it stands; a fix still on its way stays +not-done, its PR or issue in the note. Three words say the distance: +**split** — the same feasible region and optimum under a different +statement: several `where:` blocks, or a bookkeeping difference the note +names · **open** — not stated yet · **out** — never stated, deliberately: +emitted only under the keyword, scope or version the note names. A name carrying `{k}` or `{s}` stands for the family PyPSA numbers per segment or scenario. -A name that carries `{k}` or `{s}` stands for the family that PyPSA numbers per -segment or per scenario. - -Each rung's banner below states what PyPSA solved its reference network to. What -an engine makes of the same rung is that engine's own record: the objective, the -prices, and the two linopy models compared label for label. lpspec certifies -itself against these rungs under `differential/pypsa/` in its own tree. +Each rung's banner below states what PyPSA solved its reference network +to. What an engine makes of the same rung — the objective and prices across +the fence, and the two linopy models label for label — is that engine's own +record: lpspec certifies itself against these rungs under +`differential/pypsa/` in its own tree. > Every rung's network is `spine.build()` plus the rung's own `n.add` calls, data inline; a keyword not passed is PyPSA's default. A banner states what PyPSA solved the rung to; how an engine binds the network to the file, and what it makes of it, is that engine's own record. @@ -420,8 +413,8 @@ def build(): ### Rung 5 — global constraints -PyPSA names all of these `GlobalConstraint-{name}`. The type and the comparator -are both data, so each type becomes three blocks here, one per sense. +`GlobalConstraint-{name}` for all; the type and the comparator are data, so +each type is three blocks by sense. | PyPSA type | status | note | | ------------------------------------- | ----------- | ------------------------------------------------- | @@ -612,7 +605,7 @@ def build(): | [`{c}-com-p-lower/upper`](#generator-com-p-lower) | done | | | [`{c}-*-p-fixed-upper`](#generator-status-p-fixed-upper) | done | status, start and stop each at most one, as explicit rows | | [`{c}-com-transition-start-up/shut-down`](#generator-com-transition-start-up) | done | the state carried into a snapshot is a cased quantity, so the first snapshot needs no block of its own | -| [`{c}-com-up-time`, `-down-time`](#generator-com-up-time) | done | `sum_back(within=min_up_time)` | +| [`{c}-com-up-time`, `-down-time`](#generator-com-up-time) | done | `sum_back(window=min_up_time)` | | [`{c}-com-status-*-must_stay_up`](#generator-com-status-min_up_time_must_stay_up) | done | the window is a prep mask — `position()` takes a literal, not a parameter | | [`stand_by_cost`, `start_up_cost`, `shut_down_cost`](#objective) | done | | | [`{c}-com-p-before/-current/-partly-*`](pypsa_linearized_uc.md) | done | rung 12, a file of its own | @@ -825,10 +818,9 @@ def build(): ### Rung 11 — ac-dc-meshed -This is PyPSA's `ac_dc_meshed` example in full. It has meshed AC and DC, -extendable lines, links and generators, carriers, and a CO2 budget. It composes -every statement above. It is also the first rung that has an objective -constant. +PyPSA's `ac_dc_meshed` example, whole: meshed AC and DC, extendable lines, +links and generators, carriers, a CO2 budget. Every statement above, +composed; the first rung with an objective constant. > ✔ `pypsa 1.3.0` solves this rung's network at objective `-3474256.0405499237`, 468 rows. @@ -1115,25 +1107,18 @@ def build(): ### Rung 16 — link delay -This rung has a source feeding two sinks, over links whose energy arrives late. - -PyPSA's `delay` lags a port's delivery by a number of snapshots. `cyclic_delay` -says what happens to the flow that is still in transit at the edge of the -horizon: it either wraps around to the start, or it is lost. +A source feeding two sinks over links whose energy arrives late. PyPSA's +`delay` lags a port's delivery by a number of snapshots, and `cyclic_delay` +says whether the flow still in transit at the horizon's edge wraps to the start +or is lost. The two are a per-link number and a per-link kind, so the balance +turns them on with a `cases:` block over `shift(…, offset=Link_output_delay, +edge=…)` — one arm wrapping (`edge='wrap'`), the other vacating (`edge=0`). -So the two settings are a per-link number and a per-link kind. The balance -turns them on with a `cases:` block over -`shift(…, offset=Link_output_delay, edge=…)`. One arm wraps, with -`edge='wrap'`, and the other vacates, with `edge=0`. - -This is the one rung whose `generators` weighting is uniform, and that matters. -PyPSA measures `delay` in those units. So with a uniform column, a delay of `n` -is a shift of exactly `n` snapshot positions, which a positional `shift` -reproduces. - -Under a non-uniform column, PyPSA resamples by elapsed time rather than by -position. That is a shift which varies along the snapshot axis, and it is above -what `shift` can state (#299). +This is the one rung whose `generators` weighting is uniform. PyPSA measures +`delay` in those units, so a uniform column makes a delay of `n` a shift of +exactly `n` snapshot positions, which a positional `shift` reproduces. Under a +non-uniform column PyPSA resamples by elapsed time rather than by position — a +shift that varies along the snapshot axis, above what `shift` states (#299). | PyPSA | status | note | | ------------------------- | ------ | ----------------------------------------------- | @@ -1212,14 +1197,10 @@ def build(): ## Refusals -Where PyPSA refuses to build, parity means refusing here too. - -None of these is a gap in the language. Each one is a data check that has not -been made yet. Where each check should live is one open question, and the -candidates are the language, data preparation, and the harness. - -The line numbers below are pinned to pypsa 1.3.0, which is the version the -records above come from. +Where PyPSA refuses to build, parity means refusing too. None is a language +gap; each is a data check not made yet, and where it should live — language, +data prep, or harness — is one open question. Line numbers are pinned pypsa +1.3.0, the version the records above are from. | PyPSA raises | on | here | note | | -------------------------------------------- | ------------------------------------------------- | ----------------------- | ---- | @@ -1229,9 +1210,9 @@ records above come from. | `NotImplementedError`, `global_constraints.py:457` | depletion with period weightings `!= 1` | out | | | `ValueError`/`RuntimeError`, losses | `s_nom_max = inf`; secant cap | out | | -The harness on the lpspec side reads the duals and the solutions back. -`marginal_price` is the balance dual over `w_objective`. `mu_upper` is the -concatenation of the regime blocks. `p0` and `p1` are derived from `Link-p`. +Duals and solutions are read back by the harness on the lpspec side: +`marginal_price` is the balance dual over `w_objective`, `mu_upper` the +concatenation of the regime blocks, `p0`/`p1` derived from `Link-p`. ## The file @@ -1243,9 +1224,9 @@ The model a plain `n.optimize()` builds, stated in one file. Every declaration i | Symbol | Meaning | |---|---| | $`\mathcal{T}`$ | index $`t`$ — `snapshot` — dispatch periods | -| $`\mathcal{N}`$ | index $`n`$ — `bus` — network nodes | +| $`\mathcal{N}`$ | index $`n`$ — `bus` with $`\mathrm{Generator\_bus}: \mathcal{G} \to \mathcal{N},\ \mathrm{Link\_bus0}: \mathcal{L} \to \mathcal{N},\ \mathrm{Link\_output\_bus}: \mathcal{O} \to \mathcal{N},\ \mathrm{Load\_bus}: \mathcal{D} \to \mathcal{N},\ \mathrm{StorageUnit\_bus}: \mathcal{S} \to \mathcal{N},\ \mathrm{Line\_bus0}: \mathcal{K} \to \mathcal{N},\ \mathrm{Line\_bus1}: \mathcal{K} \to \mathcal{N},\ \mathrm{Store\_bus}: \mathcal{V} \to \mathcal{N}`$ — network nodes | | $`\mathcal{G}`$ | index $`g`$ — `generator` with $`\mathrm{Generator\_bus}: \mathcal{G} \to \mathcal{N}`$ — generating units, each on one bus | -| $`\mathcal{L}`$ | index $`l`$ — `link` with $`\mathrm{Link\_bus0}: \mathcal{L} \to \mathcal{N}`$ — controllable connections, each from one bus to the buses it delivers to | +| $`\mathcal{L}`$ | index $`l`$ — `link` with $`\mathrm{Link\_bus0}: \mathcal{L} \to \mathcal{N},\ \mathrm{Link\_output\_link}: \mathcal{O} \to \mathcal{L}`$ — controllable connections, each from one bus to the buses it delivers to | | $`\mathcal{O}`$ | index $`o`$ — `link_output` with $`\mathrm{Link\_output\_link}: \mathcal{O} \to \mathcal{L},\ \mathrm{Link\_output\_bus}: \mathcal{O} \to \mathcal{N}`$ — a link's output ports, one label per port a link declares — PyPSA's `bus1`, `bus2`, … columns read long, so a link of any number of output ports is one term in the balance, data prep | | $`\mathcal{D}`$ | index $`d`$ — `load` with $`\mathrm{Load\_bus}: \mathcal{D} \to \mathcal{N}`$ — demands, each on one bus | | $`\mathcal{S}`$ | index $`s`$ — `storage_unit` with $`\mathrm{StorageUnit\_bus}: \mathcal{S} \to \mathcal{N}`$ — storage units, dispatch and store behind one bus connection | @@ -1874,7 +1855,7 @@ Generator_com_up_time: up time's, which the must-stay-up mask carries dims: [snapshot, generator] where: Generator_committable AND Generator_min_up_time > 0 AND position(snapshot) > 0 - expression: sum_back(Generator_start_up, over=snapshot, within=Generator_min_up_time) <= Generator_status + expression: sum_back(Generator_start_up, along=snapshot, window=Generator_min_up_time) <= Generator_status ``` ```math @@ -1890,7 +1871,7 @@ Generator_com_down_time: description: "`Generator-com-down-time` — a unit stopped within its own minimum down time is still off" dims: [snapshot, generator] where: Generator_committable AND Generator_min_down_time > 0 AND position(snapshot) > 0 - expression: sum_back(Generator_shut_down, over=snapshot, within=Generator_min_down_time) <= 1 - Generator_status + expression: sum_back(Generator_shut_down, along=snapshot, window=Generator_min_down_time) <= 1 - Generator_status ``` ```math @@ -2448,7 +2429,7 @@ Link_p_ramp_limit_up: optimize builds no row either dims: [snapshot, link] where: Link_ramp_limit_up - expression: Link_p - shift(Link_p, over=snapshot, offset=1) <= Link_ramp_limit_up * Link_p_nom_effective + expression: Link_p - shift(Link_p, along=snapshot, offset=1) <= Link_ramp_limit_up * Link_p_nom_effective ``` ```math @@ -2464,7 +2445,7 @@ Link_p_ramp_limit_down: description: "`Link-p-ramp_limit_down` — a link lowers flow no faster than its limit of the build" dims: [snapshot, link] where: Link_ramp_limit_down - expression: shift(Link_p, over=snapshot, offset=1) - Link_p <= Link_ramp_limit_down * Link_p_nom_effective + expression: shift(Link_p, along=snapshot, offset=1) - Link_p <= Link_ramp_limit_down * Link_p_nom_effective ``` ```math @@ -3129,7 +3110,7 @@ Generator_previous_status: dims: [snapshot, generator] cases: opening: { when: "position(snapshot) == 0", expression: Generator_status_initial } - otherwise: shift(Generator_status, over=snapshot, offset=1) + otherwise: shift(Generator_status, along=snapshot, offset=1) ``` ```math @@ -3147,7 +3128,7 @@ Generator_previous_p: dims: [snapshot, generator] cases: opening: { when: "position(snapshot) == 0", expression: 0 } - otherwise: shift(Generator_p, over=snapshot, offset=1) + otherwise: shift(Generator_p, along=snapshot, offset=1) ``` ```math @@ -3243,11 +3224,11 @@ StorageUnit_charge_carried_in: cases: cyclic: when: StorageUnit_cyclic_state_of_charge - expression: StorageUnit_retention * shift(StorageUnit_state_of_charge, over=snapshot, offset=1, edge='wrap') + expression: StorageUnit_retention * shift(StorageUnit_state_of_charge, along=snapshot, offset=1, edge='wrap') opening: when: not StorageUnit_cyclic_state_of_charge AND position(snapshot) == 0 expression: StorageUnit_state_of_charge_initial - otherwise: StorageUnit_retention * shift(StorageUnit_state_of_charge, over=snapshot, offset=1) + otherwise: StorageUnit_retention * shift(StorageUnit_state_of_charge, along=snapshot, offset=1) ``` ```math @@ -3267,11 +3248,11 @@ Store_energy_carried_in: cases: cyclic: when: Store_e_cyclic - expression: Store_retention * shift(Store_e, over=snapshot, offset=1, edge='wrap') + expression: Store_retention * shift(Store_e, along=snapshot, offset=1, edge='wrap') opening: when: not Store_e_cyclic AND position(snapshot) == 0 expression: Store_e_initial - otherwise: Store_retention * shift(Store_e, over=snapshot, offset=1) + otherwise: Store_retention * shift(Store_e, along=snapshot, offset=1) ``` ```math @@ -3293,8 +3274,8 @@ Link_output_arrival: cases: wrapping: when: Link_output_cyclic_delay - expression: shift(at(Link_p, by=Link_output_link) * Link_efficiency, over=snapshot, offset=Link_output_delay, edge='wrap') - otherwise: shift(at(Link_p, by=Link_output_link) * Link_efficiency, over=snapshot, offset=Link_output_delay, edge=0) + expression: shift(at(Link_p, by=Link_output_link) * Link_efficiency, along=snapshot, offset=Link_output_delay, edge='wrap') + otherwise: shift(at(Link_p, by=Link_output_link) * Link_efficiency, along=snapshot, offset=Link_output_delay, edge=0) ``` ```math diff --git a/docs/examples/pypsa_linearized_uc.md b/docs/examples/pypsa_linearized_uc.md index 94d9978e..bad04134 100644 --- a/docs/examples/pypsa_linearized_uc.md +++ b/docs/examples/pypsa_linearized_uc.md @@ -5,17 +5,15 @@ SPDX-License-Identifier: CC-BY-4.0 # PyPSA, the relaxed commitment -This is rung 12 of [PyPSA in one file](pypsa.md). It states -`n.optimize(linearized_unit_commitment=True)` on rungs 1 and 7, in a file of its -own. The model's description below says why it has its own file. Its network is -the shared spine plus the script's own additions. +Rung 12 of [PyPSA in one file](pypsa.md): `n.optimize(linearized_unit_commitment=True)`, stated on rungs 1 and 7 in a +file of its own — the model's description below says why. Its network is the spine plus the script's own additions. ## Rung 12 — linearized unit commitment | PyPSA | status | note | | --- | --- | --- | | [`Generator-status`, `-start_up`, `-shut_down`](#variable-domains) | done | shares in [0, 1], not binaries | -| [`Generator-com-p-before`](#generator-com-p-before) | done | used where a start and a stop cost the same. It is a boolean from data preparation | +| [`Generator-com-p-before`](#generator-com-p-before) | done | where start and stop cost the same — a data-prep bool | | [`Generator-com-p-current`](#generator-com-p-current) | done | | | [`Generator-com-partly-start-up`](#generator-com-partly-start-up) | done | | | [`Generator-com-partly-shut-down`](#generator-com-partly-shut-down) | done | | @@ -100,9 +98,9 @@ The relaxed class of a plain `n.optimize()`: `linearized_unit_commitment`, state | Symbol | Meaning | |---|---| | $`\mathcal{T}`$ | index $`t`$ — `snapshot` — dispatch periods | -| $`\mathcal{N}`$ | index $`n`$ — `bus` — network nodes | +| $`\mathcal{N}`$ | index $`n`$ — `bus` with $`\mathrm{Generator\_bus}: \mathcal{G} \to \mathcal{N},\ \mathrm{Link\_bus0}: \mathcal{L} \to \mathcal{N},\ \mathrm{Link\_output\_bus}: \mathcal{O} \to \mathcal{N},\ \mathrm{Load\_bus}: \mathcal{D} \to \mathcal{N}`$ — network nodes | | $`\mathcal{G}`$ | index $`g`$ — `generator` with $`\mathrm{Generator\_bus}: \mathcal{G} \to \mathcal{N}`$ — generating units, each on one bus | -| $`\mathcal{L}`$ | index $`l`$ — `link` with $`\mathrm{Link\_bus0}: \mathcal{L} \to \mathcal{N}`$ — controllable connections, each from one bus to the buses it delivers to | +| $`\mathcal{L}`$ | index $`l`$ — `link` with $`\mathrm{Link\_bus0}: \mathcal{L} \to \mathcal{N},\ \mathrm{Link\_output\_link}: \mathcal{O} \to \mathcal{L}`$ — controllable connections, each from one bus to the buses it delivers to | | $`\mathcal{O}`$ | index $`o`$ — `link_output` with $`\mathrm{Link\_output\_link}: \mathcal{O} \to \mathcal{L},\ \mathrm{Link\_output\_bus}: \mathcal{O} \to \mathcal{N}`$ — a link's output ports, one label per port a link declares — PyPSA's `bus1`, `bus2`, … columns read long, so a link of any number of output ports is one term in the balance, data prep | | $`\mathcal{D}`$ | index $`d`$ — `load` with $`\mathrm{Load\_bus}: \mathcal{D} \to \mathcal{N}`$ — demands, each on one bus | @@ -332,7 +330,7 @@ Generator_com_up_time: up time's, which the must-stay-up mask carries dims: [snapshot, generator] where: Generator_committable AND Generator_min_up_time > 0 AND position(snapshot) > 0 - expression: sum_back(Generator_start_up, over=snapshot, within=Generator_min_up_time) <= Generator_status + expression: sum_back(Generator_start_up, along=snapshot, window=Generator_min_up_time) <= Generator_status ``` ```math @@ -348,7 +346,7 @@ Generator_com_down_time: description: "`Generator-com-down-time` — a unit stopped within its own minimum down time is still off" dims: [snapshot, generator] where: Generator_committable AND Generator_min_down_time > 0 AND position(snapshot) > 0 - expression: sum_back(Generator_shut_down, over=snapshot, within=Generator_min_down_time) <= 1 - Generator_status + expression: sum_back(Generator_shut_down, along=snapshot, window=Generator_min_down_time) <= 1 - Generator_status ``` ```math @@ -487,8 +485,8 @@ Generator_com_p_before: dims: [snapshot, generator] where: Generator_committable AND Generator_partly_tightened expression: >- - shift(Generator_p, over=snapshot, offset=1) - - Generator_ramp_limit_shut_down * Generator_p_nom * shift(Generator_status, over=snapshot, offset=1) + shift(Generator_p, along=snapshot, offset=1) + - Generator_ramp_limit_shut_down * Generator_p_nom * shift(Generator_status, along=snapshot, offset=1) - (Generator_p_max_pu * Generator_p_nom - Generator_ramp_limit_shut_down * Generator_p_nom) * (Generator_status - Generator_start_up) <= 0 ``` @@ -525,9 +523,9 @@ Generator_com_partly_start_up: dims: [snapshot, generator] where: Generator_committable AND Generator_partly_tightened expression: >- - Generator_p - shift(Generator_p, over=snapshot, offset=1) + Generator_p - shift(Generator_p, along=snapshot, offset=1) - (Generator_p_min_pu * Generator_p_nom + Generator_ramp_limit_up * Generator_p_nom) * Generator_status - + Generator_p_min_pu * Generator_p_nom * shift(Generator_status, over=snapshot, offset=1) + + Generator_p_min_pu * Generator_p_nom * shift(Generator_status, along=snapshot, offset=1) + (Generator_p_min_pu * Generator_p_nom + Generator_ramp_limit_up * Generator_p_nom - Generator_ramp_limit_start_up * Generator_p_nom) * Generator_start_up <= 0 ``` @@ -546,8 +544,8 @@ Generator_com_partly_shut_down: dims: [snapshot, generator] where: Generator_committable AND Generator_partly_tightened expression: >- - shift(Generator_p, over=snapshot, offset=1) - Generator_p - - Generator_ramp_limit_shut_down * Generator_p_nom * shift(Generator_status, over=snapshot, offset=1) + shift(Generator_p, along=snapshot, offset=1) - Generator_p + - Generator_ramp_limit_shut_down * Generator_p_nom * shift(Generator_status, along=snapshot, offset=1) + (Generator_ramp_limit_shut_down * Generator_p_nom - Generator_ramp_limit_down * Generator_p_nom) * Generator_status - (Generator_p_min_pu * Generator_p_nom + Generator_ramp_limit_down * Generator_p_nom - Generator_ramp_limit_shut_down * Generator_p_nom) * Generator_start_up <= 0 @@ -567,7 +565,7 @@ Generator_previous_status: dims: [snapshot, generator] cases: opening: { when: "position(snapshot) == 0", expression: Generator_status_initial } - otherwise: shift(Generator_status, over=snapshot, offset=1) + otherwise: shift(Generator_status, along=snapshot, offset=1) ``` ```math @@ -585,7 +583,7 @@ Generator_previous_p: dims: [snapshot, generator] cases: opening: { when: "position(snapshot) == 0", expression: 0 } - otherwise: shift(Generator_p, over=snapshot, offset=1) + otherwise: shift(Generator_p, along=snapshot, offset=1) ``` ```math diff --git a/docs/examples/pypsa_losses.md b/docs/examples/pypsa_losses.md index eab0d2f2..1dc42940 100644 --- a/docs/examples/pypsa_losses.md +++ b/docs/examples/pypsa_losses.md @@ -5,11 +5,8 @@ SPDX-License-Identifier: CC-BY-4.0 # PyPSA, the lossy lines -This is rung 13 of [PyPSA in one file](pypsa.md). It states -`n.optimize(transmission_losses={'mode': 'tangents', 'segments': K})` on the -lines of rung 6, in a file of its own. The model's description below says why it -has its own file. Its network is the shared spine plus the script's own -additions. +Rung 13 of [PyPSA in one file](pypsa.md): `n.optimize(transmission_losses={'mode': 'tangents', 'segments': K})`, stated on rung 6's lines in a +file of its own — the model's description below says why. Its network is the spine plus the script's own additions. ## Rung 13 — transmission losses @@ -17,11 +14,11 @@ additions. | --- | --- | --- | | [`Line-loss`](#variable-domains) | done | | | [`Line-fix-s-*`, `Line-ext-s-*`](#line-fix-s-lower) | done | the loss counted against the rating | -| [`Bus-nodal_balance`](#bus-nodal_balance) | done | half of each incident line's loss, at either end | -| [`Line-loss_upper`](#line-loss_upper) | done | `loss_max` comes from data preparation | -| [`Line-loss_tangents-{k}-1`](#line-loss_tangents-k-1) | split | PyPSA names one row per segment. Here it is one block over the dimension | +| [`Bus-nodal_balance`](#bus-nodal_balance) | done | half of each incident line's loss at either end | +| [`Line-loss_upper`](#line-loss_upper) | done | `loss_max` is data prep | +| [`Line-loss_tangents-{k}-1`](#line-loss_tangents-k-1) | split | PyPSA names a row per segment; one block over the dimension | | [`Line-loss_tangents-{k}--1`](#line-loss_tangents-k--1) | split | | -| `Line-loss_secants-*` | out | the secant mode solves for its own segment count | +| `Line-loss_secants-*` | out | the secant mode solves for its segment count | > ✔ `pypsa 1.3.0` solves this rung's network at objective `10645.295879552297`, 150 rows. @@ -84,9 +81,9 @@ The lossy class of a plain `n.optimize()`: `transmission_losses` in its tangent | Symbol | Meaning | |---|---| | $`\mathcal{T}`$ | index $`t`$ — `snapshot` — dispatch periods | -| $`\mathcal{N}`$ | index $`n`$ — `bus` — network nodes | +| $`\mathcal{N}`$ | index $`n`$ — `bus` with $`\mathrm{Generator\_bus}: \mathcal{G} \to \mathcal{N},\ \mathrm{Line\_bus0}: \mathcal{K} \to \mathcal{N},\ \mathrm{Line\_bus1}: \mathcal{K} \to \mathcal{N},\ \mathrm{Link\_bus0}: \mathcal{L} \to \mathcal{N},\ \mathrm{Link\_output\_bus}: \mathcal{O} \to \mathcal{N},\ \mathrm{Load\_bus}: \mathcal{D} \to \mathcal{N}`$ — network nodes | | $`\mathcal{G}`$ | index $`g`$ — `generator` with $`\mathrm{Generator\_bus}: \mathcal{G} \to \mathcal{N}`$ — generating units, each on one bus | -| $`\mathcal{L}`$ | index $`l`$ — `link` with $`\mathrm{Link\_bus0}: \mathcal{L} \to \mathcal{N}`$ — controllable connections, each from one bus to the buses it delivers to | +| $`\mathcal{L}`$ | index $`l`$ — `link` with $`\mathrm{Link\_bus0}: \mathcal{L} \to \mathcal{N},\ \mathrm{Link\_output\_link}: \mathcal{O} \to \mathcal{L}`$ — controllable connections, each from one bus to the buses it delivers to | | $`\mathcal{O}`$ | index $`o`$ — `link_output` with $`\mathrm{Link\_output\_link}: \mathcal{O} \to \mathcal{L},\ \mathrm{Link\_output\_bus}: \mathcal{O} \to \mathcal{N}`$ — a link's output ports, one label per port a link declares — PyPSA's `bus1`, `bus2`, … columns read long, so a link of any number of output ports is one term in the balance, data prep | | $`\mathcal{K}`$ | index $`k`$ — `line` with $`\mathrm{Line\_bus0}: \mathcal{K} \to \mathcal{N},\ \mathrm{Line\_bus1}: \mathcal{K} \to \mathcal{N}`$ — passive branches, each between two buses, their flow set by impedance | | $`\mathcal{C}`$ | index $`c`$ — `cycle` — independent cycles of the passive network graph — the cycle basis, data prep | diff --git a/docs/examples/pypsa_multi_period.md b/docs/examples/pypsa_multi_period.md index e1334e8d..84177cd7 100644 --- a/docs/examples/pypsa_multi_period.md +++ b/docs/examples/pypsa_multi_period.md @@ -5,23 +5,18 @@ SPDX-License-Identifier: CC-BY-4.0 # PyPSA, the multi-period class -This is rung 15 of [PyPSA in one file](pypsa.md). It states -`n.optimize(multi_investment_periods=True)` on rungs 1 and 3, in a file of its -own. The model's description below says why it has its own file. - -Its network is a whole network rather than the shared spine. It has eight -snapshots over two investment periods, and the script carries the build years -and the lifetimes. +Rung 15 of [PyPSA in one file](pypsa.md): `n.optimize(multi_investment_periods=True)`, stated on rungs 1 and 3 in a +file of its own — the model's description below says why. Its network is a whole one: eight snapshots over two investment periods, build years and lifetimes on the script. ## Rung 15 — investment periods, with a growth limit | PyPSA | status | note | | --- | --- | --- | -| [`Generator-p`](#variable-domains) | done | built where the generator stands in the snapshot's period. That is `active`, which comes from data preparation | +| [`Generator-p`](#variable-domains) | done | where the generator stands in the snapshot's period — `active`, data prep | | [`Generator-fix-p-*`, `-ext-p-*`, `-ext-p_nom-*`](#generator-fix-p-lower) | done | rungs 1 and 3, masked by `active` | -| [`Carrier-growth_limit`](#carrier-growth_limit) | done | counted in the first period that a build stands in, with `edge=0` at the first period | -| [objective](#objective) | done | the period weight applies to operation. Capacity is counted once per period it stands in | -| `StorageUnit-energy_balance` per period, ramps at period starts | out | `shift(…, by=snapshot_period)` can state these. They are a later rung | +| [`Carrier-growth_limit`](#carrier-growth_limit) | done | counted in the first period a build stands in; `edge=0` at the first period | +| [objective](#objective) | done | period weight on operation; capacity once per period it stands in | +| `StorageUnit-energy_balance` per period, ramps at period starts | out | `shift(…, by=snapshot_period)` has them; a later rung | > ✔ `pypsa 1.3.0` solves this rung's network at objective `12747.19109626398`, 80 rows. @@ -121,13 +116,13 @@ The multi-period class of a plain `n.optimize()`: `multi_investment_periods`, st | Symbol | Meaning | |---|---| | $`\mathcal{T}`$ | index $`t`$ — `snapshot` with $`\mathrm{snapshot\_period}: \mathcal{T} \to \mathcal{Y}`$ — dispatch periods, positions across every investment period | -| $`\mathcal{Y}`$ | index $`y`$ — `period` — investment periods — PyPSA's `investment_periods` | -| $`\mathcal{N}`$ | index $`n`$ — `bus` — network nodes | +| $`\mathcal{Y}`$ | index $`y`$ — `period` with $`\mathrm{snapshot\_period}: \mathcal{T} \to \mathcal{Y}`$ — investment periods — PyPSA's `investment_periods` | +| $`\mathcal{N}`$ | index $`n`$ — `bus` with $`\mathrm{Generator\_bus}: \mathcal{G} \to \mathcal{N},\ \mathrm{Link\_bus0}: \mathcal{L} \to \mathcal{N},\ \mathrm{Link\_output\_bus}: \mathcal{O} \to \mathcal{N},\ \mathrm{Load\_bus}: \mathcal{D} \to \mathcal{N}`$ — network nodes | | $`\mathcal{G}`$ | index $`g`$ — `generator` with $`\mathrm{Generator\_carrier}: \mathcal{G} \to \mathcal{C},\ \mathrm{Generator\_bus}: \mathcal{G} \to \mathcal{N}`$ — generating units, each on one bus | -| $`\mathcal{L}`$ | index $`l`$ — `link` with $`\mathrm{Link\_bus0}: \mathcal{L} \to \mathcal{N}`$ — controllable connections, each from one bus to the buses it delivers to | +| $`\mathcal{L}`$ | index $`l`$ — `link` with $`\mathrm{Link\_bus0}: \mathcal{L} \to \mathcal{N},\ \mathrm{Link\_output\_link}: \mathcal{O} \to \mathcal{L}`$ — controllable connections, each from one bus to the buses it delivers to | | $`\mathcal{O}`$ | index $`o`$ — `link_output` with $`\mathrm{Link\_output\_link}: \mathcal{O} \to \mathcal{L},\ \mathrm{Link\_output\_bus}: \mathcal{O} \to \mathcal{N}`$ — a link's output ports, one label per port a link declares — PyPSA's `bus1`, `bus2`, … columns read long, so a link of any number of output ports is one term in the balance, data prep | | $`\mathcal{D}`$ | index $`d`$ — `load` with $`\mathrm{Load\_bus}: \mathcal{D} \to \mathcal{N}`$ — demands, each on one bus | -| $`\mathcal{C}`$ | index $`c`$ — `carrier` — energy carriers, what a growth limit is set per | +| $`\mathcal{C}`$ | index $`c`$ — `carrier` with $`\mathrm{Generator\_carrier}: \mathcal{G} \to \mathcal{C}`$ — energy carriers, what a growth limit is set per | #### Parameters @@ -344,7 +339,7 @@ Carrier_growth_limit: where: Carrier_max_growth expression: >- sum(Generator_p_nom_ext * Generator_first_active, by=Generator_carrier) - - shift(sum(Generator_p_nom_ext * Generator_first_active, by=Generator_carrier), over=period, offset=1, edge=0) + - shift(sum(Generator_p_nom_ext * Generator_first_active, by=Generator_carrier), along=period, offset=1, edge=0) * Carrier_max_relative_growth <= Carrier_max_growth ``` diff --git a/docs/examples/pypsa_quadratic.md b/docs/examples/pypsa_quadratic.md index c4455fd5..a7b8be2f 100644 --- a/docs/examples/pypsa_quadratic.md +++ b/docs/examples/pypsa_quadratic.md @@ -5,18 +5,16 @@ SPDX-License-Identifier: CC-BY-4.0 # PyPSA, the quadratic class -This is rung 10 of [PyPSA in one file](pypsa.md). It states PyPSA's -`marginal_cost_quadratic` on the transport surface of rung 1, in a file of its -own. The model's description below says why it has its own file. - -Its reference network starts from the same shared spine, `data/base/`, which is -shown once on [the rung ladder's page](pypsa.md#index). +Rung 10 of [PyPSA in one file](pypsa.md): PyPSA's `marginal_cost_quadratic`, +stated on rung 1's transport surface in a file of its own — the model's +description below says why. Its reference network starts from the same shared +spine, `data/base/`, shown once on [the rung ladder's page](pypsa.md#index). ## Rung 10 — quadratic costs | PyPSA | status | note | | --------------------------------------- | ------ | -------------------------------------------------------------------------- | -| [`marginal_cost_quadratic`](#objective) | done | degree 2 in the objective, on Generator and Link here. PyPSA also carries it on storage units and stores, which is one more term each of the same shape | +| [`marginal_cost_quadratic`](#objective) | done | degree 2 in the objective; Generator and Link here — PyPSA also carries it on storage units and stores, one more term each of the same shape | > ✔ `pypsa 1.3.0` solves this rung's network at objective `12587.437500000098`, 60 rows. @@ -72,9 +70,9 @@ The quadratic class of a plain `n.optimize()`: PyPSA's `marginal_cost_quadratic` | Symbol | Meaning | |---|---| | $`\mathcal{T}`$ | index $`t`$ — `snapshot` — dispatch periods | -| $`\mathcal{N}`$ | index $`n`$ — `bus` — network nodes | +| $`\mathcal{N}`$ | index $`n`$ — `bus` with $`\mathrm{Generator\_bus}: \mathcal{G} \to \mathcal{N},\ \mathrm{Link\_bus0}: \mathcal{L} \to \mathcal{N},\ \mathrm{Link\_output\_bus}: \mathcal{O} \to \mathcal{N},\ \mathrm{Load\_bus}: \mathcal{D} \to \mathcal{N}`$ — network nodes | | $`\mathcal{G}`$ | index $`g`$ — `generator` with $`\mathrm{Generator\_bus}: \mathcal{G} \to \mathcal{N}`$ — generating units, each on one bus | -| $`\mathcal{L}`$ | index $`l`$ — `link` with $`\mathrm{Link\_bus0}: \mathcal{L} \to \mathcal{N}`$ — controllable connections, each from one bus to the buses it delivers to | +| $`\mathcal{L}`$ | index $`l`$ — `link` with $`\mathrm{Link\_bus0}: \mathcal{L} \to \mathcal{N},\ \mathrm{Link\_output\_link}: \mathcal{O} \to \mathcal{L}`$ — controllable connections, each from one bus to the buses it delivers to | | $`\mathcal{O}`$ | index $`o`$ — `link_output` with $`\mathrm{Link\_output\_link}: \mathcal{O} \to \mathcal{L},\ \mathrm{Link\_output\_bus}: \mathcal{O} \to \mathcal{N}`$ — a link's output ports, one label per port a link declares — PyPSA's `bus1`, `bus2`, … columns read long, so a link of any number of output ports is one term in the balance, data prep | | $`\mathcal{D}`$ | index $`d`$ — `load` with $`\mathrm{Load\_bus}: \mathcal{D} \to \mathcal{N}`$ — demands, each on one bus | diff --git a/docs/examples/pypsa_stochastic.md b/docs/examples/pypsa_stochastic.md index acda4b02..d6b97cfc 100644 --- a/docs/examples/pypsa_stochastic.md +++ b/docs/examples/pypsa_stochastic.md @@ -5,22 +5,19 @@ SPDX-License-Identifier: CC-BY-4.0 # PyPSA, the two-stage class -This is rung 14 of [PyPSA in one file](pypsa.md). It states -`n.set_scenarios(...)` together with `n.set_risk_preference(alpha, omega)`, on -rungs 1 and 3, in a file of its own. The model's description below says why it -has its own file. Its network is the shared spine plus the script's own -additions. +Rung 14 of [PyPSA in one file](pypsa.md): `n.set_scenarios(...)` with `n.set_risk_preference(alpha, omega)`, stated on rungs 1 and 3 in a +file of its own — the model's description below says why. Its network is the spine plus the script's own additions. ## Rung 14 — two-stage stochastic, with CVaR | PyPSA | status | note | | --- | --- | --- | -| [`Generator-p`, `Link-p`](#variable-domains) | done | over `scenario`. `Generator-p_nom` is not, because it is chosen once | +| [`Generator-p`, `Link-p`](#variable-domains) | done | over `scenario`; `Generator-p_nom` is not — chosen once | | [`Generator-fix-p-*`, `-ext-p-*`, `Link-fix-p-*`, `Bus-nodal_balance`](#generator-fix-p-lower) | done | rungs 1 and 3, over `scenario` | | [`CVaR-a`, `CVaR-theta`, `CVaR`](#variable-domains) | done | | -| [`CVaR-excess-{s}`](#cvar-excess-s) | split | PyPSA names one row per scenario. Here it is one block over the dimension | -| [`CVaR-def`](#cvar-def) | done | `1 / (1 - alpha)` comes from data preparation | -| [objective](#objective) | done | capacity once. Operation is weighted `(1 - omega)` in expectation, and `omega` at the tail | +| [`CVaR-excess-{s}`](#cvar-excess-s) | split | PyPSA names a row per scenario; one block over the dimension | +| [`CVaR-def`](#cvar-def) | done | `1 / (1 - alpha)` is data prep | +| [objective](#objective) | done | capacity once; operation `(1 - omega)` in expectation, `omega` at the tail | > ✔ `pypsa 1.3.0` solves this rung's network at objective `9267.386666666665`, 87 rows. @@ -72,9 +69,9 @@ The two-stage class of a plain `n.optimize()`: a network with scenarios, stated |---|---| | $`\mathcal{S}`$ | index $`s`$ — `scenario` — the futures dispatch is chosen in, each with a weight | | $`\mathcal{T}`$ | index $`t`$ — `snapshot` — dispatch periods | -| $`\mathcal{N}`$ | index $`n`$ — `bus` — network nodes | +| $`\mathcal{N}`$ | index $`n`$ — `bus` with $`\mathrm{Generator\_bus}: \mathcal{G} \to \mathcal{N},\ \mathrm{Link\_bus0}: \mathcal{L} \to \mathcal{N},\ \mathrm{Link\_output\_bus}: \mathcal{O} \to \mathcal{N},\ \mathrm{Load\_bus}: \mathcal{D} \to \mathcal{N}`$ — network nodes | | $`\mathcal{G}`$ | index $`g`$ — `generator` with $`\mathrm{Generator\_bus}: \mathcal{G} \to \mathcal{N}`$ — generating units, each on one bus | -| $`\mathcal{L}`$ | index $`l`$ — `link` with $`\mathrm{Link\_bus0}: \mathcal{L} \to \mathcal{N}`$ — controllable connections, each from one bus to the buses it delivers to | +| $`\mathcal{L}`$ | index $`l`$ — `link` with $`\mathrm{Link\_bus0}: \mathcal{L} \to \mathcal{N},\ \mathrm{Link\_output\_link}: \mathcal{O} \to \mathcal{L}`$ — controllable connections, each from one bus to the buses it delivers to | | $`\mathcal{O}`$ | index $`o`$ — `link_output` with $`\mathrm{Link\_output\_link}: \mathcal{O} \to \mathcal{L},\ \mathrm{Link\_output\_bus}: \mathcal{O} \to \mathcal{N}`$ — a link's output ports, one label per port a link declares — PyPSA's `bus1`, `bus2`, … columns read long, so a link of any number of output ports is one term in the balance, data prep | | $`\mathcal{D}`$ | index $`d`$ — `load` with $`\mathrm{Load\_bus}: \mathcal{D} \to \mathcal{N}`$ — demands, each on one bus | diff --git a/docs/reference/language/absence.md b/docs/reference/language/absence.md index 84a8f433..478fc0db 100644 --- a/docs/reference/language/absence.md +++ b/docs/reference/language/absence.md @@ -28,12 +28,12 @@ This page says what the mask means for the rows that are built. ## What creates absence -| Construct | What is absent | -| -------------------------------------------- | ---------------------------------------------------------------- | -| `where:` on a variable | the variable, at the masked coordinates | -| `where:` on a constraint | the row | -| `shift(x, over=d, offset=n)` without `edge=` | the vacated edge coordinate ([shift](operators.md#shift)) | -| a label a lookup does not map | that label's group membership ([lookups](dimensions.md#lookups)) | +| Construct | What is absent | +| --------------------------------------------- | -------------------------------------------------------------------- | +| `where:` on a variable | the variable, at the masked coordinates | +| `where:` on a constraint | the row | +| `shift(x, along=d, offset=n)` without `edge=` | the vacated edge coordinate ([shift](operators.md#shift)) | +| a label a relation does not map | that label's group membership ([relations](dimensions.md#relations)) | Nothing else creates absence. **A missing parameter row is not absence.** A sparse table is a compressed dense table, and a missing row reads as the value @@ -86,13 +86,13 @@ there instead, write `where: rel_max` on the constraint. Every operator falls on one side of the line, and one question decides which: does an output slot stand for several input slots, or for one? -| Operator | An output slot reads | An absent input | -| ------------------------------- | ------------------------------- | ------------------------------------ | -| `sum(x, over=d)` | every position along `d` | is one summand fewer; the row stands | -| `sum(x, by=lookup)` | every member of the group | is one summand fewer; the row stands | -| `sum_back(x, over=d, within=w)` | the positions the window covers | is one summand fewer; the row stands | -| `shift(x, over=d, offset=n)` | one position, `n` back | _is_ the output, so it spreads | -| `at(x, by=lookup)` | one position, through the map | _is_ the output, so it spreads | +| Operator | An output slot reads | An absent input | +| -------------------------------- | ------------------------------- | ------------------------------------ | +| `sum(x, over=d)` | every position along `d` | is one summand fewer; the row stands | +| `sum(x, by=relation)` | every member of the group | is one summand fewer; the row stands | +| `sum_back(x, along=d, window=w)` | the positions the window covers | is one summand fewer; the row stands | +| `shift(x, along=d, offset=n)` | one position, `n` back | _is_ the output, so it spreads | +| `at(x, by=relation)` | one position, through the map | _is_ the output, so it spreads | The three summing operators put several slots into one, so a missing slot gives a shorter sum and the row survives. A window that reaches past the start of its @@ -164,6 +164,6 @@ separate not-a-number. | ---------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------- | | the row kept, the masked variable read as zero | `absence: zero` on the variable | | the row dropped where a parameter has no data | `where: p` on the constraint | -| a vacated shift position to contribute | `shift(x, over=d, offset=n, edge=0)` | +| a vacated shift position to contribute | `shift(x, along=d, offset=n, edge=0)` | | to test whether a variable exists here | its bare name in a `where` | | a bound only where the data has one | supply the bound, because `inf` is a value, or mask the variable. These are different models, so the language infers neither | diff --git a/docs/reference/language/declarations.md b/docs/reference/language/declarations.md index a2fe9259..d7f7f2ae 100644 --- a/docs/reference/language/declarations.md +++ b/docs/reference/language/declarations.md @@ -163,7 +163,7 @@ Two regimes of one rule are two blocks, each with a name a reader chose: ```yaml storage_balance: dims: [snapshot, storage] - expression: soc == shift(soc, over=snapshot, offset=1) * (1 - loss) + charge - discharge + expression: soc == shift(soc, along=snapshot, offset=1) * (1 - loss) + charge - discharge storage_balance_initial: dims: [snapshot, storage] diff --git a/docs/reference/language/dimensions.md b/docs/reference/language/dimensions.md index 3021c6a2..f1604225 100644 --- a/docs/reference/language/dimensions.md +++ b/docs/reference/language/dimensions.md @@ -3,13 +3,13 @@ SPDX-FileCopyrightText: math-spec contributors SPDX-License-Identifier: CC-BY-4.0 --> -# Dimensions and lookups +# Dimensions and relations A **dimension** is an axis of the model, such as `snapshot` or `generator`. -Declarations are indexed by it, and `sum` reduces along it. +Declarations are indexed by it, and `sum` reduces over it. -A **lookup** is a named map out of a dimension: one value for each of its -members. A generator's bus is a lookup, and so is a snapshot's period. +A **relation** is a named table between dimensions: a generator's bus, a +snapshot's period, or the buses a generator may connect to. ## `dimensions` @@ -48,28 +48,27 @@ same model. sort them, whether they are strings, integers or dates. [`shift`](operators.md#shift), `sum_back` and `position()` all count along this order, so an engine that sorted `snapshot` would give - `shift(p, over=snapshot, offset=1)` a different meaning. To get a + `shift(p, along=snapshot, offset=1)` a different meaning. To get a particular order, write the table in that order. 3. **A table has each coordinate at most once.** Two rows for `snapshot == 3` is an error that names `3`. The engine does not keep the last, keep the first, or add them. _At most_ once, not exactly once: a coordinate with no - row is [absence](absence.md), and absence is how a model masks. A lookup's + row is [absence](absence.md), and absence is how a model masks. A relation's table obeys the same rule. Every dimension has one list of members, and every parameter is lined up against it when the data binds. So if `load` has 8760 snapshots and `price` has 8759, the engine raises an error rather than build a model with one snapshot dropped. -## `lookups` +## `relations` -A lookup is how the network's wiring stays in the data. Which bus each generator -sits on, which two buses each line joins, which period a snapshot falls in: each -is a lookup table, and the file holds no adjacency matrix. - -Declare each lookup under its own name. `over:` names the dimension whose -members carry the value, and `into:` names the dimension the values are labels -of. [`sum(by=)` and `at(by=)`](operators.md) land terms on that target -dimension: +A relation is what makes topology _data_: a generator sits on a bus, a line has +two endpoints, a snapshot falls in a period, and no adjacency matrix or +hand-written join appears anywhere. A relation is a **table with one column per +dimension it relates**, and `key:` is the claim that makes it a map: one row per +key tuple, so the other columns are a function of the key. The declaration +fixes no direction; the operator that walks the table says which column it +consumes and which it produces. ```yaml dimensions: @@ -78,47 +77,212 @@ dimensions: line: { dtype: str } snapshot: { dtype: int } period: { dtype: int } -lookups: - gen_bus: { over: generator, into: bus } - line_from: { over: line, into: bus } # two lookups onto one dimension - line_to: { over: line, into: bus } - period_of: { over: snapshot, into: period } +relations: + gen_bus: { columns: [generator, bus], key: generator } # each generator on one bus + line_from: { columns: [line, bus], key: line } # two relations onto one dimension + line_to: { columns: [line, bus], key: line } + period_of: { columns: [snapshot, period], key: snapshot } + connection: { columns: [generator, bus] } # no key: a generator may connect to several buses ``` -| Field | | | -| ------------- | -------------------------------------------------------------------- | -------------- | -| `over` | required — the dimension whose members carry the map | | -| `into` | required — the dimension its values are labels of, other than `over` | | -| `description` | free text, never parsed | default `null` | +| Field | | | +| ------------- | ------------------------------------------------------------------------------------------------------------------------------------ | -------------- | +| `over` | required — the columns: a list of dimensions, or a mapping of column name to dimension where two columns share one ([roles](#roles)) | | +| `into` | not a field: a relation declares no direction | | +| `key` | the columns a row is identified by, one name or a list; omitted, the table is a bare relation ([below](#the-key-is-the-claim)) | default none | +| `description` | free text, never parsed | default `null` | + +Every column is over a declared dimension, and its values are checked against +that dimension's labels once data is bound — the check that makes `sum(by=)` +safe, and the reason a label set the model only ever _selects_ on is declared +as a dimension all the same: nothing is indexed by `period` above, and +`where: "period_of == 1"` ([where strings](expressions.md#where-strings)) is +how a declaration selects on it. A relation has at least two columns; a label on +one dimension is a parameter over it. A column named like a dimension is over +that dimension, so `columns: {bus: line}` is refused. + +### The key is the claim + +`key:` names the columns that are unique together. `key: generator` says the +generator column holds each label once: the table has **one row per +generator**, so the other column is a function of it. `key: [generator, period]` +says the pair holds each combination once. Neither column need be unique on its +own: a generator appears once per period, and a period once per generator. +The claim is checked at bind: a generator on two buses is refused, where a +`0`/`1` membership parameter would have said so legally and silently +([#161](https://github.com/energy-models/math-spec/issues/161)). The columns +the key determines are the relation's **value columns**. A key has one column per +dimension: it is read at its dimensions, and no frame carries a dimension +twice, so `key: [bus0, bus1]` is refused where both are over `bus`. + +Each cardinality is one declaration, and the key is the side that is one: + +| to say | write | checked at bind | +| ------------------------------------------ | ----------------------------------------------------------------------------------------------------------------- | ------------------------------------- | +| many-to-one, each generator on one bus | `{columns: [generator, bus], key: generator}` | one row per generator | +| one-to-many, a bus and its generators | the same table: `sum(p, by=gen_bus)` collects a bus's generators, `at(price, by=gen_bus)` reads a generator's bus | the same | +| many-to-many, a generator on several buses | `{columns: [generator, bus]}`, no key | nothing: a row exists, or it does not | +| one-to-one | not a claim the language has: a key is one set of columns, so the other side stays many | | + +The key is also what decides which walks the table admits: + +| the walk | needs | because | +| ------------------------------- | ---------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------- | +| `sum(x, by=l, over=a, into=b)` | the key **not** wholly inside the columns the operand fixes — the `into` columns and the columns joined on | a sum adds its rows up; walked to the key it finds one per coordinate, which is a read | +| `at(x, by=l, over=a, into=b)` | a key inside the columns the operand fixes — the `into` columns and the columns joined on | a read is one value per coordinate, or it is not a read | +| `shift`, `sum_back`, `position` | a key column over the dimension walked | a coordinate is in one group, or it has no neighbour | +| `where: "l == 'north'"` | a key, and the column compared a value column | a comparison is one value per coordinate | +| `where: l` (bare) | nothing | a row exists, or it does not | + +A bare relation — no `key:` — is walked by `sum` alone, with both ends named, +and tested by a bare `where`. That is what a many-to-many relation can say, +and all it can say. + +### A walk names its ends + +Every operator that takes `by=` walks the table between two of its columns: +`over=` the column **consumed**, `into=` the column **produced**, and every other +**key** column **joined on** — the operand carries its dimension and the +result keeps it. A value column not walked is not read: `ends` below, walked +from `line` to `bus1`, joins on nothing. A bare relation's columns are all +key, so all of them but the two walked are joined on. -The target must be a declared dimension, and it must differ from `over`. The -values are checked against it when the data binds, which is the check that makes -`sum(by=)` safe. +```yaml +dimensions: + generator: { dtype: str } + zone: { dtype: str } + period: { dtype: int } +relations: + zone_of: { columns: [generator, period, zone], key: [generator, period] } # a generator's zone, per period +parameters: + demand: { dims: [zone, period] } + price: { dims: [zone, period] } +variables: + p: { dims: [generator, period] } +constraints: + zone_balance: # p[generator, period] → [zone, period] + dims: [zone, period] + expression: sum(p, by=zone_of, over=generator, into=zone) >= demand + history: # p[generator, period] → [generator, zone]: the same table, walked from its other key column + dims: [generator, zone] + expression: sum(p, by=zone_of, over=period, into=zone) <= 100 + capped_revenue: # price[zone, period] → [generator, period]: the price of the zone this generator sat in that period + dims: [generator, period] + expression: at(price, by=zone_of, over=zone, into=generator) * p <= 1000 +``` -That check is also why a label set the model only ever _selects_ on is declared -as a dimension all the same. Nothing above is indexed by `period`; a declaration -selects on it with `where: "period_of == 1"` -([where strings](expressions.md#where-strings)). +**What the declaration decides, the call may leave unsaid.** Where the key has +one column and the key determines one column, the walk is the arrow the key +draws, and `sum(p, by=gen_bus)` and `at(price, by=gen_bus)` are complete: +`sum` consumes the key and produces the value, `at` consumes the value and +produces the key. Where a side has several candidates — two key columns, two +value columns — the call names it, and the refusal lists the candidates. +`zone_of` above has two key columns, so `sum` names `over=`, while `into=zone` +could have been left out. + +**A partition walks a key column and groups by the value columns.** +`shift(x, along=d, by=l)`, `sum_back(x, along=d, by=l)` and +`position(d, by=l)` take the one key column over `d`; the other key columns +are joined on, and the group is the value tuple. `within=` names the value columns the group is made of +where the table has several: `shift(x, along=snapshot, by=cal, within=week)` +walks within weeks of a calendar declared once over `[snapshot, day, week]`, +and a value column not named is not read. + +The rules, each decided at load with a refusal naming the rewrite: + +- **`over=` and `into=` name columns of the relation `by=` names**, one each or a + list each, and no column on both sides. `into=` is refused without a `by=`, + since a column needs the table that holds it. `over=` without one names a + dimension of the operand instead, which is `sum(p, over=period)`. + `sum(p, by=gen_bt, into=[bus, technology])` lands one table with two value + columns on the product `bus × technology` in one join; + `sum(p, by=zone_of, over=[generator, period])` consumes both key columns + at once, which is `sum(sum(p, by=zone_of, over=generator), over=period)` + said once; `at(tech_cap, by=gen_bt, over=[bus, technology])` reads a + two-column slot at each generator. +- **The operand carries every joined column's dimension, each once.** The map + is read at the key columns not walked, so there is no reading it at a + coordinate that lacks them; two joined columns over one dimension have + nothing to tell apart. +- **A produced dimension the operand already carries is joined on too.** + `sum(load * p, by=gen_bus)` with `load[snapshot, bus]` restricts each term to + the row where the generator's bus is the row's bus — a masked sum, which is + what the join says. +- **`at` reads one value.** Its key lies inside `into=` and the joined columns, + or the call is refused; a bare relation is never read by `at`. +- **`sum` adds its rows up.** So the reverse holds: a `sum` whose key lies inside + `into=` and the joined columns finds one row per coordinate and adds up + nothing, which is a read — it is refused toward `at`. `sum` walks to a value + column; `at` walks to the key. +- **A partition walks the one key column over the dimension it walks, and + groups by the value columns `within=` names** — all of them where it names + none. `within=` naming a key column is refused, and a bare relation + partitions nothing. The group may hold two columns over one dimension, a + pair of buses say: a partition lands nothing, so nothing needs the + dimension twice. +- **A `by=` list walks each relation by its declared arrow.** `by=[a, b]` is one + grouping, so no column keyword has anything to name; every relation in it + consumes the same dimension, joins on its own other columns, and no two + produce the same dimension. +- **A `where` comparison reads a value column of a keyed relation at its key.** + `zone_of == 'north'` reads the one value column; `ends.bus0 != ends.bus1` + names the columns where there are several. The frame carries the key's + dimensions, and two relations compared have keys over the same dimensions and + columns over one. A bare name — `where: gen_bus` — tests that a row exists: + at the key for a keyed relation, at every column for a bare relation. +- **Every column is over a declared dimension, every column name is distinct, + the key names columns the relation has, and does not name all of them.** + +**Every relation name joins the flat namespace**, so a relation may not shadow a +dimension. `generator`'s map onto `bus` is `gen_bus`, never a second `bus`. + +### Roles + +A list under `columns:` names each column after its dimension. Two columns over +one dimension need names of their own, and the mapping form gives them: -A partial lookup is legal. A label the map leaves out belongs to no group, so a -generator can sit on no bus and a line can have one open end. `sum(by=)` places -such a label's terms nowhere. A value that names no label of the target is an -error. +```yaml +relations: + ends: { columns: { line: line, bus0: bus, bus1: bus }, key: line } # a line's two ends, one table + rep_of: { columns: { snapshot: snapshot, rep: snapshot }, key: snapshot } # the representative snapshot +``` + +`sum(f, by=ends, over=line, into=bus1) - sum(f, by=ends, over=line, into=bus0)` +is the nodal balance through one table where two relations did it before, and +`where: "ends.bus0 != ends.bus1"` excludes a self-loop by comparing two of its +columns. + +`rep_of` relates a dimension to itself, which is how a clustered year is run +on a few typical days: every snapshot names the one that stands for it. +Nothing changes in the rules — `snapshot` is consumed and `rep` produced, both +over one dimension, so the frame is unchanged through `sum(by=)` and `at(by=)` +alike: -Several lookups may group at once. `sum(x, by=[gen_bus, gen_tech])` groups -through both maps in one reduction and lands on `bus` and `technology`. Every -lookup in the list must be `over:` the same dimension, and each must target a -different one. A member that either map leaves out belongs to no group. +```yaml +constraints: + representative: # every snapshot takes its representative's value + dims: [snapshot] + expression: p == at(p, by=rep_of) + weighted: # the snapshots a representative stands for, summed onto it + dims: [snapshot] + expression: sum(p, by=rep_of) <= 100 +``` -Every lookup name joins the flat namespace, so a lookup may not shadow a -dimension, and that includes its own target. The map from `generator` onto `bus` -is called `gen_bus`, never a second `bus`. +**A self-map is directional exactly as far as its key says.** `key: snapshot` +makes `rep` a function of `snapshot`, so the arrow runs from a snapshot to its +representative: `at` reads along it and `sum` collects against it, the inverse +of a many-to-one map being one-to-many, reachable as a grouping and never as +a function. Two steps along the arrow are two nested calls. Without a key the +same two columns are an undirected relation — a neighbour table — which `sum` +walks either way and nothing reads. Selecting the representatives themselves, +the rows where the map is the identity, is not a comparison the language has, +since a relation is never compared to a dimension; declare a `bool` parameter +for them. ### How the map is supplied -The map is a source key like any other, under the lookup's own name. It carries -two columns, each named after the dimension it holds: the `over` dimension, and -the target: +`gen_bus` is a source key like any other, carrying one column per column +declared, named after the column: ```python sources = { @@ -127,40 +291,44 @@ sources = { } ``` -A partial map is exactly the rows it has: `g3` appears in no row, so `g3` sits -on no bus. A null in the value column is refused, because a missing row already -says the same thing. A key that matches no label of `over` is an error rather -than a new member. - -Values are never inferred from the parameters that use the target. If they were, -a mistyped label would extend the label set instead of being rejected. +**A partial map is the rows it has.** `g3` is in no row, so `g3` sits on no +bus — absence is the absent row, exactly as it is for a parameter, and a null +in any column is refused for saying both at once. A keyed table holds one row +per key tuple, and a value matching no label of its column's dimension is a +typo rather than a new member. Values are never inferred from the parameters +that use a dimension: inferring would let a mistyped label extend the label +set instead of being rejected. -The map touches no table but its own, so you can add a lookup to a model the way -you add a parameter. The index of the `over` dimension may carry other columns, -but a column named after the lookup is refused rather than read. +Supplying it this way touches no table but its own, which is what a caller who +did not generate the index needs: a model can be extended with a relation the +same way it can be extended with a parameter. **A column of a dimension's index +named after a relation is refused** rather than read — an index may carry any +other extra, and this one would be a map read by accident. -## Dimension, lookup or parameter? +## Dimension, relation or parameter? Every column of data is one of the three. What decides which is what the math does with the column, not what the column holds: -| The column… | is declared as | because | -| -------------------------------------------------------------------------------------- | ------------------------------------- | --------------------------------------------------------------------------------------------- | -| is an axis: something is indexed by it, or an aggregation lands terms on it | a `dimension` | its members are the coordinate set every table over it is reindexed onto | -| has one value per member of a dimension and points at another — a generator's bus | a `lookup` into that dimension | it is a map that `sum(by=)` and `at(by=)` walk, and its values are checked against the target | -| is a label set the model only selects on or counts within — a period, a season, a zone | a `dimension`, and a `lookup` into it | the membership check is worth one line and one member list | -| scales terms — a coefficient, a bound, an offset | a `parameter` (`float` or `int`) | arithmetic is over numbers ([dtype](declarations.md#parameters)) | -| is a per-row attribute the math only selects on — a fuel, a constraint's sense | a `str` parameter | it names rows rather than scaling them, and no set is declared to check its values against | -| is a mask | a `bool` parameter | a bare name in a `where` is its own answer | +| The column… | is declared as | because | +| ------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------- | +| is an axis: something is indexed by it, or an aggregation lands terms on it | a `dimension` | its members are the coordinate set every table over it is reindexed onto | +| has one value per member of a dimension, or per tuple of several — a generator's bus, a line's two ends, a generator's zone by period | a `relation` with that `key` | it is a map every operator walks, and its values are checked against the dimensions they name | +| relates members of two dimensions many-to-many, with nothing to weigh — which buses a generator may connect to | a `relation` with no key | `sum` walks it with both ends named, and a bare `where` tests it. Nothing reads it, because there is no one value to read | +| relates members of two dimensions many-to-many, with a weight per pair — a link's efficiency to each bus, a cycle's lines | a `parameter` over both | the weight is the data, its row set is the relation, and the aggregation is `sum(w * x, over=a)` | +| is a label set the model only selects on or counts within — a period, a season, a zone | a `dimension`, and a keyed `relation` onto it | the membership check is worth one line and one member list | +| scales terms — a coefficient, a bound, an offset | a `parameter` (`float` or `int`) | arithmetic is over numbers ([dtype](declarations.md#parameters)) | +| is a per-row attribute the math only selects on — a fuel, a constraint's sense | a `str` parameter | it names rows rather than scaling them, and no set is declared to check its values against | +| is a mask | a `bool` parameter | a bare name in a `where` is its own answer | Two rules follow from the table. If `b` has one value per `a`, then `b` is a -**lookup** over `a`, and not a dimension: a `dims:` product over two +**relation** keyed by `a`, and not a dimension: a `dims` product over two dimensions that depend on each other, cut back with a mask, is the shape that -`lookups` replaces. +`relations` replaces. And everything under `dimensions:` is an axis. A dimension is never legal where a value belongs, because it is a coordinate space and not data. To use a dimension's coordinates as data, declare a parameter over it. `python -m math_spec check` advises on a declared dimension that nothing is -indexed by, nothing aggregates into and no lookup targets +indexed by, nothing aggregates into and no relation has a column over ([errors](errors.md#what-advice-warns-about)). diff --git a/docs/reference/language/errors.md b/docs/reference/language/errors.md index 8b7bf518..7354836a 100644 --- a/docs/reference/language/errors.md +++ b/docs/reference/language/errors.md @@ -40,7 +40,7 @@ contrast, prints to stderr and exits 1. | `kind` | The file has… | The advice says… | | --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------- | -| `never-an-axis` | a dimension nothing is indexed by, nothing aggregates into and no lookup targets | remove it, or keep it knowingly if its declarations are still to come | +| `never-an-axis` | a dimension nothing is indexed by, nothing aggregates into and no relation targets | remove it, or keep it knowingly if its declarations are still to come | | `unbounded` | a variable that no constraint uses, whose objective term pushes it towards a bound it does not have. `slack` with `bounds.lower: -inf` and a `+slack` term in a `minimize` objective | give it a finite bound, or the constraint that was meant to define it | A variable of the second kind runs to infinity for every dataset there is. A diff --git a/docs/reference/language/expressions.md b/docs/reference/language/expressions.md index 320bf4d4..095b2db4 100644 --- a/docs/reference/language/expressions.md +++ b/docs/reference/language/expressions.md @@ -79,7 +79,7 @@ A name is a letter or an underscore, followed by letters, digits or underscores. A declaration keyed by anything else is a load error, because no expression could write that key. -One flat namespace covers dimensions, lookups, parameters, variables, named +One flat namespace covers dimensions, relations, parameters, variables, named expressions, macros and the built-in operators. A collision is a load error that names both declarations. There is no shadowing: with shadowing, a new parameter named `snapshot` would silently change what `where: "snapshot > 0"` means. @@ -87,18 +87,18 @@ named `snapshot` would silently change what `where: "snapshot > 0"` means. Position decides which kinds of name are legal, and the kind of every name is fixed at load: -| Position | Legal kinds | -| --------------------------------------- | ------------------------------------------------------------------------------------------------------------ | -| expression (`p * cost`) | a variable, or a parameter whose values are numbers ([dtype](declarations.md#parameters)) | -| dimension argument (`over=`) | a dimension | -| lookup argument (`by=` on `sum` / `at`) | a lookup, and never a dimension | -| `where` string | a parameter, variable, dimension or lookup ([where strings](#where-strings)) | -| `bounds.lower` / `bounds.upper` | a parameter name, or a number | -| the `edge` key of `shift` | `'wrap'` in quotes, or a bare number. Never a dimension | -| `dual` argument (`dual(c)`) | a constraint. It resolves against the constraints alone ([reported](reported.md#reading-a-constraints-dual)) | +| Position | Legal kinds | +| ----------------------------------------- | ------------------------------------------------------------------------------------------------------------ | +| expression (`p * cost`) | a variable, or a parameter whose values are numbers ([dtype](declarations.md#parameters)) | +| dimension argument (`over=`, `along=`) | a dimension | +| relation argument (`by=` on `sum` / `at`) | a relation, and never a dimension. `over=` and `into=` name its columns | +| `where` string | a parameter, variable, dimension or relation ([where strings](#where-strings)) | +| `bounds.lower` / `bounds.upper` | a parameter name, or a number | +| the `edge` key of `shift` | `'wrap'` in quotes, or a bare number. Never a dimension | +| `dual` argument (`dual(c)`) | a constraint. It resolves against the constraints alone ([reported](reported.md#reading-a-constraints-dual)) | A bare word in the value of a keyword argument is a name to resolve. That is why -`wrap` is quoted: `shift(x, over=wrap, edge='wrap')` reads one way, even in a +`wrap` is quoted: `shift(x, along=wrap, edge='wrap')` reads one way, even in a model with a dimension called `wrap`. `edge` is the one keyword whose _key_ is fixed rather than naming a dimension, so a dimension called `edge` changes nothing. @@ -120,19 +120,19 @@ A parameter and a variable both declare `dims`, and every dimension argument is name-checked. So **the dimension set of every expression is known before any data binds**: -| Node | Dim set | Error | -| ------------------------------- | -------------------------------------------- | ----------------------------------------------------------------------------------------------------------------- | -| number | `{}` | | -| parameter / variable | its `dims` | | -| `-x`, `+x` | `dims(x)` | | -| `a + b`, `a * b`, `a / b` | `dims(a) ∪ dims(b)` | | -| `sum(x)` | `{}` | error if `dims(x)` is already empty | -| `sum(x, over=d)` | `dims(x) − {d}` | error if `d ∉ dims(x)` | -| `sum(x, by=l)` | `(dims(x) − {over(l)}) ∪ {into(l)}` | error if `over(l) ∉ dims(x)`, or if `into(l)` is already in `dims(x)` | -| `sum(x, by=[l, m])` | `(dims(x) − {over(l)}) ∪ {into(l), into(m)}` | the same errors, plus an error if `l` and `m` are over different dimensions, or if they target the same dimension | -| `at(x, by=l)` | `(dims(x) − {into(l)}) ∪ {over(l)}` | error if `into(l) ∉ dims(x)`, or if `over(l)` is already in `dims(x)` | -| `shift(x, over=d, offset=n)` | `dims(x)` | error if `d ∉ dims(x)` | -| `sum_back(x, over=d, within=n)` | `dims(x)` | error if `d ∉ dims(x)` | +| Node | Dim set | Error | +| -------------------------------- | ----------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| number | `{}` | | +| parameter / variable | its `dims` | | +| `-x`, `+x` | `dims(x)` | | +| `a + b`, `a * b`, `a / b` | `dims(a) ∪ dims(b)` | | +| `sum(x)` | `{}` | error if `dims(x)` is already empty | +| `sum(x, over=d)` | `dims(x) − {d}` | error if `d ∉ dims(x)` | +| `sum(x, by=l)` | `(dims(x) − from(l)) ∪ into(l)` | error if `from(l) ⊄ dims(x)`, if a joined column's dimension is not in `dims(x)`, or if `l`'s key lies inside the columns `into=` names and the joined columns — that walk is a read, which is `at`'s | +| `sum(x, by=[l, m])` | `(dims(x) − from(l)) ∪ into(l) ∪ into(m)` | the same errors, plus an error if `l` and `m` consume different dimensions, or if they produce the same one | +| `at(x, by=l)` | `(dims(x) − from(l)) ∪ into(l)` | error if `from(l) ⊄ dims(x)`, if a joined column's dimension is not, or if `l` has no key inside the columns `into=` names | +| `shift(x, along=d, offset=n)` | `dims(x)` | error if `d ∉ dims(x)` | +| `sum_back(x, along=d, window=n)` | `dims(x)` | error if `d ∉ dims(x)` | A binary operator takes the **union** of the two dimension sets, so an outer product is allowed wherever the declaration's own dimensions cover the result. @@ -165,20 +165,20 @@ POSITION ::= "position" "(" NAME [ "," "by" "=" NAME ] ")" QUOTED ::= "'" chars "'" | '"' chars '"' ``` -| Written as | Names a… | Meaning | -| -------------------------------- | -------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `name` (bare) | parameter | The value is defined here. A `bool` is its own answer. A `str` is defined wherever the table has a row. A number has to have a row and be finite, so `0.0` counts and `inf` does not | -| `name` (bare) | variable | The variable exists at this coordinate | -| `name` (bare) | lookup | The label maps somewhere. A lookup may be [partial](dimensions.md#lookups), and this selects the labels that do map | -| `name` (bare) | dimension | A load error. It would be true everywhere. Compare it against something instead | -| `name OP value` | parameter | Element-wise, and a null compares false. The right-hand side is a literal, or a bare name read as a string label | -| `name OP value` | dimension | A filter on the frame's own coordinate column | -| `name OP value` | lookup | A filter on the lookup's value, so the `over` dimension has to be in the frame. A null compares false | -| `name OP name` | two lookups | Legal only where both lookups are over the same dimension and into the same dimension. `from != to` excludes a self-loop | -| `position(name) OP i` | dimension | Where the row sits along the dimension's own order. `0` is first, and a negative number counts from the end | -| `position(name, by=lookup) OP i` | a dimension and a lookup over it | The same, counted within each group the lookup makes | -| `AND` `OR` `NOT` | — | Case-insensitive. `NOT` binds tighter than `AND`, and `AND` tighter than `OR` | -| `True` / `False` | — | Literals, folded at load wherever they stand. `True` is the same as no `where`; `False` gives a declaration with no rows. `x AND False` folds to `False`, and `NOT NOT x` to `x`. A [case `when:`](#the-rules-that-keep-the-cases-apart) is the one place a mask that folds to a literal is refused | +| Written as | Names a… | Meaning | +| ----------------------------------------- | ---------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `name` (bare) | parameter | The value is defined here. A `bool` is its own answer. A `str` is defined wherever the table has a row. A number has to have a row and be finite, so `0.0` counts and `inf` does not | +| `name` (bare) | variable | The variable exists at this coordinate | +| `name` (bare) | relation | A row exists: at the key for a keyed relation, at every column for a bare relation. A relation may be [partial](dimensions.md#relations), and this selects the labels that do map | +| `name` (bare) | dimension | A load error. It would be true everywhere. Compare it against something instead | +| `name OP value` | parameter | Element-wise, and a null compares false. The right-hand side is a literal, or a bare name read as a string label | +| `name OP value` | dimension | A filter on the frame's own coordinate column | +| `name OP value`, `name.col OP value` | relation | A filter on a value column of a keyed relation, read at its key, so the key's dimensions have to be in the frame. Name the column where the key determines several. A null compares false | +| `name OP name`, `name.a OP name.b` | two relation columns | Legal only where both relations are keyed over the same dimensions and both columns are over one dimension. `ends.bus0 != ends.bus1` excludes a self-loop | +| `position(name) OP i` | dimension | Where the row sits along the dimension's own order. `0` is first, and a negative number counts from the end | +| `position(name, by=relation[, within=c])` | a dimension and a relation keyed over it | The same, counted within each group the relation's value columns make | +| `AND` `OR` `NOT` | — | Case-insensitive. `NOT` binds tighter than `AND`, and `AND` tighter than `OR` | +| `True` / `False` | — | Literals, folded at load wherever they stand. `True` is the same as no `where`; `False` gives a declaration with no rows. `x AND False` folds to `False`, and `NOT NOT x` to `x`. A [case `when:`](#the-rules-that-keep-the-cases-apart) is the one place a mask that folds to a literal is refused | The dimensions of the mask must not exceed the frame it sits in. A bare name that is not declared is a load error. @@ -212,9 +212,11 @@ dimension does not carry compares equal to nothing, so the mask is false there rather than an error. Comparing two parameters, or two dimensions, is not in the language. Precompute a -boolean parameter instead. Two lookups are the exception, where both lookups -share both ends: over one dimension they are two columns of one table, and into -one dimension they draw from one label set. +boolean parameter instead. Two relation columns are the exception, where the two +relations are keyed over the same dimensions and the two columns are over one +dimension. Keyed alike, they are two columns of one key table, so the comparison +filters that table rather than joining two. Over one dimension they draw from one +label set, so a match is possible at all. ### `position()` @@ -244,15 +246,15 @@ row. coordinate occupies is an error when the data binds, not an empty mask, because seeding no row is the failure the clause was written to prevent. -`by=` counts inside each group that a lookup makes. That gives one seeded row per +`by=` counts inside each group that a relation makes. That gives one seeded row per period, however long each period is: ```yaml dimensions: snapshot: { dtype: int } period: { dtype: int } -lookups: - period_of: { over: snapshot, into: period } +relations: + period_of: { columns: [snapshot, period], key: snapshot } parameters: soc_initial: { dims: [period] } variables: @@ -264,8 +266,9 @@ constraints: expression: soc == at(soc_initial, by=period_of) ``` -The lookup must be over the dimension being counted. A coordinate the lookup -sends nowhere is in no group. A group shorter than the position is an error when +The relation must have a key column over the dimension being counted, and its +value columns are the groups. A coordinate the relation sends nowhere is in no +group. A group shorter than the position is an error when the data binds, for the same reason as above. ## Named expressions @@ -324,12 +327,12 @@ expressions: boundary: when: "committable and position(snapshot) == 0" expression: status_initial - otherwise: shift(status, over=snapshot, offset=1) + otherwise: shift(status, along=snapshot, offset=1) constraints: ramp_up: dims: [snapshot, generator] expression: >- - p - shift(p, over=snapshot, offset=1, edge=0) + p - shift(p, along=snapshot, offset=1, edge=0) <= ramp_limit * previous_status + start_up_limit * (1 - previous_status) ``` diff --git a/docs/reference/language/file.md b/docs/reference/language/file.md index 30b1ce16..68b3aec9 100644 --- a/docs/reference/language/file.md +++ b/docs/reference/language/file.md @@ -11,7 +11,7 @@ and `description`. Any subset of the ten is accepted. | Key | | | ------------- | ------------------------------------------------------------------------------------------------------------------- | | `dimensions` | the axes ([dimensions](dimensions.md)) | -| `lookups` | named maps out of a dimension ([lookups](dimensions.md#lookups)) | +| `relations` | named relations between dimensions ([relations](dimensions.md#relations)) | | `parameters` | the data the model expects ([declarations](declarations.md)) | | `variables` | what the solver decides | | `constraints` | the rules those decisions obey | diff --git a/docs/reference/language/index.md b/docs/reference/language/index.md index 2430fbe8..d2146542 100644 --- a/docs/reference/language/index.md +++ b/docs/reference/language/index.md @@ -49,7 +49,7 @@ message that names the fix. These ten rules are what it checks. | 1 | A file has ten declaration keys, plus `version` and `description`. A key the schema does not know is refused, with the nearest valid key named: `boundz` → `bounds`. | [File shape](file.md) | | 2 | Everything that can be checked without data is checked when the file loads. | [Errors](errors.md) | | 3 | Every name is declared once. A parameter and a dimension both called `snapshot` is refused, and the message names both lines. | [Names](expressions.md#name-resolution) | -| 4 | Where a name may stand depends on what it is. A dimension may follow `over=`, and may not be multiplied: `p * snapshot` is refused, because `snapshot` is an axis and not a column of numbers. | [Names](expressions.md#name-resolution) | +| 4 | Where a name may stand depends on what it is. A dimension may follow `over=` or `along=`, and may not be multiplied: `p * snapshot` is refused, because `snapshot` is an axis and not a column of numbers. | [Names](expressions.md#name-resolution) | | 5 | `a + b` carries the dimensions of `a` and of `b` together. A constraint's expression must carry **exactly** its `dims`. The objective must carry none. A `where` or a bound may carry fewer dimensions than its declaration, never more. | [How dimensions combine](expressions.md#how-dimensions-combine) | | 6 | A variable's `where:` deletes the variable at the masked coordinates. There is no column there, not a column fixed at zero. A constraint's `where:` deletes the row. | [Absence](absence.md) | | 7 | A deleted variable takes its row with it: `x + y >= 1` has no row where `y` is deleted. Inside a `sum` it is one term fewer, and the row stays. So `sum(x + y)` and `sum(x) + sum(y)` are different constraints. | [Absence](absence.md#how-absence-travels) | @@ -62,7 +62,7 @@ message that names the fix. These ten rules are what it checks. | | | | ----------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------- | | [File shape](file.md) | the ten keys, `version`, `description`, and how the YAML is read | -| [Dimensions and lookups](dimensions.md) | the axes, and the maps from one axis onto another | +| [Dimensions and relations](dimensions.md) | the axes, and the maps from one axis onto another | | [Parameters, variables, constraints and the objective](declarations.md) | the four blocks that carry the math | | [Expressions](expressions.md) | the arithmetic grammar and the `where` grammar, where each kind of name may stand, and how dimensions combine | | [Reported expressions](reported.md) | named quantities that no constraint or objective uses, which you read back after a solve | diff --git a/docs/reference/language/operators.md b/docs/reference/language/operators.md index 2ab05482..757aeb11 100644 --- a/docs/reference/language/operators.md +++ b/docs/reference/language/operators.md @@ -15,18 +15,21 @@ model can never depend on what a caller registered. A composition of them goes i | -------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- | | `sum(array)` | Every dimension that `array` carries collapses. The result is a scalar | | `sum(array, over=dim)` | `dim` collapses. `array` must carry `dim` | -| `sum(array, by=lookup)` | The dimension that the lookup is over collapses onto the dimension it maps into | -| `sum(array, by=[lookup, …])` | The same, onto every dimension that the lookups map into. All the lookups must be over the same dimension | -| `at(array, by=lookup)` | The dimension that the lookup maps into is replaced by the dimension it is over | -| `shift(array, over=dim, offset=n)` | The value `n` positions earlier along `dim`. The vacated edge is **absent** | -| `shift(array, over=dim, offset=n, edge='wrap')` | The value `n` positions earlier, counted cyclically, so nothing is vacated | -| `shift(array, over=dim, offset=n, edge=v)` | The value `n` positions earlier, with the number `v` standing where the edge was vacated | -| `shift(array, over=dim, offset=p, edge=…)` | `p` is an integer parameter, so each entity is reached by its own offset. Declared over what a `by=` groups into, it gives one lag per group | -| `shift(array, over=dim, offset=n, by=lookup)` | The translation walks inside each group that the lookup makes. Neighbours, edges and a wrap all belong to that group | -| `sum_back(array, over=dim, within=n)` | The sum of the last `n` positions along `dim`, ending at the position being written | -| `sum_back(array, over=dim, within=p)` | `p` is an integer parameter, so each entity gets its own window length | -| `sum_back(array, over=dim, within=p, edge='wrap')` | The window reaches around the axis, instead of stopping short at its start | -| `sum_back(array, over=dim, within=n, by=lookup)` | The window stays inside each group that the lookup makes | +| `sum(array, by=relation)` | The relation's key column collapses onto its value column | +| `sum(array, by=[relation, …])` | The same, onto every relation's value column. All the relations must consume the same dimension | +| `sum(array, by=relation, over=a, into=b)` | Column `a` collapses onto column `b`. The other key columns are joined on, so the array carries them and the result keeps them. Walked to the key, where each coordinate finds one row, it is a read — that is `at`'s | +| `sum(array, by=relation, over=[a, …], into=[b, …])` | The same with several columns on either side: consumed together, landed on a product | +| `at(array, by=relation)` | The relation's value column is replaced by its key column | +| `at(array, by=relation, over=a, into=b)` | Column `a` is replaced by column `b`, one value per coordinate, so the key lies in `b` and the joined columns. Either may be a list | +| `shift(array, along=dim, offset=n)` | The value `n` positions earlier along `dim`. The vacated edge is **absent** | +| `shift(array, along=dim, offset=n, edge='wrap')` | The value `n` positions earlier, counted cyclically, so nothing is vacated | +| `shift(array, along=dim, offset=n, edge=v)` | The value `n` positions earlier, with the number `v` standing where the edge was vacated | +| `shift(array, along=dim, offset=p, edge=…)` | `p` is an integer parameter, so each entity is reached by its own offset. Declared over what a `by=` groups into, it gives one lag per group | +| `shift(array, along=dim, offset=n, by=relation[, within=c])` | The translation walks inside each group that the relation makes. Neighbours, edges and a wrap all belong to that group | +| `sum_back(array, along=dim, window=n)` | The sum of the last `n` positions along `dim`, ending at the position being written | +| `sum_back(array, along=dim, window=p)` | `p` is an integer parameter, so each entity gets its own window length | +| `sum_back(array, along=dim, window=p, edge='wrap')` | The window reaches around the axis, instead of stopping short at its start | +| `sum_back(array, along=dim, window=n, by=relation)` | The window stays inside each group that the relation makes | `array` is any expression with the right dimension set, so each operator reads a parameter as readily as a variable. Dimension arguments are name-checked at load, @@ -40,22 +43,23 @@ so `sum(p, over=snapshto)` is an error rather than a silent no-op. `sum(x)` names no dimension and reduces every dimension `x` carries, so its result is a scalar. It is `sum(sum(x, over=a), over=b)` written once. -An operand that is already scalar, and an `over=` naming a dimension the operand -does not carry, are both errors rather than no-ops. +An operand that is already scalar, and a `over=` naming a dimension the +operand does not carry, are both errors rather than no-ops. -`sum(x, by=l)` sums along a [lookup](dimensions.md#lookups) and lands the result -on the dimension the lookup maps into. A nodal balance is one `sum(by=)` per -kind of component, and the network's wiring stays in the lookup tables: +`sum(x, by=l)` sums through a [relation](dimensions.md#relations) and lands the result +on the column it walks to: the value column, where the key draws the arrow, or +the one `into=` names. A nodal balance is one `sum(by=)` per kind of component, +and the network's wiring stays in the relations: ```yaml dimensions: bus: { dtype: str } generator: { dtype: str } line: { dtype: str } -lookups: - gen_bus: { over: generator, into: bus } - line_from: { over: line, into: bus } - line_to: { over: line, into: bus } +relations: + gen_bus: { columns: [generator, bus], key: generator } + line_from: { columns: [line, bus], key: line } + line_to: { columns: [line, bus], key: line } parameters: load: { dims: [bus] } variables: @@ -71,34 +75,45 @@ constraints: == load ``` -The same `f` is summed twice through two lookups, once as inflow and once as +The same `f` is summed twice through two relations, once as inflow and once as outflow, with no adjacency matrix and no join written by hand. -Give **at most one** of `over=` and `by=`. A lookup carries its own dimensions, -so `by=` leaves `over=` nothing to add. +`by=` and `over=` compose: `by=` names the table and `over=` names what +leaves the frame, so a call may give both, either, or neither. -The lookup's values are the group labels, checked against the target dimension +`over=` and `into=` say [which columns the walk runs between](dimensions.md#a-walk-names-its-ends) +where the declaration leaves a choice. Every other key column is joined on, so +the operand carries it, the sum keeps it, and each group is one coordinate of +it. A value column that is not walked is not read. A bare relation, one with no +`key:`, is summed with both ends named, and a row it holds twice counts twice. + +The relation's values are the group labels, checked against their own dimension when the data binds. A group with no members contributes nothing, and a member -whose lookup value is null belongs to no group. An empty group is a value rather +whose relation value is null belongs to no group. An empty group is a value rather than a gap: on the constant side of a comparison it reads as zero, where a coordinate the data never covered is refused. See [absence](absence.md). ## `at` -`at(x, by=l)` walks the same lookup table the other way. `sum(by=)` consumes the -dimension the lookup is over and produces the target. `at` consumes the target -and produces the dimension the lookup is over: it reads one coarse value once for -each fine label that points at it. +`at(x, by=l)` walks the same relation the other way. `sum(by=)` consumes the +key column and produces the value column. `at` consumes the value column and +produces the key column: it reads one coarse value once for each fine label that +points at it. `over=` and `into=` name the two columns where the key leaves a +choice. A read is one value per coordinate, so the relation's key must lie inside +`into=` and the columns joined on, and a bare relation is never read by `at`. `at` reads a variable as readily as a parameter. One decision taken per bus, read once by every line that touches the bus, is `at(decision, by=line_bus)`. -A fine label whose lookup value is null reads nothing, and its row is absent. -That matches the null group in `sum(by=)`. +A fine label whose relation value is null reads nothing, and its row is absent. +That matches the null group in `sum(by=)`. Through a relation with a +[column joined on](dimensions.md#a-walk-names-its-ends) `at` reads the coarse +value at the row's own coordinate of that column, which is the price of the zone +this generator sat in that period. ## `sum_back` -`sum_back(x, over=d, within=n)` is the sum of the last `n` positions along `d`, +`sum_back(x, along=d, window=n)` is the sum of the last `n` positions along `d`, ending at the position being written. It states a minimum up time, a rolling budget or a delivery horizon. A width of `1` is `x` itself. @@ -120,12 +135,12 @@ variables: constraints: stays_up_its_own_time: dims: [unit, hour] - expression: sum_back(started, over=hour, within=min_up) <= on + expression: sum_back(started, along=hour, window=min_up) <= on objective: { sense: minimize, expression: sum(on) } ``` -`within=` takes a number or the name of an integer parameter, and never an +`window=` takes a number or the name of an integer parameter, and never an expression. With a parameter, each entity gets a window of its own length. A fixed width can be written as a run of `shift`s; a width that is a column cannot. Two rules hold for a named width, and breaking either is a load error: @@ -142,13 +157,13 @@ a load error, because the expression can add a constant itself. `edge='wrap'` makes the window reach around the axis, which a representative period that repeats asks for. -`by=` keeps the window inside each group that a lookup makes, so no window -reaches out of its own group. The lookup obeys the rules given for +`by=` keeps the window inside each group that a relation makes, so no window +reaches out of its own group. The relation obeys the rules given for [`shift(by=)`](#a-translation-that-stops-at-each-groups-edge). ## `shift` -`shift(x, over=d, offset=n)` moves values along one dimension by `n` positions, +`shift(x, along=d, offset=n)` moves values along one dimension by `n` positions, counted in the dimension's **declared order**. The value at each coordinate becomes the value that stood `n` places before it. Only the values move, and the coordinates stay in place. `edge=` says what stands where nothing moved in. @@ -166,7 +181,7 @@ variables: constraints: storage_balance: dims: [snapshot, storage] - expression: soc == shift(soc, over=snapshot, offset=1, edge='wrap') + charge * eta - discharge + expression: soc == shift(soc, along=snapshot, offset=1, edge='wrap') + charge * eta - discharge ``` `edge='wrap'` makes a battery cyclic without a boundary condition written out: @@ -192,13 +207,13 @@ Two rules hold across all three: - **A bare `shift` over an expression with no variable is a load error.** A parameter's missing row is a zero coefficient, so there is no absence for the vacated slot to carry, and inventing one would turn - `x <= shift(dt, over=t, offset=1)` into `x <= 0`. The error names the + `x <= shift(dt, along=t, offset=1)` into `x <= 0`. The error names the rewrites: `edge='wrap'`, `edge=0`, or `edge=0` together with a `where` that excludes the vacated coordinate. A `where` on its own does not lift the refusal, and `edge=0` on its own leaves a row at that coordinate bounded by zero. -`shift` reads parameters too. `shift(dt, over=t, offset=1, edge=0)` is the +`shift` reads parameters too. `shift(dt, along=t, offset=1, edge=0)` is the previous snapshot's duration, without a pre-shifted copy of the table. ### A translation that stops at each group's edge @@ -211,8 +226,8 @@ investment period or a representative day: dimensions: snapshot: { dtype: int } season: { dtype: str } -lookups: - season_of: { over: snapshot, into: season } +relations: + season_of: { columns: [snapshot, season], key: snapshot } parameters: inflow: { dims: [snapshot] } variables: @@ -220,7 +235,7 @@ variables: constraints: season_balance: dims: [snapshot] - expression: soc == shift(soc, over=snapshot, offset=1, edge='wrap', by=season_of) + inflow + expression: soc == shift(soc, along=snapshot, offset=1, edge='wrap', by=season_of) + inflow objective: { sense: minimize, expression: sum(soc) } ``` @@ -229,11 +244,12 @@ coordinate of each group is vacated and its row drops. `edge='wrap'` closes each group onto its own last coordinate, which a store that returns to its starting level every period asks for. `edge=v` puts `v` at the edge of each group. -`by=` takes a lookup over the dimension being walked, and groups a row by the -group of that dimension it is in. The lookup's target is what a named `offset=` -may vary over, so each group is reached by its own offset. +`by=` takes a relation with a key column over the dimension being walked, and the +group is the value columns: all of them, or the ones `within=` names, so one +calendar table serves `within=day` and `within=week` alike. The group columns are +what a named `offset=` may vary over, so each group is reached by its own offset. -A coordinate the lookup sends nowhere is in no group, so it reaches nothing, and +A coordinate the relation sends nowhere is in no group, so it reaches nothing, and no `edge=` speaks for it. Its row drops under `edge=0` exactly as it does bare. Without `by=`, `edge='wrap'` wraps the whole axis: the last coordinate of the @@ -259,7 +275,7 @@ variables: constraints: arrives_after_its_lead: dims: [technology, month] - expression: shift(order, over=month, offset=lead, edge=0) >= demand + expression: shift(order, along=month, offset=lead, edge=0) >= demand objective: { sense: minimize, expression: sum(order) } ``` @@ -273,7 +289,7 @@ names its rewrite: a lag. - **The parameter varies only over dimensions where the shift can read it.** Those are the dimensions of the shifted expression, and the dimension a `by=` - lookup groups into. + relation groups into. A named offset may be bare. Its vacated positions differ per entity, and they are absent exactly as a numeric offset's are. A bare `shift` over an expression @@ -285,7 +301,7 @@ backwards says so where the data is read. ### A lag that differs per group `offset=` may name a parameter declared over the dimension that a -[`by=`](#a-translation-that-stops-at-each-groups-edge) lookup groups into. Then +[`by=`](#a-translation-that-stops-at-each-groups-edge) relation groups into. Then every snapshot of a period moves by that period's own lead time, and no coordinate reaches out of its own group: @@ -293,8 +309,8 @@ coordinate reaches out of its own group: dimensions: snapshot: { dtype: int } period: { dtype: int } -lookups: - period_of: { over: snapshot, into: period } +relations: + period_of: { columns: [snapshot, period], key: snapshot } parameters: lead: { dims: [period], dtype: int } demand: { dims: [snapshot] } @@ -305,7 +321,7 @@ variables: constraints: arrives_after_its_periods_lead: dims: [snapshot] - expression: shift(order, over=snapshot, offset=lead, by=period_of, edge=0) >= demand + expression: shift(order, along=snapshot, offset=lead, by=period_of, edge=0) >= demand objective: { sense: minimize, expression: sum(order) } ``` @@ -328,25 +344,25 @@ language prints on [Every construct, as math](../notation.md). |---|---| | `sum(array)` | $`\sum_{t \in \mathcal{T},\ g \in \mathcal{G}} p_{t,g} \le \mathrm{budget}`$ | | `sum(array, over=dim)` | $`\sum_{g \in \mathcal{G}} p_{t,g} \le \mathrm{limit}_{t} \qquad \forall\, t \in \mathcal{T}`$ | -| `sum(array, by=lookup)` | $`\sum_{g \in \mathcal{G} \,:\, \mathrm{gen\_bus}(g) = b} p_{t,g} \le \mathrm{limit}_{t,b} \qquad \forall\, t \in \mathcal{T},\ b \in \mathcal{B}`$ | -| `sum(array, by=[lookup, …])` | $`\sum_{g \in \mathcal{G} \,:\, \mathrm{gen\_bus}(g) = b \wedge \mathrm{gen\_tech}(g) = e} p_{t,g} \le \mathrm{limit}_{t,b,e} \qquad \forall\, t \in \mathcal{T},\ b \in \mathcal{B},\ e \in \mathcal{E}`$ | -| `at(array, by=lookup)` | $`p_{t} \le \mathrm{cap}_{\mathrm{period\_of}(t)} \qquad \forall\, t \in \mathcal{T}`$ | -| `shift(array, over=dim, offset=n)` | $`p_{t} \le p_{t - 1} \qquad \forall\, t \in \mathcal{T}`$ | -| `shift(array, over=dim, offset=n, edge='wrap')` | $`p_{t} \le p_{t \ominus 1} \qquad \forall\, t \in \mathcal{T}`$ | -| `shift(array, over=dim, offset=n, edge=v)` | $`p_{t} \le p_{t \boxminus_{0} 1} \qquad \forall\, t \in \mathcal{T}`$ | -| `shift(array, over=dim, offset=p, edge=…)` | $`\mathit{order}_{t,m \boxminus_{0} \mathrm{lead}} \ge \mathrm{demand}_{t,m} \qquad \forall\, t \in \mathcal{T},\ m \in \mathcal{M}`$ | -| `shift(array, over=dim, offset=n, by=lookup)` | $`p_{t} \le p_{t \ominus^{\mathrm{season\_of}(t)} 1} \qquad \forall\, t \in \mathcal{T}`$ | -| `sum_back(array, over=dim, within=n)` | $`\sum_{h' \in \mathcal{H} \,:\, 0 \le h - h' < 3} \mathit{started}_{u,h'} \le \mathit{on}_{u,h} \qquad \forall\, u \in \mathcal{U},\ h \in \mathcal{H}`$ | -| `sum_back(array, over=dim, within=p)` | $`\sum_{h' \in \mathcal{H} \,:\, 0 \le h - h' < \mathrm{min\_up}} \mathit{started}_{u,h'} \le \mathit{on}_{u,h} \qquad \forall\, u \in \mathcal{U},\ h \in \mathcal{H}`$ | -| `sum_back(array, over=dim, within=p, edge='wrap')` | $`\sum_{h' \in \mathcal{H} \,:\, 0 \le h \ominus h' < \mathrm{min\_up}} \mathit{started}_{u,h'} \le \mathit{on}_{u,h} \qquad \forall\, u \in \mathcal{U},\ h \in \mathcal{H}`$ | -| `sum_back(array, over=dim, within=n, by=lookup)` | $`\sum_{h' \in \mathcal{H} \,:\, 0 \le h -^{\mathrm{day\_of}(h)} h' < 3} \mathit{started}_{u,h'} \le \mathit{on}_{u,h} \qquad \forall\, u \in \mathcal{U},\ h \in \mathcal{H}`$ | +| `sum(array, by=relation)` | $`\sum_{g \in \mathcal{G} \,:\, \mathrm{gen\_bus}(g) = b} p_{t,g} \le \mathrm{limit}_{t,b} \qquad \forall\, t \in \mathcal{T},\ b \in \mathcal{B}`$ | +| `sum(array, by=[relation, …])` | $`\sum_{g \in \mathcal{G} \,:\, \mathrm{gen\_bus}(g) = b \wedge \mathrm{gen\_tech}(g) = e} p_{t,g} \le \mathrm{limit}_{t,b,e} \qquad \forall\, t \in \mathcal{T},\ b \in \mathcal{B},\ e \in \mathcal{E}`$ | +| `at(array, by=relation)` | $`p_{t} \le \mathrm{cap}_{\mathrm{period\_of}(t)} \qquad \forall\, t \in \mathcal{T}`$ | +| `shift(array, along=dim, offset=n)` | $`p_{t} \le p_{t - 1} \qquad \forall\, t \in \mathcal{T}`$ | +| `shift(array, along=dim, offset=n, edge='wrap')` | $`p_{t} \le p_{t \ominus 1} \qquad \forall\, t \in \mathcal{T}`$ | +| `shift(array, along=dim, offset=n, edge=v)` | $`p_{t} \le p_{t \boxminus_{0} 1} \qquad \forall\, t \in \mathcal{T}`$ | +| `shift(array, along=dim, offset=p, edge=…)` | $`\mathit{order}_{t,m \boxminus_{0} \mathrm{lead}} \ge \mathrm{demand}_{t,m} \qquad \forall\, t \in \mathcal{T},\ m \in \mathcal{M}`$ | +| `shift(array, along=dim, offset=n, by=relation)` | $`p_{t} \le p_{t \ominus^{\mathrm{season\_of}(t)} 1} \qquad \forall\, t \in \mathcal{T}`$ | +| `sum_back(array, along=dim, window=n)` | $`\sum_{h' \in \mathcal{H} \,:\, 0 \le h - h' < 3} \mathit{started}_{u,h'} \le \mathit{on}_{u,h} \qquad \forall\, u \in \mathcal{U},\ h \in \mathcal{H}`$ | +| `sum_back(array, along=dim, window=p)` | $`\sum_{h' \in \mathcal{H} \,:\, 0 \le h - h' < \mathrm{min\_up}} \mathit{started}_{u,h'} \le \mathit{on}_{u,h} \qquad \forall\, u \in \mathcal{U},\ h \in \mathcal{H}`$ | +| `sum_back(array, along=dim, window=p, edge='wrap')` | $`\sum_{h' \in \mathcal{H} \,:\, 0 \le h \ominus h' < \mathrm{min\_up}} \mathit{started}_{u,h'} \le \mathit{on}_{u,h} \qquad \forall\, u \in \mathcal{U},\ h \in \mathcal{H}`$ | +| `sum_back(array, along=dim, window=n, by=relation)` | $`\sum_{h' \in \mathcal{H} \,:\, 0 \le h -^{\mathrm{day\_of}(h)} h' < 3} \mathit{started}_{u,h'} \le \mathit{on}_{u,h} \qquad \forall\, u \in \mathcal{U},\ h \in \mathcal{H}`$ | | `dual(constraint)` | $`\mathit{price}_{t} = \lambda_{\mathrm{balance},t} \qquad \forall\, t \in \mathcal{T}`$ | $`t \ominus k`$ denotes cyclic translation: index $`t-k`$ taken modulo the size of the dimension (`roll`). Plain $`t-k`$ (`shift`) has no wraparound — terms translated past the edge are simply absent. $`t \boxminus_{v} k`$ denotes translation with $`v`$ standing where index $`t-k`$ leaves the dimension (`shift(edge=v)`), so the row at that boundary is built and carries $`v`$ rather than being dropped. -$`t \ominus^{\mathrm{lookup}(t)} k`$ denotes a translation counted inside the group a lookup puts $`t`$ in (`shift(by=lookup)`), so a term never crosses out of its own group. +$`t \ominus^{\mathrm{relation}(t)} k`$ denotes a translation counted inside the group a relation puts $`t`$ in (`shift(by=relation)`), so a term never crosses out of its own group. Regenerate with `pixi run python -m tools.spec_math`. diff --git a/docs/reference/language/piecewise.md b/docs/reference/language/piecewise.md index f6b01992..305e15e8 100644 --- a/docs/reference/language/piecewise.md +++ b/docs/reference/language/piecewise.md @@ -134,12 +134,12 @@ binds. `method` varies one thing: how the weights are restricted once they exist. -| `method` | What it adds | | -| ----------------------- | ------------------------------------------------------------------------------ | -------------------------------------------------------------- | -| `adjacency` _(default)_ | a binary per segment, and `lam <= seg + shift(seg, over=bp, offset=1, edge=0)` | the curve, built | -| `sos2` | an [`sos:`](#sos) block over the same weights | the curve, stated for a solver that branches on the set itself | -| `convex` | nothing | the hull, which is a pure linear program | -| `lp` | no weights at all: one row per segment line, plus two rows holding the domain | the curve as its own lines | +| `method` | What it adds | | +| ----------------------- | ------------------------------------------------------------------------------- | -------------------------------------------------------------- | +| `adjacency` _(default)_ | a binary per segment, and `lam <= seg + shift(seg, along=bp, offset=1, edge=0)` | the curve, built | +| `sos2` | an [`sos:`](#sos) block over the same weights | the curve, stated for a solver that branches on the set itself | +| `convex` | nothing | the hull, which is a pure linear program | +| `lp` | no weights at all: one row per segment line, plus two rows holding the domain | the curve as its own lines | `adjacency` and `sos2` state the same restriction and reach the same optimum. They differ in what the solver is handed, so which is faster is a property of the diff --git a/docs/reference/language/reading.md b/docs/reference/language/reading.md index 3b6abcc6..d546b5ff 100644 --- a/docs/reference/language/reading.md +++ b/docs/reference/language/reading.md @@ -158,7 +158,7 @@ either way, so nothing later would tell you. program.separability['bp'].windowable # False tied = program.separability['generator'].coupled["constraint 'target'"] tied.partition(' — ')[0] # 'sums over generator' -'sum_back(within=n)' in tied # True +'sum_back(window=n)' in tied # True ``` Every declared axis has an entry, and the report is walked once and held, like @@ -172,11 +172,11 @@ expansion introduced is named under the declaration the expansion emitted. state the caller seeds, and a grouping is windowed along the dimension it groups into. The report names the change and never applies it. - `undecided` lists each read whose reach only the data can say. Each entry is a - `Reach`, carrying the declaration, the parameter or lookup it reads, and the + `Reach`, carrying the declaration, the parameter or relation it reads, and the kind of read: an `offset` from a parameter, a `partition` a shift is grouped by, or a `coordinate` read through `at()`. A caller that holds the data reads the smallest value of each named parameter and hands it to `resolved`, which - returns the report with those reads decided. A reach that a lookup decides is + returns the report with those reads decided. A reach that a relation decides is not a number, so it stays undecided. - `restarts` names each declaration that counts a `position()` along the axis, because a window restarts that count at its first row. diff --git a/docs/reference/notation.md b/docs/reference/notation.md index 10d99b65..e37ca6b6 100644 --- a/docs/reference/notation.md +++ b/docs/reference/notation.md @@ -5,34 +5,37 @@ SPDX-License-Identifier: CC-BY-4.0 # Every construct, as math -Every construct the language has is printed here, beside the math the -typesetter gives it. Look up the notation for one construct, or read the whole -notation at once: two constructs that mean different things print differently, -and every symbol below appears first in the legend that defines it. - -The page is generated by `pixi run python -m tools.notation`, almost all of it -from one model, -[`tests/typesetting/golden/model.yaml`](https://github.com/energy-models/math-spec/blob/main/tests/typesetting/golden/model.yaml). -That model is not a sensible optimisation problem. It is the one file that -carries every construct at once, and tests in `tests/typesetting/` hold it to -the language: every operator a format spells, every kind of node the parsers -produce, and every line of the code that walks them. So a construct added to the -language either arrives on this page, or CI fails. The curves are the exception, -and use one real model per `method:`. - -For what an operator _does_, read [Operators](language/operators.md), which -prints the same math with one row per call. For models written to be read, start -with the [examples](../examples/index.md). - -The symbols below are **derived** from the names in the file, which is what a -model prints with no setup, so you see $\mathit{load}_{t}$ rather than $\ell_t$. -A [symbol table](typeset.md#symbol-tables) replaces every symbol, and changes -nothing else on this page. +[Typesetting](typeset.md) prints a model the way a paper prints it. This page +prints _all_ of it: every construct the language has, beside the math the +typesetter gives it, so the notation can be read as the one system it has to +be — two constructs that mean different things looking different, a symbol +introduced where it is defined and used where it is meant. + +It is generated by `pixi run python -m tools.notation`, almost all of it from one +model: +[`tests/typesetting/golden/model.yaml`](https://github.com/energy-models/math-spec/blob/main/tests/typesetting/golden/model.yaml), +which is not a sensible optimisation problem and is not trying to be: it is the +one file that carries every construct at once, and three checks in +`tests/typesetting/test_typeset.py` hold it to the language — every operator a format +spells, every node kind the parsers produce, every line of the walk. So _every_ +here is asserted rather than promised, and a construct added to the language +arrives on this page or CI goes red. The curves are the exception, one real +model per `method:`, for the reason the section gives. + +Two things this page is not. It is not the operator reference — what each +operator _does_ is [Operators](language/operators.md), which renders the same +math one row per call shape. And it is not a tutorial: the models under +`examples/` are the ones written to be read. + +The symbols are the **derived** ones, taken with no symbol table, because that +is what a model prints with no setup — $\mathit{load}_{t}$ rather than +$\ell_t$. A [symbol table](typeset.md#symbol-tables) replaces them wholesale +and changes nothing else on this page. ### The legend -A dimension, a lookup and a parameter declare no equation; what they print is the legend every model opens with. +A dimension, a relation and a parameter declare no equation; what they print is the legend every model opens with. ```yaml dimensions: @@ -43,12 +46,16 @@ dimensions: season: { dtype: str } technology: { dtype: str } -lookups: - gen_bus: { over: generator, into: bus } - gen_tech: { over: generator, into: technology } # a second map out of `generator`, to group through both at once - zone_of: { over: bus, into: zone } - area_of: { over: bus, into: zone } # a second map into the same set, to compare against - season_of: { over: snapshot, into: season } +relations: + gen_bus: { columns: [generator, bus], key: generator } + gen_tech: { columns: [generator, technology], key: generator } # a second map out of `generator`, to group through both at once + zone_of: { columns: [bus, zone], key: bus } + area_of: { columns: [bus, zone], key: bus } # a second map into the same set, to compare against + season_of: { columns: [snapshot, season], key: snapshot } + gen_zone: { columns: [generator, snapshot, zone], key: [generator, snapshot] } # a map keyed by two dimensions: a call walks one and joins on the other + rep_of: { columns: { snapshot: snapshot, rep: snapshot }, key: snapshot } # a map into its own dimension: the representative snapshot + connection: { columns: [generator, bus] } # a bare relation, no key: many-to-many, walked only by sum with both ends named + gen_bt: { columns: [generator, bus, technology], key: generator } # one table with two value columns, walked to both at once parameters: p_max: { dims: [generator] } @@ -69,12 +76,12 @@ parameters: | Symbol | Meaning | |---|---| -| $`\mathcal{T}`$ | index $`t`$ — `snapshot` (`int` coordinates) with $`\mathrm{season\_of}: \mathcal{T} \to \mathcal{S}`$ | -| $`\mathcal{G}`$ | index $`g`$ — `generator` with $`\mathrm{gen\_bus}: \mathcal{G} \to \mathcal{B},\ \mathrm{gen\_tech}: \mathcal{G} \to \mathcal{E}`$ | -| $`\mathcal{B}`$ | index $`b`$ — `bus` with $`\mathrm{zone\_of}: \mathcal{B} \to \mathcal{Z},\ \mathrm{area\_of}: \mathcal{B} \to \mathcal{Z}`$ | -| $`\mathcal{Z}`$ | index $`z`$ — `zone` | -| $`\mathcal{S}`$ | index $`s`$ — `season` | -| $`\mathcal{E}`$ | index $`e`$ — `technology` | +| $`\mathcal{T}`$ | index $`t`$ — `snapshot` (`int` coordinates) with $`\mathrm{season\_of}: \mathcal{T} \to \mathcal{S},\ \mathrm{gen\_zone}: \mathcal{G} \times \mathcal{T} \to \mathcal{Z},\ \mathrm{rep\_of}: \mathcal{T} \to \mathcal{T}`$ | +| $`\mathcal{G}`$ | index $`g`$ — `generator` with $`\mathrm{gen\_bus}: \mathcal{G} \to \mathcal{B},\ \mathrm{gen\_tech}: \mathcal{G} \to \mathcal{E},\ \mathrm{gen\_zone}: \mathcal{G} \times \mathcal{T} \to \mathcal{Z},\ \mathrm{connection} \subseteq \mathcal{G} \times \mathcal{B},\ \mathrm{gen\_bt}: \mathcal{G} \to \mathcal{B} \times \mathcal{E}`$ | +| $`\mathcal{B}`$ | index $`b`$ — `bus` with $`\mathrm{gen\_bus}: \mathcal{G} \to \mathcal{B},\ \mathrm{zone\_of}: \mathcal{B} \to \mathcal{Z},\ \mathrm{area\_of}: \mathcal{B} \to \mathcal{Z},\ \mathrm{connection} \subseteq \mathcal{G} \times \mathcal{B},\ \mathrm{gen\_bt}: \mathcal{G} \to \mathcal{B} \times \mathcal{E}`$ | +| $`\mathcal{Z}`$ | index $`z`$ — `zone` with $`\mathrm{zone\_of}: \mathcal{B} \to \mathcal{Z},\ \mathrm{area\_of}: \mathcal{B} \to \mathcal{Z},\ \mathrm{gen\_zone}: \mathcal{G} \times \mathcal{T} \to \mathcal{Z}`$ | +| $`\mathcal{S}`$ | index $`s`$ — `season` with $`\mathrm{season\_of}: \mathcal{T} \to \mathcal{S}`$ | +| $`\mathcal{E}`$ | index $`e`$ — `technology` with $`\mathrm{gen\_tech}: \mathcal{G} \to \mathcal{E},\ \mathrm{gen\_bt}: \mathcal{G} \to \mathcal{B} \times \mathcal{E}`$ | #### Parameters @@ -123,11 +130,11 @@ $`t \ominus k`$ denotes cyclic translation: index $`t-k`$ taken modulo the size $`t \boxminus_{v} k`$ denotes translation with $`v`$ standing where index $`t-k`$ leaves the dimension (`shift(edge=v)`), so the row at that boundary is built and carries $`v`$ rather than being dropped. -$`t \ominus^{\mathrm{lookup}(t)} k`$ denotes a translation counted inside the group a lookup puts $`t`$ in (`shift(by=lookup)`), so a term never crosses out of its own group. The two modifiers take different slots — the group above, the fill below — so $`t \boxminus_{v}^{\mathrm{lookup}(t)} k`$ is both at once. +$`t \ominus^{\mathrm{relation}(t)} k`$ denotes a translation counted inside the group a relation puts $`t`$ in (`shift(by=relation)`), so a term never crosses out of its own group. The two modifiers take different slots — the group above, the fill below — so $`t \boxminus_{v}^{\mathrm{relation}(t)} k`$ is both at once. $`\mathrm{pos}(t)`$ denotes where index $`t`$ sits along its dimension's own order — the order `shift` walks, not the order labels sort in — counted from $`0`$. The index itself stays the coordinate, so $`t`$ compares against labels and $`\mathrm{pos}(t)`$ against positions. -$`\mathrm{pos}_{\mathrm{lookup}(t)}(t)`$ counts within the group a lookup puts $`t`$ in: the subscript names the map, $`\mathcal{T}_{\mathrm{lookup}(t)}`$ is the group it lands in, and that group has a first position of its own. +$`\mathrm{pos}_{\mathrm{relation}(t)}(t)`$ counts within the group a relation puts $`t`$ in: the subscript names the map, $`\mathcal{T}_{\mathrm{relation}(t)}`$ is the group it lands in, and that group has a first position of its own. $`\lvert \mathcal{T} \rvert`$ denotes the size of the set being counted along, and a position counted from the end prints against it — $`\lvert \mathcal{T} \rvert - 1`$ is the last position, one less than the size because the first is $`0`$. @@ -178,7 +185,7 @@ p_{t,g} \le \mathrm{startup\_cost}_{t,g} \qquad \forall\, t \in \mathcal{T},\ g #### `balance` -sum over a lookup +sum over a relation ```yaml balance: @@ -197,7 +204,7 @@ roll (cyclic) and shift (acyclic) in one equation ```yaml ramp: dims: [snapshot, generator] - expression: p - shift(p, over=snapshot, offset=1, edge='wrap') <= shift(p, over=snapshot, offset=1) + p_max + expression: p - shift(p, along=snapshot, offset=1, edge='wrap') <= shift(p, along=snapshot, offset=1) + p_max ``` ```math @@ -212,8 +219,8 @@ the two translations `ramp` leaves out: a fill, and forwards edges: dims: [snapshot, generator] expression: >- - shift(p, over=snapshot, offset=1, edge=0) - <= shift(p, over=snapshot, offset=-1, edge=0) + p_max + shift(p, along=snapshot, offset=1, edge=0) + <= shift(p, along=snapshot, offset=-1, edge=0) + p_max ``` ```math @@ -227,7 +234,7 @@ the cyclic translation forwards, which is a fourth symbol again ```yaml ahead: dims: [snapshot, generator] - expression: p <= shift(p, over=snapshot, offset=-1, edge='wrap') + expression: p <= shift(p, along=snapshot, offset=-1, edge='wrap') ``` ```math @@ -241,7 +248,7 @@ two steps of one policy are one step; a zero step is none at all ```yaml composed: dims: [snapshot, generator] - expression: shift(shift(p, over=snapshot, offset=1), over=snapshot, offset=1) <= shift(p_max, over=generator, offset=0) + expression: shift(shift(p, along=snapshot, offset=1), along=snapshot, offset=1) <= shift(p_max, along=generator, offset=0) ``` ```math @@ -255,7 +262,7 @@ a named offset under a numbered one stays two steps, not their sum ```yaml uncomposed: dims: [snapshot, generator] - expression: shift(shift(p, over=snapshot, offset=lead, edge=0), over=snapshot, offset=1) <= p_max + expression: shift(shift(p, along=snapshot, offset=lead, edge=0), along=snapshot, offset=1) <= p_max ``` ```math @@ -269,7 +276,7 @@ two dimensions translated at one leaf, each with its own policy ```yaml crossed: dims: [snapshot, generator] - expression: shift(shift(p, over=snapshot, offset=1, edge='wrap'), over=generator, offset=-1) <= p_max + expression: shift(shift(p, along=snapshot, offset=1, edge='wrap'), along=generator, offset=-1) <= p_max ``` ```math @@ -283,7 +290,7 @@ an offset the data carries, so it prints as a symbol rather than a number ```yaml lead_time: dims: [snapshot, generator] - expression: shift(p, over=snapshot, offset=lead, edge=0) <= p_max + expression: shift(p, along=snapshot, offset=lead, edge=0) <= p_max ``` ```math @@ -292,12 +299,12 @@ p_{t \boxminus_{0} \mathrm{lead},g} \le \mathrm{p}^{\mathrm{max}}_{g} \qquad \fo #### `in_season` -a translation partitioned by a lookup: the group rides on the operator +a translation partitioned by a relation: the group rides on the operator ```yaml in_season: dims: [snapshot, generator] - expression: p <= shift(p, over=snapshot, offset=1, edge='wrap', by=season_of) + expression: p <= shift(p, along=snapshot, offset=1, edge='wrap', by=season_of) ``` ```math @@ -311,7 +318,7 @@ the same group, with a fill: each season's opening row is kept and given a zero ```yaml held_in_season: dims: [snapshot, generator] - expression: p <= shift(p, over=snapshot, offset=1, edge=0, by=season_of) + expression: p <= shift(p, along=snapshot, offset=1, edge=0, by=season_of) ``` ```math @@ -325,7 +332,7 @@ a trailing window of fixed width ```yaml window: dims: [snapshot, generator] - expression: sum_back(on, over=snapshot, within=3) <= units + expression: sum_back(on, along=snapshot, window=3) <= units ``` ```math @@ -339,7 +346,7 @@ the same window, its width in the data and its edge wrapped ```yaml history: dims: [snapshot, generator] - expression: sum_back(on, over=snapshot, within=min_up, edge='wrap') <= units + expression: sum_back(on, along=snapshot, window=min_up, edge='wrap') <= units ``` ```math @@ -348,12 +355,12 @@ history: #### `seasonal_window` -a window partitioned by a lookup: the group rides on the operator +a window partitioned by a relation: the group rides on the operator ```yaml seasonal_window: dims: [snapshot, generator] - expression: sum_back(on, over=snapshot, within=3, by=season_of) <= units + expression: sum_back(on, along=snapshot, window=3, by=season_of) <= units ``` ```math @@ -362,7 +369,7 @@ seasonal_window: #### `pullback` -at(), which re-indexes through a lookup instead of an offset +at(), which re-indexes through a relation instead of an offset ```yaml pullback: @@ -374,6 +381,92 @@ pullback: \mathit{spill}_{t} \le \mathrm{zone\_cap}_{\mathrm{zone\_of}(b)} \qquad \forall\, t \in \mathcal{T},\ b \in \mathcal{B} ``` +#### `grouped_once` + +one table walked to two value columns: the domain carries a condition per column + +```yaml +grouped_once: + dims: [snapshot, bus, technology] + expression: sum(p, by=gen_bt, into=[bus, technology]) <= tech_cap +``` + +```math +\sum_{g \in \mathcal{G} \,:\, \mathrm{gen\_bt.bus}(g) = b \wedge \mathrm{gen\_bt.technology}(g) = e} p_{t,g} \le \mathrm{tech\_cap}_{b,e} \qquad \forall\, t \in \mathcal{T},\ b \in \mathcal{B},\ e \in \mathcal{E} +``` + +#### `pulled_back_once` + +its adjoint, reading one slot through two columns of one table + +```yaml +pulled_back_once: + dims: [generator] + expression: units <= at(tech_cap, by=gen_bt, over=[bus, technology]) +``` + +```math +\mathit{units}_{g} \le \mathrm{tech\_cap}_{\mathrm{gen\_bt.bus}(g),\mathrm{gen\_bt.technology}(g)} \qquad \forall\, g \in \mathcal{G} +``` + +#### `within_bus` + +a partition grouped by one named value column of a two-value table, and a position within both + +```yaml +within_bus: + dims: [generator] + where: "position(generator, by=gen_bt, within=[bus, technology]) == 0" + expression: units <= shift(units, along=generator, offset=1, edge=0, by=gen_bt, within=bus) +``` + +```math +\mathit{units}_{g} \le \mathit{units}_{g \boxminus_{0}^{\mathrm{gen\_bt.bus}(g)} 1} \qquad \forall\, g \in \mathcal{G} \,:\, \mathrm{pos}_{\left( \mathrm{gen\_bt.bus}(g),\ \mathrm{gen\_bt.technology}(g) \right)}(g) = 0 +``` + +#### `relational` + +a sum through a bare relation: the domain is a row of the relation rather than a function's value + +```yaml +relational: + dims: [snapshot, bus] + expression: sum(p, by=connection, over=generator, into=bus) <= load +``` + +```math +\sum_{g \in \mathcal{G} \,:\, \left( g,\ b \right) \in \mathrm{connection}} p_{t,g} \le \mathrm{load}_{t,b} \qquad \forall\, t \in \mathcal{T},\ b \in \mathcal{B} +``` + +#### `connected` + +a bare relation as a where: the row of the frame has to be a member of the relation + +```yaml +connected: + dims: [snapshot, generator, bus] + where: "connection" + expression: p <= load +``` + +```math +p_{t,g} \le \mathrm{load}_{t,b} \qquad \forall\, t \in \mathcal{T},\ g \in \mathcal{G},\ b \in \mathcal{B} \,:\, \left( g,\ b \right) \in \mathrm{connection} +``` + +#### `representative` + +a map into its own dimension, walked both ways: the frame is unchanged and the index is primed + +```yaml +representative: + dims: [snapshot] + expression: sum(spill, by=rep_of) <= at(spill, by=rep_of) +``` + +```math +\sum_{t' \in \mathcal{T} \,:\, \mathrm{rep\_of}(t') = t} \mathit{spill}_{t'} \le \mathit{spill}_{\mathrm{rep\_of}(t)} \qquad \forall\, t \in \mathcal{T} +``` + #### `grouped_twice` one grouping through two maps: the domain carries both conditions @@ -402,6 +495,49 @@ pulled_back_twice: \mathit{units}_{g} \le \mathrm{tech\_cap}_{\mathrm{gen\_bus}(g),\mathrm{gen\_tech}(g)} \qquad \forall\, g \in \mathcal{G} ``` +#### `zonal` + +a grouping through a two-key map, walked along one key: the condition reads the other, and the row keeps it + +```yaml +zonal: + dims: [snapshot, zone] + expression: sum(p, by=gen_zone, over=generator) <= zone_cap +``` + +```math +\sum_{g \in \mathcal{G} \,:\, \mathrm{gen\_zone}(g,\ t) = z} p_{t,g} \le \mathrm{zone\_cap}_{z} \qquad \forall\, t \in \mathcal{T},\ z \in \mathcal{Z} +``` + +#### `zonal_history` + +the same table walked along its other key + +```yaml +zonal_history: + dims: [generator, zone] + expression: sum(p, by=gen_zone, over=snapshot) <= zone_cap +``` + +```math +\sum_{t \in \mathcal{T} \,:\, \mathrm{gen\_zone}(g,\ t) = z} p_{t,g} \le \mathrm{zone\_cap}_{z} \qquad \forall\, g \in \mathcal{G},\ z \in \mathcal{Z} +``` + +#### `zonal_pullback` + +its adjoint, reading the slot the row's own snapshot puts the generator in + +```yaml +zonal_pullback: + dims: [snapshot, generator] + where: "gen_zone == 'north' AND position(generator, by=gen_zone) == 0" + expression: p <= at(spill * zone_cap, by=gen_zone, into=generator) +``` + +```math +p_{t,g} \le \mathit{spill}_{t} \cdot \mathrm{zone\_cap}_{\mathrm{gen\_zone}(g,\ t)} \qquad \forall\, t \in \mathcal{T},\ g \in \mathcal{G} \,:\, \mathrm{gen\_zone}(g,\ t) = \text{'}\mathrm{north}\text{'} \wedge \mathrm{pos}_{\mathrm{gen\_zone}(g,\ t)}(g) = 0 +``` + #### `arithmetic` division, both unary signs, a sign beside a sign, floats with and without an exponent, bracketing @@ -494,7 +630,7 @@ last: #### `northern` -a lookup compared to a label, to another lookup, and to nothing +a relation compared to a label, to another relation, and to nothing ```yaml northern: diff --git a/examples/commitment.yaml b/examples/commitment.yaml index 21fb82e8..07e59081 100644 --- a/examples/commitment.yaml +++ b/examples/commitment.yaml @@ -44,7 +44,7 @@ expressions: boundary: when: "committable and position(snapshot) == 0" expression: status_initial - otherwise: shift(status, over=snapshot, offset=1) + otherwise: shift(status, along=snapshot, offset=1) constraints: power_balance: @@ -64,7 +64,7 @@ constraints: `ramp_limit`, a unit starting up to `start_up_limit`. dims: [snapshot, generator] expression: >- - p - shift(p, over=snapshot, offset=1, edge=0) + p - shift(p, along=snapshot, offset=1, edge=0) <= ramp_limit * previous_status + start_up_limit * (1 - previous_status) objective: diff --git a/examples/operators/at.yaml b/examples/operators/at.yaml index ff381077..fed0f5d2 100644 --- a/examples/operators/at.yaml +++ b/examples/operators/at.yaml @@ -3,15 +3,15 @@ # SPDX-License-Identifier: MIT description: >- - The adjoint of the membership reduction — `at(array, by=lookup)` reads one + The adjoint of the membership reduction — `at(array, by=relation)` reads one coarse value once per fine label pointing at it. dimensions: snapshot: { dtype: int } period: { dtype: int } -lookups: - period_of: { over: snapshot, into: period } +relations: + period_of: { columns: [snapshot, period], key: snapshot } parameters: cap: { dims: [period] } diff --git a/examples/operators/shift.yaml b/examples/operators/shift.yaml index 1851a45f..e1eefb51 100644 --- a/examples/operators/shift.yaml +++ b/examples/operators/shift.yaml @@ -17,6 +17,6 @@ variables: constraints: no_faster_than_before: dims: [snapshot] - expression: p <= shift(p, over=snapshot, offset=1) + expression: p <= shift(p, along=snapshot, offset=1) objective: { sense: minimize, expression: sum(p) } diff --git a/examples/operators/shift_by_parameter.yaml b/examples/operators/shift_by_parameter.yaml index 6913471b..796f74cd 100644 --- a/examples/operators/shift_by_parameter.yaml +++ b/examples/operators/shift_by_parameter.yaml @@ -23,6 +23,6 @@ variables: constraints: arrives_after_its_lead: dims: [technology, month] - expression: shift(order, over=month, offset=lead, edge=0) >= demand + expression: shift(order, along=month, offset=lead, edge=0) >= demand objective: { sense: minimize, expression: sum(order) } diff --git a/examples/operators/shift_edge.yaml b/examples/operators/shift_edge.yaml index d015aa2f..8eb699eb 100644 --- a/examples/operators/shift_edge.yaml +++ b/examples/operators/shift_edge.yaml @@ -17,6 +17,6 @@ variables: constraints: no_faster_than_before: dims: [snapshot] - expression: p <= shift(p, over=snapshot, offset=1, edge=0) + expression: p <= shift(p, along=snapshot, offset=1, edge=0) objective: { sense: minimize, expression: sum(p) } diff --git a/examples/operators/shift_partitioned.yaml b/examples/operators/shift_partitioned.yaml index e0c0cdaa..5a12f35a 100644 --- a/examples/operators/shift_partitioned.yaml +++ b/examples/operators/shift_partitioned.yaml @@ -10,8 +10,8 @@ dimensions: snapshot: { dtype: int } season: { dtype: str } -lookups: - season_of: { over: snapshot, into: season } +relations: + season_of: { columns: [snapshot, season], key: snapshot } variables: p: @@ -21,6 +21,6 @@ variables: constraints: no_faster_than_before_in_season: dims: [snapshot] - expression: p <= shift(p, over=snapshot, offset=1, edge='wrap', by=season_of) + expression: p <= shift(p, along=snapshot, offset=1, edge='wrap', by=season_of) objective: { sense: minimize, expression: sum(p) } diff --git a/examples/operators/shift_wrap.yaml b/examples/operators/shift_wrap.yaml index 4283ccca..10c9e7b0 100644 --- a/examples/operators/shift_wrap.yaml +++ b/examples/operators/shift_wrap.yaml @@ -17,6 +17,6 @@ variables: constraints: no_faster_than_before: dims: [snapshot] - expression: p <= shift(p, over=snapshot, offset=1, edge='wrap') + expression: p <= shift(p, along=snapshot, offset=1, edge='wrap') objective: { sense: minimize, expression: sum(p) } diff --git a/examples/operators/sum_back.yaml b/examples/operators/sum_back.yaml index 0f31e586..2be35392 100644 --- a/examples/operators/sum_back.yaml +++ b/examples/operators/sum_back.yaml @@ -24,6 +24,6 @@ variables: constraints: stays_up_its_own_time: dims: [unit, hour] - expression: sum_back(started, over=hour, within=3) <= on + expression: sum_back(started, along=hour, window=3) <= on objective: { sense: minimize, expression: sum(on) } diff --git a/examples/operators/sum_back_by_parameter.yaml b/examples/operators/sum_back_by_parameter.yaml index 05bf5923..3a0998b6 100644 --- a/examples/operators/sum_back_by_parameter.yaml +++ b/examples/operators/sum_back_by_parameter.yaml @@ -24,6 +24,6 @@ variables: constraints: stays_up_its_own_time: dims: [unit, hour] - expression: sum_back(started, over=hour, within=min_up) <= on + expression: sum_back(started, along=hour, window=min_up) <= on objective: { sense: minimize, expression: sum(on) } diff --git a/examples/operators/sum_back_partitioned.yaml b/examples/operators/sum_back_partitioned.yaml index 0f748894..650ce044 100644 --- a/examples/operators/sum_back_partitioned.yaml +++ b/examples/operators/sum_back_partitioned.yaml @@ -12,8 +12,8 @@ dimensions: hour: { dtype: int } day: { dtype: str } -lookups: - day_of: { over: hour, into: day } +relations: + day_of: { columns: [hour, day], key: hour } variables: started: @@ -26,6 +26,6 @@ variables: constraints: stays_up_inside_its_day: dims: [unit, hour] - expression: sum_back(started, over=hour, within=3, by=day_of) <= on + expression: sum_back(started, along=hour, window=3, by=day_of) <= on objective: { sense: minimize, expression: sum(on) } diff --git a/examples/operators/sum_back_wrap.yaml b/examples/operators/sum_back_wrap.yaml index 504e0f42..eb436e08 100644 --- a/examples/operators/sum_back_wrap.yaml +++ b/examples/operators/sum_back_wrap.yaml @@ -24,6 +24,6 @@ variables: constraints: stays_up_its_own_time: dims: [unit, hour] - expression: sum_back(started, over=hour, within=min_up, edge='wrap') <= on + expression: sum_back(started, along=hour, window=min_up, edge='wrap') <= on objective: { sense: minimize, expression: sum(on) } diff --git a/examples/operators/sum_by.yaml b/examples/operators/sum_by.yaml index a36b221f..2ce09cce 100644 --- a/examples/operators/sum_by.yaml +++ b/examples/operators/sum_by.yaml @@ -3,8 +3,8 @@ # SPDX-License-Identifier: MIT description: >- - The membership reduction — `sum(array, by=lookup)` lands the result on the - dimension the lookup maps into, which is what makes topology data rather than + The membership reduction — `sum(array, by=relation)` lands the result on the + column the relation is walked to, which is what makes topology data rather than structure. dimensions: @@ -12,8 +12,8 @@ dimensions: generator: { dtype: str } bus: { dtype: str } -lookups: - gen_bus: { over: generator, into: bus } +relations: + gen_bus: { columns: [generator, bus], key: generator } parameters: limit: { dims: [snapshot, bus] } diff --git a/examples/operators/sum_by_lookups.yaml b/examples/operators/sum_by_relations.yaml similarity index 67% rename from examples/operators/sum_by_lookups.yaml rename to examples/operators/sum_by_relations.yaml index f1b47f2f..56508677 100644 --- a/examples/operators/sum_by_lookups.yaml +++ b/examples/operators/sum_by_relations.yaml @@ -3,8 +3,8 @@ # SPDX-License-Identifier: MIT description: >- - Grouping through several maps at once — `sum(array, by=[lookup, …])` lands - the result on every dimension the lookups map into, which is one grouping + Grouping through several maps at once — `sum(array, by=[relation, …])` lands + the result on every dimension the relations map into, which is one grouping rather than a composition of two: the generator dimension is consumed once. dimensions: @@ -13,9 +13,9 @@ dimensions: bus: { dtype: str } technology: { dtype: str } -lookups: - gen_bus: { over: generator, into: bus } - gen_tech: { over: generator, into: technology } +relations: + gen_bus: { columns: [generator, bus], key: generator } + gen_tech: { columns: [generator, technology], key: generator } parameters: limit: { dims: [snapshot, bus, technology] } diff --git a/examples/pypsa.yaml b/examples/pypsa.yaml index 0f15cde4..26f9b5e5 100644 --- a/examples/pypsa.yaml +++ b/examples/pypsa.yaml @@ -419,46 +419,46 @@ parameters: description: one where the store is in the row's carrier-and-bus set — data prep; one outside it has no row dims: [global_constraint, store] -lookups: +relations: Generator_bus: description: the bus a generator sits on - over: generator - into: bus + columns: [generator, bus] + key: generator Link_bus0: description: the bus a link leaves - over: link - into: bus + columns: [link, bus] + key: link Link_output_link: description: the link an output port belongs to - over: link_output - into: link + columns: [link_output, link] + key: link_output Link_output_bus: description: >- the bus an output port delivers to — PyPSA's `bus1`, `bus2`, … columns. A link of three output ports is three labels here rather than a third - lookup, so the file states any number of them - over: link_output - into: bus + relation, so the file states any number of them + columns: [link_output, bus] + key: link_output Load_bus: description: the bus a load sits on - over: load - into: bus + columns: [load, bus] + key: load StorageUnit_bus: description: the bus a storage unit sits on - over: storage_unit - into: bus + columns: [storage_unit, bus] + key: storage_unit Line_bus0: description: the bus a line's flow is measured at - over: line - into: bus + columns: [line, bus] + key: line Line_bus1: description: the bus at a line's other end - over: line - into: bus + columns: [line, bus] + key: line Store_bus: description: the bus a store sits on - over: store - into: bus + columns: [store, bus] + key: store variables: Generator_p: @@ -571,7 +571,7 @@ expressions: dims: [snapshot, generator] cases: opening: { when: "position(snapshot) == 0", expression: Generator_status_initial } - otherwise: shift(Generator_status, over=snapshot, offset=1) + otherwise: shift(Generator_status, along=snapshot, offset=1) Generator_previous_p: description: >- the output a generator carries into a snapshot — nothing at the start of @@ -580,7 +580,7 @@ expressions: dims: [snapshot, generator] cases: opening: { when: "position(snapshot) == 0", expression: 0 } - otherwise: shift(Generator_p, over=snapshot, offset=1) + otherwise: shift(Generator_p, along=snapshot, offset=1) Generator_p_nom_effective: description: the build a generator's limits are taken against — the chosen one where it is extendable, the given one otherwise dims: [generator] @@ -631,11 +631,11 @@ expressions: cases: cyclic: when: StorageUnit_cyclic_state_of_charge - expression: StorageUnit_retention * shift(StorageUnit_state_of_charge, over=snapshot, offset=1, edge='wrap') + expression: StorageUnit_retention * shift(StorageUnit_state_of_charge, along=snapshot, offset=1, edge='wrap') opening: when: not StorageUnit_cyclic_state_of_charge AND position(snapshot) == 0 expression: StorageUnit_state_of_charge_initial - otherwise: StorageUnit_retention * shift(StorageUnit_state_of_charge, over=snapshot, offset=1) + otherwise: StorageUnit_retention * shift(StorageUnit_state_of_charge, along=snapshot, offset=1) Store_energy_carried_in: description: >- the energy a store opens a snapshot with — its last snapshot's less @@ -646,11 +646,11 @@ expressions: cases: cyclic: when: Store_e_cyclic - expression: Store_retention * shift(Store_e, over=snapshot, offset=1, edge='wrap') + expression: Store_retention * shift(Store_e, along=snapshot, offset=1, edge='wrap') opening: when: not Store_e_cyclic AND position(snapshot) == 0 expression: Store_e_initial - otherwise: Store_retention * shift(Store_e, over=snapshot, offset=1) + otherwise: Store_retention * shift(Store_e, along=snapshot, offset=1) Link_output_arrival: description: >- what a link delivers to an output port at a snapshot — its flow after the @@ -663,8 +663,8 @@ expressions: cases: wrapping: when: Link_output_cyclic_delay - expression: shift(at(Link_p, by=Link_output_link) * Link_efficiency, over=snapshot, offset=Link_output_delay, edge='wrap') - otherwise: shift(at(Link_p, by=Link_output_link) * Link_efficiency, over=snapshot, offset=Link_output_delay, edge=0) + expression: shift(at(Link_p, by=Link_output_link) * Link_efficiency, along=snapshot, offset=Link_output_delay, edge='wrap') + otherwise: shift(at(Link_p, by=Link_output_link) * Link_efficiency, along=snapshot, offset=Link_output_delay, edge=0) primary_energy: description: >- what a `primary_energy` row totals — weighted generator energy, less @@ -842,12 +842,12 @@ constraints: up time's, which the must-stay-up mask carries dims: [snapshot, generator] where: Generator_committable AND Generator_min_up_time > 0 AND position(snapshot) > 0 - expression: sum_back(Generator_start_up, over=snapshot, within=Generator_min_up_time) <= Generator_status + expression: sum_back(Generator_start_up, along=snapshot, window=Generator_min_up_time) <= Generator_status Generator_com_down_time: description: "`Generator-com-down-time` — a unit stopped within its own minimum down time is still off" dims: [snapshot, generator] where: Generator_committable AND Generator_min_down_time > 0 AND position(snapshot) > 0 - expression: sum_back(Generator_shut_down, over=snapshot, within=Generator_min_down_time) <= 1 - Generator_status + expression: sum_back(Generator_shut_down, along=snapshot, window=Generator_min_down_time) <= 1 - Generator_status Generator_com_status_must_stay_up: description: "`Generator-com-status-min_up_time_must_stay_up` — a unit still serving the up time it brought in stays on" dims: [snapshot, generator] @@ -1075,12 +1075,12 @@ constraints: optimize builds no row either dims: [snapshot, link] where: Link_ramp_limit_up - expression: Link_p - shift(Link_p, over=snapshot, offset=1) <= Link_ramp_limit_up * Link_p_nom_effective + expression: Link_p - shift(Link_p, along=snapshot, offset=1) <= Link_ramp_limit_up * Link_p_nom_effective Link_p_ramp_limit_down: description: "`Link-p-ramp_limit_down` — a link lowers flow no faster than its limit of the build" dims: [snapshot, link] where: Link_ramp_limit_down - expression: shift(Link_p, over=snapshot, offset=1) - Link_p <= Link_ramp_limit_down * Link_p_nom_effective + expression: shift(Link_p, along=snapshot, offset=1) - Link_p <= Link_ramp_limit_down * Link_p_nom_effective StorageUnit_ext_p_dispatch_lower: description: "`StorageUnit-ext-p_dispatch-lower` — dispatch is non-negative" dims: [snapshot, storage_unit] diff --git a/examples/pypsa_linearized_uc.yaml b/examples/pypsa_linearized_uc.yaml index 7135641c..ebf1d99e 100644 --- a/examples/pypsa_linearized_uc.yaml +++ b/examples/pypsa_linearized_uc.yaml @@ -117,30 +117,30 @@ parameters: dims: [generator] dtype: bool -lookups: +relations: Generator_bus: description: the bus a generator sits on - over: generator - into: bus + columns: [generator, bus] + key: generator Link_bus0: description: the bus a link leaves - over: link - into: bus + columns: [link, bus] + key: link Link_output_link: description: the link an output port belongs to - over: link_output - into: link + columns: [link_output, link] + key: link_output Link_output_bus: description: >- the bus an output port delivers to — PyPSA's `bus1`, `bus2`, … columns. A link of three output ports is three labels here rather than a third - lookup, so the file states any number of them - over: link_output - into: bus + relation, so the file states any number of them + columns: [link_output, bus] + key: link_output Load_bus: description: the bus a load sits on - over: load - into: bus + columns: [load, bus] + key: load variables: Generator_p: @@ -182,7 +182,7 @@ expressions: dims: [snapshot, generator] cases: opening: { when: "position(snapshot) == 0", expression: Generator_status_initial } - otherwise: shift(Generator_status, over=snapshot, offset=1) + otherwise: shift(Generator_status, along=snapshot, offset=1) Generator_previous_p: description: >- the output a generator carries into a snapshot — nothing at the start of @@ -191,7 +191,7 @@ expressions: dims: [snapshot, generator] cases: opening: { when: "position(snapshot) == 0", expression: 0 } - otherwise: shift(Generator_p, over=snapshot, offset=1) + otherwise: shift(Generator_p, along=snapshot, offset=1) constraints: Generator_fix_p_lower: @@ -250,12 +250,12 @@ constraints: up time's, which the must-stay-up mask carries dims: [snapshot, generator] where: Generator_committable AND Generator_min_up_time > 0 AND position(snapshot) > 0 - expression: sum_back(Generator_start_up, over=snapshot, within=Generator_min_up_time) <= Generator_status + expression: sum_back(Generator_start_up, along=snapshot, window=Generator_min_up_time) <= Generator_status Generator_com_down_time: description: "`Generator-com-down-time` — a unit stopped within its own minimum down time is still off" dims: [snapshot, generator] where: Generator_committable AND Generator_min_down_time > 0 AND position(snapshot) > 0 - expression: sum_back(Generator_shut_down, over=snapshot, within=Generator_min_down_time) <= 1 - Generator_status + expression: sum_back(Generator_shut_down, along=snapshot, window=Generator_min_down_time) <= 1 - Generator_status Generator_com_status_must_stay_up: description: "`Generator-com-status-min_up_time_must_stay_up` — a unit still serving the up time it brought in stays on" dims: [snapshot, generator] @@ -317,8 +317,8 @@ constraints: dims: [snapshot, generator] where: Generator_committable AND Generator_partly_tightened expression: >- - shift(Generator_p, over=snapshot, offset=1) - - Generator_ramp_limit_shut_down * Generator_p_nom * shift(Generator_status, over=snapshot, offset=1) + shift(Generator_p, along=snapshot, offset=1) + - Generator_ramp_limit_shut_down * Generator_p_nom * shift(Generator_status, along=snapshot, offset=1) - (Generator_p_max_pu * Generator_p_nom - Generator_ramp_limit_shut_down * Generator_p_nom) * (Generator_status - Generator_start_up) <= 0 Generator_com_p_current: @@ -333,9 +333,9 @@ constraints: dims: [snapshot, generator] where: Generator_committable AND Generator_partly_tightened expression: >- - Generator_p - shift(Generator_p, over=snapshot, offset=1) + Generator_p - shift(Generator_p, along=snapshot, offset=1) - (Generator_p_min_pu * Generator_p_nom + Generator_ramp_limit_up * Generator_p_nom) * Generator_status - + Generator_p_min_pu * Generator_p_nom * shift(Generator_status, over=snapshot, offset=1) + + Generator_p_min_pu * Generator_p_nom * shift(Generator_status, along=snapshot, offset=1) + (Generator_p_min_pu * Generator_p_nom + Generator_ramp_limit_up * Generator_p_nom - Generator_ramp_limit_start_up * Generator_p_nom) * Generator_start_up <= 0 Generator_com_partly_shut_down: @@ -343,8 +343,8 @@ constraints: dims: [snapshot, generator] where: Generator_committable AND Generator_partly_tightened expression: >- - shift(Generator_p, over=snapshot, offset=1) - Generator_p - - Generator_ramp_limit_shut_down * Generator_p_nom * shift(Generator_status, over=snapshot, offset=1) + shift(Generator_p, along=snapshot, offset=1) - Generator_p + - Generator_ramp_limit_shut_down * Generator_p_nom * shift(Generator_status, along=snapshot, offset=1) + (Generator_ramp_limit_shut_down * Generator_p_nom - Generator_ramp_limit_down * Generator_p_nom) * Generator_status - (Generator_p_min_pu * Generator_p_nom + Generator_ramp_limit_down * Generator_p_nom - Generator_ramp_limit_shut_down * Generator_p_nom) * Generator_start_up <= 0 diff --git a/examples/pypsa_losses.yaml b/examples/pypsa_losses.yaml index 256099a9..81e03bc6 100644 --- a/examples/pypsa_losses.yaml +++ b/examples/pypsa_losses.yaml @@ -106,38 +106,38 @@ parameters: description: demand dims: [snapshot, load] -lookups: +relations: Generator_bus: description: the bus a generator sits on - over: generator - into: bus + columns: [generator, bus] + key: generator Line_bus0: description: the bus a line's flow is measured at - over: line - into: bus + columns: [line, bus] + key: line Line_bus1: description: the bus at a line's other end - over: line - into: bus + columns: [line, bus] + key: line Link_bus0: description: the bus a link leaves - over: link - into: bus + columns: [link, bus] + key: link Link_output_link: description: the link an output port belongs to - over: link_output - into: link + columns: [link_output, link] + key: link_output Link_output_bus: description: >- the bus an output port delivers to — PyPSA's `bus1`, `bus2`, … columns. A link of three output ports is three labels here rather than a third - lookup, so the file states any number of them - over: link_output - into: bus + relation, so the file states any number of them + columns: [link_output, bus] + key: link_output Load_bus: description: the bus a load sits on - over: load - into: bus + columns: [load, bus] + key: load variables: Generator_p: diff --git a/examples/pypsa_multi_period.yaml b/examples/pypsa_multi_period.yaml index 0c283619..854b4da9 100644 --- a/examples/pypsa_multi_period.yaml +++ b/examples/pypsa_multi_period.yaml @@ -105,38 +105,38 @@ parameters: description: demand dims: [snapshot, load] -lookups: +relations: snapshot_period: description: the investment period a snapshot falls in - over: snapshot - into: period + columns: [snapshot, period] + key: snapshot Generator_carrier: description: the carrier a generator converts from - over: generator - into: carrier + columns: [generator, carrier] + key: generator Generator_bus: description: the bus a generator sits on - over: generator - into: bus + columns: [generator, bus] + key: generator Link_bus0: description: the bus a link leaves - over: link - into: bus + columns: [link, bus] + key: link Link_output_link: description: the link an output port belongs to - over: link_output - into: link + columns: [link_output, link] + key: link_output Link_output_bus: description: >- the bus an output port delivers to — PyPSA's `bus1`, `bus2`, … columns. A link of three output ports is three labels here rather than a third - lookup, so the file states any number of them - over: link_output - into: bus + relation, so the file states any number of them + columns: [link_output, bus] + key: link_output Load_bus: description: the bus a load sits on - over: load - into: bus + columns: [load, bus] + key: load variables: Generator_p: @@ -216,7 +216,7 @@ constraints: where: Carrier_max_growth expression: >- sum(Generator_p_nom_ext * Generator_first_active, by=Generator_carrier) - - shift(sum(Generator_p_nom_ext * Generator_first_active, by=Generator_carrier), over=period, offset=1, edge=0) + - shift(sum(Generator_p_nom_ext * Generator_first_active, by=Generator_carrier), along=period, offset=1, edge=0) * Carrier_max_relative_growth <= Carrier_max_growth diff --git a/examples/pypsa_quadratic.yaml b/examples/pypsa_quadratic.yaml index 0155c09a..61f9b057 100644 --- a/examples/pypsa_quadratic.yaml +++ b/examples/pypsa_quadratic.yaml @@ -74,30 +74,30 @@ parameters: description: demand dims: [snapshot, load] -lookups: +relations: Generator_bus: description: the bus a generator sits on - over: generator - into: bus + columns: [generator, bus] + key: generator Link_bus0: description: the bus a link leaves - over: link - into: bus + columns: [link, bus] + key: link Link_output_link: description: the link an output port belongs to - over: link_output - into: link + columns: [link_output, link] + key: link_output Link_output_bus: description: >- the bus an output port delivers to — PyPSA's `bus1`, `bus2`, … columns. A link of three output ports is three labels here rather than a third - lookup, so the file states any number of them - over: link_output - into: bus + relation, so the file states any number of them + columns: [link_output, bus] + key: link_output Load_bus: description: the bus a load sits on - over: load - into: bus + columns: [load, bus] + key: load variables: Generator_p: diff --git a/examples/pypsa_stochastic.yaml b/examples/pypsa_stochastic.yaml index 62f82042..38cca7e2 100644 --- a/examples/pypsa_stochastic.yaml +++ b/examples/pypsa_stochastic.yaml @@ -92,30 +92,30 @@ parameters: description: demand dims: [scenario, snapshot, load] -lookups: +relations: Generator_bus: description: the bus a generator sits on - over: generator - into: bus + columns: [generator, bus] + key: generator Link_bus0: description: the bus a link leaves - over: link - into: bus + columns: [link, bus] + key: link Link_output_link: description: the link an output port belongs to - over: link_output - into: link + columns: [link_output, link] + key: link_output Link_output_bus: description: >- the bus an output port delivers to — PyPSA's `bus1`, `bus2`, … columns. A link of three output ports is three labels here rather than a third - lookup, so the file states any number of them - over: link_output - into: bus + relation, so the file states any number of them + columns: [link_output, bus] + key: link_output Load_bus: description: the bus a load sits on - over: load - into: bus + columns: [load, bus] + key: load variables: Generator_p: diff --git a/mkdocs.yml b/mkdocs.yml index db063afc..bfd56786 100644 --- a/mkdocs.yml +++ b/mkdocs.yml @@ -43,7 +43,7 @@ nav: - reference/language/index.md - File shape: reference/language/file.md - Parameters, variables, constraints and the objective: reference/language/declarations.md - - Dimensions and lookups: reference/language/dimensions.md + - Dimensions and relations: reference/language/dimensions.md - Expressions: reference/language/expressions.md - Reported expressions: reference/language/reported.md - Operators: reference/language/operators.md diff --git a/schema/math-spec.schema.json b/schema/math-spec.schema.json index 5465841e..ae5c8378 100644 --- a/schema/math-spec.schema.json +++ b/schema/math-spec.schema.json @@ -79,7 +79,7 @@ }, "DimensionBlock": { "additionalProperties": false, - "description": "A declared dimension, and the dtype its coordinates must be.\n\nA dimension is an axis and nothing else: it declares that the axis exists\nand what its coordinates are typed as, never which coordinates there are \u2014\nthose are data, and arrive at bind time. The maps its members carry \u2014 a\ngenerator's bus, a snapshot's period \u2014 are top-level ``lookups:``\n(:class:`LookupBlock`), keyed by their own name.", + "description": "A declared dimension, and the dtype its coordinates must be.\n\nA dimension is an axis and nothing else: it declares that the axis exists\nand what its coordinates are typed as, never which coordinates there are \u2014\nthose are data, and arrive at bind time. The maps its members carry \u2014 a\ngenerator's bus, a snapshot's period \u2014 are top-level ``relations:``\n(:class:`RelationBlock`), keyed by their own name.", "properties": { "description": { "anyOf": [ @@ -216,38 +216,6 @@ "title": "ExpressionCase", "type": "object" }, - "LookupBlock": { - "additionalProperties": false, - "description": "A named single-valued map out of one dimension ``into:`` another.\n\nIts values are labels of ``into``, which is what ``sum(by=)`` and\n``at(by=)`` land terms on::\n\n lookups:\n bus_of: {over: generator, into: bus}\n\nThe map itself is data, and arrives at bind time under the lookup's name.", - "properties": { - "description": { - "anyOf": [ - { - "type": "string" - }, - { - "type": "null" - } - ], - "default": null, - "title": "Description" - }, - "into": { - "title": "Into", - "type": "string" - }, - "over": { - "title": "Over", - "type": "string" - } - }, - "required": [ - "over", - "into" - ], - "title": "LookupBlock", - "type": "object" - }, "MacroBlock": { "additionalProperties": false, "description": "A parameterised expression template, defined in the YAML itself.\n\nLanguage, not code: formals (``args`` positional, ``kwargs`` keyword)\nshadow model names inside the template, and every call site expands into\ncore AST before either backend sees the expression.", @@ -480,6 +448,67 @@ } ] }, + "RelationBlock": { + "additionalProperties": false, + "description": "A named relation between dimensions, and the key it is single-valued per.\n\n``columns:`` is the table's columns \u2014 a list of dimensions, or a mapping\nof column name to dimension where two columns share one. ``key:`` names the\ncolumns each row is identified by, and is the claim the language checks\nat bind: one row per key tuple, so the other columns are a function of\nit. Without a key the table is a bare relation::\n\n relations:\n gen_bus: {columns: [generator, bus], key: generator}\n zone_of: {columns: [generator, period, zone], key: [generator, period]}\n rep_of: {columns: {snapshot: snapshot, rep: snapshot}, key: snapshot}\n connection: {columns: [entity, bus]}\n\nAn operator walks the table in the direction the call names\n(``over=``, ``into=``), joining on the other key columns; the\ndeclaration fixes no direction. The map itself is data, and arrives at bind\ntime under the relation's name, one column per role.", + "properties": { + "columns": { + "anyOf": [ + { + "type": "string" + }, + { + "items": { + "type": "string" + }, + "type": "array" + }, + { + "additionalProperties": { + "type": "string" + }, + "type": "object" + } + ], + "title": "Columns" + }, + "description": { + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Description" + }, + "key": { + "anyOf": [ + { + "type": "string" + }, + { + "items": { + "type": "string" + }, + "type": "array" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Key" + } + }, + "required": [ + "columns" + ], + "title": "RelationBlock", + "type": "object" + }, "SosBlock": { "additionalProperties": false, "description": "A special-ordered set over one dimension of one variable.\n\nOne set per coordinate of the variable's ``dims`` minus ``over``; the\nmembers are the variable's *existing* coordinates along ``over``, in that\ndimension's declared order, and ``big_m`` is the optional cap a consumer\nthat reformulates the set puts on its linking rows.\n\n``type: 1`` admits at most one nonzero member, ``type: 2`` at most two,\nand those two consecutive. Unlike every other block this one declares no\nmath to read off ``A``: it is a *set*, carried to a consumer that has the\nconcept and reformulated for one that does not.", @@ -639,14 +668,6 @@ "title": "Expressions", "type": "object" }, - "lookups": { - "additionalProperties": { - "$ref": "#/$defs/LookupBlock" - }, - "default": {}, - "title": "Lookups", - "type": "object" - }, "macros": { "additionalProperties": { "$ref": "#/$defs/MacroBlock" @@ -682,6 +703,14 @@ "title": "Piecewise", "type": "object" }, + "relations": { + "additionalProperties": { + "$ref": "#/$defs/RelationBlock" + }, + "default": {}, + "title": "Relations", + "type": "object" + }, "sos": { "additionalProperties": { "$ref": "#/$defs/SosBlock" diff --git a/src/math_spec/_expression_parser.py b/src/math_spec/_expression_parser.py index 642a9f6a..132a5182 100644 --- a/src/math_spec/_expression_parser.py +++ b/src/math_spec/_expression_parser.py @@ -22,7 +22,7 @@ if TYPE_CHECKING: from collections.abc import Callable, Iterator, Mapping - from math_spec.program import WhereNode + from math_spec.program import Walk, WhereNode #: The relation a comparison may carry — the three an expression may be #: written with, which is what a constraint's sense is read off. @@ -114,17 +114,21 @@ def shown(self) -> str: @dataclass(frozen=True) -class LookupNode: - """A resolved reference to one or more declared lookups, legal only in a kwarg value. - - ``dimension`` is the one every lookup is over — what ``sum`` consumes and - ``at`` produces — and ``into`` the targets, one per name in the order - written; ``sum(x, by=[gen_bus, gen_tech])`` is one grouping, not two. +class RelationNode: + """A resolved ``by=`` — one or more relations, each with the walk the call takes through it. + + ``dimensions`` is the fine side every walk shares — what ``sum`` consumes + and ``at`` produces — and ``into`` the coarse dims, in the order the + names and their columns are written, which ``sum`` produces and ``at`` + consumes; ``sum(x, by=[gen_bus, gen_tech])`` is one grouping, not two. + The roles joined on are the operand's to carry, and the operator passes + them through. """ names: tuple[str, ...] - dimension: str + dimensions: tuple[str, ...] into: tuple[str, ...] + walks: tuple[Walk, ...] = () @property def shown(self) -> str: @@ -240,7 +244,7 @@ class DefinitionNode: | ParameterNode | DualNode | DimensionNode - | LookupNode + | RelationNode | EdgeNode | KeywordNode | UnaryOperatorNode @@ -272,10 +276,10 @@ def shown(names: tuple[str, ...]) -> str: # Node groups #: A resolved reference the language admits only as an operator kwarg *value*: -#: ``sum(x, over=d)``, ``sum(x, by=l)``, ``shift(..., edge='wrap')``. None of +#: ``sum(x, along=d)``, ``sum(x, by=l)``, ``shift(..., edge='wrap')``. None of #: the three is data, so none may stand in arithmetic — which is why the passes #: that walk a value position refuse them together. -KwargNode = DimensionNode | LookupNode | EdgeNode +KwargNode = DimensionNode | RelationNode | EdgeNode #: What resolution rewrites away: a bare name, whose kind only the schema #: knows, and the two kwarg-only literals its kwarg consumes. Meeting one diff --git a/src/math_spec/_where_parser.py b/src/math_spec/_where_parser.py index ea873445..c0819cad 100644 --- a/src/math_spec/_where_parser.py +++ b/src/math_spec/_where_parser.py @@ -51,12 +51,13 @@ class UnresolvedComparisonNode: @dataclass(frozen=True) class UnresolvedPositionNode: - """``position(dim[, by=lookup]) i`` before the names are checked; ``resolution.py`` types it.""" + """``position(dim[, by=relation[, within=columns]]) i`` before the names are checked; ``resolution.py`` types it.""" dimension: str op: PredicateOperator position: int by: str | None = None + into: tuple[str, ...] | None = None #: What resolution rewrites away on the where side — the three nodes whose @@ -76,10 +77,11 @@ class _Quoted(str): def _position_comparison(tokens: pp.ParseResults) -> UnresolvedPositionNode: - """``position(dim[, by=lookup]) i`` off the tokens the grammar captured.""" - *call, op, at = tokens - dimension, by = call[0], call[1] if len(call) > 1 else None - return UnresolvedPositionNode(str(dimension), op, at, None if by is None else str(by)) + """``position(dim[, by=relation[, within=columns]]) i`` off the tokens the grammar captured.""" + dimension, *call, op, at = tokens + by = str(call[0]) if call else None + into = tuple(str(token) for token in call[1]) if len(call) > 1 else None + return UnresolvedPositionNode(str(dimension), op, at, by, into) def _comparison(tokens: pp.ParseResults) -> UnresolvedComparisonNode: @@ -112,7 +114,12 @@ def _build_where_grammar() -> pp.ParserElement: lambda t: _Quoted(t[0]) ) - grouped_by = pp.Suppress(',') + pp.Suppress(pp.Keyword('by')) + pp.Suppress('=') + name + column = pp.Regex(rf'{NAME}(\.{NAME})?') + columns = name | (pp.Suppress('[') + pp.DelimitedList(name) + pp.Suppress(']')) + grouped_within = pp.Group(pp.Suppress(',') + pp.Suppress(pp.Keyword('within')) + pp.Suppress('=') + columns) + grouped_by = ( + pp.Suppress(',') + pp.Suppress(pp.Keyword('by')) + pp.Suppress('=') + name + pp.Optional(grouped_within) + ) comparator = pp.one_of(list(get_args(PredicateOperator))) position_call = ( @@ -120,7 +127,7 @@ def _build_where_grammar() -> pp.ParserElement: ) position_comparison = (position_call + comparator + position).set_parse_action(_position_comparison) - comparison = (name + comparator + (number | quoted | name)).set_parse_action(_comparison) + comparison = (column + comparator + (number | quoted | column)).set_parse_action(_comparison) # pyrefly: ignore[implicit-any-lambda] existence = name.copy().set_parse_action(lambda t: UnresolvedNameNode(t[0])) @@ -187,7 +194,7 @@ def _named_rewrite(text: str, loc: int) -> str | None: #: is a test the file could carry as data instead, which is the language's own #: answer before the general one. _DEEP_REWRITE = ( - 'Declare a parameter or lookup carrying part of the test and name that here, or split the ' + 'Declare a parameter or relation carrying part of the test and name that here, or split the ' 'declaration into two, each masked by one half.' ) diff --git a/src/math_spec/_yaml.py b/src/math_spec/_yaml.py index c1274916..696175b6 100644 --- a/src/math_spec/_yaml.py +++ b/src/math_spec/_yaml.py @@ -8,7 +8,7 @@ language whose scalars are user data; both are fixed here: - **1.2 booleans.** ``on``/``off``/``yes``/``no``/``y``/``n`` are ordinary - names in this language — a country code as a dimension, a mode as a lookup. + names in this language — a country code as a dimension, a mode as a relation. YAML 1.1 resolves them to ``True``/``False``; only ``true``/``false`` are booleans here, which is the YAML 1.2 core schema. - **Duplicate keys.** 1.1 lets the last one win silently, discarding a diff --git a/src/math_spec/advice.py b/src/math_spec/advice.py index 9e90b5a6..910478db 100644 --- a/src/math_spec/advice.py +++ b/src/math_spec/advice.py @@ -43,22 +43,22 @@ def advice(model: str | Path | dict[str, Any] | Spec | Program) -> tuple[Advice, def _never_an_axis(program: Program) -> list[Advice]: """One piece of advice per dimension nothing reaches. - A dimension a lookup targets is reached: its members are the labels the - map's values are checked against, and a ``where`` selects on them, so it - is in use even where nothing is indexed by it. + A dimension a relation has a column over is reached: its members are the + labels that column is checked against, and a ``where`` selects on them, + so it is in use even where nothing is indexed by it. """ reached: set[str] = set() for declaration in (*program.parameters.values(), *program.variables.values(), *program.constraints.values()): reached.update(declaration.dims) reached |= _produced_axes(program) - reached |= {lk.target for _, lk in program.lookups} + reached |= {dim for lk in program.relations.values() for dim in lk.dims} return [ Advice( 'never-an-axis', name, f"dimension '{name}' is never used: nothing is indexed by it, nothing " - f'aggregates into it, and no lookup targets it. Remove it — or keep it ' + f'aggregates into it, and no relation has a column over it. Remove it — or keep it ' f'knowingly, if the declarations that use it are still to be written.', ) for name in program.dimensions @@ -76,5 +76,5 @@ def _produced_axes(program: Program) -> set[str]: if isinstance(node, GroupSum): axes.update(node.into) elif isinstance(node, At): - axes.add(node.over) + axes.update(node.over) return axes diff --git a/src/math_spec/dimensions.py b/src/math_spec/dimensions.py index 3c07aed5..dbb5e8a8 100644 --- a/src/math_spec/dimensions.py +++ b/src/math_spec/dimensions.py @@ -27,10 +27,10 @@ EdgeNode, FunctionCallNode, KwargNode, - LookupNode, NumberNode, ParameterNode, ParsedNode, + RelationNode, UnaryOperatorNode, UnresolvedNode, VariableNode, @@ -42,12 +42,12 @@ from math_spec.program import ( DimensionComparisonNode, DimensionPositionNode, - LookupComparisonNode, - LookupDefinedNode, - LookupPairComparisonNode, Mask, ParameterComparisonNode, ParameterDefinedNode, + RelationComparisonNode, + RelationDefinedNode, + RelationPairComparisonNode, VariableDefinedNode, ) @@ -134,7 +134,7 @@ def _dims_call(node: FunctionCallNode, schema: Spec, context: str) -> frozenset[ def _sum_dims(node: FunctionCallNode, inner: frozenset[str], schema: Spec, context: str) -> frozenset[str]: - """``sum`` reduces a dim away, or through a lookup into the dim it maps *into*.""" + """``sum`` reduces a dim away, or walks relations: the consumed dim goes, the produced dims arrive, the joined stay.""" by = node.kwargs.get('by') if by is None and 'over' not in node.kwargs: if not inner: @@ -145,38 +145,32 @@ def _sum_dims(node: FunctionCallNode, inner: frozenset[str], schema: Spec, conte ) return frozenset() if by is None: - over = node.kwargs['over'] - assert isinstance(over, DimensionNode) - if over.name not in inner: - raise DimensionError(_not_carried(context, f'sum(over={over.name})', inner, 'drop the sum, or fix the dim')) - return inner - {over.name} - - assert isinstance(by, LookupNode) - if by.dimension not in inner: + consumed = node.kwargs['over'] + assert isinstance(consumed, DimensionNode) + if consumed.name not in inner: + raise DimensionError( + _not_carried(context, f'sum(over={consumed.name})', inner, 'drop the sum, or fix the dim') + ) + return inner - {consumed.name} + + assert isinstance(by, RelationNode) + if missing := sorted(set(by.dimensions) - inner): raise DimensionError( _not_carried( context, - f"sum(by={by.shown}) consumes '{by.dimension}', the dim it maps out of,", + f'sum(by={by.shown}) consumes {missing}, the dims it walks from,', inner, 'drop the sum, or fix the dim', ) ) - collides = sorted(set(by.into) & (inner - {by.dimension})) - if collides: - raise DimensionError( - f'{context}: sum(by={by.shown}) targets {collides}, ' - f'which the expression already carries ({sorted(inner)}). The result would ' - f"need {collides} twice — once as the operand's own dim and once as the " - f'group it is placed into. Sum over one of the two first, ' - f'or group into a dimension the operand does not have.' - ) - return (inner - {by.dimension}) | set(by.into) + _check_joined(f'sum(by={by.shown})', by, inner, context) + return (inner - set(by.dimensions)) | set(by.into) def _at_dims(node: FunctionCallNode, inner: frozenset[str], schema: Spec, context: str) -> frozenset[str]: - """``at`` is the adjoint of ``sum(by=)``: it consumes the dim a lookup maps *into* and produces the one it is over.""" + """``at`` is the adjoint of ``sum(by=)``: it consumes the dims the walks produce and produces the one they consume.""" by = node.kwargs['by'] - assert isinstance(by, LookupNode) + assert isinstance(by, RelationNode) absent = sorted(set(by.into) - inner) if absent: raise DimensionError( @@ -185,25 +179,19 @@ def _at_dims(node: FunctionCallNode, inner: frozenset[str], schema: Spec, contex f'{sorted(inner)}). A pullback needs the coarse dims to read *from* — ' f'sum is the direction that produces them.' ) - if by.dimension in inner - set(by.into): - raise DimensionError( - f'{context}: at(by={by.shown}) places terms onto ' - f"'{by.dimension}', which the expression already carries ({sorted(inner)}). " - f"The result would need '{by.dimension}' twice — once as the operand's own " - f'dim and once as the dim it is spread onto. Sum over one of the two first.' - ) - return (inner - set(by.into)) | {by.dimension} + _check_joined(f'at(by={by.shown})', by, inner, context) + return (inner - set(by.into)) | set(by.dimensions) def _translation_dims(node: FunctionCallNode, inner: frozenset[str], schema: Spec, context: str) -> frozenset[str]: """``shift`` and ``sum_back`` keep every dim, and their amount, edge and partition are checked here.""" - over = node.kwargs['over'] + over = node.kwargs['along'] assert isinstance(over, DimensionNode) if over.name not in inner: raise DimensionError( _not_carried( context, - f'{node.name}(over={over.name})', + f'{node.name}(along={over.name})', inner, f'walk a dim the operand carries, or drop the {node.name}', ) @@ -213,22 +201,36 @@ def _translation_dims(node: FunctionCallNode, inner: frozenset[str], schema: Spe _check_edge(node, context) partition = node.kwargs.get('by') if partition is not None: - assert isinstance(partition, LookupNode) + assert isinstance(partition, RelationNode) if len(partition.names) > 1: raise DimensionError( - f'{context}: {node.name}(over={over.name}, by={partition.shown}) partitions by ' - f'several lookups at once. A partition says which rows are neighbours rather than ' - f'which group a term lands in, so it names one lookup — partition by a lookup whose ' + f'{context}: {node.name}(along={over.name}, by={partition.shown}) partitions by ' + f'several relations at once. A partition says which rows are neighbours rather than ' + f'which group a term lands in, so it names one relation — partition by a relation whose ' f'values already distinguish them.' ) - if partition.dimension != over.name: + _check_joined(f'{node.name}(along={over.name}, by={partition.shown})', partition, inner, context) + return inner + + +def _check_joined(call: str, by: RelationNode, inner: frozenset[str], context: str) -> None: + """The columns a walk joins on are read at their dimensions, so the operand carries every one, each once.""" + for walk in by.walks: + dims = walk.joined_dims + if missing := sorted(set(dims) - inner): raise DimensionError( - f'{context}: {node.name}(over={over.name}, by={partition.shown}) walks ' - f"'{over.name}' but groups by a lookup over '{partition.dimension}'. No row of " - f"'{over.name}' carries it, so no coordinate has a neighbour inside a group — " - f"partition by a lookup over '{over.name}'." + f'{context}: {call} joins on {missing} (columns {[r for r in walk.joined if walk.dim(r) in missing]} ' + f"of '{walk.name}'), which the expression does not carry (dims {sorted(inner)}). A relation is " + f'walked between two of its columns and read at the others — index the operand by them, or ' + f'walk between different columns.' + ) + twice = sorted({d for d in dims if dims.count(d) > 1 or d in by.dimensions}) + if twice: + raise DimensionError( + f"{context}: {call} joins '{walk.name}' on {twice} through more than one column, and the operand " + f'carries each dimension once. Give the walk different columns, or a relation whose joined ' + f'columns are over distinct dimensions.' ) - return inner #: The dim rule of each built-in, by name. @@ -289,7 +291,7 @@ def _whole(node: ArithmeticNode, minimum: float) -> bool: def _check_amount_form(node: FunctionCallNode, context: str) -> None: - """An ``offset=`` or ``within=`` is a whole number in the operator's range, or a parameter name.""" + """An ``offset=`` or ``window=`` is a whole number in the operator's range, or a parameter name.""" kwarg, amount = _amount_of(node) if isinstance(amount, ParameterNode) or _whole(amount, _AMOUNTS[node.name].minimum): return @@ -373,8 +375,8 @@ def _shift_over_data_message(context: str) -> str: return ( f'{context}: shift() over a variable-free expression leaves vacated positions with no ' f'value, and inventing one is what silently pinned a bound to zero. Say which you mean:\n' - f" shift(x, over=d, offset=n, edge='wrap') the dimension really is cyclic\n" - f' shift(x, over=d, offset=n, edge=0) the vacated positions contribute zero\n' + f" shift(x, along=d, offset=n, edge='wrap') the dimension really is cyclic\n" + f' shift(x, along=d, offset=n, edge=0) the vacated positions contribute zero\n' f' ...and a where: excluding them the vacated rows should not exist at all\n' f'A where: alone does not lift this — it is decided on the expression, before any mask ' f'is read — and edge=0 alone leaves a row whose bound is that zero.' @@ -382,7 +384,7 @@ def _shift_over_data_message(context: str) -> str: def _check_named_amount(node: FunctionCallNode, over: str, inner: frozenset[str], schema: Spec, context: str) -> None: - """The rules that hold of an ``offset=`` or ``within=`` naming a parameter; a literal breaks none of them.""" + """The rules that hold of an ``offset=`` or ``window=`` naming a parameter; a literal breaks none of them.""" kwarg, amount = _amount_of(node) words = _AMOUNTS[node.name] if isinstance(amount, UnaryOperatorNode) and isinstance(amount.operand, ParameterNode): @@ -407,13 +409,17 @@ def _check_named_amount(node: FunctionCallNode, over: str, inner: frozenset[str] f"— declare '{amount.name}' over dims '{over}' is not one of." ) partition = node.kwargs.get('by') - groups = frozenset(partition.into) if isinstance(partition, LookupNode) else frozenset() + groups = ( + frozenset(partition.walks[0].dim(v) for v in partition.walks[0].produced) + if isinstance(partition, RelationNode) + else frozenset() + ) if stray := sorted(frozenset(declared.dims) - inner - groups): raise DimensionError( f'{context}: {node.name}({kwarg}={amount.name}) reads its {words.noun} at the coordinate it ' f"walks, but '{amount.name}' varies over {stray}, which that coordinate does not carry " f'(dims {sorted(inner)}). A dim the coordinate does not have is no coordinate at all — ' - f"declare '{amount.name}' over dims the expression carries, or group by a lookup into " + f"declare '{amount.name}' over dims the expression carries, or group by a relation into " f'one of {stray}, so that each group is reached by its own {words.noun}.' ) @@ -520,8 +526,8 @@ def _check_where_dims( noun = 'variable' case DimensionComparisonNode() | DimensionPositionNode(): noun = 'dimension' - case LookupComparisonNode() | LookupPairComparisonNode() | LookupDefinedNode(): - noun = 'lookup' + case RelationComparisonNode() | RelationPairComparisonNode() | RelationDefinedNode(): + noun = 'relation' case _: assert_never(atom) raise DimensionError( diff --git a/src/math_spec/exclusivity.py b/src/math_spec/exclusivity.py index 450aad92..75ddcb45 100644 --- a/src/math_spec/exclusivity.py +++ b/src/math_spec/exclusivity.py @@ -26,14 +26,14 @@ BooleanLiteralNode, DimensionComparisonNode, DimensionPositionNode, - LookupComparisonNode, - LookupDefinedNode, - LookupPairComparisonNode, Mask, NotNode, OrNode, ParameterComparisonNode, ParameterDefinedNode, + RelationComparisonNode, + RelationDefinedNode, + RelationPairComparisonNode, TypedPredicateNode, VariableDefinedNode, ) @@ -133,18 +133,20 @@ class Subject: ``kind`` separates the namespaces that could otherwise collide: a dimension's coordinates and its *rank* are two subjects over one name, and - a rank is further split by the ``by=`` lookup it is counted within. + a rank is further split by the ``by=`` relation it is counted within. """ - kind: Literal['param', 'dim', 'rank', 'lookup', 'lookup_pair', 'variable'] + kind: Literal['param', 'dim', 'rank', 'relation', 'relation_pair', 'variable'] name: str qualifier: str | None = None + #: A rank's group columns: two positions by one relation into different columns are two subjects. + group: tuple[str, ...] = () def __str__(self) -> str: if self.kind == 'rank': within = f' within {self.qualifier}' if self.qualifier else '' return f'the position of {self.name}{within}' - if self.kind == 'lookup_pair': + if self.kind == 'relation_pair': return f'{self.name} vs {self.qualifier}' return self.name @@ -191,15 +193,15 @@ def _observe(node: TypedPredicateNode, subject: Subject, values: set[Any], dtype """ if isinstance(node, DimensionPositionNode): values.add(node.position) - elif isinstance(node, LookupPairComparisonNode): + elif isinstance(node, RelationPairComparisonNode): if node.op not in ('==', '!='): msg = ( - f'{subject} is ordered with {node.op!r}, and two lookups carry no order ' + f'{subject} is ordered with {node.op!r}, and two relations carry no order ' f'against each other — compare them with == or !=, or precompute the ' f'ordering as a boolean parameter and test that' ) raise Undecidable(msg) - elif isinstance(node, ParameterComparisonNode | DimensionComparisonNode | LookupComparisonNode): + elif isinstance(node, ParameterComparisonNode | DimensionComparisonNode | RelationComparisonNode): if node.op not in ('==', '!=') and dtypes.get(subject.name) not in _ORDERED_DTYPES: msg = ( f'{subject} has dtype {dtypes.get(subject.name)!r} and is ordered with ' @@ -218,12 +220,14 @@ def _subject_of(node: TypedPredicateNode) -> Subject: return Subject('variable', name) case DimensionComparisonNode(name=name): return Subject('dim', name) - case DimensionPositionNode(name=name, by=by): - return Subject('rank', name, by) - case LookupDefinedNode(name=name) | LookupComparisonNode(name=name): - return Subject('lookup', name) - case LookupPairComparisonNode(name=name, other=other): - return Subject('lookup_pair', name, other) + case DimensionPositionNode(name=name, partition=partition): + if partition is None: + return Subject('rank', name) + return Subject('rank', name, partition.name, partition.produced) + case RelationDefinedNode(name=name) | RelationComparisonNode(name=name): + return Subject('relation', name) + case RelationPairComparisonNode(name=name, other=other): + return Subject('relation_pair', name, other) case _: assert_never(node) @@ -236,7 +240,7 @@ def _cells_for(subject: Subject, values: set[Any], dtypes: Mapping[str, Declared """ if subject.kind == 'rank': return _rank_cells(subject, cast('set[int]', values)) - if subject.kind in ('lookup_pair', 'variable'): + if subject.kind in ('relation_pair', 'variable'): return [True, False] dtype = dtypes.get(subject.name) if dtype == 'bool': @@ -362,7 +366,7 @@ def _rank_cells(subject: Subject, positions_seen: set[int]) -> list[Cell]: def _shown(subject: Subject, value: Cell) -> str: if subject.kind == 'rank': return str(value) - if subject.kind == 'lookup_pair': + if subject.kind == 'relation_pair': return 'equal' if value else 'different' if isinstance(value, Special): return {Special.NULL: 'absent', Special.OTHER: 'anything else'}.get(value, value.value) @@ -397,17 +401,17 @@ def _atom(node: TypedPredicateNode, cell: dict[Subject, Cell], grid: _Grid) -> b subject = grid.subjects[id(node)] value = cell[subject] match node: - case ParameterDefinedNode() | LookupDefinedNode(): + case ParameterDefinedNode() | RelationDefinedNode(): if isinstance(value, bool): return value return value not in (Special.NULL, Special.POS_INF, Special.NEG_INF) case VariableDefinedNode(): return bool(value) - case LookupPairComparisonNode(op=op): + case RelationPairComparisonNode(op=op): return bool(value) if op == '==' else not value case DimensionPositionNode(op=op, position=position): return _compare(value, op, position) - case ParameterComparisonNode(op=op, value=literal) | LookupComparisonNode(op=op, value=literal): + case ParameterComparisonNode(op=op, value=literal) | RelationComparisonNode(op=op, value=literal): if value is Special.NULL: return False return _compare(value, op, literal) diff --git a/src/math_spec/lowering.py b/src/math_spec/lowering.py index d1df7af4..a78fab20 100644 --- a/src/math_spec/lowering.py +++ b/src/math_spec/lowering.py @@ -26,9 +26,9 @@ EdgeNode, FunctionCallNode, KwargNode, - LookupNode, NumberNode, ParameterNode, + RelationNode, UnaryOperatorNode, UnresolvedNode, VariableNode, @@ -143,7 +143,9 @@ def lower_program(expanded: _ExpandedSpec) -> program.Program: dimensions = { dname: program.DimensionDeclaration( tuple( - program.LookupDeclaration(lname, lk.into) for lname, lk in expanded.lookups.items() if lk.over == dname + program.RelationDeclaration(lname, lk.pairs, lk.keys) + for lname, lk in expanded.relations.items() + if dname in lk.dims ), ddef.dtype, ) @@ -257,7 +259,7 @@ def _cases(self, node: CasesNode) -> program.Cases: return program.Cases(tuple(regions)) def sum(self, node: FunctionCallNode) -> program.ExpressionNode: - """``sum(x)``, ``sum(x, over=d)`` or ``sum(x, by=lookup)``. + """``sum(x)``, ``sum(x, over=d)`` or ``sum(x, by=relation)``. Two program nodes under one surface verb: reducing a dim away and reducing it *into* another are different relational shapes, so ``by=`` decides which @@ -268,55 +270,50 @@ def sum(self, node: FunctionCallNode) -> program.ExpressionNode: if by_node is None and 'over' not in node.kwargs: return program.Sum(operand, tuple(sorted(dims_of(node.args[0], self.schema, self.context)))) if by_node is None: - over_node = node.kwargs['over'] - assert isinstance(over_node, DimensionNode), 'resolution refuses an over= that is not a dimension' - return program.Sum(operand, (over_node.name,)) - assert isinstance(by_node, LookupNode), 'resolution refuses a by= that is not a lookup' - return program.GroupSum(operand, over=by_node.dimension, coordinate=by_node.names, into=by_node.into) + consumed = node.kwargs['over'] + assert isinstance(consumed, DimensionNode), 'resolution refuses a over= that is not a dimension' + return program.Sum(operand, (consumed.name,)) + assert isinstance(by_node, RelationNode), 'resolution refuses a by= that is not a relation' + return program.GroupSum(operand, walks=by_node.walks) def at(self, node: FunctionCallNode) -> program.ExpressionNode: - """``at(x, by=lookup)`` — the adjoint of :meth:`sum`'s ``by=`` form.""" + """``at(x, by=relation)`` — the adjoint of :meth:`sum`'s ``by=`` form.""" by_node = node.kwargs['by'] - assert isinstance(by_node, LookupNode), 'resolution refuses a by= that is not a lookup' - return program.At( - self.expr(node.args[0]), - over=by_node.dimension, - coordinate=by_node.names, - into=by_node.into, - ) + assert isinstance(by_node, RelationNode), 'resolution refuses a by= that is not a relation' + return program.At(self.expr(node.args[0]), walks=by_node.walks) def sum_back(self, node: FunctionCallNode) -> program.ExpressionNode: - """``sum_back(x, over=d, within=w)`` — a trailing window along one dimension. + """``sum_back(x, along=d, window=w)`` — a trailing window along one dimension. - *within* is an integer literal of at least one, or a parameter naming a + *window* is an integer literal of at least one, or a parameter naming a per-entity width, which the language holds to the two rules that make it mean one thing before this is reached. - ``by=`` names the lookup the window stops at the edges of, and rides on + ``by=`` names the relation the window stops at the edges of, and rides on the node the way it rides on a translation — the dim rules have already - held it to one lookup over the walked dimension. + held it to one relation over the walked dimension. """ - over_node = node.kwargs['over'] - assert isinstance(over_node, DimensionNode), 'resolution refuses an over= that is not a dimension' - within_node = node.kwargs['within'] + over_node = node.kwargs['along'] + assert isinstance(over_node, DimensionNode), 'resolution refuses an along= that is not a dimension' + window_node = node.kwargs['window'] operand = self.expr(node.args[0]) wrap = isinstance(node.kwargs.get('edge'), EdgeNode) width: int | str - if isinstance(within_node, ParameterNode): - width = within_node.name + if isinstance(window_node, ParameterNode): + width = window_node.name else: - assert isinstance(within_node, NumberNode), 'a within= that is neither is refused at load' - width = int(within_node.value) + assert isinstance(window_node, NumberNode), 'a window= that is neither is refused at load' + width = int(window_node.value) return program.Window(operand, over_node.name, width=width, wrap=wrap, partition=_partition_of(node)) def shift(self, node: FunctionCallNode) -> program.ExpressionNode: - """``shift(x, over=d, offset=n)`` — the value at *t - offset* along one dim. + """``shift(x, along=d, offset=n)`` — the value at *t - offset* along one dim. What the vacated positions contribute is ``edge=``'s to say, and the language has already held it to the keyword or a number. """ - over_node = node.kwargs['over'] - assert isinstance(over_node, DimensionNode), 'resolution refuses an over= that is not a dimension' + over_node = node.kwargs['along'] + assert isinstance(over_node, DimensionNode), 'resolution refuses an along= that is not a dimension' by_node = node.kwargs['offset'] operand = self.expr(node.args[0]) edge = node.kwargs.get('edge') @@ -345,18 +342,18 @@ def shift(self, node: FunctionCallNode) -> program.ExpressionNode: } -def _partition_of(node: FunctionCallNode) -> str | None: - """The lookup a translation walks inside, if the call names one. +def _partition_of(node: FunctionCallNode) -> program.Walk | None: + """The walk a translation partitions by, if the call names a relation. - That it is a *single* lookup, and one *over the translated dimension*, is + That it is a *single* relation, walked *along the translated dimension*, is checked with the other dim rules (``math_spec.dimensions``), where a model is refused before any data is read. """ by_node = node.kwargs.get('by') if by_node is None: return None - assert isinstance(by_node, LookupNode) - return by_node.names[0] + assert isinstance(by_node, RelationNode) + return by_node.walks[0] def _bound_expression(value: float | str) -> program.ExpressionNode: diff --git a/src/math_spec/model.py b/src/math_spec/model.py index b7dadb76..0ce239d2 100644 --- a/src/math_spec/model.py +++ b/src/math_spec/model.py @@ -90,7 +90,7 @@ def _reject_unknown_keys(cls, data: Any) -> Any: ParameterDtype = Literal['float', 'int', 'bool', 'str'] #: What a *name* a where comparison tests may be — a parameter's dtype or a -#: dimension's, since a lookup's is its target's. The union rather than either +#: dimension's, since a relation's is its target's. The union rather than either #: half, because a mask names all three kinds and reads the dtype the same way. DeclaredDtype = ParameterDtype | DimensionDtype @@ -153,24 +153,64 @@ def _also_written_as( return {'anyOf': [dict(generated), shorthand]} -class LookupBlock(_StrictBlock): - """A named single-valued map out of one dimension ``into:`` another. +class RelationBlock(_StrictBlock): + """A named relation between dimensions, and the key it is single-valued per. - Its values are labels of ``into``, which is what ``sum(by=)`` and - ``at(by=)`` land terms on:: + ``columns:`` is the table's columns — a list of dimensions, or a mapping + of column name to dimension where two columns share one. ``key:`` names the + columns each row is identified by, and is the claim the language checks + at bind: one row per key tuple, so the other columns are a function of + it. Without a key the table is a bare relation:: - lookups: - bus_of: {over: generator, into: bus} + relations: + gen_bus: {columns: [generator, bus], key: generator} + zone_of: {columns: [generator, period, zone], key: [generator, period]} + rep_of: {columns: {snapshot: snapshot, rep: snapshot}, key: snapshot} + connection: {columns: [entity, bus]} - The map itself is data, and arrives at bind time under the lookup's name. + An operator walks the table in the direction the call names + (``over=``, ``into=``), joining on the other key columns; the + declaration fixes no direction. The map itself is data, and arrives at bind + time under the relation's name, one column per role. """ - _label: ClassVar[str] = 'a lookup declaration' + _label: ClassVar[str] = 'a relation declaration' - over: str - into: str + columns: str | list[str] | dict[str, str] + key: str | list[str] | None = None description: str | None = None + @property + def pairs(self) -> tuple[tuple[str, str], ...]: + """``(role, dimension)`` per column in declared order — a list names each role after its dimension. + + The program calls the same thing :attr:`~math_spec.program.RelationDeclaration.columns`; + here that name belongs to the field, which is what the file wrote. + """ + if isinstance(self.columns, dict): + return tuple(self.columns.items()) + return tuple((d, d) for d in ((self.columns,) if isinstance(self.columns, str) else self.columns)) + + @property + def roles(self) -> tuple[str, ...]: + return tuple(role for role, _ in self.pairs) + + @property + def dims(self) -> tuple[str, ...]: + return tuple(dim for _, dim in self.pairs) + + @property + def keys(self) -> tuple[str, ...]: + """The key roles, however ``key:`` was written; empty for a bare relation.""" + if self.key is None: + return () + return (self.key,) if isinstance(self.key, str) else tuple(self.key) + + @property + def values(self) -> tuple[str, ...]: + """The roles the key determines; empty where there is no key.""" + return tuple(role for role in self.roles if role not in self.keys) if self.key is not None else () + class DimensionBlock(_StrictBlock): """A declared dimension, and the dtype its coordinates must be. @@ -178,8 +218,8 @@ class DimensionBlock(_StrictBlock): A dimension is an axis and nothing else: it declares that the axis exists and what its coordinates are typed as, never which coordinates there are — those are data, and arrive at bind time. The maps its members carry — a - generator's bus, a snapshot's period — are top-level ``lookups:`` - (:class:`LookupBlock`), keyed by their own name. + generator's bus, a snapshot's period — are top-level ``relations:`` + (:class:`RelationBlock`), keyed by their own name. """ _label: ClassVar[str] = 'a dimension declaration' @@ -665,7 +705,7 @@ class Spec(_StrictBlock): #: ``description:`` takes. The typeset document opens with it. description: str | None = None dimensions: dict[str, DimensionBlock] = {} - lookups: dict[str, LookupBlock] = {} + relations: dict[str, RelationBlock] = {} parameters: dict[str, ParameterBlock] = {} variables: dict[str, VariableBlock] = {} constraints: dict[str, ConstraintBlock] = {} @@ -675,9 +715,9 @@ class Spec(_StrictBlock): piecewise: dict[str, PiecewiseBlock] = {} sos: dict[str, SosBlock] = {} - def lookups_of(self, dimension: str) -> dict[str, str]: - """The lookups over *dimension*: name -> the dim they map into.""" - return {n: lk.into for n, lk in self.lookups.items() if lk.over == dimension} + def relations_of(self, dimension: str) -> dict[str, RelationBlock]: + """The relations with a column over *dimension*, by name.""" + return {n: lk for n, lk in self.relations.items() if dimension in lk.dims} @classmethod @override @@ -755,7 +795,7 @@ def _validate_references(self) -> Spec: errors = [ *self._name_collisions(), *self._frame_dimensions(), - *self._lookup_targets(), + *self._relation_targets(), *self._bound_names(), *self._sos_shapes(), ] @@ -767,7 +807,7 @@ def _name_collisions(self) -> Iterator[str]: """A name is declared once, and never as a built-in operator.""" kinds: list[tuple[str, Iterable[str]]] = [ ('dimension', self.dimensions), - ('lookup', self.lookups), + ('relation', self.relations), ('parameter', self.parameters), ('variable', self.variables), ('named expression', self.expressions), @@ -806,20 +846,51 @@ def _frame_dimensions(self) -> Iterator[str]: if count > 1 ) - def _lookup_targets(self) -> Iterator[str]: - """A lookup is over a declared dimension and maps into a different declared one.""" - for lname, lk in self.lookups.items(): - if lk.over not in self.dimensions: - yield (undeclared_dimension('Lookup', lname, lk.over)) - if lk.into is not None: - if lk.into not in self.dimensions: - yield ( - f"Lookup '{lname}' targets undeclared dimension '{lk.into}'. " - f"Declare it under 'dimensions:' — the target is what the " - f'lookup values are checked against.' - ) - elif lk.into == lk.over: - yield (f"Lookup '{lname}' maps '{lk.over}' into itself. A lookup maps into a different dimension.") + def _relation_targets(self) -> Iterator[str]: + """A relation has at least two columns over declared dimensions, each role once, and a key that is a proper subset of them.""" + for lname, lk in self.relations.items(): + if len(lk.pairs) < 2: + yield ( + f"Relation '{lname}' has {len(lk.pairs)} column(s). A relation relates dimensions, so 'columns:' " + f'names at least two — a label on one dimension is a parameter over it.' + ) + yield from ( + f"Relation '{lname}' names dimension '{d}' twice under 'columns:'. Give the two columns roles: " + f'columns: {{{d}0: {d}, {d}1: {d}}}.' + for d, count in Counter(lk.dims).items() + if count > 1 and not isinstance(lk.columns, dict) + ) + yield from ( + undeclared_dimension('Relation', lname, d) for d in dict.fromkeys(lk.dims) if d not in self.dimensions + ) + yield from ( + f"Relation '{lname}' names column '{role}' after dimension '{role}', but the column is over " + f"'{dim}'. A column named like a dimension is read as over it — name it after what it holds." + for role, dim in lk.pairs + if role in self.dimensions and role != dim + ) + yield from ( + f"Relation '{lname}' has key column '{k}', which is not one of its columns {list(lk.roles)}." + for k in lk.keys + if k not in lk.roles + ) + yield from ( + f"Relation '{lname}' names '{k}' twice under 'key:'. A key names each column once." + for k, count in Counter(lk.keys).items() + if count > 1 + ) + yield from ( + f"Relation '{lname}' has two key columns over '{d}' ({[k for k in lk.keys if dict(lk.pairs)[k] == d]}). " + f'A key is read at its dimensions, and no frame carries a dimension twice — key the table by ' + f'one column over each, or leave one of them a value column.' + for d, count in Counter(dict(lk.pairs)[k] for k in lk.keys if k in lk.roles).items() + if count > 1 + ) + if lk.key is not None and set(lk.keys) >= set(lk.roles): + yield ( + f"Relation '{lname}' has every column in its key, so the key determines nothing. Leave one " + f'column out of it, or drop the key for a bare relation.' + ) def _bound_names(self) -> Iterator[str]: """A named bound is a numeric parameter.""" diff --git a/src/math_spec/operators.py b/src/math_spec/operators.py index 5c149a2a..0771a7bb 100644 --- a/src/math_spec/operators.py +++ b/src/math_spec/operators.py @@ -23,26 +23,32 @@ class Builtin: Keyword arguments come in four kinds, and the kind decides what resolution turns the value into: ``dimension_kwargs`` name a dimension - (``sum(x, over=generator)``); ``lookup_kwargs`` name a lookup, which + (``sum(x, over=generator)``); ``relation_kwargs`` name a relation, which carries its own dimensions, so it needs no sibling kwarg; ``edge_kwargs`` take a closed keyword or a number; ``required_value_kwargs`` are ordinary values that must be present — a number, never a name to resolve (``shift(..., offset=1)``). Every operator takes one positional argument, the expression; every - dimension or lookup it names arrives in a kwarg *value*, which is what + dimension or relation it names arrives in a kwarg *value*, which is what lets a macro pass one as a formal. ``usage`` is the wording every refusal quotes back. """ usage: str dimension_kwargs: tuple[str, ...] = () - lookup_kwargs: tuple[str, ...] = () - #: Kwargs of which the call carries at most one — ``sum`` takes ``over=`` - #: (reduce the dim away) or ``by=`` (reduce it into the lookup's target), - #: never both, and neither means every dim the operand carries. Members are - #: excluded from the required set; their kind still comes from the tuples - #: above. + relation_kwargs: tuple[str, ...] = () + #: Kwargs naming a column of the relation ``by=`` names — ``over=`` and + #: ``into=`` — which resolution folds into the relation's walk. + role_kwargs: tuple[str, ...] = () + #: Kwargs naming a dimension on their own and a column of the relation where + #: ``by=`` names one. ``sum(x, over=generator)`` reduces the dimension + #: away; ``sum(x, by=l, over=c)`` names the column the walk consumes. + #: One meaning — what leaves the frame — read in the namespace ``by=`` + #: decides. + dimension_or_role_kwargs: tuple[str, ...] = () + #: Kwargs of which the call carries at most one. Members are excluded from + #: the required set; their kind still comes from the tuples above. at_most_one_of: tuple[str, ...] = () edge_kwargs: tuple[str, ...] = () required_value_kwargs: tuple[str, ...] = () @@ -54,52 +60,69 @@ class Builtin: def required(self) -> frozenset[str]: """Every keyword the call must carry.""" return ( - (frozenset(self.dimension_kwargs) | frozenset(self.lookup_kwargs) | frozenset(self.required_value_kwargs)) + (frozenset(self.dimension_kwargs) | frozenset(self.relation_kwargs) | frozenset(self.required_value_kwargs)) - frozenset(self.at_most_one_of) - frozenset(self.optional_kwargs) ) - def kind_of(self, kwarg: str) -> Literal['dimension', 'lookup', 'edge', 'value']: - """What resolution turns the value of *kwarg* into: a dimension, a lookup, an edge policy, or a plain value.""" + def kind_of( + self, kwarg: str, *, with_relation: bool = False + ) -> Literal['dimension', 'relation', 'role', 'edge', 'value']: + """What resolution turns the value of *kwarg* into. + + A dimension, a relation, a column of it, an edge policy, or a plain value. + *with_relation* says whether the call carries a ``by=``, which is what + decides the kind of a :attr:`dimension_or_role_kwargs` member. + """ + if kwarg in self.dimension_or_role_kwargs: + return 'role' if with_relation else 'dimension' if kwarg in self.dimension_kwargs: return 'dimension' - if kwarg in self.lookup_kwargs: - return 'lookup' + if kwarg in self.relation_kwargs: + return 'relation' + if kwarg in self.role_kwargs: + return 'role' if kwarg in self.edge_kwargs: return 'edge' return 'value' -#: The closed operator set. ``by=`` is the one keyword that addresses a lookup, -#: and a lookup carries its own dimensions, so no sibling kwarg restates them. +#: The closed operator set. ``by=`` is the one keyword that addresses a relation, +#: and a relation carries its own dimensions, so no sibling kwarg restates them. #: On ``shift`` and ``sum_back`` it partitions the axis the operator walks: it -#: says which rows are neighbours, not which group a term lands in. +#: says which rows are neighbours, not which group a term lands in, and +#: ``within=`` names the columns whose values that group is read from. BUILTINS: dict[str, Builtin] = { 'sum': Builtin( - 'sum(), sum(, over=) or sum(, by=)', - dimension_kwargs=('over',), - lookup_kwargs=('by',), - at_most_one_of=('over', 'by'), + 'sum(), sum(, over=) or sum(, by=[, over=, into=])', + relation_kwargs=('by',), + role_kwargs=('into',), + dimension_or_role_kwargs=('over',), + optional_kwargs=('by', 'over', 'into'), ), 'at': Builtin( - 'at(, by=)', - lookup_kwargs=('by',), + 'at(, by=[, over=, into=])', + relation_kwargs=('by',), + role_kwargs=('over', 'into'), + optional_kwargs=('over', 'into'), ), 'sum_back': Builtin( - "sum_back(, over=, within=[, edge='wrap'][, by=])", - dimension_kwargs=('over',), - lookup_kwargs=('by',), - required_value_kwargs=('within',), + "sum_back(, along=, window=[, edge='wrap'][, by=[, within=]])", + dimension_kwargs=('along',), + relation_kwargs=('by',), + role_kwargs=('within',), + required_value_kwargs=('window',), edge_kwargs=('edge',), - optional_kwargs=('by',), + optional_kwargs=('by', 'within'), ), 'shift': Builtin( - "shift(, over=, offset=[, edge='wrap'|][, by=])", - dimension_kwargs=('over',), - lookup_kwargs=('by',), + "shift(, along=, offset=[, edge='wrap'|][, by=[, within=]])", + dimension_kwargs=('along',), + relation_kwargs=('by',), + role_kwargs=('within',), required_value_kwargs=('offset',), edge_kwargs=('edge',), - optional_kwargs=('by',), + optional_kwargs=('by', 'within'), ), 'dual': Builtin('dual()'), } @@ -128,7 +151,7 @@ def call_shape_error(name: str, positional: int, kwargs: Iterable[str]) -> str | if len(keys & set(builtin.at_most_one_of)) > 1: alternatives = ' or '.join(f'{k}=' for k in builtin.at_most_one_of) return ( - f'{name}() takes at most one of {alternatives} — a lookup carries ' + f'{name}() takes at most one of {alternatives} — a relation carries ' f'its own dimensions, so by= leaves over= nothing to add.\n' f'Write: {builtin.usage}' ) diff --git a/src/math_spec/piecewise.py b/src/math_spec/piecewise.py index 7123c4f0..1df23845 100644 --- a/src/math_spec/piecewise.py +++ b/src/math_spec/piecewise.py @@ -206,7 +206,7 @@ def _weights(self) -> None: self._constraint( self.adjacency, [*self.frame, d], - f'{self.lam} <= {self.seg} + shift({self.seg}, over={d}, offset=1, edge=0)', + f'{self.lam} <= {self.seg} + shift({self.seg}, along={d}, offset=1, edge=0)', ) def _gate_rows(self) -> tuple[tuple[str, str | None, str], ...]: @@ -247,8 +247,8 @@ def _segment_lines(self) -> None: x_link, y_link = self.pw.curve d = self.pw.over mask = self.mask - run = f'({x_link.values} - shift({x_link.values}, over={d}, offset=1, edge=0))' - rise = f'({y_link.values} - shift({y_link.values}, over={d}, offset=1, edge=0))' + run = f'({x_link.values} - shift({x_link.values}, along={d}, offset=1, edge=0))' + rise = f'({y_link.values} - shift({y_link.values}, along={d}, offset=1, edge=0))' interior = f'{mask} AND NOT {self.starts}' if mask else f'position({d}) != 0' self._constraint( self.chord, diff --git a/src/math_spec/program.py b/src/math_spec/program.py index 5478a31a..7567aa09 100644 --- a/src/math_spec/program.py +++ b/src/math_spec/program.py @@ -66,10 +66,6 @@ 'GroupSum', 'Increasing', 'LastOf', - 'LookupComparisonNode', - 'LookupDeclaration', - 'LookupDefinedNode', - 'LookupPairComparisonNode', 'Mask', 'MaskOf', 'Multiply', @@ -90,6 +86,10 @@ 'QuadraticPosition', 'Reach', 'Region', + 'RelationComparisonNode', + 'RelationDeclaration', + 'RelationDefinedNode', + 'RelationPairComparisonNode', 'Separability', 'SosDeclaration', 'Sum', @@ -100,6 +100,7 @@ 'VariableDeclaration', 'VariableDefinedNode', 'VariableDomain', + 'Walk', 'WhereNode', 'Window', 'carries_variable', @@ -262,34 +263,61 @@ class Sum(Expression): @dataclass(frozen=True) class GroupSum(Expression): - """Sum ``operand`` through coordinates declared on dim ``over``. - - ``coordinate`` names lookups carried by dim ``over`` whose values are - labels of the matching dim in ``into``; the result replaces ``over`` with - all of them. The two tuples are the same length and their order pairs - them: several coordinates are one grouping into a product of targets, - consumed in a single join. + """Sum ``operand`` through relations, consuming the dims ``over`` and producing ``into``. + + ``walks`` says, per relation, which columns are consumed, which produced + and which joined on, and is the one fact the node holds: ``coordinate`` + names the relations, ``over`` is the dims every walk consumes and ``into`` + the dims they produce, in walk order, so that several coordinates are + one grouping into a product of targets, consumed in a single join. The + result replaces every dim in ``over`` with every dim in ``into``. The + join keys on the consumed columns and every joined column, and on a + produced column too where the operand already carries its dimension. """ operand: ExpressionNode - over: str - coordinate: tuple[str, ...] - into: tuple[str, ...] + walks: tuple[Walk, ...] + + @property + def coordinate(self) -> tuple[str, ...]: + return tuple(walk.name for walk in self.walks) + + @property + def over(self) -> tuple[str, ...]: + return self.walks[0].consumed_dims + + @property + def into(self) -> tuple[str, ...]: + return tuple(dim for walk in self.walks for dim in walk.produced_dims) @dataclass(frozen=True) class At(Expression): - """Read ``operand`` through a lookup — the adjoint of :class:`GroupSum`. - - Same mapping table, walked the other way: ``GroupSum`` consumes ``over`` - and produces ``into``, this consumes ``into`` and produces ``over``. The - join fans out, many ``over`` labels sharing one ``into`` tuple. + """Read ``operand`` through relations — the adjoint of :class:`GroupSum`. + + Same tables, walked the other way: this consumes the dims in ``into`` and + produces the dims in ``over``, one value per coordinate because every + walk reads value columns at a key the operand fixes + (``Walk.is_function_read``). The join fans out, many ``over`` tuples + sharing one ``into`` tuple — at each coordinate of the joined columns, + which the operand carries and the result keeps. As on + :class:`GroupSum`, ``walks`` is the fact and the three are read off it. """ operand: ExpressionNode - over: str - coordinate: tuple[str, ...] - into: tuple[str, ...] + walks: tuple[Walk, ...] + + @property + def coordinate(self) -> tuple[str, ...]: + return tuple(walk.name for walk in self.walks) + + @property + def over(self) -> tuple[str, ...]: + return self.walks[0].produced_dims + + @property + def into(self) -> tuple[str, ...]: + return tuple(dim for walk in self.walks for dim in walk.consumed_dims) @dataclass(frozen=True) @@ -304,10 +332,12 @@ class Translate(Expression): ``offset`` is an integer, or the name of an integer parameter that does not depend on ``dimension`` and carries its sign in the values. - ``partition`` names a lookup over ``dimension``, and the translation then - happens inside each group it makes: the neighbour is the one before in - the same group, the edge is the group's, and a wrap closes each group onto - itself. A coordinate the lookup sends nowhere reaches nothing. + ``partition`` is a relation walked along ``dimension`` — its consumed + column is a key over that dimension, its produced columns are the group — + and the translation then happens inside each group: the neighbour is the + one before in the same group, the edge is the group's, and a wrap closes + each group onto itself. A coordinate the relation sends nowhere reaches + nothing. """ operand: ExpressionNode @@ -315,7 +345,7 @@ class Translate(Expression): offset: int | str wrap: bool fill: float | None = None - partition: str | None = None + partition: Walk | None = None @dataclass(frozen=True) @@ -334,16 +364,16 @@ class Window(Expression): ``wrap`` says whether the window reaches around the start of the axis instead of stopping short at it, and is stated on every node. - ``partition`` names a lookup over that dimension, and the window then stops + ``partition`` names a relation over that dimension, and the window then stops at each group's edge. Positions are counted inside the group, so a - coordinate the lookup places nowhere reaches nothing — not even itself. + coordinate the relation places nowhere reaches nothing — not even itself. """ operand: ExpressionNode dimension: str width: int | str wrap: bool - partition: str | None = None + partition: Walk | None = None @dataclass(frozen=True) @@ -437,41 +467,104 @@ def children(expression: ExpressionNode) -> tuple[ExpressionNode, ...]: # -------------------------------------------------------------------------- -class LookupDeclaration(NamedTuple): - """One declared lookup over a dimension. +class RelationDeclaration(NamedTuple): + """One declared relation: a relation over its ``columns``, single-valued per ``key``. - Its values are labels of ``target``, checked for containment once the dim - tables exist — which keeps a mistyped label from silently dropping its - terms in the join that places them — and it is what ``sum(by=)`` lands - terms on. + ``columns`` binds each role to its dimension in the order the table + carries them; ``key`` is the roles a row is identified by, empty for a + bare relation. Every value is checked at bind to be a label of its + column's dimension, and a keyed table to have one row per key tuple — + which keeps a mistyped label from silently dropping its terms in the join + that places them, and is what lets ``at`` read one value. """ name: str - target: str + columns: tuple[tuple[str, str], ...] + key: tuple[str, ...] = () + + @property + def roles(self) -> tuple[str, ...]: + return tuple(role for role, _ in self.columns) + + @property + def dims(self) -> tuple[str, ...]: + return tuple(dim for _, dim in self.columns) + + @property + def values(self) -> tuple[str, ...]: + """The roles the key determines.""" + return tuple(role for role in self.roles if role not in self.key) + + def dim(self, role: str) -> str: + return dict(self.columns)[role] + + +class Walk(NamedTuple): + """One relation as an operator walks it — which columns are consumed, which produced, which joined on. + + ``consumed``, ``produced`` and ``joined`` are *roles* — column names of + ``relation``, which binds every role to its dimension and names the key. + ``joined`` is the key roles not walked (every role, for a bare relation): + the join keys on them, and a value role not walked is not read. For a + partition (``shift``, ``sum_back``, ``position``) ``consumed`` is the key + role over the dimension walked and ``produced`` the value roles that make + the group — every value role unless the call named some with ``within=``. + """ + + relation: RelationDeclaration + consumed: tuple[str, ...] + produced: tuple[str, ...] + joined: tuple[str, ...] + + @property + def name(self) -> str: + return self.relation.name + + @property + def key(self) -> tuple[str, ...]: + return self.relation.key + + @property + def roles(self) -> tuple[str, ...]: + return self.relation.roles + + @property + def values(self) -> tuple[str, ...]: + return self.relation.values + + def dim(self, role: str) -> str: + """The dimension *role* is bound to.""" + return self.relation.dim(role) + + @property + def consumed_dims(self) -> tuple[str, ...]: + return tuple(self.dim(role) for role in self.consumed) + + @property + def produced_dims(self) -> tuple[str, ...]: + return tuple(self.dim(role) for role in self.produced) + + @property + def joined_dims(self) -> tuple[str, ...]: + return tuple(self.dim(role) for role in self.joined) + + @property + def is_function_read(self) -> bool: + """Whether the walk reads one value per coordinate: the key lies inside what is fixed.""" + return bool(self.key) and set(self.key) <= {*self.joined, *self.produced} @dataclass(frozen=True) class DimensionDeclaration: - """A dimension and the lookups its labels carry.""" + """A dimension and the relations with a column over it.""" - lookups: tuple[LookupDeclaration, ...] = () + relations: tuple[RelationDeclaration, ...] = () #: What the labels are, as the file declares them. A dimension is read from #: whatever table carries it, so the declared type is what that column is #: checked against — the same claim ``ParameterDeclaration.dtype`` makes #: about a value column, one axis over. dtype: DimensionDtype = 'str' - @property - def targets(self) -> dict[str, str]: - """Each map over the dimension, to the dimension its values are labels of. - - The question every consumer of a ``by=`` asks, and asked here so it has - one answer: an operator grouping through a lookup names the target as - the dim it lands on, and a partition array is named for it so an amount - declared over the group's own dim can be read through it. - """ - return {lk.name: lk.target for lk in self.lookups} - @dataclass(frozen=True) class MaskOf: @@ -737,10 +830,10 @@ class Reach: Attributes: label: The declaration reading, as the lowering's messages label it. - name: The parameter or lookup that says how far. + name: The parameter or relation that says how far. kind: An ``offset`` is a parameter's values, which :meth:`Separability.resolved` folds in; a ``partition`` and a - ``coordinate`` are a lookup's groups, which it does not. + ``coordinate`` are a relation's groups, which it does not. """ label: str @@ -778,7 +871,7 @@ class Separability: applied. undecided: Each read along the axis whose reach only data can say — a named offset, a partition whose groups a window may cut, a read - through a lookup at a coordinate the data chooses. + through a relation at a coordinate the data chooses. :meth:`resolved` folds a parameter's values in. restarts: Each declaration counting a position along the axis, which a window restarts at its first row. Whether that is wanted — a seed @@ -811,7 +904,7 @@ def resolved(self, least: Mapping[str, int]) -> Separability: Args: least: Parameter name to the least of its values. A reach through - a lookup — a partition, a coordinate — cannot be folded this + a relation — a partition, a coordinate — cannot be folded this way and stays undecided, as does a parameter left out. Raises: @@ -894,9 +987,9 @@ def dimension(self, name: str) -> DimensionDeclaration: return _declared(self.dimensions, name, 'dimension') @property - def lookups(self) -> tuple[tuple[str, LookupDeclaration], ...]: - """Every lookup in the program, with the dimension it is over.""" - return tuple((dimension, lk) for dimension, d in self.dimensions.items() for lk in d.lookups) + def relations(self) -> dict[str, RelationDeclaration]: + """Every relation in the program by name, each once — a relation keyed by two dimensions sits under both.""" + return {lk.name: lk for d in self.dimensions.values() for lk in d.relations} def parameter(self, name: str) -> ParameterDeclaration: return _declared(self.parameters, name, 'parameter') @@ -1056,45 +1149,61 @@ class DimensionComparisonNode: class DimensionPositionNode: """Compare where a row sits along a dimension against a position — ``position(snapshot) == 0``. - Both sides are integers, negative counting from the end. With ``by`` the - position is counted within each group the lookup makes. + Both sides are integers, negative counting from the end. With a + ``partition`` the position is counted within each group the relation makes, + walked as :class:`Translate` walks one: its consumed column is the key + column over ``name``, the group is its produced columns, and its joined + columns are the other key columns, whose dimensions the frame carries. """ name: str op: PredicateOperator position: int - by: str | None = None + partition: Walk | None = None @dataclass(frozen=True) -class LookupComparisonNode: - """Compare a lookup's values against a literal — ``period_of == 2030``. +class RelationComparisonNode: + """Compare one value column of a keyed relation against a literal — ``period_of == 2030``. - ``over`` is the dimension the lookup maps out of. + ``column`` is the role read, and ``dims`` the dimensions of the key + columns: the leaf is read at them, one value per coordinate. """ name: str - over: str + column: str op: PredicateOperator value: float | str | datetime.date + dims: tuple[str, ...] = () @dataclass(frozen=True) -class LookupPairComparisonNode: - """Compare two lookups over one dimension — ``from != to``, row by row on that dimension's table.""" +class RelationPairComparisonNode: + """Compare a value column of one keyed relation with one of another — ``from_bus != to_bus`` — row by row on the key. + + Both keys are over the same ``dims``, and the two columns are over one + dimension, so a match is possible at all. + """ name: str + column: str other: str - over: str + other_column: str op: PredicateOperator + dims: tuple[str, ...] = () @dataclass(frozen=True) -class LookupDefinedNode: - """True where the named lookup has a value — a null says the label belongs to no group.""" +class RelationDefinedNode: + """True where the relation has a row at the frame's coordinates. + + ``dims`` is what the frame supplies: the key's dimensions for a keyed + relation, whose row is then the one the key finds; every column's for a + bare relation, where a row is the whole tuple. + """ name: str - over: str + dims: tuple[str, ...] = () @dataclass(frozen=True) @@ -1125,9 +1234,9 @@ class OrNode: | VariableDefinedNode | ParameterComparisonNode | DimensionComparisonNode - | LookupComparisonNode - | LookupPairComparisonNode - | LookupDefinedNode + | RelationComparisonNode + | RelationPairComparisonNode + | RelationDefinedNode | NotNode | AndNode | OrNode @@ -1142,9 +1251,9 @@ class OrNode: | VariableDefinedNode | DimensionComparisonNode | DimensionPositionNode - | LookupComparisonNode - | LookupPairComparisonNode - | LookupDefinedNode + | RelationComparisonNode + | RelationPairComparisonNode + | RelationDefinedNode ) #: The boolean connectives — the only where nodes carrying other where nodes, @@ -1190,20 +1299,23 @@ def _atom_dims(atom: TypedPredicateNode) -> frozenset[str]: """One leaf's dims — the rule :attr:`Mask.dims` is the union of. A parameter or variable leaf carries its own dims off the declaration; a - comparison on a dimension is read through that dimension, and a lookup - through the dimension it maps out of — the dim it leaves, not the one it - lands in. Separate from the union because the load-time frame check - reports per leaf. Closed by ``assert_never``: a predicate node added - without a reading is a type error here, at the one place that has to grow - a branch, rather than a wrong dim set at the first model to use it. + comparison on a dimension is read through that dimension, and a relation + through the dimensions of the columns it is read at — its key for a + comparison, every column for a bare existence. + Separate from the union because the load-time frame check reports per + leaf. Closed by ``assert_never``: a predicate node added without a reading + is a type error here, at the one place that has to grow a branch, rather + than a wrong dim set at the first model to use it. """ match atom: case ParameterComparisonNode() | ParameterDefinedNode() | VariableDefinedNode(): return frozenset(atom.dims) - case DimensionComparisonNode() | DimensionPositionNode(): + case DimensionComparisonNode(): return frozenset({atom.name}) - case LookupComparisonNode() | LookupPairComparisonNode() | LookupDefinedNode(): - return frozenset({atom.over}) + case DimensionPositionNode(): + return frozenset({atom.name, *(atom.partition.joined_dims if atom.partition is not None else ())}) + case RelationComparisonNode() | RelationPairComparisonNode() | RelationDefinedNode(): + return frozenset(atom.dims) case _: assert_never(atom) @@ -1212,7 +1324,7 @@ def _atom_names(atom: TypedPredicateNode) -> frozenset[str]: """One leaf's declarations, its dimension apart — the rule :attr:`Mask.names_read` is the union of. A comparison on a dimension names no declaration — a coordinate is not - data to feed — and a lookup pair names both maps it compares. + data to feed — and a relation pair names both maps it compares. ``assert_never``-closed for the reason :func:`_atom_dims` is: a predicate node added without a reading is a type error at this one branch rather than a name silently dropped at the first model to use it. @@ -1222,11 +1334,11 @@ def _atom_names(atom: TypedPredicateNode) -> frozenset[str]: ParameterComparisonNode() | ParameterDefinedNode() | VariableDefinedNode() - | LookupComparisonNode() - | LookupDefinedNode() + | RelationComparisonNode() + | RelationDefinedNode() ): return frozenset({atom.name}) - case LookupPairComparisonNode(): + case RelationPairComparisonNode(): return frozenset({atom.name, atom.other}) case DimensionComparisonNode() | DimensionPositionNode(): return frozenset() @@ -1316,7 +1428,7 @@ def conjuncts(self) -> tuple[WhereNode, ...]: @property def names_read(self) -> frozenset[str]: - """The parameters, lookups and variables the mask names.""" + """The parameters, relations and variables the mask names.""" return frozenset(name for atom in self.atoms for name in _atom_names(atom)) @property diff --git a/src/math_spec/resolution.py b/src/math_spec/resolution.py index 4af355b0..e70d1aab 100644 --- a/src/math_spec/resolution.py +++ b/src/math_spec/resolution.py @@ -30,12 +30,12 @@ FunctionCallNode, KeywordNode, KwargNode, - LookupNode, NameListNode, NameNode, NumberNode, ParameterNode, ParsedNode, + RelationNode, UnaryOperatorNode, VariableNode, case_context, @@ -65,16 +65,19 @@ BooleanLiteralNode, DimensionComparisonNode, DimensionPositionNode, - LookupComparisonNode, - LookupDefinedNode, - LookupPairComparisonNode, Mask, NotNode, OrNode, ParameterComparisonNode, ParameterDefinedNode, + PredicateOperator, + RelationComparisonNode, + RelationDeclaration, + RelationDefinedNode, + RelationPairComparisonNode, TypedPredicateNode, VariableDefinedNode, + Walk, WhereNode, ) @@ -87,7 +90,7 @@ #: What a name a file may write turns out to be. Answered by #: :meth:`Namespace.kind`, so a pass reading a name switches over this rather #: than over the stores it would otherwise have to try in order. -DeclarationKind = Literal['variable', 'parameter', 'dimension', 'lookup'] +DeclarationKind = Literal['variable', 'parameter', 'dimension', 'relation'] class Namespace: @@ -96,14 +99,14 @@ class Namespace: A name has one kind: model.py refuses one declared under two sections. """ - __slots__ = ('constraints', 'dimensions', 'dtypes', 'leaf_dims', 'lookups', 'parameters', 'variables') + __slots__ = ('constraints', 'dimensions', 'dtypes', 'leaf_dims', 'parameters', 'relations', 'variables') def __init__( self, variables: Iterable[str], parameters: Iterable[str], dimensions: Iterable[str], - lookups: Mapping[str, tuple[str, str]], + relations: Mapping[str, RelationDeclaration], dtypes: Mapping[str, DeclaredDtype], leaf_dims: Mapping[str, tuple[str, ...]], constraints: Iterable[str], @@ -115,31 +118,27 @@ def __init__( #: never reaches them, so a model may name a constraint after a variable. #: Consulted only in ``dual()``'s argument position. self.constraints = frozenset(constraints) - #: name -> declared dtype, for dimensions, parameters and lookups alike; + #: name -> declared dtype, for dimensions, parameters and relations alike; #: what a where comparison checks its literal against. self.dtypes: dict[str, DeclaredDtype] = dict(dtypes) - #: lookup name -> ``(over, into)``. - self.lookups: dict[str, tuple[str, str]] = dict(lookups) + #: relation name -> its columns and key, as declared. + self.relations: dict[str, RelationDeclaration] = dict(relations) #: parameter or variable name -> the dims it is read through — #: parameters by their ``dims``, variables by their frame. Stamped onto - #: each leaf a where names, the way a lookup leaf carries ``over``. + #: each leaf a where names, the way a relation leaf carries ``over``. self.leaf_dims: dict[str, tuple[str, ...]] = dict(leaf_dims) @classmethod def of(cls, schema: Spec) -> Namespace: - """Build the namespace of *schema*, the whole of what a file may name. - - A lookup's values are labels of its target, so its dtype is the target's. - """ + """Build the namespace of *schema*, the whole of what a file may name.""" return cls( schema.variables, schema.parameters, schema.dimensions, - {n: (lk.over, lk.into) for n, lk in schema.lookups.items()}, + {n: RelationDeclaration(n, lk.pairs, lk.keys) for n, lk in schema.relations.items()}, { **{p: pd.dtype for p, pd in schema.parameters.items()}, **{d: dd.dtype for d, dd in schema.dimensions.items()}, - **{n: schema.dimensions[lk.into].dtype for n, lk in schema.lookups.items()}, }, { **{p: tuple(pd.dims) for p, pd in schema.parameters.items()}, @@ -156,17 +155,13 @@ def kind(self, name: str) -> DeclarationKind | None: return 'parameter' if name in self.dimensions: return 'dimension' - if name in self.lookups: - return 'lookup' + if name in self.relations: + return 'relation' return None - def over_of(self, lookup: str) -> str: - """The dimension *lookup* maps out of.""" - return self.lookups[lookup][0] - - def into_of(self, lookup: str) -> str: - """The dimension *lookup*'s values are labels of.""" - return self.lookups[lookup][1] + def shape_of(self, relation: str) -> RelationDeclaration: + """The columns and key of *relation*, as declared.""" + return self.relations[relation] def unknown(self, name: str, context: str, *, allow_dims: bool, formals: Iterable[str] = ()) -> str: """The refusal for a *name* declared nowhere, listing what it could have been. @@ -175,14 +170,14 @@ def unknown(self, name: str, context: str, *, allow_dims: bool, formals: Iterabl name: The name the file wrote. context: The declaration it was found in. allow_dims: Whether a dimension would have been accepted there. It marks a - where string, which reads a lookup as readily as a parameter, so the - listing carries the lookups too; an expression, where a lookup is not a + where string, which reads a relation as readily as a parameter, so the + listing carries the relations too; an expression, where a relation is not a value, lists the variables instead. formals: A macro's formals, listed first when there are any. """ shown: list[tuple[str, Iterable[str]]] = [('Formals', formals)] if formals else [] shown += ( - [('Parameters', self.parameters), ('Dimensions', self.dimensions), ('Lookups', self.lookups)] + [('Parameters', self.parameters), ('Dimensions', self.dimensions), ('Relations', self.relations)] if allow_dims else [('Variables', self.variables), ('Parameters', self.parameters)] ) @@ -290,7 +285,7 @@ def where_of(text: str | None, ns: Namespace, context: str, self_variable: str | def names_in(value: ArithmeticNode) -> tuple[str, ...]: - """The names a lookup kwarg carries: one bare, several bracketed, none otherwise.""" + """The names a relation kwarg carries: one bare, several bracketed, none otherwise.""" if isinstance(value, NameNode): return (value.name,) return value.names if isinstance(value, NameListNode) else () @@ -394,7 +389,7 @@ def expression(self, node: ParsedNode) -> ParsedNode: def _arith(self, node: ArithmeticNode, *, amount: bool = False) -> ArithmeticNode: """One arithmetic node typed. - *amount* marks an ``offset=``/``within=`` value, whose dtype rule is + *amount* marks an ``offset=``/``window=`` value, whose dtype rule is ``dimensions._check_named_amount``'s and stricter than "a number", so the numeric check here stands aside for it. A quoted keyword or a name list in arithmetic arrives through a macro formal bound to one. @@ -426,7 +421,7 @@ def _arith(self, node: ArithmeticNode, *, amount: bool = False) -> ArithmeticNod assert_never(node) def _name(self, node: NameNode, *, amount: bool) -> ArithmeticNode: - """A bare name as the variable or parameter it declares; a dimension or lookup is not a value.""" + """A bare name as the variable or parameter it declares; a dimension or relation is not a value.""" match self.ns.kind(node.name): case 'variable': return VariableNode(node.name) @@ -445,10 +440,10 @@ def _name(self, node: NameNode, *, amount: bool) -> ArithmeticNode: f'declare a parameter over it.' ) return node - case 'lookup': + case 'relation': self.errors.append( - f"{self.context}: '{node.name}' is a lookup, and a lookup is structure " - f'rather than data, so it is not a value in an expression. A lookup ' + f"{self.context}: '{node.name}' is a relation, and a relation is structure " + f'rather than data, so it is not a value in an expression. A relation ' f'appears in a helper (sum(x, by={node.name})) and in a where — to ' f'carry numbers along this dimension, declare a parameter over it.' ) @@ -470,14 +465,23 @@ def _call(self, node: FunctionCallNode) -> ArithmeticNode: return node if shape_error is not None else self._dual(node) args = tuple(self._arith(a) for a in node.args) kwargs: dict[str, ArithmeticNode] = {} + with_relation = any(k in node.kwargs for k in builtin.relation_kwargs) + roles = {k: v for k, v in node.kwargs.items() if builtin.kind_of(k, with_relation=with_relation) == 'role'} + if roles and 'by' not in node.kwargs: + self.errors.append( + f'{self.context}: {node.name}({", ".join(f"{k}=" for k in roles)}) names a column of a relation, ' + f'and no by= names the relation. Write {builtin.usage}' + ) for key, value in node.kwargs.items(): - match builtin.kind_of(key): + match builtin.kind_of(key, with_relation=with_relation): case 'edge': kwargs[key] = self._edge(value, node.name) case 'dimension': kwargs[key] = self._dim_ref(value, node.name, key) - case 'lookup': - kwargs[key] = self._lookup_ref(value, node.name, key) + case 'relation': + kwargs[key] = self._relation_ref(value, node.name, key, roles, node.kwargs.get('along')) + case 'role': + pass case 'value': kwargs[key] = self._amount(value, node.name, key) return FunctionCallNode(node.name, args, kwargs) @@ -492,7 +496,7 @@ def _cases(self, node: CasesNode) -> CasesNode: return CasesNode(node.name, tuple(arms)) def _amount(self, value: ArithmeticNode, operator: str, key: str) -> ArithmeticNode: - """``offset=`` or ``within=``: a number or a parameter name, never an expression. + """``offset=`` or ``window=``: a number or a parameter name, never an expression. Closed so that :func:`math_spec.dimensions._check_named_amount` sees every parameter an amount carries. @@ -561,65 +565,260 @@ def _dual(self, node: FunctionCallNode) -> ArithmeticNode: return node return DualNode(value.name) - def _lookup_ref(self, value: ArithmeticNode, operator: str, key: str) -> ArithmeticNode: - """An operator kwarg whose *value* must name lookups. - - A lookup carries its own dimensions, so nothing else in the call is - consulted: the names alone decide both the dim the operator consumes and - the ones it produces. A bracketed list is one grouping through several - maps at once rather than a composition of groupings, so its members must - share the dim they are over and must not target the same dim twice. + def _relation_ref( + self, + value: ArithmeticNode, + operator: str, + key: str, + roles: Mapping[str, ArithmeticNode], + over: ArithmeticNode | None, + ) -> ArithmeticNode: + """An operator's ``by=``, with the ``over=`` and ``into=`` that say how each relation is walked. + + A relation carries its own dimensions, so the call names columns rather + than dims: ``over=`` the column consumed, ``into=`` the column produced, + every other key column joined on — a value column not walked is not + read, and a bare relation's columns are all key. Where the declaration + leaves one choice + — a key of one column, a value of one column — the call may leave it + unsaid. A bracketed list is one grouping through several tables at + once rather than a composition of groupings, so its members walk the + same dimension, take their defaults, and must not produce the same + dim twice. """ names = names_in(value) if not names: - self.errors.append(f'{self.context}: {operator}({key}=...) must name a lookup.') + self.errors.append(f'{self.context}: {operator}({key}=...) must name a relation.') return value - ns = self.ns - named = [self._not_a_lookup(name, operator, key) for name in names] - if any(problem is not None for problem in named): - self.errors.extend(problem for problem in named if problem is not None) + problems = [p for p in (self._not_a_relation(n, operator, key) for n in names) if p is not None] + if problems: + self.errors.extend(problems) return value - - over = {ns.over_of(name) for name in names} - if len(over) > 1: + if len(names) > 1 and roles: self.errors.append( - f'{self.context}: {operator}({key}={shown(names)}) groups through lookups over ' - f'different dimensions ({", ".join(f"{n} over {ns.over_of(n)}" for n in names)}). ' - f'One grouping consumes one dimension, so every lookup in the list must be ' - f'over the same one — group through them in turn instead, one call each.' + f'{self.context}: {operator}({key}={shown(names)}, {", ".join(f"{k}=" for k in roles)}): a list ' + f'walks each relation by its declared key and value, so a column keyword has nothing to name. ' + f'Name one relation, or declare one table with the columns of both.' ) return value + named = {k: self._role_name(v, operator, k) for k, v in roles.items()} + if any(r is None for r in named.values()): + return value + if operator in ('shift', 'sum_back'): + over_dim = over.name if isinstance(over, NameNode | DimensionNode) else None + walks = [self._partition_walk(n, operator, over_dim, named.get('within')) for n in names] + else: + walks = [self._walk(n, operator, named.get('over'), named.get('into')) for n in names] + if any(w is None for w in walks): + return value + resolved = [w for w in walks if w is not None] + + partition = operator in ('shift', 'sum_back') + + def fine_of(w: Walk) -> tuple[str, ...]: + return w.produced_dims if operator == 'at' else w.consumed_dims - targets = tuple(ns.into_of(name) for name in names) - repeated = sorted({t for t in targets if targets.count(t) > 1}) + def coarse_of(w: Walk) -> tuple[str, ...]: + """The dims the call lands on — none for a partition, whose produced columns are a group, not a frame.""" + if partition: + return () + return w.consumed_dims if operator == 'at' else w.produced_dims + + fine = {frozenset(fine_of(w)) for w in resolved} + if len(fine) > 1: + self.errors.append( + f'{self.context}: {operator}({key}={shown(names)}) groups through relations along different ' + f'dimensions ({", ".join(f"{w.name} along {sorted(fine_of(w))}" for w in resolved)}). One grouping ' + f'consumes one set of dimensions, so every relation in the list must walk the same — group through ' + f'them in turn instead, one call each.' + ) + return value + coarse = tuple(dim for w in resolved for dim in coarse_of(w)) + repeated = sorted({t for t in coarse if coarse.count(t) > 1}) if repeated: self.errors.append( - f'{self.context}: {operator}({key}={shown(names)}) targets {repeated} more than once. ' - f'Each lookup in the list produces its own dimension, so two that land on the ' - f'same one would need it twice — drop one, or group into a dimension of its own.' + f'{self.context}: {operator}({key}={shown(names)}) produces {repeated} more than once. ' + f'Each column walked to produces its own dimension, so two that land on the ' + f'same one would need it twice — drop one.' ) return value + return RelationNode(names, dimensions=fine_of(resolved[0]), into=coarse, walks=tuple(resolved)) - return LookupNode(names, dimension=next(iter(over)), into=targets) + def _role_name(self, value: ArithmeticNode, operator: str, key: str) -> tuple[str, ...] | None: + """``over=`` or ``into=`` as the column names it must be — one bare name, or a bracketed list of them.""" + if isinstance(value, NameNode): + return (value.name,) + if isinstance(value, NameListNode): + return value.names + self.errors.append( + f'{self.context}: {operator}({key}=...) names columns of the relation — a bare name, or a list of them.' + ) + return None + + def _walk( + self, + name: str, + operator: str, + from_roles: tuple[str, ...] | None, + into_roles: tuple[str, ...] | None, + ) -> Walk | None: + """How ``sum`` or ``at`` walks relation *name*, from the columns the call named and the declaration's defaults. + + The call consumes one or more columns and produces one or more; a + side it leaves unsaid is taken from the declaration where it has + exactly one candidate, and refused with the candidates otherwise. + ``at`` needs the walk single-valued and ``sum`` needs it not: a sum + that walks to the key has one term per coordinate and adds up nothing, + which is a read, so it is refused toward ``at``. + """ + ns, context = self.ns, self.context + shape = ns.shape_of(name) + call = f'{operator}(by={name})' + if not ( + self._known_roles(name, call, from_roles, 'over') and self._known_roles(name, call, into_roles, 'into') + ): + return None + + forward = operator == 'sum' + if from_roles is None: + side = shape.key if forward else shape.values + default = self._default_role(name, call, 'over', side, 'key' if forward else 'value') + if default is None: + return None + from_roles = (default,) + if into_roles is None: + side = shape.values if forward else shape.key + default = self._default_role(name, call, 'into', side, 'value' if forward else 'key') + if default is None: + return None + into_roles = (default,) + if both := sorted(set(from_roles) & set(into_roles)): + self.errors.append( + f'{context}: {call}: over= and into= both name {both}, and a walk goes between two sets of columns.' + ) + return None + for kwarg, roles in (('over', from_roles), ('into', into_roles)): + dims = [shape.dim(r) for r in roles] + if shared := sorted({d for d in dims if dims.count(d) > 1}): + self.errors.append( + f'{context}: {call}: {kwarg}={list(roles)} names two columns over {shared}, and the operand ' + f'carries each dimension once, so nothing says which column its coordinate is read at. Walk ' + f'between columns over distinct dimensions.' + ) + return None + joined = tuple(r for r in (shape.key or shape.roles) if r not in from_roles and r not in into_roles) + walk = Walk(shape, from_roles, into_roles, joined) + if not forward and not walk.is_function_read: + self.errors.append( + f"{context}: {call}: at reads one value per coordinate, and '{name}' is not single-valued in " + f'{list(from_roles)} at the columns the operand fixes ({[*into_roles, *joined]}) — its key is ' + f'{list(shape.key)}. Declare a key those columns contain, or read the other way.' + ) + return None + if forward and walk.is_function_read: + self.errors.append( + f'{context}: {call}: this sum walks to the key {list(shape.key)}, so each coordinate has one ' + f"term and nothing is added up — that is a read, which is at()'s. Write " + f'at(..., by={name}, over={list(from_roles)}, into={list(into_roles)}), or sum toward ' + f'a value column.' + ) + return None + return walk + + def _known_roles(self, name: str, call: str, roles: tuple[str, ...] | None, kwarg: str) -> bool: + """Whether every role *kwarg* names is a column of relation *name*, each once; the refusal otherwise.""" + shape = self.ns.shape_of(name) + for role in roles or (): + if role not in shape.roles: + self.errors.append( + f"{self.context}: {call}: {kwarg}={role} names no column of '{name}', whose columns are " + f'{list(shape.roles)}.' + ) + return False + if roles is not None and len(set(roles)) < len(roles): + self.errors.append(f'{self.context}: {call}: {kwarg}={list(roles)} names a column twice.') + return False + return True + + def _partition_walk( + self, name: str, operator: str, walked_dim: str | None, within_roles: tuple[str, ...] | None + ) -> Walk | None: + """How a partition (``shift``, ``sum_back``, ``position``) walks relation *name* along *walked_dim*. + + It walks the one key column over that dimension (a key has one column + per dimension), joins on the other key columns and groups by the value + columns *within_roles* names — every value column where the call names + none. ``None`` where the dimension is not one (already refused), the + relation has no key column over it, or ``within=`` names a column that is + not a value column. + """ + context = self.context + shape = self.ns.shape_of(name) + call = f'{operator}(by={name})' + if walked_dim is None or not self._known_roles(name, call, within_roles, 'within'): + return None + if not shape.key: + self.errors.append( + f"{context}: {call}: '{name}' declares no key, so no coordinate is in exactly one group. " + f'Declare key: on the relation, naming the column {operator} walks.' + ) + return None + over_keys = [r for r in shape.key if shape.dim(r) == walked_dim] + if not over_keys: + self.errors.append( + f"{context}: {call}: '{name}' has no key column over '{walked_dim}' — its key is " + f'{list(shape.key)} — and a partition walks a key column over the dimension it groups.' + ) + return None + if keyed := [r for r in within_roles or () if r in shape.key]: + self.errors.append( + f"{context}: {call}: within={keyed} names a key column of '{name}', and a partition groups by " + f'value columns — its value columns are {list(shape.values)}.' + ) + return None + (walked,) = over_keys + joined = tuple(r for r in shape.key if r != walked) + return Walk(shape, (walked,), shape.values if within_roles is None else within_roles, joined) + + def _default_role(self, name: str, call: str, kwarg: str, side: tuple[str, ...], what: str) -> str | None: + """The one column *side* offers, or the refusal naming what the call has to choose from.""" + if len(side) == 1: + return side[0] + shape = self.ns.shape_of(name) + if not shape.key: + self.errors.append( + f"{self.context}: {call}: '{name}' declares no key, so nothing says which column {call.split('(', maxsplit=1)[0]} " + f'walks. Name both: {kwarg}= among {list(shape.roles)} — or declare key: on the relation.' + ) + return None + self.errors.append( + f"{self.context}: {call}: '{name}' has {len(side)} {what} columns ({list(side)}), and the call has to say " + f'which {kwarg}= names.' + ) + return None - def _not_a_lookup(self, name: str, operator: str, key: str) -> str | None: - """Why *name* is not a lookup; ``None`` where it is one.""" + def _not_a_relation(self, name: str, operator: str, key: str) -> str | None: + """Why *name* is not a relation; ``None`` where it is one.""" ns, context = self.ns, self.context - if name in ns.lookups: + if name in ns.relations: return None if name in ns.dimensions: - into_here = sorted(n for n, (_, into) in ns.lookups.items() if into == name) - hint = f" Lookups into '{name}': {into_here}" if into_here else f" No lookup maps into '{name}'." + over_here = sorted(n for n, shape in ns.relations.items() if name in dict(shape.columns).values()) + hint = ( + f" Relations with a column over '{name}': {over_here}" + if over_here + else f" No relation has a column over '{name}'." + ) return ( f"{context}: {operator}({key}={name}): '{name}' is a dimension, and " - f'{key}= takes a lookup — the named map out of a dimension.\n{hint}' + f'{key}= takes a relation — the named map out of a dimension.\n{hint}' ) return ( - f'{context}: {operator}({key}={name}) does not name a lookup. ' - f'{did_you_mean(name, ns.lookups, label="Lookups")}\n' - f"Declare it under 'lookups:' — {name}: {{over: , into: }}.' + f'{context}: {operator}({key}={name}) does not name a relation. ' + f'{did_you_mean(name, ns.relations, label="Relations")}\n' + f"Declare it under 'relations:' — {name}: {{over: [], key: }}.' ) # -- where strings ----------------------------------------------------- @@ -647,7 +846,7 @@ def _child(self, node: WhereNode | UnresolvedWhereNode) -> WhereNode: return cast('WhereNode', self.where(node)) def _where_name(self, node: UnresolvedNameNode) -> WhereNode | UnresolvedWhereNode: - """A bare name: a parameter's or lookup's definedness, or a variable's existence.""" + """A bare name: a parameter's or relation's definedness, or a variable's existence.""" ns, context = self.ns, self.context kind = ns.kind(node.name) if kind is None: @@ -662,8 +861,17 @@ def _where_name(self, node: UnresolvedNameNode) -> WhereNode | UnresolvedWhereNo f'name is true at every coordinate — the mask has no effect. ' f'Remove it, or compare it: where: "{node.name} > 0".' ) - case 'lookup': - return LookupDefinedNode(node.name, ns.over_of(node.name)) + case 'relation': + shape = ns.shape_of(node.name) + dims = tuple(shape.dim(k) for k in shape.key) if shape.key else tuple(dim for _, dim in shape.columns) + if len(set(dims)) < len(dims): + self.errors.append( + f"{context}: '{node.name}' has two columns over one dimension ({list(shape.roles)}), so a " + f'bare name cannot say which the frame supplies. Compare a column: ' + f'{node.name}.{shape.values[0] if shape.values else shape.roles[-1]} == ....' + ) + return node + return RelationDefinedNode(node.name, dims) case 'variable': if node.name == self.self_variable: self.errors.append( @@ -676,7 +884,7 @@ def _where_name(self, node: UnresolvedNameNode) -> WhereNode | UnresolvedWhereNo return node def _position(self, node: UnresolvedPositionNode) -> DimensionPositionNode | UnresolvedPositionNode: - """``position(dim[, by=lookup]) i``: the name a dimension, ``by=`` a lookup over it.""" + """``position(dim[, by=relation[, within=columns]]) i``: the name a dimension, ``by=`` a relation keyed over it.""" ns, context = self.ns, self.context if node.dimension not in ns.dimensions: self.errors.append( @@ -686,44 +894,62 @@ def _position(self, node: UnresolvedPositionNode) -> DimensionPositionNode | Unr ) return node if node.by is None: - return DimensionPositionNode(node.dimension, node.op, node.position, node.by) + return DimensionPositionNode(node.dimension, node.op, node.position) call = f'position({node.dimension}, by={node.by})' - if ns.kind(node.by) != 'lookup': + if ns.kind(node.by) != 'relation': self.errors.append( f"{context}: '{call}' groups by '{node.by}', which is {_declared_as(ns, node.by)}. " - f'``by=`` takes a lookup over that dimension. ' - f'{did_you_mean(node.by, ns.lookups, label="Lookups")}' + f'``by=`` takes a relation with a key column over that dimension. ' + f'{did_you_mean(node.by, ns.relations, label="Relations")}' ) return node - over = ns.over_of(node.by) - if over != node.dimension: - self.errors.append( - f"{context}: '{call}' counts positions along '{node.dimension}' but groups by a " - f"lookup over '{over}'. No row of '{node.dimension}' carries it, so there is no " - f"position within a group to name — group by a lookup over '{node.dimension}'." - ) + walk = self._partition_walk(node.by, 'position', node.dimension, node.into) + if walk is None: return node - return DimensionPositionNode(node.dimension, node.op, node.position, node.by) + return DimensionPositionNode(node.dimension, node.op, node.position, walk) def _comparison(self, node: UnresolvedComparisonNode) -> WhereNode | UnresolvedWhereNode: - """``name literal``, or the one structural form ``lookup lookup``.""" + """``name literal``, or the one structural form ``relation relation``.""" ns, context = self.ns, self.context value = node.value - if not node.quoted and isinstance(value, str) and (rhs_kind := ns.kind(value)) is not None: - if rhs_kind == 'lookup' and ns.kind(node.name) == 'lookup': - if (refusal := _lookup_pair_error(context, node, value, ns)) is not None: - self.errors.append(refusal) - return node - return LookupPairComparisonNode(node.name, value, ns.over_of(node.name), node.op) - self.errors.append(_declared_rhs_error(context, node, value, rhs_kind)) - return node + left_name, _, left_column = node.name.partition('.') + if not node.quoted and isinstance(value, str): + right_name, _, right_column = value.partition('.') + if (rhs_kind := ns.kind(right_name)) is not None: + if rhs_kind == 'relation' and ns.kind(left_name) == 'relation': + left = self._relation_column(left_name, left_column or None, node.name, node.op) + right = self._relation_column(right_name, right_column or None, value, node.op) + if left is None or right is None: + return node + if (refusal := _relation_pair_error(context, node, value, ns, left, right)) is not None: + self.errors.append(refusal) + return node + dims = tuple(ns.shape_of(left_name).dim(k) for k in ns.shape_of(left_name).key) + return RelationPairComparisonNode(left_name, left, right_name, right, node.op, dims) + self.errors.append(_declared_rhs_error(context, node, value, rhs_kind)) + return node - kind = ns.kind(node.name) + kind = ns.kind(left_name) if kind is None: - self.errors.append(ns.unknown(node.name, context, allow_dims=True)) + self.errors.append(ns.unknown(left_name, context, allow_dims=True)) + return node + if left_column and kind != 'relation': + self.errors.append( + f"{context}: '{node.name}' reads a column of '{left_name}', which is {_declared_as(ns, left_name)}. " + f'Only a relation has columns.' + ) return node - if kind in ('parameter', 'dimension', 'lookup'): - typed = self._typed_literal(node, ns.dtypes[node.name]) + column = None + dtype: DeclaredDtype | None = None + if kind == 'relation': + column = self._relation_column(left_name, left_column or None, node.name, node.op) + if column is None: + return node + dtype = ns.dtypes[ns.shape_of(left_name).dim(column)] + elif kind in ('parameter', 'dimension'): + dtype = ns.dtypes[left_name] + if dtype is not None: + typed = self._typed_literal(node, dtype) if typed is None: return node value = typed @@ -731,19 +957,58 @@ def _comparison(self, node: UnresolvedComparisonNode) -> WhereNode | UnresolvedW match kind: case 'parameter': assert not isinstance(value, datetime.date) - return ParameterComparisonNode(node.name, node.op, value, ns.leaf_dims[node.name]) + return ParameterComparisonNode(left_name, node.op, value, ns.leaf_dims[left_name]) case 'dimension': - return DimensionComparisonNode(node.name, node.op, value) - case 'lookup': - return LookupComparisonNode(node.name, ns.over_of(node.name), node.op, value) + return DimensionComparisonNode(left_name, node.op, value) + case 'relation': + assert column is not None + shape = ns.shape_of(left_name) + return RelationComparisonNode(left_name, column, node.op, value, tuple(shape.dim(k) for k in shape.key)) case 'variable': self.errors.append( - f"{context}: where references variable '{node.name}'. A where " + f"{context}: where references variable '{left_name}'. A where " f'mask is built before variables exist — it may test parameters ' f'and dimension coordinates only.' ) return node + def _relation_column(self, name: str, column: str | None, spelling: str, op: PredicateOperator) -> str | None: + """The value column a where-comparison on relation *name* reads, or the refusal. + + A comparison reads one value per coordinate, so the relation is keyed + and the column is one the key determines; unsaid, it is the one value + column where there is exactly one. + """ + ns, context = self.ns, self.context + shape = ns.shape_of(name) + if not shape.key: + self.errors.append( + f"{context}: '{spelling}' compares a column of '{name}', which declares no key, so it has no one " + f'value per coordinate to compare. Declare key: on the relation, or test the bare name — ' + f"'{name}' — for whether a row exists." + ) + return None + if column is None: + if len(shape.values) != 1: + self.errors.append( + f"{context}: '{spelling}': '{name}' has {len(shape.values)} value columns ({list(shape.values)}), " + f'so say which the comparison reads: {name}.{shape.values[0] if shape.values else "..."}.' + ) + return None + return shape.values[0] + if column not in shape.roles: + self.errors.append( + f"{context}: '{spelling}': '{column}' is not a column of '{name}', whose columns are {list(shape.roles)}." + ) + return None + if column in shape.key: + self.errors.append( + f"{context}: '{spelling}': '{column}' is a key column of '{name}', which the frame supplies rather " + f"than reads. Compare the frame's own coordinate — {shape.dim(column)} {op} ... — or a value column." + ) + return None + return column + def _typed_literal( self, node: UnresolvedComparisonNode, dtype: DeclaredDtype ) -> float | str | datetime.date | None: @@ -868,12 +1133,12 @@ def _declared_rhs_error(context: str, node: UnresolvedComparisonNode, value: str f'{context}: {comparison} compares against variable {value!r}. ' f'A where mask is built before variables exist.' ) - if kind == 'lookup': + if kind == 'relation': return ( - f'{context}: {comparison} compares {node.name!r} against lookup {value!r}, and a ' - f'lookup is structure rather than data — every other comparison tests a name ' - f'against a literal. A lookup on the right-hand side is the one exception, and ' - f'only where the left-hand side is a lookup sharing its dimension and its target.' + f'{context}: {comparison} compares {node.name!r} against relation {value!r}, and a ' + f'relation is structure rather than data — every other comparison tests a name ' + f'against a literal. A relation on the right-hand side is the one exception, and ' + f'only where the left-hand side is a relation sharing its dimension and its target.' ) return ( f'{context}: {comparison} compares against dimension {value!r}, which the RHS reads ' @@ -883,28 +1148,32 @@ def _declared_rhs_error(context: str, node: UnresolvedComparisonNode, value: str ) -def _lookup_pair_error(context: str, node: UnresolvedComparisonNode, other: str, ns: Namespace) -> str | None: - """Why two lookups may not be compared, or ``None`` where they may. +def _relation_pair_error( + context: str, node: UnresolvedComparisonNode, other: str, ns: Namespace, left: str, right: str +) -> str | None: + """Why two relation columns may not be compared, or ``None`` where they may. - They must map out of the same dimension, or no row carries both; and into - the same one, or no value of one is ever a value of the other. Both wrong + Both relations are read at their keys, so the keys must be over the same + dimensions or no row carries both; and the two columns must be over one + dimension, or no value of one is ever a value of the other. Both wrong answers are silent, and a build's data library decides which one. """ comparison = f"'{node.name} {node.op} {other}'" - left_over, right_over = ns.over_of(node.name), ns.over_of(other) - if left_over != right_over: + left_name, right_name = node.name.partition('.')[0], other.partition('.')[0] + ls, rs = ns.shape_of(left_name), ns.shape_of(right_name) + left_keys, right_keys = {ls.dim(k) for k in ls.key}, {rs.dim(k) for k in rs.key} + if left_keys != right_keys: return ( - f'{context}: {comparison} compares lookups over different dimensions ' - f"('{left_over}' and '{right_over}') — there is no row carrying both, so the " - f'comparison has nothing to test. Two lookups may be compared only where they ' - f'map out of the same dimension.' + f'{context}: {comparison} compares relations keyed over different dimensions ' + f"('{left_name}' by {sorted(left_keys)}, '{right_name}' by {sorted(right_keys)}) — there is no row " + f'carrying both, so the comparison has nothing to test. Two relations may be compared only ' + f'where their keys are over the same dimensions.' ) - left, right = ns.into_of(node.name), ns.into_of(other) - if left != right: + if ls.dim(left) != rs.dim(right): return ( - f"{context}: {comparison} compares '{node.name}' (mapping into '{left}') with " - f"'{other}' (mapping into '{right}'). No value of one is ever a value of the other, so " - f'the predicate can only mask everything out. Two lookups may be compared only ' - f'where they map into the same dimension.' + f"{context}: {comparison} compares '{node.name}' (a column over '{ls.dim(left)}') with " + f"'{other}' (a column over '{rs.dim(right)}'). No value of one is ever a value of the other, so " + f'the predicate can only mask everything out. Two columns may be compared only ' + f'where they are over the same dimension.' ) return None diff --git a/src/math_spec/separability.py b/src/math_spec/separability.py index 64a2ac12..018a2ddd 100644 --- a/src/math_spec/separability.py +++ b/src/math_spec/separability.py @@ -81,19 +81,20 @@ def waits_on(dimension: str, label: str, name: str, kind: Literal['offset', 'par 'coupled', dimension, label, - f'sums over {dimension} — a rolling sum_back(within=n) windows, a total over the horizon does not', + f'sums over {dimension} — a rolling sum_back(window=n) windows, a total over the horizon does not', ) elif isinstance(node, GroupSum): - report( - 'coupled', - node.over, - label, - f'groups {node.over} into {", ".join(node.into)} — window that dimension instead, or cut only at the group edges', - ) + for dimension in node.over: + report( + 'coupled', + dimension, + label, + f'groups {dimension} into {", ".join(node.into)} — window that dimension instead, or cut only at the group edges', + ) elif isinstance(node, At): for dimension in node.into: - for lookup in node.coordinate: - waits_on(dimension, label, lookup, 'coordinate') + for relation in node.coordinate: + waits_on(dimension, label, relation, 'coordinate') elif isinstance(node, (Translate, Window)): dimension = node.dimension if node.wrap: @@ -106,7 +107,7 @@ def waits_on(dimension: str, label: str, name: str, kind: Literal['offset', 'par ) continue if node.partition is not None: - waits_on(dimension, label, node.partition, 'partition') + waits_on(dimension, label, node.partition.name, 'partition') if isinstance(node, Window): continue if isinstance(node.offset, str): diff --git a/src/math_spec/typesetting/format.py b/src/math_spec/typesetting/format.py index 489bee45..a0bedfb5 100644 --- a/src/math_spec/typesetting/format.py +++ b/src/math_spec/typesetting/format.py @@ -51,6 +51,7 @@ 'edge_plus', 'times', 'maps_to', + 'subset_of', 'reals', 'integers', 'binary_set', @@ -92,6 +93,7 @@ 'edge_plus': (r'\boxplus', 'plus.square'), 'times': (r'\times', 'times'), 'maps_to': (r'\to', 'arrow.r'), + 'subset_of': (r'\subseteq', 'subset.eq'), 'reals': (r'\mathbb{R}', 'RR'), 'integers': (r'\mathbb{Z}', 'ZZ'), 'binary_set': (r'\{0, 1\}', '{0, 1}'), diff --git a/src/math_spec/typesetting/walk.py b/src/math_spec/typesetting/walk.py index 4569052f..bb883a6e 100644 --- a/src/math_spec/typesetting/walk.py +++ b/src/math_spec/typesetting/walk.py @@ -26,9 +26,9 @@ EdgeNode, FunctionCallNode, KwargNode, - LookupNode, NumberNode, ParameterNode, + RelationNode, UnaryOperatorNode, UnresolvedNode, VariableNode, @@ -39,15 +39,15 @@ BooleanLiteralNode, DimensionComparisonNode, DimensionPositionNode, - LookupComparisonNode, - LookupDefinedNode, - LookupPairComparisonNode, Mask, NotNode, OrNode, ParameterComparisonNode, ParameterDefinedNode, PredicateOperator, + RelationComparisonNode, + RelationDefinedNode, + RelationPairComparisonNode, VariableDefinedNode, WhereNode, ) @@ -55,9 +55,10 @@ if TYPE_CHECKING: import datetime - from collections.abc import Iterable + from collections.abc import Iterable, Mapping - from math_spec.model import SosBlock, _ExpandedSpec + from math_spec.model import RelationBlock, SosBlock, _ExpandedSpec + from math_spec.program import Walk as RelationWalk from math_spec.typesetting.format import Format from math_spec.typesetting.symbols import Symbols @@ -273,9 +274,43 @@ def _translation(self, step: _Step) -> str: self.noticed.grouped = True return self.format.superscript(operator, step.within) - def _lookup(self, name: str, index: str) -> str: - """A coordinate map applied to an index: ``bus(g)``.""" - return self.format.apply(self.format.upright(name), index) + def _relation_read(self, walk: RelationWalk, at: Mapping[str, str], read: str) -> str: + """A relation's column *read* as a function at the columns *at* fixes: ``bus(g)``, ``zone_of(g, p)`` or ``ends.bus0(l)``. + + *at* maps each key role to the index it is read at. The function is + named after the relation alone where the key determines one column, and + after the column read otherwise. + """ + name = walk.name if len(walk.values) == 1 else f'{walk.name}.{read}' + return self.format.apply(self.format.upright(name), self.format.joined([at[k] for k in walk.key], '')) + + def _relation_member(self, walk: RelationWalk, at: Mapping[str, str]) -> str: + """A relation read as a relation: ``(g, b) ∈ gen_bus``, every column in declared order at the index *at* gives it.""" + row = self.format.parenthesise(self.format.joined([at[r] for r in walk.roles], '')) + return f'{row} {self._op("in")} {self.format.upright(walk.name)}' + + def _value_read(self, name: str, column: str, ctx: _Context) -> str: + """A keyed relation's value *column* read at the frame's own indices of its key: ``period_of(t)``.""" + lk = self.schema.relations[name] + keyed = self.format.joined([ctx.subscript(dict(lk.pairs)[k]) for k in lk.keys], '') + return self.format.apply(self._column(name, column, len(lk.values) == 1), keyed) + + def _position_group(self, node: DimensionPositionNode, ctx: _Context) -> str: + """The group a grouped position counts within: the relation's group columns read at the row's key.""" + assert node.partition is not None + walk = node.partition + keyed = self.format.joined([ctx.subscript(walk.dim(k)) for k in walk.key], '') + single = len(walk.values) == 1 + reads = [self.format.apply(self._column(walk.name, column, single), keyed) for column in walk.produced] + return self._tuple(reads) + + def _tuple(self, reads: list[str]) -> str: + """Several reads as one group label: the read alone where there is one, a bracketed tuple otherwise.""" + return reads[0] if len(reads) == 1 else self.format.parenthesise(self.format.joined(reads, '')) + + def _column(self, name: str, column: str, single: bool) -> str: + """The function a keyed relation's value *column* is: the relation's own name where it has one value column.""" + return self.format.upright(name if single else f'{name}.{column}') def _context(self, frame: Iterable[str] = ()) -> _Context: return _Context(self, bound=tuple(frame)) @@ -387,7 +422,7 @@ def _call(self, node: FunctionCallNode, ctx: _Context) -> tuple[str, int]: which, since the call does not. """ if node.name == 'shift': - dim = node.kwargs['over'] + dim = node.kwargs['along'] assert isinstance(dim, DimensionNode) step = self._step(_amount(node.kwargs['offset']), node.kwargs.get('edge')) self.noticed.policies.add(step.policy) @@ -395,7 +430,7 @@ def _call(self, node: FunctionCallNode, ctx: _Context) -> tuple[str, int]: return self._arithmetic(node.args[0], ctx.translated(dim.name, step)) if node.name == 'sum_back': - over = node.kwargs['over'] + over = node.kwargs['along'] assert isinstance(over, DimensionNode) policy = 'wrap' if isinstance(node.kwargs.get('edge'), EdgeNode) else 'plain' step = _Step(1, policy, within=self._group(node.kwargs.get('by'), over.name)) @@ -404,33 +439,36 @@ def _call(self, node: FunctionCallNode, ctx: _Context) -> tuple[str, int]: lag = f'{ctx.subscript(over.name)} {self._translation(step)} {source}' domain = ( f'{source} {self._op("in")} {self.symbols.set[over.name]} {self._op("such_that")} ' - f'0 {self._op("le")} {lag} {self._op("lt")} {self._width(node.kwargs["within"])}' + f'0 {self._op("le")} {lag} {self._op("lt")} {self._width(node.kwargs["window"])}' ) body = self._reduction_body(node.args[0], inner) return self.format.summation(domain, body), _PRECEDENCE['+'] if node.name == 'at': by = node.kwargs['by'] - assert isinstance(by, LookupNode) - for name, into in zip(by.names, by.into, strict=True): - ctx = ctx.pulled_back(into, self._lookup(name, ctx.subscript(by.dimension))) + assert isinstance(by, RelationNode) + outer = ctx + for walk in by.walks: + at = {r: outer.subscript(walk.dim(r)) for r in (*walk.produced, *walk.joined)} + for read in walk.consumed: + ctx = ctx.pulled_back(walk.dim(read), self._relation_read(walk, at, read)) return self._arithmetic(node.args[0], ctx) if (by := node.kwargs.get('by')) is not None: - assert isinstance(by, LookupNode) - dummy, inner = ctx.reducing(by.dimension) - conditions = [ - f'{self._lookup(name, dummy)} {self._op("equal")} {ctx.subscript(into)}' - for name, into in zip(by.names, by.into, strict=True) - ] + assert isinstance(by, RelationNode) + dummies: dict[str, str] = {} + inner = ctx + for d in by.dimensions: + dummies[d], inner = inner.reducing(d) + conditions = [c for walk in by.walks for c in self._grouping(walk, dummies, ctx)] domain = ( - f'{self._membership(by.dimension, dummy)} {self._op("such_that")} ' - f'{self.format.joined(conditions, self._op("and"))}' + f'{self.format.joined([self._membership(d, dummies[d]) for d in by.dimensions], "")} ' + f'{self._op("such_that")} {self.format.joined(conditions, self._op("and"))}' ) - elif (over := node.kwargs.get('over')) is not None: - assert isinstance(over, DimensionNode) - dummy, inner = ctx.reducing(over.name) - domain = self._membership(over.name, dummy) + elif (consumed := node.kwargs.get('over')) is not None: + assert isinstance(consumed, DimensionNode) + dummy, inner = ctx.reducing(consumed.name) + domain = self._membership(consumed.name, dummy) else: memberships = [] inner = ctx @@ -440,6 +478,22 @@ def _call(self, node: FunctionCallNode, ctx: _Context) -> tuple[str, int]: domain = self.format.joined(memberships, '') return self.format.summation(domain, self._reduction_body(node.args[0], inner)), _PRECEDENCE['+'] + def _grouping(self, walk: RelationWalk, dummies: Mapping[str, str], ctx: _Context) -> list[str]: + """The conditions a grouped sum's domain carries for one walk: each produced column as a function equal to its target, or one row in the relation. + + The function form holds where the key lies inside the consumed and + joined columns — one value per summand — and the relation form is the + reading that is always right. + """ + at = { + **{r: dummies[walk.dim(r)] for r in walk.consumed}, + **{r: ctx.subscript(walk.dim(r)) for r in walk.joined}, + } + targets = {r: ctx.subscript(walk.dim(r)) for r in walk.produced} + if walk.key and set(walk.key) <= set(at) and not set(walk.produced) & set(walk.key): + return [f'{self._relation_read(walk, at, r)} {self._op("equal")} {targets[r]}' for r in walk.produced] + return [self._relation_member(walk, {**at, **targets})] + def _group(self, by: ArithmeticNode | None, dim: str) -> str: """A ``by=`` as the superscript its translation operator carries. @@ -449,11 +503,13 @@ def _group(self, by: ArithmeticNode | None, dim: str) -> str: """ if by is None: return '' - assert isinstance(by, LookupNode) - return self._lookup(by.names[0], self.symbols.index[dim]) + assert isinstance(by, RelationNode) + walk = by.walks[0] + at = {r: self.symbols.index[walk.dim(r)] for r in (*walk.consumed, *walk.joined)} + return self._tuple([self._relation_read(walk, at, r) for r in walk.produced]) def _width(self, node: ArithmeticNode) -> str: - """``sum_back``'s ``within=``: a number, or a parameter's own symbol. + """``sum_back``'s ``window=``: a number, or a parameter's own symbol. Unsubscripted where it is named, as a translation's named offset is: the symbol identifies the parameter and the legend carries its dims, @@ -532,24 +588,28 @@ def _where(self, node: WhereNode, ctx: _Context) -> tuple[str, int]: ) if isinstance(node, DimensionPositionNode): - grouping = None if node.by is None else self._lookup(node.by, ctx.subscript(node.name)) + grouping = None if node.partition is None else self._position_group(node, ctx) place = self._position(ctx.subscript(node.name), grouping) ordinal = self._ordinal(node.name, node.position, grouping) return f'{place} {self._op(_PREDICATES[node.op])} {ordinal}', comparison - if isinstance(node, LookupComparisonNode): - applied = self._lookup(node.name, ctx.subscript(node.over)) + if isinstance(node, RelationComparisonNode): + applied = self._value_read(node.name, node.column, ctx) return f'{applied} {self._op(_PREDICATES[node.op])} {self._literal(node.value)}', comparison - if isinstance(node, LookupPairComparisonNode): - index = ctx.subscript(node.over) - left = self._lookup(node.name, index) - right = self._lookup(node.other, index) + if isinstance(node, RelationPairComparisonNode): + left = self._value_read(node.name, node.column, ctx) + right = self._value_read(node.other, node.other_column, ctx) return f'{left} {self._op(_PREDICATES[node.op])} {right}', comparison - if isinstance(node, LookupDefinedNode): - applied = self._lookup(node.name, ctx.subscript(node.over)) - return f'{applied} {self.format.prose(" is defined")}', comparison + if isinstance(node, RelationDefinedNode): + lk = self.schema.relations[node.name] + if lk.keys: + keyed = self.format.joined([ctx.subscript(dict(lk.pairs)[k]) for k in lk.keys], '') + applied = self.format.apply(self.format.upright(node.name), keyed) + return f'{applied} {self.format.prose(" is defined")}', comparison + row = self.format.parenthesise(self.format.joined([ctx.subscript(d) for d in lk.dims], '')) + return f'{row} {self._op("in")} {self.format.upright(node.name)}', comparison if isinstance(node, NotNode): return ( @@ -824,25 +884,30 @@ def _over(self, dims: list[str]) -> str: product = self.format.joined([self.symbols.set[d] for d in dims], self._op('times')) return f' over {self.format.math(product)}' + def _signature(self, name: str, lk: RelationBlock) -> str: + """A relation in the legend: a function from its key sets to its value sets, or a relation inside the product.""" + columns = dict(lk.pairs) + + def product(roles: Iterable[str]) -> str: + return self.format.joined([self.symbols.set[columns[r]] for r in roles], self._op('times')) + + if lk.keys: + return f'{self.format.upright(name)}: {product(lk.keys)} {self._op("maps_to")} {product(lk.values)}' + return f'{self.format.upright(name)} {self._op("subset_of")} {product(lk.roles)}' + def _coords(self, dim: str, noticed: Noticed) -> str: - """The dimension's carried structure: each lookup as the map it is (``bus_of: G ↦ B``). + """The dimension's carried structure: each relation with a column over it, as the map or relation it is. The dtype is named only where an equation compared the index against a number, the one place "position 3" and "the coordinate 3" are both readings of a line. """ - targeted = self.schema.lookups_of(dim) + carried = self.schema.relations_of(dim) clauses = [] if dim in noticed.numeric_coordinates: clauses.append(f' ({self.format.mono(self.schema.dimensions[dim].dtype)} coordinates)') - if targeted: - maps = self.format.joined( - [ - f'{self.format.upright(c)}: {self.symbols.set[dim]} {self._op("maps_to")} {self.symbols.set[target]}' - for c, target in targeted.items() - ], - '', - ) + if carried: + maps = self.format.joined([self._signature(c, lk) for c, lk in carried.items()], '') clauses.append(f' with {self.format.math(maps)}') return ''.join(clauses) @@ -885,11 +950,11 @@ def translation_notes(self, noticed: Noticed) -> list[str]: f'at that boundary is built and carries {self.format.math("v")} rather than being dropped.' ) if noticed.grouped: - applied = self._lookup('lookup', 't') + applied = self.format.apply(self.format.upright('relation'), 't') counted = self.format.math(f't {self.format.superscript(self._op("cyclic_minus"), applied)} k') note = ( - f'{counted} denotes a translation counted inside the group a lookup puts {self.format.math("t")} ' - f'in ({self.format.mono("shift(by=lookup)")}), so a term never crosses out of its own group.' + f'{counted} denotes a translation counted inside the group a relation puts {self.format.math("t")} ' + f'in ({self.format.mono("shift(by=relation)")}), so a term never crosses out of its own group.' ) if 'edge' in noticed.policies: both = self.format.superscript(self.format.subscript(self._op('edge_minus'), ['v']), applied) @@ -914,11 +979,11 @@ def position_notes(self, noticed: Noticed) -> list[str]: f'labels and {place} against positions.' ) if 'grouped' in noticed.positions: - applied = self._lookup('lookup', 't') + applied = self.format.apply(self.format.upright('relation'), 't') grouped = self.format.math(self.format.apply(self.format.subscript(self._op('position'), [applied]), 't')) group = self.format.math(self.format.subscript(self.format.script('T'), [applied])) notes.append( - f'{grouped} counts within the group a lookup puts {self.format.math("t")} in: the subscript names ' + f'{grouped} counts within the group a relation puts {self.format.math("t")} in: the subscript names ' f'the map, {group} is the group it lands in, and that group has a first position of its own.' ) if 'from_end' in noticed.positions: diff --git a/src/math_spec/validation.py b/src/math_spec/validation.py index f8316600..e4017957 100644 --- a/src/math_spec/validation.py +++ b/src/math_spec/validation.py @@ -343,21 +343,24 @@ def _check_template_names( for arg in node.args: _check_template_names(arg, context, ns, formals, errors) for kwarg, value in node.kwargs.items(): - match builtin.kind_of(kwarg) if builtin else 'value': + with_relation = builtin is not None and any(k in node.kwargs for k in builtin.relation_kwargs) + match builtin.kind_of(kwarg, with_relation=with_relation) if builtin else 'value': case 'dimension': if isinstance(value, NameNode) and value.name not in ns.dimensions | formals: errors.append( f'{context}: {node.name}({kwarg}={value.name}) does not name a ' f'declared dimension or a formal of this macro.' ) - case 'lookup': + case 'relation': errors.extend( - f'{context}: {node.name}({kwarg}={one}) does not name a lookup or a formal of this macro.' + f'{context}: {node.name}({kwarg}={one}) does not name a relation or a formal of this macro.' for one in names_in(value) - if one not in formals and ns.kind(one) != 'lookup' + if one not in formals and ns.kind(one) != 'relation' ) case 'value': _check_template_names(value, context, ns, formals, errors) + case 'role': + pass case 'edge': pass # a keyword or a number: nothing in it to name return diff --git a/tests/fixtures.py b/tests/fixtures.py index 593c573f..f389bc84 100644 --- a/tests/fixtures.py +++ b/tests/fixtures.py @@ -36,13 +36,13 @@ 'objective': {'sense': 'minimize', 'expression': 'sum(p * cost)'}, } -#: Two dimensions, a lookup between them, a numeric, a scalar, a boolean and a +#: Two dimensions, a relation between them, a numeric, a scalar, a boolean and a #: label parameter, a variable on each frame — one declaration of every kind #: a rule can name, and no objective, so a test adds what it judges. `p` and `r` #: share no dimension, which is what a rule about *different* dims needs. SMALL_MODEL: dict[str, Any] = { 'dimensions': {'g': {'dtype': 'str'}, 'h': {'dtype': 'str'}}, - 'lookups': {'lk': {'over': 'g', 'into': 'h'}}, + 'relations': {'lk': {'columns': ['g', 'h'], 'key': 'g'}}, 'parameters': { 'c': {'dims': ['g']}, 'k': {'dims': []}, diff --git a/tests/fixtures/every_program_node.yaml b/tests/fixtures/every_program_node.yaml index fd26b5bf..9b668b70 100644 --- a/tests/fixtures/every_program_node.yaml +++ b/tests/fixtures/every_program_node.yaml @@ -7,8 +7,8 @@ dimensions: t: { dtype: int } g: { dtype: str } zone: { dtype: str } -lookups: - zone_of: { over: g, into: zone } +relations: + zone_of: { columns: [g, zone], key: g } parameters: cost: { dims: [g] } load: { dims: [zone] } @@ -23,7 +23,7 @@ expressions: dims: [t, g] cases: opening: { when: "position(t) == 0", expression: 0 } - otherwise: shift(p, over=t, offset=1, edge='wrap') + otherwise: shift(p, along=t, offset=1, edge='wrap') shadow: dual(grouped) # a Dual leaf, the one node only an entry the math never reads lowers to constraints: regional: @@ -43,10 +43,10 @@ constraints: expression: "p - at(q, by=zone_of) <= 0" translated: dims: [t, g] - expression: "p - shift(p, over=t, offset=lead, edge='wrap') <= 0" + expression: "p - shift(p, along=t, offset=lead, edge='wrap') <= 0" windowed: dims: [t, g] - expression: "sum_back(p, over=t, within=width) <= 10" + expression: "sum_back(p, along=t, window=width) <= 10" objective: sense: minimize expression: "sum(p * cost)" diff --git a/tests/test_advice.py b/tests/test_advice.py index f57a84fc..be7b5707 100644 --- a/tests/test_advice.py +++ b/tests/test_advice.py @@ -29,8 +29,8 @@ objective={'sense': 'minimize', 'expression': 'sum(p * c)'}, ) -#: The same with the lookup gone, so nothing reaches ``h`` at all. -UNREACHED = override(TARGET_ONLY, lookups={}) +#: The same with the relation gone, so nothing reaches ``h`` at all. +UNREACHED = override(TARGET_ONLY, relations={}) def test_a_dimension_nothing_reaches_is_named(): @@ -42,7 +42,7 @@ def test_a_dimension_nothing_reaches_is_named(): @pytest.mark.parametrize( 'patch', [ - pytest.param({}, id='targeted-by-a-lookup'), + pytest.param({}, id='targeted-by-a-relation'), pytest.param( {'constraints': {'cap': {'dims': ['h'], 'expression': 'sum(p, by=lk) <= k'}}}, id='grouping-into-it', @@ -52,7 +52,7 @@ def test_a_dimension_nothing_reaches_is_named(): ) def test_a_dimension_something_reaches_is_in_use(patch): assert not advice(override(TARGET_ONLY, **patch)), ( - 'a dimension a lookup targets, a declaration indexes or a grouping lands on is in use' + 'a dimension a relation targets, a declaration indexes or a grouping lands on is in use' ) diff --git a/tests/test_boundedness.py b/tests/test_boundedness.py index 07ae4dba..db36028f 100644 --- a/tests/test_boundedness.py +++ b/tests/test_boundedness.py @@ -45,7 +45,7 @@ def _notes(**patch) -> list[str]: pytest.param({'objective.expression': 'sum(-3 * v, over=g)'}, 'upper', id='a-negative-literal-flips-it'), pytest.param({'objective.expression': 'sum(v / 2, over=g)'}, 'lower', id='a-literal-divisor-keeps-it'), pytest.param( - {'objective.expression': 'sum(shift(v, over=g, offset=1), over=g)'}, + {'objective.expression': 'sum(shift(v, along=g, offset=1), over=g)'}, 'lower', id='an-operator-argument-keeps-it', ), @@ -92,9 +92,9 @@ def test_nothing_is_claimed_where_the_file_does_not_decide_it(patch): #: built-in arrives with a case of its own. THROUGH_EACH_OPERATOR = { 'sum': {'objective.expression': 'sum(v, over=g)'}, - 'shift': {'objective.expression': 'sum(shift(v, over=g, offset=1), over=g)'}, - 'sum_back': {'objective.expression': 'sum(sum_back(v, over=g, within=2), over=g)'}, - # `at` reads onto the lookup's source, so the variable it drives is on `h` + 'shift': {'objective.expression': 'sum(shift(v, along=g, offset=1), over=g)'}, + 'sum_back': {'objective.expression': 'sum(sum_back(v, along=g, window=2), over=g)'}, + # `at` reads onto the relation's source, so the variable it drives is on `h` 'at': {'variables.u': {'dims': ['h']}, 'objective.expression': 'sum(at(u, by=lk), over=g)'}, } diff --git a/tests/test_degree.py b/tests/test_degree.py index 05a4fca7..84d6e2c9 100644 --- a/tests/test_degree.py +++ b/tests/test_degree.py @@ -85,7 +85,7 @@ def test_the_objective_takes_degree_two(text): pytest.param('(p * q) * (p * q)', 'this product is degree 4', id='a-quartic'), pytest.param('sum(p, over=g) * sum(q, over=g)', 'outer product', id='two-reductions'), pytest.param('(p + q) * (p + q)', 'outer product', id='two-sums-of-variables'), - pytest.param('sum_back(p, over=g, within=1) * (p - q)', 'outer product', id='a-window-against-a-difference'), + pytest.param('sum_back(p, along=g, window=1) * (p - q)', 'outer product', id='a-window-against-a-difference'), ], ) def test_degree_two_is_one_term_against_one_term_and_no_higher(text, fragment): diff --git a/tests/test_dimensions.py b/tests/test_dimensions.py index a14c87fb..d1ddc19d 100644 --- a/tests/test_dimensions.py +++ b/tests/test_dimensions.py @@ -11,7 +11,7 @@ import pytest from math_spec.dimensions import DimensionError, _check_where_dims, dims_of -from math_spec.program import LookupPairComparisonNode, Mask +from math_spec.program import Mask, RelationPairComparisonNode from math_spec.resolution import Namespace, expression_of, where_of from math_spec.validation import to_spec from tests.fixtures import override, schema_of @@ -30,15 +30,23 @@ 'snapshot': {'dtype': 'int'}, 'generator': {'dtype': 'str'}, 'bus': {'dtype': 'str'}, + 'zone': {'dtype': 'str'}, }, - 'lookups': { - 'gen_bus': {'over': 'generator', 'into': 'bus'}, - 'snap_bus': {'over': 'snapshot', 'into': 'bus'}, + 'relations': { + 'gen_bus': {'columns': ['generator', 'bus'], 'key': 'generator'}, + 'snap_bus': {'columns': ['snapshot', 'bus'], 'key': 'snapshot'}, + 'gen_zone': {'columns': ['generator', 'snapshot', 'zone'], 'key': ['generator', 'snapshot']}, + 'rep_of': {'columns': {'snapshot': 'snapshot', 'rep': 'snapshot'}, 'key': 'snapshot'}, + 'gen_bz': {'columns': ['generator', 'bus', 'zone'], 'key': 'generator'}, + 'pair': {'columns': {'g': 'generator', 'b0': 'bus', 'b1': 'bus'}, 'key': 'g'}, }, 'parameters': { 'p_max': {'dims': ['generator']}, 'cost': {'dims': ['generator']}, 'load': {'dims': ['snapshot', 'bus']}, + 'zone_cap': {'dims': ['zone']}, + 'zone_load': {'dims': ['snapshot', 'zone']}, + 'bz': {'dims': ['bus', 'zone']}, 'spinup': {'dims': ['generator'], 'dtype': 'int'}, 'horizon': {'dims': ['snapshot'], 'dtype': 'int'}, 'bus_lead': {'dims': ['bus'], 'dtype': 'int'}, @@ -86,20 +94,95 @@ def namespace() -> Namespace: ('sum(p, over=generator)', {'snapshot'}), ('sum(p * cost, over=generator)', {'snapshot'}), ('sum(p, by=gen_bus)', {'snapshot', 'bus'}), - ("shift(p, over=snapshot, offset=1, edge='wrap')", {'snapshot', 'generator'}), - ("shift(p, over=snapshot, offset=spinup, edge='wrap')", {'snapshot', 'generator'}), - ('sum_back(p, over=snapshot, within=spinup)', {'snapshot', 'generator'}), + ("shift(p, along=snapshot, offset=1, edge='wrap')", {'snapshot', 'generator'}), + ("shift(p, along=snapshot, offset=spinup, edge='wrap')", {'snapshot', 'generator'}), + ('sum_back(p, along=snapshot, window=spinup)', {'snapshot', 'generator'}), pytest.param( - "shift(p, over=snapshot, offset=bus_lead, edge='wrap', by=snap_bus)", + "shift(p, along=snapshot, offset=bus_lead, edge='wrap', by=snap_bus)", {'snapshot', 'generator'}, id='a-by-makes-an-offset-over-another-dim-readable-one-lag-per-group', ), pytest.param( - 'sum_back(p, over=snapshot, within=bus_lead, by=snap_bus)', + 'sum_back(p, along=snapshot, window=bus_lead, by=snap_bus)', {'snapshot', 'generator'}, id='a-by-makes-a-width-over-another-dim-readable-one-window-per-group', ), pytest.param('p + 1', {'snapshot', 'generator'}, id='a-scalar-broadcasts'), + pytest.param( + 'sum(load * p, by=gen_bus)', + {'snapshot', 'bus'}, + id='a-produced-dim-the-operand-already-carries-is-joined-on-so-the-walk-is-a-masked-sum', + ), + pytest.param( + 'sum(p, by=gen_zone, over=generator)', + {'snapshot', 'zone'}, + id='a-two-key-relation-consumes-the-key-it-walks-and-keeps-the-other', + ), + pytest.param( + 'sum(p, by=gen_zone, over=snapshot)', + {'generator', 'zone'}, + id='the-same-table-walked-along-its-other-key', + ), + pytest.param( + 'at(zone_load, by=gen_zone, into=generator)', + {'snapshot', 'generator'}, + id='its-pullback-keeps-the-joined-key-too', + ), + pytest.param( + "shift(p, along=generator, offset=1, edge='wrap', by=gen_zone)", + {'snapshot', 'generator'}, + id='a-partition-along-one-key-joined-on-the-other', + ), + pytest.param( + "shift(p, along=generator, offset=1, edge='wrap', by=gen_bz, within=bus)", + {'snapshot', 'generator'}, + id='a-partition-grouped-by-one-value-column-of-a-two-value-table', + ), + pytest.param( + 'sum_back(p, along=generator, window=2, by=gen_bz, within=[bus, zone])', + {'snapshot', 'generator'}, + id='a-window-grouped-by-both-value-columns-named', + ), + pytest.param( + "shift(p, along=generator, offset=1, edge='wrap', by=pair)", + {'snapshot', 'generator'}, + id='a-partition-grouped-by-two-columns-over-one-dimension-lands-nothing', + ), + pytest.param( + 'sum(p, by=gen_bus, over=generator)', {'snapshot', 'bus'}, id='the-dot-is-legal-on-a-one-key-relation' + ), + pytest.param( + 'sum(p, by=gen_bz, into=[bus, zone])', + {'snapshot', 'bus', 'zone'}, + id='a-to-list-lands-on-a-product-from-one-table', + ), + pytest.param( + 'at(bz, by=gen_bz, over=[bus, zone])', + {'generator'}, + id='a-from-list-reads-two-value-columns-at-once', + ), + pytest.param( + 'sum(p, by=gen_zone, over=[generator, snapshot])', + {'zone'}, + id='a-from-list-consumes-two-key-columns-at-once', + ), + pytest.param( + 'sum(p, by=gen_bz, into=bus)', + {'snapshot', 'bus'}, + id='a-value-column-not-walked-is-not-read', + ), + pytest.param( + 'sum(p, by=gen_bz, over=generator, into=bus)', + {'snapshot', 'bus'}, + id='by-and-over-compose', + ), + pytest.param('sum(p, by=rep_of)', {'snapshot', 'generator'}, id='a-map-into-its-own-dimension-keeps-the-frame'), + pytest.param('at(p, by=rep_of)', {'snapshot', 'generator'}, id='and-so-does-its-pullback'), + pytest.param( + "shift(p, along=snapshot, offset=1, edge='wrap', by=rep_of)", + {'snapshot', 'generator'}, + id='a-partition-into-its-own-dimension', + ), ], ) def test_dim_inference(expr, expected): @@ -130,7 +213,7 @@ def test_a_bare_name_reaches_the_variable_a_dual_the_same_named_constraint(): pytest.param( 'sum(p, over=bus)', r'sum\(over=bus\) but the expression has dims', - id='sum-over-an-absent-dim-is-an-error-not-a-noop', + id='sum-consuming-an-absent-dim-is-an-error-not-a-noop', ), pytest.param( 'sum(sum(p))', @@ -139,54 +222,64 @@ def test_a_bare_name_reaches_the_variable_a_dual_the_same_named_constraint(): ), pytest.param( 'sum(load, by=gen_bus)', - r"sum\(by=gen_bus\) consumes 'generator', the dim it maps out of", + r"sum\(by=gen_bus\) consumes \['generator'\], the dims it walks from", id='sum-requires-the-grouped-dim', ), pytest.param( - 'sum(load * p, by=gen_bus)', - 'already carries', - id='sum-into-a-dim-the-operand-already-carries', - ), - pytest.param( - "shift(cost, over=snapshot, offset=1, edge='wrap')", - r'shift\(over=snapshot\) but the expression has dims', + "shift(cost, along=snapshot, offset=1, edge='wrap')", + r'shift\(along=snapshot\) but the expression has dims', id='shift-requires-the-dim', ), pytest.param( - "shift(p, over=snapshot, offset=cost, edge='wrap')", + "shift(p, along=snapshot, offset=cost, edge='wrap')", r'declared dtype: float', id='a-named-offset-is-integral-58', ), pytest.param( - "shift(p, over=snapshot, offset=horizon, edge='wrap')", + "shift(p, along=snapshot, offset=horizon, edge='wrap')", r'varies along the axis it walks is a permutation rather than a lag', id='a-named-offset-does-not-span-the-axis-it-walks', ), pytest.param( - 'sum_back(p, over=snapshot, within=cost)', + 'sum_back(p, along=snapshot, window=cost)', r'declared dtype: float', id='a-named-width-is-integral', ), pytest.param( - 'sum_back(p, over=snapshot, within=horizon)', + 'sum_back(p, along=snapshot, window=horizon)', r'no longer "the last n"', id='a-named-width-does-not-span-the-summed-axis', ), pytest.param( - "shift(p, over=snapshot, offset=-spinup, edge='wrap')", + "shift(p, along=snapshot, offset=-spinup, edge='wrap')", r'negates a named offset', id='a-named-offset-is-not-negated-at-the-call-62', ), pytest.param( - 'sum_back(p, over=snapshot, within=-spinup)', + 'sum_back(p, along=snapshot, window=-spinup)', r'which way a window reaches is the operator', id='a-named-width-has-no-direction-to-negate', ), pytest.param( - "shift(p, over=snapshot, offset=bus_lead, edge='wrap')", + "shift(p, along=snapshot, offset=bus_lead, edge='wrap')", r"varies over \['bus'\], which that coordinate does not carry", id='a-named-offset-is-read-where-the-expression-has-a-coordinate', ), + pytest.param( + 'sum(cost, by=gen_zone, over=generator)', + r"sum\(by=gen_zone\) joins on \['snapshot'\]", + id='a-grouped-sum-needs-the-keys-it-joins-on', + ), + pytest.param( + 'at(zone_cap, by=gen_zone, into=generator)', + r"at\(by=gen_zone\) joins on \['snapshot'\]", + id='a-pullback-needs-the-keys-it-joins-on', + ), + pytest.param( + "shift(cost, along=generator, offset=1, edge='wrap', by=gen_zone)", + r"by=gen_zone\) joins on \['snapshot'\]", + id='a-partition-needs-the-keys-it-joins-on', + ), ], ) def test_an_ill_dimensioned_expression_is_rejected(expr, match): @@ -268,27 +361,27 @@ def _refused(self, expression: str) -> str: ('expression', 'fragment'), [ pytest.param( - 'p <= shift(cap, over=g, offset=1)', + 'p <= shift(cap, along=g, offset=1)', 'leaves vacated positions with no value', id='a-shift-over-data-with-no-edge', ), pytest.param( - 'p <= shift(p, over=t, offset=lead)', + 'p <= shift(p, along=t, offset=lead)', 'per-entity offset cannot say yet', id='a-named-offset-with-no-edge', ), pytest.param( - 'p <= shift(p, over=t, offset=1, edge=2)', + 'p <= shift(p, along=t, offset=1, edge=2)', 'only fill=0 is representable', id='a-nonzero-edge-over-a-variable', ), pytest.param( - 'p <= sum_back(p, over=t, within=2, edge=0)', + 'p <= sum_back(p, along=t, window=2, edge=0)', "takes 'wrap' or nothing", id='a-numeric-edge-on-a-window', ), pytest.param( - "p <= shift(p, over=t, offset=1.5, edge='wrap')", + "p <= shift(p, along=t, offset=1.5, edge='wrap')", 'must be a whole number', id='a-fractional-amount', ), @@ -304,7 +397,7 @@ def test_a_literal_width_below_one_is_refused_by_to_spec(self, width): The sign was stripped before the `at least 1` comparison, so `-2` was tested as `2` and reached lowering, which asserted (#222). """ - assert 'at least 1' in self._refused(f'p <= sum_back(p, over=t, within={width})') + assert 'at least 1' in self._refused(f'p <= sum_back(p, along=t, window={width})') def test_a_zero_step_vacates_nothing_and_needs_no_edge(self): """`shift(x, offset=0)` reaches every coordinate from itself. @@ -313,7 +406,7 @@ def test_a_zero_step_vacates_nothing_and_needs_no_edge(self): literal zero vacates none, so there is nothing for an `edge=` to answer for. A *named* offset may be zero in the data and is not known here. """ - to_spec(override(self.BASE, **{'constraints.k.expression': 'p <= shift(cap, over=g, offset=0)'})) + to_spec(override(self.BASE, **{'constraints.k.expression': 'p <= shift(cap, along=g, offset=0)'})) # --------------------------------------------------------------------------- @@ -327,7 +420,21 @@ def test_a_zero_step_vacates_nothing_and_needs_no_edge(self): pytest.param('p_max > 0', {'generator'}, id='a-parameter-through-its-own-dims'), pytest.param('snapshot == 0', {'snapshot'}, id='a-dimension-through-itself'), pytest.param('position(snapshot) == 0', {'snapshot'}, id='a-position-through-the-axis-it-counts'), - pytest.param('snap_bus == "b1"', {'snapshot'}, id='a-lookup-through-the-dim-it-maps-out-of'), + pytest.param('snap_bus == "b1"', {'snapshot'}, id='a-relation-through-the-dim-it-maps-out-of'), + pytest.param('gen_zone == "z1"', {'generator', 'snapshot'}, id='a-two-key-relation-through-both-keys'), + pytest.param('gen_zone', {'generator', 'snapshot'}, id='a-bare-two-key-relation-the-same'), + pytest.param('rep_of == 3', {'snapshot'}, id='a-map-into-its-own-dimension-through-its-key'), + pytest.param('position(snapshot, by=rep_of) == 0', {'snapshot'}, id='a-position-within-a-representative'), + pytest.param( + 'position(generator, by=gen_zone) == 0', + {'generator', 'snapshot'}, + id='a-position-within-a-group-of-a-two-key-relation-reads-both-keys', + ), + pytest.param( + 'position(generator, by=gen_bz, within=zone) == 0', + {'generator'}, + id='a-position-within-one-named-value-column-reads-the-key', + ), pytest.param('p_max > 0 AND snapshot == 0', {'generator', 'snapshot'}, id='a-conjunction-reads-both-sides'), pytest.param('NOT p_max > 0', {'generator'}, id='a-negation-reads-what-it-negates'), pytest.param('False', set(), id='a-literal-reads-nothing'), @@ -368,8 +475,8 @@ def test_the_frame_check_and_the_reading_walk_the_same_leaves(namespace): pytest.param('p_max > 0', {'p_max'}, id='a-parameter-comparison-names-the-parameter'), pytest.param('spinup', {'spinup'}, id='a-parameter-bare-names-the-parameter'), pytest.param('p', {'p'}, id='a-variable-bare-names-the-variable'), - pytest.param('snap_bus == "b1"', {'snap_bus'}, id='a-lookup-comparison-names-the-lookup'), - pytest.param('gen_bus', {'gen_bus'}, id='a-lookup-bare-names-the-lookup'), + pytest.param('snap_bus == "b1"', {'snap_bus'}, id='a-relation-comparison-names-the-relation'), + pytest.param('gen_bus', {'gen_bus'}, id='a-relation-bare-names-the-relation'), pytest.param('snapshot == 0', set(), id='a-dimension-names-nothing-it-is-a-coordinate'), pytest.param('position(snapshot) == 0', set(), id='a-position-names-nothing'), pytest.param('p_max > 0 AND snapshot == 0', {'p_max'}, id='a-conjunction-drops-the-dimension-side'), @@ -386,12 +493,12 @@ def test_a_predicate_names_the_declarations_its_leaves_test(namespace, predicate assert where.names_read == expected -def test_names_read_takes_both_sides_of_a_lookup_pair(): +def test_names_read_takes_both_sides_of_a_relation_pair(): """The one leaf that names two declarations — two maps compared on the dimension they share. - BASE has one lookup per dimension, so the pair is built directly rather than + BASE has one relation per dimension, so the pair is built directly rather than resolved from a predicate string. """ - where = LookupPairComparisonNode('from_bus', 'to_bus', 'line', '!=') + where = RelationPairComparisonNode('from_bus', 'bus', 'to_bus', 'bus', '!=', ('line',)) - assert Mask(where).names_read == {'from_bus', 'to_bus'}, 'a lookup pair names both maps it compares' + assert Mask(where).names_read == {'from_bus', 'to_bus'}, 'a relation pair names both maps it compares' diff --git a/tests/test_exclusivity.py b/tests/test_exclusivity.py index b91cad3f..3b62be64 100644 --- a/tests/test_exclusivity.py +++ b/tests/test_exclusivity.py @@ -35,7 +35,7 @@ 'storage': {}, 'period': {'dtype': 'int'}, }, - 'lookups': {'period_of': {'over': 'snapshot', 'into': 'period'}}, + 'relations': {'period_of': {'columns': ['snapshot', 'period'], 'key': 'snapshot'}}, 'parameters': { 'cyclic': {'dims': ['storage'], 'dtype': 'bool'}, 'committable': {'dims': ['storage'], 'dtype': 'bool'}, diff --git a/tests/test_expansion.py b/tests/test_expansion.py index 2b8bc944..b79bc6e6 100644 --- a/tests/test_expansion.py +++ b/tests/test_expansion.py @@ -212,19 +212,19 @@ def test_macro_collisions_rejected(patch, match): id='a-comparison-in-a-template', ), pytest.param( - {'lag': {'args': ['x'], 'template': 'shift(x, over=snapshot, offset=nope)'}}, + {'lag': {'args': ['x'], 'template': 'shift(x, along=snapshot, offset=nope)'}}, r"Macro 'lag'.*'nope' not found", id='a-typo-in-an-amount', ), pytest.param( {'grouped': {'args': ['x'], 'template': 'sum(x, by=nope)'}}, - r"Macro 'grouped'.*sum\(by=nope\) does not name a lookup", - id='a-typo-in-a-lookup-kwarg', + r"Macro 'grouped'.*sum\(by=nope\) does not name a relation", + id='a-typo-in-a-relation-kwarg', ), pytest.param( {'grouped': {'args': ['x'], 'template': 'sum(x, by=[nope, also])'}}, - r"Macro 'grouped'.*sum\(by=nope\) does not name a lookup", - id='a-typo-in-a-lookup-list', + r"Macro 'grouped'.*sum\(by=nope\) does not name a relation", + id='a-typo-in-a-relation-list', ), ], ) diff --git a/tests/test_lowering.py b/tests/test_lowering.py index f743319f..7ea96867 100644 --- a/tests/test_lowering.py +++ b/tests/test_lowering.py @@ -36,7 +36,6 @@ ExpressionNode, Footprint, GroupSum, - LookupDeclaration, Mask, Multiply, Negate, @@ -48,9 +47,11 @@ Power, Program, Region, + RelationDeclaration, Sum, Translate, Variable, + Walk, Window, children, divisor_parameters, @@ -81,15 +82,22 @@ 'constraints': {'c': {'dims': [], 'expression': 'sum(p, over=g) >= 1'}}, } -#: `fixtures.SMALL_MODEL` plus a second lookup and a per-entity +#: `lk` and `lk2` as `sum` walks them: key consumed, value produced, nothing joined. +LK = RelationDeclaration('lk', (('g', 'g'), ('h', 'h')), ('g',)) +LK2 = RelationDeclaration('lk2', (('g', 'g'), ('z', 'z')), ('g',)) +LK_WALK = Walk(LK, ('g',), ('h',), ()) +LK2_WALK = Walk(LK2, ('g',), ('z',), ()) +AT_BUS = RelationDeclaration('at_bus', (('g', 'g'), ('bus', 'bus')), ('g',)) + +#: `fixtures.SMALL_MODEL` plus a second relation and a per-entity #: offset. Which node a construct becomes is mostly a claim about the dim it #: consumes and the dim it lands on, and stating that needs a third dimension -#: and two lookups over one of them. +#: and two relations over one of them. SHAPES_MODEL = override( SMALL_MODEL, **{ 'dimensions.z': {'dtype': 'str'}, - 'lookups.lk2': {'over': 'g', 'into': 'z'}, + 'relations.lk2': {'columns': ['g', 'z'], 'key': 'g'}, 'parameters.lead': {'dims': ['g'], 'dtype': 'int'}, }, ) @@ -163,7 +171,7 @@ def test_a_file_with_no_objective_lowers_to_no_sense(): def test_a_literal_amount_resolves_to_one_signed_number(dispatch_schema): """`offset=-1` parses as a unary minus over `1`; after resolution it is `-1`, for every reader alike.""" ns = Namespace.of(dispatch_schema) - node = expression_of('shift(p, over=snapshot, offset=-1, edge=+2)', dispatch_schema, ns, 't') + node = expression_of('shift(p, along=snapshot, offset=-1, edge=+2)', dispatch_schema, ns, 't') assert isinstance(node, FunctionCallNode) assert (node.kwargs['offset'], node.kwargs['edge']) == (NumberNode(-1.0), NumberNode(2.0)) @@ -274,7 +282,7 @@ def test_a_lowered_where_is_a_mask_that_answers_from_its_root(dispatch_program): {'g', 'h'}, 2, 2, - id='a-lookup-is-read-at-the-dim-it-maps-out-of-and-a-position-at-its-own', + id='a-relation-is-read-at-the-dim-it-maps-out-of-and-a-position-at-its-own', ), pytest.param('p', 'k > 0', set(), 1, 1, id='a-scalar-parameter-is-read-at-no-coordinate'), pytest.param( @@ -404,58 +412,71 @@ def test_a_power_lowers_to_a_node_of_its_own(dispatch_schema): pytest.param('sum(q, over=h)', Sum(Variable('q'), ('h',)), id='an-over-consumes-the-dim-it-names'), pytest.param( 'sum(p, by=lk)', - GroupSum(Variable('p'), over='g', coordinate=('lk',), into=('h',)), + GroupSum(Variable('p'), walks=(LK_WALK,)), id='a-grouped-sum-names-the-dim-it-consumes-and-the-one-it-lands-on', ), pytest.param( 'sum(p, by=[lk])', - GroupSum(Variable('p'), over='g', coordinate=('lk',), into=('h',)), + GroupSum(Variable('p'), walks=(LK_WALK,)), id='a-one-element-list-is-the-plain-form', ), pytest.param( 'sum(p, by=[lk, lk2])', - GroupSum(Variable('p'), over='g', coordinate=('lk', 'lk2'), into=('h', 'z')), + GroupSum(Variable('p'), walks=(LK_WALK, LK2_WALK)), id='two-coordinates-are-one-grouping-with-paired-tuples', ), pytest.param( 'at(r, by=lk)', - At(Variable('r'), over='g', coordinate=('lk',), into=('h',)), + At(Variable('r'), walks=(Walk(LK, ('h',), ('g',), ()),)), id='a-pullback-walks-the-same-table-back', ), pytest.param( - "shift(p, over=g, offset=1, edge='wrap')", + "shift(p, along=g, offset=1, edge='wrap')", Translate(Variable('p'), 'g', offset=1, wrap=True, fill=None), id='a-wrapping-translation-fills-nothing', ), pytest.param( - 'shift(p, over=g, offset=-2, edge=0)', + 'shift(p, along=g, offset=-2, edge=0)', Translate(Variable('p'), 'g', offset=-2, wrap=False, fill=0.0), id='a-lead-is-a-negative-offset-and-the-edge-is-what-it-fills-with', ), pytest.param( - 'shift(p, over=g, offset=lead, edge=0)', + 'shift(p, along=g, offset=lead, edge=0)', Translate(Variable('p'), 'g', offset='lead', wrap=False, fill=0.0), id='a-named-offset-crosses-as-the-parameter-name', ), pytest.param( - 'shift(p, over=g, offset=1, by=lk, edge=0)', - Translate(Variable('p'), 'g', offset=1, wrap=False, fill=0.0, partition='lk'), - id='a-translation-stops-at-the-edges-of-the-lookup-it-names', + 'shift(p, along=g, offset=1, by=lk, edge=0)', + Translate( + Variable('p'), + 'g', + offset=1, + wrap=False, + fill=0.0, + partition=Walk(LK, ('g',), ('h',), ()), + ), + id='a-translation-stops-at-the-edges-of-the-relation-it-names', ), pytest.param( - 'sum_back(p, over=g, within=3)', + 'sum_back(p, along=g, window=3)', Window(Variable('p'), 'g', width=3, wrap=False), id='a-window-is-one-node-rather-than-a-fold-of-translations', ), pytest.param( - 'sum_back(p, over=g, within=k)', + 'sum_back(p, along=g, window=k)', Window(Variable('p'), 'g', width='k', wrap=False), id='a-named-width-crosses-as-the-parameter-name', ), pytest.param( - 'sum_back(p, over=g, within=2, by=lk)', - Window(Variable('p'), 'g', width=2, wrap=False, partition='lk'), - id='a-window-stops-at-the-edges-of-the-lookup-it-names', + 'sum_back(p, along=g, window=2, by=lk)', + Window( + Variable('p'), + 'g', + width=2, + wrap=False, + partition=Walk(LK, ('g',), ('h',), ()), + ), + id='a-window-stops-at-the-edges-of-the-relation-it-names', ), ], ) @@ -465,6 +486,66 @@ def test_a_construct_lowers_to_its_node(shapes_schema, expression, expected): assert lowered == expected, 'the whole frozen node, so no field is asserted by omission' +def test_a_relation_lowers_with_the_walk_each_call_takes(): + """Every node reading a relation carries its columns, its key and the walk, so a consumer joins on the right columns.""" + program = to_program( + { + 'dimensions': {'snapshot': {'dtype': 'int'}, 'generator': {}, 'zone': {}}, + 'relations': {'zone_of': {'columns': ['generator', 'snapshot', 'zone'], 'key': ['generator', 'snapshot']}}, + 'parameters': {'price': {'dims': ['snapshot', 'zone']}}, + 'variables': { + 'p': {'dims': ['snapshot', 'generator'], 'where': "zone_of == 'A' AND zone_of"}, + 'first': {'dims': ['snapshot', 'generator'], 'where': 'position(generator, by=zone_of) == 0'}, + }, + 'constraints': { + 'zonal': {'dims': ['snapshot', 'zone'], 'expression': 'sum(p, by=zone_of, over=generator) <= 1'}, + 'priced': { + 'dims': ['snapshot', 'generator'], + 'expression': 'p <= at(price, by=zone_of, into=generator)', + }, + 'history': { + 'dims': ['generator', 'zone'], + 'expression': 'sum(p, by=zone_of, over=snapshot) <= 1', + }, + }, + } + ) + + columns = (('generator', 'generator'), ('snapshot', 'snapshot'), ('zone', 'zone')) + declared = RelationDeclaration('zone_of', columns, ('generator', 'snapshot')) + assert program.dimension('generator').relations == (declared,), 'the relation sits under its first column' + assert program.dimension('zone').relations == (declared,), 'and under its last' + assert program.relations == {'zone_of': declared}, 'and once in the program' + zonal = program.constraints['zonal'].lhs + assert zonal == GroupSum(Variable('p'), walks=(Walk(declared, ('generator',), ('zone',), ('snapshot',)),)), ( + 'a grouped sum names the column it consumes, the one it produces and the one it joins on' + ) + assert isinstance(zonal, GroupSum) + assert (zonal.over, zonal.into, zonal.coordinate) == (('generator',), ('zone',), ('zone_of',)), ( + 'the dims a consumer reads are read off the walk' + ) + assert program.constraints['history'].lhs == GroupSum( + Variable('p'), walks=(Walk(declared, ('snapshot',), ('zone',), ('generator',)),) + ), 'the same table walked from its other key column' + priced = program.constraints['priced'].rhs + assert priced == At(Parameter('price'), walks=(Walk(declared, ('zone',), ('generator',), ('snapshot',)),)), ( + 'and its adjoint consumes the value column and produces the key column' + ) + assert isinstance(priced, At) + assert (priced.over, priced.into) == (('generator',), ('zone',)), ( + 'an at produces the fine dims and consumes the coarse' + ) + p_where = program.variable('p').where + assert p_where is not None + assert [(type(a).__name__, a.dims) for a in p_where.atoms] == [ + ('RelationComparisonNode', ('generator', 'snapshot')), + ('RelationDefinedNode', ('generator', 'snapshot')), + ], 'a comparison and an existence are both read at the key of a keyed relation' + first_where = program.variable('first').where + assert first_where is not None + assert first_where.dims == {'generator', 'snapshot'}, 'a position within a group is read at every key column' + + def test_a_binary_variable_lowers_to_a_binary_domain(): program = to_program(schema_of(DISPATCH_YAML, **{'variables.p.domain': 'binary', 'variables.p.bounds': {}})) assert program.variable('p').domain == 'binary' @@ -473,7 +554,8 @@ def test_a_binary_variable_lowers_to_a_binary_domain(): def test_a_divisor_under_a_pullback_is_still_named(): """`children` has to descend through every node, or a refusal loses its name.""" quotient = Divide(Variable('x'), Parameter('rate')) - pulled = At(quotient, over='flow', coordinate=('component',), into=('component',)) + component_of = RelationDeclaration('component_of', (('flow', 'flow'), ('component', 'component')), ('flow',)) + pulled = At(quotient, walks=(Walk(component_of, ('component',), ('flow',), ()),)) assert divisor_parameters(pulled) == frozenset({'rate'}), 'the walk descends through `At`' assert divisor_parameters(Sum(pulled, ('flow',))) == frozenset({'rate'}), 'and through a `Sum` over it' @@ -513,8 +595,8 @@ def test_a_quotient_is_found_whole_so_its_two_halves_stay_paired(): Power(Parameter('c'), Constant(2.0)): 'one-to-one', Divide(Variable('p'), Parameter('c')): 'one-to-one', Sum(Variable('p'), ('g',)): 'many-to-one', - GroupSum(Variable('p'), over='g', coordinate=('at_bus',), into=('bus',)): 'many-to-one', - At(Variable('p'), over='g', coordinate=('at_bus',), into=('bus',)): 'one-to-one', + GroupSum(Variable('p'), walks=(Walk(AT_BUS, ('g',), ('bus',), ()),)): 'many-to-one', + At(Variable('p'), walks=(Walk(AT_BUS, ('bus',), ('g',), ()),)): 'one-to-one', Translate(Variable('p'), 't', offset=1, wrap=False, fill=0.0): 'one-to-one', Window(Variable('p'), 't', width=2, wrap=False): 'one-to-many', Cases((Region(Mask(ParameterDefinedNode('c', ('g',))), Variable('p')),)): 'one-to-one', @@ -535,7 +617,7 @@ def test_a_node_answers_its_fan_in(node, expected): assert fan_in(node) == expected -def test_a_lookup_names_the_dimension_its_values_label(): +def test_a_relation_names_the_dimension_its_values_label(): """Five sites asked this and each walked for it; the plan answers it once.""" program = Program( parameters={}, @@ -543,19 +625,20 @@ def test_a_lookup_names_the_dimension_its_values_label(): constraints={}, objective=None, dimensions={ - 'snapshot': DimensionDeclaration((LookupDeclaration('season_of', 'season'),)), - 'generator': DimensionDeclaration((LookupDeclaration('at_bus', 'bus'),)), + 'snapshot': DimensionDeclaration( + (RelationDeclaration('season_of', (('snapshot', 'snapshot'), ('season', 'season')), ('snapshot',)),) + ), + 'generator': DimensionDeclaration( + (RelationDeclaration('at_bus', (('generator', 'generator'), ('bus', 'bus')), ('generator',)),) + ), }, ) - assert program.dimension('snapshot').targets == {'season_of': 'season'}, ( + assert [lk.name for lk in program.dimension('snapshot').relations] == ['season_of'], ( 'one dimension names its own maps and no other dimension' ) - assert program.dimension('generator').targets == {'at_bus': 'bus'}, 'and the same for the second' - assert [(d, lk.name) for d, lk in program.lookups] == [ - ('snapshot', 'season_of'), - ('generator', 'at_bus'), - ], 'every map with the dimension it is over, in declaration order' + assert program.dimension('snapshot').relations[0].values == ('season',), 'and the map says what its key determines' + assert list(program.relations) == ['season_of', 'at_bus'], 'every map once, by name, in declaration order' def test_an_unknown_dimension_is_a_near_miss_rather_than_an_empty_declaration(): @@ -686,7 +769,7 @@ def test_a_dimension_carries_the_dtype_its_labels_are_checked_against(): 'always_on': {'when': 'not committable', 'expression': 1}, 'boundary': {'when': 'committable and position(t) == 0', 'expression': 'initial'}, }, - 'otherwise': 'shift(status, over=t, offset=1)', + 'otherwise': 'shift(status, along=t, offset=1)', } }, 'constraints': {'no_restart': {'dims': ['t', 'g'], 'expression': 'status - previous <= 1'}}, @@ -787,7 +870,9 @@ def test_a_cased_expression_is_readable_by_the_name_the_file_wrote(): ), pytest.param({}, False, id='nothing-reads-it'), pytest.param( - {'expressions.ratio': 'spend / sum(p, over=g)'}, False, id='only-an-entry-the-math-never-reads-inlines-it' + {'expressions.ratio': 'spend / sum(p, over=g)'}, + False, + id='only-an-entry-the-math-never-reads-inlines-it', ), ], ) diff --git a/tests/test_piecewise.py b/tests/test_piecewise.py index 57e3cfd6..48d71537 100644 --- a/tests/test_piecewise.py +++ b/tests/test_piecewise.py @@ -344,8 +344,8 @@ def test_a_link_reading_a_dual_entry_is_refused(): @pytest.mark.parametrize( ('activity', 'match'), [ - pytest.param('at(u_unit, by=unit_of)', 'is not a declared variable', id='a-pullback-through-a-lookup'), - pytest.param('shift(u, over=snapshot, offset=1)', 'is not a declared variable', id='a-shifted-gate'), + pytest.param('at(u_unit, by=unit_of)', 'is not a declared variable', id='a-pullback-through-a-relation'), + pytest.param('shift(u, along=snapshot, offset=1)', 'is not a declared variable', id='a-shifted-gate'), pytest.param('u * 2', 'is not a declared variable', id='an-arithmetic-gate'), ], ) diff --git a/tests/test_pypsa_references.py b/tests/test_pypsa_references.py index fbe86c64..3e1fcabd 100644 --- a/tests/test_pypsa_references.py +++ b/tests/test_pypsa_references.py @@ -114,7 +114,7 @@ def test_a_file_of_its_own_shares_its_declarations_with_the_base(page: str): """A keyword file restates the base surface; a shared name keeps its PyPSA name and its dtype, or it has drifted.""" own = SPECS[page] drifted = [] - for section in ('parameters', 'lookups', 'variables', 'constraints'): + for section in ('parameters', 'relations', 'variables', 'constraints'): theirs, ours = getattr(BASE, section), getattr(own, section) for name in set(theirs) & set(ours): if (theirs[name].description or '').split(' — ')[0] != (ours[name].description or '').split(' — ')[0]: diff --git a/tests/test_separability.py b/tests/test_separability.py index 33b95204..3cceaddc 100644 --- a/tests/test_separability.py +++ b/tests/test_separability.py @@ -25,7 +25,7 @@ BASE: dict[str, Any] = { 'dimensions': {'h': {'dtype': 'int'}, 'u': {'dtype': 'str'}, 'zone': {'dtype': 'str'}, 'day': {'dtype': 'int'}}, - 'lookups': {'zone_of': {'over': 'u', 'into': 'zone'}, 'day_of': {'over': 'h', 'into': 'day'}}, + 'relations': {'zone_of': {'columns': ['u', 'zone'], 'key': 'u'}, 'day_of': {'columns': ['h', 'day'], 'key': 'h'}}, 'parameters': { 'cost': {'dims': ['u']}, 'budget': {'dims': []}, @@ -50,12 +50,12 @@ def _rows(expression: str, *, dims: list[str] | None = None, **block: Any) -> di [ pytest.param(_rows('p >= 0'), 0, id='pointwise-needs-no-overlap'), pytest.param( - _rows('p >= shift(p, over=h, offset=1, edge=0)'), 0, id='a-shift-behind-is-the-edge-and-asks-nothing' + _rows('p >= shift(p, along=h, offset=1, edge=0)'), 0, id='a-shift-behind-is-the-edge-and-asks-nothing' ), - pytest.param(_rows('p >= shift(p, over=h, offset=-2, edge=0)'), 2, id='a-negative-shift-reads-ahead'), - pytest.param(_rows('sum_back(p, over=h, within=4) >= 0'), 0, id='a-trailing-window-reads-behind-only'), - pytest.param(_rows('sum_back(p, over=h, within=width) >= 0'), 0, id='and-so-does-one-of-a-width-from-data'), - pytest.param(_rows('p >= shift(p, over=u, offset=-1, edge=0)'), 0, id='a-shift-along-another-axis-is-nothing'), + pytest.param(_rows('p >= shift(p, along=h, offset=-2, edge=0)'), 2, id='a-negative-shift-reads-ahead'), + pytest.param(_rows('sum_back(p, along=h, window=4) >= 0'), 0, id='a-trailing-window-reads-behind-only'), + pytest.param(_rows('sum_back(p, along=h, window=width) >= 0'), 0, id='and-so-does-one-of-a-width-from-data'), + pytest.param(_rows('p >= shift(p, along=u, offset=-1, edge=0)'), 0, id='a-shift-along-another-axis-is-nothing'), ], ) def test_a_separable_model_reports_the_lookahead_a_window_needs(patch, ahead): @@ -68,7 +68,7 @@ def test_a_separable_model_reports_the_lookahead_a_window_needs(patch, ahead): ('patch', 'fragment'), [ pytest.param(_rows('sum(p, over=h) <= budget', dims=['u']), 'sums over h', id='a-budget-over-the-horizon'), - pytest.param(_rows("p >= shift(p, over=h, offset=1, edge='wrap')"), 'wraps around h', id='a-cyclic-shift'), + pytest.param(_rows("p >= shift(p, along=h, offset=1, edge='wrap')"), 'wraps around h', id='a-cyclic-shift'), ], ) def test_a_model_the_axis_ties_together_names_what_ties_it(patch, fragment): @@ -82,19 +82,19 @@ def test_a_model_the_axis_ties_together_names_what_ties_it(patch, fragment): ('patch', 'reach'), [ pytest.param( - _rows('p >= shift(p, over=h, offset=1, by=day_of, edge=0)'), + _rows('p >= shift(p, along=h, offset=1, by=day_of, edge=0)'), Reach("constraint 'k'", 'day_of', 'partition'), id='a-shift-inside-groups', ), pytest.param( - _rows('p >= shift(p, over=h, offset=width, edge=0)'), + _rows('p >= shift(p, along=h, offset=width, edge=0)'), Reach("constraint 'k'", 'width', 'offset'), id='an-offset-from-data', ), ], ) def test_a_reach_only_data_can_say_names_what_says_it(patch, reach): - """The verdict names the parameter or lookup and what it stands as, rather + """The verdict names the parameter or relation and what it stands as, rather than refusing the model, so a driver holding the data knows what to read and `resolved` knows how to fold it.""" verdict = _verdict(**patch) @@ -106,10 +106,12 @@ def test_a_reach_only_data_can_say_names_what_says_it(patch, reach): @pytest.mark.parametrize( ('patch', 'least', 'ahead'), [ - pytest.param(_rows('p >= shift(p, over=h, offset=width, edge=0)'), 1, 0, id='offsets-behind-ask-nothing'), - pytest.param(_rows('p >= shift(p, over=h, offset=width, edge=0)'), -3, 3, id='offsets-ahead-read-by-the-least'), + pytest.param(_rows('p >= shift(p, along=h, offset=width, edge=0)'), 1, 0, id='offsets-behind-ask-nothing'), pytest.param( - _rows('shift(p, over=h, offset=width, edge=0) + shift(p, over=h, offset=-1, edge=0) + p >= 0'), + _rows('p >= shift(p, along=h, offset=width, edge=0)'), -3, 3, id='offsets-ahead-read-by-the-least' + ), + pytest.param( + _rows('shift(p, along=h, offset=width, edge=0) + shift(p, along=h, offset=-1, edge=0) + p >= 0'), -3, 3, id='a-folded-value-widens-what-the-model-reads-on-its-own', @@ -124,40 +126,40 @@ def test_a_named_reach_resolves_to_the_lookahead_its_values_need(patch, least, a assert verdict.ahead == ahead, 'the value decides the lookahead, by its sign' -def test_resolving_keeps_the_static_reach_and_what_a_lookup_decides(): +def test_resolving_keeps_the_static_reach_and_what_a_relation_decides(): verdict = _verdict( constraints={ - 'fixed': {'dims': ['h', 'u'], 'expression': 'p >= shift(p, over=h, offset=-2, edge=0)'}, - 'named': {'dims': ['h', 'u'], 'expression': 'p >= shift(p, over=h, offset=width, edge=0)'}, - 'grouped': {'dims': ['h', 'u'], 'expression': 'p >= shift(p, over=h, offset=1, by=day_of, edge=0)'}, + 'fixed': {'dims': ['h', 'u'], 'expression': 'p >= shift(p, along=h, offset=-2, edge=0)'}, + 'named': {'dims': ['h', 'u'], 'expression': 'p >= shift(p, along=h, offset=width, edge=0)'}, + 'grouped': {'dims': ['h', 'u'], 'expression': 'p >= shift(p, along=h, offset=1, by=day_of, edge=0)'}, } ).resolved({'width': -1}) assert verdict.ahead == 2, 'a folded value never narrows what the model reads on its own' assert verdict.undecided == (Reach("constraint 'grouped'", 'day_of', 'partition'),), ( - 'a reach a lookup decides is not a value and stays undecided' + 'a reach a relation decides is not a value and stays undecided' ) assert not verdict.windowable, 'so the axis is still not windowable' def test_resolving_a_name_nothing_waits_on_is_refused(): with pytest.raises(KeyError, match="'depth' is not a parameter an undecided reach along 'h' waits on"): - _verdict(**_rows('p >= shift(p, over=h, offset=width, edge=0)')).resolved({'depth': 0}) + _verdict(**_rows('p >= shift(p, along=h, offset=width, edge=0)')).resolved({'depth': 0}) -def test_a_read_through_a_lookup_is_undecided_on_the_axis_it_reads(): - """`at(cap, by=zone_of)` reads `zone` at whatever coordinate the lookup - chooses, so how far that reaches along `zone` is the lookup's data to say.""" +def test_a_read_through_a_relation_is_undecided_on_the_axis_it_reads(): + """`at(cap, by=zone_of)` reads `zone` at whatever coordinate the relation + chooses, so how far that reaches along `zone` is the relation's data to say.""" verdict = _verdict('zone', **_rows('p - at(cap, by=zone_of) <= 0')) - assert not verdict.windowable and not verdict.coupled, 'undecided until the lookup binds' + assert not verdict.windowable and not verdict.coupled, 'undecided until the relation binds' assert verdict.undecided == (Reach("constraint 'k'", 'zone_of', 'coordinate'),), ( - 'the report names the lookup a driver has to read' + 'the report names the relation a driver has to read' ) def test_a_coupling_names_the_change_that_would_lift_it(): coupled = _verdict(**_rows('sum(p, over=h) <= budget', dims=['u'])).coupled["constraint 'k'"] - assert 'sum_back(within=n)' in coupled, 'a horizon total becomes a rolling one' - wrapped = _verdict(**_rows("p >= shift(p, over=h, offset=1, edge='wrap')")).coupled["constraint 'k'"] + assert 'sum_back(window=n)' in coupled, 'a horizon total becomes a rolling one' + wrapped = _verdict(**_rows("p >= shift(p, along=h, offset=1, edge='wrap')")).coupled["constraint 'k'"] assert 'position(h) == 0' in wrapped, 'a wrap becomes an opening-state seed' @@ -189,7 +191,7 @@ def test_a_position_inside_a_cased_region_is_found(): 'prev': { 'dims': ['h', 'u'], 'cases': {'opening': {'when': 'position(h) == 0', 'expression': 0}}, - 'otherwise': 'shift(p, over=h, offset=1, edge=0)', + 'otherwise': 'shift(p, along=h, offset=1, edge=0)', } }, **_rows('p - prev <= 1'), @@ -200,8 +202,8 @@ def test_a_position_inside_a_cased_region_is_found(): def test_the_lookahead_is_the_widest_reach_of_any_block(): verdict = _verdict( constraints={ - 'near': {'dims': ['h', 'u'], 'expression': 'p >= shift(p, over=h, offset=-1, edge=0)'}, - 'far': {'dims': ['h', 'u'], 'expression': 'p >= shift(p, over=h, offset=-5, edge=0)'}, + 'near': {'dims': ['h', 'u'], 'expression': 'p >= shift(p, along=h, offset=-1, edge=0)'}, + 'far': {'dims': ['h', 'u'], 'expression': 'p >= shift(p, along=h, offset=-5, edge=0)'}, } ) assert verdict.ahead == 5, 'one window must see past its last row as far as any block reads' diff --git a/tests/test_validation.py b/tests/test_validation.py index 2fcc591a..cca1d1c8 100644 --- a/tests/test_validation.py +++ b/tests/test_validation.py @@ -85,8 +85,8 @@ class TestValidateExpressions: ), pytest.param( {'constraints': {'cap': {'dims': ['g'], 'where': 'lkk', 'expression': 'p <= c'}}}, - ("'lkk' not found", "Lookups: ['lk']"), - id='a-mistyped-lookup-in-a-where-lists-the-lookups', + ("'lkk' not found", "Relations: ['lk']"), + id='a-mistyped-relation-in-a-where-lists-the-relations', ), ], ) @@ -183,7 +183,7 @@ def test_an_unreferenced_nonlinear_entry_loads_and_is_reported(self): def _kwarg_model(expression: str, dims: list[str] | None = None) -> dict[str, Any]: - """A model over (snapshot, generator), with `zone` a lookup into `bus`. + """A model over (snapshot, generator), with `zone` a relation into `bus`. `zone` deliberately targets a dim `p` does *not* carry: grouping into one it already has needs that dim twice, which is its own error. @@ -196,7 +196,7 @@ def _kwarg_model(expression: str, dims: list[str] | None = None) -> dict[str, An 'bus': {'dtype': 'str'}, 'generator': {'dtype': 'str'}, }, - 'lookups': {'zone': {'over': 'generator', 'into': 'bus'}}, + 'relations': {'zone': {'columns': ['generator', 'bus'], 'key': 'generator'}}, 'parameters': {'load': {'dims': ['snapshot']}}, 'variables': {'p': {'dims': ['snapshot', 'generator']}}, 'constraints': {'c': {'dims': ['snapshot'] if dims is None else dims, 'expression': expression}}, @@ -284,16 +284,16 @@ class TestDimensionKwargs: pytest.param('sum(p, over=snapshto) == load', ('silent no-op', 'sum(over=snapshto)'), id='sum-over-typo'), pytest.param( 'sum(p, by=bus) == load', - ("'bus' is a dimension, and by= takes a lookup",), + ("'bus' is a dimension, and by= takes a relation",), id='by-names-a-dimension', ), pytest.param( 'sum(p, by=zne) == load', - ('does not name a lookup', "Did you mean 'zone'?"), - id='by-lookup-typo', + ('does not name a relation', "Did you mean 'zone'?"), + id='by-relation-typo', ), pytest.param( - 'shift(p, over=snapshto, offset=1) == load', + 'shift(p, along=snapshto, offset=1) == load', ('does not name a declared dimension',), id='shift-over-typo', ), @@ -310,11 +310,11 @@ def test_a_dim_kwarg_typo_is_rejected(self, expression, fragments): pytest.param('sum(p, over=generator) == load', ['snapshot'], id='a-sum'), pytest.param('sum(p, by=zone) == load', ['snapshot', 'bus'], id='a-grouped-sum'), pytest.param( - "shift(p, over=snapshot, offset=1, edge='wrap') == load", + "shift(p, along=snapshot, offset=1, edge='wrap') == load", ['snapshot', 'generator'], id='a-wrapping-shift', ), - pytest.param('shift(p, over=snapshot, offset=1) == load', ['snapshot', 'generator'], id='a-bare-shift'), + pytest.param('shift(p, along=snapshot, offset=1) == load', ['snapshot', 'generator'], id='a-bare-shift'), ], ) def test_declared_dimensions_still_pass(self, expression, dims): @@ -324,7 +324,11 @@ def test_macro_formals_are_not_mistaken_for_dimensions(self): """A formal in a dim position is legal inside the template body.""" _schema( macros={ - 'ws': {'args': ['array', 'weights'], 'kwargs': ['over'], 'template': 'sum(array * weights, over=over)'} + 'ws': { + 'args': ['array', 'weights'], + 'kwargs': ['over'], + 'template': 'sum(array * weights, over=over)', + } }, objective={'sense': 'minimize', 'expression': 'ws(p, c, over=g)'}, ) @@ -400,7 +404,7 @@ def test_a_named_amount_keeps_its_own_sentence(self): _schema( **{ 'parameters.lag': {'dims': [], 'dtype': 'str'}, - 'objective': {'expression': "sum(shift(p, over=g, offset=lag, edge='wrap'))"}, + 'objective': {'expression': "sum(shift(p, along=g, offset=lag, edge='wrap'))"}, } ) @@ -423,13 +427,13 @@ def test_the_version_gates_no_behaviour(self): assert _schema().model_dump(exclude={'version'}) == _schema(version=0).model_dump(exclude={'version'}) -#: `position(dim)` needs a lookup over *that* dimension, so one over it and one into it. +#: `position(dim)` needs a relation over *that* dimension, so one over it and one into it. POSITION_SCHEMA = to_spec( { 'dimensions': {'snapshot': {'dtype': 'int'}, 'period': {'dtype': 'int'}}, - 'lookups': { - 'period_of': {'over': 'snapshot', 'into': 'period'}, - 'starts_at': {'over': 'period', 'into': 'snapshot'}, + 'relations': { + 'period_of': {'columns': ['snapshot', 'period'], 'key': 'snapshot'}, + 'starts_at': {'columns': ['period', 'snapshot'], 'key': 'period'}, }, 'parameters': {'load': {'dims': ['snapshot']}}, 'variables': {'p': {'dims': ['snapshot']}}, @@ -440,8 +444,8 @@ def test_the_version_gates_no_behaviour(self): class TestPositionResolves: """`position(dim)` — the conversion #32 put on the left-hand side. - A `by=` has to be a lookup over *that* dimension, and that is the whole - test: a lookup over anything else carries no row for a position to be a + A `by=` has to be a relation over *that* dimension, and that is the whole + test: a relation over anything else carries no row for a position to be a position in. """ @@ -460,17 +464,17 @@ def test_it_resolves(self, mask: str, position: int, by: str | None): assert isinstance(node, DimensionPositionNode) assert node.name == 'snapshot' assert node.position == position - assert node.by == by + assert (node.partition.name if node.partition is not None else None) == by @pytest.mark.parametrize( ('mask', 'fragments'), [ ('position(load) == 0', ["counts along a dimension's coordinates", "'load' is a parameter"]), ('position(nope) == 0', ["'nope' is not declared"]), - ('position(snapshot, by=load) == 0', ['groups by', '``by=`` takes a lookup']), - ('position(snapshot, by=starts_at) == 0', ["along 'snapshot'", "lookup over 'period'"]), + ('position(snapshot, by=load) == 0', ['groups by', '``by=`` takes a relation']), + ('position(snapshot, by=starts_at) == 0', ["no key column over 'snapshot'", "its key is ['period']"]), ], - ids=['a parameter', 'undeclared', 'by= is not a lookup', 'by= is over another dim'], + ids=['a parameter', 'undeclared', 'by= is not a relation', 'by= is over another dim'], ) def test_it_refuses(self, mask: str, fragments: list[str]): with pytest.raises(LanguageError) as excinfo: @@ -566,25 +570,172 @@ class TestRulesDecidedWithoutData: id='sos-big-m-infinite', ), pytest.param( - {'lookups.tag': {'over': 'g', 'dtype': 'str'}}, - ("unknown key 'dtype' in a lookup declaration. Valid keys: description, into, over.",), - id='lookup-with-a-dtype-of-its-own', + {'relations.tag': {'columns': 'g', 'dtype': 'str'}}, + ("unknown key 'dtype' in a relation declaration. Valid keys: columns, description, key.",), + id='relation-with-a-dtype-of-its-own', + ), + pytest.param({'relations.tag': {'columns': 'g'}}, ('has 1 column(s)',), id='relation-with-one-column'), + pytest.param( + {'relations.lk.columns': 'z'}, ("references undeclared dimension 'z'",), id='relation-over-undeclared' + ), + pytest.param( + {'relations.lk.key': 'z'}, + ("has key column 'z', which is not one of its columns",), + id='relation-key-not-a-column', + ), + pytest.param( + {'relations.lk.key': ['g', 'h']}, ('has every column in its key',), id='relation-keyed-by-every-column' + ), + pytest.param( + {'relations.pair': {'columns': {'g0': 'g', 'g1': 'g', 'h': 'h'}, 'key': ['g0', 'g1']}}, + ("has two key columns over 'g' (['g0', 'g1'])", 'no frame carries a dimension twice'), + id='relation-keyed-twice-over-one-dimension', + ), + pytest.param( + {'relations.odd': {'columns': {'h': 'g', 'x': 'h'}, 'key': 'h'}}, + ("names column 'h' after dimension 'h', but the column is over 'g'",), + id='relation-column-named-after-a-dimension-it-is-not-over', + ), + pytest.param( + {'relations.lk.columns': ['g', 'z']}, + ("references undeclared dimension 'z'",), + id='relation-key-undeclared', + ), + pytest.param( + {'relations.lk.columns': ['g', 'g']}, + ("names dimension 'g' twice under 'columns:'", 'columns: {g0: g, g1: g}'), + id='relation-naming-a-dim-twice-without-roles', + ), + pytest.param({'relations.lk.columns': []}, ('has 0 column(s)',), id='relation-with-no-columns'), + pytest.param( + { + 'dimensions.z': {}, + 'relations.lk': {'columns': ['g', 'z', 'h'], 'key': ['g', 'z']}, + 'variables.q.dims': ['g', 'h', 'z'], + 'objective': {'expression': 'sum(sum(q, by=lk))'}, + }, + ("'lk' has 2 key columns (['g', 'z']), and the call has to say which over= names",), + id='by-a-two-key-relation-without-from', + ), + pytest.param( + {'objective': {'expression': 'sum(sum(p, by=lk, over=z))'}}, + ("over=z names no column of 'lk', whose columns are ['g', 'h']",), + id='from-a-column-the-relation-lacks', + ), + pytest.param( + {'objective': {'expression': 'sum(sum(p, by=lk, over=h, into=h))'}}, + ("over= and into= both name ['h']",), + id='from-and-to-the-same-column', + ), + pytest.param( + { + 'dimensions.z': {}, + 'relations.lz': {'columns': ['g', 'h', 'z'], 'key': 'g'}, + 'objective': {'expression': 'sum(sum(p, by=lz, into=[h, h]))'}, + }, + ("into=['h', 'h'] names a column twice",), + id='a-to-list-naming-a-column-twice', + ), + pytest.param( + { + 'dimensions.z': {}, + 'relations.lz': {'columns': ['g', 'h', 'z'], 'key': 'g'}, + 'objective': {'expression': 'sum(sum(p, by=lz, over=[g, h], into=h))'}, + }, + ("over= and into= both name ['h']",), + id='a-from-list-overlapping-to', + ), + pytest.param( + { + 'relations.lz': {'columns': {'g': 'g', 'h0': 'h', 'h1': 'h'}, 'key': 'g'}, + 'objective': {'expression': 'sum(sum(p, by=lz, over=[h0, h1], into=g))'}, + }, + ("over=['h0', 'h1'] names two columns over ['h'], and the operand carries each dimension once",), + id='a-from-list-naming-two-columns-over-one-dimension', + ), + pytest.param( + {'objective': {'expression': 'sum(shift(p, along=g, offset=1, edge=0, by=lk, over=g))'}}, + ( + "shift() expects shift(, along=, offset=[, edge='wrap'|]" + '[, by=[, within=]])', + ), + id='a-partition-takes-no-from', ), - pytest.param({'lookups.tag': {'over': 'g'}}, ('lookups.tag.into: Field required',), id='lookup-no-into'), pytest.param( - {'lookups.lk.over': 'z'}, ("references undeclared dimension 'z'",), id='lookup-over-undeclared' + { + 'dimensions.z': {}, + 'relations.lz': {'columns': ['g', 'h', 'z'], 'key': 'g'}, + 'objective': {'expression': 'sum(shift(p, along=g, offset=1, edge=0, by=lz, within=g))'}, + }, + ("within=['g'] names a key column of 'lz', and a partition groups by value columns",), + id='a-partition-grouped-within-a-key-column', ), - pytest.param({'lookups.lk.into': 'z'}, ("targets undeclared dimension 'z'",), id='lookup-into-undeclared'), - pytest.param({'lookups.lk.into': 'g'}, ("maps 'g' into itself",), id='lookup-into-itself'), pytest.param( - {'lookups.g': {'over': 'h', 'into': 'g'}}, - ("Lookup 'g' collides with the dimension",), - id='lookup-named-after-a-dimension', + {'variables.q.where': 'position(g, by=lk, within=z) == 0'}, + ("within=z names no column of 'lk', whose columns are ['g', 'h']",), + id='position-within-a-column-the-relation-lacks', ), pytest.param( - {'lookups.lk.values': {'a': 'x'}}, - ("unknown key 'values' in a lookup declaration", 'Valid keys'), - id='a-lookup-declaring-its-map', + {'objective': {'expression': 'sum(sum(p, into=g))'}}, + ('names a column of a relation, and no by= names the relation',), + id='into-without-by', + ), + pytest.param( + {'relations.rel': {'columns': ['g', 'h']}, 'objective': {'expression': 'sum(sum(p, by=rel))'}}, + ("'rel' declares no key, so nothing says which column sum walks",), + id='a-bare-relation-needs-both-ends-named', + ), + pytest.param( + { + 'relations.rel': {'columns': ['g', 'h']}, + 'objective': {'expression': 'sum(at(r, by=rel, over=h, into=g))'}, + }, + ("at reads one value per coordinate, and 'rel' is not single-valued",), + id='at-through-a-bare-relation', + ), + pytest.param( + {'objective': {'expression': 'sum(sum(q, by=lk, over=h, into=g))'}}, + ("this sum walks to the key ['g']", 'that is a read, which is', 'at(..., by=lk'), + id='a-sum-that-walks-to-the-key-is-a-read', + ), + pytest.param( + { + 'relations.rel': {'columns': ['g', 'h']}, + 'objective': {'expression': 'sum(shift(p, along=g, offset=1, edge=0, by=rel))'}, + }, + ("'rel' declares no key, so no coordinate is in exactly one group",), + id='a-partition-through-a-bare-relation', + ), + pytest.param( + {'relations.rel': {'columns': ['g', 'h']}, 'variables.q.where': "rel == 'x'"}, + ("compares a column of 'rel', which declares no key",), + id='where-compares-a-bare-relation', + ), + pytest.param( + {'variables.q.where': "lk.g == 'x'"}, + ( + "'g' is a key column of 'lk', which the frame supplies rather than reads", + 'g == ...', + ), + id='where-compares-a-key-column', + ), + pytest.param( + { + 'relations.pair': {'columns': {'g0': 'g', 'g1': 'g'}}, + 'variables.q.where': 'pair', + }, + ('has two columns over one dimension', 'Compare a column'), + id='where-bare-name-of-a-relation-with-two-columns-over-one-dim', + ), + pytest.param( + {'relations.g': {'columns': ['h', 'g'], 'key': 'h'}}, + ("Relation 'g' collides with the dimension",), + id='relation-named-after-a-dimension', + ), + pytest.param( + {'relations.lk.values': {'a': 'x'}}, + ("unknown key 'values' in a relation declaration", 'Valid keys'), + id='a-relation-declaring-its-map', ), pytest.param( {'variables.p.absence': 'zero'}, ('absence: zero needs a `where:`',), id='absence-without-a-mask' @@ -670,50 +821,77 @@ class TestRulesDecidedWithoutData: ), pytest.param( {'objective': {'expression': 'sum(lk + p, over=g)'}}, - ("'lk' is a lookup, and a lookup is structure",), - id='a-lookup-as-a-value', + ("'lk' is a relation, and a relation is structure",), + id='a-relation-as-a-value', ), pytest.param( - {'objective': {'expression': 'sum(shift(p, over=g, offset=1, edge=wrap), over=g)'}}, + {'objective': {'expression': 'sum(shift(p, along=g, offset=1, edge=wrap), over=g)'}}, ('is a bare name where a keyword belongs',), id='a-bare-edge-keyword', ), pytest.param( - {'objective': {'expression': "sum(shift(p, over=g, offset=1, edge='foo'), over=g)"}}, + {'objective': {'expression': "sum(shift(p, along=g, offset=1, edge='foo'), over=g)"}}, ("edge='foo') is not an edge policy",), id='an-edge-policy-that-is-not-one', ), - pytest.param( - {'objective': {'expression': 'sum(p, over=g, by=lk)'}}, ('at most one of',), id='over-and-by-together' - ), pytest.param( { 'parameters.off': {'dims': [], 'dtype': 'int'}, - 'objective': {'expression': 'sum(shift(p, over=g, offset=off + 0), over=g)'}, + 'objective': {'expression': 'sum(shift(p, along=g, offset=off + 0), over=g)'}, }, ('shift(offset=) takes a number or the name of an integer parameter', 'Precompute it as a parameter'), id='an-amount-that-is-an-expression', ), pytest.param( - {'objective': {'expression': 'sum(sum_back(p, over=g, within=2 * 1), over=g)'}}, - ('sum_back(within=) takes a number or the name of an integer parameter',), + {'objective': {'expression': 'sum(sum_back(p, along=g, window=2 * 1), over=g)'}}, + ('sum_back(window=) takes a number or the name of an integer parameter',), id='a-width-that-is-an-expression', ), pytest.param( - {'objective': {'expression': 'sum(shift(p, over=g, offset=1, edge=1 + 1), over=g)'}}, + {'objective': {'expression': 'sum(shift(p, along=g, offset=1, edge=1 + 1), over=g)'}}, ('shift(edge=) is an expression, and an edge is the keyword',), id='an-edge-that-is-an-expression', ), pytest.param( - {'lookups.hk': {'over': 'h', 'into': 'g'}, 'objective': {'expression': 'sum(sum(q, by=[lk, hk]))'}}, - ('groups through lookups over different dimensions',), - id='by-lookups-over-different-dimensions', + { + 'relations.hk': {'columns': ['h', 'g'], 'key': 'h'}, + 'objective': {'expression': 'sum(sum(q, by=[lk, hk]))'}, + }, + ('groups through relations along different dimensions',), + id='by-relations-over-different-dimensions', ), pytest.param( {'objective': {'expression': 'sum(sum(p, by=[lk, lk]))'}}, - ("targets ['h'] more than once",), + ("produces ['h'] more than once",), id='by-the-same-target-twice', ), + pytest.param( + { + 'dimensions.z': {}, + 'relations.lz': {'columns': ['h', 'z'], 'key': 'h'}, + 'objective': {'expression': 'sum(sum(q, by=[lk, lz]))'}, + }, + ('groups through relations along different dimensions',), + id='by-relations-walking-different-dimensions', + ), + pytest.param( + { + 'dimensions.z': {}, + 'relations.lz': {'columns': ['g', 'z', 'h'], 'key': ['g', 'z']}, + 'objective': {'expression': 'sum(sum(q, by=[lk, lz], over=g))'}, + }, + ('a list walks each relation by its declared key and value, so a column keyword has nothing to name',), + id='by-a-list-with-from', + ), + pytest.param( + { + 'dimensions.z': {}, + 'relations.lz': {'columns': ['g', 'z', 'h'], 'key': ['g', 'z']}, + 'variables.q.where': 'lk != lz', + }, + ('compares relations keyed over different dimensions',), + id='where-two-relations-with-different-keys', + ), pytest.param( {'variables.p.where': 'c > flag'}, ('compares two parameters',), id='where-against-a-parameter' ), @@ -724,8 +902,8 @@ class TestRulesDecidedWithoutData: ), pytest.param( {'variables.p.where': 'c > lk'}, - ('against lookup', 'structure rather than data'), - id='where-against-a-lookup', + ('against relation', 'structure rather than data'), + id='where-against-a-relation', ), pytest.param( {'variables.p.where': 'c > h'}, @@ -795,7 +973,7 @@ def test_an_empty_list_survives_the_round_trip(self): def test_an_empty_section_is_not_written(self): written = to_spec(DISPATCH_MODEL).to_yaml() - assert 'lookups' not in written and 'macros' not in written, 'a section declaring nothing says nothing' + assert 'relations' not in written and 'macros' not in written, 'a section declaring nothing says nothing' def test_a_default_is_written_out_and_an_absence_is_not(self): written = to_spec(DISPATCH_MODEL).to_dict() @@ -1071,7 +1249,7 @@ class TestADeclarationIsNamed: 'section', [ 'dimensions', - 'lookups', + 'relations', 'parameters', 'variables', 'expressions', @@ -1084,7 +1262,7 @@ class TestADeclarationIsNamed: def test_a_name_no_expression_could_write_is_refused(self, section: str, name: str): declarations: dict[str, Any] = { 'dimensions': {'dtype': 'str'}, - 'lookups': {'over': 'g', 'into': 'h'}, + 'relations': {'columns': ['g', 'h'], 'key': 'g'}, 'parameters': {'dims': ['g']}, 'variables': {'dims': ['g']}, 'expressions': {'expression': 'c'}, diff --git a/tests/typesetting/golden/latex.out b/tests/typesetting/golden/latex.out index 4b434197..9a9cadd3 100644 --- a/tests/typesetting/golden/latex.out +++ b/tests/typesetting/golden/latex.out @@ -9,12 +9,12 @@ \paragraph{Sets} \begin{description} -\item[{$\mathcal{T}$}] index $t$ --- \texttt{snapshot} (\texttt{int} coordinates) with $\mathrm{season\_of}: \mathcal{T} \to \mathcal{S}$ -\item[{$\mathcal{G}$}] index $g$ --- \texttt{generator} with $\mathrm{gen\_bus}: \mathcal{G} \to \mathcal{B},\ \mathrm{gen\_tech}: \mathcal{G} \to \mathcal{E}$ -\item[{$\mathcal{B}$}] index $b$ --- \texttt{bus} with $\mathrm{zone\_of}: \mathcal{B} \to \mathcal{Z},\ \mathrm{area\_of}: \mathcal{B} \to \mathcal{Z}$ -\item[{$\mathcal{Z}$}] index $z$ --- \texttt{zone} -\item[{$\mathcal{S}$}] index $s$ --- \texttt{season} -\item[{$\mathcal{E}$}] index $e$ --- \texttt{technology} +\item[{$\mathcal{T}$}] index $t$ --- \texttt{snapshot} (\texttt{int} coordinates) with $\mathrm{season\_of}: \mathcal{T} \to \mathcal{S},\ \mathrm{gen\_zone}: \mathcal{G} \times \mathcal{T} \to \mathcal{Z},\ \mathrm{rep\_of}: \mathcal{T} \to \mathcal{T}$ +\item[{$\mathcal{G}$}] index $g$ --- \texttt{generator} with $\mathrm{gen\_bus}: \mathcal{G} \to \mathcal{B},\ \mathrm{gen\_tech}: \mathcal{G} \to \mathcal{E},\ \mathrm{gen\_zone}: \mathcal{G} \times \mathcal{T} \to \mathcal{Z},\ \mathrm{connection} \subseteq \mathcal{G} \times \mathcal{B},\ \mathrm{gen\_bt}: \mathcal{G} \to \mathcal{B} \times \mathcal{E}$ +\item[{$\mathcal{B}$}] index $b$ --- \texttt{bus} with $\mathrm{gen\_bus}: \mathcal{G} \to \mathcal{B},\ \mathrm{zone\_of}: \mathcal{B} \to \mathcal{Z},\ \mathrm{area\_of}: \mathcal{B} \to \mathcal{Z},\ \mathrm{connection} \subseteq \mathcal{G} \times \mathcal{B},\ \mathrm{gen\_bt}: \mathcal{G} \to \mathcal{B} \times \mathcal{E}$ +\item[{$\mathcal{Z}$}] index $z$ --- \texttt{zone} with $\mathrm{zone\_of}: \mathcal{B} \to \mathcal{Z},\ \mathrm{area\_of}: \mathcal{B} \to \mathcal{Z},\ \mathrm{gen\_zone}: \mathcal{G} \times \mathcal{T} \to \mathcal{Z}$ +\item[{$\mathcal{S}$}] index $s$ --- \texttt{season} with $\mathrm{season\_of}: \mathcal{T} \to \mathcal{S}$ +\item[{$\mathcal{E}$}] index $e$ --- \texttt{technology} with $\mathrm{gen\_tech}: \mathcal{G} \to \mathcal{E},\ \mathrm{gen\_bt}: \mathcal{G} \to \mathcal{B} \times \mathcal{E}$ \end{description} \paragraph{Parameters} @@ -61,11 +61,11 @@ \noindent $t \boxminus_{v} k$ denotes translation with $v$ standing where index $t-k$ leaves the dimension (\texttt{shift(edge=v)}), so the row at that boundary is built and carries $v$ rather than being dropped. -\noindent $t \ominus^{\mathrm{lookup}(t)} k$ denotes a translation counted inside the group a lookup puts $t$ in (\texttt{shift(by=lookup)}), so a term never crosses out of its own group. The two modifiers take different slots --- the group above, the fill below --- so $t \boxminus_{v}^{\mathrm{lookup}(t)} k$ is both at once. +\noindent $t \ominus^{\mathrm{relation}(t)} k$ denotes a translation counted inside the group a relation puts $t$ in (\texttt{shift(by=relation)}), so a term never crosses out of its own group. The two modifiers take different slots --- the group above, the fill below --- so $t \boxminus_{v}^{\mathrm{relation}(t)} k$ is both at once. \noindent $\mathrm{pos}(t)$ denotes where index $t$ sits along its dimension's own order --- the order \texttt{shift} walks, not the order labels sort in --- counted from $0$. The index itself stays the coordinate, so $t$ compares against labels and $\mathrm{pos}(t)$ against positions. -\noindent $\mathrm{pos}_{\mathrm{lookup}(t)}(t)$ counts within the group a lookup puts $t$ in: the subscript names the map, $\mathcal{T}_{\mathrm{lookup}(t)}$ is the group it lands in, and that group has a first position of its own. +\noindent $\mathrm{pos}_{\mathrm{relation}(t)}(t)$ counts within the group a relation puts $t$ in: the subscript names the map, $\mathcal{T}_{\mathrm{relation}(t)}$ is the group it lands in, and that group has a first position of its own. \noindent $\lvert \mathcal{T} \rvert$ denotes the size of the set being counted along, and a position counted from the end prints against it --- $\lvert \mathcal{T} \rvert - 1$ is the last position, one less than the size because the first is $0$. @@ -92,8 +92,17 @@ \text{history} && \sum_{t' \in \mathcal{T} \,:\, 0 \le t \ominus t' < \mathrm{min\_up}} \mathit{on}_{t',g} & \le \mathit{units}_{g} && \forall\, t \in \mathcal{T},\ g \in \mathcal{G} \\ \text{seasonal\_window} && \sum_{t' \in \mathcal{T} \,:\, 0 \le t -^{\mathrm{season\_of}(t)} t' < 3} \mathit{on}_{t',g} & \le \mathit{units}_{g} && \forall\, t \in \mathcal{T},\ g \in \mathcal{G} \\ \text{pullback} && \mathit{spill}_{t} & \le \mathrm{zone\_cap}_{\mathrm{zone\_of}(b)} && \forall\, t \in \mathcal{T},\ b \in \mathcal{B} \\ +\text{grouped\_once} && \sum_{g \in \mathcal{G} \,:\, \mathrm{gen\_bt.bus}(g) = b \wedge \mathrm{gen\_bt.technology}(g) = e} p_{t,g} & \le \mathrm{tech\_cap}_{b,e} && \forall\, t \in \mathcal{T},\ b \in \mathcal{B},\ e \in \mathcal{E} \\ +\text{pulled\_back\_once} && \mathit{units}_{g} & \le \mathrm{tech\_cap}_{\mathrm{gen\_bt.bus}(g),\mathrm{gen\_bt.technology}(g)} && \forall\, g \in \mathcal{G} \\ +\text{within\_bus} && \mathit{units}_{g} & \le \mathit{units}_{g \boxminus_{0}^{\mathrm{gen\_bt.bus}(g)} 1} && \forall\, g \in \mathcal{G} \,:\, \mathrm{pos}_{\left( \mathrm{gen\_bt.bus}(g),\ \mathrm{gen\_bt.technology}(g) \right)}(g) = 0 \\ +\text{relational} && \sum_{g \in \mathcal{G} \,:\, \left( g,\ b \right) \in \mathrm{connection}} p_{t,g} & \le \mathrm{load}_{t,b} && \forall\, t \in \mathcal{T},\ b \in \mathcal{B} \\ +\text{connected} && p_{t,g} & \le \mathrm{load}_{t,b} && \forall\, t \in \mathcal{T},\ g \in \mathcal{G},\ b \in \mathcal{B} \,:\, \left( g,\ b \right) \in \mathrm{connection} \\ +\text{representative} && \sum_{t' \in \mathcal{T} \,:\, \mathrm{rep\_of}(t') = t} \mathit{spill}_{t'} & \le \mathit{spill}_{\mathrm{rep\_of}(t)} && \forall\, t \in \mathcal{T} \\ \text{grouped\_twice} && \sum_{g \in \mathcal{G} \,:\, \mathrm{gen\_bus}(g) = b \wedge \mathrm{gen\_tech}(g) = e} p_{t,g} & \le \mathrm{tech\_cap}_{b,e} && \forall\, t \in \mathcal{T},\ b \in \mathcal{B},\ e \in \mathcal{E} \\ \text{pulled\_back\_twice} && \mathit{units}_{g} & \le \mathrm{tech\_cap}_{\mathrm{gen\_bus}(g),\mathrm{gen\_tech}(g)} && \forall\, g \in \mathcal{G} \\ +\text{zonal} && \sum_{g \in \mathcal{G} \,:\, \mathrm{gen\_zone}(g,\ t) = z} p_{t,g} & \le \mathrm{zone\_cap}_{z} && \forall\, t \in \mathcal{T},\ z \in \mathcal{Z} \\ +\text{zonal\_history} && \sum_{t \in \mathcal{T} \,:\, \mathrm{gen\_zone}(g,\ t) = z} p_{t,g} & \le \mathrm{zone\_cap}_{z} && \forall\, g \in \mathcal{G},\ z \in \mathcal{Z} \\ +\text{zonal\_pullback} && p_{t,g} & \le \mathit{spill}_{t} \cdot \mathrm{zone\_cap}_{\mathrm{gen\_zone}(g,\ t)} && \forall\, t \in \mathcal{T},\ g \in \mathcal{G} \,:\, \mathrm{gen\_zone}(g,\ t) = \text{'}\mathrm{north}\text{'} \wedge \mathrm{pos}_{\mathrm{gen\_zone}(g,\ t)}(g) = 0 \\ \text{arithmetic} && \sum_{g \in \mathcal{G}} \left( \frac{p_{t,g}}{2} - \mathrm{cost}_{g} + 10^{-5} \cdot p_{t,g} + 2.5 \times 10^{-7} \cdot \mathrm{cost}_{g} + 0.5 \cdot p_{t,g} \right) & \ge -\left( \sum_{g \in \mathcal{G}} p_{t,g} \right) \cdot \left( -3 \right) && \forall\, t \in \mathcal{T} \\ \text{total} && \sum_{t \in \mathcal{T},\ g \in \mathcal{G}} p_{t,g} & \le \mathrm{budget} \\ \text{scalar} && \mathit{units}_{g} & \le \mathrm{budget} && \forall\, g \in \mathcal{G} \,:\, \mathrm{cost}_{g} \text{ is defined} \\ diff --git a/tests/typesetting/golden/markdown.out b/tests/typesetting/golden/markdown.out index fa7ee0a5..a34712bd 100644 --- a/tests/typesetting/golden/markdown.out +++ b/tests/typesetting/golden/markdown.out @@ -6,12 +6,12 @@ every character a notation escapes, set as text: link\_to, 100% & \#1 costs \$5 | Symbol | Meaning | |---|---| -| $`\mathcal{T}`$ | index $`t`$ — `snapshot` (`int` coordinates) with $`\mathrm{season\_of}: \mathcal{T} \to \mathcal{S}`$ | -| $`\mathcal{G}`$ | index $`g`$ — `generator` with $`\mathrm{gen\_bus}: \mathcal{G} \to \mathcal{B},\ \mathrm{gen\_tech}: \mathcal{G} \to \mathcal{E}`$ | -| $`\mathcal{B}`$ | index $`b`$ — `bus` with $`\mathrm{zone\_of}: \mathcal{B} \to \mathcal{Z},\ \mathrm{area\_of}: \mathcal{B} \to \mathcal{Z}`$ | -| $`\mathcal{Z}`$ | index $`z`$ — `zone` | -| $`\mathcal{S}`$ | index $`s`$ — `season` | -| $`\mathcal{E}`$ | index $`e`$ — `technology` | +| $`\mathcal{T}`$ | index $`t`$ — `snapshot` (`int` coordinates) with $`\mathrm{season\_of}: \mathcal{T} \to \mathcal{S},\ \mathrm{gen\_zone}: \mathcal{G} \times \mathcal{T} \to \mathcal{Z},\ \mathrm{rep\_of}: \mathcal{T} \to \mathcal{T}`$ | +| $`\mathcal{G}`$ | index $`g`$ — `generator` with $`\mathrm{gen\_bus}: \mathcal{G} \to \mathcal{B},\ \mathrm{gen\_tech}: \mathcal{G} \to \mathcal{E},\ \mathrm{gen\_zone}: \mathcal{G} \times \mathcal{T} \to \mathcal{Z},\ \mathrm{connection} \subseteq \mathcal{G} \times \mathcal{B},\ \mathrm{gen\_bt}: \mathcal{G} \to \mathcal{B} \times \mathcal{E}`$ | +| $`\mathcal{B}`$ | index $`b`$ — `bus` with $`\mathrm{gen\_bus}: \mathcal{G} \to \mathcal{B},\ \mathrm{zone\_of}: \mathcal{B} \to \mathcal{Z},\ \mathrm{area\_of}: \mathcal{B} \to \mathcal{Z},\ \mathrm{connection} \subseteq \mathcal{G} \times \mathcal{B},\ \mathrm{gen\_bt}: \mathcal{G} \to \mathcal{B} \times \mathcal{E}`$ | +| $`\mathcal{Z}`$ | index $`z`$ — `zone` with $`\mathrm{zone\_of}: \mathcal{B} \to \mathcal{Z},\ \mathrm{area\_of}: \mathcal{B} \to \mathcal{Z},\ \mathrm{gen\_zone}: \mathcal{G} \times \mathcal{T} \to \mathcal{Z}`$ | +| $`\mathcal{S}`$ | index $`s`$ — `season` with $`\mathrm{season\_of}: \mathcal{T} \to \mathcal{S}`$ | +| $`\mathcal{E}`$ | index $`e`$ — `technology` with $`\mathrm{gen\_tech}: \mathcal{G} \to \mathcal{E},\ \mathrm{gen\_bt}: \mathcal{G} \to \mathcal{B} \times \mathcal{E}`$ | #### Parameters @@ -60,11 +60,11 @@ $`t \ominus k`$ denotes cyclic translation: index $`t-k`$ taken modulo the size $`t \boxminus_{v} k`$ denotes translation with $`v`$ standing where index $`t-k`$ leaves the dimension (`shift(edge=v)`), so the row at that boundary is built and carries $`v`$ rather than being dropped. -$`t \ominus^{\mathrm{lookup}(t)} k`$ denotes a translation counted inside the group a lookup puts $`t`$ in (`shift(by=lookup)`), so a term never crosses out of its own group. The two modifiers take different slots — the group above, the fill below — so $`t \boxminus_{v}^{\mathrm{lookup}(t)} k`$ is both at once. +$`t \ominus^{\mathrm{relation}(t)} k`$ denotes a translation counted inside the group a relation puts $`t`$ in (`shift(by=relation)`), so a term never crosses out of its own group. The two modifiers take different slots — the group above, the fill below — so $`t \boxminus_{v}^{\mathrm{relation}(t)} k`$ is both at once. $`\mathrm{pos}(t)`$ denotes where index $`t`$ sits along its dimension's own order — the order `shift` walks, not the order labels sort in — counted from $`0`$. The index itself stays the coordinate, so $`t`$ compares against labels and $`\mathrm{pos}(t)`$ against positions. -$`\mathrm{pos}_{\mathrm{lookup}(t)}(t)`$ counts within the group a lookup puts $`t`$ in: the subscript names the map, $`\mathcal{T}_{\mathrm{lookup}(t)}`$ is the group it lands in, and that group has a first position of its own. +$`\mathrm{pos}_{\mathrm{relation}(t)}(t)`$ counts within the group a relation puts $`t`$ in: the subscript names the map, $`\mathcal{T}_{\mathrm{relation}(t)}`$ is the group it lands in, and that group has a first position of its own. $`\lvert \mathcal{T} \rvert`$ denotes the size of the set being counted along, and a position counted from the end prints against it — $`\lvert \mathcal{T} \rvert - 1`$ is the last position, one less than the size because the first is $`0`$. @@ -172,6 +172,42 @@ p_{t,g} \le p_{t \boxminus_{0}^{\mathrm{season\_of}(t)} 1,g} \qquad \forall\, t \mathit{spill}_{t} \le \mathrm{zone\_cap}_{\mathrm{zone\_of}(b)} \qquad \forall\, t \in \mathcal{T},\ b \in \mathcal{B} ``` +**`grouped_once`** + +```math +\sum_{g \in \mathcal{G} \,:\, \mathrm{gen\_bt.bus}(g) = b \wedge \mathrm{gen\_bt.technology}(g) = e} p_{t,g} \le \mathrm{tech\_cap}_{b,e} \qquad \forall\, t \in \mathcal{T},\ b \in \mathcal{B},\ e \in \mathcal{E} +``` + +**`pulled_back_once`** + +```math +\mathit{units}_{g} \le \mathrm{tech\_cap}_{\mathrm{gen\_bt.bus}(g),\mathrm{gen\_bt.technology}(g)} \qquad \forall\, g \in \mathcal{G} +``` + +**`within_bus`** + +```math +\mathit{units}_{g} \le \mathit{units}_{g \boxminus_{0}^{\mathrm{gen\_bt.bus}(g)} 1} \qquad \forall\, g \in \mathcal{G} \,:\, \mathrm{pos}_{\left( \mathrm{gen\_bt.bus}(g),\ \mathrm{gen\_bt.technology}(g) \right)}(g) = 0 +``` + +**`relational`** + +```math +\sum_{g \in \mathcal{G} \,:\, \left( g,\ b \right) \in \mathrm{connection}} p_{t,g} \le \mathrm{load}_{t,b} \qquad \forall\, t \in \mathcal{T},\ b \in \mathcal{B} +``` + +**`connected`** + +```math +p_{t,g} \le \mathrm{load}_{t,b} \qquad \forall\, t \in \mathcal{T},\ g \in \mathcal{G},\ b \in \mathcal{B} \,:\, \left( g,\ b \right) \in \mathrm{connection} +``` + +**`representative`** + +```math +\sum_{t' \in \mathcal{T} \,:\, \mathrm{rep\_of}(t') = t} \mathit{spill}_{t'} \le \mathit{spill}_{\mathrm{rep\_of}(t)} \qquad \forall\, t \in \mathcal{T} +``` + **`grouped_twice`** ```math @@ -184,6 +220,24 @@ p_{t,g} \le p_{t \boxminus_{0}^{\mathrm{season\_of}(t)} 1,g} \qquad \forall\, t \mathit{units}_{g} \le \mathrm{tech\_cap}_{\mathrm{gen\_bus}(g),\mathrm{gen\_tech}(g)} \qquad \forall\, g \in \mathcal{G} ``` +**`zonal`** + +```math +\sum_{g \in \mathcal{G} \,:\, \mathrm{gen\_zone}(g,\ t) = z} p_{t,g} \le \mathrm{zone\_cap}_{z} \qquad \forall\, t \in \mathcal{T},\ z \in \mathcal{Z} +``` + +**`zonal_history`** + +```math +\sum_{t \in \mathcal{T} \,:\, \mathrm{gen\_zone}(g,\ t) = z} p_{t,g} \le \mathrm{zone\_cap}_{z} \qquad \forall\, g \in \mathcal{G},\ z \in \mathcal{Z} +``` + +**`zonal_pullback`** + +```math +p_{t,g} \le \mathit{spill}_{t} \cdot \mathrm{zone\_cap}_{\mathrm{gen\_zone}(g,\ t)} \qquad \forall\, t \in \mathcal{T},\ g \in \mathcal{G} \,:\, \mathrm{gen\_zone}(g,\ t) = \text{'}\mathrm{north}\text{'} \wedge \mathrm{pos}_{\mathrm{gen\_zone}(g,\ t)}(g) = 0 +``` + **`arithmetic`** ```math diff --git a/tests/typesetting/golden/model.yaml b/tests/typesetting/golden/model.yaml index e1ebe59d..74eb6930 100644 --- a/tests/typesetting/golden/model.yaml +++ b/tests/typesetting/golden/model.yaml @@ -20,12 +20,16 @@ dimensions: season: { dtype: str } technology: { dtype: str } -lookups: - gen_bus: { over: generator, into: bus } - gen_tech: { over: generator, into: technology } # a second map out of `generator`, to group through both at once - zone_of: { over: bus, into: zone } - area_of: { over: bus, into: zone } # a second map into the same set, to compare against - season_of: { over: snapshot, into: season } +relations: + gen_bus: { columns: [generator, bus], key: generator } + gen_tech: { columns: [generator, technology], key: generator } # a second map out of `generator`, to group through both at once + zone_of: { columns: [bus, zone], key: bus } + area_of: { columns: [bus, zone], key: bus } # a second map into the same set, to compare against + season_of: { columns: [snapshot, season], key: snapshot } + gen_zone: { columns: [generator, snapshot, zone], key: [generator, snapshot] } # a map keyed by two dimensions: a call walks one and joins on the other + rep_of: { columns: { snapshot: snapshot, rep: snapshot }, key: snapshot } # a map into its own dimension: the representative snapshot + connection: { columns: [generator, bus] } # a bare relation, no key: many-to-many, walked only by sum with both ends named + gen_bt: { columns: [generator, bus, technology], key: generator } # one table with two value columns, walked to both at once parameters: p_max: { dims: [generator] } @@ -102,56 +106,86 @@ constraints: starts: # names the cased expression: its symbol prints here, its block once below dims: [snapshot, generator] expression: p <= startup_cost - balance: # sum over a lookup + balance: # sum over a relation dims: [snapshot, bus] expression: sum(p, by=gen_bus) + spill - slack == load ramp: # roll (cyclic) and shift (acyclic) in one equation dims: [snapshot, generator] - expression: p - shift(p, over=snapshot, offset=1, edge='wrap') <= shift(p, over=snapshot, offset=1) + p_max + expression: p - shift(p, along=snapshot, offset=1, edge='wrap') <= shift(p, along=snapshot, offset=1) + p_max edges: # the two translations `ramp` leaves out: a fill, and forwards dims: [snapshot, generator] expression: >- - shift(p, over=snapshot, offset=1, edge=0) - <= shift(p, over=snapshot, offset=-1, edge=0) + p_max + shift(p, along=snapshot, offset=1, edge=0) + <= shift(p, along=snapshot, offset=-1, edge=0) + p_max ahead: # the cyclic translation forwards, which is a fourth symbol again dims: [snapshot, generator] - expression: p <= shift(p, over=snapshot, offset=-1, edge='wrap') + expression: p <= shift(p, along=snapshot, offset=-1, edge='wrap') composed: # two steps of one policy are one step; a zero step is none at all dims: [snapshot, generator] - expression: shift(shift(p, over=snapshot, offset=1), over=snapshot, offset=1) <= shift(p_max, over=generator, offset=0) + expression: shift(shift(p, along=snapshot, offset=1), along=snapshot, offset=1) <= shift(p_max, along=generator, offset=0) uncomposed: # a named offset under a numbered one stays two steps, not their sum dims: [snapshot, generator] - expression: shift(shift(p, over=snapshot, offset=lead, edge=0), over=snapshot, offset=1) <= p_max + expression: shift(shift(p, along=snapshot, offset=lead, edge=0), along=snapshot, offset=1) <= p_max crossed: # two dimensions translated at one leaf, each with its own policy dims: [snapshot, generator] - expression: shift(shift(p, over=snapshot, offset=1, edge='wrap'), over=generator, offset=-1) <= p_max + expression: shift(shift(p, along=snapshot, offset=1, edge='wrap'), along=generator, offset=-1) <= p_max lead_time: # an offset the data carries, so it prints as a symbol rather than a number dims: [snapshot, generator] - expression: shift(p, over=snapshot, offset=lead, edge=0) <= p_max - in_season: # a translation partitioned by a lookup: the group rides on the operator + expression: shift(p, along=snapshot, offset=lead, edge=0) <= p_max + in_season: # a translation partitioned by a relation: the group rides on the operator dims: [snapshot, generator] - expression: p <= shift(p, over=snapshot, offset=1, edge='wrap', by=season_of) + expression: p <= shift(p, along=snapshot, offset=1, edge='wrap', by=season_of) held_in_season: # the same group, with a fill: each season's opening row is kept and given a zero dims: [snapshot, generator] - expression: p <= shift(p, over=snapshot, offset=1, edge=0, by=season_of) + expression: p <= shift(p, along=snapshot, offset=1, edge=0, by=season_of) window: # a trailing window of fixed width dims: [snapshot, generator] - expression: sum_back(on, over=snapshot, within=3) <= units + expression: sum_back(on, along=snapshot, window=3) <= units history: # the same window, its width in the data and its edge wrapped dims: [snapshot, generator] - expression: sum_back(on, over=snapshot, within=min_up, edge='wrap') <= units - seasonal_window: # a window partitioned by a lookup: the group rides on the operator + expression: sum_back(on, along=snapshot, window=min_up, edge='wrap') <= units + seasonal_window: # a window partitioned by a relation: the group rides on the operator dims: [snapshot, generator] - expression: sum_back(on, over=snapshot, within=3, by=season_of) <= units - pullback: # at(), which re-indexes through a lookup instead of an offset + expression: sum_back(on, along=snapshot, window=3, by=season_of) <= units + pullback: # at(), which re-indexes through a relation instead of an offset dims: [snapshot, bus] expression: spill <= at(zone_cap, by=zone_of) + grouped_once: # one table walked to two value columns: the domain carries a condition per column + dims: [snapshot, bus, technology] + expression: sum(p, by=gen_bt, into=[bus, technology]) <= tech_cap + pulled_back_once: # its adjoint, reading one slot through two columns of one table + dims: [generator] + expression: units <= at(tech_cap, by=gen_bt, over=[bus, technology]) + within_bus: # a partition grouped by one named value column of a two-value table, and a position within both + dims: [generator] + where: "position(generator, by=gen_bt, within=[bus, technology]) == 0" + expression: units <= shift(units, along=generator, offset=1, edge=0, by=gen_bt, within=bus) + relational: # a sum through a bare relation: the domain is a row of the relation rather than a function's value + dims: [snapshot, bus] + expression: sum(p, by=connection, over=generator, into=bus) <= load + connected: # a bare relation as a where: the row of the frame has to be a member of the relation + dims: [snapshot, generator, bus] + where: "connection" + expression: p <= load + representative: # a map into its own dimension, walked both ways: the frame is unchanged and the index is primed + dims: [snapshot] + expression: sum(spill, by=rep_of) <= at(spill, by=rep_of) grouped_twice: # one grouping through two maps: the domain carries both conditions dims: [snapshot, bus, technology] expression: sum(p, by=[gen_bus, gen_tech]) <= tech_cap pulled_back_twice: # its adjoint, reading one slot through a pair of labels dims: [generator] expression: units <= at(tech_cap, by=[gen_bus, gen_tech]) + zonal: # a grouping through a two-key map, walked along one key: the condition reads the other, and the row keeps it + dims: [snapshot, zone] + expression: sum(p, by=gen_zone, over=generator) <= zone_cap + zonal_history: # the same table walked along its other key + dims: [generator, zone] + expression: sum(p, by=gen_zone, over=snapshot) <= zone_cap + zonal_pullback: # its adjoint, reading the slot the row's own snapshot puts the generator in + dims: [snapshot, generator] + where: "gen_zone == 'north' AND position(generator, by=gen_zone) == 0" + expression: p <= at(spill * zone_cap, by=gen_zone, into=generator) arithmetic: # division, both unary signs, a sign beside a sign, floats with and without an exponent, bracketing dims: [snapshot] expression: >- @@ -176,7 +210,7 @@ constraints: dims: [snapshot, generator] where: "position(snapshot) == -1 OR position(snapshot, by=season_of) == -1" expression: on == 0 - northern: # a lookup compared to a label, to another lookup, and to nothing + northern: # a relation compared to a label, to another relation, and to nothing dims: [snapshot, bus] where: "zone_of == 'north' AND zone_of != area_of AND zone_of" expression: slack <= load diff --git a/tests/typesetting/golden/typst.out b/tests/typesetting/golden/typst.out index 7f286e31..6f3500dd 100644 --- a/tests/typesetting/golden/typst.out +++ b/tests/typesetting/golden/typst.out @@ -4,12 +4,12 @@ every character a notation escapes, set as text: link\_to, 100% & \#1 costs \$5 {net} \~ ^ \\ \*star\* \@ref \, and `a_name` in backticks set in code == Sets -/ $cal(T)$: index $t$ --- `snapshot` (`int` coordinates) with $upright("season_of"): cal(T) arrow.r cal(S)$ -/ $cal(G)$: index $g$ --- `generator` with $upright("gen_bus"): cal(G) arrow.r cal(B), upright("gen_tech"): cal(G) arrow.r cal(E)$ -/ $cal(B)$: index $b$ --- `bus` with $upright("zone_of"): cal(B) arrow.r cal(Z), upright("area_of"): cal(B) arrow.r cal(Z)$ -/ $cal(Z)$: index $z$ --- `zone` -/ $cal(S)$: index $s$ --- `season` -/ $cal(E)$: index $e$ --- `technology` +/ $cal(T)$: index $t$ --- `snapshot` (`int` coordinates) with $upright("season_of"): cal(T) arrow.r cal(S), upright("gen_zone"): cal(G) times cal(T) arrow.r cal(Z), upright("rep_of"): cal(T) arrow.r cal(T)$ +/ $cal(G)$: index $g$ --- `generator` with $upright("gen_bus"): cal(G) arrow.r cal(B), upright("gen_tech"): cal(G) arrow.r cal(E), upright("gen_zone"): cal(G) times cal(T) arrow.r cal(Z), upright("connection") subset.eq cal(G) times cal(B), upright("gen_bt"): cal(G) arrow.r cal(B) times cal(E)$ +/ $cal(B)$: index $b$ --- `bus` with $upright("gen_bus"): cal(G) arrow.r cal(B), upright("zone_of"): cal(B) arrow.r cal(Z), upright("area_of"): cal(B) arrow.r cal(Z), upright("connection") subset.eq cal(G) times cal(B), upright("gen_bt"): cal(G) arrow.r cal(B) times cal(E)$ +/ $cal(Z)$: index $z$ --- `zone` with $upright("zone_of"): cal(B) arrow.r cal(Z), upright("area_of"): cal(B) arrow.r cal(Z), upright("gen_zone"): cal(G) times cal(T) arrow.r cal(Z)$ +/ $cal(S)$: index $s$ --- `season` with $upright("season_of"): cal(T) arrow.r cal(S)$ +/ $cal(E)$: index $e$ --- `technology` with $upright("gen_tech"): cal(G) arrow.r cal(E), upright("gen_bt"): cal(G) arrow.r cal(B) times cal(E)$ == Parameters / $upright("p")^(upright("max"))$: `p_max` over $cal(G)$ @@ -49,11 +49,11 @@ $t minus.o k$ denotes cyclic translation: index $t-k$ taken modulo the size of t $t minus.square_(v) k$ denotes translation with $v$ standing where index $t-k$ leaves the dimension (`shift(edge=v)`), so the row at that boundary is built and carries $v$ rather than being dropped. -$t minus.o^(upright("lookup")(t)) k$ denotes a translation counted inside the group a lookup puts $t$ in (`shift(by=lookup)`), so a term never crosses out of its own group. The two modifiers take different slots --- the group above, the fill below --- so $t minus.square_(v)^(upright("lookup")(t)) k$ is both at once. +$t minus.o^(upright("relation")(t)) k$ denotes a translation counted inside the group a relation puts $t$ in (`shift(by=relation)`), so a term never crosses out of its own group. The two modifiers take different slots --- the group above, the fill below --- so $t minus.square_(v)^(upright("relation")(t)) k$ is both at once. $upright("pos")(t)$ denotes where index $t$ sits along its dimension's own order --- the order `shift` walks, not the order labels sort in --- counted from $0$. The index itself stays the coordinate, so $t$ compares against labels and $upright("pos")(t)$ against positions. -$upright("pos")_(upright("lookup")(t))(t)$ counts within the group a lookup puts $t$ in: the subscript names the map, $cal(T)_(upright("lookup")(t))$ is the group it lands in, and that group has a first position of its own. +$upright("pos")_(upright("relation")(t))(t)$ counts within the group a relation puts $t$ in: the subscript names the map, $cal(T)_(upright("relation")(t))$ is the group it lands in, and that group has a first position of its own. $abs(cal(T))$ denotes the size of the set being counted along, and a position counted from the end prints against it --- $abs(cal(T)) - 1$ is the last position, one less than the size because the first is $0$. @@ -79,8 +79,17 @@ $ upright("budgeted") & italic("spend")_(t) & <= upright("budget") & forall t in upright("history") & sum_(t' in cal(T) colon 0 <= t minus.o t' < upright("min_up")) italic("on")_(t',g) & <= italic("units")_(g) & forall t in cal(T), g in cal(G) \ upright("seasonal_window") & sum_(t' in cal(T) colon 0 <= t -^(upright("season_of")(t)) t' < 3) italic("on")_(t',g) & <= italic("units")_(g) & forall t in cal(T), g in cal(G) \ upright("pullback") & italic("spill")_(t) & <= upright("zone_cap")_(upright("zone_of")(b)) & forall t in cal(T), b in cal(B) \ + upright("grouped_once") & sum_(g in cal(G) colon upright("gen_bt.bus")(g) = b and upright("gen_bt.technology")(g) = e) p_(t,g) & <= upright("tech_cap")_(b,e) & forall t in cal(T), b in cal(B), e in cal(E) \ + upright("pulled_back_once") & italic("units")_(g) & <= upright("tech_cap")_(upright("gen_bt.bus")(g),upright("gen_bt.technology")(g)) & forall g in cal(G) \ + upright("within_bus") & italic("units")_(g) & <= italic("units")_(g minus.square_(0)^(upright("gen_bt.bus")(g)) 1) & forall g in cal(G) colon upright("pos")_((upright("gen_bt.bus")(g), upright("gen_bt.technology")(g)))(g) = 0 \ + upright("relational") & sum_(g in cal(G) colon (g, b) in upright("connection")) p_(t,g) & <= upright("load")_(t,b) & forall t in cal(T), b in cal(B) \ + upright("connected") & p_(t,g) & <= upright("load")_(t,b) & forall t in cal(T), g in cal(G), b in cal(B) colon (g, b) in upright("connection") \ + upright("representative") & sum_(t' in cal(T) colon upright("rep_of")(t') = t) italic("spill")_(t') & <= italic("spill")_(upright("rep_of")(t)) & forall t in cal(T) \ upright("grouped_twice") & sum_(g in cal(G) colon upright("gen_bus")(g) = b and upright("gen_tech")(g) = e) p_(t,g) & <= upright("tech_cap")_(b,e) & forall t in cal(T), b in cal(B), e in cal(E) \ upright("pulled_back_twice") & italic("units")_(g) & <= upright("tech_cap")_(upright("gen_bus")(g),upright("gen_tech")(g)) & forall g in cal(G) \ + upright("zonal") & sum_(g in cal(G) colon upright("gen_zone")(g, t) = z) p_(t,g) & <= upright("zone_cap")_(z) & forall t in cal(T), z in cal(Z) \ + upright("zonal_history") & sum_(t in cal(T) colon upright("gen_zone")(g, t) = z) p_(t,g) & <= upright("zone_cap")_(z) & forall g in cal(G), z in cal(Z) \ + upright("zonal_pullback") & p_(t,g) & <= italic("spill")_(t) dot upright("zone_cap")_(upright("gen_zone")(g, t)) & forall t in cal(T), g in cal(G) colon upright("gen_zone")(g, t) = upright("'north'") and upright("pos")_(upright("gen_zone")(g, t))(g) = 0 \ upright("arithmetic") & sum_(g in cal(G)) (frac(p_(t,g), 2) - upright("cost")_(g) + 10^(-5) dot p_(t,g) + 2.5 times 10^(-7) dot upright("cost")_(g) + 0.5 dot p_(t,g)) & >= -(sum_(g in cal(G)) p_(t,g)) dot (-3) & forall t in cal(T) \ upright("total") & sum_(t in cal(T), g in cal(G)) p_(t,g) & <= upright("budget") \ upright("scalar") & italic("units")_(g) & <= upright("budget") & forall g in cal(G) colon upright("cost")_(g) upright(" is defined") \ diff --git a/tests/typesetting/test_walk.py b/tests/typesetting/test_walk.py index 0db831e7..47694c0b 100644 --- a/tests/typesetting/test_walk.py +++ b/tests/typesetting/test_walk.py @@ -108,14 +108,14 @@ def test_a_mask_reads_as_definedness_unless_its_parameter_is_boolean( def _storage(shift: str) -> dict[str, object]: - """A state-of-charge balance, `soc == shift(soc, over=snapshot, )`: one model per translation policy. + """A state-of-charge balance, `soc == shift(soc, along=snapshot, )`: one model per translation policy. No parameter, so it is also the model the "given" convention has nothing to say about. """ return { 'dimensions': {'snapshot': {'dtype': 'int'}}, 'variables': {'soc': {'dims': ['snapshot'], 'bounds': {'lower': 0, 'upper': 100}}}, - 'constraints': {'balance': {'dims': ['snapshot'], 'expression': f'soc == shift(soc, over=snapshot, {shift})'}}, + 'constraints': {'balance': {'dims': ['snapshot'], 'expression': f'soc == shift(soc, along=snapshot, {shift})'}}, } @@ -156,12 +156,12 @@ def test_a_fill_and_a_group_take_the_operators_two_slots(name: FormatName, fmt: """ model = { 'dimensions': {'snapshot': {'dtype': 'int'}, 'season': {'dtype': 'str'}}, - 'lookups': {'season_of': {'over': 'snapshot', 'into': 'season'}}, + 'relations': {'season_of': {'columns': ['snapshot', 'season'], 'key': 'snapshot'}}, 'variables': {'p': {'dims': ['snapshot'], 'bounds': {'lower': 0}}}, 'constraints': { 'held': { 'dims': ['snapshot'], - 'expression': 'p <= shift(p, over=snapshot, offset=1, edge=0, by=season_of)', + 'expression': 'p <= shift(p, along=snapshot, offset=1, edge=0, by=season_of)', } }, 'objective': {'sense': 'minimize', 'expression': 'sum(p)'}, @@ -179,7 +179,7 @@ def test_a_fill_and_a_group_take_the_operators_two_slots(name: FormatName, fmt: def test_a_translation_under_a_pullback_survives_it(name: FormatName, fmt: Format): """``at`` and ``shift`` both re-index at the leaf, and the leaf has one subscript. - Whoever wrote it last used to win: ``at(shift(cap, over=period, offset=1, + Whoever wrote it last used to win: ``at(shift(cap, along=period, offset=1, edge=0), by=period_of)`` printed `cap_{period_of(t)}`, dropping a translation the plan builds. The subscript is a composition, so it renders as one. @@ -189,13 +189,13 @@ def test_a_translation_under_a_pullback_survives_it(name: FormatName, fmt: Forma 'snapshot': {'dtype': 'int'}, 'period': {'dtype': 'int'}, }, - 'lookups': {'period_of': {'over': 'snapshot', 'into': 'period'}}, + 'relations': {'period_of': {'columns': ['snapshot', 'period'], 'key': 'snapshot'}}, 'parameters': {'cap': {'dims': ['period']}}, 'variables': {'p': {'dims': ['snapshot'], 'bounds': {'lower': 0}}}, 'constraints': { 'within': { 'dims': ['snapshot'], - 'expression': 'p <= at(shift(cap, over=period, offset=1, edge=0), by=period_of)', + 'expression': 'p <= at(shift(cap, along=period, offset=1, edge=0), by=period_of)', } }, } @@ -219,7 +219,7 @@ def test_translations_that_disagree_at_the_edge_do_not_merge(name: FormatName, f 'constraints': { 'b': { 'dims': ['snapshot'], - 'expression': "soc <= shift(shift(soc, over=snapshot, offset=1, edge='wrap'), over=snapshot, offset=1)", + 'expression': "soc <= shift(shift(soc, along=snapshot, offset=1, edge='wrap'), along=snapshot, offset=1)", } }, } @@ -265,16 +265,16 @@ def test_a_negative_fill_prints(name: FormatName, fmt: Format): 'dimensions': {'g': {}}, 'parameters': {'cap': {'dims': ['g']}}, 'variables': {'p': {'dims': ['g']}}, - 'constraints': {'k': {'dims': ['g'], 'expression': 'p <= shift(cap, over=g, offset=1, edge=-1)'}}, + 'constraints': {'k': {'dims': ['g'], 'expression': 'p <= shift(cap, along=g, offset=1, edge=-1)'}}, } assert fmt.operators['edge_minus'] in typeset(model, name, legend=False) def _selected(mask: str) -> dict[str, Any]: - """One constraint carrying *mask*, over a dimension a lookup groups.""" + """One constraint carrying *mask*, over a dimension a relation groups.""" return { 'dimensions': {'snapshot': {'dtype': 'int'}, 'season': {'dtype': 'str'}}, - 'lookups': {'season_of': {'over': 'snapshot', 'into': 'season'}}, + 'relations': {'season_of': {'columns': ['snapshot', 'season'], 'key': 'snapshot'}}, 'variables': {'soc': {'dims': ['snapshot'], 'bounds': {'lower': 0}}}, 'constraints': {'seed': {'dims': ['snapshot'], 'where': mask, 'expression': 'soc == 0'}}, } @@ -636,11 +636,11 @@ def test_every_operator_probe_renders(path, name: FormatName, fmt: Format): # --------------------------------------------------------------------------- -#: Two frames over generators, a lookup onto buses and a boolean mask — what the +#: Two frames over generators, a relation onto buses and a boolean mask — what the #: scope and bracketing cases are written against. BUSES = { 'dimensions': {'snapshot': {'dtype': 'int'}, 'generator': {'dtype': 'str'}, 'bus': {'dtype': 'str'}}, - 'lookups': {'bus_of': {'over': 'generator', 'into': 'bus'}}, + 'relations': {'bus_of': {'columns': ['generator', 'bus'], 'key': 'generator'}}, 'parameters': {'load': {'dims': ['snapshot']}, 'k': {'dims': []}, 'flag': {'dims': ['snapshot'], 'dtype': 'bool'}}, 'variables': {'p': {'dims': ['snapshot', 'generator']}, 'q': {'dims': ['snapshot', 'generator']}}, } @@ -661,7 +661,7 @@ def _row(expression: str, where: str | None = None, **patch: object) -> str: pytest.param( 'p == at(sum(q, by=bus_of), by=bus_of)', r"\sum_{g' \in \mathcal{G} \,:\, \mathrm{bus\_of}(g') = \mathrm{bus\_of}(g)} q_{t,g'}", - id='grouped-by-a-lookup', + id='grouped-by-a-relation', ), pytest.param('p == q - sum(q, over=generator)', r"\sum_{g' \in \mathcal{G}} q_{t,g'}", id='over-the-whole-dim'), ], diff --git a/tools/notation.py b/tools/notation.py index de0b0af0..affe89ef 100644 --- a/tools/notation.py +++ b/tools/notation.py @@ -47,7 +47,7 @@ BEGIN, END = '', '' #: The blocks that declare math, in the order the page walks them, and the -#: heading each gets. ``dimensions``, ``lookups`` and ``parameters`` are absent +#: heading each gets. ``dimensions``, ``relations`` and ``parameters`` are absent #: on purpose: they declare no equation, and what they print is the legend, #: which the page shows once as a legend rather than a row at a time. SECTIONS = { @@ -172,11 +172,11 @@ def legend(rendered: str) -> str: #: What the legend is made of. No equation comes from these, so they are shown #: once, together, above the tables they turn into. -DECLARED = ('dimensions', 'lookups', 'parameters') +DECLARED = ('dimensions', 'relations', 'parameters') def preamble(text: str) -> str: - """The fixture's ``dimensions``/``lookups``/``parameters`` blocks, verbatim.""" + """The fixture's ``dimensions``/``relations``/``parameters`` blocks, verbatim.""" blocks = [] for name in DECLARED: body = text[text.index(f'\n{name}:') + 1 :] @@ -190,7 +190,7 @@ def block() -> str: rendered = to_markdown(MODEL, numbered=False) parts = [ '### The legend', - 'A dimension, a lookup and a parameter declare no equation; what they ' + 'A dimension, a relation and a parameter declare no equation; what they ' 'print is the legend every model opens with.', f'```yaml\n{preamble(MODEL.read_text())}\n```', legend(rendered), diff --git a/tools/spec_math.py b/tools/spec_math.py index 48518510..3285753c 100644 --- a/tools/spec_math.py +++ b/tools/spec_math.py @@ -29,18 +29,18 @@ OPERATORS = { 'sum(array)': 'sum_all', 'sum(array, over=dim)': 'sum', - 'sum(array, by=lookup)': 'sum_by', - 'sum(array, by=[lookup, …])': 'sum_by_lookups', - 'at(array, by=lookup)': 'at', - 'shift(array, over=dim, offset=n)': 'shift', - "shift(array, over=dim, offset=n, edge='wrap')": 'shift_wrap', - 'shift(array, over=dim, offset=n, edge=v)': 'shift_edge', - 'shift(array, over=dim, offset=p, edge=…)': 'shift_by_parameter', - 'shift(array, over=dim, offset=n, by=lookup)': 'shift_partitioned', - 'sum_back(array, over=dim, within=n)': 'sum_back', - 'sum_back(array, over=dim, within=p)': 'sum_back_by_parameter', - "sum_back(array, over=dim, within=p, edge='wrap')": 'sum_back_wrap', - 'sum_back(array, over=dim, within=n, by=lookup)': 'sum_back_partitioned', + 'sum(array, by=relation)': 'sum_by', + 'sum(array, by=[relation, …])': 'sum_by_relations', + 'at(array, by=relation)': 'at', + 'shift(array, along=dim, offset=n)': 'shift', + "shift(array, along=dim, offset=n, edge='wrap')": 'shift_wrap', + 'shift(array, along=dim, offset=n, edge=v)': 'shift_edge', + 'shift(array, along=dim, offset=p, edge=…)': 'shift_by_parameter', + 'shift(array, along=dim, offset=n, by=relation)': 'shift_partitioned', + 'sum_back(array, along=dim, window=n)': 'sum_back', + 'sum_back(array, along=dim, window=p)': 'sum_back_by_parameter', + "sum_back(array, along=dim, window=p, edge='wrap')": 'sum_back_wrap', + 'sum_back(array, along=dim, window=n, by=relation)': 'sum_back_partitioned', 'dual(constraint)': 'dual', }