From f15933724ebfb97f1934b78043a7c8ba34d426dd Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 16 Sep 2026 08:51:49 +0000 Subject: [PATCH 1/4] docs(language): a walk through a relation is stated once, in the three verbs the loader uses The relations reference opens its walk section with the definition: a walk consumes one column, produces another, and joins on every other key column; `sum` consumes a key column and produces a value column, `at` consumes a value column and produces the key; name a column only where the relation offers two. The example that follows spells `over=zone` on `at` and annotates each constraint with the three roles, and the refusals quoted are the three that teach the rule. The walk-needs table under "The key is the claim", the rules list that restated it, the partition paragraph and the `where` bullet are folded into `Walks`, `Partitions` and a link to the where-strings table that owns those forms. The field table names `columns:` rather than the retired `over:`. The operators page keeps the null-value and variable facts for `at` and links here for the walk. Sentence measure (docs-writing script) on dimensions.md: 96 sentences, median 16 words, 18 over 25, from 85 sentences, median 22, 30 over 25. Verified without pixi, which the proxy blocks: prettier --check clean, typos clean, `mkdocs build --strict` clean with the Python inventory dropped from a temporary config copy since docs.python.org is blocked, and `pytest tests/test_docs.py tests/test_reading_page.py` gives 34 passed. The example YAML loads through `to_spec`. Not run: `reuse lint`, the full suite, and `compile-tex`, which no `examples/` change needs. Co-Authored-By: Claude Fable 5.1 Claude-Session: https://claude.ai/code/session_019vg9UDLgdaFiu9gtbqAj7P --- docs/reference/language/dimensions.md | 216 ++++++++++++-------------- docs/reference/language/operators.md | 29 ++-- 2 files changed, 113 insertions(+), 132 deletions(-) diff --git a/docs/reference/language/dimensions.md b/docs/reference/language/dimensions.md index 2bada4d0..420904d9 100644 --- a/docs/reference/language/dimensions.md +++ b/docs/reference/language/dimensions.md @@ -62,13 +62,13 @@ engine raises an error rather than build a model with one snapshot dropped. ## `relations` -A relation is what makes topology _data_: a generator sits on a bus, a line has -two endpoints, a snapshot falls in a period, and no adjacency matrix or -hand-written join appears anywhere. A relation is a **table with one column per -dimension it relates**, and `key:` is the claim that makes it a map: one row per -key tuple, so the other columns are a function of the key. The declaration -fixes no direction; the operator that walks the table says which column it -consumes and which it produces. +A relation is a **table with one column per dimension it relates**: a +generator's bus, a snapshot's period, or the buses a generator may connect to. +It is what makes topology data, so no adjacency matrix and no hand-written join +appears anywhere. `key:` names the columns that identify a row. With a key the +table is a map, and the other columns are a function of the key. The +declaration fixes no direction. The operator that walks the table says which +column it consumes and which it produces. ```yaml dimensions: @@ -85,21 +85,23 @@ relations: connection: { columns: [generator, bus] } # no key: a generator may connect to several buses ``` -| Field | | | -| ------------- | ------------------------------------------------------------------------------------------------------------------------------------ | -------------- | -| `over` | required — the columns: a list of dimensions, or a mapping of column name to dimension where two columns share one ([roles](#roles)) | | -| `into` | not a field: a relation declares no direction | | -| `key` | the columns a row is identified by, one name or a list; omitted, the table is a bare relation ([below](#the-key-is-the-claim)) | default none | -| `description` | free text, never parsed | default `null` | - -Every column is over a declared dimension, and its values are checked against -that dimension's labels once data is bound — the check that makes `sum(by=)` -safe, and the reason a label set the model only ever _selects_ on is declared -as a dimension all the same: nothing is indexed by `period` above, and -`where: "period_of == 1"` ([where strings](expressions.md#where-strings)) is -how a declaration selects on it. A relation has at least two columns; a label on -one dimension is a parameter over it. A column named like a dimension is over -that dimension, so `columns: {bus: line}` is refused. +| Field | | | +| ------------- | --------------------------------------------------------------------------------------------------------------------------- | -------------- | +| `columns` | required — a list of dimensions, or a mapping of column name to dimension where two columns share one ([roles](#roles)) | | +| `key` | the columns that identify a row, one name or a list; omitted, the table is a bare relation ([below](#the-key-is-the-claim)) | default none | +| `description` | free text, never parsed | default `null` | + +A relation has at least two columns, each over a declared dimension, and each +column name is distinct. A column named like a dimension is over that +dimension, so `columns: {bus: line}` is refused. The key names columns the +relation has, and not all of them. A relation name may not shadow a dimension: +`generator`'s map onto `bus` is `gen_bus`, never a second `bus`. + +A column's values are checked against its dimension's labels when data is +bound. That check is what makes `sum(by=)` safe, and it is why a label set the +model only selects on is declared as a dimension all the same. Nothing above is +indexed by `period`, and `where: "period_of == 1"` +([where strings](expressions.md#where-strings)) selects on it. ### The key is the claim @@ -107,13 +109,12 @@ that dimension, so `columns: {bus: line}` is refused. generator column holds each label once: the table has **one row per generator**, so the other column is a function of it. `key: [generator, period]` says the pair holds each combination once. Neither column need be unique on its -own: a generator appears once per period, and a period once per generator. -The claim is checked at bind: a generator on two buses is refused, where a -`0`/`1` membership parameter would have said so legally and silently +own: a generator appears once per period, and a period once per generator. The +claim is checked at bind: a generator on two buses is refused, where a `0`/`1` +membership parameter would have said so legally and silently ([#161](https://github.com/energy-models/math-spec/issues/161)). The columns -the key determines are the relation's **value columns**. A key has one column per -dimension: it is read at its dimensions, and no frame carries a dimension -twice, so `key: [bus0, bus1]` is refused where both are over `bus`. +the key determines are the relation's **value columns**. A key has one column +per dimension, so `key: [bus0, bus1]` is refused where both are over `bus`. Each cardinality is one declaration, and the key is the side that is one: @@ -124,28 +125,21 @@ Each cardinality is one declaration, and the key is the side that is one: | many-to-many, a generator on several buses | `{columns: [generator, bus]}`, no key | nothing: a row exists, or it does not | | one-to-one | not a claim the language has: a key is one set of columns, so the other side stays many | | -The key is also what decides which walks the table admits: +A bare relation, one with no `key:`, is walked by `sum` alone, with both ends +named, and tested by a bare `where`. That is what a many-to-many relation can +say, and all it can say. -| the walk | needs | because | -| ------------------------------- | ---------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------- | -| `sum(x, by=l, over=a, into=b)` | the key **not** wholly inside the columns the operand fixes — the `into` columns and the columns joined on | a sum adds its rows up; walked to the key it finds one per coordinate, which is a read | -| `at(x, by=l, over=a, into=b)` | a key inside the columns the operand fixes — the `into` columns and the columns joined on | a read is one value per coordinate, or it is not a read | -| `shift`, `sum_back`, `position` | a key column over the dimension walked | a coordinate is in one group, or it has no neighbour | -| `where: "l == 'north'"` | a key, and the column compared a value column | a comparison is one value per coordinate | -| `where: l` (bare) | nothing | a row exists, or it does not | +### Walks -A bare relation — no `key:` — is walked by `sum` alone, with both ends named, -and tested by a bare `where`. That is what a many-to-many relation can say, -and all it can say. +A walk consumes one column of a relation, produces another, and joins on every +other key column. The operand carries each joined dimension. The result keeps +it, and keeps every dimension the relation does not name. -### A walk names its ends +`sum` consumes a key column and produces a value column. `at` consumes a value +column and produces the key. -Every operator that takes `by=` walks the table between two of its columns: -`over=` the column **consumed**, `into=` the column **produced**, and every other -**key** column **joined on** — the operand carries its dimension and the -result keeps it. A value column not walked is not read: `ends` below, walked -from `line` to `bus1`, joins on nothing. A bare relation's columns are all -key, so all of them but the two walked are joined on. +`over=` names the column consumed and `into=` the column produced. Name a column +only where the relation offers two. ```yaml dimensions: @@ -160,81 +154,75 @@ parameters: variables: p: { dims: [generator, period] } constraints: - zone_balance: # p[generator, period] → [zone, period] + zone_balance: # consumes generator, joins on period, produces zone: [generator, period] → [zone, period] dims: [zone, period] - expression: sum(p, by=zone_of, over=generator, into=zone) >= demand - history: # p[generator, period] → [generator, zone]: the same table, walked from its other key column + expression: sum(p, by=zone_of, over=generator) >= demand + history: # consumes period, joins on generator, produces zone: [generator, period] → [generator, zone] dims: [generator, zone] - expression: sum(p, by=zone_of, over=period, into=zone) <= 100 - capped_revenue: # price[zone, period] → [generator, period]: the price of the zone this generator sat in that period + expression: sum(p, by=zone_of, over=period) <= 100 + capped_revenue: # consumes zone, joins on period, produces generator: [zone, period] → [generator, period] dims: [generator, period] expression: at(price, by=zone_of, over=zone, into=generator) * p <= 1000 ``` -**What the declaration decides, the call may leave unsaid.** Where the key has -one column and the key determines one column, the walk is the arrow the key -draws, and `sum(p, by=gen_bus)` and `at(price, by=gen_bus)` are complete: -`sum` consumes the key and produces the value, `at` consumes the value and -produces the key. Where a side has several candidates — two key columns, two -value columns — the call names it, and the refusal lists the candidates. -`zone_of` above has two key columns, so `sum` names `over=`, while `into=zone` -could have been left out. - -**A partition walks a key column and groups by the value columns.** -`shift(x, along=d, by=l)`, `sum_back(x, along=d, by=l)` and -`position(d, by=l)` take the one key column over `d`; the other key columns -are joined on, and the group is the value tuple. `within=` names the value columns the group is made of -where the table has several: `shift(x, along=snapshot, by=cal, within=week)` -walks within weeks of a calendar declared once over `[snapshot, day, week]`, -and a value column not named is not read. - -The rules, each decided at load with a refusal naming the rewrite: - -- **`over=` and `into=` name columns of the relation `by=` names**, one each or a - list each, and no column on both sides. `into=` is refused without a `by=`, - since a column needs the table that holds it. `over=` without one names a - dimension of the operand instead, which is `sum(p, over=period)`. - `sum(p, by=gen_bt, into=[bus, technology])` lands one table with two value - columns on the product `bus × technology` in one join; - `sum(p, by=zone_of, over=[generator, period])` consumes both key columns - at once, which is `sum(sum(p, by=zone_of, over=generator), over=period)` - said once; `at(tech_cap, by=gen_bt, over=[bus, technology])` reads a - two-column slot at each generator. -- **The operand carries every joined column's dimension, each once.** The map - is read at the key columns not walked, so there is no reading it at a - coordinate that lacks them; two joined columns over one dimension have - nothing to tell apart. -- **A produced dimension the operand already carries is joined on too.** +`zone_of` has one value column, so `into=zone` on `sum` and `over=zone` on `at` +may be left out. It has two key columns, so `sum` names the one it consumes, +because `over=generator` and `over=period` are different constraints, and `at` +names the one it produces. `period` is joined on either way. With one key column +and one value column, `sum(p, by=gen_bus)` and `at(price, by=gen_bus)` are +complete. A column left out where the relation offers two is refused, and the +message lists the candidates: + +``` +sum(by=zone_of): 'zone_of' has 2 key columns (['generator', 'period']), and the call has to say which over= names. +``` + +`capped_revenue` reads the price of the zone this generator sat in that period. +The typesetter prints it as $`\mathrm{price}_{\mathrm{zone\_of}(g,\ e),e}`$, +and the joined `period` is the second subscript. + +- **Either keyword takes a list.** `sum(p, by=gen_bt, into=[bus, technology])` + lands on the product `bus × technology` in one join. + `sum(p, by=zone_of, over=[generator, period])` consumes both key columns at + once. `at(tech_cap, by=gen_bt, over=[bus, technology])` reads a two-column + slot at each generator. +- **A produced dimension the operand already carries is joined on.** `sum(load * p, by=gen_bus)` with `load[snapshot, bus]` restricts each term to - the row where the generator's bus is the row's bus — a masked sum, which is - what the join says. -- **`at` reads one value.** Its key lies inside `into=` and the joined columns, - or the call is refused; a bare relation is never read by `at`. -- **`sum` adds its rows up.** So the reverse holds: a `sum` whose key lies inside - `into=` and the joined columns finds one row per coordinate and adds up - nothing, which is a read — it is refused toward `at`. `sum` walks to a value - column; `at` walks to the key. -- **A partition walks the one key column over the dimension it walks, and - groups by the value columns `within=` names** — all of them where it names - none. `within=` naming a key column is refused, and a bare relation - partitions nothing. The group may hold two columns over one dimension, a - pair of buses say: a partition lands nothing, so nothing needs the - dimension twice. -- **A `by=` list walks each relation by its declared arrow.** `by=[a, b]` is one - grouping, so no column keyword has anything to name; every relation in it - consumes the same dimension, joins on its own other columns, and no two - produce the same dimension. -- **A `where` comparison reads a value column of a keyed relation at its key.** - `zone_of == 'north'` reads the one value column; `ends.bus0 != ends.bus1` - names the columns where there are several. The frame carries the key's - dimensions, and two relations compared have keys over the same dimensions and - columns over one. A bare name — `where: gen_bus` — tests that a row exists: - at the key for a keyed relation, at every column for a bare relation. -- **Every column is over a declared dimension, every column name is distinct, - the key names columns the relation has, and does not name all of them.** - -**Every relation name joins the flat namespace**, so a relation may not shadow a -dimension. `generator`'s map onto `bus` is `gen_bus`, never a second `bus`. + the row where the generator's bus is the row's bus, which is a masked sum. +- **A value column that is not walked is not read.** `ends`, walked from `line` + to `bus1`, joins on nothing ([roles](#roles)). +- **`by=[a, b]` is one grouping.** Each relation walks by its declared key and + value, so no column keyword has anything to name. Every relation in the list + consumes the same dimension, and no two produce the same one. +- **`into=` needs a `by=`**, since a column needs the table that holds it. + `over=` without one names a dimension of the operand, which is + `sum(p, over=period)`. + +The loader refuses a walk that is not one, and the message names the rewrite: + +| refused | message | +| ------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `at` on a bare relation | `at(by=connection): at reads one value per coordinate, and 'connection' is not single-valued in ['bus'] at the columns the operand fixes (['generator']) — its key is []. Declare a key those columns contain, or read the other way.` | +| a `sum` that consumes no key column | `sum(by=zone_of): this sum walks to the key ['generator', 'period'], so each coordinate has one term and nothing is added up — that is a read, which is at()'s. Write at(..., by=zone_of, over=['zone'], into=['generator']), or sum toward a value column.` | +| an operand missing a joined dimension | `at(by=zone_of) joins on ['period'] (columns ['period'] of 'zone_of'), which the expression does not carry (dims ['zone']). A relation is walked between two of its columns and read at the others — index the operand by them, or walk between different columns.` | + +### Partitions + +`shift(x, along=d, by=l)`, `sum_back(x, along=d, by=l)` and `position(d, by=l)` +walk the one key column over `d`, join on the other key columns, and group by +the value columns. The frame does not change: the group says which rows are +neighbours, and nothing lands anywhere. + +`within=` names the value columns the group is made of where the table has +several. `shift(x, along=snapshot, by=cal, within=week)` walks within weeks of a +calendar declared once over `[snapshot, day, week]`, and a value column not +named is not read. The group may hold two columns over one dimension, a pair of +buses say, since a partition lands nothing. `within=` naming a key column is +refused, and a bare relation partitions nothing. + +A `where` string reads a relation too: a value column at its key, two columns +of one table compared, or a bare name that tests a row exists +([where strings](expressions.md#where-strings)). ### Roles diff --git a/docs/reference/language/operators.md b/docs/reference/language/operators.md index 757aeb11..31040a01 100644 --- a/docs/reference/language/operators.md +++ b/docs/reference/language/operators.md @@ -78,14 +78,11 @@ constraints: The same `f` is summed twice through two relations, once as inflow and once as outflow, with no adjacency matrix and no join written by hand. -`by=` and `over=` compose: `by=` names the table and `over=` names what -leaves the frame, so a call may give both, either, or neither. - -`over=` and `into=` say [which columns the walk runs between](dimensions.md#a-walk-names-its-ends) -where the declaration leaves a choice. Every other key column is joined on, so -the operand carries it, the sum keeps it, and each group is one coordinate of -it. A value column that is not walked is not read. A bare relation, one with no -`key:`, is summed with both ends named, and a row it holds twice counts twice. +`sum(by=)` consumes a key column and produces a value column. `over=` and +`into=` name them where the relation offers two ([walks](dimensions.md#walks)), +and every other key column is joined on, so each group is one coordinate of it. +A bare relation, one with no `key:`, is summed with both ends named, and a row +it holds twice counts twice. The relation's values are the group labels, checked against their own dimension when the data binds. A group with no members contributes nothing, and a member @@ -95,21 +92,17 @@ coordinate the data never covered is refused. See [absence](absence.md). ## `at` -`at(x, by=l)` walks the same relation the other way. `sum(by=)` consumes the -key column and produces the value column. `at` consumes the value column and -produces the key column: it reads one coarse value once for each fine label that -points at it. `over=` and `into=` name the two columns where the key leaves a -choice. A read is one value per coordinate, so the relation's key must lie inside -`into=` and the columns joined on, and a bare relation is never read by `at`. +`at(x, by=l)` walks the same relation the other way. It consumes a value column +and produces the key, so it reads one coarse value once for each fine label that +points at it, and a bare relation is never read by `at`. `over=` and `into=` +name the columns where the relation offers two, and every other key column is +read at the row's own coordinate ([walks](dimensions.md#walks)). `at` reads a variable as readily as a parameter. One decision taken per bus, read once by every line that touches the bus, is `at(decision, by=line_bus)`. A fine label whose relation value is null reads nothing, and its row is absent. -That matches the null group in `sum(by=)`. Through a relation with a -[column joined on](dimensions.md#a-walk-names-its-ends) `at` reads the coarse -value at the row's own coordinate of that column, which is the price of the zone -this generator sat in that period. +That matches the null group in `sum(by=)`. ## `sum_back` From dea5a331e0cc19d8392f683cc520001c0e7e7133 Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 16 Sep 2026 08:56:52 +0000 Subject: [PATCH 2/4] docs(language): a walk consumes one or more columns, as the list forms already say Co-Authored-By: Claude Fable 5.1 Claude-Session: https://claude.ai/code/session_019vg9UDLgdaFiu9gtbqAj7P --- docs/reference/language/dimensions.md | 10 +++++----- 1 file changed, 5 insertions(+), 5 deletions(-) diff --git a/docs/reference/language/dimensions.md b/docs/reference/language/dimensions.md index 420904d9..41b3145d 100644 --- a/docs/reference/language/dimensions.md +++ b/docs/reference/language/dimensions.md @@ -131,12 +131,12 @@ say, and all it can say. ### Walks -A walk consumes one column of a relation, produces another, and joins on every -other key column. The operand carries each joined dimension. The result keeps -it, and keeps every dimension the relation does not name. +A walk consumes one or more columns of a relation, produces one or more, and +joins on every other key column. The operand carries each joined dimension. The +result keeps it, and keeps every dimension the relation does not name. -`sum` consumes a key column and produces a value column. `at` consumes a value -column and produces the key. +`sum` consumes key columns and produces value columns. `at` consumes value +columns and produces the key. `over=` names the column consumed and `into=` the column produced. Name a column only where the relation offers two. From 8b6c2dd1eb8eb61dc725a14ec9b451e6a5039db0 Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 16 Sep 2026 08:57:49 +0000 Subject: [PATCH 3/4] docs(language): the example says which column each operator leaves unsaid, one operator at a time Co-Authored-By: Claude Fable 5.1 Claude-Session: https://claude.ai/code/session_019vg9UDLgdaFiu9gtbqAj7P --- docs/reference/language/dimensions.md | 15 ++++++++------- 1 file changed, 8 insertions(+), 7 deletions(-) diff --git a/docs/reference/language/dimensions.md b/docs/reference/language/dimensions.md index 41b3145d..dc073e0f 100644 --- a/docs/reference/language/dimensions.md +++ b/docs/reference/language/dimensions.md @@ -165,13 +165,14 @@ constraints: expression: at(price, by=zone_of, over=zone, into=generator) * p <= 1000 ``` -`zone_of` has one value column, so `into=zone` on `sum` and `over=zone` on `at` -may be left out. It has two key columns, so `sum` names the one it consumes, -because `over=generator` and `over=period` are different constraints, and `at` -names the one it produces. `period` is joined on either way. With one key column -and one value column, `sum(p, by=gen_bus)` and `at(price, by=gen_bus)` are -complete. A column left out where the relation offers two is refused, and the -message lists the candidates: +`zone_of` has one value column, `zone`. It is the only column `sum` can produce, +so `sum` leaves `into=zone` unsaid. It is the only column `at` can consume, so +`at` may leave `over=zone` unsaid too. `zone_of` has two key columns, and there +the call chooses: `sum` names the one it consumes, because `over=generator` and +`over=period` are different constraints, and `at` names the one it produces. +`period` is joined on either way. With one key column and one value column, +`sum(p, by=gen_bus)` and `at(price, by=gen_bus)` are complete. A column left out +where the relation offers two is refused, and the message lists the candidates: ``` sum(by=zone_of): 'zone_of' has 2 key columns (['generator', 'period']), and the call has to say which over= names. From f652f36c4252b58ef0763a380337f34d57c64322 Mon Sep 17 00:00:00 2001 From: Claude Date: Wed, 16 Sep 2026 09:02:03 +0000 Subject: [PATCH 4/4] docs(language): the relations page says each rule in things before abstractions, and keeps rationale out Twelve passages rewritten on dimensions.md: the bind-check paragraph under the field table, the cardinality lead-in that called the key the one side, the list and masked-sum bullets, the self-map paragraph, the data-supply section, and the closing table's row on selected-on label sets. Two pieces of rationale leave the page: the comparison to a 0/1 membership parameter, and the history of the nodal balance through two relations. One sentence that repeated rule 1 of where the members come from is cut. Sentence measure on dimensions.md: 104 sentences, median 16 words, 13 over 25, from 96, 16 and 18 at the previous commit. Verified from the uv environment: prettier and typos clean, mkdocs build --strict clean with the Python inventory dropped, and tests/test_docs.py plus tests/test_reading_page.py give 34 passed. The rewritten self-map claim is checked against the loader: comparing the self-map's value column to its key column is refused in every spelling. Co-Authored-By: Claude Fable 5.1 Claude-Session: https://claude.ai/code/session_019vg9UDLgdaFiu9gtbqAj7P --- docs/reference/language/dimensions.md | 136 +++++++++++++------------- 1 file changed, 66 insertions(+), 70 deletions(-) diff --git a/docs/reference/language/dimensions.md b/docs/reference/language/dimensions.md index dc073e0f..44f3d469 100644 --- a/docs/reference/language/dimensions.md +++ b/docs/reference/language/dimensions.md @@ -93,15 +93,16 @@ relations: A relation has at least two columns, each over a declared dimension, and each column name is distinct. A column named like a dimension is over that -dimension, so `columns: {bus: line}` is refused. The key names columns the -relation has, and not all of them. A relation name may not shadow a dimension: +dimension, so `columns: {bus: line}` is refused. A key names columns of the +relation, and never all of them. A relation name may not shadow a dimension: `generator`'s map onto `bus` is `gen_bus`, never a second `bus`. A column's values are checked against its dimension's labels when data is -bound. That check is what makes `sum(by=)` safe, and it is why a label set the -model only selects on is declared as a dimension all the same. Nothing above is -indexed by `period`, and `where: "period_of == 1"` -([where strings](expressions.md#where-strings)) selects on it. +bound, so a mistyped bus is refused rather than summed into a group of its own. +That is why a label set the model only selects on is still declared as a +dimension. Nothing above is indexed by `period`, but `period` is declared, so +`where: "period_of == 1"` ([where strings](expressions.md#where-strings)) +compares against a checked label. ### The key is the claim @@ -110,20 +111,19 @@ generator column holds each label once: the table has **one row per generator**, so the other column is a function of it. `key: [generator, period]` says the pair holds each combination once. Neither column need be unique on its own: a generator appears once per period, and a period once per generator. The -claim is checked at bind: a generator on two buses is refused, where a `0`/`1` -membership parameter would have said so legally and silently +claim is checked at bind, so a generator on two buses is refused ([#161](https://github.com/energy-models/math-spec/issues/161)). The columns the key determines are the relation's **value columns**. A key has one column per dimension, so `key: [bus0, bus1]` is refused where both are over `bus`. -Each cardinality is one declaration, and the key is the side that is one: +Each cardinality is one declaration: | to say | write | checked at bind | | ------------------------------------------ | ----------------------------------------------------------------------------------------------------------------- | ------------------------------------- | | many-to-one, each generator on one bus | `{columns: [generator, bus], key: generator}` | one row per generator | | one-to-many, a bus and its generators | the same table: `sum(p, by=gen_bus)` collects a bus's generators, `at(price, by=gen_bus)` reads a generator's bus | the same | | many-to-many, a generator on several buses | `{columns: [generator, bus]}`, no key | nothing: a row exists, or it does not | -| one-to-one | not a claim the language has: a key is one set of columns, so the other side stays many | | +| one-to-one | not a claim the language has: a key is one set of columns, and nothing checks the other side | | A bare relation, one with no `key:`, is walked by `sum` alone, with both ends named, and tested by a bare `where`. That is what a many-to-many relation can @@ -171,8 +171,9 @@ so `sum` leaves `into=zone` unsaid. It is the only column `at` can consume, so the call chooses: `sum` names the one it consumes, because `over=generator` and `over=period` are different constraints, and `at` names the one it produces. `period` is joined on either way. With one key column and one value column, -`sum(p, by=gen_bus)` and `at(price, by=gen_bus)` are complete. A column left out -where the relation offers two is refused, and the message lists the candidates: +`sum(p, by=gen_bus)` and `at(price, by=gen_bus)` need neither keyword. A column +left out where the relation offers two is refused, and the message lists the +candidates: ``` sum(by=zone_of): 'zone_of' has 2 key columns (['generator', 'period']), and the call has to say which over= names. @@ -185,21 +186,23 @@ and the joined `period` is the second subscript. - **Either keyword takes a list.** `sum(p, by=gen_bt, into=[bus, technology])` lands on the product `bus × technology` in one join. `sum(p, by=zone_of, over=[generator, period])` consumes both key columns at - once. `at(tech_cap, by=gen_bt, over=[bus, technology])` reads a two-column - slot at each generator. -- **A produced dimension the operand already carries is joined on.** - `sum(load * p, by=gen_bus)` with `load[snapshot, bus]` restricts each term to - the row where the generator's bus is the row's bus, which is a masked sum. -- **A value column that is not walked is not read.** `ends`, walked from `line` - to `bus1`, joins on nothing ([roles](#roles)). -- **`by=[a, b]` is one grouping.** Each relation walks by its declared key and - value, so no column keyword has anything to name. Every relation in the list - consumes the same dimension, and no two produce the same one. -- **`into=` needs a `by=`**, since a column needs the table that holds it. - `over=` without one names a dimension of the operand, which is - `sum(p, over=period)`. - -The loader refuses a walk that is not one, and the message names the rewrite: + once. `at(tech_cap, by=gen_bt, over=[bus, technology])` reads `tech_cap` at + each generator's bus and technology together. +- **A produced dimension the operand already carries is joined on.** In + `sum(load * p, by=gen_bus)` with `load[snapshot, bus]`, the walk produces + `bus` and `load` already carries it. So each generator's term is read at the + bus the generator sits on, and the sum lands there. +- **A value column that is not walked is not read.** + `sum(f, by=ends, over=line, into=bus1)` reads `bus1` and ignores `bus0` + ([roles](#roles)). +- **`by=[a, b]` is one grouping onto what `a` and `b` produce together.** Each + relation is walked from its key to its value, so `over=` and `into=` have + nothing to name. The relations consume the same dimension, and no two produce + the same one. +- **`into=` needs a `by=`**, because a column belongs to a table. `over=` + without a `by=` names a dimension, as in `sum(p, over=period)`. + +Three refusals draw the line, and each message names the rewrite: | refused | message | | ------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | @@ -217,9 +220,9 @@ neighbours, and nothing lands anywhere. `within=` names the value columns the group is made of where the table has several. `shift(x, along=snapshot, by=cal, within=week)` walks within weeks of a calendar declared once over `[snapshot, day, week]`, and a value column not -named is not read. The group may hold two columns over one dimension, a pair of -buses say, since a partition lands nothing. `within=` naming a key column is -refused, and a bare relation partitions nothing. +named is not read. The group may be two columns over one dimension, such as a +line's two buses, because a partition produces no dimension. `within=` naming a +key column is refused, and a bare relation partitions nothing. A `where` string reads a relation too: a value column at its key, two columns of one table compared, or a bare name that tests a row exists @@ -237,15 +240,14 @@ relations: ``` `sum(f, by=ends, over=line, into=bus1) - sum(f, by=ends, over=line, into=bus0)` -is the nodal balance through one table where two relations did it before, and -`where: "ends.bus0 != ends.bus1"` excludes a self-loop by comparing two of its -columns. +is a nodal balance through one table: flow arriving at `bus1` less flow leaving +`bus0`. `where: "ends.bus0 != ends.bus1"` excludes a line whose two ends are +one bus. -`rep_of` relates a dimension to itself, which is how a clustered year is run -on a few typical days: every snapshot names the one that stands for it. -Nothing changes in the rules — `snapshot` is consumed and `rep` produced, both -over one dimension, so the frame is unchanged through `sum(by=)` and `at(by=)` -alike: +`rep_of` relates a dimension to itself. That is how a clustered year runs on a +few typical days: every snapshot names the snapshot that stands for it. The +rules are the same. `sum` consumes `snapshot` and produces `rep`, both over +`snapshot`, so the frame is `[snapshot]` before the walk and after it: ```yaml constraints: @@ -257,21 +259,19 @@ constraints: expression: sum(p, by=rep_of) <= 100 ``` -**A self-map is directional exactly as far as its key says.** `key: snapshot` -makes `rep` a function of `snapshot`, so the arrow runs from a snapshot to its -representative: `at` reads along it and `sum` collects against it, the inverse -of a many-to-one map being one-to-many, reachable as a grouping and never as -a function. Two steps along the arrow are two nested calls. Without a key the -same two columns are an undirected relation — a neighbour table — which `sum` -walks either way and nothing reads. Selecting the representatives themselves, -the rows where the map is the identity, is not a comparison the language has, -since a relation is never compared to a dimension; declare a `bool` parameter -for them. +**The key directs a self-map.** `key: snapshot` makes `rep` a function of +`snapshot`. `at` reads each snapshot's representative, and `sum` collects onto +a representative the snapshots it stands for. Two steps along the map are two +nested calls. Without a key the same two columns are a neighbour table, which +`sum` walks either way and nothing reads. The rows where a snapshot is its own +representative cannot be selected with a `where`, because a value column is +never compared to the frame's own coordinate. Declare a `bool` parameter for +them. ### How the map is supplied -`gen_bus` is a source key like any other, carrying one column per column -declared, named after the column: +The data for `gen_bus` arrives under the key `gen_bus`, as a table with one +column per declared column, named after it: ```python sources = { @@ -281,18 +281,15 @@ sources = { ``` **A partial map is the rows it has.** `g3` is in no row, so `g3` sits on no -bus — absence is the absent row, exactly as it is for a parameter, and a null -in any column is refused for saying both at once. A keyed table holds one row -per key tuple, and a value matching no label of its column's dimension is a -typo rather than a new member. Values are never inferred from the parameters -that use a dimension: inferring would let a mistyped label extend the label -set instead of being rejected. - -Supplying it this way touches no table but its own, which is what a caller who -did not generate the index needs: a model can be extended with a relation the -same way it can be extended with a parameter. **A column of a dimension's index -named after a relation is refused** rather than read — an index may carry any -other extra, and this one would be a map read by accident. +bus. Absence is the missing row, as it is for a parameter. A null in any column +is refused, because a row that is present and empty says both at once. A keyed +table holds one row per key tuple. A value that matches no label of its +dimension is refused as a typo, never added as a member. + +A relation's table stands on its own, so a model gains a relation the way it +gains a parameter: one more table, and no change to the others. **A column named +after a relation inside a dimension's table is refused.** That table may carry +other extra columns, but this one would be a map read by accident. ## Dimension, relation or parameter? @@ -305,19 +302,18 @@ does with the column, not what the column holds: | has one value per member of a dimension, or per tuple of several — a generator's bus, a line's two ends, a generator's zone by period | a `relation` with that `key` | it is a map every operator walks, and its values are checked against the dimensions they name | | relates members of two dimensions many-to-many, with nothing to weigh — which buses a generator may connect to | a `relation` with no key | `sum` walks it with both ends named, and a bare `where` tests it. Nothing reads it, because there is no one value to read | | relates members of two dimensions many-to-many, with a weight per pair — a link's efficiency to each bus, a cycle's lines | a `parameter` over both | the weight is the data, its row set is the relation, and the aggregation is `sum(w * x, over=a)` | -| is a label set the model only selects on or counts within — a period, a season, a zone | a `dimension`, and a keyed `relation` onto it | the membership check is worth one line and one member list | +| is a label set the model only selects on or counts within — a period, a season, a zone | a `dimension`, and a keyed `relation` onto it | its labels are checked, at the cost of one line and one table | | scales terms — a coefficient, a bound, an offset | a `parameter` (`float` or `int`) | arithmetic is over numbers ([dtype](declarations.md#parameters)) | | is a per-row attribute the math only selects on — a fuel, a constraint's sense | a `str` parameter | it names rows rather than scaling them, and no set is declared to check its values against | | is a mask | a `bool` parameter | a bare name in a `where` is its own answer | -Two rules follow from the table. If `b` has one value per `a`, then `b` is a -**relation** keyed by `a`, and not a dimension: a `dims` product over two -dimensions that depend on each other, cut back with a mask, is the shape that -`relations` replaces. +Two rules follow. If `b` has one value per `a`, declare `b` as a **relation** +keyed by `a`, not as a dimension. Two dimensions that depend on each other, +declared as a `dims` product and cut back with a mask, are one relation. -And everything under `dimensions:` is an axis. A dimension is never legal where -a value belongs, because it is a coordinate space and not data. To use a -dimension's coordinates as data, declare a parameter over it. +Everything under `dimensions:` is an axis. A dimension is never legal where a +value belongs, because it is a coordinate space and not data. To use a +dimension's labels as data, declare a parameter over it. `python -m math_spec check` advises on a declared dimension that nothing is indexed by, nothing aggregates into and no relation has a column over ([errors](errors.md#what-advice-warns-about)).