Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
96 changes: 81 additions & 15 deletions docs/reference/language/dimensions.md
Original file line number Diff line number Diff line change
Expand Up @@ -85,13 +85,13 @@ lookups:
period_of: { over: snapshot, into: period }
```

| Field | | |
| ------------- | -------------------------------------------------------------------- | -------------- |
| `over` | required — the dimension whose members carry the map | |
| `into` | required — the dimension its values are labels of, other than `over` | |
| `description` | free text, never parsed | default `null` |
| Field | | |
| ------------- | ------------------------------------------------------------------------------------------------------------------------------- | -------------- |
| `over` | required — the map's key dimensions: one, or a list in the order the table carries them ([below](#keyed-by-several-dimensions)) | |
| `into` | required — the dimension its values are labels of, which is not a key | |
| `description` | free text, never parsed | default `null` |

The target must be a declared dimension, and it must differ from `over`. The
The target must be a declared dimension, and it must be none of the keys. The
values are checked against it when the data binds, which is the check that makes
`sum(by=)` safe.

Expand All @@ -107,13 +107,66 @@ error.

Several lookups may group at once. `sum(x, by=[gen_bus, gen_tech])` groups
through both maps in one reduction and lands on `bus` and `technology`. Every
lookup in the list must be `over:` the same dimension, and each must target a
lookup in the list must walk the same dimension, and each must target a
different one. A member that either map leaves out belongs to no group.

Every lookup name joins the flat namespace, so a lookup may not shadow a
dimension, and that includes its own target. The map from `generator` onto `bus`
is called `gen_bus`, never a second `bus`.

### Keyed by several dimensions

A map keyed by one dimension gives every generator one zone for the whole
model. A generator whose bidding zone changes by period needs a second key, and
`over:` takes a list of them:

```yaml
dimensions:
generator: { dtype: str }
zone: { dtype: str }
period: { dtype: int }
lookups:
zone_of: { over: [generator, period], into: zone }
parameters:
demand: { dims: [zone, period] }
variables:
p: { foreach: [generator, period] }
constraints:
zone_balance:
foreach: [zone, period]
expression: sum(p, by=zone_of.generator) >= demand
```

A call walks one key and joins on the rest. The dot says which:
`by=zone_of.generator` consumes `generator`, produces `zone`, and joins on
`period`. So `sum(p, by=zone_of.generator)` takes `p[generator, period]` to
`[zone, period]`, and `at(price, by=zone_of.generator)` reads
`price[zone, period]` back at `[generator, period]`, which is the price of the
zone this generator sat in that period. The same table walked along its other
key is a different sum: `sum(p, by=zone_of.period)` takes `p` to
`[generator, zone]`, each generator's output over the periods it spent in each
zone.

Six rules follow, and the loader decides each of them before any data binds:

- **The dot names a key.** Write it wherever the lookup has more than one key.
Without it the call is refused, because the operator cannot know which key it
consumes. With one key the dot is redundant and legal, so `by=gen_bus` and
`by=gen_bus.generator` are the same call.
- **The operand carries every key but the one walked.** The map is read at
those keys, so there is no reading it at a coordinate that lacks them.
- **The walked key is the walked dimension.** `shift(x, over=d, by=l.k)`,
`sum_back(x, over=d, by=l.k)` and `position(d, by=l.k)` need `k` to be `d`.
Each groups the rows of `d` within one coordinate of the other keys.
- **A `by=` list walks one dimension.** `by=[a.k, b.k]` is one grouping, so
every lookup in it names the same key dimension. Each joins on its own other
keys.
- **A `where` reads every key.** `zone_of == 'north'`, a bare `zone_of` and
`zone_of != area_of` are filters on the key table, so the frame carries all of
a lookup's keys, and two lookups compared carry the same keys.
- **Each key is a declared dimension, named once.** The target is not one of
them.

### How the map is supplied

The map is a source key like any other, under the lookup's own name. It carries
Expand All @@ -132,6 +185,10 @@ on no bus. A null in the value column is refused, because a missing row already
says the same thing. A key that matches no label of `over` is an error rather
than a new member.

A map with several keys carries one column per key, named after its dimension,
and is single-valued per key tuple. A generator in two zones in one period is
refused, where a `0`/`1` membership parameter says it legally and silently.

Values are never inferred from the parameters that use the target. If they were,
a mistyped label would extend the label set instead of being rejected.

Expand All @@ -144,14 +201,23 @@ but a column named after the lookup is refused rather than read.
Every column of data is one of the three. What decides which is what the math
does with the column, not what the column holds:

| The column… | is declared as | because |
| -------------------------------------------------------------------------------------- | ------------------------------------- | --------------------------------------------------------------------------------------------- |
| is an axis: something is indexed by it, or an aggregation lands terms on it | a `dimension` | its members are the coordinate set every table over it is reindexed onto |
| has one value per member of a dimension and points at another — a generator's bus | a `lookup` into that dimension | it is a map that `sum(by=)` and `at(by=)` walk, and its values are checked against the target |
| is a label set the model only selects on or counts within — a period, a season, a zone | a `dimension`, and a `lookup` into it | the membership check is worth one line and one member list |
| scales terms — a coefficient, a bound, an offset | a `parameter` (`float` or `int`) | arithmetic is over numbers ([dtype](declarations.md#parameters)) |
| is a per-row attribute the math only selects on — a fuel, a constraint's sense | a `str` parameter | it names rows rather than scaling them, and no set is declared to check its values against |
| is a mask | a `bool` parameter | a bare name in a `where` is its own answer |
| The column… | is declared as | because |
| ----------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------- |
| is an axis: something is indexed by it, or an aggregation lands terms on it | a `dimension` | its members are the coordinate set every table over it is reindexed onto |
| has one value per member of a dimension, or per tuple of several, and points at another — a generator's bus, a generator's zone by period | a `lookup` into that dimension | it is a map that `sum(by=)` and `at(by=)` walk, and its values are checked against the target |
| relates members of two dimensions many-to-many — a link's several buses with their efficiencies, a cycle's lines | a `parameter` over both | `bool` where it only selects, numeric where it weights. The aggregation is `sum(w * x, over=a)`, and a pair the table lacks is absent |
| is a label set the model only selects on or counts within — a period, a season, a zone | a `dimension`, and a `lookup` into it | the membership check is worth one line and one member list |
| scales terms — a coefficient, a bound, an offset | a `parameter` (`float` or `int`) | arithmetic is over numbers ([dtype](declarations.md#parameters)) |
| is a per-row attribute the math only selects on — a fuel, a constraint's sense | a `str` parameter | it names rows rather than scaling them, and no set is declared to check its values against |
| is a mask | a `bool` parameter | a bare name in a `where` is its own answer |

**A many-to-many relation is a parameter, weighted or not.** Pairs alone are a
`bool` parameter, written `connection: {dims: [entity, bus], dtype: bool}` and
read with `where: connection`. Pairs with a weight are a numeric one, and the
aggregation needs no lookup: `sum(efficiency * p, over=entity)` lands on `bus`,
because `efficiency[entity, bus]` has a row exactly where the pair exists. A
lookup is the single-valued case, where the language checks a claim a parameter
cannot make.

Two rules follow from the table. If `b` has one value per `a`, then `b` is a
**lookup** over `a`, and not a dimension: a `foreach` product over two
Expand Down
6 changes: 3 additions & 3 deletions docs/reference/language/expressions.md
Original file line number Diff line number Diff line change
Expand Up @@ -91,7 +91,7 @@ fixed at load:
| --------------------------------------- | ------------------------------------------------------------------------------------------------------------ |
| expression (`p * cost`) | a variable, or a parameter whose values are numbers ([dtype](declarations.md#parameters)) |
| dimension argument (`over=`) | a dimension |
| lookup argument (`by=` on `sum` / `at`) | a lookup, and never a dimension |
| lookup argument (`by=` on `sum` / `at`) | a lookup, dotted with the key it walks where it has several, and never a dimension |
| `where` string | a parameter, variable, dimension or lookup ([where strings](#where-strings)) |
| `bounds.lower` / `bounds.upper` | a parameter name, or a number |
| the `edge` key of `shift` | `'wrap'` in quotes, or a bare number. Never a dimension |
Expand Down Expand Up @@ -172,8 +172,8 @@ QUOTED ::= "'" chars "'" | '"' chars '"'
| `name` (bare) | dimension | A load error. It would be true everywhere. Compare it against something instead |
| `name OP value` | parameter | Element-wise, and a null compares false. The right-hand side is a literal, or a bare name read as a string label |
| `name OP value` | dimension | A filter on the frame's own coordinate column |
| `name OP value` | lookup | A filter on the lookup's value, so the `over` dimension has to be in the frame. A null compares false |
| `name OP name` | two lookups | Legal only where both lookups are over the same dimension and into the same dimension. `from != to` excludes a self-loop |
| `name OP value` | lookup | A filter on the lookup's value, so [every key](dimensions.md#keyed-by-several-dimensions) has to be in the frame. A null compares false |
| `name OP name` | two lookups | Legal only where both lookups have the same keys and map into the same dimension. `from != to` excludes a self-loop |
| `position(name) OP i` | dimension | Where the row sits along the dimension's own order. `0` is first, and a negative number counts from the end |
| `position(name, by=lookup) OP i` | a dimension and a lookup over it | The same, counted within each group the lookup makes |
| `AND` `OR` `NOT` | — | Case-insensitive. `NOT` binds tighter than `AND`, and `AND` tighter than `OR` |
Expand Down
9 changes: 9 additions & 0 deletions docs/reference/language/operators.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,7 @@ model can never depend on what a caller registered. A composition of them goes i
| `sum(array, over=dim)` | `dim` collapses. `array` must carry `dim` |
| `sum(array, by=lookup)` | The dimension that the lookup is over collapses onto the dimension it maps into |
| `sum(array, by=[lookup, …])` | The same, onto every dimension that the lookups map into. All the lookups must be over the same dimension |
| `sum(array, by=lookup.key)` | A lookup keyed by several dimensions is walked along the named key. The others are joined on, so the array carries them and the result keeps them |
| `at(array, by=lookup)` | The dimension that the lookup maps into is replaced by the dimension it is over |
| `shift(array, over=dim, offset=n)` | The value `n` positions earlier along `dim`. The vacated edge is **absent** |
| `shift(array, over=dim, offset=n, edge='wrap')` | The value `n` positions earlier, counted cyclically, so nothing is vacated |
Expand Down Expand Up @@ -77,6 +78,10 @@ outflow, with no adjacency matrix and no join written by hand.
Give **at most one** of `over=` and `by=`. A lookup carries its own dimensions,
so `by=` leaves `over=` nothing to add.

A lookup [keyed by several dimensions](dimensions.md#keyed-by-several-dimensions)
is walked along the key the dot names and joined on the others. The operand
carries them, the sum keeps them, and each group is one coordinate of them.

The lookup's values are the group labels, checked against the target dimension
when the data binds. A group with no members contributes nothing, and a member
whose lookup value is null belongs to no group. An empty group is a value rather
Expand All @@ -96,6 +101,10 @@ once by every line that touches the bus, is `at(decision, by=line_bus)`.
A fine label whose lookup value is null reads nothing, and its row is absent.
That matches the null group in `sum(by=)`.

Through a lookup [keyed by several dimensions](dimensions.md#keyed-by-several-dimensions)
`at` reads the coarse value at the row's own coordinate of the keys not walked.
That is the price of the zone this generator sat in that period.

## `sum_back`

`sum_back(x, over=d, within=n)` is the sum of the last `n` positions along `d`,
Expand Down
Loading
Loading