Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
21 commits
Select commit Hold shift + click to select a range
449424e
docs(about): a relation is an indicator, and a sum through it is a co…
claude Sep 22, 2026
5449764
feat(language): a relation call is described as the columns it joins …
claude Sep 22, 2026
0b24fc2
feat(program): a join names the dims of the key columns it keeps
claude Sep 22, 2026
e73d24a
docs(language): the last relation prose says join and group-by
claude Sep 22, 2026
2223c7c
Merge origin/main into the join and group-by rename
claude Sep 22, 2026
53ba4a2
feat(program): a sum through a relation is a sum over a join (#607)
FBumann Sep 22, 2026
5cbb7ac
chore: a rule the package wrote twice is written once
claude Sep 22, 2026
638a2bf
chore: a where string's bare name and quoted label are the expression…
claude Sep 22, 2026
3d7584d
fix(language): a macro template nothing calls is held to every rule a…
claude Sep 22, 2026
a57176d
chore(typeset): a legend section is titled the way an equation sectio…
claude Sep 22, 2026
b279760
refactor(program): a comparison of expressions is one node before and…
claude Sep 22, 2026
968f1f9
refactor(language): each named expression is resolved once, and every…
claude Sep 22, 2026
a70ba12
chore(resolution): one function reads an expression string into its t…
claude Sep 23, 2026
c509d8d
refactor(language): a piecewise block is checked once at load, and th…
claude Sep 23, 2026
4145de7
chore: a shape the package already has is not declared a second time
claude Sep 23, 2026
ec9089b
refactor(language): an expression resolves straight into the program'…
claude Sep 23, 2026
fb58db9
chore(schema): the published schema carries the macro block's current…
claude Sep 23, 2026
3270e48
refactor(program): an assumption is the class named Assumption
claude Sep 23, 2026
1e4282d
feat(program): sum(by=) and at() are one join node, told apart by whe…
claude Sep 23, 2026
0455e93
feat(program): a sum through a relation is a sum over the axes its jo…
claude Sep 23, 2026
5f85607
Merge origin/main into the join branch
claude Sep 23, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
196 changes: 196 additions & 0 deletions docs/about/relations-as-linear-maps.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,196 @@
<!--
SPDX-FileCopyrightText: math-spec contributors
SPDX-License-Identifier: CC-BY-4.0
-->

# Relations as linear maps

This page says what a [relation](../reference/language/relations.md) is in the
language of linear algebra. It then shows that the join and group-by on that
page are one computation with the sum a paper prints. Read it if "joined on"
and "grouped by" read as database words and you want the math they stand for.

```yaml
dimensions:
snapshot: { dtype: int }
generator: { dtype: str }
zone: { dtype: str }
relations:
gen_zone: { key: [generator, snapshot], values: zone }
parameters:
zone_cap: { dims: [snapshot, zone] }
variables:
p: { dims: [snapshot, generator] }
constraints:
zonal:
dims: [snapshot, zone]
expression: sum(p, by=gen_zone, over=generator, into=zone) <= zone_cap
looked_up:
dims: [snapshot, generator]
where: gen_zone
expression: p <= at(zone_cap, by=gen_zone, over=zone, into=generator)
objective:
sense: minimize
expression: sum(p)
```

## A relation is an indicator

A relation with columns over the dimensions $`D_1, \dots, D_n`$ is a set of
rows, so it is a subset $`R \subseteq D_1 \times \dots \times D_n`$. Its
**indicator** $`\mathbf{1}_R`$ is $`1`$ at a row of the table and $`0`$
everywhere else. `key:` is a claim about the shape of that set: one row per key
tuple. So a relation with `key: K` and `values: V` is the graph of a function
$`f: K \to V`$, and

```math
\mathbf{1}_R(k, v) = [\, f(k) = v \,].
```

The function is partial where a key tuple has no row. A bare relation is a
subset and nothing more. Above, `gen_zone` is the graph of
$`f: \mathcal{G} \times \mathcal{T} \to \mathcal{Z}`$. The
[data contract](../reference/language/relations.md#the-data-contract) makes it
one: the loader checks one row per key tuple when the data binds.

## A join and group-by is a contraction

`sum(p, by=gen_zone, over=generator, into=zone)` is the product of two arrays,
summed over the one index they share and the call names:

```math
y_{t,z} = \sum_{g} \mathbf{1}_R(g, t, z) \cdot p_{t,g} = \sum_{g \,:\, f(g,\,t) = z} p_{t,g}
```

The right-hand form is what the typesetter
[prints](../reference/notation.md). The left-hand form is a tensor contraction,
and the two questions the relations page asks of a column are the two
positions an index can take in it:

| The relations page says | In the formula |
| ------------------------ | ----------------------------------------------------------- |
| joined on, `over=` | $`g`$ is on both factors and summed. It leaves. |
| grouped by, `into=` | $`z`$ is on the indicator alone and not summed. It arrives. |
| joined on and grouped by | $`t`$ is on both factors and not summed. It stays. |
| neither | a value column not in the formula. It is not read. |

**The result carries the free indices.** Those are the indices of the operand
and the indicator together, less the summed one, which is
`(dims(x) − joined) ∪ grouped`. That is the rule in the [expressions
reference](../reference/language/expressions.md#how-dimensions-combine), and
`_join_dims` in `src/math_spec/dimensions.py` computes it. The three refusals
beside it are the three things the formula needs:

- **The operand carries every column joined on**, or there is nothing to match.
- **The operand carries every unnamed key column.** Otherwise $`t`$ would sit
on the indicator alone, which is the grouped position, and the call names a
grouped column with `into=`.
- **The operand carries no column grouped by.** Otherwise $`z`$ would sit on
both factors and not be summed. That matches the two occurrences instead of
adding an axis. The language makes you write that match outside the
operator: `load * sum(p, by=gen_bus, over=generator, into=bus)`.

So "joined on" and "grouped by" are the two positions of the summation
convention, applied to one product. Nothing on the relations page is a
separate rule.

**The unnamed key column makes the matrix block-diagonal.** At each $`t`$,
$`M_t[z, g] = \mathbf{1}_R(g, t, z)`$ is a $`|\mathcal{Z}| \times |\mathcal{G}|`$
matrix of zeros and ones, and $`y_t = M_t\, p_t`$. A relation keyed by one
column has one block.

## A join alone is the transpose

`at(zone_cap, by=gen_zone, over=zone, into=generator)` joins the same table
with the other index bound:

```math
w_{t,g} = \sum_{z} \mathbf{1}_R(g, t, z) \cdot \mathrm{zone\_cap}_{t,z} = \mathrm{zone\_cap}_{t,\, f(g,\,t)}
```

Same indicator, same contraction, so the same free-index rule gives the result's
dimensions. As a matrix it is $`M_t^{\mathsf{T}}`$. Because $`R`$ is the graph
of $`f`$, the sum over $`z`$ has exactly one term where $`(g, t)`$ is in the
domain of $`f`$, and none elsewhere. So the contraction is the composition
$`\mathrm{zone\_cap} \circ f`$, which is a pullback.

**That one term is the whole difference between `at` and `sum`**, and
resolution decides it from the key alone. Where the columns a call groups by
hold the whole key, every group is one row, and the group-by adds nothing.
`_join` in `src/math_spec/resolution.py` names this `one_row_per_group`. A
`sum` with one row per group adds up nothing and is refused toward `at`. An
`at` with several rows per group would have several terms and is refused toward
`sum`. The
[relations page](../reference/language/relations.md#joins-and-group-bys)
quotes the message.

**The two are adjoint.** For `gen_bus: { key: generator, values: bus }`,
$`x`$ over generators and $`y`$ over buses,

```math
\langle M x, y \rangle = \sum_{b} y_b \sum_{g} \mathbf{1}_R(g, b)\, x_g = \sum_{g} x_g \sum_{b} \mathbf{1}_R(g, b)\, y_b = \langle x, M^{\mathsf{T}} y \rangle,
```

which is why the program lowers `at` to a `Join` node and `sum(by=)` to the
same `Join` under a `Sum`. The join keeps the column it sums away as an axis
named for the relation's column, `gen_zone.generator`, and the `Sum` stands
over that axis. So a map into its own dimension, which drops and adds one
dimension, still has two axes between the join and the sum. A bare relation has
the same matrix without the functional claim. A column of $`M`$ may hold several
ones, so the sum fans out and no group is one row. That is why `at` through a bare relation is refused.

## The join and the aggregate

The language fixes the formula, and an engine decides how to evaluate it. A
table stores $`\mathbf{1}_R`$ as its support: the rows where it is $`1`$.
Multiplying a matrix stored that way by a vector takes two steps.

1. **Pair each row of the table with each row of the operand that agrees on
every shared index**, here $`g`$ and $`t`$. That is an inner equi-join on
the columns joined on. Each pair is one nonzero product
$`1 \cdot p_{t,g}`$, relabelled by the column grouped by, $`z`$.
2. **Add the pairs that agree on the free indices**, here $`(t, z)`$. That is a
group-by on the result's dimensions with a sum.

lpspec, the reference engine, runs exactly these two steps. `join_relation` in
`src/lpspec/relational/engines/polars/relations.py` is step 1. It runs one inner
join on the dimensions joined on. A select then drops the dimensions not
grouped by and renames the landing column to the dimension grouped by. Step 2
is the terminal aggregate in `assembly.py`, a `group_by` over the row's
coordinates with `sum`, run once per constraint after every term has landed.
Until then a sum of linear terms is a list of terms, and adding is
concatenation. `at` runs step 1 against the same table and needs no step 2:
one row per group means no two pairs land on one coordinate. Where several
$`(g, t)`$ share one $`z`$ the join fans out, which is a column of
$`M^{\mathsf{T}}`$ holding several ones.

So the two pictures are one. The join is the multiplication by an entry of
$`\mathbf{1}_R`$, which is $`1`$ or absent. The group-by is the $`\sum`$.

## Where the built model departs from the map

Two positions of the formula have no variable to build, and there the model a
consumer builds is not the matrix.

- **An empty group is the empty sum.** A zone no generator maps to at $`t`$ has
$`y_{t,z} = 0`$, and the row reads $`0 \le \mathrm{zone\_cap}_{t,z}`$. It
names no variable, so an engine [does not build
it](../reference/language/absence.md#rows-with-no-variable-terms) and reports
the omission. With `>=` the omitted row would have been infeasible.
- **Off the domain there is no value.** $`\mathrm{zone\_cap} \circ f`$ is
undefined where $`f`$ is, so `at` is absent there and [absence
spreads](../reference/language/absence.md#how-absence-travels) to the row.
`where: gen_zone` on `looked_up` writes the domain of $`f`$ on the page, so a
reader sees which rows exist without opening the data.

## Partitions and tests

The other two uses of a relation do not contract against $`\mathbf{1}_R`$.

- **A partition steps inside a fibre.** `shift(x, along=snapshot, offset=1,
by=season_of, within=season)` reads the neighbour $`t'`$ of $`t`$ with
$`f(t') = f(t)`$. The fibres of $`f`$ partition the axis, and the frame does
not change.
- **A test is the indicator itself.** A relation's name in a `where` evaluates
$`\mathbf{1}_R`$ at the frame's own coordinate, and keeps the coordinate where
it is $`1`$.
4 changes: 2 additions & 2 deletions docs/about/what-counts-as-language.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,8 +29,8 @@ Four rules follow from the test:
in the renderer.
- The set of operators is fixed. A tool cannot add a `roll` that the others
do not know.
- Each operator has one rule for the dimensions it produces. `sum(p, by=gen_bus, over=generator, into=bus)`
lands on `bus` for every program.
- Each operator has one rule for the dimensions of its result. `sum(p, by=gen_bus, over=generator, into=bus)`
groups by `bus` for every program.
- Degree is decided when the file loads. Whether `x * y` is allowed does not
depend on which engine builds the model.

Expand Down
14 changes: 7 additions & 7 deletions docs/contributing.md
Original file line number Diff line number Diff line change
Expand Up @@ -107,13 +107,13 @@ suffix says which layer:
A node names the operation, not the verb a file writes. One verb can resolve
to two nodes, so the file's spelling cannot decide the name.

| File verb | Node | What the node names |
| ------------------ | ----------- | ------------------------------ |
| `sum(over=)` | `Sum` | dims removed from the result |
| `sum(by=)` | `GroupSum` | a sum through a relation |
| `at(by=)` | `Pullback` | a read through a relation |
| `shift(along=)` | `Translate` | a re-index along one dimension |
| `sum_back(along=)` | `WindowSum` | a sum over a trailing window |
| File verb | Node | What the node names |
| ------------------ | ------------------- | ---------------------------------------------------- |
| `sum(over=)` | `Sum` | dims removed from the result |
| `sum(by=)` | `Sum` over a `Join` | a join, and the sum over the axes it opens |
| `at(by=)` | `Join` | a join whose groups are one row, with no sum over it |
| `shift(along=)` | `Translate` | a re-index along one dimension |
| `sum_back(along=)` | `WindowSum` | a sum over a trailing window |

Nothing is abbreviated.

Expand Down
12 changes: 6 additions & 6 deletions docs/examples/operators.md
Original file line number Diff line number Diff line change
Expand Up @@ -75,7 +75,7 @@ $`\sum_{g \in \mathcal{G}} p_{t,g} \le \mathrm{limit}_{t} \qquad \forall\, t \in

```yaml
description: >-
The membership reduction — `sum(array, by=relation, over=a, into=b)` lands the result on the
The membership reduction — `sum(array, by=relation, over=a, into=b)` groups the result by the
column the relation is read to, which is what makes topology data rather than
structure.

Expand Down Expand Up @@ -112,8 +112,8 @@ $`\sum_{g \in \mathcal{G} \,:\, \mathrm{gen\_bus}(g) = b} p_{t,g} \le \mathrm{li
```yaml
description: >-
A call that names its ends — `sum(array, by=relation, over=a, into=b)`
consumes column `a` and lands on column `b`, and the other key column is
joined on, so each zone's total is taken per period.
joins on column `a` and groups by column `b`, and the other key column is
joined on and kept, so each zone's total is taken per period.

dimensions:
generator: { dtype: str }
Expand Down Expand Up @@ -148,8 +148,8 @@ $`\sum_{g \in \mathcal{G} \,:\, \mathrm{zone\_of}(g,\ e) = z} p_{g,e} \ge \mathr
```yaml
description: >-
A call with several columns at each end — `sum(array, by=relation, over=[a, …], into=[b, …])`
consumes both key columns at once and lands on the product of both value
columns in one join.
joins on both key columns at once and groups by both value columns in one
join.

dimensions:
generator: { dtype: str }
Expand Down Expand Up @@ -184,7 +184,7 @@ $`\sum_{g \in \mathcal{G},\ e \in \mathcal{E} \,:\, \mathrm{slot\_of.bus}(g,\ e)

```yaml
description: >-
The adjoint of the membership reduction — `at(array, by=relation, over=a, into=b)` reads one
The same join with no group-by — `at(array, by=relation, over=a, into=b)` reads one
coarse value once per fine label pointing at it.

dimensions:
Expand Down
2 changes: 1 addition & 1 deletion docs/howto/declare-a-column.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,7 @@ what the math does with the column.

| The column… | is declared as | because |
| ------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------- | ------------------------------------------------------------------------------------------------------------------------- |
| is an axis: something is indexed by it, or an aggregation lands terms on it | a `dimension` | its members are the coordinate set every table over it is reindexed onto |
| is an axis: something is indexed by it, or a grouping groups terms by it | a `dimension` | its members are the coordinate set every table over it is reindexed onto |
| has one value per member of a dimension, or per tuple of several — a generator's bus, a line's two ends, a generator's zone by period | a `relation` with that `key` | it is a map every operator reads, and its values are checked against the dimensions they name |
| relates members of two dimensions many-to-many, with nothing to weigh — which buses a generator may connect to | a bare `relation`, with no `values:` | `sum` reads it with both ends named, and a bare `where` tests it. Nothing reads it, because there is no one value to read |
| relates members of two dimensions many-to-many, with a weight per pair — a link's efficiency to each bus, a cycle's lines | a `parameter` over both | the weight is the data, its row set is the relation, and the aggregation is `sum(w * x, over=a)` |
Expand Down
24 changes: 12 additions & 12 deletions docs/reference/language/expressions.md
Original file line number Diff line number Diff line change
Expand Up @@ -85,18 +85,18 @@ after a variable. The objective has no name at all.

The dimension set of every expression is known before any data binds:

| Node | Dim set | Error |
| -------------------------------- | --------------------------------- | -------------------------------------------------------------------------------- |
| number | `{}` | |
| parameter / variable | its `dims` | |
| `-x`, `+x` | `dims(x)` | |
| `a + b`, `a * b`, `a / b` | `dims(a) ∪ dims(b)` | |
| `sum(x)` | `{}` | error if `dims(x)` is already empty |
| `sum(x, over=d)` | `dims(x) − {d}` | error if `d ∉ dims(x)` |
| `sum(x, by=l, over=a, into=b)` | `(dims(x) − consumed) ∪ produced` | the refusals under [how a relation is used](relations.md#how-a-relation-is-used) |
| `at(x, by=l, over=a, into=b)` | `(dims(x) − consumed) ∪ produced` | the same |
| `shift(x, along=d, offset=n)` | `dims(x)` | error if `d ∉ dims(x)` |
| `sum_back(x, along=d, window=n)` | `dims(x)` | error if `d ∉ dims(x)` |
| Node | Dim set | Error |
| -------------------------------- | ------------------------------ | -------------------------------------------------------------------------------- |
| number | `{}` | |
| parameter / variable | its `dims` | |
| `-x`, `+x` | `dims(x)` | |
| `a + b`, `a * b`, `a / b` | `dims(a) ∪ dims(b)` | |
| `sum(x)` | `{}` | error if `dims(x)` is already empty |
| `sum(x, over=d)` | `dims(x) − {d}` | error if `d ∉ dims(x)` |
| `sum(x, by=l, over=a, into=b)` | `(dims(x) − joined) ∪ grouped` | the refusals under [how a relation is used](relations.md#how-a-relation-is-used) |
| `at(x, by=l, over=a, into=b)` | `(dims(x) − joined) ∪ grouped` | the same |
| `shift(x, along=d, offset=n)` | `dims(x)` | error if `d ∉ dims(x)` |
| `sum_back(x, along=d, window=n)` | `dims(x)` | error if `d ∉ dims(x)` |

A binary operator takes the **union** of the two dimension sets, so an outer
product is allowed. The declaration's own dimensions are its **frame**, and a
Expand Down
Loading
Loading