Skip to content

chore: #275's three lookup proposals, side by side on five models - #453

Closed
FBumann wants to merge 13 commits into
mainfrom
claude/mathspec-proposals-comparison-ebn7dv
Closed

FBumann wants to merge 13 commits into
mainfrom
claude/mathspec-proposals-comparison-ebn7dv

Conversation

@FBumann

@FBumann FBumann commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

Prompt: "We need a document or PR that showcases the 3 discussed proposals from #275. How can we create a document that helps the maintainers of mathspec decide which of the proposals to take? I want use 2-4 math problems/formulations and showcase how each would look in either proposal." — "A published html artifact, but as part of the repo. In a dedicated branch" — "Lets find easy to understand use cases for the different use cases of relations and explain it like im 5 years old" — "Can we have it as a regular html file and link it in the pr description instead?"

Note

The following content was generated by AI.

This is a review surface, not a proposal to merge. It is a draft so the models and the prose can take line comments. Two of the three proposals close as not planned once #275 is decided, and this directory closes with them.

Read the page

comparison/index.html is a standalone HTML document. It needs no server and no build step.

git fetch origin claude/mathspec-proposals-comparison-ebn7dv
git switch claude/mathspec-proposals-comparison-ebn7dv
open comparison/index.html      # xdg-open on Linux, start on Windows

GitHub shows the file as source rather than rendering it. To read it without a checkout, a raw-HTML viewer renders it from the branch: raw.githack.com. The page fetches MathJax and two typefaces from public CDNs; with no network the equations stay as TeX source and everything else reads as it should.

What this changes

Adds comparison/, which writes one page comparing #428, #433 and #437 on five models. Each model is written three times, once per proposal, and every file is loaded on the branch that proposes it. Four of the five load under all three proposals. The fifth loads under one.

Nothing on the page is typed by hand except the prose. comparison/verify.py makes a worktree per branch, loads every model in a child interpreter under that branch's src/, and writes evidence.json. comparison/build.py writes index.html from it. So every YAML block, every list of dimensions, every equation and every refusal message is the branch's own output.

Two findings came out of building it, and both are on the page verbatim:

  1. A dtype: bool flag cannot be multiplied. The stand-in for an unweighted relation is a table of ones, multiplied into the sum, so the ones have to be declared dtype: int. The message says A flag masks rather than scales.
  2. The rewrite feat(language): a lookup may map into its own dimension, so a representative snapshot is sayable #436's description prescribes does not load. It sends an undirected neighbour relation to "a parameter over [bus, bus]". Measured on feat(language): a lookup may map into its own dimension, so a representative snapshot is sayable #436's own head: Parameter 'adjacent' names dimension 'region' twice. A frame is a product of distinct dimensions. That is the fifth model, and the one place a column is empty rather than clumsy.
The five models, and what each proposal does with them
Model per: #428 keys and a dot #433 relations #437
A zone that changes by period one table one table one table
A cap across a stay in a zone two declarations one table one table
Nodal balance over lines two tables two tables one table
A cap per bus and technology two tables two tables one table
Neighbouring regions refused refused one table

Outside the five, a masked sum — the produced dimension already carried — is refused by #428 and #433 and accepted by #437.

The page also carries a primer of seven cards, one per kind of pairing, in everyday terms first and then in a model. Each says which proposals spell it and how.

Files
Verified

pixi is not installable in this environment, so the gates ran from a uv venv on Python 3.13 with the repository installed, and from the pinned tool versions where they were reachable.

  • All 15 models reloaded on their branches after every change: 13 load, 2 refuse, and the refusals are the two the page prints.
  • ruff check and ruff format --check clean on comparison/ (ruff 0.14.0).
  • prettier --check clean on comparison/**/*.md, with the repository's own .prettierrc.yaml.
  • reuse lint passes on the whole tree: 240 / 240 files with copyright and licence information.
  • No trailing whitespace and a final newline in every file the branch adds, which is what trailing-whitespace-fixer and end-of-file-fixer check.
  • The page opened from the file system in a headless Chromium: standards mode, no script errors, and all 15 comparison panels laid out.
  • The prose was measured with the sentence script in .claude/skills/docs-writing, adapted to read the built page's prose blocks: n 193, avg 11.0, median 10, one sentence over 25 words.

Not run: pixi run ci as a whole, and inside it pytest, pyrefly, typos, taplo, zizmor, docs-build and compile-tex. pyrefly reads project-includes = ["src/math_spec"], so it does not reach this directory; docs-build does not either, since the page is not under docs/ and has no nav entry. One thing could not be checked here: this sandbox's proxy blocks cdnjs, so MathJax was never fetched and the equations were only seen as TeX source.

Why

#275 has three open pull requests and exactly one will be merged. Each PR body argues its own case, and the capability table in #437 is the only place they are compared — by the author of one of them. This directory compares them on models instead, with the loader on each branch as the witness, so a claim about what a proposal can say is a file that loads or a message that says why not.

The commit type departs from the table in AGENTS.md: the diff touches pyproject.toml, which the table sends to build, but the change is a decision aid and the two config lines exist only so lint passes on it. Both types hide from the changelog, so this names the change honestly rather than telling a changelog reader a build story.

🤖 Generated with Claude Code

https://claude.ai/code/session_01W4kdrj2n7tESNrgYmfXkAq

Four modelling problems, written three times each — under `per:` (#428),
keys and a dot (#433), and relations (#437) — and every file loaded on the
branch that proposes it. `comparison/verify.py` makes the evidence and
`comparison/build.py` writes `comparison/index.html` from it, so the YAML on
the page, the frames under it and the refusal messages are the loader's own
output rather than hand-typed.

`pyproject.toml` gains one `per-file-ignores` entry, the same exemption
`tools/` already has: these two scripts are generators that print what they
wrote, and the page prints a set as the script capital the typesetter uses.

Nothing here is proposed for `main`. Two of the three proposals close as not
planned once the choice is made, and this comparison closes with them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W4kdrj2n7tESNrgYmfXkAq
…e only #437 can say

Adds a primer above the models: seven cards, each one kind of pairing in
everyday terms first and then in a model, with the spelling and what each
proposal does with it. Every `says` entry is measured — the file is under
`models/` or `probes/` and the branch's own answer is in `evidence.json`.

Adds a fifth model, `models/p5`: which regions are neighbours. Both columns
are regions, so the fallback for a relation — a parameter over the pair — is
refused, and the two left-hand panels print that refusal instead of a frame.
`build.py` renders a refused model as evidence rather than treating it as a
failure.

Two measured findings the primer prints verbatim:

- a `dtype: bool` flag cannot be multiplied, so the table of ones that stands
  in for an unweighted relation has to be declared `dtype: int`;
- #436's description sends an undirected neighbour relation to "a parameter
  over `[bus, bus]`", and that file does not load, on #436's own branch or any
  other.

`verify.py` gains #436's branch, which is not a fourth proposal but the draft
stacked on #433: a claim about the self-map is a claim about #433 with #436.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W4kdrj2n7tESNrgYmfXkAq
Applies `.claude/skills/docs-writing` to the page's prose, which was written
before the skill was read.

- Every heading names a subject. `Three ways to say a lookup that a single
  arrow cannot` is `Three lookup proposals`; `Nodal balance, where a line has
  two ends` is `Nodal balance over lines`; four more lose the second clause
  after a dash.
- The page opens with its purpose in one sentence, and says who needs it,
  rather than opening on #422's history.
- One word per concept. `lookup` is the construct throughout, where the page
  said `map`, `relation` and `table` for the same thing. `map` survives only
  in `self-map` and in the legend's own reading.
- `lookup`, `relation` and `frame` are glossed at first use, each in one
  clause with a concrete instance.
- Fragments become sentences with a subject and a finite verb, in the kind
  cards and the footer. Six sentences that joined two independent clauses with
  a dash or a colon are split.
- Bold lead-ins are claims, so the three questions under the primer read as
  the three answers.

Measured with the skill's sentence script, adapted to read the built page's
prose blocks: n 193, avg 11.0, median 10, one sentence over 25 words.

Not run: `pixi run docs-build` and `pixi run lint`, which do not reach this
directory — the page is not in `docs/` and has no nav entry. `ruff check` and
`ruff format --check` pass on `comparison/`, and all 15 models were reloaded
on their branches after the rewrite.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W4kdrj2n7tESNrgYmfXkAq
Three things the pre-commit jobs would have rewritten on a pull request, where
`fail_on_changes: always` turns a rewrite into a failure:

- `verify.py` writes `evidence.json`, and JSON carries no comment, so the file
  needs a `REUSE.toml` entry as the schema and the lockfile do.
- `build.py` emitted trailing whitespace, which `trailing-whitespace-fixer`
  strips from every text file. The generator now writes the page the way the
  hook wants it, rather than leaving the hook to rewrite generated output.
- `prettier` reformatted the README's table.

The README and the page footer also stop saying the branch is not proposed for
`main`, and say what its pull request is instead: a review surface, not a
proposal to merge.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W4kdrj2n7tESNrgYmfXkAq
`comparison/index.html` was written as a fragment for a host that wrapped it,
so it had no doctype, `<html>`, `<head>` or `<body>`, and it borrowed the
host's charset, viewport and reset. It now carries all of them, and a browser
renders it in standards mode from the file system with no server and no build
step.

The reset the host used to supply is now in the page: `color-scheme`, an image
cap and the `[hidden]` rule. The README gains a section on opening the file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W4kdrj2n7tESNrgYmfXkAq
#437's head moved from 75840fd to 684fff5, which renamed the partition group
keyword to `within=` and `sum_back`'s length to `window=`. `verify.py` reloaded
all 15 models on the new head: 13 load, 2 refuse, and every frame is what it
was. No model here walks a partition, so nothing the page shows changed.

Two quoted numbers were stale and now are not: #437's call syntax beyond `by=`
is `from=, into=, within=`, and its diff against main is +2108 −802.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W4kdrj2n7tESNrgYmfXkAq
#437's head moved twice since the page last read it: `consume=`/`produce=`
replaced `over=` on `sum` and `from=`/`into=` on `sum` and `at` (cf7ca1b), and
the declaration's `over:` became `columns:` (1baa338).

The five relations models and three probes are rewritten in that language and
reloaded on the new head: 13 load, 2 refuse, and every frame is what it was.
The kind cards, two verdicts and the cost rows follow, and the diff against
main is +2325 −1000.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W4kdrj2n7tESNrgYmfXkAq
The coloured marks are gone. A filled, hollow or struck dot asked a reader to
decode a legend, and it collapsed two proposals into one cell wherever they
agreed. Each proposal now has a row of its own — its colour, its name, its
answer in full — in all three places the page compared them:

- the five problems, where three side-by-side columns become three full-width
  rows, so a file is read at full width instead of in a third of it;
- the capability comparison, where a grid of marks becomes a block per
  capability with three rows under it, each spelled out even where two say the
  same thing word for word;
- the kind cards in the primer, whose three marked lines become three rows.

The masked sum follows: it was a panel for #437 beside one refusal shared by
the other two, and is now three rows like everything else.

One component does all of it, so the page has one way of showing a proposal
rather than three. The view switch goes with the columns — every proposal is on
the page at once now — and with it the only script the page had beyond MathJax.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W4kdrj2n7tESNrgYmfXkAq
The kinds of pairing sat in a two-column grid, where a card with one code
block left a hole between the file and the rows, and the rows' own container
showed its rule colour as a grey bar above them. Each kind is now one
full-width row: the plain words beside the file, and the three proposals
underneath across the whole width. At phone width every proposal row stacks
its name above its answer instead of holding a 13rem column.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W4kdrj2n7tESNrgYmfXkAq
The primer showed one spelling per kind of pairing and left the others to a
phrase, so "two lookups" was a claim the reader could not see. Each proposal's
row now holds the file that proposal writes, in full, including the two rows
that write the same thing. The block key is part of the quote, so a pairing
written as a lookup and a pairing written as a parameter are told apart on
sight. Every line quoted is a line of a file that was loaded, and the last kind
moved from a hand-written parameter to the cap that models/p4 declares.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W4kdrj2n7tESNrgYmfXkAq
A raw-HTML viewer caches a branch URL and can serve a page older than the
branch head, and nothing on the page told the reader which copy they had. The
header and the footer now carry the time build.py wrote the page, and the
README says how to force a fresh copy.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W4kdrj2n7tESNrgYmfXkAq
The two-slots card argued from a seesaw, whose seats are symmetric, and the
card for a number on a pairing had drifted to a cap over bus and technology,
which is a dense table rather than a pairing. The first is now a journey, which
has ends that differ, and each row carries the trade the shapes make: two
lookups let a line leave an end open, and one table with roles asserts both
ends, since a row carries every column and a null is refused. The second is an
efficiency over generator and bus again, now measured — probes/efficiency_pair
loads on every branch, so the quoted line is a line of a file that was loaded.

The evidence is regenerated, so the relations head it reports is e7a95bc.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W4kdrj2n7tESNrgYmfXkAq
The trade-off card still said a call names its ends with from= and into=, and
counted two keywords where there are three. The rel_no_from probe still
declared its lookup with over:, so #437 refused it at the declaration and the
page printed a message that did not match the heading above it — the probe now
uses columns:, and the refusal is the one the heading promises: the call has to
say which consume= names. A keyless relation's columns: line is marked as
carrying the proposal, which is the one declaration line that was not. The
primer's intro calls its entries rows, which is what they became.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W4kdrj2n7tESNrgYmfXkAq

FBumann commented Sep 15, 2026

Copy link
Copy Markdown
Contributor Author

Prompt: "#437 was merged / Close all related issues and stale PRs"

Note

The following content was generated by AI.

Closed having done its job. This carried #275's three proposals side by side on five loadable models so the choice could be made on what each could say rather than on how each read. #437 won and merged as 3284926, released in v0.0.0-alpha.89; #428 and #433 are closed.

It was never for main — a decision aid, like #454's comparison-gams/, and the decision is made. The branch goes with it.

Worth knowing if anyone reaches for it again: the models here are written in the pre-#477 spelling, so they use lookups:, consume= and produce= where main now takes relations:, over= and into=, and along= on shift and sum_back. They no longer load.


Generated by Claude Code

@FBumann FBumann closed this Sep 15, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area: relations relations and dimensions: the relation design

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants