Skip to content

docs(benchmarks): the benchmark charts read a published JSON table of measurements and label our engine specsolve - #1851

Merged
FBumann merged 5 commits into
mainfrom
claude/exciting-ritchie-ap020w
Oct 6, 2026
Merged

FBumann merged 5 commits into
mainfrom
claude/exciting-ritchie-ap020w

Conversation

@FBumann

@FBumann FBumann commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

Prompt: Could we radically simplify and improve it by using tanstack charts 1.0.0 ?

Prompt: And leverage the descriptive nature fully! We should improve our data model around it probably! And use json etc — I would like to have a clean, simple data model which is used by tanstack charts, without too much data manipulation. — we shoud name it specsolve!

Note

The following content was generated by AI.

bench.plot writes benchmarks-scaling.json: tidy rows that TanStack Charts marks read by field name. A JSON spec on the page declares the charts. This replaces about 250 lines of hand-written SVG, and the page labels our engine specsolve. The rows hold the numbers re-measured on 0.6.0 in #1850.

The data model and the spec

One row per model, sink, ladder, rung and library. A measurement carries numbers. A refusal carries the budget label instead:

{"model": "dispatch", "sink": "gurobi", "ladder": "length", "rung": "xs", "variables": 10000, "library": "gurobipy-loop", "wall_s": 0.0235, "wall_q1_s": 0.0225, "wall_q3_s": 0.0239, "peak_gb": 0.1998}
{"model": "fleet", "sink": "gurobi", "ladder": "length", "rung": "m", "variables": 1200000, "library": "pyomo", "refused": ">30 s"}

variables is the rung's declared size (results.nominal). It equals the column count in every committed cell, so every library shares one x per rung. The page passes these rows straight to areaY (y1: wall_q1_s, y2: wall_q3_s) and lineY (y: wall_s, z/color: library). The only manipulation is filtering the rows to one facet. library is the harness arm name as it is.

The page spec is <script type="application/json" id="spec">. It holds facet: ["model", "sink"], the x field, one entry per metric with its y, band and format, the initial filter, and the series colour domain and range. The control buttons name a metric or a filter value from it. About 60 lines of JS turn the spec into TanStack definitions. The rest of the JS is the table, the legend and the controls. Formats are named in the spec because a function is not JSON.

bench.plot no longer edits the HTML. The page is purely hand-written, and the JSON is the generated file.

Loading TanStack without a bundler

dist/ is unbundled ESM with bare d3-* imports. An import map points @tanstack/charts/ at the raw dist/ files on jsdelivr, pinned to 1.0.0, so all the entry points share one module graph. d3-shape and d3-scale map to their pinned +esm builds, and d3-scale provides scaleLog. The six entry points the page imports pull 47 files, about 330 KB before compression, and 48 CDN requests in total. MathJax already loads from unpkg, so a runtime CDN has precedent here. Nothing joins the pixi or Python dependency sets.

What was verified, and what was not
  • Parity: I compared every cell of the rows bench.plot writes from the results committed in docs: the published benchmarks are re-measured on specsolve 0.6.0 #1850 with the const DATA that docs: the published benchmarks are re-measured on specsolve 0.6.0 #1850 wrote into the old page on main (79258b1). All 112 cells are identical, measurements and refusals both, with polars read as specsolve.
  • Browser: headless Chromium against docs/about/ and against the strict-built site/about/, in light and dark. 5 panels, the tooltip, the metric switch, the width ladder and hiding a library all work, with no console errors.
  • Not checked: I could not load jsdelivr itself, because the session proxy blocks it. The test served the same files from the npm tarball: the raw dist/ unchanged, and an esbuild bundle in place of each d3 +esm URL. The first look at the deployed page is the check for that.
  • Gates: on the branch rebased onto main, pixi run lint, format-check and typecheck are clean, the bench harness tests pass (266), and docs-build --strict passes and ships the JSON beside the page. CI on the previous head was green.
Mutation table for plot.rows

tools.mutate took the line deletions, with --tests bench/test_harness.py and CI=1 to pass the load guard: the run times nothing, and the load was my own earlier test run. Hand mutations on the refusal condition used the same three precautions. Line numbers are from before the name map was removed.

mutation result
the emit-phase filter (plot.py:72) caught
the no-peak filter (plot.py:73) caught
the ladder-rung filter (plot.py:74) caught by the new test_a_rung_neither_ladder_plots_is_left_out. It was green before that test.
the size-on-axis check dropped (line 84, by hand) caught
the already-measured check dropped (line 84, by hand) caught by the new test_a_rung_measured_past_a_ceiling_is_a_measurement_not_a_refusal. It was green before that test.
the 'error' not in r filter It survived deletion. results.py says nothing produces an error record, so it is deleted.
Coverage moved, defaults departed from, and what was not done

🤖 Generated with Claude Code

https://claude.ai/code/session_01K42KwZJtqPPhf5hDzUD91o

@read-the-docs-community

read-the-docs-community Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Documentation build overview

📚 specsolve | 🛠️ Build #34973521 | 📁 Comparing 49cf2be against latest (aa90239)

  🔍 Preview build  

3 files changed
± about/benchmarks-scaling.html
± about/benchmarks/index.html
± about/changelog/index.html

FBumann commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator Author

Note

The following content was generated by AI.

suite fails on two tests in tests/test_resolution_parity.py. The cause is polars 2.0.0, which was published to PyPI today at 11:44 UTC. This PR does not cause it.

  • The failure: test_the_empty_coordinate_builds_on_both_lanes and test_a_masked_scalar_variable_takes_its_row_with_it[masked-in] get an empty frame from dual('budget_row') for a constraint row with no dimensions.
  • Why every fresh solve gets 2.0: the requirement is polars>=1.30 and pixi.lock is gitignored.
  • Why it is not this PR's: the diff touches nothing under src/ or tests/. The same 2 tests fail on an untouched main (aa90239) with polars 2.0.0. With polars 1.44.2 in the same environment, all 30 tests in that file pass. linopy returns the dual for that row correctly (1.0), so the empty frame comes from specsolve's own read.
  • Fix: none exists yet. I did not port one into this PR, because a dependency cap is outside its scope. A stopgap is polars>=1.30,<2 in pyproject.toml (both the dependency list and the pixi table). The lasting fix is to read a dimensionless row's duals in the way polars 2.0 requires.
Reproduction
pixi run pytest tests/test_resolution_parity.py -q   # polars 2.0.0: 2 failed, 28 passed
uv pip install --target /tmp/pl1 'polars==1.44.2'
PYTHONPATH=/tmp/pl1 pixi run pytest tests/test_resolution_parity.py -q   # 30 passed

Generated by Claude Code

@codspeed

codspeed Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Merging this PR will not alter performance

✅ 67 untouched benchmarks
⏩ 307 skipped benchmarks1


Comparing claude/exciting-ritchie-ap020w (49cf2be) with main (85d864a)2

Open in CodSpeed

Footnotes

  1. 307 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

  2. No successful run was found on main (79258b1) during the generation of this report, so 85d864a was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

claude added 4 commits October 6, 2026 15:53
… measurements and label our engine specsolve

bench.plot writes docs/about/benchmarks-scaling.json, one row per model,
sink, ladder, rung and library, instead of rewriting a line of the page.
The page declares its charts as a JSON spec over those fields and draws
them with TanStack Charts 1.0.0 from jsdelivr, replacing the hand-written
SVG renderer.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K42KwZJtqPPhf5hDzUD91o
…ilter nothing reaches is gone

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K42KwZJtqPPhf5hDzUD91o
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K42KwZJtqPPhf5hDzUD91o
…n 0.6.0

bench.plot is re-run on the results #1850 committed. Those results name the
engine's arm specsolve, so the map that renamed lpspec goes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K42KwZJtqPPhf5hDzUD91o
@FBumann
FBumann added this pull request to stack #1859 October 6, 2026 20:36
@FBumann
FBumann merged commit f222a52 into main Oct 6, 2026
15 checks passed
@FBumann
FBumann deleted the claude/exciting-ritchie-ap020w branch October 6, 2026 20:50
FBumann added a commit that referenced this pull request Oct 6, 2026
…he sizes a budget refused, and compares every library with it in one table (#1855)

> **Prompt:** I would rather like you to propose a better plotting for
our benchmarks.
>
> **Prompt:** Absolute numbers are important! But maybe we can add a
toggle for log or linear scaling? … We should highlight specsolve! — A
as the lead + E is what i want. — I would like to have the axes both
scaled properly — Could we add ticked lines as extrapolations? — Use a
blue to green color for the heatmap — We could use our own brands
violet…? — Should we unselect the pure matrix ones? They are merely an
oracle and probably not comparable…? — Yes, style them neutral. And
probably also in the heatmap…?

> [!NOTE]
> The following content was generated by AI.

The benchmark page leads with absolute cost against size. specsolve is a
thick line in the site's brand violet; the other libraries are thinner,
and the two hand-written matrix arms are one neutral "matrix floor". One
switch flips both axes between log–log and linear–linear, and dashed
lines project the sizes a budget refused. A single ratio table, shaded
blue to green, replaces the per-panel tables.

<details><summary>What the page draws</summary>

- **Panels.** One panel per model and solver: each library's cost
against size, with the median line and the Q1–Q3 band. specsolve's line
is 3.5 px and the libraries' lines are 2 px. specsolve is drawn last and
named above its last measured point. Every measured point is a dot
ringed in the surface colour.
- **Matrix floor.** `gurobipy-matrix` and `highspy-matrix` are a
hand-written CSR handed to the solver's bulk API, the limit a modelling
library can approach. They are drawn as one thin grey line with no band,
behind one legend entry, and they never share a panel. In the table they
are one unshaded "matrix floor" column at the right, because being
faster than specsolve is what a floor is for. A note on the page says
what the floor is. `gurobipy-loop` stays a library.
- **Axes.** One switch for log–log or linear–linear. Mixed axes are left
out: linear time over log size draws linear growth as an upward curve.
- **Projections.** A dashed line goes from a library's last measurement
to each size its time or memory budget refused, at the log–log slope of
the last measured step, floored at 1. A hollow ring marks each projected
size. The axes scale to the measurements, so a steep projection is
clipped at the top. The tooltip shows `≈ value (projected)`.
- **Table.** specsolve's own cost, then each library's cost divided by
specsolve's at the same size: green where specsolve is faster, blue
where the other library is, grey within 10%. A cell's hover gives the
library's own number. A refused cell keeps its `>30 s` or `>16 GB` label
and gives the projection on hover.
- **Spec.** The JSON spec gains `highlight`, `floor` (its label and
members), and the `ink` and `paint` colour lists. The ratios, the
projections, the floor and the highlighting all read from it.

</details>

<details><summary>Colour</summary>

- **Palette.** specsolve brand violet (`#6d28d9`; in dark mode
`#9670e8`, because the site's `#a78bfa` is too light for a chart line),
linopy gold, pyomo blue, gurobipy-loop red (Gurobi's colour), and the
matrix floor neutral grey (`#6b6a61`; in dark mode `#a3a196`).
- **Validation.** The four library colours pass the dataviz skill's
validator in both themes, with worst adjacent CVD ΔE 26.0 in light mode
and 24.7 in dark (protan). The libraries other than specsolve are each
hue mixed into the surface, 70% in light mode and 85% in dark.
- **Brand colours.** linopy, Pyomo and HiGHS have no signature colour I
could verify.
- **Table.** The blue–green pair passes deutan (ΔE ≈ 20) but is weak for
tritanopia (ΔE ≈ 5). Every cell prints its ratio, which carries the
reading there.

</details>

<details><summary>What was verified, and what was not</summary>

- **Browser:** headless Chromium in light and dark, with the jsdelivr
files served from the npm tarball as in #1851. 5 panels and 20 table
rows. Switching to linear–linear, the width ladder and peak RSS, and
hiding pyomo, gives 1 panel, 4 rows, and each refused cell's projection
on hover. No console errors.
- **Gates:** `pixi run docs-build` (strict) passes and ships the page
and its JSON. The diff touches no Python. CI was green on the previous
head.
- **Not checked:** loading from the real jsdelivr. The Read the Docs
preview of this PR is the first real load.

</details>

<details><summary>Stack and what was not done</summary>

- **Base:** stacked on #1851, which carries the JSON data (the numbers
re-measured on 0.6.0) and the spec this page reads.
- **Not done:** no change to `bench/plot.py` or the data model. The
per-panel absolute tables are gone; the specsolve column carries
absolute values, and every other value is in the cell hovers and chart
tooltips. The panels stay three to a row.

</details>

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01K42KwZJtqPPhf5hDzUD91o

---------

Co-authored-by: Claude <noreply@anthropic.com>
FBumann added a commit that referenced this pull request Oct 6, 2026
…of the docs, in the site's theme (#1860)

> **Prompt:** could we include the charts etc into the regular docs? Or
does it need to be a standalone html page like it it right now? — Yes
merge them

> [!NOTE]
> The following content was generated by AI.

The benchmark charts move from the standalone `benchmarks-scaling.html`
into `about/benchmarks.md`. That page now also holds the method, so
there is one "Benchmarks" entry in the nav. The charts follow the site's
light/dark toggle and work with instant navigation.

<details><summary>How the charts live inside the docs</summary>

- **Import map.** `overrides/main.html` (`theme.custom_dir`) adds the
TanStack/d3 import map to the head of every page. It has to be every
page because instant navigation never re-renders the head: a reader who
reaches the benchmark page from another page keeps the first page's
head. A map costs nothing until a module asks for it.
- **Script.** `docs/javascripts/benchmarks.js` is loaded through
`extra_javascript` as `type: module`. On each `document$` emission it
destroys the previous page's charts. It then imports TanStack on demand,
only when the page has a `#benchmark-charts` block, so no other page
fetches the library. It stops if the reader has navigated away while the
library loads.
- **Page.** The controls, panels, table and JSON spec are one HTML block
in the Markdown. The prose around them is Markdown: a purpose line, the
model list, "Reading the charts", then the method sections that were
`benchmarks.md`.
- **Theme.** `docs/stylesheets/benchmarks.css` takes surfaces and text
from the theme's `--md-*` variables and switches the series and table
colours under `[data-md-color-scheme="slate"]`, so the block follows the
site's toggle. The standalone page only followed the operating-system
setting. The panels sit two to a row in the narrower content column.
- **Data.** `benchmarks-scaling.json` is renamed `benchmarks.json`,
after its page. `bench/plot.py`, the published-benchmark workflow
artifact list, `docs/README.md` and `bench/README.md` follow.
- **Links.** The README card, the landing-page card icon selector in
`extra.css`, the About index row and the two in-page links now point at
`about/benchmarks/`. `bench/report.py` drops its link to the old chart
page.

</details>

<details><summary>Prose</summary>

- The old chart page said "a quick cell here took 84 rounds and a slow
one 9", which contradicts the method page's current "every measurement
gets the same nine rounds, pinned". That sentence is dropped, and the
median section now opens with the nine-round rule.
- The explanation of the length/width switch was an HTML comment on the
old page. It is now a bullet in "Reading the charts".
- The About index row said the page measures "build and solve cost
against linopy". It measures build cost only, against four other ways of
writing the model, and now says so.
- Measured with the docs-writing script: 86 sentences, median 17 words,
18 over 25. The old method page: 54, median 17, 11 over 25.

</details>

<details><summary>What was verified, and what was not</summary>

- **Gates:** `pixi run check` passes on 53ad43d (4999 passed). `pixi
run docs-build` (strict), `pixi run docs-test` (112) and
`tests/test_docs_site.py` + `tests/test_doc_examples.py` (140) pass. The
bench harness tests pass (266).
- **Browser, on the built site:** headless Chromium, with the jsdelivr
files served from the npm tarball. Each check was run twice, once on a
direct load and once after arriving by instant navigation from the About
index.
  - **The page** draws 5 panels and 20 table rows.
- **Each toggle is asserted by its effect.** Width gives 1 panel and 4
rows. Peak RSS gives axes in GB. Linear–linear gives a zero tick. Hiding
pyomo in the legend removes its table column. The pressed state follows
each click.
- **The dots read their computed colours.** Every measured dot has a
ring in the surface colour, and every projection ring is filled with the
surface colour.
- **Leaving** for another page leaves no charts or tooltips behind, and
Back redraws them.
- **No other page** fetched from jsdelivr before the benchmark page was
opened.
  - **Both themes** follow the site scheme.
- **Console errors** came only from MathJax, Mermaid, fonts and the
GitHub API, whose hosts the test blocks. None came from the charts.
- **Two defects shipped in earlier commits of this PR,** and both came
from names carried over from the standalone page.
- **Dead toggles.** The click handler still matched the old `.seg`
class. My first browser test counted panels and rows, which stay the
same across the scale and metric toggles, so it missed this. e9a71bb
fixes it.
- **Dots without rings, and black projection rings.** The dots used
`--surface-1`, a token that the docs stylesheet does not define.
53ad43d switches them to `--surface`.
- The toggle and dot-colour checks above are the ones that would have
caught these defects.
- **Not checked:** the real jsdelivr load, as in #1851 and #1855.

</details>

<details><summary>What was not done</summary>

- **No redirect from the old URL.** `about/benchmarks-scaling.html` now
returns 404, under the no-compatibility rule.
- **A stale anchor left alone.** `bench/README.md` links a
`#the-density-sweep-and-a-claim-it-refuses` anchor that predates this PR
and no longer exists. It is unchanged here.
- **The label, tooltip and polish changes** discussed in the session are
left for a follow-up.

</details>

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01K42KwZJtqPPhf5hDzUD91o

---------

Co-authored-by: Claude <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants