Skip to content

docs(benchmarks): the benchmark page highlights specsolve, projects the sizes a budget refused, and compares every library with it in one table - #1855

Merged
FBumann merged 6 commits into
claude/exciting-ritchie-ap020wfrom
claude/benchmark-chart-design
Oct 6, 2026
Merged

FBumann merged 6 commits into
claude/exciting-ritchie-ap020wfrom
claude/benchmark-chart-design

Conversation

@FBumann

@FBumann FBumann commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

Prompt: I would rather like you to propose a better plotting for our benchmarks.

Prompt: Absolute numbers are important! But maybe we can add a toggle for log or linear scaling? … We should highlight specsolve! — A as the lead + E is what i want. — I would like to have the axes both scaled properly — Could we add ticked lines as extrapolations? — Use a blue to green color for the heatmap — We could use our own brands violet…? — Should we unselect the pure matrix ones? They are merely an oracle and probably not comparable…? — Yes, style them neutral. And probably also in the heatmap…?

Note

The following content was generated by AI.

The benchmark page leads with absolute cost against size. specsolve is a thick line in the site's brand violet; the other libraries are thinner, and the two hand-written matrix arms are one neutral "matrix floor". One switch flips both axes between log–log and linear–linear, and dashed lines project the sizes a budget refused. A single ratio table, shaded blue to green, replaces the per-panel tables.

What the page draws
  • Panels. One panel per model and solver: each library's cost against size, with the median line and the Q1–Q3 band. specsolve's line is 3.5 px and the libraries' lines are 2 px. specsolve is drawn last and named above its last measured point. Every measured point is a dot ringed in the surface colour.
  • Matrix floor. gurobipy-matrix and highspy-matrix are a hand-written CSR handed to the solver's bulk API, the limit a modelling library can approach. They are drawn as one thin grey line with no band, behind one legend entry, and they never share a panel. In the table they are one unshaded "matrix floor" column at the right, because being faster than specsolve is what a floor is for. A note on the page says what the floor is. gurobipy-loop stays a library.
  • Axes. One switch for log–log or linear–linear. Mixed axes are left out: linear time over log size draws linear growth as an upward curve.
  • Projections. A dashed line goes from a library's last measurement to each size its time or memory budget refused, at the log–log slope of the last measured step, floored at 1. A hollow ring marks each projected size. The axes scale to the measurements, so a steep projection is clipped at the top. The tooltip shows ≈ value (projected).
  • Table. specsolve's own cost, then each library's cost divided by specsolve's at the same size: green where specsolve is faster, blue where the other library is, grey within 10%. A cell's hover gives the library's own number. A refused cell keeps its >30 s or >16 GB label and gives the projection on hover.
  • Spec. The JSON spec gains highlight, floor (its label and members), and the ink and paint colour lists. The ratios, the projections, the floor and the highlighting all read from it.
Colour
  • Palette. specsolve brand violet (#6d28d9; in dark mode #9670e8, because the site's #a78bfa is too light for a chart line), linopy gold, pyomo blue, gurobipy-loop red (Gurobi's colour), and the matrix floor neutral grey (#6b6a61; in dark mode #a3a196).
  • Validation. The four library colours pass the dataviz skill's validator in both themes, with worst adjacent CVD ΔE 26.0 in light mode and 24.7 in dark (protan). The libraries other than specsolve are each hue mixed into the surface, 70% in light mode and 85% in dark.
  • Brand colours. linopy, Pyomo and HiGHS have no signature colour I could verify.
  • Table. The blue–green pair passes deutan (ΔE ≈ 20) but is weak for tritanopia (ΔE ≈ 5). Every cell prints its ratio, which carries the reading there.
What was verified, and what was not
  • Browser: headless Chromium in light and dark, with the jsdelivr files served from the npm tarball as in docs(benchmarks): the benchmark charts read a published JSON table of measurements and label our engine specsolve #1851. 5 panels and 20 table rows. Switching to linear–linear, the width ladder and peak RSS, and hiding pyomo, gives 1 panel, 4 rows, and each refused cell's projection on hover. No console errors.
  • Gates: pixi run docs-build (strict) passes and ships the page and its JSON. The diff touches no Python. CI was green on the previous head.
  • Not checked: loading from the real jsdelivr. The Read the Docs preview of this PR is the first real load.
Stack and what was not done

🤖 Generated with Claude Code

https://claude.ai/code/session_01K42KwZJtqPPhf5hDzUD91o

@read-the-docs-community

read-the-docs-community Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Documentation build overview

📚 specsolve | 🛠️ Build #34979039 | 📁 Comparing 1fb99c2 against latest (79258b1)

  🔍 Preview build  

2 files changed
± about/benchmarks-scaling.html
± about/changelog/index.html

@codspeed

codspeed Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Merging this PR will not alter performance

✅ 67 untouched benchmarks
⏩ 307 skipped benchmarks1


Comparing claude/benchmark-chart-design (1fb99c2) with claude/exciting-ritchie-ap020w (49cf2be)

Open in CodSpeed

Footnotes

  1. 307 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

@FBumann
FBumann added this pull request to stack #1859 October 6, 2026 20:36
claude added 6 commits October 6, 2026 22:37
…he sizes a budget refused, and compares every library with it in one table

Each panel draws specsolve as a thick red line over thin pastel lines for the
other libraries, with one switch between log-log and linear-linear axes. A
dashed line continues a library past the size its budget refused, at the growth
rate of its last measured step and never slower than linear. One table replaces
the per-panel tables: specsolve's own cost beside every other library's ratio
to it, shaded blue to green, with the projection on a refused cell's hover.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K42KwZJtqPPhf5hDzUD91o
…ze is a hollow ring, and specsolve's line is named

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K42KwZJtqPPhf5hDzUD91o
…so they read as colours beside specsolve's

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K42KwZJtqPPhf5hDzUD91o
…robipy's matrix line takes Gurobi's red

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K42KwZJtqPPhf5hDzUD91o
…eutral matrix floor rather than as rival libraries

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K42KwZJtqPPhf5hDzUD91o
@FBumann
FBumann force-pushed the claude/benchmark-chart-design branch from 7953b5f to 1fb99c2 Compare October 6, 2026 20:37
@FBumann
FBumann merged commit 99462fe into main Oct 6, 2026
13 checks passed
@FBumann
FBumann deleted the claude/benchmark-chart-design branch October 6, 2026 20:50
FBumann added a commit that referenced this pull request Oct 6, 2026
… measurements and label our engine specsolve (#1851)

> **Prompt:** Could we radically simplify and improve it by using
tanstack charts 1.0.0 ?
>
> **Prompt:** And leverage the descriptive nature fully! We should
improve our data model around it probably! And use json etc — I would
like to have a clean, simple data model which is used by tanstack
charts, without too much data manipulation. — we shoud name it
specsolve!

> [!NOTE]
> The following content was generated by AI.

`bench.plot` writes `benchmarks-scaling.json`: tidy rows that TanStack
Charts marks read by field name. A JSON spec on the page declares the
charts. This replaces about 250 lines of hand-written SVG, and the page
labels our engine `specsolve`. The rows hold the numbers re-measured on
0.6.0 in #1850.

<details><summary>The data model and the spec</summary>

One row per model, sink, ladder, rung and library. A measurement carries
numbers. A refusal carries the budget label instead:

```json
{"model": "dispatch", "sink": "gurobi", "ladder": "length", "rung": "xs", "variables": 10000, "library": "gurobipy-loop", "wall_s": 0.0235, "wall_q1_s": 0.0225, "wall_q3_s": 0.0239, "peak_gb": 0.1998}
{"model": "fleet", "sink": "gurobi", "ladder": "length", "rung": "m", "variables": 1200000, "library": "pyomo", "refused": ">30 s"}
```

`variables` is the rung's declared size (`results.nominal`). It equals
the column count in every committed cell, so every library shares one x
per rung. The page passes these rows straight to `areaY` (`y1:
wall_q1_s`, `y2: wall_q3_s`) and `lineY` (`y: wall_s`, `z/color:
library`). The only manipulation is filtering the rows to one facet.
`library` is the harness arm name as it is.

The page spec is `<script type="application/json" id="spec">`. It holds
`facet: ["model", "sink"]`, the `x` field, one entry per metric with its
`y`, `band` and `format`, the initial `filter`, and the series colour
domain and range. The control buttons name a metric or a filter value
from it. About 60 lines of JS turn the spec into TanStack definitions.
The rest of the JS is the table, the legend and the controls. Formats
are named in the spec because a function is not JSON.

`bench.plot` no longer edits the HTML. The page is purely hand-written,
and the JSON is the generated file.

</details>

<details><summary>Loading TanStack without a bundler</summary>

`dist/` is unbundled ESM with bare `d3-*` imports. An import map points
`@tanstack/charts/` at the raw `dist/` files on jsdelivr, pinned to
`1.0.0`, so all the entry points share one module graph. `d3-shape` and
`d3-scale` map to their pinned `+esm` builds, and `d3-scale` provides
`scaleLog`. The six entry points the page imports pull 47 files, about
330 KB before compression, and 48 CDN requests in total. MathJax already
loads from unpkg, so a runtime CDN has precedent here. Nothing joins the
pixi or Python dependency sets.

</details>

<details><summary>What was verified, and what was not</summary>

- **Parity:** I compared every cell of the rows `bench.plot` writes from
the results committed in #1850 with the `const DATA` that #1850 wrote
into the old page on `main` (79258b1). All 112 cells are identical,
measurements and refusals both, with `polars` read as `specsolve`.
- **Browser:** headless Chromium against `docs/about/` and against the
strict-built `site/about/`, in light and dark. 5 panels, the tooltip,
the metric switch, the width ladder and hiding a library all work, with
no console errors.
- **Not checked:** I could not load jsdelivr itself, because the session
proxy blocks it. The test served the same files from the npm tarball:
the raw `dist/` unchanged, and an esbuild bundle in place of each d3
`+esm` URL. The first look at the deployed page is the check for that.
- **Gates:** on the branch rebased onto `main`, `pixi run lint`,
`format-check` and `typecheck` are clean, the bench harness tests pass
(266), and `docs-build --strict` passes and ships the JSON beside the
page. CI on the previous head was green.

</details>

<details><summary>Mutation table for <code>plot.rows</code></summary>

`tools.mutate` took the line deletions, with `--tests
bench/test_harness.py` and `CI=1` to pass the load guard: the run times
nothing, and the load was my own earlier test run. Hand mutations on the
refusal condition used the same three precautions. Line numbers are from
before the name map was removed.

| mutation | result |
|---|---|
| the emit-phase filter (`plot.py:72`) | **caught** |
| the no-peak filter (`plot.py:73`) | **caught** |
| the ladder-rung filter (`plot.py:74`) | **caught** by the new
`test_a_rung_neither_ladder_plots_is_left_out`. It was green before that
test. |
| the size-on-axis check dropped (line 84, by hand) | **caught** |
| the already-measured check dropped (line 84, by hand) | **caught** by
the new
`test_a_rung_measured_past_a_ceiling_is_a_measurement_not_a_refusal`. It
was green before that test. |
| the `'error' not in r` filter | It survived deletion. `results.py`
says nothing produces an `error` record, so it is deleted. |

</details>

<details><summary>Coverage moved, defaults departed from, and what was
not done</summary>

- **Tests:** the five tests on `plot.series` and `plot.panels` now
assert the same claims on `plot.rows`, through a `_refused(records,
library)` helper. `_plotted` is gone.
- **Type:** the diff touches `bench/`, `docs/`, `.github/` and
`pyproject.toml`. By the table, the topmost row is `docs/`. The
changelog line is added.
- **Stacked on this:** #1855 redesigns the page's charts on top of this
data model.
- **Not done:** no new measurements here. I also left the footer's stale
references to `bench/run.py` and `docs/benchmarks.md` as they were. They
predate this PR.

</details>

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01K42KwZJtqPPhf5hDzUD91o

---------

Co-authored-by: Claude <noreply@anthropic.com>
FBumann added a commit that referenced this pull request Oct 6, 2026
…of the docs, in the site's theme (#1860)

> **Prompt:** could we include the charts etc into the regular docs? Or
does it need to be a standalone html page like it it right now? — Yes
merge them

> [!NOTE]
> The following content was generated by AI.

The benchmark charts move from the standalone `benchmarks-scaling.html`
into `about/benchmarks.md`. That page now also holds the method, so
there is one "Benchmarks" entry in the nav. The charts follow the site's
light/dark toggle and work with instant navigation.

<details><summary>How the charts live inside the docs</summary>

- **Import map.** `overrides/main.html` (`theme.custom_dir`) adds the
TanStack/d3 import map to the head of every page. It has to be every
page because instant navigation never re-renders the head: a reader who
reaches the benchmark page from another page keeps the first page's
head. A map costs nothing until a module asks for it.
- **Script.** `docs/javascripts/benchmarks.js` is loaded through
`extra_javascript` as `type: module`. On each `document$` emission it
destroys the previous page's charts. It then imports TanStack on demand,
only when the page has a `#benchmark-charts` block, so no other page
fetches the library. It stops if the reader has navigated away while the
library loads.
- **Page.** The controls, panels, table and JSON spec are one HTML block
in the Markdown. The prose around them is Markdown: a purpose line, the
model list, "Reading the charts", then the method sections that were
`benchmarks.md`.
- **Theme.** `docs/stylesheets/benchmarks.css` takes surfaces and text
from the theme's `--md-*` variables and switches the series and table
colours under `[data-md-color-scheme="slate"]`, so the block follows the
site's toggle. The standalone page only followed the operating-system
setting. The panels sit two to a row in the narrower content column.
- **Data.** `benchmarks-scaling.json` is renamed `benchmarks.json`,
after its page. `bench/plot.py`, the published-benchmark workflow
artifact list, `docs/README.md` and `bench/README.md` follow.
- **Links.** The README card, the landing-page card icon selector in
`extra.css`, the About index row and the two in-page links now point at
`about/benchmarks/`. `bench/report.py` drops its link to the old chart
page.

</details>

<details><summary>Prose</summary>

- The old chart page said "a quick cell here took 84 rounds and a slow
one 9", which contradicts the method page's current "every measurement
gets the same nine rounds, pinned". That sentence is dropped, and the
median section now opens with the nine-round rule.
- The explanation of the length/width switch was an HTML comment on the
old page. It is now a bullet in "Reading the charts".
- The About index row said the page measures "build and solve cost
against linopy". It measures build cost only, against four other ways of
writing the model, and now says so.
- Measured with the docs-writing script: 86 sentences, median 17 words,
18 over 25. The old method page: 54, median 17, 11 over 25.

</details>

<details><summary>What was verified, and what was not</summary>

- **Gates:** `pixi run check` passes on 53ad43d (4999 passed). `pixi
run docs-build` (strict), `pixi run docs-test` (112) and
`tests/test_docs_site.py` + `tests/test_doc_examples.py` (140) pass. The
bench harness tests pass (266).
- **Browser, on the built site:** headless Chromium, with the jsdelivr
files served from the npm tarball. Each check was run twice, once on a
direct load and once after arriving by instant navigation from the About
index.
  - **The page** draws 5 panels and 20 table rows.
- **Each toggle is asserted by its effect.** Width gives 1 panel and 4
rows. Peak RSS gives axes in GB. Linear–linear gives a zero tick. Hiding
pyomo in the legend removes its table column. The pressed state follows
each click.
- **The dots read their computed colours.** Every measured dot has a
ring in the surface colour, and every projection ring is filled with the
surface colour.
- **Leaving** for another page leaves no charts or tooltips behind, and
Back redraws them.
- **No other page** fetched from jsdelivr before the benchmark page was
opened.
  - **Both themes** follow the site scheme.
- **Console errors** came only from MathJax, Mermaid, fonts and the
GitHub API, whose hosts the test blocks. None came from the charts.
- **Two defects shipped in earlier commits of this PR,** and both came
from names carried over from the standalone page.
- **Dead toggles.** The click handler still matched the old `.seg`
class. My first browser test counted panels and rows, which stay the
same across the scale and metric toggles, so it missed this. e9a71bb
fixes it.
- **Dots without rings, and black projection rings.** The dots used
`--surface-1`, a token that the docs stylesheet does not define.
53ad43d switches them to `--surface`.
- The toggle and dot-colour checks above are the ones that would have
caught these defects.
- **Not checked:** the real jsdelivr load, as in #1851 and #1855.

</details>

<details><summary>What was not done</summary>

- **No redirect from the old URL.** `about/benchmarks-scaling.html` now
returns 404, under the no-compatibility rule.
- **A stale anchor left alone.** `bench/README.md` links a
`#the-density-sweep-and-a-claim-it-refuses` anchor that predates this PR
and no longer exists. It is unchanged here.
- **The label, tooltip and polish changes** discussed in the session are
left for a follow-up.

</details>

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01K42KwZJtqPPhf5hDzUD91o

---------

Co-authored-by: Claude <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants