Skip to content

docs: the published benchmarks are re-measured on specsolve 0.6.0 - #1850

Merged
FBumann merged 5 commits into
mainfrom
claude/gifted-darwin-b7918s
Oct 6, 2026
Merged

FBumann merged 5 commits into
mainfrom
claude/gifted-darwin-b7918s

Conversation

@FBumann

@FBumann FBumann commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

Prompt: The benchmark ran. Can you update our published benchmark with a PR?

Note

The following content was generated by AI.

The benchmark pages now show run 37298997947, measured on aa902393 (0.6.0). It covers the same five cases as the page it replaces. The page's sentences about the run are recomputed from the new files, and bench/reproduce.py is pinned to that run.

What changed
  • Results. Five cases: highs dispatch and fleet, and gurobi dispatch, transport and fleet. The other three cases were stopped by the memory watchdog, as on the previous run (the OOM record dies with the run artifact instead of landing beside the numbers it explains #1498).
  • Where the files come from. The network policy of the session blocks the artifact host, so the files come from the eight per-case artifacts, uploaded by hand. Every results file is byte-identical in each artifact that holds it. results-gurobi-fleet holds all ten files.
  • casualties.json is not committed. It was never committed, and bench/results.py skips it. Its three entries are in the page text.
  • Generated content. The tables come from pixi run -e bench report and the chart data from pixi run -e bench plot. Their output, fingerprint included, is the same as the run's own report and plot steps.
  • Hand-written sentences, recomputed. Before I used each method on the new files, I checked that it reproduces the old sentence exactly from main's results.
    • Mean against median: 14 cells are above 1.10x, against 4 before. The worst is 3.41x, against 1.21x before.
    • The cell the median flips: now transport/w1 on gurobi against gurobipy-loop, in specsolve's favour. Before, it was dispatch/s against linopy, against specsolve. With it goes the the gurobi sink alternates between a fast and a slow build, round after round #1288 sentence: specsolve's rounds on that cell are now 88 to 94 ms.
    • The three casualties: 24.6, 26.3 and 23.9 GB.
    • Removed: the sentence on why storage survived on gurobi earlier. The previous run already contradicted it.
  • bench/reproduce.py. test_the_lock_installs_what_the_published_numbers_were_taken_on failed on the new files, because the lock installed lpspec at 8d27e88b. Changes:
  • main is merged in, with fix(deps): specsolve installs a polars older than 2.0, on which a row's dual can come back empty and a model can take nine times the memory #1852's polars cap. The cap was ported here first, and the merge makes it identical to main. The only conflict was in CHANGELOG.md, which now keeps both lines.
Found, and not fixed here
  • A cold first round. In 13 cells, the first of the nine rounds of an xs cell on specsolve, linopy or pyomo is 1.9 to 22.6 times the median of the other eight. Pyomo dispatch/xs on highs takes 1524 ms, then about 67 ms. main's results from 14 September have no such round. Published medians do not move, because that round is always the largest of nine. The q3 of a band may move a little. _rounds had warmup_rounds=0 at the previous run too, so the cause is elsewhere in what changed after 8d27e88b. I did not look further.
  • The 84-rounds sentence. benchmarks-scaling.html says "a quick cell here took 84 rounds and a slow one 9". That sentence is hand-written and was already false, because rounds are pinned at 9.
  • linopy's commit. The run does not record which linopy commit it measured, so the lock pins today's master (f665a260). The guard checks only the versions the run records.
Gates
  • pixi run check, before the main merge, on polars 1.44.2: 4996 passed, 564 skipped, 1 xfailed. After the merge: tools.changelog check, tests/test_changelog.py and tests/test_tooling_pins.py pass.
  • pytest bench/test_harness.py: 262 passed, 16 skipped. This is test-bench, including test_the_report_renders_from_the_committed_results and the lock guard.
  • pixi run docs-build, the command Read the Docs runs: clean.
  • uv lock --script bench/reproduce.py --check: clean. pixi run lint and format-check: clean.
  • Not run: uv run --locked bench/reproduce.py itself.

🤖 Generated with Claude Code

https://claude.ai/code/session_018nKGXp8uKnMeJACjFYwbdY

@read-the-docs-community

read-the-docs-community Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Documentation build overview

📚 specsolve | 🛠️ Build #34972480 | 📁 Comparing 3040c5f against latest (aa90239)

  🔍 Preview build  

2 files changed
± about/benchmarks/index.html
± about/changelog/index.html

FBumann commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator Author

Note

The following content was generated by AI.

Two signals on this PR come from polars 2.0.0, not from this PR: the red suite job and CodSpeed's memory regression. The fix is #1852, which caps polars at <2. It is ported here as dce68af8, which changes pyproject.toml and uv.lock. The port adds nothing once #1852 is on main.

polars 2.0.0 was released on PyPI on 2026-10-06 at 11:51 UTC. pyproject.toml asked only for polars>=1.30, and there is no lock, so CI and CodSpeed both installed 2.0.0.

  • suite. Two tests fail:

    • tests/test_resolution_parity.py::test_the_empty_coordinate_builds_on_both_lanes
    • tests/test_resolution_parity.py::test_a_masked_scalar_variable_takes_its_row_with_it[masked-in]

    In both, dual('budget_row') returns 0 rows where the tests expect 1.

  • CodSpeed. The storage/s specsolve cells go from about 17 MB to about 776 MB. On the same tree, test_emit[storage-s-specsolve-highs] with --benchmark-memory measures:

polars peak_bytes RSS
1.44.2 119–123 MB 234 MB
2.0.0 1091–1146 MB 1072 MB
How it was checked
  • Tree 2de8fbc3, with linopy master at f665a260, which is what CI installs.
  • The test run and the memory run were each taken twice, with only polars swapped between them.
  • The memory figures are three rounds of one local run. They are not comparable in absolute terms with CodSpeed's instrument, only in direction and size.
  • After the port, pixi run check passes on polars 1.44.2: 4996 passed, 564 skipped, 1 xfailed.

Generated by Claude Code

@codspeed

codspeed Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Merging this PR will not alter performance

✅ 67 untouched benchmarks
⏩ 307 skipped benchmarks1


Comparing claude/gifted-darwin-b7918s (3040c5f) with main (331a839)2

Open in CodSpeed

Footnotes

  1. 307 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

  2. No successful run was found on main (85d864a) during the generation of this report, so 331a839 was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

claude added 2 commits October 6, 2026 15:14
…'s dual can come back empty and a model can take nine times the memory

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018nKGXp8uKnMeJACjFYwbdY
@FBumann FBumann mentioned this pull request Oct 6, 2026
@FBumann
FBumann merged commit e13ece8 into main Oct 6, 2026
15 checks passed
@FBumann
FBumann deleted the claude/gifted-darwin-b7918s branch October 6, 2026 15:51
FBumann added a commit that referenced this pull request Oct 6, 2026
> **Prompt:** cut the 0.6.1 release

> [!NOTE]
> The following content was generated by AI.

Merging this releases 0.6.1, a patch release with one change: the polars
`<2` cap (#1852). The diff is `CHANGELOG.md` alone.

<details><summary>What the section says, and what was checked</summary>

- `## Upcoming version` became `## 0.6.1 (2026-10-06)`, with a short
paragraph and the one PR line. A new, empty `## Upcoming version` is
above it.
- **Why a patch.** No import, file or archive breaks. `LAYOUT` is
unchanged, because `main` since `v0.6.0` is #1852 alone, which changes
`pyproject.toml` and `uv.lock`. The ceiling only refuses polars 2.
- **Not included:** #1850, the benchmark refresh, which is still open.
- `python -m tools.changelog check` prints `releases 0.6.1 on merge`.
`python -m tools.changelog notes 0.6.1` prints the section.
`tests/test_changelog.py`: 39 passed.
- On merge, `release.yaml` tags `v0.6.1`, opens the GitHub release and
builds. Then it waits for a reviewer to approve the `pypi` environment
before the upload.

</details>

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_018nKGXp8uKnMeJACjFYwbdY

---
_Generated by [Claude
Code](https://claude.ai/code/session_018nKGXp8uKnMeJACjFYwbdY)_

Co-authored-by: Claude <noreply@anthropic.com>
FBumann pushed a commit that referenced this pull request Oct 6, 2026
…n 0.6.0

bench.plot is re-run on the results #1850 committed. Those results name the
engine's arm specsolve, so the map that renamed lpspec goes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K42KwZJtqPPhf5hDzUD91o
FBumann added a commit that referenced this pull request Oct 6, 2026
… measurements and label our engine specsolve (#1851)

> **Prompt:** Could we radically simplify and improve it by using
tanstack charts 1.0.0 ?
>
> **Prompt:** And leverage the descriptive nature fully! We should
improve our data model around it probably! And use json etc — I would
like to have a clean, simple data model which is used by tanstack
charts, without too much data manipulation. — we shoud name it
specsolve!

> [!NOTE]
> The following content was generated by AI.

`bench.plot` writes `benchmarks-scaling.json`: tidy rows that TanStack
Charts marks read by field name. A JSON spec on the page declares the
charts. This replaces about 250 lines of hand-written SVG, and the page
labels our engine `specsolve`. The rows hold the numbers re-measured on
0.6.0 in #1850.

<details><summary>The data model and the spec</summary>

One row per model, sink, ladder, rung and library. A measurement carries
numbers. A refusal carries the budget label instead:

```json
{"model": "dispatch", "sink": "gurobi", "ladder": "length", "rung": "xs", "variables": 10000, "library": "gurobipy-loop", "wall_s": 0.0235, "wall_q1_s": 0.0225, "wall_q3_s": 0.0239, "peak_gb": 0.1998}
{"model": "fleet", "sink": "gurobi", "ladder": "length", "rung": "m", "variables": 1200000, "library": "pyomo", "refused": ">30 s"}
```

`variables` is the rung's declared size (`results.nominal`). It equals
the column count in every committed cell, so every library shares one x
per rung. The page passes these rows straight to `areaY` (`y1:
wall_q1_s`, `y2: wall_q3_s`) and `lineY` (`y: wall_s`, `z/color:
library`). The only manipulation is filtering the rows to one facet.
`library` is the harness arm name as it is.

The page spec is `<script type="application/json" id="spec">`. It holds
`facet: ["model", "sink"]`, the `x` field, one entry per metric with its
`y`, `band` and `format`, the initial `filter`, and the series colour
domain and range. The control buttons name a metric or a filter value
from it. About 60 lines of JS turn the spec into TanStack definitions.
The rest of the JS is the table, the legend and the controls. Formats
are named in the spec because a function is not JSON.

`bench.plot` no longer edits the HTML. The page is purely hand-written,
and the JSON is the generated file.

</details>

<details><summary>Loading TanStack without a bundler</summary>

`dist/` is unbundled ESM with bare `d3-*` imports. An import map points
`@tanstack/charts/` at the raw `dist/` files on jsdelivr, pinned to
`1.0.0`, so all the entry points share one module graph. `d3-shape` and
`d3-scale` map to their pinned `+esm` builds, and `d3-scale` provides
`scaleLog`. The six entry points the page imports pull 47 files, about
330 KB before compression, and 48 CDN requests in total. MathJax already
loads from unpkg, so a runtime CDN has precedent here. Nothing joins the
pixi or Python dependency sets.

</details>

<details><summary>What was verified, and what was not</summary>

- **Parity:** I compared every cell of the rows `bench.plot` writes from
the results committed in #1850 with the `const DATA` that #1850 wrote
into the old page on `main` (79258b1). All 112 cells are identical,
measurements and refusals both, with `polars` read as `specsolve`.
- **Browser:** headless Chromium against `docs/about/` and against the
strict-built `site/about/`, in light and dark. 5 panels, the tooltip,
the metric switch, the width ladder and hiding a library all work, with
no console errors.
- **Not checked:** I could not load jsdelivr itself, because the session
proxy blocks it. The test served the same files from the npm tarball:
the raw `dist/` unchanged, and an esbuild bundle in place of each d3
`+esm` URL. The first look at the deployed page is the check for that.
- **Gates:** on the branch rebased onto `main`, `pixi run lint`,
`format-check` and `typecheck` are clean, the bench harness tests pass
(266), and `docs-build --strict` passes and ships the JSON beside the
page. CI on the previous head was green.

</details>

<details><summary>Mutation table for <code>plot.rows</code></summary>

`tools.mutate` took the line deletions, with `--tests
bench/test_harness.py` and `CI=1` to pass the load guard: the run times
nothing, and the load was my own earlier test run. Hand mutations on the
refusal condition used the same three precautions. Line numbers are from
before the name map was removed.

| mutation | result |
|---|---|
| the emit-phase filter (`plot.py:72`) | **caught** |
| the no-peak filter (`plot.py:73`) | **caught** |
| the ladder-rung filter (`plot.py:74`) | **caught** by the new
`test_a_rung_neither_ladder_plots_is_left_out`. It was green before that
test. |
| the size-on-axis check dropped (line 84, by hand) | **caught** |
| the already-measured check dropped (line 84, by hand) | **caught** by
the new
`test_a_rung_measured_past_a_ceiling_is_a_measurement_not_a_refusal`. It
was green before that test. |
| the `'error' not in r` filter | It survived deletion. `results.py`
says nothing produces an `error` record, so it is deleted. |

</details>

<details><summary>Coverage moved, defaults departed from, and what was
not done</summary>

- **Tests:** the five tests on `plot.series` and `plot.panels` now
assert the same claims on `plot.rows`, through a `_refused(records,
library)` helper. `_plotted` is gone.
- **Type:** the diff touches `bench/`, `docs/`, `.github/` and
`pyproject.toml`. By the table, the topmost row is `docs/`. The
changelog line is added.
- **Stacked on this:** #1855 redesigns the page's charts on top of this
data model.
- **Not done:** no new measurements here. I also left the footer's stale
references to `bench/run.py` and `docs/benchmarks.md` as they were. They
predate this PR.

</details>

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01K42KwZJtqPPhf5hDzUD91o

---------

Co-authored-by: Claude <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants