Repository navigation
docs: the published benchmarks are re-measured on specsolve 0.6.0 - #1850
Conversation
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018nKGXp8uKnMeJACjFYwbdY
… was measured on Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018nKGXp8uKnMeJACjFYwbdY
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018nKGXp8uKnMeJACjFYwbdY
|
Note The following content was generated by AI. Two signals on this PR come from polars 2.0.0, not from this PR: the red polars 2.0.0 was released on PyPI on 2026-10-06 at 11:51 UTC.
How it was checked
Generated by Claude Code |
Merging this PR will not alter performance
Comparing Footnotes
|
…'s dual can come back empty and a model can take nine times the memory Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018nKGXp8uKnMeJACjFYwbdY
…b7918s # Conflicts: # CHANGELOG.md
> **Prompt:** cut the 0.6.1 release > [!NOTE] > The following content was generated by AI. Merging this releases 0.6.1, a patch release with one change: the polars `<2` cap (#1852). The diff is `CHANGELOG.md` alone. <details><summary>What the section says, and what was checked</summary> - `## Upcoming version` became `## 0.6.1 (2026-10-06)`, with a short paragraph and the one PR line. A new, empty `## Upcoming version` is above it. - **Why a patch.** No import, file or archive breaks. `LAYOUT` is unchanged, because `main` since `v0.6.0` is #1852 alone, which changes `pyproject.toml` and `uv.lock`. The ceiling only refuses polars 2. - **Not included:** #1850, the benchmark refresh, which is still open. - `python -m tools.changelog check` prints `releases 0.6.1 on merge`. `python -m tools.changelog notes 0.6.1` prints the section. `tests/test_changelog.py`: 39 passed. - On merge, `release.yaml` tags `v0.6.1`, opens the GitHub release and builds. Then it waits for a reviewer to approve the `pypi` environment before the upload. </details> 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_018nKGXp8uKnMeJACjFYwbdY --- _Generated by [Claude Code](https://claude.ai/code/session_018nKGXp8uKnMeJACjFYwbdY)_ Co-authored-by: Claude <noreply@anthropic.com>
…n 0.6.0 bench.plot is re-run on the results #1850 committed. Those results name the engine's arm specsolve, so the map that renamed lpspec goes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01K42KwZJtqPPhf5hDzUD91o
… measurements and label our engine specsolve (#1851) > **Prompt:** Could we radically simplify and improve it by using tanstack charts 1.0.0 ? > > **Prompt:** And leverage the descriptive nature fully! We should improve our data model around it probably! And use json etc — I would like to have a clean, simple data model which is used by tanstack charts, without too much data manipulation. — we shoud name it specsolve! > [!NOTE] > The following content was generated by AI. `bench.plot` writes `benchmarks-scaling.json`: tidy rows that TanStack Charts marks read by field name. A JSON spec on the page declares the charts. This replaces about 250 lines of hand-written SVG, and the page labels our engine `specsolve`. The rows hold the numbers re-measured on 0.6.0 in #1850. <details><summary>The data model and the spec</summary> One row per model, sink, ladder, rung and library. A measurement carries numbers. A refusal carries the budget label instead: ```json {"model": "dispatch", "sink": "gurobi", "ladder": "length", "rung": "xs", "variables": 10000, "library": "gurobipy-loop", "wall_s": 0.0235, "wall_q1_s": 0.0225, "wall_q3_s": 0.0239, "peak_gb": 0.1998} {"model": "fleet", "sink": "gurobi", "ladder": "length", "rung": "m", "variables": 1200000, "library": "pyomo", "refused": ">30 s"} ``` `variables` is the rung's declared size (`results.nominal`). It equals the column count in every committed cell, so every library shares one x per rung. The page passes these rows straight to `areaY` (`y1: wall_q1_s`, `y2: wall_q3_s`) and `lineY` (`y: wall_s`, `z/color: library`). The only manipulation is filtering the rows to one facet. `library` is the harness arm name as it is. The page spec is `<script type="application/json" id="spec">`. It holds `facet: ["model", "sink"]`, the `x` field, one entry per metric with its `y`, `band` and `format`, the initial `filter`, and the series colour domain and range. The control buttons name a metric or a filter value from it. About 60 lines of JS turn the spec into TanStack definitions. The rest of the JS is the table, the legend and the controls. Formats are named in the spec because a function is not JSON. `bench.plot` no longer edits the HTML. The page is purely hand-written, and the JSON is the generated file. </details> <details><summary>Loading TanStack without a bundler</summary> `dist/` is unbundled ESM with bare `d3-*` imports. An import map points `@tanstack/charts/` at the raw `dist/` files on jsdelivr, pinned to `1.0.0`, so all the entry points share one module graph. `d3-shape` and `d3-scale` map to their pinned `+esm` builds, and `d3-scale` provides `scaleLog`. The six entry points the page imports pull 47 files, about 330 KB before compression, and 48 CDN requests in total. MathJax already loads from unpkg, so a runtime CDN has precedent here. Nothing joins the pixi or Python dependency sets. </details> <details><summary>What was verified, and what was not</summary> - **Parity:** I compared every cell of the rows `bench.plot` writes from the results committed in #1850 with the `const DATA` that #1850 wrote into the old page on `main` (79258b1). All 112 cells are identical, measurements and refusals both, with `polars` read as `specsolve`. - **Browser:** headless Chromium against `docs/about/` and against the strict-built `site/about/`, in light and dark. 5 panels, the tooltip, the metric switch, the width ladder and hiding a library all work, with no console errors. - **Not checked:** I could not load jsdelivr itself, because the session proxy blocks it. The test served the same files from the npm tarball: the raw `dist/` unchanged, and an esbuild bundle in place of each d3 `+esm` URL. The first look at the deployed page is the check for that. - **Gates:** on the branch rebased onto `main`, `pixi run lint`, `format-check` and `typecheck` are clean, the bench harness tests pass (266), and `docs-build --strict` passes and ships the JSON beside the page. CI on the previous head was green. </details> <details><summary>Mutation table for <code>plot.rows</code></summary> `tools.mutate` took the line deletions, with `--tests bench/test_harness.py` and `CI=1` to pass the load guard: the run times nothing, and the load was my own earlier test run. Hand mutations on the refusal condition used the same three precautions. Line numbers are from before the name map was removed. | mutation | result | |---|---| | the emit-phase filter (`plot.py:72`) | **caught** | | the no-peak filter (`plot.py:73`) | **caught** | | the ladder-rung filter (`plot.py:74`) | **caught** by the new `test_a_rung_neither_ladder_plots_is_left_out`. It was green before that test. | | the size-on-axis check dropped (line 84, by hand) | **caught** | | the already-measured check dropped (line 84, by hand) | **caught** by the new `test_a_rung_measured_past_a_ceiling_is_a_measurement_not_a_refusal`. It was green before that test. | | the `'error' not in r` filter | It survived deletion. `results.py` says nothing produces an `error` record, so it is deleted. | </details> <details><summary>Coverage moved, defaults departed from, and what was not done</summary> - **Tests:** the five tests on `plot.series` and `plot.panels` now assert the same claims on `plot.rows`, through a `_refused(records, library)` helper. `_plotted` is gone. - **Type:** the diff touches `bench/`, `docs/`, `.github/` and `pyproject.toml`. By the table, the topmost row is `docs/`. The changelog line is added. - **Stacked on this:** #1855 redesigns the page's charts on top of this data model. - **Not done:** no new measurements here. I also left the footer's stale references to `bench/run.py` and `docs/benchmarks.md` as they were. They predate this PR. </details> 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01K42KwZJtqPPhf5hDzUD91o --------- Co-authored-by: Claude <noreply@anthropic.com>
Note
The following content was generated by AI.
The benchmark pages now show run 37298997947, measured on
aa902393(0.6.0). It covers the same five cases as the page it replaces. The page's sentences about the run are recomputed from the new files, andbench/reproduce.pyis pinned to that run.What changed
highsdispatch and fleet, andgurobidispatch, transport and fleet. The other three cases were stopped by the memory watchdog, as on the previous run (the OOM record dies with the run artifact instead of landing beside the numbers it explains #1498).results-gurobi-fleetholds all ten files.casualties.jsonis not committed. It was never committed, andbench/results.pyskips it. Its three entries are in the page text.pixi run -e bench reportand the chart data frompixi run -e bench plot. Their output, fingerprint included, is the same as the run's ownreportandplotsteps.main's results.transport/w1ongurobiagainst gurobipy-loop, in specsolve's favour. Before, it wasdispatch/sagainst linopy, against specsolve. With it goes the the gurobi sink alternates between a fast and a slow build, round after round #1288 sentence: specsolve's rounds on that cell are now 88 to 94 ms.storagesurvived ongurobiearlier. The previous run already contradicted it.bench/reproduce.py.test_the_lock_installs_what_the_published_numbers_were_taken_onfailed on the new files, because the lock installedlpspecat8d27e88b. Changes:specsolve[gurobi]ataa90239393, and linopy frommaster, because thelinopyextra is gone (refactor(api): specsolve no longer builds a linopy model, and its linopy extra and LaneError are gone #1755).lpspec 0.0.1a61numbers is removed.mainis merged in, with fix(deps): specsolve installs a polars older than 2.0, on which a row's dual can come back empty and a model can take nine times the memory #1852's polars cap. The cap was ported here first, and the merge makes it identical tomain. The only conflict was inCHANGELOG.md, which now keeps both lines.Found, and not fixed here
xscell on specsolve, linopy or pyomo is 1.9 to 22.6 times the median of the other eight. Pyomodispatch/xsonhighstakes 1524 ms, then about 67 ms.main's results from 14 September have no such round. Published medians do not move, because that round is always the largest of nine. The q3 of a band may move a little._roundshadwarmup_rounds=0at the previous run too, so the cause is elsewhere in what changed after8d27e88b. I did not look further.benchmarks-scaling.htmlsays "a quick cell here took 84 rounds and a slow one 9". That sentence is hand-written and was already false, because rounds are pinned at 9.master(f665a260). The guard checks only the versions the run records.Gates
pixi run check, before themainmerge, on polars 1.44.2: 4996 passed, 564 skipped, 1 xfailed. After the merge:tools.changelog check,tests/test_changelog.pyandtests/test_tooling_pins.pypass.pytest bench/test_harness.py: 262 passed, 16 skipped. This istest-bench, includingtest_the_report_renders_from_the_committed_resultsand the lock guard.pixi run docs-build, the command Read the Docs runs: clean.uv lock --script bench/reproduce.py --check: clean.pixi run lintandformat-check: clean.uv run --locked bench/reproduce.pyitself.🤖 Generated with Claude Code
https://claude.ai/code/session_018nKGXp8uKnMeJACjFYwbdY