Skip to content

fix(deps): specsolve installs a polars older than 2.0, on which a row's dual can come back empty and a model can take nine times the memory - #1852

Merged
FBumann merged 2 commits into
mainfrom
claude/polars-below-2
Oct 6, 2026
Merged

FBumann merged 2 commits into
mainfrom
claude/polars-below-2

Conversation

@FBumann

@FBumann FBumann commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator

Prompt: Do a pr that fixes the polars 2 issue. Cap it for now!

Note

The following content was generated by AI.

specsolve now requires polars>=1.30,<2. On polars 2.0.0, released on 2026-10-06, two tests get an empty dual where a row has one, and storage/s takes about nine times the memory. Making specsolve work on polars 2 is not part of this PR.

What fails on polars 2.0.0

Both results come from the same tree, with only polars swapped (1.44.2 against 2.0.0):

  • Suite. Two tests fail: dual('budget_row') returns 0 rows where 1 is expected.

    • tests/test_resolution_parity.py::test_the_empty_coordinate_builds_on_both_lanes
    • tests/test_resolution_parity.py::test_a_masked_scalar_variable_takes_its_row_with_it[masked-in]

    This is how CI went red on docs: the published benchmarks are re-measured on specsolve 0.6.0 #1850, and it fails the same way on every PR and on main.

  • Memory. test_emit[storage-s-specsolve-highs] with --benchmark-memory, three rounds of one local run:

polars peak_bytes RSS
1.44.2 119–123 MB 234 MB
2.0.0 1091–1146 MB 1072 MB

CodSpeed reported the same direction on #1850: the storage/s specsolve cells went from about 17 MB to about 776 MB.

What changed, and gates
  • pyproject.toml. polars>=1.30 becomes polars>=1.30,<2. tests/test_tooling_pins.py already accepts a ,< ceiling. The floor of 1.30, and the floors environment that pins it, do not change.
  • uv.lock. Relocked with uv lock. Besides the polars specifier, the relock brings in three pins that pyproject.toml on main already declares and the lock had not caught up with: ruff 0.16.9, zensical 0.0.66 and pymdown-extensions 12.1. I did not edit the lock by hand to keep them out.
  • CHANGELOG.md. The line for this PR.
  • The cap works. uv pip compile of polars>=1.30 resolves 2.0.0. With ,<2 it resolves 1.44.2. A fresh pixi environment in this worktree also resolves 1.44.2.
  • Gates. pixi run check: 4996 passed, 564 skipped, 1 xfailed, the two tests above included. This was run in the fresh environment, which has polars 1.44.2.
  • Not run. test-floors, because the floors pins do not change. docs-build, because no page changes.
  • Not done. No test asserts the ceiling. With the ceiling removed, a fresh install resolves 2.0.0 and the suite fails on the two tests above, which is the guard. A test that pins <2 would have to be deleted with the cap.
  • Type. fix(deps), as in fix(deps): an installed specsolve release keeps the mathspec minor version it was released against #1816: the ceiling changes what the package accepts.

🤖 Generated with Claude Code

https://claude.ai/code/session_018nKGXp8uKnMeJACjFYwbdY


Generated by Claude Code

claude added 2 commits October 6, 2026 15:13
…'s dual can come back empty and a model can take nine times the memory

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018nKGXp8uKnMeJACjFYwbdY
…'s dual can come back empty and a model can take nine times the memory

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018nKGXp8uKnMeJACjFYwbdY
@FBumann
FBumann enabled auto-merge (squash) October 6, 2026 15:17
@FBumann
FBumann merged commit 85d864a into main Oct 6, 2026
17 of 23 checks passed
@FBumann
FBumann deleted the claude/polars-below-2 branch October 6, 2026 15:19
@codspeed

codspeed Bot commented Oct 6, 2026

Copy link
Copy Markdown

Merging this PR will not alter performance

✅ 67 untouched benchmarks
⏩ 307 skipped benchmarks1


Comparing claude/polars-below-2 (8c7b794) with main (331a839)2

Open in CodSpeed

Footnotes

  1. 307 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

  2. No successful run was found on main (aa90239) during the generation of this report, so 331a839 was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

@FBumann FBumann mentioned this pull request Oct 6, 2026
FBumann added a commit that referenced this pull request Oct 6, 2026
)

> **Prompt:** The benchmark ran. Can you update our published benchmark
with a PR?

> [!NOTE]
> The following content was generated by AI.

The benchmark pages now show run
[37298997947](https://github.com/fluxopt/specsolve/actions/runs/37298997947),
measured on `aa902393` (0.6.0). It covers the same five cases as the
page it replaces. The page's sentences about the run are recomputed from
the new files, and `bench/reproduce.py` is pinned to that run.

<details><summary>What changed</summary>

- **Results.** Five cases: `highs` dispatch and fleet, and `gurobi`
dispatch, transport and fleet. The other three cases were stopped by the
memory watchdog, as on the previous run (#1498).
- **Where the files come from.** The network policy of the session
blocks the artifact host, so the files come from the eight per-case
artifacts, uploaded by hand. Every results file is byte-identical in
each artifact that holds it. `results-gurobi-fleet` holds all ten files.
- **`casualties.json` is not committed.** It was never committed, and
`bench/results.py` skips it. Its three entries are in the page text.
- **Generated content.** The tables come from `pixi run -e bench report`
and the chart data from `pixi run -e bench plot`. Their output,
fingerprint included, is the same as the run's own `report` and `plot`
steps.
- **Hand-written sentences, recomputed.** Before I used each method on
the new files, I checked that it reproduces the old sentence exactly
from `main`'s results.
- **Mean against median:** 14 cells are above 1.10x, against 4 before.
The worst is 3.41x, against 1.21x before.
- **The cell the median flips:** now `transport/w1` on `gurobi` against
gurobipy-loop, in specsolve's favour. Before, it was `dispatch/s`
against linopy, against specsolve. With it goes the #1288 sentence:
specsolve's rounds on that cell are now 88 to 94 ms.
  - **The three casualties:** 24.6, 26.3 and 23.9 GB.
- **Removed:** the sentence on why `storage` survived on `gurobi`
earlier. The previous run already contradicted it.
- **`bench/reproduce.py`.**
`test_the_lock_installs_what_the_published_numbers_were_taken_on` failed
on the new files, because the lock installed `lpspec` at `8d27e88b`.
Changes:
- The script now installs `specsolve[gurobi]` at `aa90239393`, and
linopy from `master`, because the `linopy` extra is gone (#1755).
  - The lock moves pandas to 3.0.6, as measured.
  - The docstring line about the `lpspec 0.0.1a61` numbers is removed.
- **`main` is merged in, with #1852's polars cap.** The cap was ported
here first, and the merge makes it identical to `main`. The only
conflict was in `CHANGELOG.md`, which now keeps both lines.

</details>

<details><summary>Found, and not fixed here</summary>

- **A cold first round.** In 13 cells, the first of the nine rounds of
an `xs` cell on specsolve, linopy or pyomo is 1.9 to 22.6 times the
median of the other eight. Pyomo `dispatch/xs` on `highs` takes 1524 ms,
then about 67 ms. `main`'s results from 14 September have no such round.
Published medians do not move, because that round is always the largest
of nine. The q3 of a band may move a little. `_rounds` had
`warmup_rounds=0` at the previous run too, so the cause is elsewhere in
what changed after `8d27e88b`. I did not look further.
- **The 84-rounds sentence.** `benchmarks-scaling.html` says "a quick
cell here took 84 rounds and a slow one 9". That sentence is
hand-written and was already false, because rounds are pinned at 9.
- **linopy's commit.** The run does not record which linopy commit it
measured, so the lock pins today's `master` (`f665a260`). The guard
checks only the versions the run records.

</details>

<details><summary>Gates</summary>

- `pixi run check`, before the `main` merge, on polars 1.44.2: 4996
passed, 564 skipped, 1 xfailed. After the merge: `tools.changelog
check`, `tests/test_changelog.py` and `tests/test_tooling_pins.py` pass.
- `pytest bench/test_harness.py`: 262 passed, 16 skipped. This is
`test-bench`, including
`test_the_report_renders_from_the_committed_results` and the lock guard.
- `pixi run docs-build`, the command Read the Docs runs: clean.
- `uv lock --script bench/reproduce.py --check`: clean. `pixi run lint`
and `format-check`: clean.
- **Not run:** `uv run --locked bench/reproduce.py` itself.

</details>

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_018nKGXp8uKnMeJACjFYwbdY

---------

Co-authored-by: Claude <noreply@anthropic.com>
FBumann added a commit that referenced this pull request Oct 6, 2026
> **Prompt:** cut the 0.6.1 release

> [!NOTE]
> The following content was generated by AI.

Merging this releases 0.6.1, a patch release with one change: the polars
`<2` cap (#1852). The diff is `CHANGELOG.md` alone.

<details><summary>What the section says, and what was checked</summary>

- `## Upcoming version` became `## 0.6.1 (2026-10-06)`, with a short
paragraph and the one PR line. A new, empty `## Upcoming version` is
above it.
- **Why a patch.** No import, file or archive breaks. `LAYOUT` is
unchanged, because `main` since `v0.6.0` is #1852 alone, which changes
`pyproject.toml` and `uv.lock`. The ceiling only refuses polars 2.
- **Not included:** #1850, the benchmark refresh, which is still open.
- `python -m tools.changelog check` prints `releases 0.6.1 on merge`.
`python -m tools.changelog notes 0.6.1` prints the section.
`tests/test_changelog.py`: 39 passed.
- On merge, `release.yaml` tags `v0.6.1`, opens the GitHub release and
builds. Then it waits for a reviewer to approve the `pypi` environment
before the upload.

</details>

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_018nKGXp8uKnMeJACjFYwbdY

---
_Generated by [Claude
Code](https://claude.ai/code/session_018nKGXp8uKnMeJACjFYwbdY)_

Co-authored-by: Claude <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants