Skip to content

feat(deps): specsolve runs on polars 2.0 and requires it - #1857

Open
FBumann wants to merge 4 commits into
mainfrom
claude/friendly-dirac-t7g8rd
Open

FBumann wants to merge 4 commits into
mainfrom
claude/friendly-dirac-t7g8rd

Conversation

@FBumann

@FBumann FBumann commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator

Prompt: Lets work on the bump to polars 2.0.0 in specsolve. We should leverage their new engine things and check performance issues!

Note

The following content was generated by AI.

specsolve now requires polars>=2.0, and lifts the <2 cap from #1852. On polars 2.0.0 it reads a scalar declaration correctly again, and it builds every ladder rung in the memory and time it took on 1.44.2. Three polars 2 changes caused the breakage.

What broke on polars 2.0, and the fix for each
  1. A scalar declaration read back empty. On polars 2, select() with no columns gives a frame of height 0. Every reader of a declaration with no dimensions did select(*dims).with_columns(share). So primal, dual and activity came back empty, and evaluate('dual(row)') read 0.0 where the answer is 2.0. This is the dual failure in fix(deps): specsolve installs a polars older than 2.0, on which a row's dual can come back empty and a model can take nine times the memory #1852. The fix is Labelled.valued, which attaches the share first and then projects.
  2. Memory grew up to 9× on storage. polars 2 adds cost-based join reordering (QueryOptFlags.join_order). On the cyclic shift, it moves the constraint rows' join between the shift's two joins, keyed on store alone. That intermediate holds every pair of snapshots per store. The same thing reproduces in plain polars with no specsolve involved: 20,000 result rows peak at 557 MB with reordering and 66 MB without. specsolve orders its joins by hand, so every collect now goes through relational.collect.collected, which runs with join_order=False. An architecture test refuses a direct .collect() / collect_all() anywhere else in src/.
  3. Builds were up to 38% slower on transport. The streaming cross join now picks which side it buffers from the table sizes (perf: Use stats to decide cross join buffering side pola-rs/polars#29270). The coordinate product therefore lost the label order that Scope.product relied on, and each label frame had to be sorted (7M rows: 0.36 s). Each cross join now passes maintain_order='left_right' and folds forward, so the order is guaranteed rather than incidental.

The floor. join_order does not exist in polars 1.x, so the floor moves to 2.0. A version branch was the alternative. floors now pins polars==2.0. The pypsa-parity line moves to polars>=2.0; pypsa 1.3.0 resolves with polars 2.0.0.

What polars 2 gives us. The engine changes are mostly free speed. With the three fixes, dispatch/l builds 12% faster and transport/l 5% faster than on 1.44.2, with lower peaks. Out-of-core spilling is on by default in polars 2, and no code here needs it. The new knobs (POLARS_OOC_*, hot-table size, runtime join filters, engine affinity) did not move storage/s at all; only join_order did.

Benchmarks: build seconds (minimum) / peak RSS (median), 3 rounds, fresh process per cell

Method: bench.arms.specsolve.build_and_emit('highs', …) (build and HiGHS load, no solve) in a fresh process per cell, with configurations alternated cell by cell on an idle 4-core machine. main is 79258b1. main-2.0.0 is main with polars swapped (installed --no-deps past the cap). A watchdog kills a process above 10 GB RSS.

cell main / 1.44.2 main / 2.0.0 this PR / 2.0.0
dispatch/2xs 0.017s / 188 MB 0.015s / 187 MB 0.016s / 187 MB
dispatch/xs 0.018s / 191 MB 0.022s / 190 MB 0.020s / 190 MB
dispatch/s 0.030s / 212 MB 0.033s / 210 MB 0.033s / 210 MB
dispatch/m 0.134s / 398 MB 0.103s / 394 MB 0.121s / 390 MB
dispatch/l 0.928s / 1521 MB 0.747s / 1417 MB 0.816s / 1418 MB
fleet/2xs 0.112s / 189 MB 0.121s / 188 MB 0.106s / 188 MB
fleet/xs 0.118s / 193 MB 0.120s / 192 MB 0.109s / 192 MB
fleet/s 0.134s / 225 MB 0.150s / 222 MB 0.138s / 222 MB
fleet/m 0.376s / 492 MB 0.393s / 477 MB 0.386s / 483 MB
fleet/l 2.790s / 2265 MB 2.625s / 2161 MB 2.595s / 2160 MB
transport/2xs 0.024s / 190 MB 0.025s / 189 MB 0.025s / 189 MB
transport/xs 0.028s / 194 MB 0.030s / 194 MB 0.030s / 194 MB
transport/s 0.052s / 227 MB 0.060s / 230 MB 0.052s / 226 MB
transport/m 0.253s / 524 MB 0.266s / 561 MB 0.225s / 518 MB
transport/l 2.134s / 1776 MB 2.560s / 1895 MB 2.024s / 1705 MB
transport/w10 0.052s / 230 MB 0.060s / 233 MB 0.052s / 229 MB
transport/w100 0.252s / 533 MB 0.296s / 552 MB 0.249s / 530 MB
transport/w1000 2.380s / 1758 MB 2.949s / 1963 MB 2.256s / 1764 MB
storage/2xs 0.036s / 189 MB 0.040s / 188 MB 0.037s / 188 MB
storage/xs 0.041s / 193 MB 0.051s / 204 MB 0.043s / 192 MB
storage/s 0.064s / 222 MB 0.295s / 988 MB 0.074s / 220 MB
storage/m 0.270s / 490 MB killed above 10 GB 0.307s / 477 MB
storage/l 2.489s / 1871 MB killed above 10 GB 2.593s / 1751 MB
storage/w10 0.067s / 223 MB 0.095s / 298 MB 0.069s / 221 MB
storage/w100 0.288s / 480 MB 0.501s / 1212 MB 0.273s / 473 MB
storage/w1000 2.435s / 1815 MB 8.301s / 8819 MB 2.373s / 1727 MB

storage/m (+14%) and storage/l (+4%) are the cells where this PR builds slower than 1.44.2. An earlier run of the same cells gave 0.297 s and 2.533 s against 0.307 s and 2.615 s, so I read both as noise. This is not the pixi run ladder harness and not CodSpeed. xl/2xl were not run.

Mutation table: hand mutations, full suite with -x

Each mutation is a substitution rather than a deletion, so it was taken by hand: committed tree, git checkout -- restore, __pycache__ dropped on both sides, tree checked clean afterwards.

mutation result tree
labels.py valued: select the dims before attaching the share caught: test_a_declaration_with_no_dimensions_reads_back_its_one_value[primal] clean
collect.py: QueryOptFlags(), so join reordering is on caught: test_a_join_runs_in_the_order_it_was_written clean
scope.py product: no maintain_order caught: test_a_coordinate_product_arrives_in_label_order clean
assumptions.py: one bare .collect() caught: test_every_frame_is_collected_through_one_function clean

The scalar-reader test was written first as a strict xfail. It failed on polars 2.0.0 ([] == [{'value': 10.0}]) and XPASSed on 1.44.2. The marker came off with the fix.

Gates, and what was not done
  • Gates. ruff check ., ruff format --check . and pyrefly are clean. The suite on polars 2.0.0 gives 5002 passed, 564 skipped, 1 xfailed; before the fixes it gave the two failures of fix(deps): specsolve installs a polars older than 2.0, on which a row's dual can come back empty and a model can take nine times the memory #1852. pixi run -e floors test-floors on polars 2.0.0 gives 3271 passed. pixi run -e docs docs-build is strict and clean.
  • uv.lock. Relocked with uv lock (uv 0.11.32). Besides polars 2.0.0, it drops platform markers on the textual / rich / jinja2 dependents; I did not hand-edit them out.
  • Docs. The relational/collect.py row in docs/about/architecture.md and the "Collect engine" entry in docs/reference/internals.md now say every collect goes through collected, with join reordering off.
  • Not done.
    • The reordering bug is not reported upstream to pola-rs. A standalone reproducer is in tests/test_collect.py::_shift_shaped.
    • sink_parquet calls do not go through collected. Their frames are results with no join chain.
    • A separate bug exists on 1.44.2 and 2.0.0 alike. An expression that reads a scalar variable (2 * slack) raises ColumnNotFoundError: __unit__, and so does save() with such an expression declared. It is left for its own PR.
  • Departures. The floor is raised rather than kept at 1.30, because the join_order flag needs polars 2.

🤖 Generated with Claude Code

https://claude.ai/code/session_01EpzN6XKNByxsfi3GhX3v3D


Generated by Claude Code

claude added 3 commits October 6, 2026 16:24
polars 2.0 changed three things specsolve leaned on:

- A select of no columns is a frame of height 0, so every reader of a
  declaration with no dimensions came back empty, and an expression
  reading its dual read 0.0. Labelled.valued attaches the share before it
  projects.
- Cost-based join reordering put a constraint's rows between a shift's two
  joins, keyed on one dimension of the two. Every collect now goes through
  relational.collect.collected, with join_order off.
- The streaming cross join picks its buffered side from the table sizes,
  so the coordinate product lost label order and had to be sorted. Each
  cross join now maintains left-then-right order.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EpzN6XKNByxsfi3GhX3v3D
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EpzN6XKNByxsfi3GhX3v3D
@codspeed

codspeed Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Merging this PR will regress 6 benchmarks

⚠️ Different runtime environments detected

Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.

Open the report in CodSpeed to investigate

⚡ 15 improved benchmarks
❌ 6 regressed benchmarks
✅ 46 untouched benchmarks
⏩ 307 skipped benchmarks1

Warning

Please fix the performance issues or acknowledge them on CodSpeed.

Performance Changes

Benchmark BASE HEAD Efficiency
❌ test_emit[fleet-s-specsolve-lp] 26.2 MB 32.3 MB -18.75%
❌ test_emit[commitment-s-specsolve-lp] 2.6 MB 3.2 MB -18.22%
❌ test_emit[nodal-s-specsolve-lp] 6 MB 7 MB -12.97%
❌ test_emit[storage-s-specsolve-lp] 21.7 MB 24.7 MB -11.94%
❌ test_read[commitment-s-specsolve-parquet] 695.3 KB 775.5 KB -10.34%
❌ test_read[storage-s-specsolve-parquet] 2.6 MB 2.8 MB -7.4%
⚡ test_window[sector-s-specsolve-highs-shape] 5.3 MB 4.1 MB +28.67%
⚡ test_window[sector-s-specsolve-highs-coefficient] 5.3 MB 4.1 MB +28.09%
⚡ test_window[sector-s-specsolve-highs-cold] 5.3 MB 4.1 MB +27.79%
⚡ test_window[sector-s-specsolve-highs-one] 5.3 MB 4.2 MB +26.3%
⚡ test_emit[sector-s-specsolve-highs] 5.3 MB 4.2 MB +26.26%
⚡ test_window[sector-s-specsolve-highs-values] 5.3 MB 4.2 MB +24.91%
⚡ test_emit[sector-s-specsolve-lp] 5.3 MB 4.7 MB +14.71%
⚡ test_window[storage-s-specsolve-highs-one] 13.8 MB 13 MB +5.85%
⚡ test_window[storage-s-specsolve-highs-values] 13.8 MB 13 MB +5.85%
⚡ test_window[fleet-s-specsolve-highs-values] 17.4 MB 16.5 MB +5.57%
⚡ test_window[fleet-s-specsolve-highs-one] 17.4 MB 16.5 MB +5.57%
⚡ test_emit[storage-s-specsolve-highs] 16.9 MB 16 MB +5.49%
⚡ test_window[commitment-s-specsolve-highs-values] 1.6 MB 1.5 MB +5.05%
⚡ test_window[commitment-s-specsolve-highs-one] 1.6 MB 1.5 MB +5.05%
... ... ... ... ...

ℹ️ Only the first 20 benchmarks are displayed. Go to the app to view all benchmarks.

Tip

Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.


Comparing claude/friendly-dirac-t7g8rd (94340ef) with main (7af455a)2

Open in CodSpeed

Footnotes

  1. 307 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

  2. No successful run was found on main (4a6a998) during the generation of this report, so 7af455a was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

@read-the-docs-community

read-the-docs-community Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Documentation build overview

📚 specsolve | 🛠️ Build #34980930 | 📁 Comparing 94340ef against latest (79258b1)

  🔍 Preview build  

6 files changed · ± 5 modified · - 1 deleted

± Modified

- Deleted

@FBumann
FBumann removed this pull request from stack #1866 October 6, 2026 22:36
@FBumann
FBumann added this pull request to stack #1867 October 6, 2026 22:36

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants