Skip to content

perf(engine): a small model builds faster, on the in-memory engine - #1863

Open
FBumann wants to merge 5 commits into
claude/friendly-dirac-t7g8rdfrom
perf/engine-by-size
Open

FBumann wants to merge 5 commits into
claude/friendly-dirac-t7g8rdfrom
perf/engine-by-size

Conversation

@FBumann

@FBumann FBumann commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

Prompt: Let's continue working on using polars v2. Can we leverage the out of core spillage? And other code paths that are faster now? Or otherwise improve performance?

Prompt, on the memory result below: Memory isn't an issue at those sizes though…

Note

The following content was generated by AI.

A build whose largest declaration has fewer than 250,000 coordinates now collects on polars' in-memory engine. Builds at 2xs are 23–35% faster, and a 16-slice sweep is 4–17% faster. No measured build is slower beyond noise. CodSpeed's heap benchmark puts sector/s at 4.1 → 6.1 MB. Its wall-time report flags m and l cells that a local A/B does not reproduce (see below).

Stacked on #1857. Merge that first.

Why, and what changed
  • Why. The streaming engine starts up for every query, and a small build is many small queries. In polars 2.0 engine="auto" resolves to streaming, so polars does not make this choice itself (storage/2xs builds in 0.069 s on auto, 0.076 s on streaming, 0.045 s in-memory).
  • relational/collect.py.
    • engine_for(coordinates) is the one decision. It answers in-memory below IN_MEMORY_BELOW, or where this polars has no streaming engine, and streaming otherwise.
    • collecting_on(engine) is the block every collected() inside runs in.
    • streaming_available() is the probe that was collect_engine(). The new name says what it answers: since this PR, it no longer names the engine every collect uses.
    • Outside a build, a collect takes engine_for of an unknown size, which is streaming as before.
  • Engine.build runs the assembly inside collecting_on(engine_for(_largest(...))). _largest is the coordinate count of the largest variable or constraint, unmasked, which is known once the sources are attached. That with clause and that helper are the whole change on the engine side.
  • Not added. The decision is not stored on the built model, and no caller can choose the engine. Nothing reads either yet. Sized reads, or a user-chosen engine, would add them where they are read.
  • Where the gain is. It comes from the build alone. Moving source reading or result reads to in-memory as well changed 2xs solve-and-read by noise only (storage: 64 ms none, 52 ms build only, 51 ms everywhere).
  • The threshold. At 1,000,000, storage/m (400k coordinates) built 14% slower in memory. Every rung that gained by a clear margin has at most about 120k coordinates, and the mixed results start at 400k. At 250,000, no measured cell regresses in time.
  • Docs. The relational/collect.py row in docs/about/architecture.md and the "Collect engine" glossary entry say so.
  • Housekeeping. The first commit carried a local bench/.cache symlink by mistake. A later commit removes it, so it is not in this PR's diff.
Wall time: CodSpeed's macro run, and a local A/B of the cells it flags

CodSpeed's wall-time run on 54c7907 against 94340ef reports 37 regressions in all (the 7 memory ones above included) and 76 improvements. Its 13 largest wall-time regressions are all at m or l:

  • Most of them run the same code on both sides. sector/m and nodal/m have 1.2M coordinates in their largest declaration, so they stream with or without this PR, and every test_read runs the same path.
  • No wall-time baseline was measured on 94340ef. That commit's Wall time (macro runner) job was skipped, so CodSpeed compared against a baseline taken elsewhere.

The local A/B used pytest bench --cases sector nodal commitment dispatch transport --sizes 2xs m l --arms specsolve --sinks highs lp --builds 0 on #1857's head 94340ef and on 54c7907. As on CodSpeed's runner, it runs every 2xs cell in the same process before m and l. Three alternations on an idle 4-core machine; the minimum per cell:

cell CodSpeed, base → head local, #1857 → this PR
test_read[commitment-m-specsolve-frames] 5.4 → 7.4 ms 3.3 → 3.4 ms (+3.5%)
test_emit[sector-m-specsolve-highs] 163.8 → 212.3 ms 113.3 → 115.9 ms (+2.3%)
test_read[nodal-l-specsolve-frames] 24.2 → 31.1 ms 37.9 → 36.8 ms (-3.0%)
test_read[commitment-l-specsolve-frames] 8.6 → 10.7 ms 9.9 → 9.5 ms (-3.4%)
test_read[dispatch-m-specsolve-frames] 5.4 → 6.7 ms 6.5 → 6.1 ms (-5.4%)
test_read[sector-l-specsolve-frames] 19.3 → 23.3 ms 25.6 → 27.6 ms (+7.6%)
test_window[nodal-m-specsolve-highs-cold] 177.6 → 213.6 ms 132.6 → 115.4 ms (-13.0%)
test_window[sector-m-specsolve-highs-one] 235.9 → 280.4 ms 138.9 → 124.9 ms (-10.1%)
test_window[sector-m-specsolve-highs-coefficient] 186.6 → 219.8 ms 110.5 → 106.0 ms (-4.1%)
test_read[transport-l-specsolve-frames] 39.2 → 45.6 ms 56.6 → 56.1 ms (-0.9%)
test_emit[nodal-m-specsolve-lp] 275.2 → 316.3 ms 179.6 → 171.7 ms (-4.4%)
test_window[nodal-m-specsolve-highs-one] 328.5 → 376.8 ms 176.4 → 183.1 ms (+3.8%)
test_window[sector-m-specsolve-highs-values] 241.4 → 276.6 ms 132.2 → 132.7 ms (+0.4%)
  • The m and l cells. Locally they move by −13% to +7.6%, in both directions. A run without the 2xs cells first gave −6.5% to +6.9%.
  • The 2xs cells in the same runs. test_emit is 3–20% faster and test_window 2–18% faster; test_read moves by −7% to +19%, on a path this PR does not change.
  • A clean comparison. CodSpeed would need a wall-time run on feat(deps): specsolve runs on polars 2.0 and requires it #1857's head as the baseline. Neither run here is CodSpeed's macro runner, whose absolute times differ from this machine's.
Memory: CodSpeed's heap benchmark

CodSpeed's memory run on 54c7907 against 94340ef reports 7 regressions and 5 improvements:

  • Regressions. All seven are sector at s, on test_emit and every test_window variant: peak heap 4.1–4.7 MB → 6.1–6.2 MB. sector/s's largest declaration has 120,000 coordinates, so it now builds in memory, and the in-memory engine materialises each join's result where streaming takes it in chunks. sector crosses a sparse portfolio with dense carriers, so its join results are the largest relative to the model.
  • Improvements. profiled/s read 7.6 → 6.7 MB, nodal/s windows 4.5 → 4.1 MB, transport/s windows 14.7 → 13.6 MB.
  • Not reproduced locally. memray, which --benchmark-memory uses, counts the system allocator, and polars allocates through its bundled jemalloc. memray reads 96.3 MB on sector/s for both trees, so it can neither confirm nor refute CodSpeed's figure. How the extra heap grows between 120k coordinates and the threshold is not measured.
  • The maintainer's view. The second prompt above is their answer to this result.
Benchmarks: build seconds (minimum) / peak RSS (median), fresh process per cell, polars 2.0.0

Method: bench.arms.specsolve.build_and_emit('highs', …) (build and HiGHS load, no solve), with configurations alternated cell by cell on an idle 4-core machine. The base is #1857's head df06eb3. These numbers were taken on the PR's first commit. The later restructures choose the same engine for every cell.

Threshold at 250k, 5 rounds:

cell #1857 this PR
dispatch/s 0.033s / 194 MB 0.033s / 193 MB
dispatch/m 0.112s / 373 MB 0.109s / 379 MB
transport/s 0.049s / 209 MB 0.040s / 203 MB
transport/m 0.243s / 504 MB 0.239s / 502 MB
storage/s 0.067s / 204 MB 0.053s / 204 MB
storage/m 0.265s / 460 MB 0.265s / 459 MB
fleet/s 0.135s / 206 MB 0.101s / 203 MB
fleet/m 0.320s / 465 MB 0.284s / 437 MB

2xs, xs and l, 3 rounds (taken with the threshold at 1M; at these rungs the two thresholds choose the same engine):

cell #1857 this PR
dispatch/2xs 0.017s 0.012s
dispatch/xs 0.020s 0.016s
transport/2xs 0.026s 0.017s
transport/xs 0.027s 0.023s
storage/2xs 0.040s 0.031s
storage/xs 0.041s 0.037s
fleet/2xs 0.111s 0.079s
fleet/xs 0.115s 0.084s
dispatch/l 0.806s 0.729s
transport/l 2.356s 2.173s
storage/l 2.729s 2.847s
fleet/l 3.391s 3.180s

The l rungs stream in both columns. Their spread (−10% to +4%) is the noise floor of that run, not an effect of this PR.

The workloads small builds serve, best of 7:

cell solve_over ×16, #1857 → this PR solve and read every primal, #1857 → this PR
dispatch/2xs 0.609 → 0.584 s 34.8 → 32.3 ms
dispatch/xs 0.887 → 0.848 s 53.2 → 43.3 ms
transport/2xs 1.082 → 0.979 s 61.9 → 55.9 ms
transport/xs 1.900 → 1.638 s 107.0 → 91.0 ms
storage/2xs 1.273 → 1.116 s 74.8 → 62.7 ms
storage/xs 2.731 → 2.481 s 146.7 → 146.8 ms
fleet/2xs 2.227 → 1.846 s 134.7 → 101.5 ms
fleet/xs 2.735 → 2.413 s 161.5 → 140.2 ms

Not the pixi run ladder harness. xl and 2xl were not run.

Coverage, mutation table, gates
  • Coverage. The suite's models are small, so its builds now take the in-memory engine. test_every_port_builds_the_same_model_on_either_engine builds every port on both engines and compares each hand-off frame (obj sorted, since it has no order contract) and row_starts. The comparison is to 1e-12, not to the bit. On transport_dantzig, transport_pwl and osemosys_utopia, the two engines round a * b / c differently, by at most 7e-15.
  • Saved answers. Because of that rounding, a model at most one unit in the last place apart has another contents digest. A small model's answer saved under feat(deps): specsolve runs on polars 2.0 and requires it #1857 (streamed) can be refused when it is rebuilt here (in memory). The same model always picks the same engine, so a save and a rebuild on one version agree. LAYOUT is not raised: nothing written to disk changes shape, and archive.py already says two polars versions may digest one table differently.
  • Mutation table. Each mutation was taken by hand on 373f41c, a substitution on a committed tree, with the full suite run with -x:
mutation result tree
engine.py: the build runs outside collecting_on caught: test_a_build_collects_on_the_engine_its_largest_declaration_sizes[small] clean
engine.py: the size ignored (engine_for(0)) caught: test_a_build_collects_on_the_engine_its_largest_declaration_sizes[at-the-threshold] clean
collect.py: >= becomes > in engine_for caught: test_the_engine_is_sized_to_the_model[at] clean
collect.py: the streaming probe ignored in engine_for caught: test_the_probe_sees_the_refusal clean
collect.py: collecting_on ignored caught: test_a_named_engine_holds_inside_its_block_only clean
Measured and not done
  • Out-of-core spilling (on by default in polars 2) does not lower specsolve's peak. storage/l peaks at 1.69–1.70 GB under the default budget and under 1000 MB and 300 MB. At 300 MB, dispatch/l rises from about 1.4 GB to 1.56 GB and slows down. The peak is the frames a build materialises and HiGHS's copy of them, and spilling reaches neither.
  • jemalloc dirty_decay_ms and background_thread do not move the peak either.
  • Reads still stream. At l that is right: reading every primal, dual and activity takes 0.065 s streamed and 0.202 s in memory (transport), and 0.139 s against 0.358 s (fleet). At 2xs, reading on the build's engine measured −16% to +5% across cases. It is left for a stacked follow-up.
  • HiGHS's own load (passModel) is the largest single cost of a large build: 2–4 s at l, of which about 2.4 s is a cold first call that allocates its memory. It is outside polars and was not changed. Loading the matrix column-wise is slower than row-wise at every l rung, before the cost of producing the column order.

🤖 Generated with Claude Code

https://claude.ai/code/session_01EpzN6XKNByxsfi3GhX3v3D

claude added 2 commits October 6, 2026 21:25
The streaming engine starts up for every query, and a small build is many
small queries. A build whose largest declaration has fewer than 250,000
coordinates now collects on the in-memory engine; a larger one streams as
before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EpzN6XKNByxsfi3GhX3v3D
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EpzN6XKNByxsfi3GhX3v3D
@read-the-docs-community

read-the-docs-community Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Documentation build overview

📚 specsolve | 🛠️ Build #34980949 | 📁 Comparing 54c7907 against latest (7af455a)

  🔍 Preview build  

3 files changed
± about/architecture/index.html
± about/changelog/index.html
± reference/internals/index.html

@FBumann
FBumann added this pull request to stack #1864 October 6, 2026 21:36
claude added 2 commits October 6, 2026 21:53
…ilt model

engine_for(coordinates) is the decision, on(engine) the block every collect
of the build runs in, and BuiltModel.engine the one place the choice lives.
Also drops bench/.cache, a local symlink committed by mistake.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EpzN6XKNByxsfi3GhX3v3D
…s what it does

streaming_available() is the probe, engine_for() the decision, and
collecting_on() the block a build collects in. BuiltModel.engine and the
Assembly parameter go: nothing read them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EpzN6XKNByxsfi3GhX3v3D
@FBumann
FBumann removed this pull request from stack #1864 October 6, 2026 22:22
@FBumann
FBumann added this pull request to stack #1866 October 6, 2026 22:22
@codspeed

codspeed Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Merging this PR will regress 37 benchmarks

⚡ 76 improved benchmarks
❌ 37 regressed benchmarks
✅ 155 untouched benchmarks
⏩ 106 skipped benchmarks1

Warning

Please fix the performance issues or acknowledge them on CodSpeed.

Performance Changes

Mode Benchmark BASE HEAD Efficiency
❌ Memory test_window[sector-s-specsolve-highs-shape] 4.1 MB 6.1 MB -32.71%
❌ Memory test_window[sector-s-specsolve-highs-coefficient] 4.1 MB 6.1 MB -32.56%
❌ Memory test_window[sector-s-specsolve-highs-cold] 4.1 MB 6.1 MB -32.49%
❌ Memory test_emit[sector-s-specsolve-highs] 4.2 MB 6.1 MB -31.91%
❌ Memory test_window[sector-s-specsolve-highs-one] 4.2 MB 6.1 MB -31.68%
❌ Memory test_window[sector-s-specsolve-highs-values] 4.2 MB 6.1 MB -31.58%
❌ WallTime test_read[commitment-m-specsolve-frames] 5.4 ms 7.4 ms -26.14%
❌ Memory test_emit[sector-s-specsolve-lp] 4.7 MB 6.2 MB -24.52%
❌ WallTime test_emit[sector-m-specsolve-highs] 163.8 ms 212.3 ms -22.83%
❌ WallTime test_read[nodal-l-specsolve-frames] 24.2 ms 31.1 ms -22.29%
❌ WallTime test_read[commitment-l-specsolve-frames] 8.6 ms 10.7 ms -20.17%
❌ WallTime test_read[dispatch-m-specsolve-frames] 5.4 ms 6.7 ms -18.5%
❌ WallTime test_read[sector-l-specsolve-frames] 19.3 ms 23.3 ms -17.33%
❌ WallTime test_window[nodal-m-specsolve-highs-cold] 177.6 ms 213.6 ms -16.86%
❌ WallTime test_window[sector-m-specsolve-highs-one] 235.9 ms 280.4 ms -15.85%
❌ WallTime test_window[sector-m-specsolve-highs-coefficient] 186.6 ms 219.8 ms -15.08%
❌ WallTime test_read[transport-l-specsolve-frames] 39.2 ms 45.6 ms -13.95%
❌ WallTime test_emit[nodal-m-specsolve-lp] 275.2 ms 316.3 ms -12.97%
❌ WallTime test_window[nodal-m-specsolve-highs-one] 328.5 ms 376.8 ms -12.81%
❌ WallTime test_window[sector-m-specsolve-highs-values] 241.4 ms 276.6 ms -12.73%
... ... ... ... ... ...

ℹ️ Only the first 20 benchmarks are displayed. Go to the app to view all benchmarks.

Tip

Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.


Comparing perf/engine-by-size (54c7907) with claude/friendly-dirac-t7g8rd (94340ef)

Open in CodSpeed

Footnotes

  1. 106 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

…to perf/engine-by-size

# Conflicts:
#	CHANGELOG.md
@FBumann
FBumann removed this pull request from stack #1866 October 6, 2026 22:36
@FBumann
FBumann added this pull request to stack #1867 October 6, 2026 22:36
@FBumann FBumann added the trigger:bench Run the benchmark comparison on this PR; removed when the run finishes label Oct 6, 2026
@github-actions github-actions Bot removed the trigger:bench Run the benchmark comparison on this PR; removed when the run finishes label Oct 6, 2026

FBumann commented Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator Author

Note

The following content was generated by AI.

bench was cancelled by its 60-minute limit during the base pass, before this PR's code ran. The base pass (94340ef) took 59 min 51 s, so there is no 0002 result to compare.

Where the hour goes, and a proposed fix
  • In the CI log, 37 of the 60 minutes are one stretch: test_emit at s/m for every arm (22:39:44 → 23:16:26).
  • The same stretch locally, pytest bench --sizes s m --cases dispatch nodal transport --benchmark-memory -k "test_emit and not test_sweep" --durations=20, ran 41 min. pyomo at m is about 32 of them:
    • transport-m-pyomo-highs: 1026 s
    • dispatch-m-pyomo-highs: 432 s
    • dispatch-m-pyomo-lp: 239 s
    • transport-m-pyomo-lp: 231 s
  • The slowest specsolve cell in that stretch is transport-m-specsolve-lp, at 18.5 s with its memory pass.
  • Not this PR's. No specsolve code runs in the pyomo cells. The job runs SELECT twice under timeout-minutes: 60, and one pass nearly fills that. A re-run would time out the same way, so none was started.
  • Proposed fix, not made here.
    • The gate (--fail-on peak:20%) compares memray's peak for every arm. Between the base and head passes, though, only the specsolve arm runs different code.
    • So .github/workflows/bench.yml could pass --arms specsolve in SELECT without losing what the gate can catch. Raising timeout-minutes is the other option.
    • Either is a separate ci: change.

Generated by Claude Code

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants