Repository navigation
Conversation
The streaming engine starts up for every query, and a small build is many small queries. A build whose largest declaration has fewer than 250,000 coordinates now collects on the in-memory engine; a larger one streams as before. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EpzN6XKNByxsfi3GhX3v3D
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EpzN6XKNByxsfi3GhX3v3D
Documentation build overview
3 files changed± about/architecture/index.html± about/changelog/index.html± reference/internals/index.html |
…ilt model engine_for(coordinates) is the decision, on(engine) the block every collect of the build runs in, and BuiltModel.engine the one place the choice lives. Also drops bench/.cache, a local symlink committed by mistake. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EpzN6XKNByxsfi3GhX3v3D
…s what it does streaming_available() is the probe, engine_for() the decision, and collecting_on() the block a build collects in. BuiltModel.engine and the Assembly parameter go: nothing read them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EpzN6XKNByxsfi3GhX3v3D
Merging this PR will regress 37 benchmarks
Warning Please fix the performance issues or acknowledge them on CodSpeed. Performance Changes
Tip Investigate this regression by commenting Comparing Footnotes
|
…to perf/engine-by-size # Conflicts: # CHANGELOG.md
|
Note The following content was generated by AI.
Where the hour goes, and a proposed fix
Generated by Claude Code |
Note
The following content was generated by AI.
A build whose largest declaration has fewer than 250,000 coordinates now collects on polars' in-memory engine. Builds at
2xsare 23–35% faster, and a 16-slice sweep is 4–17% faster. No measured build is slower beyond noise. CodSpeed's heap benchmark putssector/sat 4.1 → 6.1 MB. Its wall-time report flagsmandlcells that a local A/B does not reproduce (see below).Stacked on #1857. Merge that first.
Why, and what changed
engine="auto"resolves to streaming, so polars does not make this choice itself (storage/2xsbuilds in 0.069 s onauto, 0.076 s on streaming, 0.045 s in-memory).relational/collect.py.engine_for(coordinates)is the one decision. It answers in-memory belowIN_MEMORY_BELOW, or where this polars has no streaming engine, and streaming otherwise.collecting_on(engine)is the block everycollected()inside runs in.streaming_available()is the probe that wascollect_engine(). The new name says what it answers: since this PR, it no longer names the engine every collect uses.engine_forof an unknown size, which is streaming as before.Engine.buildruns the assembly insidecollecting_on(engine_for(_largest(...)))._largestis the coordinate count of the largest variable or constraint, unmasked, which is known once the sources are attached. Thatwithclause and that helper are the whole change on the engine side.2xssolve-and-read by noise only (storage: 64 ms none, 52 ms build only, 51 ms everywhere).storage/m(400k coordinates) built 14% slower in memory. Every rung that gained by a clear margin has at most about 120k coordinates, and the mixed results start at 400k. At 250,000, no measured cell regresses in time.relational/collect.pyrow indocs/about/architecture.mdand the "Collect engine" glossary entry say so.bench/.cachesymlink by mistake. A later commit removes it, so it is not in this PR's diff.Wall time: CodSpeed's macro run, and a local A/B of the cells it flags
CodSpeed's wall-time run on
54c7907against94340efreports 37 regressions in all (the 7 memory ones above included) and 76 improvements. Its 13 largest wall-time regressions are all atmorl:sector/mandnodal/mhave 1.2M coordinates in their largest declaration, so they stream with or without this PR, and everytest_readruns the same path.94340ef. That commit'sWall time (macro runner)job was skipped, so CodSpeed compared against a baseline taken elsewhere.The local A/B used
pytest bench --cases sector nodal commitment dispatch transport --sizes 2xs m l --arms specsolve --sinks highs lp --builds 0on #1857's head94340efand on54c7907. As on CodSpeed's runner, it runs every2xscell in the same process beforemandl. Three alternations on an idle 4-core machine; the minimum per cell:test_read[commitment-m-specsolve-frames]test_emit[sector-m-specsolve-highs]test_read[nodal-l-specsolve-frames]test_read[commitment-l-specsolve-frames]test_read[dispatch-m-specsolve-frames]test_read[sector-l-specsolve-frames]test_window[nodal-m-specsolve-highs-cold]test_window[sector-m-specsolve-highs-one]test_window[sector-m-specsolve-highs-coefficient]test_read[transport-l-specsolve-frames]test_emit[nodal-m-specsolve-lp]test_window[nodal-m-specsolve-highs-one]test_window[sector-m-specsolve-highs-values]mandlcells. Locally they move by −13% to +7.6%, in both directions. A run without the2xscells first gave −6.5% to +6.9%.2xscells in the same runs.test_emitis 3–20% faster andtest_window2–18% faster;test_readmoves by −7% to +19%, on a path this PR does not change.Memory: CodSpeed's heap benchmark
CodSpeed's memory run on
54c7907against94340efreports 7 regressions and 5 improvements:sectorats, ontest_emitand everytest_windowvariant: peak heap 4.1–4.7 MB → 6.1–6.2 MB.sector/s's largest declaration has 120,000 coordinates, so it now builds in memory, and the in-memory engine materialises each join's result where streaming takes it in chunks.sectorcrosses a sparse portfolio with dense carriers, so its join results are the largest relative to the model.profiled/sread 7.6 → 6.7 MB,nodal/swindows 4.5 → 4.1 MB,transport/swindows 14.7 → 13.6 MB.--benchmark-memoryuses, counts the system allocator, and polars allocates through its bundled jemalloc. memray reads 96.3 MB onsector/sfor both trees, so it can neither confirm nor refute CodSpeed's figure. How the extra heap grows between 120k coordinates and the threshold is not measured.Benchmarks: build seconds (minimum) / peak RSS (median), fresh process per cell, polars 2.0.0
Method:
bench.arms.specsolve.build_and_emit('highs', …)(build and HiGHS load, no solve), with configurations alternated cell by cell on an idle 4-core machine. The base is #1857's headdf06eb3. These numbers were taken on the PR's first commit. The later restructures choose the same engine for every cell.Threshold at 250k, 5 rounds:
2xs,xsandl, 3 rounds (taken with the threshold at 1M; at these rungs the two thresholds choose the same engine):The
lrungs stream in both columns. Their spread (−10% to +4%) is the noise floor of that run, not an effect of this PR.The workloads small builds serve, best of 7:
solve_over×16, #1857 → this PRNot the
pixi run ladderharness.xland2xlwere not run.Coverage, mutation table, gates
test_every_port_builds_the_same_model_on_either_enginebuilds every port on both engines and compares each hand-off frame (objsorted, since it has no order contract) androw_starts. The comparison is to1e-12, not to the bit. Ontransport_dantzig,transport_pwlandosemosys_utopia, the two engines rounda * b / cdifferently, by at most 7e-15.contentsdigest. A small model's answer saved under feat(deps): specsolve runs on polars 2.0 and requires it #1857 (streamed) can be refused when it is rebuilt here (in memory). The same model always picks the same engine, so a save and a rebuild on one version agree.LAYOUTis not raised: nothing written to disk changes shape, andarchive.pyalready says two polars versions may digest one table differently.373f41c, a substitution on a committed tree, with the full suite run with-x:engine.py: the build runs outsidecollecting_ontest_a_build_collects_on_the_engine_its_largest_declaration_sizes[small]engine.py: the size ignored (engine_for(0))test_a_build_collects_on_the_engine_its_largest_declaration_sizes[at-the-threshold]collect.py:>=becomes>inengine_fortest_the_engine_is_sized_to_the_model[at]collect.py: the streaming probe ignored inengine_fortest_the_probe_sees_the_refusalcollect.py:collecting_onignoredtest_a_named_engine_holds_inside_its_block_only54c7907(the merge ofmainthrough feat(deps): specsolve runs on polars 2.0 and requires it #1857):ruff check .,ruff format --check .andpyrefly(0 errors) are clean.Measured and not done
storage/lpeaks at 1.69–1.70 GB under the default budget and under 1000 MB and 300 MB. At 300 MB,dispatch/lrises from about 1.4 GB to 1.56 GB and slows down. The peak is the frames a build materialises and HiGHS's copy of them, and spilling reaches neither.dirty_decay_msandbackground_threaddo not move the peak either.lthat is right: reading every primal, dual and activity takes 0.065 s streamed and 0.202 s in memory (transport), and 0.139 s against 0.358 s (fleet). At2xs, reading on the build's engine measured −16% to +5% across cases. It is left for a stacked follow-up.passModel) is the largest single cost of a large build: 2–4 s atl, of which about 2.4 s is a cold first call that allocates its memory. It is outside polars and was not changed. Loading the matrix column-wise is slower than row-wise at everylrung, before the cost of producing the column order.🤖 Generated with Claude Code
https://claude.ai/code/session_01EpzN6XKNByxsfi3GhX3v3D