Repository navigation
mcpp test re-does dependency-closure planning on every invocation — 15s to 2.5min of silent work before anything compiles #529
Description
Activity
- added a commit that references this issue
on Aug 30, 2026 Validated 2026.8.30.2 on the workspace the original numbers came from, as the design document requested.
Setup: same machine, same MCPP_HOME, only the mcpp binary swapped (2026.8.28.2 release vs 2026.8.30.2 release, both from GitHub Releases, checksums verified). For each member we first let the new binary do a full rebuild under its own fingerprint (real compilation observed in the output), then measured back-to-back warm
mcpp test -p <member>runs. The baseline was re-measured the same hour with the old binary, same warm state.Warm
mcpp test -p, wall clock:- small leaf member (11 tests): 18.2 s before, 1.1 s after
- our largest-closure member (15 tests; a path dependency pulls in our biggest module-interface set): 109 s before, 2.7 s after
All tests pass under both binaries. The headline symptom we reported is gone on our tree; the remaining one-to-three-second floor is consistent with the items the design document lists as not shipped (no test fast path,
-pdisabling the fast path, per-member replanning under--workspace), and that floor is fine for us.One correction to our original report: we attributed the silent phase to dependency-closure replanning. Your root cause — the post-link verdicts being wiped by the per-invocation regeneration of resolution.json — is the correct mechanism; the scaling we observed matches it just as well. Thanks for the fix. From our side this issue can be closed.
Follow-up with the full-workspace number. We switched the tree's engine to the 2026.8.30.2 release binary (payloads untouched; the one-time full rebuild that follows from the version participating in the build fingerprint took about 35 minutes, as expected). A warm full-workspace
mcpp testacross our 7 members / 108 tests now completes in 24 seconds wall:workspace result ok. 7 member(s); 108 passed; 0 failed; finished in 23.62sUnder 2026.8.28.2 the same run took about 17 minutes and could not finish inside a 10-minute foreground window at all. That closes the whole gap we reported, not just the per-member floor. Nothing further from our side.
Fixed in 2026.8.30.2 (PR #538). Closing on the re-measurement you posted above rather than on ours.
Root cause, recorded here because the shape recurs: both post-link ELF passes were written with a read-back keyed on the artifact's stat, and the record they read lived in
resolution.json— whichprepare_buildrewrites from a fresh json object at the start of every run. Every invocation therefore began by deleting the memo its own backend was about to look for. Within a single invocation the read-back worked, which is why a one-command profile could not see it.What landed:
- The verdicts move to
.mcpp-runtime-verdicts.json, which survives that rewrite;resolution.jsonkeeps publishing a copy. - Pruning now asks whether the artifact left the DISK, not whether it left the current command's plan.
buildandtestshare one output directory with different link-unit sets and were deleting each other's records. - The record's invalidation key gains the SubOS farm stamp and
MCPP_ALLOW_HOST_LIBS. Making a memo durable creates a staleness obligation that did not exist while the answer was recomputed on every run.
Regression test:
tests/e2e/323_post_link_record_durability.sh. It asserts across processes AND across commands, because neither half is visible to a single repeated command, and it asserts the record's CONTENT rather than a wall-clock number — the timing is a machine property, the record is what makes the memo correct.Analysis and measurements:
.agents/docs/2026-08-30-issues-527-529-535-537-analysis-and-design.md.- The verdicts move to
Version: mcpp 2026.8.28.2, llvm@22.1.8, x86_64-linux-gnu, 8 cores.
On a workspace where every target is already built and cached,
mcpp test -p <member>still spends most of its wall time in a single silent phase before test compilation
starts. Nothing is recompiled — the phase produces no artefacts at all.
For our workspace the full test sweep costs ~17 minutes, and this phase is nearly all
of it.
mcpp buildover the same workspace, fully cached, finishes in 0.7s.What the time is spent on
Timestamping each output line of a repeat run (member and test names elided):
Three observations about that 15.6s window:
than it under the member's
target/and under$MCPP_HOME/build-cacheyields5
.jsonfiles and 1.ninjafile. No.pcm, no.o, no executables.ps --ppid <mcpp-pid>every 1.5s returns nothingfor the whole window, and
/proc/<pid>/wchanreads0throughout — mcpp isrunning on-CPU in-process. Neither clang nor ninja is spawned.
gives the same time both times, and interleaving other members does not change it.
How it scales
Not with the number of cached dependency units — with the module-interface surface of
the whole dependency closure, including workspace
pathdependencies (which are notcounted in the
(N units)line):pathdep on a large memberD is the interesting row: same cached-unit count as B, ~10x the time, and the only
difference is a
pathdependency on a member with a few hundred module interfaces.C measured 60s and 133s for the identical command on different runs, which looks like
page-cache sensitivity — consistent with the phase reading a lot of files.
Things that don't help
--profile dev(testdefaults torelease,builddefaults todev)--cache localWhy it matters
It puts a floor on the edit-test loop that is independent of how small the edit was.
For us it is the difference between a 7-minute and a 17-minute verification pass, and
it scales with workspace size rather than with change size, so it gets worse over time.
If the planning result were memoised — keyed on the resolved manifest set and the
toolchain, invalidated when either changes — a warm repeat invocation could skip it
entirely, the same way
mcpp buildalready does.Happy to put together a minimal reproducing workspace (a member with a large
pathdependency seems to be enough to show it) if that would help.