Skip to content

mcpp-language-server 0.0.10: the module cache is bounded, visible, and sweepable without a restart - #39

Merged
Sunrisepeak merged 31 commits into
mainfrom
cache-growth-root-fix
Oct 3, 2026
Merged

Sunrisepeak merged 31 commits into
mainfrom
cache-growth-root-fix

Conversation

@Sunrisepeak

Copy link
Copy Markdown
Owner

Plan: .agents/docs/2026-10-02-cache-growth-root-fix-plan.md (C-7…C-13, X-6) · 5 review rounds recorded in the document

The problem

clangd's copy-on-read BMI copies die with the process that holds them, its only GC waits three days
and reads atime. One real machine (Windows 11, clangd 23.1.0, GalTranslPP) reached 64.36 GiB in
26.5 hours
: 97.1% copies, up to 127 in one command directory, and three orphaned instance
directories holding 94.5%. mcppls owned the module locks (RD12) but nothing else in that tree.

What this does

Bounded (C-7, C-8, C-9, X-6)

  • Every clangd generation start sweeps the versioned copies a dead generation left — background,
    bounded by the new generation's start so nothing it writes can be touched, and every published
    <module>.pcm kept. One file copy is what re-reading costs, not a rebuild.
  • orchestrator::cache is the one implementation: sweep, budget (mcppls.cache.maxBytes 4G per
    workspace, cache.totalBytes 16G global), classified report. Copies and dead instance
    directories go first; what still does not fit is reported, never taken from the published BMIs.
  • Every instance writes instance.json with a heartbeat into the directory it works in; the lease
    tick renames dead directories aside and the background removes them. A live guest is protected by
    its own heartbeat even when the owner's lease is gone.
  • Windows finally asks the process table: process_identity/process_alive/cpu_seconds have real
    answers there, so a crash-and-restart within the lease expiry no longer spawns a second instance.

Visible (C-13)

  • One status item: the cache is a segment under a character budget (mcppls.statusBar.maxLength,
    default 36). Hover = read-only card; click = the hub, a QuickPick in five groups (cache / sweep /
    maintenance / logs / open source) with a codicon per entry and one primary action,
    Sweep the Module Cache, with an eye button for the dry run. No webview, no sidebar — the E2E
    locks webviewPanelCount at zero.
  • mcppls.sweepCache (S3 5.8) removes what no engine holds without stopping one; mcppls cache
    reports classified numbers (canonical/copies/instances/trash) and --prune grew
    --dry-run, --older-than, --max-size. The old "each module may be stored twice" hint is gone.
  • A read-only prompt for a local agent (mcppls cache --prompt agent|issue or one click from the
    hub) carries the environment facts, the commands to look at and the hypotheses this case taught.
    Nothing is uploaded, nowhere, by anyone.

Record

  • Spec five-piece together: S3 4 field + §5.7 + §5.8 (18 new rules), cache-budget fixture,
    traceability.json — validate.py: 0 failures, 289 rules traced.
  • WA-CLANGD-011 registered; Upstream defects register (clangd, mcpp) #24 gets the UP-24 comment (upstream left unfiled, mcppls
    works around it deliberately).
  • Version 0.0.10 via version --set; CHANGELOG, docs (en + zh), design.md RD18–RD20.

Prove it

  • mcpp test 34/34 (incl. new test_cache.cpp: copy predicate + same-directory guard, sweep
    bounds, budgets, instance heartbeat/grace/rename, prompts; test_process.cpp X-6).
  • mcpp build --target x86_64-windows-gnu green (the Win32 process tables compile).
  • Extension: tsc clean, 101 unit tests, E2E added (cacheHub.test.ts).
  • mcpp run -p devtools -- check all green.

Task map (plan §9)

Commit Plan tasks
feat(platform) T1 (X-6)
fix(cache) T2, T3 (C-7, C-8, C-10, C-11)
feat(cache) T4, T5, T7 (C-9, C-13.1, spec)
feat(editors) T6, T8 (C-13.2–13.5, WA-011, version, docs)

owner_gone() could never answer on windows -- process_alive() and cpu_seconds() returned nullopt
there, so a server that crashed and was restarted within the lease expiry always read as a second
instance and started cold in a private directory. The process tables are now asked directly
(OpenProcess/GetProcessTimes for identity and cpu time, the exit code for liveness), a process_self()
names this process for the lease, and macOS reads the kernel's birth time instead of shelling out to
ps for it. The system headers stay confined to process.cpp; fs also learns a failure-visible remove
and a touched-file reading of the file clock, which the cache sweep needs next.
…cache stays under a budget

clangd's copy-on-read copies die with the process that holds them: 26.5 hours of crash loop grew a
cache to 64.36 GiB, 97.1% copies, up to 127 in one command directory. mcppls owns
<cdb>/.cache/clangd the way it already owns its module locks (RD12): before clangd starts -- at
server start and at every in-session restart -- a background sweep removes the files whose name
carries the versioned stamp AND whose canonical BMI sits beside them, keeping everything younger
than the new generation's start, so a copy the starting clangd is about to write is never touched.

The sweep, the report and the budget share one predicate and one implementation
(orchestrator::cache): sweep_copies for the copies, enforce_budget for mcppls.cache.maxBytes
(4G a workspace) and cache.totalBytes (16G over all), report for the classified numbers. The CLI's
cache command renders from the same functions: the report is classified (canonical / copies /
instances / trash) and the 'each module may be stored twice' guess is gone, prune grew --dry-run,
--older-than, --max-size and an instances-only report, and --prompt prints the read-only
troubleshooting prompt the server renders too. engine.clangd tells its host when a generation
starts (cache_sweep_due) and when the live one did (generation_started_at), which the interactive
sweep's bound is built on.
…the wire

Every instance writes instance.json with a heartbeat into the directory it works in: the owner beside
the lease, a guest inside instances/<token>/. The lease tick renews it for guests too, and its cheap
half reaps what died -- stat a few files, rename the dead directories aside, remove them in the
background, so a 60 GiB leftover costs the tick nothing. A live guest is protected by its own
heartbeat even when the owner's lease is gone, which the old prune read the wrong way; the premise
is now liveness, not the presence of a lease file. owner_gone() reads the identity the platform now
provides, and the server records its pid on windows too.

The cache answers for itself: cxxModules/status carries an optional cache field at a 100 MB grain
(S3-4-29/30), cxxModules/cache answers the classified report read-only from a 30 s cache recomputed
after a sweep (S3-5.7), and mcppls.sweepCache removes what no engine holds -- no engine stopped, no
published BMI, one sweep at a time, a dry run over exactly the set it would have removed (S3-5.8).
The settings registry grows the bytes and count kinds and five rows (cache.maxBytes, cache.totalBytes,
cache.instanceGrace, cache.showInStatusBar, statusBar.maxLength); docs regenerate with it. Spec,
schema fixtures and traceability go together: the cache-budget fixture drives a real server through
status, report, dry-run and sweep and asserts the engine generation never changed.
…it without a restart

The status bar's one C++ Modules item carries the cache as a segment under a character budget
(mcppls.statusBar.maxLength, default 36, an icon counted as two): hover opens a read-only card -- the
size against the budget, a text bar of the four classes, the oldest copy, the last sweep -- and click
opens the cache hub, a QuickPick in five groups (cache, sweep, maintenance, logs, open source) with a
codicon on every entry and one primary action, Sweep the Module Cache, with an eye button for the dry
run. No webview, no sidebar: the E2E locks webviewPanelCount to zero. Off stays one click away, as
before.

The hub's open-source group copies a read-only troubleshooting prompt (environment facts, the commands
to look at, the hypotheses this case taught, what never to do) and opens a prefilled issue through the
issueUrl machinery, with a feedback variant that carries no crash code. The prompts are the server's
own rendering, fetched from cxxModules/cache and written to the clipboard verbatim. WA-CLANGD-011
registers the workaround and its premise; #24 carries the upstream record (UP-24). Version 0.0.10,
CHANGELOG and the docs say what is now true.
…l-linux can move

a fresh resolve today picks openkal-musl 0.20.1, which asks for openkal(-linux) 0.16.1 -- more than
the vendored 0.15.1 declares -- and the build refuses to start. the pin keeps the graph where the
vendored fixes live (the comment says when to unpin); CI's clean checkouts are the ones that
re-resolve, which is why local builds never saw it.
a plain "0.15.1" is caret: a fresh resolve today floats openkal-llvm-runtime to 0.15.4, which pulls
openkal-musl 0.20.1, which asks for openkal(-linux) 0.16.1 -- two releases above the vendored 0.15.1,
whose termux/PRoot fixes upstream has not shipped yet (vendor/README.md). Both pins are exact now, so
a clean checkout resolves the graph this repository is tested with; the comment says when to unpin.
…ependency pin

the openkal set is pinned exactly now (the root mcpp.toml says why), and the mcpp CI builds with did
not read the '=' syntax yet -- it floated past the pin and failed the resolve. the local builds that
found the pin and the CI builds that did not now run the same mcpp.
…builds members now

mcpp 2026.9.30 builds a member (mcpp build -p devtools) into the root's
target/<triple>/<fingerprint>/bin -- the member's own tools/devtools/target layout is gone, and the
find that looked there failed the job on an empty tree. the smoke steps, the repository invariants,
the cross-build and the release checks all look in the one place now.
… sdk

the cross build to aarch64-macos does not search the apple sdk's headers (openkal-musl's own
adapter is the interface), so sys/sysctl.h was not there to include. ps(1) is the tool every macos
host has that names a process by pid: lstart is the kernel's birth time, stable for one incarnation
and different for the next owner of the same pid. also removes the duplicated windows include
block the move had left behind, which failed every build with an unterminated conditional.
…rs on the real target

the first cut guarded its Win32 calls with _WIN32 and included windows.h -- both wrong on this
build's windows target, and wrong in ways a linux build cannot see: the target reaches its C library
through openkal-musl and compiles as x86_64-pc-cygwin with __CYGWIN__ undefined, so _WIN32 is not
defined at all (mcpp's own __MCPP_TARGET_WINDOWS__ is the marker) and no vendor SDK is on the
search path. everything compiled out and the identity answered nullopt -- the CI windows run was
the first machine to execute it.

the calls are now declared in process.cpp the way openkal-windows' win32.h declares its own --
exactly the seven used, dllimport and stdcall -- and the three names openkal-windows' generated
kernel32 does not carry (GetProcessTimes, OpenProcess, GetCurrentProcessId) come from this
package's own import library: port/winprocess.def, three names still bound to KERNEL32.dll, turned
into libwinprocess.a by build.mcpp with the same llvm-dlltool, under its own name so it cannot
shadow the forty-five-name kernel32 that is already on the link line. liveness asks a zero-timeout
WaitForSingleObject rather than the exit code: 259 is what a live process reports AND what a
process that exits with code 259 reports, and the wait answers the question asked.

verified under wine on the cross build (identity stable, self alive, a ghost pid not resurrected)
and on the linux and macos builds, which are unchanged.
…ool is found on every host

the versions invariant (devtools check) requires a member to take a third-party dependency through
[workspace.dependencies] with .workspace = true -- the root's direct openkal-musl pin was the one
row that did not, and the macos run of the workspace tests said so. both exact pins now sit where
the rule says, with the why.

the windows host found no llvm-dlltool: MCPP_TOOLCHAIN_DIR is not set there and PATH has none, so
the import library step fell back to the bare name and failed. the running mcpp names its toolchain
directory (mcpp::toolchain_dir(), which CI's pinned 2026.9.30.2 has; the environment variable is
still asked first, for the older tools), and the candidates carry the .exe spelling beside the
plain one.
…g rules, not against them

cmd.exe /C eats a command line that begins with a quoted path, and a '/' inside a mixed-separator
path reads as a switch -- both measured on the windows runner, which answered 'The filename,
directory name, or volume label syntax is incorrect' to the call that had worked on every unix
host. the separators are normalised to the host's own and the command goes in as one literally
quoted block (/S), which is the documented way to say 'these quotes are mine'.
…swers without a modal

the loop's event names were a table indexed by EventKind's own value, eight rows for eight kinds;
the ninth kind (cache_swept) read past it and built a string from whatever sat there -- a garbage
string_view of two megabytes and a segfault the moment the startup sweep's event reached the loop.
That crash was the E2E's initialize failure: every suite needing the server failed on it, while the
unit tests never did (nothing drove a real workspace through the loop). The table now has a row per
enum member with a comment saying why the two must move together.

The sweep command's answer also moved off showInformationMessage: the numbers that prove it are the
status bar's own and the hub's receipt, and a modal to an unasked question is exactly the
unsolicited UI the E2E holds to zero -- it is window progress now, with a quiet status-bar word
only when there was nothing to free. The E2E's status-bar assertion reads what barFor actually
writes (one 'mcppls' segment), not the off state's wording.

The full VSIX-form E2E and the conflicts scenario pass locally against the fixed server.
…ilds with -p can see it

mcpp 2026.9.30 reads a --features request against the selected member's own declaration, and the
CLion plugin's packaging step selects devtools alone (-p devtools) -- the root's clion = {} was not
reached from there, and the release check answered 'no selected workspace member declares'. The
declaration now lives in both places, with the why.
…xture that proves it runs

the CLI's two report paths brace-initialised the numbers object -- nlohmann's initializer-list
constructor wraps a returned Json in an ARRAY, and every .value() on it then aborted the command
with a 306 type error the moment any workspace existed. The unit test had caught the same shape in
its own copy and was fixed; these two hid behind a stale binary in the local run. Assignment
initialises the object it already is.

The cache-budget fixture's own assertions asked for a canonical BMI count this scenario cannot
have -- its modules are primed into module-hints and clangd writes no .pcm of its own here -- and
the fixture was not in ci.yml's per-platform lists at all, so the green CI had never run it (the
lists name every fixture; a new one joins them or joins nothing). The expectations now assert what
a healthy run of THIS scenario is: the report classifies, copies and instance directories stay at
zero across a dry run and a real sweep, and the engine is ready and answering definitions after
both. Verified green locally against the fixed server.
…aken over

the ux kill-server stage measured it: a server killed and restarted within the lease expiry
started cold in a private directory -- 222 s where 25 was the budget -- because the killed
process had not been reaped, its /proc entry still carried its start time, and the identity read
'the same live process' through the fields alone. process_alive already refused zombies by their
state letter; the identity now does too, for the same reason: 'is this the process I knew' has no
answer for a dead one. The regression test names its zombie (a shell that ignores SIGCHLD), waits
for it to exit, and asserts both the identity and the liveness refuse it.
…getpid()

the identity change read /proc/<getpid()>/stat, and under mcppls's own sandbox getpid() answers
for the sandbox -- it says 1 -- so every server's lease carried the same identity (pid 1, the
boot's start time). a killed server and its restart then compared as 'the same live process' both
ways: the zombie fix below covers the unreaped case, and this one covers the reaped case that
still failed, because pid 1 is always alive and always has that start time. the ux kill-server
stage measured the whole story twice: 222 s where 25 was allowed, identical to the digit.

/proc/self/stat names the real process inside the sandbox and out (0.0.9's own lease read it that
way, which is why this is a regression of the refactor rather than an old defect); the zombie
refusal and the named-pid path keep their own comments. verified end to end: a killed server's
lease is taken over by its restart, which writes its own pid and works in the workspace's cache.
…at cannot lose its drill-down, and a task book the agent can act on

The card re-reads as one glance: project (state dot, module and unit counts,
the real preparation progress), cache (a table, one bar a class), actions
(sweep, logs & reports, self-check) above the repository link that replaces
the footnote. The hub re-groups into overview/clean/diagnostics/feedback and
every entry now carries its behavior as data -- the v1 drill-down was dead
code because dispatch parsed icon text -- with a back button in drill-downs
and a full repaint after a sweep. The self-check prompt becomes a task book:
facts, checks with what healthy looks like, an output contract, and a bug
branch where, once the developer agrees, the agent drafts the issue itself,
shows it for approval, and never uploads anything.

The editor now speaks the display language: package.nls* for the manifest,
l10n/ bundles through a strings seam for the runtime, English source and a
zh-cn translation held to key parity by unit tests; the server's logs, CLI
and prompts stay English. cxxModules/cache gains paths.bundlesDirectory
(S3-5.7-6) so a client reveals a directory the report named, never one it
guessed; the prompt structure is S3-5.7-7. openDocumentation was referenced
by the hub but never registered -- registered.
…e hub first

Two defects the first live hover found. (1) The card's rows were joined
with single newlines, which markdown folds into one paragraph -- the card
read as a single run-on line (and part of why v1 read as 'messy'). Zones
are now blank-line separated blocks, rows inside a zone hard-broken, and
the table stands as its own block. (2) The table and bars render only from
the cxxModules/cache detail, which until now only the hub fetched: a hover
on a fresh window showed the coarse fallback with no chart at all. The
status controller now fetches the detail itself, at most once per 30 s (the
server's own report cache window) while a cache is on the status, and
re-applies the card when it lands; the coarse fallback gains a budget bar
so even it has one chart. The reset-offer tooltip keeps precedence over
the async repaint.
…whole cache zone, and honest button names

Color: hover markdown strips style attributes, but emoji squares are colored
plain text -- bars are now fixed-width runs of them, one color a class
(published blue, copies orange, instances purple, trash brown; the budget
green, yellow near it, red over it; preparation green). Every bar is eight
cells, so the chart column aligns by construction, and each row's own name
is its legend -- color is never the only carrier.

Alignment: the big number became the table's bold total row against the
budget, so the whole cache zone is ONE grid and there is no separate
headline block to drift out of line with it; the coarse fallback renders
the same shape (total row + budget bar) so the card never changes form.

Naming: 'self-check' read as the tool checking itself. The actions now say
what they do, verb first: sweep cache / open logs & reports / copy agent
prompt on the card, 'Copy the Agent Troubleshooting Prompt' in the hub,
'Copy the Agent Prompt for Cache Troubleshooting' in the palette, in both
languages.
…lly colour

The emoji squares read as cheap; the character bars the first review liked
cannot be coloured (hover markdown strips every style attribute). Markdown
does carry images, so the bar is now a self-drawn SVG data URI: the same
dot-matrix language, twelve rounded cells, filled cells in the class colour
and empty cells a 20% tint of it as the track -- real colour, muted palette,
nothing fetched (the URI is the image). The alt text is the old block-
character run, so a renderer that refuses the image degrades to the v2.1
bar instead of losing the chart.
…; reveal opens the directory itself

The emoji squares and the SVG strips both lost the review: the bar is again
a fixed-width monochrome run of block characters in a code span -- the
design the first live look liked, colour left out entirely (the state dot's
shape and the percent column carry the tiers). The project's counts and
source moved onto the title line, so the top of the card is one line, the
name clamped to a budget that keeps the whole line unwrapped.

The layout rules are now stated and held: every bar the same width; all
four class rows always present, zero or not; a fact that does not exist
drops its own part without reshuffling the rest; the coarse fallback
renders the same table shape as the full card.

And the directories: revealFileInOS SELECTS the item in its parent (on
Linux, xdg-open on the parent), so 'logs & reports' opened the cache
root's parent -- read, correctly, as the wrong directory. A directory is
now opened with vscode.env.openExternal, which opens the folder itself.
…l's width

The joined action row was the widest line in the card, stretching the
panel past the grid and leaving the table ragged-left against it -- and
its width changed with the language. One action a line, stacked under the
table like a menu: the table is the widest block in every language, the
panel keeps one width, the card reads as one centered column.
…it it

Stacking was a workaround; the line itself was the problem -- it joined
three full sentences. The card is a glance surface: Sweep / Logs /
Agent prompt (清理 / 日志 / Agent 提示词), about thirty columns with the
icons, under the table's width in either language. The full names stay
where there is room to read them: the hub and the command palette.
…gives none

Markdown hovers cannot centre, space or align: VS Code strips every style
attribute, and five rounds of markdown reshaping could not buy the
composition the review kept asking for. The card body is now one
self-drawn SVG (a data URI, nothing fetched): a centred header line
(dot, project, state, counts, source), one grid with fixed column x
positions and right-anchored numbers, twelve-cell dot-matrix bars all
starting on one edge, a divider, a bold total row against the budget, a
centred sweep line -- theme-aware ink (dark/light), redrawn when the
theme flips. The interactive part stays real markdown under the image:
the three short action links and the repository line; the image's alt is
the one-line facts, so a renderer that refuses images degrades to text.
Generation is sub-millisecond string work and rendering happens only
while the hover is open; nothing is fetched, nothing leaves the machine.
…ines align

The drawn image is gone (asked back); the card is markdown end to end
again -- the v2.4 shape: one title line, the one table with monochrome
dot-matrix bars and the bold total row, the sweep line. What the review
wanted fixed is the FOOTER: its two lines (the actions, the repository)
now align with each other -- the narrower one is padded at the front with
no-break spaces, half the measured width difference (icons count two
columns, CJK counts two, latin one), so the pair reads as one centred
block under the table in either language. Plain spaces cannot do this:
markdown folds runs of them; U+00A0 does not fold and cannot start a
code block.
No run of padding spaces: the three action names are chosen so the line
lands beside the repository line's ~49 columns in English AND in Chinese
-- 'Sweep cache / Open logs / Copy agent prompt' (52) and '清理缓存 /
打开日志 / 复制 Agent 提示词' (48) against the URL's 49. The no-break-space
pass stays as the trim: it closes the one to three columns that remain,
invisibly. 'Logs' the KEY stays -- the hub's directory drill-down still
uses it.
…e column's own text

The table was the one block out of step with the aligned footer, ~9
columns wider than it needed to be: 'Total / budget' repeated what the
value cell beside it already says (43.0 MB / 4.29 GB). 'Total' (合计)
brings the grid to the footer's width band in both languages, and the
whole card now reads as one column.
…ames where bundles land

Two small defects the deep review found. (A) A workspace cache reset left
the remembered cxxModules/cache answer riding the hover card: the card
showed the deleted cache's table until the next fetch window. The reset
now voids the remembered detail. (B) The issue prompt rendered (unknown)
for its bundle and report paths -- the facts never carried them. The
attach line now names the bundles directory, which the facts have had
since paths.bundlesDirectory landed.
@Sunrisepeak
Sunrisepeak merged commit eaf9ea4 into main Oct 3, 2026
123 of 124 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant