Skip to content

perf(terminal): batch output per flush, route it to its card, and draw on-screen cards with WebGL - #95

Merged
howdeploy merged 61 commits into
howdeploy:mainfrom
BIackFIame:perf/cpu
Sep 28, 2026
Merged

howdeploy merged 61 commits into
howdeploy:mainfrom
BIackFIame:perf/cpu

Conversation

@BIackFIame

Copy link
Copy Markdown
Contributor

Depends on

This branch is based on #94 (refactor/startup-and-dedup @ 47d229a, which merges #89, #90, #92 and #93) and should be merged after it. Its own changes are 3 commits. Each commit covers one topic and passes typecheck and its tests.

Goal

Find what burns CPU in the main process, the renderer and the GPU under terminal load and a real agent session, and fix the clear wins. Commits 1–2 do not change visible behaviour; commit 3 moves on-screen cards from xterm's DOM renderer to its WebGL renderer, within a fixed context budget.

How it was measured

The measurements use the hidden harness from scripts/bench-runtime:

  • windows are show:false, off-screen and unfocusable;
  • keychain calls are refused and safeStorage is off;
  • every run gets a fresh fake HOME, with the tool homes inside it;
  • no credentials are used.

What was recorded for each scenario window:

  • CPU per process kind, from ps;
  • IPC count and size per channel, in both directions;
  • renderer main-thread time, from Performance.getMetrics;
  • xterm DOM render batches;
  • CPU profiles of main (in-process inspector) and renderer (DevTools Profiler);
  • a census of main-process timers.

Scenarios:

  • idle8: 8 terminal cards, no output.
  • flood: 8 cards at 1 MB/s each (flood.mjs).
  • spin: 24 cards, each redrawing an agent-style spinner line every 80 ms plus an output line every ~1.5 s.
  • agent: real Claude Code on a local Ollama qwen3.5:9b, running 3 tool calls (Bash, Read, Write) and one permission prompt, with CanvasTTY's hooks active.

Each scenario was run 5 times on the base and 5 times on the branch, alternating in ABBA order. The tables show medians. The machine (M1 Max, 10 cores) was loaded by another project: 1-minute load average 2.9–9.6, median about 5.9. Treat absolute numbers as indicative; the before/after difference is the meaningful part.

What the profile shows (base)

  • Idle: about 0.6 % CPU for the whole app. One main-process timer in 15 s. No polling.
  • Flood:
    • Main: 16 % CPU, but its JavaScript is busy only about 5 % of the window, and half of that is the harness. App code per chunk is below 1 % of main: flushOutput, the PTY handler and the scrollback ring together take about 30 ms in 11 s.
    • Per chunk, main runs no status regexes and no redaction. Title status is an indexOf-based OSC parser, for claude and qwen only.
    • Renderer: 37 % CPU, almost all xterm. Parsing (print, decode, parse, scroll) and the DOM renderer (replaceChildren, createRow, style and layout) dominate.
  • Agent: main 0.6 %, renderer 3.5 %, GPU 3.1 %. Hook handling costs a few ms over the whole turn.
  • 24 busy cards: two costs are in our code:
    1. Every card subscribed to terminal:data itself, so each batch crossed the context bridge 24 times and was dropped by 23 cards. That is about 7,000 bridge calls a second, roughly 11 % of the renderer's JavaScript.
    2. Every session had its own 16 ms timer, and every batch was its own IPC message: 295 messages and 295 timers a second.
  • xterm renderer: @xterm/addon-webgl was enabled only for the focused card at zoom ≤ 1. Every other card used the DOM renderer, which is the largest renderer cost under output (commit 3).

Changes

  1. perf(renderer): route terminal output to its own card in the preload

    • The preload keeps one IPC listener and a TerminalDataRouter (src/shared/terminalDataRouter.ts).
    • A card subscribes with onData(listener, id) and is called only for its own session.
    • onData without an id still receives everything.
    • Tests: tests/terminal-data-router.test.mjs.
  2. perf(terminal): send one renderer message per output flush for the whole canvas

    • One 16 ms timer in TerminalManager flushes every session with queued output, in the order its output first arrived.
    • TerminalRendererOutbox sends the renderer's share of that flush as a single terminal:data-batch message at the end of the task.
    • Session and removal events first flush the output collected before them, so the renderer sees the manager's order.
    • Observers (agent control, Even G2, plugins) still get one event per session batch.
    • Tests: tests/terminal-output-batching.test.mjs. They check one message per window, one timer for 8 sessions, ordering on dispose, and outbox ordering.
    • docs/ARCHITECTURE.md (and zh-CN) and docs/performance-benchmark.md are updated. The bench byte counter understands the batch channel.
  3. perf(terminal): draw every on-screen card with WebGL from a bounded context pool

    • src/renderer/src/features/terminal/webglContextPool.ts: a renderer-wide pool of 10 WebGL contexts. Chromium keeps at most 16 active WebGL contexts per renderer process and silently loses the oldest past that; 10 leaves room for plugin canvas apps and for released contexts not yet collected.
    • Priority: the focused card, then on-screen area (visible part of the card in the viewport; a busy holder keeps its context against a card less than 1.25× larger), then most recent output. A holder idle for 10 s loses that edge, so a same-size card that is printing takes its slot. Every other card keeps the DOM renderer, as before.
    • No thrash: camera, layout and focus changes settle for 200 ms before contexts move; a fresh grant is kept for 1 s; output on a card without a context re-plans at most once a second and only if a slot could move. Leaving the screen, zoom above 1× or summary mode releases at once.
    • Release disposes the addon and calls WEBGL_lose_context.loseContext() so the slot is really free before GC.
    • Context loss (after xterm's own 3 s restore wait) or a failed attach: the card goes back to DOM with its buffer intact and stays there 30 s, doubling on repeats; the slot goes to the next card.
    • Returning to DOM refits the grid. WebGL snaps the cell width down to whole device pixels (8.0 vs 8.43 CSS px here), so a grid fitted while on WebGL can be too wide for DOM. A DOM-fitted grid always fits WebGL, so attaching needs no fit and swaps do not resize the PTY.
    • WorkspaceCanvas tells the pool about camera, HOME-editing and fullscreen changes; cards report moves, resizes, focus and eligibility.
    • Tests: tests/webgl-context-pool.test.mjs (16): priority, visible area, settle on pan, release on hide/ineligible/unmount, LRU and idle-holder hand-over, hysteresis, minimum hold, the budget is never exceeded, context loss with doubling backoff, no WebGL2 at all, off-screen output never re-plans, stale unregister after a remount.
    • CHANGELOG (en/ru/zh-CN) entry for focused-only WebGL replaced; docs/ARCHITECTURE.md (and zh-CN) describe the pool.

Commits 1–2 do not change the data, offsets or order the card receives, and the card still deduplicates by absolute offset.

Commits 1–2: before / after (medians of 5)

Base is 47d229a, branch is dfd4458. GPU memory was not recorded for this pair; these commits do not touch rendering (the GPU process CPU column is within noise).

Scenario Electron CPU, % of a core main renderer GPU Electron CPU-s per window output IPC msg/s renderer task ms
idle8 (15 s) 0.6 → 0.7 0.3 → 0.3 0.3 → 0.3 0 → 0.1 0.09 → 0.10 0 → 0 17 → 19
flood 8 × 1 MB/s (11 s) 62.4 → 64.5 16.2 → 18.0 36.9 → 36.7 9.3 → 8.8 6.86 → 7.09 155.7 → 39.5 3075 → 3106
spin, 24 cards (15 s) 19.7 → 16.4 (−17 %) 3.2 → 2.5 11.5 → 9.3 5.3 → 4.7 2.96 → 2.46 294.5 → 36.7 1129 → 881 (−22 %)
agent, Claude + Ollama (~70 s) 7.2 → 7.2 0.6 → 0.7 3.5 → 3.5 3.1 → 3.0 5.37 → 5.20 6.7 → 6.7 1038 → 1012

Run spread (Electron CPU %):

Scenario Base Branch
spin 18.6–22.3 14.3–20.7
flood 56.5–67.2 59.2–66.4

Flood and agent differ only within noise. Flood still sends 4× fewer messages, but its cost is xterm parsing and DOM rendering, not message count. All 10 agent turns finished the task.

In the branch's spin profile:

  • the renderer is busy 710 ms instead of 1,039 ms;
  • the per-card fan-out (about 116 ms) is replaced by a 7 ms dispatch;
  • webContents.send in main takes 6 ms instead of 37 ms.

WebGL pool (commit 3): before / after

Base is dfd4458 (commits 1–2), branch is commit 3. Same hidden harness, but the window is sized so every card is on screen (3100×1300 for 8 cards, 3100×2320 for 16 and for 24 cards, of which 16 are visible), device pixel ratio 2. 3 runs per build, alternating (owner limit), medians. 1-minute load average during these runs: 4.9–18.3, median 6.7. GPU memory is the GPU process's footprint (macOS physical footprint) at the end of the window; "IOAccel" is its IOAccelerator (graphics) part. Chromium's own app.getAppMetrics() working set is shown too.

Scenario Electron CPU % main renderer GPU renderer task ms WebGL cards GPU footprint MB GPU IOAccel MB GPU working set MB renderer footprint MB
8 visible, idle (15 s) 0.6 → 0.6 0.2 → 0.2 0.3 → 0.4 0.1 → 0 19 → 19 0 → 8 243 → 348 2 → 15 99 → 110 60 → 65
8 visible, flood 8 × 1 MB/s (11 s) 59.9 → 47.0 (−22 %) 16.9 → 19.9 37.4 → 18.5 6.5 → 8.5 3158 → 1693 0 → 8 416 → 400 25 → 35 104 → 121 268 → 225
16 visible, idle (15 s) 0.7 → 0.9 0.2 → 0.3 0.5 → 0.7 0 → 0 22 → 33 0 → 10 360 → 500 2 → 18 99 → 112 69 → 76
16 visible, 8 flooding at 1 MB/s (11 s) 64.6 → 39.3 (−39 %) 16.5 → 16.1 40.9 → 16.4 7.3 → 6.8 3480 → 1530 0 → 10 (all 8 flooding cards) 587 → 576 27 → 42 104 → 122 288 → 250
24 spinner cards, 16 visible (15 s) 28.8 → 29.7 3.2 → 3.2 16.9 → 14.7 8.9 → 11.2 1722 → 1402 0 → 10 664 → 735 54 → 53 106 → 126 140 → 122

Run spread (Electron CPU %): 8-card flood 58.5–63.4 → 45.6–47.5; 16-card flood 57.5–69.2 → 37.2–46.3; spinners 27.1–30.2 → 26.1–30.1.

  • Heavy output: the renderer halves (DOM style/layout/paint is gone for WebGL cards); main and GPU stay within noise, so the app uses 13–25 points of a core less.
  • Many small redraws (spinners): renderer −2 points, GPU +2 points, total unchanged. WebGL does not help or hurt here.
  • Idle: no CPU change. GPU memory is the cost: about +13 MB of GPU-process footprint per WebGL card at DPR 2 (+105 MB for 8, +140 MB for 10). Under flood the DOM renderer's own compositor layers cost about the same, so the totals meet.
  • The agent-session scenario was not re-run for commit 3 (owner: no Ollama, minimal load). The two runs done before that instruction are not used.

Behaviour checks in the hidden harness (16 visible cards, final build):

  • 10 of 16 cards get WebGL, the rest DOM; no "too many WebGL contexts" warnings.
  • 3 s of continuous wheel panning: 0 renderer swaps during the pan.
  • Panning all cards off screen releases all 10 contexts; panning back re-acquires 10.
  • 3 s of zoom wobbling across 1× (60 button steps): 10 releases at the first step above 1×, no re-acquire until the zoom settles, then 10.
  • Forced context loss on one card: after xterm's 3 s wait the card is on DOM with its text intact and accepts input; another card takes the slot (still 10).
  • 56 renderer swaps in the whole check caused 0 PTY resize messages.

Visual parity (same build with WebGL blocked vs not, offscreen capture of the same card): colours, bold/italic/underline, 256-colour and truecolour backgrounds, URL and OSC 8 link underline, box drawing and CJK look the same. Differences, all xterm's own and already true for the focused card before this change:

  • WebGL cells are 8.0 CSS px wide instead of 8.43 (device-pixel snapping), so text is about 5 % tighter and the right edge of the grid ends earlier (the gap shows the card background);
  • block shade characters (░▒▓) use xterm's WebGL custom glyphs;
  • inside a selection the DOM renderer draws italic text upright and inverse cells differently.
    The cursor could not be compared: the harness window is never focused, so neither renderer draws it.

Risks (commit 3)

  • GPU memory: up to about 130 MB more in the GPU process with 10 WebGL cards at DPR 2. On machines with little VRAM this is the main cost; the budget is one constant (WEBGL_CONTEXT_BUDGET).
  • Mixed renderers: with more than 10 cards on screen, or at zoom above 1×, some cards are DOM and some WebGL, and their text width differs by about 5 %.
  • Chromium's limit is per renderer process: if plugin canvas apps in the same process use many WebGL contexts, Chromium may evict a card's context. The card then falls back to DOM and backs off, as tested.
  • Zooming above 1× still drops WebGL for every card (unchanged rule, now applied to more cards).
  • A grid fitted while a card was on WebGL shrinks by a few columns when it returns to DOM (one PTY resize).

Not in this PR

  • scripts/bench-runtime runs the app with app.getAppPath() pointing at the harness folder, so agent hook helpers are not found there. Flood and pan are unaffected. Linking src/, scripts/ and build/ into the harness folder fixes it, if an agent scenario is added to the bench.

Gates on the branch head (commit 3)

  • npm run typecheck: pass.
  • Full suite (node --test --test-concurrency=2): 1149/1149.
  • npm run build: pass.
  • npm run audit:secrets: pass.
  • npm run test:even: 47/47.

scripts/bench-runtime.mjs runs the built app in hidden windows with a throw-away HOME (idle, N terminal
cards, a fixed-rate output flood in every card, 10 s after it, a canvas pan) and micro-benchmarks of the
main-process hot paths, and prints medians. docs/performance-benchmark.md explains the scenarios.
appendScrollback advanced bufferStart past a dropped chunk but left the chunk in the array until more
than 256 slots were dead, so a busy session kept up to 256 extra chunks alive (823 552 characters instead
of 240 000 with 16 KB reads). The slot is now emptied when the chunk leaves the ring.
…ible

setVisible(true) sent the renderer the whole retained scrollback (up to 240 000 characters) although the
card already had everything up to the hide point and drops what it wrote before. Crossing the 0.5 zoom
threshold did that for every card at once: 8 cards with full history sent 1.9 MB for 8 KB of new output.
The replay now carries the output produced since the card was hidden; a stretch longer than the ring still
arrives as the whole ring, so the truncation marker is unchanged.
…ting every card

AgentControlService used terminals.list() as a lookup: every status, ownership check, observe_agent and
list_agents joined the scrollback of every card on the canvas and threw it away (29 MB allocated per
observe_agent call on 20 cards with full history). It now reads metadata through getMetadata(id) and
listMetadata(); only the observed card's own scrollback is read. The plugin sessions.list answer is built
from listMetadata() as well.
…gins read

observe_agent, get_agent_result and the plugin screen ran every redaction rule over the whole 240 000-
character scrollback (about 12 ms on an M1 Max) to keep its last 8 192 or 4 000 characters.
SecretRedactionRegistry.redactTail masks a window that starts at least 16 384 characters before the tail
(more when a held value could wrap over that), at a line start, and at the header of a private-key block
still open there. A secret split by the tail's cut lies inside the window and is masked whole, as before;
tests compare redactTail with masking the whole text for secrets placed around the tail cut and the window
start.
…rst launch

The first Kimi launch ran spawnSync("kimi --help") on the main process (up to 3 s) to learn whether the CLI
takes a per-run MCP config; nothing else moved meanwhile. ProviderLaunchAdapters.warmKimiProbe runs the
same probe with execFile (same command, environment, limits and answer) 5 s after startup and after a
CLI recheck, and the launch uses that answer. A launch that comes before it still probes as before.
With the companion on, TerminalPresentation made an @xterm/headless screen for every session on its first
event and parsed all of its output, although the glasses read only the sessions shared with them. The
headless screen is now made on the first read, from the scrollback (the same text the stream carried),
and fed from then on; the answer state is still kept for every session.
The camera lives in App state, so each pointer move of a pan renders App and the workspace, and every
TerminalCard rendered again with it: the workspace passed new arrow functions and a rebuilt snap target
list on every render, and nothing was memoized. TerminalCard is now React.memo'd (snap targets compared by
value), and the workspace hands it callbacks that stay the same functions and call its latest handlers
through a ref. A card still renders when its session, zoom, focus, selection or a neighbour's bounds change.
With 8 cards a pan move rendered 8 TerminalCards (97 components); now none (48 components).
…x variant

hermes.png was the 1024 px original (574 KB) and codex.png 544 px (146 KB), drawn at 24 to about 125 CSS
px. Both are scaled down to 128 px plus a 256 px variant chosen through srcset on high-density displays:
574 KB + 146 KB become 12 + 42 KB and 12 + 37 KB. The marks are not otherwise changed; the assets README
says how they were scaled.
electron-vite leaves every bundle unminified, so each window load parsed 1.83 MB of renderer JavaScript.
The renderer build now minifies with esbuild (1.83 MB to 1.04 MB, CSS 149 KB to 126 KB); source maps stay
off and audit:secrets passes on the built output. Main and preload bundles are unchanged.
…tting

Quitting hung every card up and let Electron free the Node environment at
once. A shell that exited during that teardown made node-pty's native exit
watcher call into JavaScript that could no longer run; node-pty rethrew the
failure as a C++ Napi::Error and the app aborted (SIGABRT). The benchmark's
quit after the output flood hit it in most runs, on main as well.

TerminalManager now tracks every PTY until its exit is reported, and
waitForProcessExits() resolves once all have exited: after 2 s it sends
SIGKILL to the rest and waits 1 s more. shutdownServices() starts that wait
right after the terminal shutdown and awaits it before quitting finishes.
The exit handler also never lets an exception reach node-pty's native
callback, where a throw aborts the app too.
…built from them

Each held value became a regular expression with a wrap-gap class between
every two characters. From about 3,700 characters the pattern exceeded what
V8 accepts, and the SyntaxError broke every redaction call (agent observe
and result, plugin screens, failure details) as long as the value was held.

Held values are now found in the text with its wrap characters (whitespace,
box sides) taken out, by Knuth-Morris-Pratt per value, with a position map
back to the original text and the same 64-character gap limit. The cost is
linear in the text and the value, so the limit on a held value rises from
4,096 to 65,536 characters (PEM keys, service-account JSON). redactTail keeps
masking a window at least as wide as the longest possible match, so the tail
is still exactly what masking the whole text and cutting it gives.
Every Claude lifecycle hook ran the helper through Electron-as-Node, about
105 ms per call, and PostToolUse fires after every tool. Claude Code 2.1.281
can POST hooks itself (type "http"), so the gateway now also listens on
127.0.0.1 (random port) and Claude's UserPromptSubmit, PermissionRequest,
PostToolUse, Stop, StopFailure, SessionEnd and Notification hooks go there.
Measured with the real CLI, each hook costs about 1.4 ms instead of 105 ms.

- Claude does not send SessionStart over HTTP, so that hook keeps the helper.
- PreToolUse (base protection and plugin decisions) also stays on
  permission-gate.mjs and the 0600 socket. An HTTP hook that fails in any way
  lets the tool run: refused, 5xx, 401, timeout or a malformed answer. So do
  the sandbox proxy (auto profile), HTTP_PROXY, allowedHttpHookUrls and
  httpHookAllowedEnvVars. A loopback port is also reachable by other local
  users.
- Claude fills the session id and capability headers from the session's
  environment (allowedEnvVars), so the token never enters the --settings argv.
- The listener takes only POST requests with application/json, a loopback
  Host and no Origin, Referer or Sec-Fetch-*. Bodies are bounded to 512 KB like
  the helper's, and the request is checked against the same lease and token.
  Revoking a session revokes its capability.
- ClaudeHttpHookPolicy keeps the helper on Windows, in plugin environments, for
  the auto profile, for Claude versions older than 2.1.281 or not yet known,
  when a proxy is set, and when inline, managed, user or project settings
  enable the sandbox, restrict hook URLs or headers, or set a proxy. A request
  whose capability header arrives empty switches HTTP off for later launches.
- Plugins may no longer set allowedHttpHookUrls or httpHookAllowedEnvVars in
  their Claude settings.
- Lifecycle acks (socket and HTTP) go out before the app reacts. onSignal
  runs on setImmediate in arrival order, and a throw there no longer reaches
  the hook.
BrowserStore chained every save onto writeQueue with no catch. After one
failed write (full disk, permissions, antivirus lock on the temp file) the
queue stayed rejected, so every later save rejected with the same error
until restart. BrowserService awaits persistRuntime() before updating the
view in close/newTab/selectTab/closeTab/navigate, so those actions failed
half-way; did-navigate and popup adoption left unhandled rejections, and
applyBrowserSettings called setRestoreTabs() without awaiting it.

- BrowserStore.persist() continues after an earlier failure (same pattern
  as PluginMediaService.persist) and removes its temp file when the write
  or rename fails. Each save still reports its own error.
- BrowserService.persistRuntime() and a new clearSavedTabs() log a failed
  save and keep the in-memory state, so the tabs on screen are unchanged;
  only the copy restored at the next start is stale.
- index.ts catches the setRestoreTabs() promise.

No other Browser card behaviour changes.

Test: tests/browser-policy-store.test.mjs "BrowserStore recovers after one
failed write ..." (read-only data folder, then writable again) and a source
check that BrowserService callers cannot see the rejection.
The PreToolUse gate for Claude Code, Codex and Qwen Code always exited 0 and
printed nothing when it could not get an answer, so a missing or refused
socket, a hung or broken gateway, or an unreadable answer let a call that
base protection would deny (sudo, writes outside the project) run unchecked.

The launch now sets CANVASTTY_RUNTIME_FAIL_CLOSED=1 in the gate's own hook
command; it installs the gate only when base protection is on or a decision
plugin applies. With the flag, every way the check can fail is the CLI's
deny JSON with "CanvasTTY safety check unavailable: this tool call was not
run. Retry it, or ask the person how to proceed." The gateway marks an ask
that stands for its own failure as unavailable, so Codex and Qwen Code
(which cannot ask) deny it while Claude Code still asks the person. The
OpenCode guard fails closed the same way. "none" is now told apart from an
unreadable answer. Answered calls do the same work as before.
…ection

Base protection returned no deny for many destructive or outside-project
commands because it did not see the command or its target:

- a command after do/then/else/if/while/until/! in the same segment;
- env -i (its -i was taken as a value flag), env -S, busybox/toybox
  applets, script -c CMD and BSD script FILE CMD;
- perl -i / ruby -i (perl is an interpreter, so the sed/perl branch was
  unreachable);
- find -L/-H/-P/-O2/-f before the start folders (which also made
  find -L dir -exec rm {} + inside the project look like rm .);
- cp/mv/install/ln -t DIR, tar -C DIR -x..., --directory=, unzip -o
  (taken for 7z's -oDIR), bundled curl -fsSLo / wget -qO / -qP,
  --output=, --output-dir, and curl/wget side files;
- curl -fsSLo f URL && sh f was not download-and-run.

Wrappers now have their own value-flag tables, and each form is tested
against targets outside (denied like the plain command) and inside the
project (allowed).
…m the main renderer

About half of the handlers in registerIpc.ts checked the sender with
assertMainRenderer(); the rest trusted any sender. Among them were
settings.update, plugin preview/install/set-modules/enable/uninstall,
plugin and provider secrets, plugins.openExternal, terminal
create/restart/input/dispose and media.read. Today only the main window
loads the preload that exposes them, so this is defence in depth, but a
second renderer that gains ipcRenderer (a new window, a preload change)
would reach terminal input and secrets without any check.

- Those handlers now call assertMainRenderer(). terminal.input is a
  fire-and-forget send, so a foreign sender is dropped (isMainRenderer)
  instead of throwing inside the IPC layer.
- Home media: settings keep mediaPath only when it is an absolute path of
  at most 4096 characters, without NUL, with an image extension Home can
  show; anything else keeps the previous value. media.read (new
  homeMedia.ts) follows a symbolic link only while its target stays in
  the folder the file was chosen from, requires the target to be a
  supported image, and reads size and content from one open handle.
  The picker saves the resolved path of the chosen file.

Tests: tests/browser-ipc-security.test.mjs (every listed channel checks
its sender), tests/home-media.test.mjs (link inside the folder is read,
links to another folder or to a non-image are refused; settings reject a
relative path, a non-image path and a crafted settings file).
…md.exe pass

A .cmd/.bat provider (the npm shims for claude, codex, ...) is started as
`cmd /d /s /c "<exe> <args>"`. cmd.exe removes one level of carets when it
reads that line, and the shim then runs `node cli.js %*`, which cmd.exe
parses again. Arguments carried only one level of carets and used \" for
inner quotes, which cmd.exe does not treat as an escape. In the second pass
the quote state flipped at every \", so && and | inside an argument were
outside quotes: the Claude --settings JSON with its hook command
(`set "ELECTRON_RUN_AS_NODE=1" && "...CanvasTTY.exe" ...`) was cut at the
first && and the rest ran as a separate command.

- Arguments are caret-escaped twice (as cross-spawn does for npm shims);
  every quote is escaped at both levels, so cmd.exe never enters a quoted
  region and every operator stays escaped. The executable is unchanged.
- The MSVC quoting step doubled only one of two or more backslashes before
  a quote (lookahead regex), so `a\\"b` lost its quote. Backslash runs are
  now doubled whole.

Test: tests/provider-cli-registry.test.mjs "Windows batch arguments survive
cmd.exe and the shim's %* re-parse unchanged (cmd.exe model)": a model of
%VAR% expansion, caret/quote handling (both passes), operators outside
quotes and MSVC argv splitting; the settings JSON, operators, %PATH%, ^, !,
quotes, trailing and doubled backslashes round-trip exactly, and no
operator is outside quotes in either pass. This is a model; verification
on real Windows is still pending.
…llers

AgentControlGateway kept a receipt for every mutating request (create,
send, interrupt, choose, dismiss) so a retry with the same request id
replays the first answer. Receipts were never removed: after 4096 such
requests, including failed ones, every mutating request answered
LIMIT_REACHED until the app restarted. A refusal that wrote nothing
(BUSY, NOT_READY) was cached too, so retrying that request id replayed
BUSY forever.

- At the cap the oldest finished receipts are dropped; LIMIT_REACHED
  remains only when every kept receipt is still running.
- BUSY, NOT_READY, LIMIT_REACHED, LIFECYCLE_DISABLED and CLOSED are
  refusals before any write; their receipt is removed so the same id is
  performed again. Other results, successful or not, still replay.
- maxReceipts option (default 4096) so the cap can be tested.

Test: tests/agent-control.test.mjs "request receipts are bounded without
locking the gateway, and a refused request can be retried": with a cap of
3, five failed requests do not block a create, a recent receipt still
replays, and a BUSY send is performed on retry with the same id.
…ain process

WindowsPipeHostTransport had no 'error' listener on the host's stdin,
stdout or stderr. A relay write or the shutdown end() after the host
process died raises EPIPE on stdin; with no listener Node throws it as an
uncaught exception in the main process.

stdin and stdout errors now fail the transport the same way a FATAL frame
does (virtual sockets closed with the error, 'fatal' emitted, child
killed); stderr errors are ignored because stderr only feeds diagnostics.
Errors from a previous host process are ignored after a restart.

Test: tests/windows-pipe-host-transport.test.mjs "... host pipe errors into
a transport failure instead of an uncaught exception" (fake host; EPIPE on
each stream). Runs on any platform with platform: "win32" injected.
… failed start

RuntimeGateway and AgentControlGateway did not listen for the pipe host
transport's 'fatal' event (AgentGateway does). After the host died,
RuntimeGateway kept its endpoint and dead transport, so every later agent
launch failed with "must be started" until restart; agent control kept
publishing an endpoint nobody served. A rejected RuntimeGateway start()
also left the dead transport in place.

AgentControlGateway.start() did not clean up when a step after listen()
failed (for example writing the token file): the socket kept listening
and a second start() answered "already started".

- RuntimeGateway: on 'fatal' the transport and its sockets are dropped
  and the host is started again with the same bounded backoff as
  AgentGateway (3 attempts, 500 ms doubling); a failed start closes and
  forgets its transport; close() cancels a pending restart.
- AgentControlGateway: start() is split into openEndpoint /
  writeDiscovery / closeEndpoint. A failure closes the server or
  transport and removes the temporary socket folder, so start() can be
  called again. On 'fatal' the pipe host is restarted the same way and
  connection.json is rewritten with the new endpoint (the token file is
  written once). close() also removes the socket folder.
  connection.json is written to a temp file and renamed, so a controller
  reading it during that republish never sees it empty or half written;
  onTransportRestarted() reports the republished record.

Tests: tests/agent-runtime-gateway.test.mjs (fatal -> restart, launches
refused only while down, no restart after close; failed start drops the
transport) and tests/agent-control.test.mjs (failed start then a
successful start on the same gateway; Windows host restart republishes
the endpoint, awaited through onTransportRestarted; 20 isolated runs
pass). Windows cases use injected platform and a fake transport.
TerminalManager handed environments.wrap() `planned.env.PATH`. planned.env
is a plain-object copy of process.env; on Windows process.env is
case-insensitive but the copy keeps the key as Windows spells it, usually
"Path", so the value was undefined. EnvironmentRegistry.resolveCommand()
then found no bare program name on PATH and refused every wrapper that
answered with one (for example `wsl` or `docker`).

launchSearchPath(env, platform) returns PATH, and on Windows falls back to
the first key that matches PATH case-insensitively. POSIX is unchanged.

Test: tests/session-environments.test.mjs "the environment wrapper gets
the launch's search path even when Windows spells it Path" (injected
platform, plus the call site).
…call

OrchestrationGateway created an AbortController per request and aborted it
on `cancel` or disconnect, but never passed its signal to the handler.
The call ran to completion: a canceled spawn_agent still created and
prompted the subagent, a plugin tool kept the request open, and when the
call finished the orchestrator got the successful result instead of
CANCELED.

- OrchestrationCommandHandler.execute() takes the signal. The gateway
  answers CANCELED as soon as the signal aborts and drops a late result.
- ScopedOrchestrationHandler refuses a call that is already canceled; a
  spawn_agent canceled while its agent was starting closes that agent,
  since nobody will receive its id. Plugin tool calls have no cancel in
  the service protocol, so they are no longer waited for.

Tests: tests/orchestration-gateway.test.mjs "cancel reaches the running
command and the answer is CANCELED, not the late result" and "a
spawn_agent canceled while it was starting closes the agent it created".
… hash

PluginServiceSupervisor hashed the entry with one read and then started
`node <entry>`, which read the file again; an await (mkdir of the data
folder) sat between the two. A file replaced in that window ran as
trusted native code.

The service process now starts with `--import` of a small data: URL boot
that registers module load hooks (module.register). For the entry URL the
hook reads the file itself, checks its SHA-256 against the trusted hash
and returns exactly those bytes as the module source; a mismatch stops the
load and the process exits. The host's own check stays, so a changed file
still fails at once with "changed after it was trusted" and is not
retried. The entry keeps its path: argv[1], __filename, import.meta.url
and the format (ESM or CommonJS) are as before.

Checked by hand with the bundled Electron 43 in ELECTRON_RUN_AS_NODE mode
(.cjs, .js and .mjs entries run with the right hash and stop with a wrong
one); packaged-build fuses do not affect --import.

Tests: tests/plugin-services.test.mjs "an entry swapped after the host
checked it is not run ..." (the started "node" swaps the entry before the
real node reads it; the swapped code must not run) and "a verified entry
still runs from its own location with the guard in place" (ESM and
CommonJS, __filename and argv[1] unchanged).
RuntimeGateway accepts up to 64 sockets and only started a timer after
the first message. A connection that sent nothing stayed open for good,
so a process of the same user could open 64 idle connections and every
lifecycle and decision hook was refused until restart.

A connection now has 5 s (firstMessageTimeoutMs) to send its one message;
hook helpers write it right after connecting. The timer is cleared once
the message arrives, so a decision request still waits for its answer as
before.

OrchestrationGateway was listed with the same problem, but it already
closes an unauthenticated connection on its heartbeat sweep (15 s,
heartbeats need authentication); unchanged. AgentControlGateway has a
10 s per-connection timer; unchanged.

Test: tests/agent-runtime-gateway.test.mjs "RuntimeGateway closes a
connection that sends no message ...".
ProviderRuntimeLaunch.atomicWrite() and TerminalSessionStore.persist()
write a temp file and rename it over the target. When the rename failed
(EPERM from a locked file or antivirus on Windows, a full disk) the temp
file stayed. atomicWrite() names it with a random UUID, so every failed
launch left one more *.tmp next to the provider's settings for good.
TerminalSessionStore also created its folder without a mode.

Both now delete the temp file when the write or rename fails and rethrow;
TerminalSessionStore creates its folder with mode 0700 (an existing
folder is left as it is).

Test: tests/atomic-write-cleanup.test.mjs (a non-empty folder in place of
the target makes the rename fail; no *.tmp remains; a new store folder is
private).
The browser audit log is a hash chain that is verified when the store
opens; any line that does not parse sets integrityError, and every agent
browser mutation then fails with AUDIT_UNAVAILABLE ("retryable") until
someone deletes the log by hand. A crash or a full disk during an append
leaves exactly such a line: half a record with no newline. A failed
append in a running app also left its partial bytes, so the next record
was written onto the same line.

- On open, an active file that does not end with a newline was cut
  during a write: a whole last record only gets its newline back, a
  partial one is removed (with a warning). The chain before it is
  untouched, and anything else that does not verify still fails closed,
  as the tamper test requires.
- A failed append truncates the file back to its size before the write.

Test: tests/browser-audit-store.test.mjs "BrowserAuditStore repairs a
line torn by a crash instead of refusing every later action" (partial
record, then a whole record missing only its newline).
The audit chain hashes a canonical JSON of each record whose keys were
sorted with localeCompare. That order follows the ICU locale of the
process: with LC_ALL=sv_SE "ä" sorts after "z", with en_US before it,
and upper case sorts differently from code-unit order. A log written
under one locale could fail verification under another, which blocks
every agent browser mutation (AUDIT_UNAVAILABLE). The other canonical
JSON implementations in the repo sort by code unit.

- New records sort keys by UTF-16 code unit.
- Verification accepts a record whose hash matches either the code-unit
  or the previous localeCompare order, so existing logs keep verifying
  exactly as before and new records chain onto them. (An old record
  written under a different locale still depends on that locale, as it
  did before; new records no longer do.)
- Rotated audit files are ordered by code unit too (their names are
  ASCII, so the order is unchanged).

Tests: tests/browser-audit-store.test.mjs "the audit hash does not depend
on the system locale" (append under sv_SE, verify and append under
en_US, verify under sv_SE again, in child processes) and "records hashed
with the earlier locale-ordered keys still verify and extend the chain".
…oviders

Two lookups of cmd.exe disagreed. providerCliRegistry (batch provider
launches) takes ComSpec, then %SystemRoot%\System32\cmd.exe.
terminalLaunch (the plain terminal when no PowerShell exists) searched
PATH before SystemRoot, so a cmd.exe in a project folder or any earlier
PATH entry started as the terminal shell.

Both now use windowsCommandPromptPath(): ComSpec, then
%SystemRoot%\System32\cmd.exe, never PATH. pwsh is still found on PATH,
where its installer puts it; that is unchanged.

Test: tests/terminal-launch.test.mjs "the Windows terminal falls back to
the system cmd.exe, never one found on PATH, like provider launches"
(injected platform, environment and file checks).
pollDeviceCode() awaited fetch without a catch. One dropped connection or
the 15 s request timeout (an AbortError from the per-request controller)
left the loop; startDeviceFlow() swallowed the error, so the UI kept
showing the code until it expired and an approval on GitHub was never
picked up.

- A network error, a request timeout, a non-OK response or an unreadable
  body now backs off (interval doubled, at most 60 s, RFC 8628 3.5) and
  polls again until the code's lifetime ends. Cancelling the flow (new
  flow, sign-out) still stops it at once.
- slow_down adds 5 s for this and later polls and honours a longer
  `interval` GitHub sends with it. access_denied, expired_token and the
  flow lifetime end the poll as before.

Test: tests/github-auth.test.mjs "device polling survives network errors
and timeouts, backs off, and honours slow_down" (fake fetch and clock:
fetch failure, timeout, 503, slow_down, pending, then approval).
…the host resolved

The entry guard hashed the module whose URL equalled realpath(spec.entryPath) as the host computed it
before spawning, while the child was started with spec.entryPath and node resolves the main module
itself. An entry, or a folder above it, replaced by a symlink between the two resolutions sent the main
module to another URL, which the load hook passed to nextLoad unchecked: the untrusted module ran.

The hooks now also resolve: the one module resolved without a parent is node's main entry, and its
URL is checked with the same hash as the host's URL, so whatever file node resolves the entry to must
hold the trusted bytes, and those bytes are what runs (CommonJS entries included; node takes the source
the hook returns for them too).

Tests start the service through a "node" that changes the plugin files first and assert the untrusted
module's marker is never written: the entry replaced by a symlink (ES module and CommonJS), the entry's
folder replaced by a symlink, and a CommonJS entry replaced by content.
The probe test compares the blocking and background answers. Both used the
3 s production limit, which a loaded machine can hit while starting the stub
CLI, so the test failed on load rather than on a wrong answer. The limit is
now a parameter (default unchanged) and the test passes 30 s.

Copy link
Copy Markdown
Owner

P2: make activity-triggered WebGL replanning respect the camera settle boundary.

At b276c3f, webglContextPool.ts:166–181 has two independent scheduling paths:

  • viewportChanged() restarts the settle timer.
  • touch() schedules an activity callback that directly calls plan() after 1,000 ms.

Camera movement cannot postpone the second callback. In addition, plan() cancels the pending settle timer. Thus the stated guarantee that renderer changes wait until the camera settles does not hold when output and panning overlap.

A concrete regression scenario using the existing fake-clock harness:

  1. Fill the pool and leave another eligible card visible without a context.
  2. Let a nonfocused holder become idle.
  3. Send output to the waiting card, scheduling the activity replan.
  4. Keep moving the camera every 16 ms for longer than 1,000 ms, without a 200 ms quiet interval.

The activity callback can detach/attach renderers during that pan. This puts context creation/disposal and the DOM fallback refit back onto the interaction path. The current pan test checks movement without this pending activity callback, so it does not cover the combination.

Please make both scheduling paths honor the last camera/layout change, and add a regression asserting that no context swaps occur until the quiet interval has elapsed even when activity replanning is pending. Immediate release for loss of eligibility should remain intact.

This is a static-review finding; I have not run the harness or measured visible stutter.

Also, this branch inherits the two remaining curl parser findings from #92: #92 (comment). Please carry those fixes through #94 into this branch.

@BIackFIame

Copy link
Copy Markdown
Contributor Author

Added 782b6b3: during a wheel/trackpad pan the scene is now composited (will-change for the length of the gesture), as it already was for pointer drag and zoom; before, a wheel pan repainted every card on every frame. 16 cards, medians of 3 alternating runs: whole-app CPU 36.7% → 31.7%, GPU process 11.8% → 7.4%, paint 143 → 27 ms/s. Idle, output flood and busy-card scenarios unchanged. Zoom already did not re-rasterize cards during the gesture. Also measured and rejected (no gain beyond noise): raw bytes over a MessagePort, PTY backpressure, a 30 fps cap for unfocused cards, removing the per-card scrollbar layer, Electron switches.

…-output-dir alone write nothing

curl resets its per-transfer options at --next (-:, also inside a short
cluster), so --output-dir of one operation does not move the -o / -O files
of another. classifyFetch collected every output of the line and applied the
last --output-dir to all of them, so
`curl --output-dir ../out -o a URL --next --output-dir . -o b URL` was read as
two writes inside the project. Each operation is now resolved with its own
outputs, -O count, --remote-name-all and --output-dir, and the facts are
combined.

--output-dir without -o, -O or --remote-name-all sends the response to
standard output; it no longer counts as a write to that folder. Side files
(cookie jar, header dump, trace, ...) and real outputs are still judged.
Startup awaited the startup page (a data: URL, 105 to 165 ms on the
audit machine) before it started any service, and services took another
60 ms after it. The page load now runs next to service startup, and the
application surface replaces the page as soon as the IPC handlers exist.

A close during the page load is still a quiet quit. The application
surface aborting a page that is still loading is expected and ignored; a
real page load error on a live window is still reported as a failed
startup once services are up.
The main bundle imported every dependency when it started, including
those only a few paths need. Measured in Electron on this machine:
electron-updater 26 to 31 ms (packaged update checks only), yaml 9 to 12
ms (Hermes configs), @xterm/headless 7 to 8 ms (agent control screens and
the Even G2 view), secure-remote-password 2.5 ms (Even G2 pairing).

They now load through `lazyRequire` on first use; the load stays
synchronous, so no caller changes shape. The two Electron smoke runners
(test code behind env flags, about 1,200 lines) are dynamic imports and
become their own chunks. The main bundle import in the hidden startup
harness drops from about 88 ms to about 27 ms.

A test keeps static imports of these modules out of src/main.
The first limits read started `codex app-server` and it stayed up until
the app quit: about 58 MB, plus the git processes it starts, for the
whole session, while the renderer polled every 60 s even with its window
hidden.

- The app-server is stopped after 3 minutes without a Codex limits read
  and started again by the next read (the first read after that is a cold
  start, as on launch; the snapshot cache still covers 60 s).
- The renderer reads limits only while its document is visible and
  reads at once when it becomes visible again.

With the window visible nothing changes: the widget reads every 60 s and
the app-server stays up. Even G2 and plugin reads go through the same
service and start it on demand.
…te in it

Resolving the provider CLIs checks each provider's commands in every
search directory: 478 statSync calls on the main thread at startup (and
on every "Recheck"), most of them for per-provider install directories
that do not exist.

Each directory is now checked once per resolution, and a candidate in a
directory that is missing (or not a directory) is recorded as "missing"
without its own stat, the same answer that stat gives. Candidates in
existing directories are checked exactly as before. The child PATH
filter reuses the same per-resolution answers.

Hidden startup harness, empty profile: 478 -> 267 statSync, about
8 -> 5 ms on the main thread.
Twelve places cut newline-delimited JSON by hand, each with its own
limit handling: the two gateway decoders (512KB and 128KB, near byte for
byte copies), the runtime and agent control gateways, the Codex limits
client, the plugin service supervisor, the provider smoke RPC, the
runtime client and permission gate, and both MCP helpers (socket and
stdin).

`src/agent-runtime/ndjson.mjs` is now the only implementation. It ships
next to the helpers, so the helpers and the main process use the same
file. `NdjsonLineReader` returns complete lines and never buffers more
than `maxLineBytes` of one line; a longer line either throws (the
default) or is reported once through `onOversize` and dropped up to its
newline. `NdjsonDecoderBase` adds JSON parsing with the caller's error
factories; the two gateway decoders keep their names, limits and error
codes on top of it.

Every site keeps its limit. Two readers had no bound at all and now use
the protocol's own limit: the orchestration helper's gateway socket (a
longer line closes the connection; the gateway never sends one) and its
stdin (a longer request gets the same "exceeds 128KB" error as the
browser helper's). Readers that took string chunks now take bytes, so a
UTF-8 character split across chunks is decoded whole.
The four local gateways (agent browser, orchestration, agent runtime,
agent control) each hashed and compared capability tokens and created,
published and removed their Unix socket endpoints on their own, with
small differences: two comparisons had no length guard, one compared raw
tokens, the listen and close helpers existed twice and inline twice.

A single skeleton for all four is not worth the risk: their protocols,
lease models and Windows pipe host recovery differ. What they share now
lives in `gatewaySocket.ts`:

- `tokenDigest` / `tokenMatches`: SHA-256 of the token, constant-time
  compare with a length guard, a missing digest never matches, the
  presented digest is wiped. Agent control now keeps the digest of its
  token and compares digests like the other three.
- `makePrivateDirectory`, `listenOnEndpoint` (listen, then 0600 on the
  socket file), `closeServer`, `removeEndpoint`, and the 100-byte Unix
  socket path limit.

Endpoint names, fallback directories and each gateway's error tolerance
on removal are unchanged.
There were three: the browser catalog's (strict, but it rebuilt objects,
so integer-like keys came out in numeric order), the orchestration
catalog's (no checks at all; an undefined property came out as the bare
word `undefined`, which is not JSON), and the browser audit's
`stableJson`.

`canonicalStringify` in tool-catalog.mjs is now the only one; the
orchestration catalog re-exports it. Keys are sorted by UTF-16 code unit,
strictly as strings. Strict by default, as the browser bridge already
was: cycles, non-finite numbers, non-plain objects and undefined array
entries throw, undefined properties are left out. The browser audit uses
`lenient`, which answers those as JSON.stringify would, and verifies
records written before howdeploy#93 with `compareKeys: localeCompare`, so its
hash chain is unchanged.

A test replays 500 random JSON values through the previous audit
serializer in both key orders and requires the same text.
Nine places decided whether a path lies inside a root, in five styles:
`startsWith(root + sep)` (plugin media, the Even G2 web root),
`relative()` with `startsWith("..")` (browser policy, session trust,
command review, base protection), and `relative()` with an exact ".."
test (plugin assets, Home media, the plugin hook runner).

`isPathInside(root, candidate, { allowRoot })` in
src/agent-runtime/path-inside.mjs replaces all of them; it ships with the
hook runner, which uses it too. The rules are the strictest correct ones:

- both paths are resolved, so `..` segments count;
- a sibling sharing a prefix (`/root2`) is outside;
- a child named `..cache` is inside (the `startsWith("..")` variants
  refused it);
- a relation on another drive or share is outside;
- Windows compares case-insensitively, as its path functions do;
  elsewhere the comparison is exact, so a differently cased path on a
  case-insensitive macOS volume counts as outside unless it went through
  fs.promises.realpath.

Links are still resolved by the callers exactly where they were before.
Table tests cover POSIX, macOS and Windows paths, links out of and into
the root, and a case-insensitive volume on the real file system.
Four regexes decided which names hold secrets and had drifted apart:
the browser audit log (exact key match: `passcode`, `apiKey` or
`x-refresh-token` went into the log), what agents read back from the
browser, and the screenshot field mask twice (a TypeScript check and a
copy inside the page script that had to be kept in sync by hand).

`safety/sensitiveNames.ts` now holds one credential list with two
extensions: keys (cookies, auth headers, web storage) and form fields
(one-time codes, anything auth). The page script is built from the same
source string.

What agents read and which fields screenshots mask are unchanged (a test
compares both with the previous lists name by name). The audit log now
redacts any key containing a sensitive name, and a `name=value` string
with one, instead of its shorter exact list.
The full provider list was typed out by hand in eleven places (settings
defaults and validation, session store, runtime gateway, plugin manifest
checks, plugin IPC, radial menu, plugin frame, renderer defaults and the
radial launcher list), and the limit provider list in four. Adding a
provider meant finding all of them.

- `isProviderId` in providerCatalog.ts checks against the catalog's
  labels (typed as a Record over every id, so it cannot miss one).
- `AGENT_PROVIDERS` and the new `LIMIT_PROVIDERS` in contracts.ts are
  the only lists; `RADIAL_LAUNCHER_ITEMS` is the launcher list plus its
  three actions.
- Settings migrations keep their frozen historical lists (legacy,
  pre-Qwen, added providers): those describe old profiles, not the
  current catalog.

Order and contents of every list are unchanged; a test ties them to the
catalog.
…overlays

The Kimi overlay (agent-browser/ProviderLaunch.ts) was a renamed copy of
the Hermes one (hermesConfig.ts): the same cross-process lock with stale
owner recovery, compare-and-swap writes, atomic replacement, backups,
hashes and file helpers, about 300 lines each. The lifecycle hook
overlays (ProviderRuntimeLaunch.ts) had a third atomic write, backup
restore and file helpers.

configOverlay.ts now holds them once. Hermes and Kimi call it with their
label, so every error message reads as before; the Kimi test seams
(beforeReclaim, beforeRelease) stay. The recovery journals, their fields
and paths, and the lock file format are unchanged, so an overlay left by
an older build is still recovered.

Small differences that went away: new directories made by an atomic
write are chmod'ed to 0700 on every path (only the lifecycle overlays did
this), and the lifecycle overlays now also re-apply the file mode after
the rename, as Hermes and Kimi did.
Seven TerminalManager test files each defined their own
`availableRegistry()` and `fakeSpawner()`, the same code with small
drifts (pid base, what `write` records, whether answers are frozen).

tests/helpers/terminal.mjs has one of each. The registry answers are
frozen like the real registry's; the spawner takes the pid base and a
write observer as options and returns a PTY that can emit data and exit.
knip (with tests and scripts as entry points) found one function that
nothing calls, `byteLengthOfCanonicalJson` in tool-catalog.mjs and its
declaration, and 46 exported values that are used only inside their own
module: timeouts and limits, snap and region constants, command review
helpers, the provider CLI definitions, and three names the agent-browser
barrel re-exported for nobody. They are module-private now.

After this change knip reports no unused value export in src/. Unused
exported types (147, many of them option interfaces) are left alone.
Every card subscribed to terminal:data on its own, so the preload handed each
output batch to every card across the context bridge and all but one dropped
it by id. With 24 cards printing a spinner that was about 7 000 bridge calls a
second for 300 batches. The preload now keeps one IPC listener and a per-session
router; a card subscribes with its session id and is called only for its own
output. onData without an id still receives everything.
…ole canvas

Each session had its own 16 ms batch timer and every batch was its own IPC
message, so a canvas of busy cards cost one timer and one message per card per
window (24 cards with a spinner: ~300 messages and timers a second). One timer
now flushes every session with queued output, in the order its output first
arrived, and TerminalRendererOutbox sends the renderer's share of that flush as
a single terminal:data-batch message at the end of the task. Session and
removal events flush what was collected before them first, so the renderer
sees the manager's order. Observers still get one event per session batch.
Measured: 24 cards, 295 -> ~40 messages a second.
…ontext pool

Only the focused card used xterm's WebGL renderer; every other visible card
paid for the DOM renderer's style, layout and paint on each row it changed,
which was most of the renderer's CPU under heavy output. Cards now ask a
renderer-wide pool (webglContextPool.ts) for a context. It hands out 10,
below Chromium's 16 active WebGL contexts per renderer process: the focused
card first, then by on-screen area (a busy holder keeps its context against
a card less than 1.25x larger), then by most recent output. Cards without a
context keep the DOM renderer.

Camera, layout and focus changes settle for 200 ms before contexts move, and
a fresh grant is kept for 1 s, so panning and wheel zoom do not rebuild
contexts. A card that leaves the screen, zooms above 1x or enters summary
mode releases its context at once and loses it explicitly, so the slot is
free before GC. A lost context (or a failed attach) puts the card back on DOM
with its buffer intact for 30 s, doubling on repeats. Going back to DOM
refits the grid: WebGL snaps cells down to whole device pixels, so a grid
fitted on WebGL can be too wide for DOM; a DOM-fitted grid always fits
WebGL, so swaps do not resize the PTY.

Measured (hidden harness, medians of 3): 8 visible cards at 8 x 1 MB/s,
Electron CPU 59.9% -> 47.0%, renderer 37.4% -> 18.5%; 16 visible with 8
flooding, 64.6% -> 39.3%. Idle and 24 spinner cards unchanged. GPU process
footprint at idle +105 MB (8 cards) / +140 MB (16 cards).
The scene carried will-change only during a pointer drag and a zoom. A
wheel or trackpad pan, the usual way to move around on a Mac, translated an
uncomposited scene, so Chromium repainted every card on every frame. The
wheel pan now marks its own gesture (same 160 ms trailing settle as a zoom)
and the scene is composited while it lasts. A pan does not change the scale,
so the cached raster stays sharp; once the wheel is silent the hint drops and
the resting scene is uncomposited as before. The settle logic for both
gestures moved to gestureSettle.ts.

Measured (hidden harness, 16 visible cards, 5 s of wheel pan at 60 Hz,
medians of 3): renderer Paint 143 -> 27 ms/s, raster 9.6 -> 2.6 ms/s,
Electron CPU 36.7% -> 31.7%, GPU process 11.8% -> 7.4%. Idle, 8 x 1 MB/s
output and 24 spinner cards unchanged.
…o settle

The pool had two independent timers: viewportChanged() restarted the
settle timer, while touch() on a waiting card scheduled a plan() 1000 ms
later that nothing could postpone. That plan() also cancelled the pending
settle timer. When output and a pan overlapped, contexts were created and
disposed (and DOM-fallback cards refitted) in the middle of the pan.

Every timer now re-plans through one gate: until WEBGL_SETTLE_MS has passed
since the last viewportChanged() it reschedules itself to that boundary
instead of planning. plan() keeps the pending settle timer while the camera
is still settling. Losing eligibility still releases the context at once.
@BIackFIame

Copy link
Copy Markdown
Contributor Author

Thanks. The finding reproduced with the fake-clock harness and is fixed in one new commit. This branch also carries the two #92 curl fixes through the rebuilt #94 base.

Final head: aab91b565d4140ca09bc3c48bf42eb5e63ca9466.

The branch was replayed onto the new #94 head 8accbd1e43f76ff251cd8a76ec9dbb8b00f11465. Its own commits are unchanged apart from their SHAs, and the diff of the four earlier commits is identical to before:

  • 73dcffe (was 7784fb7)
  • c3bfee8 (was dfd4458)
  • 7f50b8b (was b276c3f)
  • d409204 (was 782b6b3)
  • aab91b5 (new)

P2: activity-triggered replanning ignored the camera settle boundary

Commit: aab91b565d4140ca09bc3c48bf42eb5e63ca9466

Root cause. WebglContextPool had two independent timers. viewportChanged() restarted the settle timer, while touch() on a waiting card scheduled a callback that called plan() directly after WEBGL_ACTIVITY_REPLAN_MS. Camera movement could not postpone that callback, and plan() also cancelled the pending settle timer. When output and a pan overlapped, the idle holder was detached and the waiting card attached in the middle of the pan. Your scenario reproduced exactly: at the old head the log showed -b +c about 1 s into a pan with a camera move every 16 ms.

Fix.

  • The pool records the time of the last viewportChanged().
  • Every timer now re-plans through one gate: the settle timer, the activity replan, the pin-expiry look, and the backoff end. Until WEBGL_SETTLE_MS has passed since the last camera or layout change, the gate reschedules itself to that boundary instead of planning.
  • plan() no longer cancels the pending settle timer while the camera is still settling, so the plan for the camera's resting position still runs.
  • update() with eligible: false still releases the context immediately, without waiting.

Tests

tests/webgl-context-pool.test.mjs:

  • "output replanning waits for the camera too: no context moves mid-pan, the swap comes once it is quiet": this is your regression.
    • Setup: budget 2, a focused holder, an idle non-focused holder b and a waiting card c. Output reaches c and schedules the activity replan.
    • During the pan: the camera moves every 16 ms for more than 1,500 ms, with c printing on every frame. The log records no attach or detach during the pan, and still none 1 ms before the quiet interval ends.
    • After the pan: once the quiet interval ends, the log is exactly -b +c.
  • "an activity replan with the camera still also defers to a later move, and loss of eligibility stays immediate":
    • A single camera move 50 ms before the activity replan is due pushes the swap to the end of the settle time.
    • A card losing eligibility mid-pan releases its context immediately.

Both tests fail at the old head and pass now. The earlier pan, idle-yield, pin, backoff and budget tests are unchanged and pass.

Typecheck and npm run build pass at aab91b5. The full suite passed 1154 of 1155 tests. The one failure was a temp-folder cleanup race (ENOTEMPTY on rmdir) in tests/session-environments.test.mjs. That file is unchanged from #94, where it passed, and it passed 3 of 3 reruns on its own.

@howdeploy
howdeploy merged commit 20f6855 into howdeploy:main Sep 28, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants