perf(terminal): batch output per flush, route it to its card, and draw on-screen cards with WebGL - #95
Conversation
scripts/bench-runtime.mjs runs the built app in hidden windows with a throw-away HOME (idle, N terminal cards, a fixed-rate output flood in every card, 10 s after it, a canvas pan) and micro-benchmarks of the main-process hot paths, and prints medians. docs/performance-benchmark.md explains the scenarios.
appendScrollback advanced bufferStart past a dropped chunk but left the chunk in the array until more than 256 slots were dead, so a busy session kept up to 256 extra chunks alive (823 552 characters instead of 240 000 with 16 KB reads). The slot is now emptied when the chunk leaves the ring.
…ible setVisible(true) sent the renderer the whole retained scrollback (up to 240 000 characters) although the card already had everything up to the hide point and drops what it wrote before. Crossing the 0.5 zoom threshold did that for every card at once: 8 cards with full history sent 1.9 MB for 8 KB of new output. The replay now carries the output produced since the card was hidden; a stretch longer than the ring still arrives as the whole ring, so the truncation marker is unchanged.
…ting every card AgentControlService used terminals.list() as a lookup: every status, ownership check, observe_agent and list_agents joined the scrollback of every card on the canvas and threw it away (29 MB allocated per observe_agent call on 20 cards with full history). It now reads metadata through getMetadata(id) and listMetadata(); only the observed card's own scrollback is read. The plugin sessions.list answer is built from listMetadata() as well.
…gins read observe_agent, get_agent_result and the plugin screen ran every redaction rule over the whole 240 000- character scrollback (about 12 ms on an M1 Max) to keep its last 8 192 or 4 000 characters. SecretRedactionRegistry.redactTail masks a window that starts at least 16 384 characters before the tail (more when a held value could wrap over that), at a line start, and at the header of a private-key block still open there. A secret split by the tail's cut lies inside the window and is masked whole, as before; tests compare redactTail with masking the whole text for secrets placed around the tail cut and the window start.
…rst launch
The first Kimi launch ran spawnSync("kimi --help") on the main process (up to 3 s) to learn whether the CLI
takes a per-run MCP config; nothing else moved meanwhile. ProviderLaunchAdapters.warmKimiProbe runs the
same probe with execFile (same command, environment, limits and answer) 5 s after startup and after a
CLI recheck, and the launch uses that answer. A launch that comes before it still probes as before.
With the companion on, TerminalPresentation made an @xterm/headless screen for every session on its first event and parsed all of its output, although the glasses read only the sessions shared with them. The headless screen is now made on the first read, from the scrollback (the same text the stream carried), and fed from then on; the answer state is still kept for every session.
The camera lives in App state, so each pointer move of a pan renders App and the workspace, and every TerminalCard rendered again with it: the workspace passed new arrow functions and a rebuilt snap target list on every render, and nothing was memoized. TerminalCard is now React.memo'd (snap targets compared by value), and the workspace hands it callbacks that stay the same functions and call its latest handlers through a ref. A card still renders when its session, zoom, focus, selection or a neighbour's bounds change. With 8 cards a pan move rendered 8 TerminalCards (97 components); now none (48 components).
…x variant hermes.png was the 1024 px original (574 KB) and codex.png 544 px (146 KB), drawn at 24 to about 125 CSS px. Both are scaled down to 128 px plus a 256 px variant chosen through srcset on high-density displays: 574 KB + 146 KB become 12 + 42 KB and 12 + 37 KB. The marks are not otherwise changed; the assets README says how they were scaled.
electron-vite leaves every bundle unminified, so each window load parsed 1.83 MB of renderer JavaScript. The renderer build now minifies with esbuild (1.83 MB to 1.04 MB, CSS 149 KB to 126 KB); source maps stay off and audit:secrets passes on the built output. Main and preload bundles are unchanged.
…tting Quitting hung every card up and let Electron free the Node environment at once. A shell that exited during that teardown made node-pty's native exit watcher call into JavaScript that could no longer run; node-pty rethrew the failure as a C++ Napi::Error and the app aborted (SIGABRT). The benchmark's quit after the output flood hit it in most runs, on main as well. TerminalManager now tracks every PTY until its exit is reported, and waitForProcessExits() resolves once all have exited: after 2 s it sends SIGKILL to the rest and waits 1 s more. shutdownServices() starts that wait right after the terminal shutdown and awaits it before quitting finishes. The exit handler also never lets an exception reach node-pty's native callback, where a throw aborts the app too.
…built from them Each held value became a regular expression with a wrap-gap class between every two characters. From about 3,700 characters the pattern exceeded what V8 accepts, and the SyntaxError broke every redaction call (agent observe and result, plugin screens, failure details) as long as the value was held. Held values are now found in the text with its wrap characters (whitespace, box sides) taken out, by Knuth-Morris-Pratt per value, with a position map back to the original text and the same 64-character gap limit. The cost is linear in the text and the value, so the limit on a held value rises from 4,096 to 65,536 characters (PEM keys, service-account JSON). redactTail keeps masking a window at least as wide as the longest possible match, so the tail is still exactly what masking the whole text and cutting it gives.
Every Claude lifecycle hook ran the helper through Electron-as-Node, about 105 ms per call, and PostToolUse fires after every tool. Claude Code 2.1.281 can POST hooks itself (type "http"), so the gateway now also listens on 127.0.0.1 (random port) and Claude's UserPromptSubmit, PermissionRequest, PostToolUse, Stop, StopFailure, SessionEnd and Notification hooks go there. Measured with the real CLI, each hook costs about 1.4 ms instead of 105 ms. - Claude does not send SessionStart over HTTP, so that hook keeps the helper. - PreToolUse (base protection and plugin decisions) also stays on permission-gate.mjs and the 0600 socket. An HTTP hook that fails in any way lets the tool run: refused, 5xx, 401, timeout or a malformed answer. So do the sandbox proxy (auto profile), HTTP_PROXY, allowedHttpHookUrls and httpHookAllowedEnvVars. A loopback port is also reachable by other local users. - Claude fills the session id and capability headers from the session's environment (allowedEnvVars), so the token never enters the --settings argv. - The listener takes only POST requests with application/json, a loopback Host and no Origin, Referer or Sec-Fetch-*. Bodies are bounded to 512 KB like the helper's, and the request is checked against the same lease and token. Revoking a session revokes its capability. - ClaudeHttpHookPolicy keeps the helper on Windows, in plugin environments, for the auto profile, for Claude versions older than 2.1.281 or not yet known, when a proxy is set, and when inline, managed, user or project settings enable the sandbox, restrict hook URLs or headers, or set a proxy. A request whose capability header arrives empty switches HTTP off for later launches. - Plugins may no longer set allowedHttpHookUrls or httpHookAllowedEnvVars in their Claude settings. - Lifecycle acks (socket and HTTP) go out before the app reacts. onSignal runs on setImmediate in arrival order, and a throw there no longer reaches the hook.
BrowserStore chained every save onto writeQueue with no catch. After one failed write (full disk, permissions, antivirus lock on the temp file) the queue stayed rejected, so every later save rejected with the same error until restart. BrowserService awaits persistRuntime() before updating the view in close/newTab/selectTab/closeTab/navigate, so those actions failed half-way; did-navigate and popup adoption left unhandled rejections, and applyBrowserSettings called setRestoreTabs() without awaiting it. - BrowserStore.persist() continues after an earlier failure (same pattern as PluginMediaService.persist) and removes its temp file when the write or rename fails. Each save still reports its own error. - BrowserService.persistRuntime() and a new clearSavedTabs() log a failed save and keep the in-memory state, so the tabs on screen are unchanged; only the copy restored at the next start is stale. - index.ts catches the setRestoreTabs() promise. No other Browser card behaviour changes. Test: tests/browser-policy-store.test.mjs "BrowserStore recovers after one failed write ..." (read-only data folder, then writable again) and a source check that BrowserService callers cannot see the rejection.
The PreToolUse gate for Claude Code, Codex and Qwen Code always exited 0 and printed nothing when it could not get an answer, so a missing or refused socket, a hung or broken gateway, or an unreadable answer let a call that base protection would deny (sudo, writes outside the project) run unchecked. The launch now sets CANVASTTY_RUNTIME_FAIL_CLOSED=1 in the gate's own hook command; it installs the gate only when base protection is on or a decision plugin applies. With the flag, every way the check can fail is the CLI's deny JSON with "CanvasTTY safety check unavailable: this tool call was not run. Retry it, or ask the person how to proceed." The gateway marks an ask that stands for its own failure as unavailable, so Codex and Qwen Code (which cannot ask) deny it while Claude Code still asks the person. The OpenCode guard fails closed the same way. "none" is now told apart from an unreadable answer. Answered calls do the same work as before.
…ection
Base protection returned no deny for many destructive or outside-project
commands because it did not see the command or its target:
- a command after do/then/else/if/while/until/! in the same segment;
- env -i (its -i was taken as a value flag), env -S, busybox/toybox
applets, script -c CMD and BSD script FILE CMD;
- perl -i / ruby -i (perl is an interpreter, so the sed/perl branch was
unreachable);
- find -L/-H/-P/-O2/-f before the start folders (which also made
find -L dir -exec rm {} + inside the project look like rm .);
- cp/mv/install/ln -t DIR, tar -C DIR -x..., --directory=, unzip -o
(taken for 7z's -oDIR), bundled curl -fsSLo / wget -qO / -qP,
--output=, --output-dir, and curl/wget side files;
- curl -fsSLo f URL && sh f was not download-and-run.
Wrappers now have their own value-flag tables, and each form is tested
against targets outside (denied like the plain command) and inside the
project (allowed).
…m the main renderer About half of the handlers in registerIpc.ts checked the sender with assertMainRenderer(); the rest trusted any sender. Among them were settings.update, plugin preview/install/set-modules/enable/uninstall, plugin and provider secrets, plugins.openExternal, terminal create/restart/input/dispose and media.read. Today only the main window loads the preload that exposes them, so this is defence in depth, but a second renderer that gains ipcRenderer (a new window, a preload change) would reach terminal input and secrets without any check. - Those handlers now call assertMainRenderer(). terminal.input is a fire-and-forget send, so a foreign sender is dropped (isMainRenderer) instead of throwing inside the IPC layer. - Home media: settings keep mediaPath only when it is an absolute path of at most 4096 characters, without NUL, with an image extension Home can show; anything else keeps the previous value. media.read (new homeMedia.ts) follows a symbolic link only while its target stays in the folder the file was chosen from, requires the target to be a supported image, and reads size and content from one open handle. The picker saves the resolved path of the chosen file. Tests: tests/browser-ipc-security.test.mjs (every listed channel checks its sender), tests/home-media.test.mjs (link inside the folder is read, links to another folder or to a non-image are refused; settings reject a relative path, a non-image path and a crafted settings file).
…md.exe pass A .cmd/.bat provider (the npm shims for claude, codex, ...) is started as `cmd /d /s /c "<exe> <args>"`. cmd.exe removes one level of carets when it reads that line, and the shim then runs `node cli.js %*`, which cmd.exe parses again. Arguments carried only one level of carets and used \" for inner quotes, which cmd.exe does not treat as an escape. In the second pass the quote state flipped at every \", so && and | inside an argument were outside quotes: the Claude --settings JSON with its hook command (`set "ELECTRON_RUN_AS_NODE=1" && "...CanvasTTY.exe" ...`) was cut at the first && and the rest ran as a separate command. - Arguments are caret-escaped twice (as cross-spawn does for npm shims); every quote is escaped at both levels, so cmd.exe never enters a quoted region and every operator stays escaped. The executable is unchanged. - The MSVC quoting step doubled only one of two or more backslashes before a quote (lookahead regex), so `a\\"b` lost its quote. Backslash runs are now doubled whole. Test: tests/provider-cli-registry.test.mjs "Windows batch arguments survive cmd.exe and the shim's %* re-parse unchanged (cmd.exe model)": a model of %VAR% expansion, caret/quote handling (both passes), operators outside quotes and MSVC argv splitting; the settings JSON, operators, %PATH%, ^, !, quotes, trailing and doubled backslashes round-trip exactly, and no operator is outside quotes in either pass. This is a model; verification on real Windows is still pending.
…llers AgentControlGateway kept a receipt for every mutating request (create, send, interrupt, choose, dismiss) so a retry with the same request id replays the first answer. Receipts were never removed: after 4096 such requests, including failed ones, every mutating request answered LIMIT_REACHED until the app restarted. A refusal that wrote nothing (BUSY, NOT_READY) was cached too, so retrying that request id replayed BUSY forever. - At the cap the oldest finished receipts are dropped; LIMIT_REACHED remains only when every kept receipt is still running. - BUSY, NOT_READY, LIMIT_REACHED, LIFECYCLE_DISABLED and CLOSED are refusals before any write; their receipt is removed so the same id is performed again. Other results, successful or not, still replay. - maxReceipts option (default 4096) so the cap can be tested. Test: tests/agent-control.test.mjs "request receipts are bounded without locking the gateway, and a refused request can be retried": with a cap of 3, five failed requests do not block a create, a recent receipt still replays, and a BUSY send is performed on retry with the same id.
…ain process WindowsPipeHostTransport had no 'error' listener on the host's stdin, stdout or stderr. A relay write or the shutdown end() after the host process died raises EPIPE on stdin; with no listener Node throws it as an uncaught exception in the main process. stdin and stdout errors now fail the transport the same way a FATAL frame does (virtual sockets closed with the error, 'fatal' emitted, child killed); stderr errors are ignored because stderr only feeds diagnostics. Errors from a previous host process are ignored after a restart. Test: tests/windows-pipe-host-transport.test.mjs "... host pipe errors into a transport failure instead of an uncaught exception" (fake host; EPIPE on each stream). Runs on any platform with platform: "win32" injected.
… failed start RuntimeGateway and AgentControlGateway did not listen for the pipe host transport's 'fatal' event (AgentGateway does). After the host died, RuntimeGateway kept its endpoint and dead transport, so every later agent launch failed with "must be started" until restart; agent control kept publishing an endpoint nobody served. A rejected RuntimeGateway start() also left the dead transport in place. AgentControlGateway.start() did not clean up when a step after listen() failed (for example writing the token file): the socket kept listening and a second start() answered "already started". - RuntimeGateway: on 'fatal' the transport and its sockets are dropped and the host is started again with the same bounded backoff as AgentGateway (3 attempts, 500 ms doubling); a failed start closes and forgets its transport; close() cancels a pending restart. - AgentControlGateway: start() is split into openEndpoint / writeDiscovery / closeEndpoint. A failure closes the server or transport and removes the temporary socket folder, so start() can be called again. On 'fatal' the pipe host is restarted the same way and connection.json is rewritten with the new endpoint (the token file is written once). close() also removes the socket folder. connection.json is written to a temp file and renamed, so a controller reading it during that republish never sees it empty or half written; onTransportRestarted() reports the republished record. Tests: tests/agent-runtime-gateway.test.mjs (fatal -> restart, launches refused only while down, no restart after close; failed start drops the transport) and tests/agent-control.test.mjs (failed start then a successful start on the same gateway; Windows host restart republishes the endpoint, awaited through onTransportRestarted; 20 isolated runs pass). Windows cases use injected platform and a fake transport.
TerminalManager handed environments.wrap() `planned.env.PATH`. planned.env is a plain-object copy of process.env; on Windows process.env is case-insensitive but the copy keeps the key as Windows spells it, usually "Path", so the value was undefined. EnvironmentRegistry.resolveCommand() then found no bare program name on PATH and refused every wrapper that answered with one (for example `wsl` or `docker`). launchSearchPath(env, platform) returns PATH, and on Windows falls back to the first key that matches PATH case-insensitively. POSIX is unchanged. Test: tests/session-environments.test.mjs "the environment wrapper gets the launch's search path even when Windows spells it Path" (injected platform, plus the call site).
…call OrchestrationGateway created an AbortController per request and aborted it on `cancel` or disconnect, but never passed its signal to the handler. The call ran to completion: a canceled spawn_agent still created and prompted the subagent, a plugin tool kept the request open, and when the call finished the orchestrator got the successful result instead of CANCELED. - OrchestrationCommandHandler.execute() takes the signal. The gateway answers CANCELED as soon as the signal aborts and drops a late result. - ScopedOrchestrationHandler refuses a call that is already canceled; a spawn_agent canceled while its agent was starting closes that agent, since nobody will receive its id. Plugin tool calls have no cancel in the service protocol, so they are no longer waited for. Tests: tests/orchestration-gateway.test.mjs "cancel reaches the running command and the answer is CANCELED, not the late result" and "a spawn_agent canceled while it was starting closes the agent it created".
… hash PluginServiceSupervisor hashed the entry with one read and then started `node <entry>`, which read the file again; an await (mkdir of the data folder) sat between the two. A file replaced in that window ran as trusted native code. The service process now starts with `--import` of a small data: URL boot that registers module load hooks (module.register). For the entry URL the hook reads the file itself, checks its SHA-256 against the trusted hash and returns exactly those bytes as the module source; a mismatch stops the load and the process exits. The host's own check stays, so a changed file still fails at once with "changed after it was trusted" and is not retried. The entry keeps its path: argv[1], __filename, import.meta.url and the format (ESM or CommonJS) are as before. Checked by hand with the bundled Electron 43 in ELECTRON_RUN_AS_NODE mode (.cjs, .js and .mjs entries run with the right hash and stop with a wrong one); packaged-build fuses do not affect --import. Tests: tests/plugin-services.test.mjs "an entry swapped after the host checked it is not run ..." (the started "node" swaps the entry before the real node reads it; the swapped code must not run) and "a verified entry still runs from its own location with the guard in place" (ESM and CommonJS, __filename and argv[1] unchanged).
RuntimeGateway accepts up to 64 sockets and only started a timer after the first message. A connection that sent nothing stayed open for good, so a process of the same user could open 64 idle connections and every lifecycle and decision hook was refused until restart. A connection now has 5 s (firstMessageTimeoutMs) to send its one message; hook helpers write it right after connecting. The timer is cleared once the message arrives, so a decision request still waits for its answer as before. OrchestrationGateway was listed with the same problem, but it already closes an unauthenticated connection on its heartbeat sweep (15 s, heartbeats need authentication); unchanged. AgentControlGateway has a 10 s per-connection timer; unchanged. Test: tests/agent-runtime-gateway.test.mjs "RuntimeGateway closes a connection that sends no message ...".
ProviderRuntimeLaunch.atomicWrite() and TerminalSessionStore.persist() write a temp file and rename it over the target. When the rename failed (EPERM from a locked file or antivirus on Windows, a full disk) the temp file stayed. atomicWrite() names it with a random UUID, so every failed launch left one more *.tmp next to the provider's settings for good. TerminalSessionStore also created its folder without a mode. Both now delete the temp file when the write or rename fails and rethrow; TerminalSessionStore creates its folder with mode 0700 (an existing folder is left as it is). Test: tests/atomic-write-cleanup.test.mjs (a non-empty folder in place of the target makes the rename fail; no *.tmp remains; a new store folder is private).
The browser audit log is a hash chain that is verified when the store
opens; any line that does not parse sets integrityError, and every agent
browser mutation then fails with AUDIT_UNAVAILABLE ("retryable") until
someone deletes the log by hand. A crash or a full disk during an append
leaves exactly such a line: half a record with no newline. A failed
append in a running app also left its partial bytes, so the next record
was written onto the same line.
- On open, an active file that does not end with a newline was cut
during a write: a whole last record only gets its newline back, a
partial one is removed (with a warning). The chain before it is
untouched, and anything else that does not verify still fails closed,
as the tamper test requires.
- A failed append truncates the file back to its size before the write.
Test: tests/browser-audit-store.test.mjs "BrowserAuditStore repairs a
line torn by a crash instead of refusing every later action" (partial
record, then a whole record missing only its newline).
The audit chain hashes a canonical JSON of each record whose keys were sorted with localeCompare. That order follows the ICU locale of the process: with LC_ALL=sv_SE "ä" sorts after "z", with en_US before it, and upper case sorts differently from code-unit order. A log written under one locale could fail verification under another, which blocks every agent browser mutation (AUDIT_UNAVAILABLE). The other canonical JSON implementations in the repo sort by code unit. - New records sort keys by UTF-16 code unit. - Verification accepts a record whose hash matches either the code-unit or the previous localeCompare order, so existing logs keep verifying exactly as before and new records chain onto them. (An old record written under a different locale still depends on that locale, as it did before; new records no longer do.) - Rotated audit files are ordered by code unit too (their names are ASCII, so the order is unchanged). Tests: tests/browser-audit-store.test.mjs "the audit hash does not depend on the system locale" (append under sv_SE, verify and append under en_US, verify under sv_SE again, in child processes) and "records hashed with the earlier locale-ordered keys still verify and extend the chain".
…oviders Two lookups of cmd.exe disagreed. providerCliRegistry (batch provider launches) takes ComSpec, then %SystemRoot%\System32\cmd.exe. terminalLaunch (the plain terminal when no PowerShell exists) searched PATH before SystemRoot, so a cmd.exe in a project folder or any earlier PATH entry started as the terminal shell. Both now use windowsCommandPromptPath(): ComSpec, then %SystemRoot%\System32\cmd.exe, never PATH. pwsh is still found on PATH, where its installer puts it; that is unchanged. Test: tests/terminal-launch.test.mjs "the Windows terminal falls back to the system cmd.exe, never one found on PATH, like provider launches" (injected platform, environment and file checks).
pollDeviceCode() awaited fetch without a catch. One dropped connection or the 15 s request timeout (an AbortError from the per-request controller) left the loop; startDeviceFlow() swallowed the error, so the UI kept showing the code until it expired and an approval on GitHub was never picked up. - A network error, a request timeout, a non-OK response or an unreadable body now backs off (interval doubled, at most 60 s, RFC 8628 3.5) and polls again until the code's lifetime ends. Cancelling the flow (new flow, sign-out) still stops it at once. - slow_down adds 5 s for this and later polls and honours a longer `interval` GitHub sends with it. access_denied, expired_token and the flow lifetime end the poll as before. Test: tests/github-auth.test.mjs "device polling survives network errors and timeouts, backs off, and honours slow_down" (fake fetch and clock: fetch failure, timeout, 503, slow_down, pending, then approval).
…the host resolved The entry guard hashed the module whose URL equalled realpath(spec.entryPath) as the host computed it before spawning, while the child was started with spec.entryPath and node resolves the main module itself. An entry, or a folder above it, replaced by a symlink between the two resolutions sent the main module to another URL, which the load hook passed to nextLoad unchecked: the untrusted module ran. The hooks now also resolve: the one module resolved without a parent is node's main entry, and its URL is checked with the same hash as the host's URL, so whatever file node resolves the entry to must hold the trusted bytes, and those bytes are what runs (CommonJS entries included; node takes the source the hook returns for them too). Tests start the service through a "node" that changes the plugin files first and assert the untrusted module's marker is never written: the entry replaced by a symlink (ES module and CommonJS), the entry's folder replaced by a symlink, and a CommonJS entry replaced by content.
The probe test compares the blocking and background answers. Both used the 3 s production limit, which a loaded machine can hit while starting the stub CLI, so the test failed on load rather than on a wrong answer. The limit is now a parameter (default unchanged) and the test passes 30 s.
|
P2: make activity-triggered WebGL replanning respect the camera settle boundary. At
Camera movement cannot postpone the second callback. In addition, A concrete regression scenario using the existing fake-clock harness:
The activity callback can detach/attach renderers during that pan. This puts context creation/disposal and the DOM fallback refit back onto the interaction path. The current pan test checks movement without this pending activity callback, so it does not cover the combination. Please make both scheduling paths honor the last camera/layout change, and add a regression asserting that no context swaps occur until the quiet interval has elapsed even when activity replanning is pending. Immediate release for loss of eligibility should remain intact. This is a static-review finding; I have not run the harness or measured visible stutter. Also, this branch inherits the two remaining curl parser findings from #92: #92 (comment). Please carry those fixes through #94 into this branch. |
|
Added 782b6b3: during a wheel/trackpad pan the scene is now composited ( |
…-output-dir alone write nothing curl resets its per-transfer options at --next (-:, also inside a short cluster), so --output-dir of one operation does not move the -o / -O files of another. classifyFetch collected every output of the line and applied the last --output-dir to all of them, so `curl --output-dir ../out -o a URL --next --output-dir . -o b URL` was read as two writes inside the project. Each operation is now resolved with its own outputs, -O count, --remote-name-all and --output-dir, and the facts are combined. --output-dir without -o, -O or --remote-name-all sends the response to standard output; it no longer counts as a write to that folder. Side files (cookie jar, header dump, trace, ...) and real outputs are still judged.
…e base for this change
Startup awaited the startup page (a data: URL, 105 to 165 ms on the audit machine) before it started any service, and services took another 60 ms after it. The page load now runs next to service startup, and the application surface replaces the page as soon as the IPC handlers exist. A close during the page load is still a quiet quit. The application surface aborting a page that is still loading is expected and ignored; a real page load error on a live window is still reported as a failed startup once services are up.
The main bundle imported every dependency when it started, including those only a few paths need. Measured in Electron on this machine: electron-updater 26 to 31 ms (packaged update checks only), yaml 9 to 12 ms (Hermes configs), @xterm/headless 7 to 8 ms (agent control screens and the Even G2 view), secure-remote-password 2.5 ms (Even G2 pairing). They now load through `lazyRequire` on first use; the load stays synchronous, so no caller changes shape. The two Electron smoke runners (test code behind env flags, about 1,200 lines) are dynamic imports and become their own chunks. The main bundle import in the hidden startup harness drops from about 88 ms to about 27 ms. A test keeps static imports of these modules out of src/main.
The first limits read started `codex app-server` and it stayed up until the app quit: about 58 MB, plus the git processes it starts, for the whole session, while the renderer polled every 60 s even with its window hidden. - The app-server is stopped after 3 minutes without a Codex limits read and started again by the next read (the first read after that is a cold start, as on launch; the snapshot cache still covers 60 s). - The renderer reads limits only while its document is visible and reads at once when it becomes visible again. With the window visible nothing changes: the widget reads every 60 s and the app-server stays up. Even G2 and plugin reads go through the same service and start it on demand.
…te in it Resolving the provider CLIs checks each provider's commands in every search directory: 478 statSync calls on the main thread at startup (and on every "Recheck"), most of them for per-provider install directories that do not exist. Each directory is now checked once per resolution, and a candidate in a directory that is missing (or not a directory) is recorded as "missing" without its own stat, the same answer that stat gives. Candidates in existing directories are checked exactly as before. The child PATH filter reuses the same per-resolution answers. Hidden startup harness, empty profile: 478 -> 267 statSync, about 8 -> 5 ms on the main thread.
Twelve places cut newline-delimited JSON by hand, each with its own limit handling: the two gateway decoders (512KB and 128KB, near byte for byte copies), the runtime and agent control gateways, the Codex limits client, the plugin service supervisor, the provider smoke RPC, the runtime client and permission gate, and both MCP helpers (socket and stdin). `src/agent-runtime/ndjson.mjs` is now the only implementation. It ships next to the helpers, so the helpers and the main process use the same file. `NdjsonLineReader` returns complete lines and never buffers more than `maxLineBytes` of one line; a longer line either throws (the default) or is reported once through `onOversize` and dropped up to its newline. `NdjsonDecoderBase` adds JSON parsing with the caller's error factories; the two gateway decoders keep their names, limits and error codes on top of it. Every site keeps its limit. Two readers had no bound at all and now use the protocol's own limit: the orchestration helper's gateway socket (a longer line closes the connection; the gateway never sends one) and its stdin (a longer request gets the same "exceeds 128KB" error as the browser helper's). Readers that took string chunks now take bytes, so a UTF-8 character split across chunks is decoded whole.
The four local gateways (agent browser, orchestration, agent runtime, agent control) each hashed and compared capability tokens and created, published and removed their Unix socket endpoints on their own, with small differences: two comparisons had no length guard, one compared raw tokens, the listen and close helpers existed twice and inline twice. A single skeleton for all four is not worth the risk: their protocols, lease models and Windows pipe host recovery differ. What they share now lives in `gatewaySocket.ts`: - `tokenDigest` / `tokenMatches`: SHA-256 of the token, constant-time compare with a length guard, a missing digest never matches, the presented digest is wiped. Agent control now keeps the digest of its token and compares digests like the other three. - `makePrivateDirectory`, `listenOnEndpoint` (listen, then 0600 on the socket file), `closeServer`, `removeEndpoint`, and the 100-byte Unix socket path limit. Endpoint names, fallback directories and each gateway's error tolerance on removal are unchanged.
There were three: the browser catalog's (strict, but it rebuilt objects, so integer-like keys came out in numeric order), the orchestration catalog's (no checks at all; an undefined property came out as the bare word `undefined`, which is not JSON), and the browser audit's `stableJson`. `canonicalStringify` in tool-catalog.mjs is now the only one; the orchestration catalog re-exports it. Keys are sorted by UTF-16 code unit, strictly as strings. Strict by default, as the browser bridge already was: cycles, non-finite numbers, non-plain objects and undefined array entries throw, undefined properties are left out. The browser audit uses `lenient`, which answers those as JSON.stringify would, and verifies records written before howdeploy#93 with `compareKeys: localeCompare`, so its hash chain is unchanged. A test replays 500 random JSON values through the previous audit serializer in both key orders and requires the same text.
Nine places decided whether a path lies inside a root, in five styles:
`startsWith(root + sep)` (plugin media, the Even G2 web root),
`relative()` with `startsWith("..")` (browser policy, session trust,
command review, base protection), and `relative()` with an exact ".."
test (plugin assets, Home media, the plugin hook runner).
`isPathInside(root, candidate, { allowRoot })` in
src/agent-runtime/path-inside.mjs replaces all of them; it ships with the
hook runner, which uses it too. The rules are the strictest correct ones:
- both paths are resolved, so `..` segments count;
- a sibling sharing a prefix (`/root2`) is outside;
- a child named `..cache` is inside (the `startsWith("..")` variants
refused it);
- a relation on another drive or share is outside;
- Windows compares case-insensitively, as its path functions do;
elsewhere the comparison is exact, so a differently cased path on a
case-insensitive macOS volume counts as outside unless it went through
fs.promises.realpath.
Links are still resolved by the callers exactly where they were before.
Table tests cover POSIX, macOS and Windows paths, links out of and into
the root, and a case-insensitive volume on the real file system.
Four regexes decided which names hold secrets and had drifted apart: the browser audit log (exact key match: `passcode`, `apiKey` or `x-refresh-token` went into the log), what agents read back from the browser, and the screenshot field mask twice (a TypeScript check and a copy inside the page script that had to be kept in sync by hand). `safety/sensitiveNames.ts` now holds one credential list with two extensions: keys (cookies, auth headers, web storage) and form fields (one-time codes, anything auth). The page script is built from the same source string. What agents read and which fields screenshots mask are unchanged (a test compares both with the previous lists name by name). The audit log now redacts any key containing a sensitive name, and a `name=value` string with one, instead of its shorter exact list.
The full provider list was typed out by hand in eleven places (settings defaults and validation, session store, runtime gateway, plugin manifest checks, plugin IPC, radial menu, plugin frame, renderer defaults and the radial launcher list), and the limit provider list in four. Adding a provider meant finding all of them. - `isProviderId` in providerCatalog.ts checks against the catalog's labels (typed as a Record over every id, so it cannot miss one). - `AGENT_PROVIDERS` and the new `LIMIT_PROVIDERS` in contracts.ts are the only lists; `RADIAL_LAUNCHER_ITEMS` is the launcher list plus its three actions. - Settings migrations keep their frozen historical lists (legacy, pre-Qwen, added providers): those describe old profiles, not the current catalog. Order and contents of every list are unchanged; a test ties them to the catalog.
…overlays The Kimi overlay (agent-browser/ProviderLaunch.ts) was a renamed copy of the Hermes one (hermesConfig.ts): the same cross-process lock with stale owner recovery, compare-and-swap writes, atomic replacement, backups, hashes and file helpers, about 300 lines each. The lifecycle hook overlays (ProviderRuntimeLaunch.ts) had a third atomic write, backup restore and file helpers. configOverlay.ts now holds them once. Hermes and Kimi call it with their label, so every error message reads as before; the Kimi test seams (beforeReclaim, beforeRelease) stay. The recovery journals, their fields and paths, and the lock file format are unchanged, so an overlay left by an older build is still recovered. Small differences that went away: new directories made by an atomic write are chmod'ed to 0700 on every path (only the lifecycle overlays did this), and the lifecycle overlays now also re-apply the file mode after the rename, as Hermes and Kimi did.
Seven TerminalManager test files each defined their own `availableRegistry()` and `fakeSpawner()`, the same code with small drifts (pid base, what `write` records, whether answers are frozen). tests/helpers/terminal.mjs has one of each. The registry answers are frozen like the real registry's; the spawner takes the pid base and a write observer as options and returns a PTY that can emit data and exit.
knip (with tests and scripts as entry points) found one function that nothing calls, `byteLengthOfCanonicalJson` in tool-catalog.mjs and its declaration, and 46 exported values that are used only inside their own module: timeouts and limits, snap and region constants, command review helpers, the provider CLI definitions, and three names the agent-browser barrel re-exported for nobody. They are module-private now. After this change knip reports no unused value export in src/. Unused exported types (147, many of them option interfaces) are left alone.
Every card subscribed to terminal:data on its own, so the preload handed each output batch to every card across the context bridge and all but one dropped it by id. With 24 cards printing a spinner that was about 7 000 bridge calls a second for 300 batches. The preload now keeps one IPC listener and a per-session router; a card subscribes with its session id and is called only for its own output. onData without an id still receives everything.
…ole canvas Each session had its own 16 ms batch timer and every batch was its own IPC message, so a canvas of busy cards cost one timer and one message per card per window (24 cards with a spinner: ~300 messages and timers a second). One timer now flushes every session with queued output, in the order its output first arrived, and TerminalRendererOutbox sends the renderer's share of that flush as a single terminal:data-batch message at the end of the task. Session and removal events flush what was collected before them first, so the renderer sees the manager's order. Observers still get one event per session batch. Measured: 24 cards, 295 -> ~40 messages a second.
…ontext pool Only the focused card used xterm's WebGL renderer; every other visible card paid for the DOM renderer's style, layout and paint on each row it changed, which was most of the renderer's CPU under heavy output. Cards now ask a renderer-wide pool (webglContextPool.ts) for a context. It hands out 10, below Chromium's 16 active WebGL contexts per renderer process: the focused card first, then by on-screen area (a busy holder keeps its context against a card less than 1.25x larger), then by most recent output. Cards without a context keep the DOM renderer. Camera, layout and focus changes settle for 200 ms before contexts move, and a fresh grant is kept for 1 s, so panning and wheel zoom do not rebuild contexts. A card that leaves the screen, zooms above 1x or enters summary mode releases its context at once and loses it explicitly, so the slot is free before GC. A lost context (or a failed attach) puts the card back on DOM with its buffer intact for 30 s, doubling on repeats. Going back to DOM refits the grid: WebGL snaps cells down to whole device pixels, so a grid fitted on WebGL can be too wide for DOM; a DOM-fitted grid always fits WebGL, so swaps do not resize the PTY. Measured (hidden harness, medians of 3): 8 visible cards at 8 x 1 MB/s, Electron CPU 59.9% -> 47.0%, renderer 37.4% -> 18.5%; 16 visible with 8 flooding, 64.6% -> 39.3%. Idle and 24 spinner cards unchanged. GPU process footprint at idle +105 MB (8 cards) / +140 MB (16 cards).
The scene carried will-change only during a pointer drag and a zoom. A wheel or trackpad pan, the usual way to move around on a Mac, translated an uncomposited scene, so Chromium repainted every card on every frame. The wheel pan now marks its own gesture (same 160 ms trailing settle as a zoom) and the scene is composited while it lasts. A pan does not change the scale, so the cached raster stays sharp; once the wheel is silent the hint drops and the resting scene is uncomposited as before. The settle logic for both gestures moved to gestureSettle.ts. Measured (hidden harness, 16 visible cards, 5 s of wheel pan at 60 Hz, medians of 3): renderer Paint 143 -> 27 ms/s, raster 9.6 -> 2.6 ms/s, Electron CPU 36.7% -> 31.7%, GPU process 11.8% -> 7.4%. Idle, 8 x 1 MB/s output and 24 spinner cards unchanged.
…o settle The pool had two independent timers: viewportChanged() restarted the settle timer, while touch() on a waiting card scheduled a plan() 1000 ms later that nothing could postpone. That plan() also cancelled the pending settle timer. When output and a pan overlapped, contexts were created and disposed (and DOM-fallback cards refitted) in the middle of the pan. Every timer now re-plans through one gate: until WEBGL_SETTLE_MS has passed since the last viewportChanged() it reschedules itself to that boundary instead of planning. plan() keeps the pending settle timer while the camera is still settling. Losing eligibility still releases the context at once.
782b6b3 to
aab91b5
Compare
|
Thanks. The finding reproduced with the fake-clock harness and is fixed in one new commit. This branch also carries the two #92 curl fixes through the rebuilt #94 base. Final head: The branch was replayed onto the new #94 head
P2: activity-triggered replanning ignored the camera settle boundaryCommit: Root cause. Fix.
Tests
Both tests fail at the old head and pass now. The earlier pan, idle-yield, pin, backoff and budget tests are unchanged and pass. Typecheck and |
Depends on
This branch is based on #94 (
refactor/startup-and-dedup@47d229a, which merges #89, #90, #92 and #93) and should be merged after it. Its own changes are 3 commits. Each commit covers one topic and passes typecheck and its tests.Goal
Find what burns CPU in the main process, the renderer and the GPU under terminal load and a real agent session, and fix the clear wins. Commits 1–2 do not change visible behaviour; commit 3 moves on-screen cards from xterm's DOM renderer to its WebGL renderer, within a fixed context budget.
How it was measured
The measurements use the hidden harness from
scripts/bench-runtime:show:false, off-screen and unfocusable;What was recorded for each scenario window:
ps;Performance.getMetrics;Scenarios:
flood.mjs).qwen3.5:9b, running 3 tool calls (Bash, Read, Write) and one permission prompt, with CanvasTTY's hooks active.Each scenario was run 5 times on the base and 5 times on the branch, alternating in ABBA order. The tables show medians. The machine (M1 Max, 10 cores) was loaded by another project: 1-minute load average 2.9–9.6, median about 5.9. Treat absolute numbers as indicative; the before/after difference is the meaningful part.
What the profile shows (base)
flushOutput, the PTY handler and the scrollback ring together take about 30 ms in 11 s.indexOf-based OSC parser, for claude and qwen only.print,decode,parse,scroll) and the DOM renderer (replaceChildren,createRow, style and layout) dominate.terminal:dataitself, so each batch crossed the context bridge 24 times and was dropped by 23 cards. That is about 7,000 bridge calls a second, roughly 11 % of the renderer's JavaScript.@xterm/addon-webglwas enabled only for the focused card at zoom ≤ 1. Every other card used the DOM renderer, which is the largest renderer cost under output (commit 3).Changes
perf(renderer): route terminal output to its own card in the preloadTerminalDataRouter(src/shared/terminalDataRouter.ts).onData(listener, id)and is called only for its own session.onDatawithout an id still receives everything.tests/terminal-data-router.test.mjs.perf(terminal): send one renderer message per output flush for the whole canvasTerminalManagerflushes every session with queued output, in the order its output first arrived.TerminalRendererOutboxsends the renderer's share of that flush as a singleterminal:data-batchmessage at the end of the task.tests/terminal-output-batching.test.mjs. They check one message per window, one timer for 8 sessions, ordering on dispose, and outbox ordering.docs/ARCHITECTURE.md(and zh-CN) anddocs/performance-benchmark.mdare updated. The bench byte counter understands the batch channel.perf(terminal): draw every on-screen card with WebGL from a bounded context poolsrc/renderer/src/features/terminal/webglContextPool.ts: a renderer-wide pool of 10 WebGL contexts. Chromium keeps at most 16 active WebGL contexts per renderer process and silently loses the oldest past that; 10 leaves room for plugin canvas apps and for released contexts not yet collected.WEBGL_lose_context.loseContext()so the slot is really free before GC.WorkspaceCanvastells the pool about camera, HOME-editing and fullscreen changes; cards report moves, resizes, focus and eligibility.tests/webgl-context-pool.test.mjs(16): priority, visible area, settle on pan, release on hide/ineligible/unmount, LRU and idle-holder hand-over, hysteresis, minimum hold, the budget is never exceeded, context loss with doubling backoff, no WebGL2 at all, off-screen output never re-plans, stale unregister after a remount.docs/ARCHITECTURE.md(and zh-CN) describe the pool.Commits 1–2 do not change the data, offsets or order the card receives, and the card still deduplicates by absolute offset.
Commits 1–2: before / after (medians of 5)
Base is
47d229a, branch isdfd4458. GPU memory was not recorded for this pair; these commits do not touch rendering (the GPU process CPU column is within noise).Run spread (Electron CPU %):
Flood and agent differ only within noise. Flood still sends 4× fewer messages, but its cost is xterm parsing and DOM rendering, not message count. All 10 agent turns finished the task.
In the branch's spin profile:
dispatch;webContents.sendin main takes 6 ms instead of 37 ms.WebGL pool (commit 3): before / after
Base is
dfd4458(commits 1–2), branch is commit 3. Same hidden harness, but the window is sized so every card is on screen (3100×1300 for 8 cards, 3100×2320 for 16 and for 24 cards, of which 16 are visible), device pixel ratio 2. 3 runs per build, alternating (owner limit), medians. 1-minute load average during these runs: 4.9–18.3, median 6.7. GPU memory is the GPU process'sfootprint(macOS physical footprint) at the end of the window; "IOAccel" is its IOAccelerator (graphics) part. Chromium's ownapp.getAppMetrics()working set is shown too.Run spread (Electron CPU %): 8-card flood 58.5–63.4 → 45.6–47.5; 16-card flood 57.5–69.2 → 37.2–46.3; spinners 27.1–30.2 → 26.1–30.1.
Behaviour checks in the hidden harness (16 visible cards, final build):
Visual parity (same build with WebGL blocked vs not, offscreen capture of the same card): colours, bold/italic/underline, 256-colour and truecolour backgrounds, URL and OSC 8 link underline, box drawing and CJK look the same. Differences, all xterm's own and already true for the focused card before this change:
The cursor could not be compared: the harness window is never focused, so neither renderer draws it.
Risks (commit 3)
WEBGL_CONTEXT_BUDGET).Not in this PR
scripts/bench-runtimeruns the app withapp.getAppPath()pointing at the harness folder, so agent hook helpers are not found there. Flood and pan are unaffected. Linkingsrc/,scripts/andbuild/into the harness folder fixes it, if an agent scenario is added to the bench.Gates on the branch head (commit 3)
npm run typecheck: pass.node --test --test-concurrency=2): 1149/1149.npm run build: pass.npm run audit:secrets: pass.npm run test:even: 47/47.