fix: startup page race, provider key paste crash, and agent access to CanvasTTY's control plane - #100
Merged
Conversation
…e settled Startup could fail with "ERR_ABORTED (-3) loading 'data:text/html...'" when the machine was busy. Since services start while the startup page loads, the application surface could start its navigation while the page was still committing or loading in its own renderer process. The application surface then committed first, and the page's ERR_ABORTED arrived afterwards, while the loadFile promise was still waiting. Electron's loadURL/loadFile promise takes the first main-frame did-fail-load it sees as its own, so the application load rejected with the startup page's abort and startup showed the failure page. Services still start while the page loads. The application surface now waits until the page load has settled, which it usually has by the time services are up. A close during the page load is still a quiet quit, and a real page error on a live window still fails startup. Measured with 40 sequential hidden launches per build, each helper process (renderer, GPU, utility) paused at random for 20 to 300 ms during the first navigations: origin/main failed 4 and 6 of 40, the same loop with this change 0 of 80. The build before "perf(startup): start services while the startup page loads" had 0 of 40. Without pauses all builds start 40 of 40; the median time to a ready window goes from 342 to 379 ms.
Pasting a MiniMax API key into Settings -> Agents -> Provider API keys turned the window black. The key field's onChange read event.currentTarget.value inside the setDrafts((current) => ...) updater. React runs that updater later, during render, whenever an earlier update of the component is still pending (a paste over a character already in the field, the second keystroke of fast typing). By then the event has finished dispatching and currentTarget is null, so the render threw "Cannot read properties of null (reading 'value')" and React unmounted the whole root. The renderer process stayed alive with an empty #root. Key length and shape are not the cause: a 120-character key pasted over a pending edit breaks the same way. The handler now reads the value while the event dispatches and passes it to the updater. No other renderer updater reads an event (checked with an AST scan over src/renderer). Tests: tests/provider-secrets-settings.test.mjs bundles the real component against a React stand-in whose updaters run after dispatch and pastes fake sk-cp-/sk-api- keys of 120 to 10 000 characters, with a trailing newline and space, over a pending edit (fails before this change with the same TypeError); a line-level guard keeps event reads out of state updaters in the renderer.
An error thrown while React renders unmounts the whole root, but the renderer process lives on, so render-process-gone never fires and the main process's crash reload never runs: the window stays black until the app is restarted. That is what the broken provider key field did. The React root now handles onUncaughtError. The first such error logs it with its component stack and reloads the application surface in place; sessions live in the main process and survive the reload, as after a renderer crash. A second one within 30 s (or with no session storage to remember the reload) shows a static page with a Reload button instead of reloading in a loop. Tests: tests/uncaught-error-recovery.test.mjs (reload, cooldown, clock change, no storage, root wiring). Checked in a hidden app: an injected render error reloads to a usable Settings screen, a second one within the cooldown shows the recovery page; forcefullyCrashRenderer still reloads through render-process-gone.
…ores and sockets Base protection now treats CanvasTTY's private data as credentials. A shell or file tool call that names the agent-control token or descriptor, the gateways' connection records, the provider and plugin secret stores, account homes, the GitHub sign-in or prepared launch runs (paths taken from the app's own userData folder, passed in by the app), or the control and runtime socket folders under the temporary folder, is refused whatever the program: readers, copies, encoders, sqlite3, recursive walks of the app folder, interpreter one-liners and heredocs, curl --unix-socket, nc -U, socat and Python sockets. The model is told calmly that agents cannot control CanvasTTY this way and to ask for an Orchestrator launch. The project, the app's settings, other sockets, an agent's own account home and the bundled control CLI are unaffected.
…hen close An unauthenticated or malformed request to the control endpoint now gets a stable INVALID_REQUEST whose message says only sessions CanvasTTY launched as orchestrators may use it and how to get one, with no protocol details, token names or paths. An HTTP request line (curl, a browser) gets a minimal 403 with the same text. Either way the connection is closed right after, instead of waiting for the idle timeout; the connection cap and the one-request-per-connection rule are unchanged. Tests cover a guessed NDJSON request, garbage and an HTTP request, and that the token file is 0600 and its folder 0700 even under umask 0 and a pre-existing loose folder.
This was referenced Sep 29, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Three small fixes found after #89–#95 were merged, in one PR so they can be reviewed together. Each topic is its own commit(s); review commit by commit.
fix(startup): load the application surface only after the startup page settled(was fix(startup): load the application surface only after the startup page settled #96)fix(settings): read a pasted provider key before its state update runsandfix(renderer): bring the window back after an uncaught render error(was fix(settings): pasting a provider API key no longer blacks out the window #97)fix(safety): refuse agent access to CanvasTTY's own tokens, secret stores and socketsandfix(agent-control): answer refused and HTTP requests with guidance, then close(was fix(safety): keep agents away from CanvasTTY's own tokens and control socket, and tell them the way in #99)Checks on the combined head: typecheck, full suite (1170 pass, 1 skipped), electron-vite build, audit:secrets.
1. Startup page race
PR description for branch
fix/startup-aborted(one commit onmain20f6855). #94 and #95 are already merged, so this is a follow-up PR againstmain, not a comment on #94 or #95.Problem
On a busy machine, startup sometimes fails and shows the startup failure page:
It started with "perf(startup): start services while the startup page loads" (#94). Since that change, services start while the startup page (a
data:URL) is still loading. When services finish first,loadFilefor the application surface starts while the startup page is still committing or loading in its own renderer process. If the application surface commits first, the page'sdid-fail-load(-3, thedata:URL) arrives afterwards, while theloadFilepromise is still waiting. Electron'sloadURL/loadFilepromise takes the first main-frame load failure it sees as its own. So the application load rejects with the startup page's abort, andstartApplicationshows the failure page.The
supersededflag from #94 did not cover this. It kept the page's own promise from reporting the abort, but the error came through the application load's promise.Change
createWindowstill starts the page load without awaiting it. The page load now resolves to the error to report, or tonull. It never rejects.startApplicationstill starts services right away. Before it loads the application surface, it waits for the page load to settle. On a normal launch the page has usually finished by the time services are up.Tests
tests/startup-lifecycle.test.mjs: new sequencing test with injected fakes. It fails onmain(the application surface loads while the page is still pending). It passes with this change and also covers a real page error and a close while waiting.npm run typecheck, full suite (--test-concurrency=2, 1155/1155),electron-vite build: pass.Launch loop
Sequential hidden launches with a throw-away HOME and the keychain stubbed. Each launch runs until the window is ready or startup fails. Plain launches fail too rarely to measure: 0 of 40 on every build, including
main. To make the race reproducible, the loop pauses a random helper process (renderer, GPU or utility) three times during the first navigations, for 20 to 300 ms each (SIGSTOP/SIGCONT). This is the kind of descheduling a loaded machine produces.Every failure has the same trace: the page's
did-fail-load -3arrives after the application surface committed, and then comes "startup failed" with thedata:URL.Cost: without pauses, the median time from launch to a ready window goes from 342 ms to 379 ms (40 launches each). The application surface now starts after the page has loaded (about 85 ms after the page load starts) instead of when services are up (about 45 ms). For reference, the parent of the #94 startup commit, where the page and services ran one after the other, measured 457 ms in the same loop. That build has different renderer code, so the comparison is only approximate.
2. Black window after pasting a provider key
Based on main (20f6855). Independent of #96; the two merge cleanly, and the packaged test build carries both.
Symptom
In Settings → Agents → Provider API keys, pasting a MiniMax key (a long
sk-cp-…/sk-api-…string) turned the whole window black, and it stayed black until the app was restarted.Root cause
The key field's change handler read the event inside a functional state updater:
React only runs an updater while the handler is still executing when the component has no update pending. When one is pending (a paste over a character already in the field, or a second keystroke before the render), React runs the updater later, during render. By then the event has finished dispatching and
currentTargetisnull. The render throwsTypeError: Cannot read properties of null (reading 'value'), and React 19 unmounts the whole root because nothing catches the error.#rootis left empty.The renderer process keeps running, so
render-process-gonenever fires and the main process's crash reload never runs. That is why the window stayed black.The key's length and shape are not the cause. A 120-character key pasted over a pending edit breaks the same way. Nothing in the renderer validates, masks or runs a regex over the key, and the main-process redaction registry is linear: registering a 100 000-character key and masking 200 KB takes under 15 ms.
Repro (before the fix)
Hidden-window harness (keychain stubbed,
--use-mock-keychain, throwaway user data), running the packaged build of main + #96. It opens Settings, focuses the MiniMax field and pastes a fake key with CDPInput.insertText:#rootwith 0 children and empty body text. The renderer console shows the TypeError from theprovider-secret__inputonChange updater. The renderer PID doesn't change,render-process-gonedoesn't fire, and nothing recovers.Fix
ProviderSecretsSettings: readevent.currentTarget.valuewhile the handler runs, then pass the value to the updater. An AST scan oversrc/rendererfound no other state updater that reads an event.onUncaughtError:Tests
tests/provider-secrets-settings.test.mjssk-cp-/sk-api-keys of 120, 200, 250, 500, 2 000 and 10 000 characters, plus one with a trailing newline and one with a trailing space, each over a pending edit.tests/uncaught-error-recovery.test.mjs: the first error reloads, a second within the cooldown shows the page, an error after the cooldown reloads again, a clock that moved backwards doesn't block the reload, no storage shows the page, and the root routes errors to the recovery.--test-concurrency=2, fake HOME), on this branch and on this branch merged with fix(startup): load the application surface only after the startup page settled #96. Typecheck,electron-vite build,audit:secretsandtest:evenpass.Checked in the packaged app
Built with
npm run packagefrom main + #96 + this branch. The harness loaded theapp.asarcontents, which matchout/byte for byte, with the window hidden, the keychain stubbed and a throwaway user data directory:forcefullyCrashRenderer()still recovers throughrender-process-gone. There is a new renderer PID, and the MiniMax field works afterwards.3. Agents trying to reach CanvasTTY's control plane
Based on main (20f6855). Independent of #96 and #97. It merges cleanly with each of them and with both together, and the full suite passes on that merge.
Incident
Someone started an OpenCode agent by hand inside a plain terminal card. CanvasTTY didn't launch it as an orchestrator, so it had no
canvastty_agentsMCP server and no control environment. It still tried to control CanvasTTY:cat.curl --unix-socketandnc -U.The gateway refused every attempt with
INVALID_REQUEST/ "Invalid or unauthenticated control request.". That message doesn't tell the agent what to do instead, so it kept guessing.What changed
1. Base protection treats CanvasTTY's private data as credentials
safety/commandFacts.tsnow records a new fact,appPrivate, andbaseProtection.tsdenies it under a new rule,app-private. That rule is checked before all the others. A shell or file tool call is refused when it names any of the following:agent-control/), the gateways' connection records and sockets (browser/runtime,lifecycle/runtime,orchestration/runtime),provider-secrets.bin,plugin-secrets/,account-homes/,github-oauth.jsonandlaunch-runs/. The app passes these paths in throughcanvasTtyPrivateData(app.getPath("userData"))→DecisionHooks→checkBaseProtection. Nothing is hard-coded per platform.ctty-control-*,ctty-runtime-*,ctty-orch-*,ctty-<uid>-*). These are recognized without the userData path.$CANVASTTY_CONTROL_CONNECTION,$CANVASTTY_*_ADDRESS,$CANVASTTY_*_CAPABILITY).The check doesn't depend on the program. It looks at every word and redirection, at paths after
=or:(--unix-socket=P,UNIX-CONNECT:P,file://), at paths quoted inside interpreter one-liners and heredocs, and at globs. It also seescdinto those folders,$(…)substitutions andbash -c. Interpreter code that builds the path from pieces is caught when the app folder's name and a private name appear together.Other rules:
grep -r,rg,find,tar,rsync…) is refused. A plainlsorcat settings.jsonof that folder is still allowed.node "$CANVASTTY_CONTROL_CLI" …,canvastty-control.mjs --connection …).2. The control gateway explains the refusal
In
AgentControlGateway, an unauthenticated or malformed NDJSON request still getsINVALID_REQUEST. Its message is now stable guidance (below), and the gateway closes the connection straight after instead of waiting for the 10 s idle timeout. An HTTP request line (curl'sGET / HTTP/1.1, a browser, an HTTP/2 preface) gets a minimalHTTP/1.1 403 Forbiddenwith the same text intext/plain, then the connection closes. The message has no protocol details, token names or paths. The 32-connection cap, the one-request-per-connection rule and the request size limit are unchanged.3. Token file permissions
Checked, and no change was needed. The token and
connection.jsonare created with mode 0600 in a folder that ischmoded to 0700, and the socket folder is 0700. A new test covers this underumask 0with a pre-existing 0777 folder.Messages shown to agents
Base protection deny:
Control endpoint (NDJSON
error.message, and the body of the HTTP 403):Scope and limits
Tests
tests/base-protection-app-private.test.mjs, table-driven, with a fake HOME and a fake userData folder:cat,head,grep,cp,base64 <,xxd,strings,sqlite3; globs;cdthen a relative read;$(cat …);bash -c;find,grep -r,rgandtarof the app folder;python3 -c,node -eand heredoc readers; a Python path built from pieces;curl --unix-socket(both spellings),nc -U,socat UNIX-CONNECT:, a PythonAF_UNIXconnect, a runtime socket in userData; the descriptor and capability variables.grep -rn plugin-secrets src, a commit message that mentions them, the app'ssettings.json,lsof the app folder, the control CLI with its descriptor, the Docker socket and other sockets,ls $TMPDIR,/tmp/ctty-notes, and a URL containingagent-control/token-1.DecisionHooksforwards the private data.tests/agent-control.test.mjs:Content-Lengthand the same text.--test-concurrency=2, fake HOME). Also 1170 passing on this branch merged with fix(startup): load the application surface only after the startup page settled #96 and fix(settings): pasting a provider API key no longer blacks out the window #97.npm run buildandaudit:secretspass.