Skip to content

fix: startup page race, provider key paste crash, and agent access to CanvasTTY's control plane - #100

Merged
howdeploy merged 5 commits into
howdeploy:mainfrom
BIackFIame:fix/post-merge-batch
Sep 29, 2026
Merged

howdeploy merged 5 commits into
howdeploy:mainfrom
BIackFIame:fix/post-merge-batch

Conversation

@BIackFIame

Copy link
Copy Markdown
Contributor

Three small fixes found after #89–#95 were merged, in one PR so they can be reviewed together. Each topic is its own commit(s); review commit by commit.

Checks on the combined head: typecheck, full suite (1170 pass, 1 skipped), electron-vite build, audit:secrets.


1. Startup page race

PR description for branch fix/startup-aborted (one commit on main 20f6855). #94 and #95 are already merged, so this is a follow-up PR against main, not a comment on #94 or #95.

Problem

On a busy machine, startup sometimes fails and shows the startup failure page:

ERR_ABORTED (-3) loading 'data:text/html;charset=utf-8,...'

It started with "perf(startup): start services while the startup page loads" (#94). Since that change, services start while the startup page (a data: URL) is still loading. When services finish first, loadFile for the application surface starts while the startup page is still committing or loading in its own renderer process. If the application surface commits first, the page's did-fail-load (-3, the data: URL) arrives afterwards, while the loadFile promise is still waiting. Electron's loadURL/loadFile promise takes the first main-frame load failure it sees as its own. So the application load rejects with the startup page's abort, and startApplication shows the failure page.

The superseded flag from #94 did not cover this. It kept the page's own promise from reporting the abort, but the error came through the application load's promise.

Change

  • createWindow still starts the page load without awaiting it. The page load now resolves to the error to report, or to null. It never rejects.
  • startApplication still starts services right away. Before it loads the application surface, it waits for the page load to settle. On a normal launch the page has usually finished by the time services are up.
  • A close during the page load is still a quiet quit. A real page error on a live window still fails startup.

Tests

  • tests/startup-lifecycle.test.mjs: new sequencing test with injected fakes. It fails on main (the application surface loads while the page is still pending). It passes with this change and also covers a real page error and a close while waiting.
  • npm run typecheck, full suite (--test-concurrency=2, 1155/1155), electron-vite build: pass.

Launch loop

Sequential hidden launches with a throw-away HOME and the keychain stubbed. Each launch runs until the window is ready or startup fails. Plain launches fail too rarely to measure: 0 of 40 on every build, including main. To make the race reproducible, the loop pauses a random helper process (renderer, GPU or utility) three times during the first navigations, for 20 to 300 ms each (SIGSTOP/SIGCONT). This is the kind of descheduling a loaded machine produces.

Build Failed startups (with pauses)
main before the #89–#95 stack (31f287d) 0 / 40
#89 head (7d6009c) 0 / 40
parent of the #94 startup commit (b57b9dd) 0 / 40
#94 startup commit (6a40740) 1 / 40
#94 head (8accbd1) 2 / 40 and 4 / 30
main (20f6855) 4 / 40 and 6 / 40
this branch 0 / 80

Every failure has the same trace: the page's did-fail-load -3 arrives after the application surface committed, and then comes "startup failed" with the data: URL.

Cost: without pauses, the median time from launch to a ready window goes from 342 ms to 379 ms (40 launches each). The application surface now starts after the page has loaded (about 85 ms after the page load starts) instead of when services are up (about 45 ms). For reference, the parent of the #94 startup commit, where the page and services ran one after the other, measured 457 ms in the same loop. That build has different renderer code, so the comparison is only approximate.


2. Black window after pasting a provider key

Based on main (20f6855). Independent of #96; the two merge cleanly, and the packaged test build carries both.

Symptom

In Settings → Agents → Provider API keys, pasting a MiniMax key (a long sk-cp-… / sk-api-… string) turned the whole window black, and it stayed black until the app was restarted.

Root cause

The key field's change handler read the event inside a functional state updater:

onChange={(event) => setDrafts((current) => ({ ...current, [secretId]: event.currentTarget.value }))}

React only runs an updater while the handler is still executing when the component has no update pending. When one is pending (a paste over a character already in the field, or a second keystroke before the render), React runs the updater later, during render. By then the event has finished dispatching and currentTarget is null. The render throws TypeError: Cannot read properties of null (reading 'value'), and React 19 unmounts the whole root because nothing catches the error. #root is left empty.

The renderer process keeps running, so render-process-gone never fires and the main process's crash reload never runs. That is why the window stayed black.

The key's length and shape are not the cause. A 120-character key pasted over a pending edit breaks the same way. Nothing in the renderer validates, masks or runs a regex over the key, and the main-process redaction registry is linear: registering a 100 000-character key and masking 200 KB takes under 15 ms.

Repro (before the fix)

Hidden-window harness (keychain stubbed, --use-mock-keychain, throwaway user data), running the packaged build of main + #96. It opens Settings, focuses the MiniMax field and pastes a fake key with CDP Input.insertText:

  • A 120-character fake key into the empty field works.
  • A 200-character fake key pasted over it leaves #root with 0 children and empty body text. The renderer console shows the TypeError from the provider-secret__input onChange updater. The renderer PID doesn't change, render-process-gone doesn't fire, and nothing recovers.

Fix

  1. ProviderSecretsSettings: read event.currentTarget.value while the handler runs, then pass the value to the updater. An AST scan over src/renderer found no other state updater that reads an event.
  2. Uncaught render errors no longer leave the window black. The React root now sets onUncaughtError:
    • The first such error is logged with its component stack, and the application surface reloads in place. Sessions live in the main process and survive the reload, as they do after a renderer crash.
    • A second error within 30 s shows a static page with a Reload button instead of reloading in a loop. The same happens when session storage can't record the reload time.
    • The page's text is in Russian and English.

Tests

  • tests/provider-secrets-settings.test.mjs
    • Bundles the real component against a React stand-in whose updaters run after dispatch, the ordering React uses when an update is pending.
    • Pastes fake sk-cp-/sk-api- keys of 120, 200, 250, 500, 2 000 and 10 000 characters, plus one with a trailing newline and one with a trailing space, each over a pending edit.
    • Checks that the field holds the key and Save is enabled, all within a 1 s bound.
    • Fails before the fix with the same TypeError.
    • A line-level guard keeps event reads out of state updaters in the renderer.
  • tests/uncaught-error-recovery.test.mjs: the first error reloads, a second within the cooldown shows the page, an error after the cooldown reloads again, a clock that moved backwards doesn't block the reload, no storage shows the page, and the root routes errors to the recovery.
  • Full suite 1162/1163 passing, 1 skipped (--test-concurrency=2, fake HOME), on this branch and on this branch merged with fix(startup): load the application surface only after the startup page settled #96. Typecheck, electron-vite build, audit:secrets and test:even pass.

Checked in the packaged app

Built with npm run package from main + #96 + this branch. The harness loaded the app.asar contents, which match out/ byte for byte, with the window hidden, the keychain stubbed and a throwaway user data directory:

  • Pasting fake keys of 120 to 10 000 characters, and ones with a trailing newline or space, keeps the window rendered. The field holds the full key and Save is enabled.
  • An injected render error reloads to a usable Settings screen. A second one within 30 s shows the recovery page.
  • forcefullyCrashRenderer() still recovers through render-process-gone. There is a new renderer PID, and the MiniMax field works afterwards.

3. Agents trying to reach CanvasTTY's control plane

Based on main (20f6855). Independent of #96 and #97. It merges cleanly with each of them and with both together, and the full suite passes on that merge.

Incident

Someone started an OpenCode agent by hand inside a plain terminal card. CanvasTTY didn't launch it as an orchestrator, so it had no canvastty_agents MCP server and no control environment. It still tried to control CanvasTTY:

  1. It read the agent-control token file in CanvasTTY's user data folder with cat.
  2. It guessed the protocol of the control socket in the temporary folder, using curl --unix-socket and nc -U.

The gateway refused every attempt with INVALID_REQUEST / "Invalid or unauthenticated control request.". That message doesn't tell the agent what to do instead, so it kept guessing.

What changed

1. Base protection treats CanvasTTY's private data as credentials

safety/commandFacts.ts now records a new fact, appPrivate, and baseProtection.ts denies it under a new rule, app-private. That rule is checked before all the others. A shell or file tool call is refused when it names any of the following:

  • Files in the app's userData folder. These are the agent-control token and descriptor (agent-control/), the gateways' connection records and sockets (browser/runtime, lifecycle/runtime, orchestration/runtime), provider-secrets.bin, plugin-secrets/, account-homes/, github-oauth.json and launch-runs/. The app passes these paths in through canvasTtyPrivateData(app.getPath("userData")) → DecisionHooks → checkBaseProtection. Nothing is hard-coded per platform.
  • The gateways' socket folders under the temporary folder (ctty-control-*, ctty-runtime-*, ctty-orch-*, ctty-<uid>-*). These are recognized without the userData path.
  • Variables that carry a descriptor, an address or a capability ($CANVASTTY_CONTROL_CONNECTION, $CANVASTTY_*_ADDRESS, $CANVASTTY_*_CAPABILITY).

The check doesn't depend on the program. It looks at every word and redirection, at paths after = or : (--unix-socket=P, UNIX-CONNECT:P, file://), at paths quoted inside interpreter one-liners and heredocs, and at globs. It also sees cd into those folders, $(…) substitutions and bash -c. Interpreter code that builds the path from pieces is caught when the app folder's name and a private name appear together.

Other rules:

  • A recursive walk of the app folder (grep -r, rg, find, tar, rsync …) is refused. A plain ls or cat settings.json of that folder is still allowed.
  • The project folder is never affected. Neither are the agent's own config folder (an account home it was launched with), other Unix sockets, or the bundled control CLI (node "$CANVASTTY_CONTROL_CLI" …, canvastty-control.mjs --connection …).

2. The control gateway explains the refusal

In AgentControlGateway, an unauthenticated or malformed NDJSON request still gets INVALID_REQUEST. Its message is now stable guidance (below), and the gateway closes the connection straight after instead of waiting for the 10 s idle timeout. An HTTP request line (curl's GET / HTTP/1.1, a browser, an HTTP/2 preface) gets a minimal HTTP/1.1 403 Forbidden with the same text in text/plain, then the connection closes. The message has no protocol details, token names or paths. The 32-connection cap, the one-request-per-connection rule and the request size limit are unchanged.

3. Token file permissions

Checked, and no change was needed. The token and connection.json are created with mode 0600 in a folder that is chmoded to 0700, and the socket folder is 0700. A new test covers this under umask 0 with a pre-existing 0777 folder.

Messages shown to agents

Base protection deny:

CanvasTTY blocked this: it reads CanvasTTY's own access tokens or secret stores, or talks to its control socket. Agents can't control CanvasTTY this way, and guessing its protocol will not work. If you need other agents, ask the person to start you from CanvasTTY's launcher with the Orchestrator role: you will then get the canvastty_agents tools (spawn_agent, list_routes, wait_for_agent and the rest). Otherwise continue your task without controlling CanvasTTY.

Control endpoint (NDJSON error.message, and the body of the HTTP 403):

CanvasTTY refused this request: this endpoint only accepts requests from sessions CanvasTTY itself launched as orchestrators, and guessing its protocol will not work. If you are an agent and need other agents, ask the person to start you from CanvasTTY's launcher with the Orchestrator role: you will then get the canvastty_agents tools (spawn_agent, list_routes, wait_for_agent and the rest).

Scope and limits

  • Base protection covers what the decision hook covers: Claude Code, Codex, Qwen Code and OpenCode cards that CanvasTTY launched while base protection was on, for shell and file-writing tools. File-read tools (Read, Grep) aren't hooked, the same as before.
  • An agent started by hand in a plain terminal card, as in the incident, has no hook at all. For that agent the gateway reply in part 2 is the guidance it gets.

Tests

  • tests/base-protection-app-private.test.mjs, table-driven, with a fake HOME and a fake userData folder:
    • 38 deny cases. cat, head, grep, cp, base64 <, xxd, strings, sqlite3; globs; cd then a relative read; $(cat …); bash -c; find, grep -r, rg and tar of the app folder; python3 -c, node -e and heredoc readers; a Python path built from pieces; curl --unix-socket (both spellings), nc -U, socat UNIX-CONNECT:, a Python AF_UNIX connect, a runtime socket in userData; the descriptor and capability variables.
    • 17 allow look-alikes. Same-named files inside the project, grep -rn plugin-secrets src, a commit message that mentions them, the app's settings.json, ls of the app folder, the control CLI with its descriptor, the Docker socket and other sockets, ls $TMPDIR, /tmp/ctty-notes, and a URL containing agent-control/token-1.
    • Beyond the tables. File tools are refused and the agent's own account home is allowed. The socket folders are known without userData. The message text has no paths. DecisionHooks forwards the private data.
  • tests/agent-control.test.mjs:
    • A guessed NDJSON request and garbage each get the guidance and a closed connection.
    • An HTTP request gets a 403 with a correct Content-Length and the same text.
    • The token-file permission test from part 3.
  • Full suite 1162/1163 passing, 1 skipped (--test-concurrency=2, fake HOME). Also 1170 passing on this branch merged with fix(startup): load the application surface only after the startup page settled #96 and fix(settings): pasting a provider API key no longer blacks out the window #97.
  • Typecheck, npm run build and audit:secrets pass.

…e settled

Startup could fail with "ERR_ABORTED (-3) loading 'data:text/html...'"
when the machine was busy. Since services start while the startup page
loads, the application surface could start its navigation while the
page was still committing or loading in its own renderer process. The
application surface then committed first, and the page's ERR_ABORTED
arrived afterwards, while the loadFile promise was still waiting.
Electron's loadURL/loadFile promise takes the first main-frame
did-fail-load it sees as its own, so the application load rejected with
the startup page's abort and startup showed the failure page.

Services still start while the page loads. The application surface now
waits until the page load has settled, which it usually has by the time
services are up. A close during the page load is still a quiet quit, and
a real page error on a live window still fails startup.

Measured with 40 sequential hidden launches per build, each helper
process (renderer, GPU, utility) paused at random for 20 to 300 ms
during the first navigations: origin/main failed 4 and 6 of 40, the
same loop with this change 0 of 80. The build before
"perf(startup): start services while the startup page loads" had 0 of
40. Without pauses all builds start 40 of 40; the median time to a
ready window goes from 342 to 379 ms.
Pasting a MiniMax API key into Settings -> Agents -> Provider API keys
turned the window black. The key field's onChange read
event.currentTarget.value inside the setDrafts((current) => ...) updater.
React runs that updater later, during render, whenever an earlier update
of the component is still pending (a paste over a character already in
the field, the second keystroke of fast typing). By then the event has
finished dispatching and currentTarget is null, so the render threw
"Cannot read properties of null (reading 'value')" and React unmounted
the whole root. The renderer process stayed alive with an empty #root.
Key length and shape are not the cause: a 120-character key pasted over
a pending edit breaks the same way.

The handler now reads the value while the event dispatches and passes
it to the updater. No other renderer updater reads an event (checked
with an AST scan over src/renderer).

Tests: tests/provider-secrets-settings.test.mjs bundles the real
component against a React stand-in whose updaters run after dispatch
and pastes fake sk-cp-/sk-api- keys of 120 to 10 000 characters, with a
trailing newline and space, over a pending edit (fails before this
change with the same TypeError); a line-level guard keeps event reads
out of state updaters in the renderer.
An error thrown while React renders unmounts the whole root, but the
renderer process lives on, so render-process-gone never fires and the
main process's crash reload never runs: the window stays black until
the app is restarted. That is what the broken provider key field did.

The React root now handles onUncaughtError. The first such error logs
it with its component stack and reloads the application surface in
place; sessions live in the main process and survive the reload, as
after a renderer crash. A second one within 30 s (or with no session
storage to remember the reload) shows a static page with a Reload
button instead of reloading in a loop.

Tests: tests/uncaught-error-recovery.test.mjs (reload, cooldown, clock
change, no storage, root wiring). Checked in a hidden app: an injected
render error reloads to a usable Settings screen, a second one within
the cooldown shows the recovery page; forcefullyCrashRenderer still
reloads through render-process-gone.
…ores and sockets

Base protection now treats CanvasTTY's private data as credentials. A shell
or file tool call that names the agent-control token or descriptor, the
gateways' connection records, the provider and plugin secret stores,
account homes, the GitHub sign-in or prepared launch runs (paths taken from
the app's own userData folder, passed in by the app), or the control and
runtime socket folders under the temporary folder, is refused whatever the
program: readers, copies, encoders, sqlite3, recursive walks of the app
folder, interpreter one-liners and heredocs, curl --unix-socket, nc -U,
socat and Python sockets. The model is told calmly that agents cannot
control CanvasTTY this way and to ask for an Orchestrator launch.

The project, the app's settings, other sockets, an agent's own account
home and the bundled control CLI are unaffected.
…hen close

An unauthenticated or malformed request to the control endpoint now gets a
stable INVALID_REQUEST whose message says only sessions CanvasTTY launched
as orchestrators may use it and how to get one, with no protocol details,
token names or paths. An HTTP request line (curl, a browser) gets a
minimal 403 with the same text. Either way the connection is closed right
after, instead of waiting for the idle timeout; the connection cap and the
one-request-per-connection rule are unchanged.

Tests cover a guessed NDJSON request, garbage and an HTTP request, and that
the token file is 0600 and its folder 0700 even under umask 0 and a
pre-existing loose folder.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants