From a5f2038c5c9ca7722fb2bba2af335a77b56d30e2 Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sat, 26 Sep 2026 19:41:33 -0700 Subject: [PATCH 01/48] docs(plan): measure transcript and catalog work before moving it off the main thread (#769) Co-Authored-By: Claude Opus 5.5 --- ...026-09-26-monitor-transcript-operations.md | 22 +++++++++++++++++++ 1 file changed, 22 insertions(+) create mode 100644 docs/plans/2026-09-26-monitor-transcript-operations.md diff --git a/docs/plans/2026-09-26-monitor-transcript-operations.md b/docs/plans/2026-09-26-monitor-transcript-operations.md new file mode 100644 index 000000000..af2c3abb4 --- /dev/null +++ b/docs/plans/2026-09-26-monitor-transcript-operations.md @@ -0,0 +1,22 @@ +# Measure transcript and catalog work before moving it off the main thread (#769, first step) + +## Problem +#769 asks to move history loads (and the since-replaced session index) off the main thread. A re-check on origin/main `b3a2c481` (comment on #769) found that every transcript read and parse still runs on main. But the recordings cannot say whether that work is what stalls main: +- **Few slow reads:** 2,436 `transcript.read` samples across all monitor runs, only 3 slow-operation incidents (1.0 s, 1.3 s, 1.8 s). +- **No link to stalls:** 14 `main-stall` incidents, none within 10 s of a slow read. Main-stall incidents carry no operation attribution. +- **Double sample:** each IPC initial history load records `transcript.read` twice. `loadInitialHistoryChunk` opens a span, ends it after path resolution (`result: 'delegated'`), then `loadInitialHistoryChunkFromFile` opens a second span with the same name for the real read. Half the samples measure path resolution only, which skews the percentiles low. +- **Blind spot:** the conversations catalog is invisible to the monitor. `sessionIndex.extractPrompts` (a synchronous parse over windows of up to 16 MiB), catalog search's prompt gathering (up to 150 rows), and discovery have no monitor operation. + +## Decisions (defaults) +- **The path-resolution span gets its own name,** `historyLoader.resolveInitialPath`. It stays in the perf journal and is not a monitor operation, so each initial load records exactly one `transcript.read`. +- **Three monitor operations** join the finite vocabulary: + - `conversations.discover`: the existing `conversations.discover` span. + - `conversations.extract`: the existing `sessionIndex.extractPrompts` span. + - `conversations.search`: a new span around search's prompt gathering. +- **Slow-operation thresholds:** extract 250 ms (one file's synchronous parse), search 1000 ms, discover 2000 ms. Picked like the existing ones: a user-visible delay, not a precise budget. +- **Discovery failure closes its span.** Today a failed discovery leaves its span open until the 10-minute sweep records a `timeout`. +- **No worker yet.** Moving work to a worker waits until the data says which path stalls main. + +## Tests +- **History loader:** one initial load through `loadInitialHistoryChunk` records exactly one `transcript.read` (red on main: two). +- **Catalog:** a search with a query records `conversations.search`, and an extract records `conversations.extract`; a failed discovery records an `error` outcome. From 6de74ba13d23d95e8753dce1805c972013758980 Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sat, 26 Sep 2026 19:43:04 -0700 Subject: [PATCH 02/48] fix(performance): count one transcript read per history load and measure the conversations catalog #769 asks to move transcript work off the main thread, but the recordings cannot attribute main stalls to it: each initial load was sampled twice (half timing only path resolution), and catalog discovery, search and prompt extraction had no monitor boundary. Measure first. Refs #769 Co-Authored-By: Claude Opus 5.5 --- ...026-09-26-monitor-transcript-operations.md | 3 +- src/main/conversations/service.system.test.ts | 22 +++++++++++++++ src/main/conversations/service.ts | 15 ++++++++++ src/main/performance/IncidentEngine.ts | 3 ++ .../sessions/historyLoader.monitor.test.ts | 28 +++++++++++++++++++ src/main/sessions/historyLoader.ts | 7 ++++- src/shared/performance/monitorPolicy.ts | 7 +++++ src/shared/performance/operationTimers.ts | 3 ++ 8 files changed, 85 insertions(+), 3 deletions(-) create mode 100644 src/main/sessions/historyLoader.monitor.test.ts diff --git a/docs/plans/2026-09-26-monitor-transcript-operations.md b/docs/plans/2026-09-26-monitor-transcript-operations.md index af2c3abb4..45ebb6fad 100644 --- a/docs/plans/2026-09-26-monitor-transcript-operations.md +++ b/docs/plans/2026-09-26-monitor-transcript-operations.md @@ -14,9 +14,8 @@ - `conversations.extract`: the existing `sessionIndex.extractPrompts` span. - `conversations.search`: a new span around search's prompt gathering. - **Slow-operation thresholds:** extract 250 ms (one file's synchronous parse), search 1000 ms, discover 2000 ms. Picked like the existing ones: a user-visible delay, not a precise budget. -- **Discovery failure closes its span.** Today a failed discovery leaves its span open until the 10-minute sweep records a `timeout`. - **No worker yet.** Moving work to a worker waits until the data says which path stalls main. ## Tests - **History loader:** one initial load through `loadInitialHistoryChunk` records exactly one `transcript.read` (red on main: two). -- **Catalog:** a search with a query records `conversations.search`, and an extract records `conversations.extract`; a failed discovery records an `error` outcome. +- **Catalog:** a search with a query records `conversations.search`, and an extract records `conversations.extract`, driven on the recorded conversations corpus. diff --git a/src/main/conversations/service.system.test.ts b/src/main/conversations/service.system.test.ts index 31750b88c..e7d1b2a25 100644 --- a/src/main/conversations/service.system.test.ts +++ b/src/main/conversations/service.system.test.ts @@ -6,6 +6,7 @@ import { ClaudeConversationSource } from './sources/claude.js' import { CodexConversationSource } from './sources/codex.js' import { OpencodeConversationSource } from './sources/opencode.js' import { ConversationService } from './service.js' +import { setMainOperationSink } from '@main/performance/operations.js' import { corpusWorktreesPorcelain, installConversationCorpus, type InstalledCorpus } from '../../../testing/support/conversations/installCorpus.js' // One corpus install per file: the install copies 450 files and costs more @@ -67,4 +68,25 @@ describe('ConversationService', () => { expect(second.rows.map(r => r.nativeId)).toEqual(first.rows.map(r => r.nativeId)) expect(s.discoveriesForTests()).toBe(1) }) + + // #769 (measure first): the catalog's discovery, search and per-file prompt + // extraction had no monitor boundary, so nothing could tell whether they + // are what stalls main. Driven on the recorded corpus. + it('records discovery, search and prompt extraction as monitor operations', async () => { + const operations: Array<{ name: string; outcome: string }> = [] + setMainOperationSink(record => { operations.push(record) }) + try { + const s = service() + const all = await s.list({ cwd: '/fixture/repo', scope: 'repository', includeChildren: true, limit: 5000 }) + const codex = all.rows.find(r => r.provider === 'codex' && r.cwd)! + await s.prompts({ provider: 'codex', nativeId: codex.nativeId, cwd: codex.cwd! }) + await s.list({ cwd: '/fixture/repo', scope: 'repository', query: 'the', limit: 50 }) + } finally { + setMainOperationSink(() => {}) + } + const names = new Set(operations.map(op => op.name)) + expect(names).toContain('conversations.discover') + expect(names).toContain('conversations.search') + expect(names).toContain('conversations.extract') + }) }) diff --git a/src/main/conversations/service.ts b/src/main/conversations/service.ts index 374a096a4..94972e12f 100644 --- a/src/main/conversations/service.ts +++ b/src/main/conversations/service.ts @@ -118,6 +118,21 @@ export class ConversationService { } private async promptTextsFor(rows: readonly Conversation[]): Promise> { + // Search's prompt gathering is the catalog's widest synchronous parse + // (up to SEARCH_PROMPT_ROWS rows of extraction on main), so it is a + // monitor boundary of its own (#769). + const span = performanceService.span('conversations.search', { rows: rows.length }) + try { + const out = await this.gatherPromptTexts(rows) + span.end({ rows: out.size }) + return out + } catch (error) { + span.fail(error) + throw error + } + } + + private async gatherPromptTexts(rows: readonly Conversation[]): Promise> { const out = new Map() const candidates = [...rows].sort((a, b) => b.lastUserActivityAt - a.lastUserActivityAt).slice(0, SEARCH_PROMPT_ROWS) await Promise.all(candidates.map(async row => { diff --git a/src/main/performance/IncidentEngine.ts b/src/main/performance/IncidentEngine.ts index 1e84f6919..a8d71e866 100644 --- a/src/main/performance/IncidentEngine.ts +++ b/src/main/performance/IncidentEngine.ts @@ -8,6 +8,9 @@ const operationThresholds: Partial> = { 'transcript.read': 1000, 'transcript.parse': 100, 'transcript.fold': 100, 'transcript.commit': 100, 'terminal.write': 250, 'persistence.serialize': 100, 'persistence.write': 2000, 'worktree.refresh': 5000, + // One file's synchronous extraction, a search's prompt gathering (many + // rows, mostly I/O-bound), and a full discovery pass (#769). + 'conversations.extract': 250, 'conversations.search': 1000, 'conversations.discover': 2000, } // Each capture retains at most 160 evidence points. Eight simultaneous scopes // bound post-trigger memory to ~1.3k points even during an app-wide storm, diff --git a/src/main/sessions/historyLoader.monitor.test.ts b/src/main/sessions/historyLoader.monitor.test.ts new file mode 100644 index 000000000..97792efff --- /dev/null +++ b/src/main/sessions/historyLoader.monitor.test.ts @@ -0,0 +1,28 @@ +import { join } from 'node:path' + +import { afterEach, expect, it, vi } from 'vitest' + +// #769 (measure first): each IPC initial history load recorded +// `transcript.read` TWICE. The outer span ended after path resolution +// ("delegated") and the file read opened a second span with the same name, +// so half the monitor's samples timed only the path lookup and the +// percentiles read low. One load must be one sample. +const fixture = join( + import.meta.dirname, + '../../../testing/fixtures/conversations/claude/projects/-fixture-repo--worktrees-feat-api-key-vault/fc475787-6395-4cda-8bd2-1faacaa18bc7.jsonl', +) +vi.mock('@main/providerSwitch/shared.js', () => ({ resolveProviderTranscriptPath: vi.fn(async () => fixture) })) +vi.mock('@providers/registry.main.js', () => ({ getMainProvider: () => ({}) })) + +const { setMainOperationSink } = await import('@main/performance/operations.js') +const { loadInitialHistoryChunk } = await import('./historyLoader.js') + +afterEach(() => setMainOperationSink(() => {})) + +it('records one transcript.read per initial history load', async () => { + const operations: Array<{ name: string }> = [] + setMainOperationSink(record => { operations.push(record) }) + const chunk = await loadInitialHistoryChunk({ kind: 'claude', cwd: '/fixture/repo', providerSessionId: 'fc475787-6395-4cda-8bd2-1faacaa18bc7', limit: 20 } as never) + expect(chunk.entries.length).toBeGreaterThan(0) + expect(operations.filter(op => op.name === 'transcript.read')).toHaveLength(1) +}) diff --git a/src/main/sessions/historyLoader.ts b/src/main/sessions/historyLoader.ts index 458af40ed..c12bf24c9 100644 --- a/src/main/sessions/historyLoader.ts +++ b/src/main/sessions/historyLoader.ts @@ -792,7 +792,12 @@ export async function loadInitialHistoryChunk( providerSource({ cwd: params.cwd, providerSessionId: params.providerSessionId, limit: params.limit }), ) } - const span = performanceService.span('historyLoader.loadInitialChunk', { + // WHY its own name (#769): this span ends once the path is resolved, and + // loadInitialHistoryChunkFromFile opens the span that times the read. Under + // the same name the monitor counted every load twice, half of them timing + // only the path lookup, which pulled transcript.read's percentiles down. + // This one stays in the perf journal and is not a monitor operation. + const span = performanceService.span('historyLoader.resolveInitialPath', { kind: params.kind, limit: params.limit, }) diff --git a/src/shared/performance/monitorPolicy.ts b/src/shared/performance/monitorPolicy.ts index 7466f2169..efc6f792b 100644 --- a/src/shared/performance/monitorPolicy.ts +++ b/src/shared/performance/monitorPolicy.ts @@ -49,6 +49,13 @@ export const MONITOR_OPERATIONS = [ 'transcript.commit', 'terminal.write', 'worktree.refresh', + // The conversations catalog (#769, measure first): its prompt extraction + // parses windows of up to 16 MiB synchronously on main, and search can run + // it for 150 rows, yet it had no monitor boundary at all, so nothing could + // say whether it is what stalls main. + 'conversations.discover', + 'conversations.extract', + 'conversations.search', 'persistence.serialize', 'persistence.write', 'orchestration.queue', diff --git a/src/shared/performance/operationTimers.ts b/src/shared/performance/operationTimers.ts index 40747b405..320dcab39 100644 --- a/src/shared/performance/operationTimers.ts +++ b/src/shared/performance/operationTimers.ts @@ -48,6 +48,9 @@ export const LEGACY_MONITOR_OPERATIONS: Readonly Date: Sat, 26 Sep 2026 20:37:29 -0700 Subject: [PATCH 03/48] docs(plan): reclaim a Claude prompt Agent Code stranded in the native composer (#1350) Co-Authored-By: Claude Opus 5.5 --- ...26-09-27-claude-stranded-prompt-reclaim.md | 39 +++++++++++++++++++ 1 file changed, 39 insertions(+) create mode 100644 docs/plans/2026-09-27-claude-stranded-prompt-reclaim.md diff --git a/docs/plans/2026-09-27-claude-stranded-prompt-reclaim.md b/docs/plans/2026-09-27-claude-stranded-prompt-reclaim.md new file mode 100644 index 000000000..102d14f2b --- /dev/null +++ b/docs/plans/2026-09-27-claude-stranded-prompt-reclaim.md @@ -0,0 +1,39 @@ +# A Claude prompt Agent Code stranded in the native composer is reclaimed, not treated as a human draft (#1350) + +## Problem +A Claude delivery writes the prompt text, then waits up to 5 s for it to become visible (absorption). While Claude is busy mid-turn it may not paint the composer. Absorption then times out, and `rollbackWrittenPrompt` watches for about 0.4 s to see our bytes before it may kill them. If it never sees them it returns `unrecoverable` (`rollback-unobserved`), and the delivery reports `absorption-timeout` / `do-not-retry` / "composer could not be recovered". + +The bytes then paint. The gate (`claudeSession.derivePromptGateState`) reads any drafted composer as `occupied: human-draft`, deliberately with no timeout, so every later delivery is refused as "occupied by a human draft". The session is unreachable for orchestration until a human clears the composer. Nothing records that the "human draft" is our own stranded prompt. + +## Evidence (recorded) +**Control history** (`~/.config/agent-code/control-history`): +- `c7eb1822`: session e41b2b15, absorption-timeout, `promptWritten: true`, "could not be recovered". +- Four `ac_agents_prompt` calls (`a0a85f4e`, `6bc085b8`, `d2ceab7a`, `5f9b16e0`) to session a835d518 were refused as "occupied by a human draft" over about 50 minutes. `draftGet` showed an empty Agent Code draft; `inputInspect` showed `nativeDraft.state: occupied`. + +**Lifecycle journal** (`incidents/runs/.../events.jsonl`), three stranding incidents: +- **a835d518:** a 1,518-byte delivery write, no submit, at 02:04:05.9 (a goal-loop continuation; the loop paused with `consecutiveDeliveryFailures: 1`). Occupied from 02:04:11.5, never ready again. +- **e41b2b15:** a 180-byte write around 03:16:44.5; occupied at 03:16:53.9, 3.8 s after the failure returned. +- **3fbf42d1** (orchestration send-prompt): a 174-byte write at 02:01:29.9; failed 02:01:35.5; occupied at 02:01:36.2. + +In all three, the composer turned occupied 0.7–3.8 s after the delivery gave up: the bytes painted after the ~0.4 s observe window. + +**Paste-debug journals** (renderer-originated deliveries only): 7 `rollback-unobserved` out of 13 rollbacks. +- Plain (22–82 chars) and paste-like (145–380 chars) alike. +- Every one reached the unobserved verdict 5.6–5.7 s after the write (the 5 s absorption timeout plus the observe window). + +## Decision: ownership proven structurally, never by comparing text +`rollbackWrittenPrompt` records why text matching cannot prove ownership (#679): the screen is viewport-clipped and wrap-lossy, and paste-like input collapses to `[Pasted text #N]`. The proof it uses instead is structural: while a delivery holds the reservation, no other writer can reach the PTY. This change extends that proof across the gap between two deliveries. +- **The mark.** `SessionManager` keeps a per-session "stranded delivery" mark. It is set when a delivery ends with `promptWritten: true, enterWritten: false` and not ok: our bytes are, or may soon be, in the native composer. +- **What clears the mark.** Every PTY writer goes through `recordInputWrite`. Any write whose origin is not `delivery` clears it: raw terminal typing, remote input, dictation, a condition answer. From then on the composer may hold someone else's text. The mark also clears on process exit or cleanup and on a successful delivery. +- **Reclaiming.** The next delivery gets `strandedComposer: true` while the mark stands. If Claude's gate then says `occupied`, the delivery holds the reservation, so nothing else can write, and everything in the composer is ours. It clears the composer with the existing verified kill loop: one Ctrl+U per PTY read, stopping on a verified-empty read, yanking back if it runs out. Then it re-awaits readiness and continues normally. +- **Fail closed.** If the kill loop cannot verify an empty composer, it yanks the text back, as the rollback does, and the delivery is refused with a message naming the stranded prompt rather than "a human draft". +- **Inspection.** `sessions.inputInspect`, via `ac_agents_input_inspect`, reports `nativeDraft.strandedDelivery: true` while the mark stands, so a caller can tell our stranded write from a human draft. +- **Out of scope.** The image-pill path still does not roll back (documented in `promptDelivery.ts`), and nothing is recovered for sessions stranded before this change. + +## Tests +- **Delivery:** a stranded mark plus an occupied composer that clears under the kill loop leads to a successful delivery. Without the mark, the same state is refused as a human draft, as today. A composer the loop cannot clear is yanked back and refused with the stranded message. +- **SessionManager:** + - A failed delivery with `promptWritten && !enterWritten` sets the mark, and the next delivery receives `strandedComposer: true`. + - A raw `write()` in between clears it, as does exit. + - `inputInspect` reports it. +- All red on main. From c41614b474bf2e44169fb71996d839ace95d4bbe Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sat, 26 Sep 2026 20:49:41 -0700 Subject: [PATCH 04/48] fix(claude): reclaim a prompt Agent Code stranded in the native composer instead of refusing it as a human draft A delivery whose bytes painted after its rollback stopped watching left them in Claude's composer, and the gate then refused every later delivery as a human draft. SessionManager now marks a session stranded while no other writer has touched the PTY since; the next delivery clears the composer under its reservation and continues. inputInspect reports the mark. Fixes #1350 Co-Authored-By: Claude Opus 5.5 --- src/control-sdk/catalog/input.ts | 6 +- .../sessionManager.strandedDelivery.test.ts | 99 +++++++++++++++++++ src/main/sessionManager.ts | 44 +++++++++ src/main/sessions/terminalControl.ts | 5 +- .../promptDelivery.strandedReclaim.test.ts | 89 +++++++++++++++++ .../claude/runtime/promptDelivery.ts | 36 ++++++- src/shared/types/providerConfig.ts | 9 ++ 7 files changed, 284 insertions(+), 4 deletions(-) create mode 100644 src/main/sessionManager.strandedDelivery.test.ts create mode 100644 src/providers/claude/runtime/promptDelivery.strandedReclaim.test.ts diff --git a/src/control-sdk/catalog/input.ts b/src/control-sdk/catalog/input.ts index 5f235cca6..500ace883 100644 --- a/src/control-sdk/catalog/input.ts +++ b/src/control-sdk/catalog/input.ts @@ -4,5 +4,9 @@ import { z } from 'zod' // No current provider port proves the full native draft, so absence of a probe // must remain unknown, never an empty string inferred from an xterm textarea. export const nativeInputOutput = z.object({ sessionId: z.string(), sessionRunId: z.string().nullable(), - backendPresent: z.boolean(), nativeDraft: z.object({ state: z.enum(['unknown', 'occupied']), text: z.null(), reason: z.string() }), + backendPresent: z.boolean(), nativeDraft: z.object({ state: z.enum(['unknown', 'occupied']), text: z.null(), reason: z.string(), + // #1350: true while the composer can hold only Agent Code's own stranded + // delivery write, which the next delivery clears. Distinguishes our + // leftover from a human draft for a caller that sees `occupied`. + strandedDelivery: z.boolean().default(false) }), inputReady: z.boolean().nullable(), readinessReason: z.string().nullable().default(null) }) diff --git a/src/main/sessionManager.strandedDelivery.test.ts b/src/main/sessionManager.strandedDelivery.test.ts new file mode 100644 index 000000000..ad5b6c2ca --- /dev/null +++ b/src/main/sessionManager.strandedDelivery.test.ts @@ -0,0 +1,99 @@ +import { afterEach, expect, it, vi } from 'vitest' + +import { SessionManager } from './sessionManager.js' + +// #1350, end to end through the manager and the real Claude delivery code. +// Recorded shape (lifecycle journal, three incidents on 2026-09-27): a +// delivery writes its prompt while Claude is busy, the text does not paint +// within absorption (5 s) plus the rollback's observe window, the delivery +// gives up "could not be recovered", and the text paints 0.7-3.8 s later. +// Claude's gate then reads it as a human draft for good. +afterEach(() => { vi.useRealTimers() }) + +const RULE = '─'.repeat(40) +const composer = (row: string) => [RULE, row, RULE].join('\n') + +function claudeLike() { + const writes: string[] = [] + // 'busy': our text is in Claude's buffer but not painted (the incident); + // 'stranded': it painted late; 'empty'; 'prompted': the next prompt shows. + let state: 'idle' | 'busy' | 'stranded' | 'empty' | 'prompted' = 'idle' + const frame = () => state === 'stranded' + ? { screen: composer('❯ an earlier prompt that painted late'), attributes: { dim: 0, inverse: 1, plain: 33 } } + : state === 'prompted' + ? { screen: composer('❯ the next task'), attributes: { dim: 0, inverse: 1, plain: 13 } } + : { screen: composer('❯'), attributes: null } + const session = { + isExited: () => false, + write: (data: string) => { + writes.push(data) + if (data === 'an earlier prompt that painted late') state = 'busy' + else if (data === '\x15' && state === 'stranded') state = 'empty' + else if (data === 'the next task') state = 'prompted' + else if (data === '\r') state = 'idle' + else if (state === 'stranded' || state === 'idle') state = 'stranded' + }, + snapshotScreen: () => frame().screen, + readComposer: () => frame(), + awaitReadyForPrompt: async () => state === 'stranded' + ? { kind: 'occupied' as const, reason: 'human-draft' as const, waitedMs: 0 } + : { kind: 'ready' as const, waitedMs: 0 }, + armPromptAcceptance: () => ({ promise: Promise.resolve({ kind: 'user' as const, acceptedAt: 1 }), cancel: vi.fn() }), + paintLate: () => { state = 'stranded' }, + } + const manager = new SessionManager() + ;(manager as unknown as { sessions: Map }).sessions.set('s1', { kind: 'claude', session }) + return { manager, session, writes } +} + +async function strand(manager: SessionManager) { + vi.useFakeTimers() + const first = manager.deliverPromptToAgent('s1', 'an earlier prompt that painted late') + await vi.advanceTimersByTimeAsync(8_000) + const result = await first + vi.useRealTimers() + return result +} + +it('reclaims its own stranded prompt on the next delivery', async () => { + const { manager, session, writes } = claudeLike() + expect(await strand(manager)).toMatchObject({ ok: false, code: 'absorption-timeout', promptWritten: true, enterWritten: false }) + expect(manager.hasStrandedDelivery('s1')).toBe(true) + session.paintLate() + + await expect(manager.deliverPromptToAgent('s1', 'the next task')).resolves.toMatchObject({ ok: true }) + expect(writes.slice(1)).toEqual(['\x15', 'the next task', '\r']) + expect(manager.hasStrandedDelivery('s1')).toBe(false) +}) + +it('treats the composer as a human draft again once anyone else has written to it', async () => { + const { manager, session, writes } = claudeLike() + await strand(manager) + session.paintLate() + // Raw terminal typing reaches the same composer; the text is no longer + // provably ours alone. + expect(manager.write('s1', 'x')).toBe(true) + expect(manager.hasStrandedDelivery('s1')).toBe(false) + + const result = await manager.deliverPromptToAgent('s1', 'the next task') + expect(result).toMatchObject({ ok: false, code: 'not-ready', promptWritten: false }) + expect(result.ok ? '' : result.message).toContain('occupied by a human draft') + expect(writes).not.toContain('\x15') +}) + +it('reports our own stranded write to input inspection', async () => { + const { terminalBackendCapabilities } = await import('@main/sessions/terminalControl.js') + const { manager } = claudeLike() + await strand(manager) + vi.spyOn(manager, 'getBackendSnapshot').mockReturnValue({ + sessionId: 's1', sessionRunId: 'run-1', kind: 'claude', cwd: '/repo', lifecycle: 'live', + input: { ready: false, reason: 'composer-occupied', revision: 2 }, + } as never) + const context = { requestId: 'inspect', caller: { kind: 'application' as const, id: 'renderer' }, owner: { kind: 'main' as const, generation: 'one' } } + const inspect = terminalBackendCapabilities(manager).find(item => item.descriptor.id === 'sessions.inputInspect')! + const result = await inspect.execute({ sessionId: 's1', cwd: '/repo', provider: 'claude' }, context) + if (!result.ok) throw new Error(JSON.stringify(result)) + const { nativeDraft } = result.value as { nativeDraft: { state: string; strandedDelivery: boolean; reason: string } } + expect(nativeDraft).toMatchObject({ state: 'occupied', strandedDelivery: true }) + expect(nativeDraft.reason).toContain('earlier Agent Code prompt') +}) diff --git a/src/main/sessionManager.ts b/src/main/sessionManager.ts index 3d463350c..429620978 100644 --- a/src/main/sessionManager.ts +++ b/src/main/sessionManager.ts @@ -587,6 +587,24 @@ export class SessionManager extends EventEmitter { // per-session critical section; without it, Enter from attempt A can submit // paste B and turn a slow operation into duplicate queue entries. private readonly promptDeliveriesInFlight = new Set() + /** + * Sessions whose native composer can hold only this app's own stranded + * write (#1350). Set when a delivery wrote prompt bytes it could not submit + * (promptWritten && !enterWritten, not ok): the bytes are, or may soon be, + * in the composer. Claude in particular can paint them seconds after the + * delivery's rollback stopped watching, and its gate then reads them as a + * human draft forever, so no later delivery could reach the session. + * + * WHY this proves ownership: every PTY writer goes through recordInputWrite, + * and any write that is not a delivery's clears the mark (raw terminal + * typing, remote input, a condition answer). While the mark stands, nothing + * but deliveries has written since, and each later delivery refused before + * writing. So the composer holds our bytes and nothing else, and the next + * delivery may clear it under its reservation (PromptDeliveryIo + * .strandedComposer). A text comparison could not prove this: the screen is + * clipped and wrap-lossy, and pastes collapse to `[Pasted text #N]`. + */ + private readonly strandedDeliveries = new Set() /** * Prompts waiting for a session that is not ready for one YET (#854). * @@ -1090,6 +1108,8 @@ export class SessionManager extends EventEmitter { this.codexCandidateObservationEdges.delete(sessionId) this.codexAttachmentObservationState.delete(sessionId) this.lastInputReadiness.delete(sessionId) + // A new process has a new, empty composer (#1350). + this.strandedDeliveries.delete(sessionId) this.lastGateEvaluation.delete(sessionId) this.spawnInfo.delete(sessionId) // Keep lastActivityAt after removal. Process telemetry can be asked about a @@ -3913,6 +3933,10 @@ export class SessionManager extends EventEmitter { data: string, origin: InputWriteOrigin, ): void { + // Before the journal check: the stranded-delivery proof must hold whether + // or not a lifecycle journal is attached. Any writer that is not a + // delivery may have put its own text in the composer (#1350). + if (origin !== 'delivery') this.strandedDeliveries.delete(sessionId) if (!this.lifecycle) return const now = Date.now() const pending = this.inputWriteCoalesce.get(sessionId) @@ -4056,6 +4080,23 @@ export class SessionManager extends EventEmitter { return true } + /** See strandedDeliveries (#1350). */ + private noteStrandedDelivery( + sessionId: string, + delivery: { ok: boolean; promptWritten?: boolean; enterWritten?: boolean }, + ): void { + if (delivery.ok) this.strandedDeliveries.delete(sessionId) + else if (delivery.promptWritten && !delivery.enterWritten) this.strandedDeliveries.add(sessionId) + // A failure that wrote nothing (refused before write) changes nothing: + // whatever the composer held before still holds, ours or not. + } + + /** Whether the session's composer can only hold our own stranded write + * (#1350); reported by sessions.inputInspect. */ + hasStrandedDelivery(sessionId: string): boolean { + return this.strandedDeliveries.has(sessionId) + } + getSessionKind(sessionId: string): SessionKind | null { return this.sessions.get(sessionId)?.kind ?? null } @@ -4952,7 +4993,9 @@ export class SessionManager extends EventEmitter { imagePaths, record, ...(options?.requireEmptyNativeComposer ? { requireEmptyNativeComposer: true } : {}), + ...(this.strandedDeliveries.has(sessionId) ? { strandedComposer: true } : {}), }) + this.noteStrandedDelivery(sessionId, delivery) finishDelivery(delivery.ok ? 'success' : 'error') // Instrumentation must never change a delivery outcome. `acceptance` is // read defensively because a provider result without it made @@ -4967,6 +5010,7 @@ export class SessionManager extends EventEmitter { finishDelivery('error') this.monitorResponses.cancel(sessionId) record?.('uncertain', { reason: 'provider-threw' }) + this.noteStrandedDelivery(sessionId, { ok: false, promptWritten, enterWritten }) return { ok: false, stage: enterWritten ? 'after-enter' : promptWritten ? 'absorption' : 'before-write', diff --git a/src/main/sessions/terminalControl.ts b/src/main/sessions/terminalControl.ts index ae4b5dc37..8c7da9104 100644 --- a/src/main/sessions/terminalControl.ts +++ b/src/main/sessions/terminalControl.ts @@ -6,7 +6,7 @@ import type { SessionManager } from '@main/sessionManager' const ownership = { cwd: z.string(), provider: z.string() } type FrozenReplay = { sessionId: string; sessionRunId: string; range: string; raw: string; capChars: number; revision: string; expires: number } -export function terminalBackendCapabilities(manager: Pick) { +export function terminalBackendCapabilities(manager: Pick & Partial>) { const snapshots = new Map() let timer: ReturnType | undefined const prune = () => { @@ -26,11 +26,12 @@ export function terminalBackendCapabilities(manager: Pick { const backend = manager.getBackendSnapshot(input.sessionId) if (backend && (backend.cwd !== input.cwd || backend.kind !== input.provider)) throw new ControlError('unavailable', 'Backend identity changed') + const stranded = Boolean(backend) && manager.hasStrandedDelivery?.(input.sessionId) === true return { sessionId: input.sessionId, sessionRunId: backend?.sessionRunId ?? null, backendPresent: Boolean(backend), inputReady: backend?.input.ready ?? null, readinessReason: backend?.input.reason ?? null, // composer-occupied is provider-owned input-readiness evidence, // not text guessed from a terminal accessibility field. Ready alone // still does not prove complete native draft emptiness. - nativeDraft: { state: backend?.input.reason === 'composer-occupied' ? 'occupied' as const : 'unknown' as const, text: null, reason: backend?.input.reason === 'composer-occupied' ? 'The provider reports an occupied composer. Resolve it through its UI; complete draft text is not exposed.' : 'The provider port does not expose a complete native composer snapshot. Readiness and the terminal accessibility input value do not prove an empty draft. agents.draftGet reads only the separate Agent Code draft.' } } + nativeDraft: { state: backend?.input.reason === 'composer-occupied' ? 'occupied' as const : 'unknown' as const, text: null, strandedDelivery: stranded, reason: backend?.input.reason === 'composer-occupied' ? (stranded ? 'The composer holds an earlier Agent Code prompt that could not be submitted; the next delivery clears it.' : 'The provider reports an occupied composer. Resolve it through its UI; complete draft text is not exposed.') : 'The provider port does not expose a complete native composer snapshot. Readiness and the terminal accessibility input value do not prove an empty draft. agents.draftGet reads only the separate Agent Code draft.' } } }, }), defineCapability({ diff --git a/src/providers/claude/runtime/promptDelivery.strandedReclaim.test.ts b/src/providers/claude/runtime/promptDelivery.strandedReclaim.test.ts new file mode 100644 index 000000000..91c83d663 --- /dev/null +++ b/src/providers/claude/runtime/promptDelivery.strandedReclaim.test.ts @@ -0,0 +1,89 @@ +import { afterEach, expect, it, vi } from 'vitest' + +import { deliverClaudePrompt } from './promptDelivery.js' +import type { PromptDeliveryIo } from '@shared/types/providerConfig.js' + +// #1350: a delivery whose absorption timed out while Claude was busy wrote its +// bytes, could not SEE them within the rollback's observe window, and gave up +// ("composer could not be recovered"). The bytes painted 0.7-3.8 s later in all +// three recorded incidents, and the gate then read them as a human draft for +// good: every later delivery was refused as "occupied by a human draft". +// SessionManager now tells the next delivery when the composer can only hold +// our own stranded write (no other writer has touched the PTY since); that +// delivery clears it under its reservation and goes on. +afterEach(() => { vi.useRealTimers() }) + +const RULE = '─'.repeat(40) +const composer = (row: string) => [RULE, row, RULE].join('\n') +type Attributes = { dim: number; inverse: number; plain: number } +type Frame = { screen: string; attributes: Attributes | null } +const STRANDED: Frame = { screen: composer('❯ an earlier prompt that painted late'), attributes: { dim: 0, inverse: 1, plain: 33 } } +const EMPTY: Frame = { screen: composer('❯'), attributes: null } + +function harness(opts: { strandedComposer?: boolean; killClears?: boolean }) { + const writes: string[] = [] + const records: string[] = [] + let state: 'stranded' | 'empty' | 'prompted' = 'stranded' + const frame = (): Frame => state === 'stranded' + ? STRANDED + : state === 'empty' ? EMPTY : { screen: composer('❯ send the next task'), attributes: { dim: 0, inverse: 1, plain: 19 } } + const io = { + sessionId: 's1', + prompt: 'send the next task', + ...(opts.strandedComposer ? { strandedComposer: true } : {}), + write: (data: string) => { + writes.push(data) + if (data === '\x15' && opts.killClears !== false) state = 'empty' + else if (data === '\x19') state = 'stranded' + else if (data === 'send the next task') state = 'prompted' + return true + }, + record: (event: string) => { records.push(event) }, + session: { + snapshotScreen: () => frame().screen, + readComposer: () => frame(), + // The gate as claudeSession derives it: any drafted composer is occupied. + awaitReadyForPrompt: vi.fn(async () => state === 'stranded' + ? { kind: 'occupied' as const, reason: 'human-draft' as const, waitedMs: 0 } + : { kind: 'ready' as const, waitedMs: 0 }), + armPromptAcceptance: () => ({ + promise: Promise.resolve({ kind: 'user' as const, acceptedAt: 1 }), + cancel: vi.fn(), + }), + }, + } as unknown as PromptDeliveryIo + return { io, writes, records } +} + +it('clears our own stranded prompt under the reservation, then delivers', async () => { + vi.useFakeTimers() + const { io, writes, records } = harness({ strandedComposer: true }) + const delivery = deliverClaudePrompt(io) + await vi.advanceTimersByTimeAsync(10_000) + await expect(delivery).resolves.toMatchObject({ ok: true }) + expect(writes).toEqual(['\x15', 'send the next task', '\r']) + expect(records).toContain('stranded-reclaimed') +}) + +it('still refuses a drafted composer it cannot prove is its own', async () => { + const { io, writes } = harness({}) + const result = await deliverClaudePrompt(io) + expect(result).toMatchObject({ ok: false, code: 'not-ready', promptWritten: false }) + expect(result.ok ? '' : result.message).toContain('occupied by a human draft') + expect(writes).toEqual([]) +}) + +it('puts the stranded prompt back and says so when it cannot clear it', async () => { + vi.useFakeTimers() + const { io, writes } = harness({ strandedComposer: true, killClears: false }) + const delivery = deliverClaudePrompt(io) + await vi.advanceTimersByTimeAsync(10_000) + const result = await delivery + expect(result).toMatchObject({ ok: false, code: 'not-ready', promptWritten: false, retrySafe: true }) + expect(result.ok ? '' : result.message).toContain('earlier prompt from Agent Code') + // Bounded kills, one yank, and never the new prompt or an Enter. + expect(writes.filter(data => data === '\x15').length).toBe(64) + expect(writes.at(-1)).toBe('\x19') + expect(writes).not.toContain('send the next task') + expect(writes).not.toContain('\r') +}) diff --git a/src/providers/claude/runtime/promptDelivery.ts b/src/providers/claude/runtime/promptDelivery.ts index 4b679b5ee..2271f4f1a 100644 --- a/src/providers/claude/runtime/promptDelivery.ts +++ b/src/providers/claude/runtime/promptDelivery.ts @@ -70,9 +70,29 @@ export async function deliverClaudePrompt( }) } if (typeof io.session.awaitReadyForPrompt === 'function') { - const ready = await io.session.awaitReadyForPrompt({ + const awaitReady = () => io.session.awaitReadyForPrompt!({ deadlineAt: Math.min(deliveryDeadlineAt, Date.now() + READY_BUDGET_MS), }) + let ready = await awaitReady() + if (ready.kind === 'occupied' && io.strandedComposer) { + // The "human draft" is our own earlier write (#1350): an earlier + // delivery's bytes painted after its rollback stopped watching, and no + // other writer has reached the PTY since (SessionManager's proof, see + // PromptDeliveryIo.strandedComposer). This delivery holds the + // reservation, so nothing can add to the composer while we clear it: + // the same structural ownership proof rollbackWrittenPrompt relies on, + // carried across the gap between the two deliveries. + const reclaimed = await killComposerToEmpty(io) + if (reclaimed !== 'cleared') { + return failure({ + stage: 'before-write', code: 'not-ready', retrySafe: true, disposition: 'retry-after-resolve', + promptWritten: false, enterWritten: false, + message: `Claude session ${io.sessionId} still has an earlier prompt from Agent Code in its composer that could not be cleared; clear it there before sending again`, + }) + } + io.record?.('stranded-reclaimed') + ready = await awaitReady() + } if (ready.kind !== 'ready') { const disposition = ready.kind === 'timeout' ? 'retry-same-session' as const @@ -487,6 +507,20 @@ async function rollbackWrittenPrompt( return 'unrecoverable' } + return killComposerToEmpty(io) +} + +/** + * Clear a composer known to hold only our bytes, verified by reading it back + * after each press. Used by the rollback above (bytes this delivery just + * wrote) and by a delivery reclaiming an earlier delivery's stranded write + * (#1350); both hold the reservation, which is what makes "only ours" true. + */ +async function killComposerToEmpty( + io: PromptDeliveryIo, +): Promise<'cleared' | 'restored' | 'unrecoverable'> { + const readComposer = (): 'empty' | 'drafted' | 'unpainted' => + classifyRollbackComposer(io.session.readComposer?.() ?? null, io.session.snapshotScreen?.() ?? '') // STEP 2 — clear, one keypress at a time. // // WHY each press must arrive alone: Claude's input tokeniser accumulates a run diff --git a/src/shared/types/providerConfig.ts b/src/shared/types/providerConfig.ts index b2c95e540..cf0b53fcf 100644 --- a/src/shared/types/providerConfig.ts +++ b/src/shared/types/providerConfig.ts @@ -436,6 +436,15 @@ export type PromptDeliveryIo = PromptDeliveryOptions & { imagePaths?: string[] /** Non-blocking forensic sink. Correctness must never await or depend on it. */ record?: (event: string, data?: Record) => void + /** + * The native composer can only hold THIS app's own stranded write (#1350). + * Set by SessionManager when an earlier delivery wrote prompt bytes it could + * not submit and no other writer has touched the PTY since. A provider may + * then clear an occupied composer under this delivery's reservation instead + * of refusing it as a human draft. Absent or false: an occupied composer is + * someone else's, exactly as before. + */ + strandedComposer?: boolean } export type PromptAcceptance = From e504a635c50aed62431ac514b84207d4da3ceb80 Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sat, 26 Sep 2026 22:15:49 -0700 Subject: [PATCH 05/48] fix(claude): bind a stranded-delivery mark to its process, wait out the late paint, never reclaim image strands Round-1 review of #1358 (steering q75): a late failure of a replaced process could mark the same-id successor and clear its human draft; a delivery inside the recorded 0.7-3.8 s paint lag skipped the reclaim; image pills are not proven safe to kill. Refs #1350 Co-Authored-By: Claude Opus 5.5 --- ...26-09-27-claude-stranded-prompt-reclaim.md | 9 +++ .../sessionManager.strandedDelivery.test.ts | 75 +++++++++++++++++++ src/main/sessionManager.ts | 47 ++++++++++-- .../promptDelivery.strandedReclaim.test.ts | 2 +- .../claude/runtime/promptDelivery.ts | 21 ++++++ src/shared/types/providerConfig.ts | 8 +- 6 files changed, 151 insertions(+), 11 deletions(-) diff --git a/docs/plans/2026-09-27-claude-stranded-prompt-reclaim.md b/docs/plans/2026-09-27-claude-stranded-prompt-reclaim.md index 102d14f2b..e1113a220 100644 --- a/docs/plans/2026-09-27-claude-stranded-prompt-reclaim.md +++ b/docs/plans/2026-09-27-claude-stranded-prompt-reclaim.md @@ -37,3 +37,12 @@ In all three, the composer turned occupied 0.7–3.8 s after the delivery gave u - A raw `write()` in between clears it, as does exit. - `inputInspect` reports it. - All red on main. + +## Round 1 review decisions (#1358) +- **a (blocker, steering q75) and c: the mark was keyed by session id only.** A delivery to process A that failed after A exited, and B had taken the same id, marked B. The next delivery then cleared B's composer, possibly a human draft. The mark now names the registry entry the delivery captured. It is set only while that entry still owns the id, and honoured only for that same entry. Test: A exits, B holds a human draft, A fails late. B gets no mark and no kill. Red on `c41614b4`. +- **c: a delivery inside the paint lag skipped the reclaim.** The recorded paint lag is 0.7–3.8 s after the failure. A delivery starting inside it saw an empty, ready composer and wrote next to the late-painting bytes. While the mark stands, the delivery now polls, still holding its reservation, for up to 8 s after the strand. If the text appears it is reclaimed; otherwise nothing of ours is there. Test: the text paints 2 s into the next delivery and is cleared first. Red on `c41614b4`. +- **a and c: image deliveries.** Whether Ctrl+U removes an image pill is unobserved, so an image delivery that strands is not marked and never reclaimed. Test added. #1350 is narrowed to text strands; the image path is a linked follow-up issue. +- **Surviving mutants:** a thrown write that strands, and the exit clearing the mark, now have a test. +- **Suspicions kept as residuals:** + - headless-internal writers (trust, resume, permission) that bypass `recordInputWrite`: no app consumer was found; + - a mark set after a throw before any bytes crossed: harmless, since the next gate reads ready and nothing is killed. diff --git a/src/main/sessionManager.strandedDelivery.test.ts b/src/main/sessionManager.strandedDelivery.test.ts index ad5b6c2ca..faab33aab 100644 --- a/src/main/sessionManager.strandedDelivery.test.ts +++ b/src/main/sessionManager.strandedDelivery.test.ts @@ -97,3 +97,78 @@ it('reports our own stranded write to input inspection', async () => { expect(nativeDraft).toMatchObject({ state: 'occupied', strandedDelivery: true }) expect(nativeDraft.reason).toContain('earlier Agent Code prompt') }) + +// Steering q75 / #1358 review a (blocker) and c: the mark belongs to the +// PROCESS whose composer holds the bytes. Sequence: a delivery writes to +// process A; A exits and B takes the same session id; a human drafts in B; +// A's delivery then fails. The late failure must neither mark B nor let the +// next delivery clear B's composer. +it('never lets a late failure of a replaced process clear the new process\'s human draft', async () => { + vi.useFakeTimers() + const { manager } = claudeLike() + const sessions = (manager as unknown as { sessions: Map }).sessions + const first = manager.deliverPromptToAgent('s1', 'an earlier prompt that painted late') + await vi.advanceTimersByTimeAsync(100) + // A exits (the manager's own cleanup) and B, same id, holds a human draft. + ;(manager as unknown as { cleanupSessionState(id: string, kind: string): void }).cleanupSessionState('s1', 'claude') + const bWrites: string[] = [] + const humanDraft = { screen: composer('❯ a human typed this'), attributes: { dim: 0, inverse: 1, plain: 18 } } + sessions.set('s1', { kind: 'claude', session: { + isExited: () => false, + write: (data: string) => { bWrites.push(data) }, + snapshotScreen: () => humanDraft.screen, + readComposer: () => humanDraft, + awaitReadyForPrompt: async () => ({ kind: 'occupied' as const, reason: 'human-draft' as const, waitedMs: 0 }), + armPromptAcceptance: () => ({ promise: new Promise(() => {}), cancel: vi.fn() }), + } }) + await vi.advanceTimersByTimeAsync(8_000) + await first + vi.useRealTimers() + expect(manager.hasStrandedDelivery('s1')).toBe(false) + + const next = await manager.deliverPromptToAgent('s1', 'the next task') + expect(next).toMatchObject({ ok: false, code: 'not-ready', promptWritten: false }) + expect(bWrites).toEqual([]) +}) + +// #1358 review c: the recorded paint lag is 0.7-3.8 s after the failure. A +// delivery that starts inside it sees an empty, ready composer; writing then +// would put its prompt after (or before) our late-painting bytes, and one +// Enter could submit both. While the mark stands it waits for the late paint +// (bounded), then reclaims. +it('waits for the stranded text to paint before writing, then clears it', async () => { + const { manager, session, writes } = claudeLike() + await strand(manager) + vi.useFakeTimers() + const next = manager.deliverPromptToAgent('s1', 'the next task') + await vi.advanceTimersByTimeAsync(2_000) + session.paintLate() + await vi.advanceTimersByTimeAsync(10_000) + await expect(next).resolves.toMatchObject({ ok: true }) + expect(writes.slice(1)).toEqual(['\x15', 'the next task', '\r']) +}) + +// #1358 reviews a and c: whether Ctrl+U removes an image pill is not +// established, so an image delivery's leftovers are never reclaimed. +it('does not mark an image delivery that stranded', async () => { + const { manager } = claudeLike() + vi.useFakeTimers() + const first = manager.deliverPromptToAgent('s1', '', ['/tmp/screenshot.png']) + await vi.advanceTimersByTimeAsync(30_000) + expect(await first).toMatchObject({ ok: false, promptWritten: true, enterWritten: false }) + vi.useRealTimers() + expect(manager.hasStrandedDelivery('s1')).toBe(false) +}) + +// #1358 review a (surviving mutant): a write that throws after bytes may have +// crossed is stranded too; review c (surviving mutant): process exit clears it. +it('marks a delivery whose write threw, and forgets it when the process exits', async () => { + const { manager, session } = claudeLike() + const write = session.write + session.write = (data: string) => { if (data === 'an earlier prompt that painted late') throw new Error('EPIPE'); write(data) } + const result = await manager.deliverPromptToAgent('s1', 'an earlier prompt that painted late') + expect(result).toMatchObject({ ok: false, code: 'transport-failed', promptWritten: true }) + expect(manager.hasStrandedDelivery('s1')).toBe(true) + ;(manager as unknown as { cleanupSessionState(id: string, kind: string): void }).cleanupSessionState('s1', 'claude') + expect(manager.hasStrandedDelivery('s1')).toBe(false) +}) diff --git a/src/main/sessionManager.ts b/src/main/sessionManager.ts index 429620978..f9825958f 100644 --- a/src/main/sessionManager.ts +++ b/src/main/sessionManager.ts @@ -604,7 +604,19 @@ export class SessionManager extends EventEmitter { * .strandedComposer). A text comparison could not prove this: the screen is * clipped and wrap-lossy, and pastes collapse to `[Pasted text #N]`. */ - private readonly strandedDeliveries = new Set() + // + // WHY keyed to the PROCESS (the registry entry), not just the session id + // (steering q75, #1358 review a): a session id outlives its process. A + // delivery can write to process A, A can exit and B take the same id, and + // A's delivery can fail only afterwards. Recording that failure against the + // id would hand B a mark for bytes B never received, and the next delivery + // would clear B's composer, which may be a human's draft. The mark + // therefore names the entry the delivery captured, is only set while that + // entry still owns the id, and is only honoured for that same entry. + // + // `at` is when the strand was recorded; the next delivery waits a bounded + // window from it for the late paint (see PromptDeliveryIo.strandedComposer). + private readonly strandedDeliveries = new Map() /** * Prompts waiting for a session that is not ready for one YET (#854). * @@ -4083,18 +4095,35 @@ export class SessionManager extends EventEmitter { /** See strandedDeliveries (#1350). */ private noteStrandedDelivery( sessionId: string, + entry: RegistryEntry, delivery: { ok: boolean; promptWritten?: boolean; enterWritten?: boolean }, + hadImages: boolean, ): void { - if (delivery.ok) this.strandedDeliveries.delete(sessionId) - else if (delivery.promptWritten && !delivery.enterWritten) this.strandedDeliveries.add(sessionId) + // The process this delivery wrote to is gone (exited, replaced under the + // same id): nothing it did says anything about the current composer. + if (this.sessions.get(sessionId) !== entry) return + if (delivery.ok) { + this.strandedDeliveries.delete(sessionId) + return + } // A failure that wrote nothing (refused before write) changes nothing: // whatever the composer held before still holds, ours or not. + if (!delivery.promptWritten || delivery.enterWritten) return + // WHY an image delivery is never marked (#1358 reviews a and c): its + // leftovers include image pills, and whether Ctrl+U removes a pill has + // not been observed on a real composer (promptDelivery.ts says the same + // for its own rollback). A reclaim could strip the text and leave pills, + // or read a pill-only composer wrongly. Such a session stays as before: + // occupied until someone clears it. + if (hadImages) return + this.strandedDeliveries.set(sessionId, { entry, at: Date.now() }) } /** Whether the session's composer can only hold our own stranded write * (#1350); reported by sessions.inputInspect. */ hasStrandedDelivery(sessionId: string): boolean { - return this.strandedDeliveries.has(sessionId) + const mark = this.strandedDeliveries.get(sessionId) + return Boolean(mark && this.sessions.get(sessionId) === mark.entry) } getSessionKind(sessionId: string): SessionKind | null { @@ -4993,9 +5022,13 @@ export class SessionManager extends EventEmitter { imagePaths, record, ...(options?.requireEmptyNativeComposer ? { requireEmptyNativeComposer: true } : {}), - ...(this.strandedDeliveries.has(sessionId) ? { strandedComposer: true } : {}), + ...(() => { + // Only for the process the strand happened in (see strandedDeliveries). + const mark = this.strandedDeliveries.get(sessionId) + return mark && mark.entry === entry ? { strandedComposer: { strandedAt: mark.at } } : {} + })(), }) - this.noteStrandedDelivery(sessionId, delivery) + this.noteStrandedDelivery(sessionId, entry, delivery, Boolean(imagePaths?.length)) finishDelivery(delivery.ok ? 'success' : 'error') // Instrumentation must never change a delivery outcome. `acceptance` is // read defensively because a provider result without it made @@ -5010,7 +5043,7 @@ export class SessionManager extends EventEmitter { finishDelivery('error') this.monitorResponses.cancel(sessionId) record?.('uncertain', { reason: 'provider-threw' }) - this.noteStrandedDelivery(sessionId, { ok: false, promptWritten, enterWritten }) + this.noteStrandedDelivery(sessionId, entry, { ok: false, promptWritten, enterWritten }, Boolean(imagePaths?.length)) return { ok: false, stage: enterWritten ? 'after-enter' : promptWritten ? 'absorption' : 'before-write', diff --git a/src/providers/claude/runtime/promptDelivery.strandedReclaim.test.ts b/src/providers/claude/runtime/promptDelivery.strandedReclaim.test.ts index 91c83d663..ccabd87a3 100644 --- a/src/providers/claude/runtime/promptDelivery.strandedReclaim.test.ts +++ b/src/providers/claude/runtime/promptDelivery.strandedReclaim.test.ts @@ -30,7 +30,7 @@ function harness(opts: { strandedComposer?: boolean; killClears?: boolean }) { const io = { sessionId: 's1', prompt: 'send the next task', - ...(opts.strandedComposer ? { strandedComposer: true } : {}), + ...(opts.strandedComposer ? { strandedComposer: { strandedAt: 0 } } : {}), write: (data: string) => { writes.push(data) if (data === '\x15' && opts.killClears !== false) state = 'empty' diff --git a/src/providers/claude/runtime/promptDelivery.ts b/src/providers/claude/runtime/promptDelivery.ts index 2271f4f1a..7e3581cf3 100644 --- a/src/providers/claude/runtime/promptDelivery.ts +++ b/src/providers/claude/runtime/promptDelivery.ts @@ -74,6 +74,22 @@ export async function deliverClaudePrompt( deadlineAt: Math.min(deliveryDeadlineAt, Date.now() + READY_BUDGET_MS), }) let ready = await awaitReady() + if (ready.kind === 'ready' && io.strandedComposer) { + // #1358 review c: our stranded bytes may not have painted yet. In the + // three recorded incidents they painted 0.7-3.8 s after the earlier + // delivery gave up. Writing now would put this prompt next to them in + // the composer, and one Enter could submit both. So, still holding the + // reservation (nobody else can type), give them a bounded window to + // appear; if they do, the gate reads occupied and they are reclaimed + // below. If they never paint, they were consumed or never landed, and + // there is nothing of ours to clear. + const windowEndsAt = Math.min(deliveryDeadlineAt, io.strandedComposer.strandedAt + STRANDED_PAINT_WINDOW_MS) + while (Date.now() < windowEndsAt) { + if (classifyRollbackComposer(io.session.readComposer?.() ?? null, io.session.snapshotScreen?.() ?? '') === 'drafted') break + await sleep(CONFIRM_POLL_INTERVAL_MS * 10) + } + ready = await awaitReady() + } if (ready.kind === 'occupied' && io.strandedComposer) { // The "human draft" is our own earlier write (#1350): an earlier // delivery's bytes painted after its rollback stopped watching, and no @@ -445,6 +461,11 @@ const KILL_KEYSTROKE_GAP_MS = 25 // cannot see it. Bounded well under the delivery deadline: this runs after a // failure, and a slow answer here delays the error the user is waiting for. const ROLLBACK_OBSERVE_ATTEMPTS = 40 +// How long after an earlier delivery stranded its bytes the next delivery +// waits for them to paint before writing its own (#1358 review c). The +// recorded late paints landed 0.7-3.8 s after the failure; this is about 2x +// the worst, and only a delivery that starts inside it ever waits. +const STRANDED_PAINT_WINDOW_MS = 8_000 const sleep = (ms: number): Promise => new Promise(resolve => { setTimeout(resolve, ms) }) diff --git a/src/shared/types/providerConfig.ts b/src/shared/types/providerConfig.ts index cf0b53fcf..fcab435e5 100644 --- a/src/shared/types/providerConfig.ts +++ b/src/shared/types/providerConfig.ts @@ -441,10 +441,12 @@ export type PromptDeliveryIo = PromptDeliveryOptions & { * Set by SessionManager when an earlier delivery wrote prompt bytes it could * not submit and no other writer has touched the PTY since. A provider may * then clear an occupied composer under this delivery's reservation instead - * of refusing it as a human draft. Absent or false: an occupied composer is - * someone else's, exactly as before. + * of refusing it as a human draft. Absent: an occupied composer is someone + * else's, exactly as before. `strandedAt` is when the earlier delivery gave + * up; its bytes may still be unpainted (they painted 0.7-3.8 s later in the + * recorded incidents), so a provider waits a bounded window from it. */ - strandedComposer?: boolean + strandedComposer?: { strandedAt: number } } export type PromptAcceptance = From 5610f2d2c667723914a1b8151a13fade74aee748 Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sat, 26 Sep 2026 22:50:29 -0700 Subject: [PATCH 06/48] fix(performance): size histograms for every operation, close failed discovery spans, pin the catalog thresholds Round-1 review of #1352: 26 operations x 4 outcomes outgrew three hard-coded 100-pair caps; a rejected discovery left its span pending until the sweep called it a timeout; the new thresholds were untested, and discovery's sat above the main-stall band. Refs #769 Co-Authored-By: Claude Opus 5.5 --- ...026-09-26-monitor-transcript-operations.md | 11 ++++++++++ src/main/conversations/service.system.test.ts | 21 ++++++++++++++++++- src/main/conversations/service.ts | 8 +++++++ src/main/performance/IncidentEngine.test.ts | 17 +++++++++++++++ src/main/performance/IncidentEngine.ts | 9 +++++--- .../performance/MonitorAggregator.test.ts | 20 ++++++++++++++++++ src/main/performance/MonitorAggregator.ts | 4 ++-- src/main/performance/MonitorHistoryStore.ts | 4 ++-- .../sessions/historyLoader.monitor.test.ts | 14 ++++++++++--- src/shared/performance/monitorPolicy.ts | 11 +++++++++- .../performance/parseMonitorSnapshot.ts | 3 ++- 11 files changed, 109 insertions(+), 13 deletions(-) diff --git a/docs/plans/2026-09-26-monitor-transcript-operations.md b/docs/plans/2026-09-26-monitor-transcript-operations.md index 45ebb6fad..312092bb2 100644 --- a/docs/plans/2026-09-26-monitor-transcript-operations.md +++ b/docs/plans/2026-09-26-monitor-transcript-operations.md @@ -19,3 +19,14 @@ ## Tests - **History loader:** one initial load through `loadInitialHistoryChunk` records exactly one `transcript.read` (red on main: two). - **Catalog:** a search with a query records `conversations.search`, and an extract records `conversations.extract`, driven on the recorded conversations corpus. + +## Round 1 review decisions (#1352) +- **a1: a discovery that rejects left its span open.** It became a `timeout` sample at the sweep. The earlier call that discovery "cannot fail" was wrong: family resolution rejects a malformed cwd. The span now closes with `fail(error)`. Test added; red before. +- **a2: the vocabulary outgrew the histogram cap.** 26 operations × 4 outcomes is 104 pairs, but the aggregator, the snapshot parser and the history store each hard-coded 100. One derived constant, `MAX_MONITOR_OPERATION_PAIRS`, now serves all three. Test: every legal pair is stored and the snapshot parses. Red before. +- **a3, b1, c2: no test pinned the three new thresholds.** Each now gets a sample just above (a named slow-operation incident) and just below (none). +- **c1: discovery's 2000 ms threshold sat above main-stall's 1000 ms band.** A 1.5 s discovery stall raised no named incident, so discovery and search are now both 1000 ms. +- **b2: the extraction assertion ignored whether prompts came back.** It now requires at least one prompt. +- **c3: the load-time test counted samples but not which span.** With the resolver taking 80 ms, the single sample must be shorter than that, so it timed the read, not the lookup. Mapping the lookup span instead fails it. +- **Declined:** + - b3 and c5, the corpus-install `beforeAll` timing out under heavy load: the hook predates this PR and CI runs it. Timeouts are never widened. + - c4, a 60 s cooldown slot per scope shared by all main-scope slow operations: pre-existing semantics, outside this PR. diff --git a/src/main/conversations/service.system.test.ts b/src/main/conversations/service.system.test.ts index e7d1b2a25..52ea3ef6f 100644 --- a/src/main/conversations/service.system.test.ts +++ b/src/main/conversations/service.system.test.ts @@ -79,7 +79,8 @@ describe('ConversationService', () => { const s = service() const all = await s.list({ cwd: '/fixture/repo', scope: 'repository', includeChildren: true, limit: 5000 }) const codex = all.rows.find(r => r.provider === 'codex' && r.cwd)! - await s.prompts({ provider: 'codex', nativeId: codex.nativeId, cwd: codex.cwd! }) + // #1352 review b: the extraction measured must be a real one. + expect((await s.prompts({ provider: 'codex', nativeId: codex.nativeId, cwd: codex.cwd! })).length).toBeGreaterThan(0) await s.list({ cwd: '/fixture/repo', scope: 'repository', query: 'the', limit: 50 }) } finally { setMainOperationSink(() => {}) @@ -89,4 +90,22 @@ describe('ConversationService', () => { expect(names).toContain('conversations.search') expect(names).toContain('conversations.extract') }) + + // #1352 review a: a discovery that rejects closes its span as an error, + // instead of leaving it pending until the sweep reports a timeout. + it('closes a failed discovery as an error, not a pending span', async () => { + const { mainOperations } = await import('@main/performance/operations.js') + const operations: Array<{ name: string; outcome: string }> = [] + setMainOperationSink(record => { operations.push(record) }) + const pendingBefore = mainOperations.size + try { + const failing = new ConversationService({ sources: [], ledger: null, listWorktrees: async () => [], claudeHistory: null }) + await expect(failing.list({ cwd: null, scope: 'repository', limit: 1 } as never)).rejects.toThrow() + } finally { + setMainOperationSink(() => {}) + } + expect(operations.filter(op => op.name === 'conversations.discover').map(op => op.outcome)).toEqual(['error']) + expect(mainOperations.size).toBe(pendingBefore) + }) }) + diff --git a/src/main/conversations/service.ts b/src/main/conversations/service.ts index 94972e12f..2b069d1d3 100644 --- a/src/main/conversations/service.ts +++ b/src/main/conversations/service.ts @@ -99,6 +99,14 @@ export class ConversationService { this.discoveries++ span.end({ rows: discovery.sources.length }) return discovery + } catch (error) { + // #1352 review a: family resolution can reject (a malformed cwd from a + // future caller; IPC validates it first today). Left open, the span + // became a `timeout` sample at the ten-minute sweep and could raise a + // slow-operation incident blaming discovery for a stall it never + // caused. Close it as the error it is. + span.fail(error) + throw error } finally { // Only one flight per key can exist (identical keys coalesce above), // so clearing by key is clearing this flight. diff --git a/src/main/performance/IncidentEngine.test.ts b/src/main/performance/IncidentEngine.test.ts index 0ccb3594d..9cf3a2c5e 100644 --- a/src/main/performance/IncidentEngine.test.ts +++ b/src/main/performance/IncidentEngine.test.ts @@ -133,4 +133,21 @@ describe('incident evidence and exclusion rules', () => { engine.loss(901, 0, 2000) expect(engine.summaries()[0]!.truncated).toBe(true) }) + + // #1352 reviews a, b, c: the three catalog thresholds are the point of the + // PR, and moving any of them left every test green. Each gets a sample just + // above (a named slow-operation incident) and just below (none). + it.each([ + ['conversations.extract', 250], + ['conversations.search', 1000], + ['conversations.discover', 1000], + ] as const)('raises a slow-operation incident for %s only above %i ms', (name, threshold) => { + const above = new IncidentEngine() + above.accept([{ kind: 'operation', at: 1000, windowId: null, sample: { kind: 'operation', name, durationMs: threshold + 1, outcome: 'success' } }], 1000, 1000) + expect(above.summaries().map(row => [row.rule, row.operation])).toEqual([['slow-operation', name]]) + const below = new IncidentEngine() + below.accept([{ kind: 'operation', at: 1000, windowId: null, sample: { kind: 'operation', name, durationMs: threshold - 1, outcome: 'success' } }], 1000, 1000) + expect(below.summaries()).toEqual([]) + }) }) + diff --git a/src/main/performance/IncidentEngine.ts b/src/main/performance/IncidentEngine.ts index a8d71e866..f92b1232d 100644 --- a/src/main/performance/IncidentEngine.ts +++ b/src/main/performance/IncidentEngine.ts @@ -8,9 +8,12 @@ const operationThresholds: Partial> = { 'transcript.read': 1000, 'transcript.parse': 100, 'transcript.fold': 100, 'transcript.commit': 100, 'terminal.write': 250, 'persistence.serialize': 100, 'persistence.write': 2000, 'worktree.refresh': 5000, - // One file's synchronous extraction, a search's prompt gathering (many - // rows, mostly I/O-bound), and a full discovery pass (#769). - 'conversations.extract': 250, 'conversations.search': 1000, 'conversations.discover': 2000, + // One file's synchronous extraction, then a search's prompt gathering and a + // full discovery pass (#769). Search and discovery sit AT the main-stall + // band (loopMax >= 1000 ms), not above it (#1352 review c): above it, a + // discovery that stalled main for 1.5 s raised a main-stall with no named + // operation incident, which is the attribution this vocabulary exists for. + 'conversations.extract': 250, 'conversations.search': 1000, 'conversations.discover': 1000, } // Each capture retains at most 160 evidence points. Eight simultaneous scopes // bound post-trigger memory to ~1.3k points even during an app-wide storm, diff --git a/src/main/performance/MonitorAggregator.test.ts b/src/main/performance/MonitorAggregator.test.ts index 267d1fc4d..726dc9405 100644 --- a/src/main/performance/MonitorAggregator.test.ts +++ b/src/main/performance/MonitorAggregator.test.ts @@ -17,4 +17,24 @@ describe('worker evidence bounds', () => { expect(JSON.stringify(snapshot)).not.toContain('session-') expect(aggregator.snapshot(11000 * 1000, 100).recent).toEqual([]) }) + + // #1352 review a: every legal (operation, outcome) pair must fit. The + // aggregator, the snapshot parser and the history store each capped at a + // literal 100, and 26 operations x 4 outcomes = 104 silently lost the last + // four pairs. + it('keeps a histogram for every legal operation and outcome, and the snapshot parses', async () => { + const { MONITOR_OPERATIONS, MONITOR_OUTCOMES } = await import('@shared/performance/monitorPolicy.js') + const { parseMonitorSnapshot } = await import('@shared/performance/parseMonitorSnapshot.js') + const aggregator = new MonitorAggregator() + for (const name of MONITOR_OPERATIONS) { + for (const outcome of MONITOR_OUTCOMES) { + aggregator.accept([{ kind: 'operation', at: 1000, windowId: null, sample: { kind: 'operation', name, outcome, durationMs: 5 } }]) + } + } + const snapshot = aggregator.snapshot(2000, 100) + expect(snapshot.operations).toHaveLength(MONITOR_OPERATIONS.length * MONITOR_OUTCOMES.length) + expect(snapshot.operations.map(entry => `${entry.name}:${entry.outcome}`)).toContain('heap.snapshot:timeout') + expect(parseMonitorSnapshot(JSON.parse(JSON.stringify(snapshot)))).not.toBeNull() + }) }) + diff --git a/src/main/performance/MonitorAggregator.ts b/src/main/performance/MonitorAggregator.ts index 88ebf4488..a9db13099 100644 --- a/src/main/performance/MonitorAggregator.ts +++ b/src/main/performance/MonitorAggregator.ts @@ -1,6 +1,6 @@ import { BoundedQueue } from '@shared/performance/boundedQueue.js' import { LatencyHistogram } from '@shared/performance/latencyHistogram.js' -import { MONITOR_POLICY } from '@shared/performance/monitorPolicy.js' +import { MAX_MONITOR_OPERATION_PAIRS, MONITOR_POLICY } from '@shared/performance/monitorPolicy.js' import type { MonitorEnvelope, MonitorMainSample, MonitorWindowSample, MonitorOperationSummary, MonitorWorkerSnapshot } from '@shared/performance/monitorSnapshot.js' /** Worker-owned evidence. Every cardinality is bounded independently of run length. */ @@ -30,7 +30,7 @@ export class MonitorAggregator { const { name, outcome, durationMs } = record.sample const key = `${name}:${outcome}` let entry = this.operations.get(key) - if (!entry && this.operations.size < 100) { + if (!entry && this.operations.size < MAX_MONITOR_OPERATION_PAIRS) { entry = { summary: { name, outcome }, histogram: new LatencyHistogram() } this.operations.set(key, entry) } diff --git a/src/main/performance/MonitorHistoryStore.ts b/src/main/performance/MonitorHistoryStore.ts index 2db7c5fea..ae087d18a 100644 --- a/src/main/performance/MonitorHistoryStore.ts +++ b/src/main/performance/MonitorHistoryStore.ts @@ -8,7 +8,7 @@ import type { MonitorIncident } from '@shared/performance/monitorIncidents.js' import { parseMonitorIncident } from '@shared/performance/parseMonitorIncident.js' import { parseMonitorHistoryPoint } from '@shared/performance/parseMonitorHistory.js' import { parseMonitorSnapshot } from '@shared/performance/parseMonitorSnapshot.js' -import { MONITOR_POLICY } from '@shared/performance/monitorPolicy.js' +import { MAX_MONITOR_OPERATION_PAIRS, MONITOR_POLICY } from '@shared/performance/monitorPolicy.js' import type { MonitorHistoryPage, MonitorHistoryPoint, MonitorHistoryResolution, MonitorHistoryStatus, MonitorReportPreview, MonitorReportResult, @@ -121,7 +121,7 @@ export class MonitorHistoryStore { // fingerprint; stringifying up to fifty full evidence sets here every // second cost ~1 MiB of JSON per second of helper CPU for no new data. if (incidents) this.pendingIncidents = incidents.slice(-INCIDENT_LIMIT) - const operationJson = JSON.stringify(snapshot.operations.slice(0, 100)) + const operationJson = JSON.stringify(snapshot.operations.slice(0, MAX_MONITOR_OPERATION_PAIRS)) if (operationJson !== this.operationFingerprint) { this.operationFingerprint = operationJson this.pendingOperations = operationJson diff --git a/src/main/sessions/historyLoader.monitor.test.ts b/src/main/sessions/historyLoader.monitor.test.ts index 97792efff..f2ed6ce89 100644 --- a/src/main/sessions/historyLoader.monitor.test.ts +++ b/src/main/sessions/historyLoader.monitor.test.ts @@ -11,7 +11,13 @@ const fixture = join( import.meta.dirname, '../../../testing/fixtures/conversations/claude/projects/-fixture-repo--worktrees-feat-api-key-vault/fc475787-6395-4cda-8bd2-1faacaa18bc7.jsonl', ) -vi.mock('@main/providerSwitch/shared.js', () => ({ resolveProviderTranscriptPath: vi.fn(async () => fixture) })) +// The resolver takes a measurable 80 ms, so the one sample can be checked to +// time the READ, not the path lookup (#1352 review c: a count alone passed +// when the lookup span was mapped instead of the read span). +const RESOLVE_MS = 80 +vi.mock('@main/providerSwitch/shared.js', () => ({ + resolveProviderTranscriptPath: vi.fn(async () => { await new Promise(resolve => setTimeout(resolve, RESOLVE_MS)); return fixture }), +})) vi.mock('@providers/registry.main.js', () => ({ getMainProvider: () => ({}) })) const { setMainOperationSink } = await import('@main/performance/operations.js') @@ -20,9 +26,11 @@ const { loadInitialHistoryChunk } = await import('./historyLoader.js') afterEach(() => setMainOperationSink(() => {})) it('records one transcript.read per initial history load', async () => { - const operations: Array<{ name: string }> = [] + const operations: Array<{ name: string; durationMs?: number }> = [] setMainOperationSink(record => { operations.push(record) }) const chunk = await loadInitialHistoryChunk({ kind: 'claude', cwd: '/fixture/repo', providerSessionId: 'fc475787-6395-4cda-8bd2-1faacaa18bc7', limit: 20 } as never) expect(chunk.entries.length).toBeGreaterThan(0) - expect(operations.filter(op => op.name === 'transcript.read')).toHaveLength(1) + const reads = operations.filter(op => op.name === 'transcript.read') as Array<{ name: string; durationMs: number }> + expect(reads).toHaveLength(1) + expect(reads[0]!.durationMs).toBeLessThan(RESOLVE_MS) }) diff --git a/src/shared/performance/monitorPolicy.ts b/src/shared/performance/monitorPolicy.ts index efc6f792b..5f3ad10c9 100644 --- a/src/shared/performance/monitorPolicy.ts +++ b/src/shared/performance/monitorPolicy.ts @@ -68,5 +68,14 @@ export const MONITOR_OPERATIONS = [ ] as const export type MonitorOperationName = typeof MONITOR_OPERATIONS[number] -export type MonitorOutcome = 'success' | 'error' | 'cancelled' | 'timeout' +export const MONITOR_OUTCOMES = ['success', 'error', 'cancelled', 'timeout'] as const +export type MonitorOutcome = typeof MONITOR_OUTCOMES[number] +/** + * How many (operation, outcome) histograms a snapshot can hold: every legal + * pair. WHY derived instead of a literal (#1352 review a): the aggregator, the + * snapshot parser and the history store each hard-coded 100, and adding three + * operations (26 x 4 = 104 pairs) made them silently drop the last pairs. A + * cap computed from the vocabulary grows with it and stays one number. + */ +export const MAX_MONITOR_OPERATION_PAIRS = MONITOR_OPERATIONS.length * MONITOR_OUTCOMES.length export type MonitorQuality = 'ok' | 'warming-up' | 'stale' | 'unsupported' | 'partial' diff --git a/src/shared/performance/parseMonitorSnapshot.ts b/src/shared/performance/parseMonitorSnapshot.ts index 39d98c8ed..e0acc3bf3 100644 --- a/src/shared/performance/parseMonitorSnapshot.ts +++ b/src/shared/performance/parseMonitorSnapshot.ts @@ -1,3 +1,4 @@ +import { MAX_MONITOR_OPERATION_PAIRS } from './monitorPolicy.js' import { parseIncidentSummary } from './parseMonitorIncident.js' import { parseMonitorHistoryStatus } from './parseMonitorHistory.js' import { parseMonitorRendererRecord } from './monitorContracts.js' @@ -25,7 +26,7 @@ export function parseMonitorSnapshot(value: unknown): MonitorWorkerSnapshot | nu || !finite(value.sampledAt) || !finite(value.workerRss) || !Array.isArray(value.windows) || value.windows.length > 64 || !Array.isArray(value.recent) || value.recent.length > 120 - || !Array.isArray(value.operations) || value.operations.length > 100) return null + || !Array.isArray(value.operations) || value.operations.length > MAX_MONITOR_OPERATION_PAIRS) return null const incidents = value.incidents === undefined ? [] : Array.isArray(value.incidents) && value.incidents.length <= 50 ? value.incidents.map(parseIncidentSummary) : null const history = value.history === undefined ? undefined : parseMonitorHistoryStatus(value.history) if (!incidents || incidents.some(row => !row) || (value.history !== undefined && !history)) return null From 2f0bf0561ec5b571dec5cfa8e0bab2de72811c1c Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sat, 26 Sep 2026 22:53:03 -0700 Subject: [PATCH 07/48] fix(claude): mark only sessions whose delivery can reclaim, pin refusals, import the control stack statically Round-1 review b of #1358. Refs #1350 Co-Authored-By: Claude Opus 5.5 --- ...26-09-27-claude-stranded-prompt-reclaim.md | 4 +++ .../sessionManager.strandedDelivery.test.ts | 28 ++++++++++++++++++- src/main/sessionManager.ts | 5 ++++ 3 files changed, 36 insertions(+), 1 deletion(-) diff --git a/docs/plans/2026-09-27-claude-stranded-prompt-reclaim.md b/docs/plans/2026-09-27-claude-stranded-prompt-reclaim.md index e1113a220..b2e41a34f 100644 --- a/docs/plans/2026-09-27-claude-stranded-prompt-reclaim.md +++ b/docs/plans/2026-09-27-claude-stranded-prompt-reclaim.md @@ -46,3 +46,7 @@ In all three, the composer turned occupied 0.7–3.8 s after the delivery gave u - **Suspicions kept as residuals:** - headless-internal writers (trust, resume, permission) that bypass `recordInputWrite`: no app consumer was found; - a mark set after a throw before any bytes crossed: harmless, since the next gate reads ready and nothing is killed. +- **b (after the first round):** + - the inspection test's dynamic import timed out under load, so it is now a static import; + - a refusal that wrote nothing is now pinned not to set a mark; + - only Claude sessions are marked, because only Claude's delivery reclaims, and inspection must not promise a reclaim no delivery performs. Test added; the mutant is killed. diff --git a/src/main/sessionManager.strandedDelivery.test.ts b/src/main/sessionManager.strandedDelivery.test.ts index faab33aab..efe392df8 100644 --- a/src/main/sessionManager.strandedDelivery.test.ts +++ b/src/main/sessionManager.strandedDelivery.test.ts @@ -1,6 +1,9 @@ import { afterEach, expect, it, vi } from 'vitest' import { SessionManager } from './sessionManager.js' +// Static, not imported inside the test (#1358 review b): a dynamic import of +// the control stack took over 5 s on a loaded machine and timed the test out. +import { terminalBackendCapabilities } from '@main/sessions/terminalControl.js' // #1350, end to end through the manager and the real Claude delivery code. // Recorded shape (lifecycle journal, three incidents on 2026-09-27): a @@ -82,7 +85,6 @@ it('treats the composer as a human draft again once anyone else has written to i }) it('reports our own stranded write to input inspection', async () => { - const { terminalBackendCapabilities } = await import('@main/sessions/terminalControl.js') const { manager } = claudeLike() await strand(manager) vi.spyOn(manager, 'getBackendSnapshot').mockReturnValue({ @@ -172,3 +174,27 @@ it('marks a delivery whose write threw, and forgets it when the process exits', ;(manager as unknown as { cleanupSessionState(id: string, kind: string): void }).cleanupSessionState('s1', 'claude') expect(manager.hasStrandedDelivery('s1')).toBe(false) }) + +// #1358 review b (surviving mutant): a delivery refused before writing leaves +// the composer as it was, so it must not create a mark. +it('does not mark a delivery that was refused before writing', async () => { + const { manager, session, writes } = claudeLike() + session.paintLate() + const result = await manager.deliverPromptToAgent('s1', 'the next task') + expect(result).toMatchObject({ ok: false, code: 'not-ready', promptWritten: false }) + expect(writes).toEqual([]) + expect(manager.hasStrandedDelivery('s1')).toBe(false) +}) + +// #1358 review b: only Claude's delivery can reclaim a stranded composer, so +// only a Claude session is marked; inspection must not promise a reclaim no +// delivery will perform. +it('does not mark a stranded delivery for a provider that cannot reclaim it', async () => { + const { manager, session } = claudeLike() + const codexLike = { ...session, write: (data: string) => { if (data.includes('earlier')) throw new Error('EPIPE') } } + ;(manager as unknown as { sessions: Map }).sessions.set('s1', { kind: 'codex', session: codexLike }) + const result = await manager.deliverPromptToAgent('s1', 'an earlier prompt that painted late') + expect(result).toMatchObject({ ok: false, promptWritten: true }) + expect(manager.hasStrandedDelivery('s1')).toBe(false) +}) + diff --git a/src/main/sessionManager.ts b/src/main/sessionManager.ts index f9825958f..0c7e2e14e 100644 --- a/src/main/sessionManager.ts +++ b/src/main/sessionManager.ts @@ -4116,6 +4116,11 @@ export class SessionManager extends EventEmitter { // or read a pill-only composer wrongly. Such a session stays as before: // occupied until someone clears it. if (hadImages) return + // Only a provider whose delivery reclaims a stranded composer (#1358 + // review b). Today that is Claude's (promptDelivery.ts); marking another + // provider would make inputInspect promise "the next delivery clears it" + // when no delivery will. A provider that adds a reclaim joins here. + if (entry.kind !== 'claude') return this.strandedDeliveries.set(sessionId, { entry, at: Date.now() }) } From 6eeac3b1dc34400caba5a44867f6426ad8b64809 Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sat, 26 Sep 2026 22:57:37 -0700 Subject: [PATCH 08/48] test(claude): exercise the stranded-mark provider guard through Pi's unknown outcome Codex writes paste+Enter atomically, so the Codex test never reached the kind check (it passed with the guard deleted). Pi's bridge 'unknown' outcome reports promptWritten && !enterWritten, which does. Co-Authored-By: Claude Opus 5.5 --- .../sessionManager.strandedDelivery.test.ts | 21 ++++++++++++++----- 1 file changed, 16 insertions(+), 5 deletions(-) diff --git a/src/main/sessionManager.strandedDelivery.test.ts b/src/main/sessionManager.strandedDelivery.test.ts index efe392df8..6420d4ba4 100644 --- a/src/main/sessionManager.strandedDelivery.test.ts +++ b/src/main/sessionManager.strandedDelivery.test.ts @@ -189,12 +189,23 @@ it('does not mark a delivery that was refused before writing', async () => { // #1358 review b: only Claude's delivery can reclaim a stranded composer, so // only a Claude session is marked; inspection must not promise a reclaim no // delivery will perform. +// +// WHY Pi and not Codex: the guard is only reachable by a non-Claude failure +// that reports promptWritten && !enterWritten. Codex writes paste + Enter in +// ONE atomic PTY write (codex/runtime/promptDelivery.ts), so any Codex failure +// after writing is already enterWritten and never reaches the kind check (a +// Codex version of this test passed with the guard deleted). Pi's bridge +// answers "unknown" when the request reached pi with no evidence back, which +// pi/runtime/promptDelivery.ts reports as exactly that pair. it('does not mark a stranded delivery for a provider that cannot reclaim it', async () => { - const { manager, session } = claudeLike() - const codexLike = { ...session, write: (data: string) => { if (data.includes('earlier')) throw new Error('EPIPE') } } - ;(manager as unknown as { sessions: Map }).sessions.set('s1', { kind: 'codex', session: codexLike }) + const { manager } = claudeLike() + const unknownOutcome = Object.assign(new Error('bridge went quiet'), { code: 'pi-terminal-unknown' }) + ;(manager as unknown as { sessions: Map }).sessions.set('s1', { kind: 'pi', session: { + isExited: () => false, + write: () => {}, + deliverPromptText: async () => { throw unknownOutcome }, + } }) const result = await manager.deliverPromptToAgent('s1', 'an earlier prompt that painted late') - expect(result).toMatchObject({ ok: false, promptWritten: true }) + expect(result).toMatchObject({ ok: false, promptWritten: true, enterWritten: false }) expect(manager.hasStrandedDelivery('s1')).toBe(false) }) - From 7ef4cecf98475021b0725b8206a76ef752657b89 Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sat, 26 Sep 2026 23:06:47 -0700 Subject: [PATCH 09/48] test(claude): pin each side of the stranded mark's process binding Mutation battery on 6eeac3b1: the write-side entry check, the strandedComposer handoff check and the exit clear each survived alone, because the other layer covered for them. A replacement registered before the old process's generation-owned cleanup now pins the read side; the map itself pins the write side and the exit clear (their only effect once the row is gone is not retaining the dead RegistryEntry). Co-Authored-By: Claude Opus 5.5 --- .../sessionManager.strandedDelivery.test.ts | 34 +++++++++++++++++++ src/main/sessionManager.ts | 8 +++++ 2 files changed, 42 insertions(+) diff --git a/src/main/sessionManager.strandedDelivery.test.ts b/src/main/sessionManager.strandedDelivery.test.ts index 6420d4ba4..373232db2 100644 --- a/src/main/sessionManager.strandedDelivery.test.ts +++ b/src/main/sessionManager.strandedDelivery.test.ts @@ -127,12 +127,42 @@ it('never lets a late failure of a replaced process clear the new process\'s hum await first vi.useRealTimers() expect(manager.hasStrandedDelivery('s1')).toBe(false) + // The write side: A's late result must not even be stored. The read side + // would ignore it, but a stored mark keeps A's dead RegistryEntry (and its + // PTY wrapper) reachable until someone writes to B. + expect((manager as unknown as { strandedDeliveries: Map }).strandedDeliveries.has('s1')).toBe(false) const next = await manager.deliverPromptToAgent('s1', 'the next task') expect(next).toMatchObject({ ok: false, code: 'not-ready', promptWritten: false }) expect(bWrites).toEqual([]) }) +// The read side of the binding. Stable ids are reused after a failed provider +// start, and the replacement can be registered BEFORE the old wrapper's exit +// fires; that late exit's cleanup is generation-owned and deliberately leaves +// the new row (and so the old mark) alone. The mark still names A's entry, so +// neither inspection nor B's next delivery may treat it as B's. Without the +// entry comparison at the strandedComposer handoff, B's human draft gets Ctrl+U. +it('never hands a mark to a replacement registered before the old process was cleaned up', async () => { + const { manager } = claudeLike() + await strand(manager) + expect(manager.hasStrandedDelivery('s1')).toBe(true) + const bWrites: string[] = [] + const humanDraft = { screen: composer('❯ a human typed this'), attributes: { dim: 0, inverse: 1, plain: 18 } } + ;(manager as unknown as { sessions: Map }).sessions.set('s1', { kind: 'claude', session: { + isExited: () => false, + write: (data: string) => { bWrites.push(data) }, + snapshotScreen: () => humanDraft.screen, + readComposer: () => humanDraft, + awaitReadyForPrompt: async () => ({ kind: 'occupied' as const, reason: 'human-draft' as const, waitedMs: 0 }), + armPromptAcceptance: () => ({ promise: new Promise(() => {}), cancel: vi.fn() }), + } }) + expect(manager.hasStrandedDelivery('s1')).toBe(false) + const next = await manager.deliverPromptToAgent('s1', 'the next task') + expect(next).toMatchObject({ ok: false, code: 'not-ready', promptWritten: false }) + expect(bWrites).toEqual([]) +}) + // #1358 review c: the recorded paint lag is 0.7-3.8 s after the failure. A // delivery that starts inside it sees an empty, ready composer; writing then // would put its prompt after (or before) our late-painting bytes, and one @@ -173,6 +203,10 @@ it('marks a delivery whose write threw, and forgets it when the process exits', expect(manager.hasStrandedDelivery('s1')).toBe(true) ;(manager as unknown as { cleanupSessionState(id: string, kind: string): void }).cleanupSessionState('s1', 'claude') expect(manager.hasStrandedDelivery('s1')).toBe(false) + // Checked on the map itself: once the row is gone the read-side entry check + // already hides the mark, so only this shows the exit clear still runs (a + // kept mark would pin the dead process's RegistryEntry). + expect((manager as unknown as { strandedDeliveries: Map }).strandedDeliveries.has('s1')).toBe(false) }) // #1358 review b (surviving mutant): a delivery refused before writing leaves diff --git a/src/main/sessionManager.ts b/src/main/sessionManager.ts index 0c7e2e14e..6b744d444 100644 --- a/src/main/sessionManager.ts +++ b/src/main/sessionManager.ts @@ -4101,6 +4101,14 @@ export class SessionManager extends EventEmitter { ): void { // The process this delivery wrote to is gone (exited, replaced under the // same id): nothing it did says anything about the current composer. + // WHY check here AND on read (hasStrandedDelivery, the strandedComposer + // handoff): the read side alone already keeps a stale mark from reaching + // a replacement, but a stored stale mark would (a) keep the dead + // RegistryEntry and its PTY wrapper reachable until the next write, and + // (b) let A's late ok/failure delete or overwrite a mark that B's own + // delivery set. The read side is still needed because a replacement can + // be registered before A's exit cleanup runs, and that cleanup is + // generation-owned, so it leaves A's mark in place. if (this.sessions.get(sessionId) !== entry) return if (delivery.ok) { this.strandedDeliveries.delete(sessionId) From 0466ebb662f88f2e7974678fc27b9851d10d8d67 Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sat, 26 Sep 2026 23:08:31 -0700 Subject: [PATCH 10/48] fix(claude): a delivery that wrote retires the standing stranded mark #1358 verification a (blocker): an image delivery that reclaimed a text mark and then stranded left the old mark standing, so the next delivery Ctrl+U'd a composer holding image pills. Any delivery that wrote now clears the mark before deciding whether to set its own. Co-Authored-By: Claude Opus 5.5 --- .../sessionManager.strandedDelivery.test.ts | 23 +++++++++++++++++++ src/main/sessionManager.ts | 15 +++++++++++- 2 files changed, 37 insertions(+), 1 deletion(-) diff --git a/src/main/sessionManager.strandedDelivery.test.ts b/src/main/sessionManager.strandedDelivery.test.ts index 373232db2..799578ad2 100644 --- a/src/main/sessionManager.strandedDelivery.test.ts +++ b/src/main/sessionManager.strandedDelivery.test.ts @@ -192,6 +192,29 @@ it('does not mark an image delivery that stranded', async () => { expect(manager.hasStrandedDelivery('s1')).toBe(false) }) +// #1358 verification a (blocker): a delivery that WROTE supersedes whatever +// mark stood. Sequence: text strands and paints; an image delivery reclaims it +// (Ctrl+U), writes its image, and strands too. The old text mark must not +// survive that, or the third delivery Ctrl+Us a composer holding image pills. +it('retires the text mark once an image delivery has written over it', async () => { + const { manager, session, writes } = claudeLike() + await strand(manager) + session.paintLate() + vi.useFakeTimers() + const image = manager.deliverPromptToAgent('s1', '', ['/tmp/screenshot.png']) + await vi.advanceTimersByTimeAsync(30_000) + expect(await image).toMatchObject({ ok: false, promptWritten: true, enterWritten: false }) + vi.useRealTimers() + expect(writes.filter(data => data === '\x15')).toHaveLength(1) + expect(manager.hasStrandedDelivery('s1')).toBe(false) + + // The image's leftovers paint; the next delivery must refuse, not reclaim. + session.paintLate() + const next = await manager.deliverPromptToAgent('s1', 'the next task') + expect(next).toMatchObject({ ok: false, code: 'not-ready', promptWritten: false }) + expect(writes.filter(data => data === '\x15')).toHaveLength(1) +}) + // #1358 review a (surviving mutant): a write that throws after bytes may have // crossed is stranded too; review c (surviving mutant): process exit clears it. it('marks a delivery whose write threw, and forgets it when the process exits', async () => { diff --git a/src/main/sessionManager.ts b/src/main/sessionManager.ts index 6b744d444..18dfe49ce 100644 --- a/src/main/sessionManager.ts +++ b/src/main/sessionManager.ts @@ -4116,7 +4116,20 @@ export class SessionManager extends EventEmitter { } // A failure that wrote nothing (refused before write) changes nothing: // whatever the composer held before still holds, ours or not. - if (!delivery.promptWritten || delivery.enterWritten) return + if (!delivery.promptWritten) return + // WHY a delivery that wrote retires the standing mark before deciding + // whether to set its own (#1358 verification a, blocker): the old mark + // described bytes this delivery has already reclaimed or written over. + // Keeping it let an image delivery that consumed a text mark and then + // stranded hand that stale mark to the next delivery, which Ctrl+U'd a + // composer holding image pills. From here the composer holds what THIS + // delivery left, so only this delivery's outcome may mark it. Deleting is + // the safe direction: no mark means the gate treats the composer as a + // human draft and refuses, as it did before #1350. + this.strandedDeliveries.delete(sessionId) + // Enter went out: the composer was submitted (or its fate is unknown), so + // it no longer provably holds only our unsubmitted text. + if (delivery.enterWritten) return // WHY an image delivery is never marked (#1358 reviews a and c): its // leftovers include image pills, and whether Ctrl+U removes a pill has // not been observed on a real composer (promptDelivery.ts says the same From 38407b03500639e272216cd0953da1ea20bba5b8 Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sat, 26 Sep 2026 23:09:19 -0700 Subject: [PATCH 11/48] test(performance): pin all-pairs history persistence and the exact catalog thresholds #1352 verification b: restoring the store's 100-entry persistence slice and raising each catalog threshold by 1 ms both survived. The store now has an all-pairs operations.json test, and the threshold cases sample exactly at the threshold. Co-Authored-By: Claude Opus 5.5 --- src/main/performance/IncidentEngine.test.ts | 8 ++++--- .../performance/MonitorHistoryStore.test.ts | 23 +++++++++++++++++++ 2 files changed, 28 insertions(+), 3 deletions(-) diff --git a/src/main/performance/IncidentEngine.test.ts b/src/main/performance/IncidentEngine.test.ts index 9cf3a2c5e..54a9a8770 100644 --- a/src/main/performance/IncidentEngine.test.ts +++ b/src/main/performance/IncidentEngine.test.ts @@ -136,14 +136,16 @@ describe('incident evidence and exclusion rules', () => { // #1352 reviews a, b, c: the three catalog thresholds are the point of the // PR, and moving any of them left every test green. Each gets a sample just - // above (a named slow-operation incident) and just below (none). + // above (a named slow-operation incident) and just below (none). The + // "above" sample sits exactly AT the threshold (the rule is >=), so moving + // a threshold up by even 1 ms fails (#1352 verification b). it.each([ ['conversations.extract', 250], ['conversations.search', 1000], ['conversations.discover', 1000], - ] as const)('raises a slow-operation incident for %s only above %i ms', (name, threshold) => { + ] as const)('raises a slow-operation incident for %s from %i ms', (name, threshold) => { const above = new IncidentEngine() - above.accept([{ kind: 'operation', at: 1000, windowId: null, sample: { kind: 'operation', name, durationMs: threshold + 1, outcome: 'success' } }], 1000, 1000) + above.accept([{ kind: 'operation', at: 1000, windowId: null, sample: { kind: 'operation', name, durationMs: threshold, outcome: 'success' } }], 1000, 1000) expect(above.summaries().map(row => [row.rule, row.operation])).toEqual([['slow-operation', name]]) const below = new IncidentEngine() below.accept([{ kind: 'operation', at: 1000, windowId: null, sample: { kind: 'operation', name, durationMs: threshold - 1, outcome: 'success' } }], 1000, 1000) diff --git a/src/main/performance/MonitorHistoryStore.test.ts b/src/main/performance/MonitorHistoryStore.test.ts index ea27de350..d792e8e6e 100644 --- a/src/main/performance/MonitorHistoryStore.test.ts +++ b/src/main/performance/MonitorHistoryStore.test.ts @@ -54,6 +54,29 @@ describe('bounded local performance history', () => { expect((await store.query(20_000, 20_000, undefined, 7)).points).toHaveLength(1) }) + // #1352 verification b: the store's persistence slice is the third place + // that capped operation histograms at a literal 100 (26 operations x 4 + // outcomes = 104). The aggregator test cannot see this path; only the + // saved operations.json shows whether every legal pair reaches history. + it('persists a histogram for every legal operation and outcome', async () => { + const { MONITOR_OPERATIONS, MONITOR_OUTCOMES } = await import('@shared/performance/monitorPolicy.js') + const { MonitorAggregator } = await import('./MonitorAggregator.js') + const aggregator = new MonitorAggregator() + for (const name of MONITOR_OPERATIONS) { + for (const outcome of MONITOR_OUTCOMES) { + aggregator.accept([{ kind: 'operation', at: 1000, windowId: null, sample: { kind: 'operation', name, outcome, durationMs: 5 } }]) + } + } + const root = await mkdtemp(join(tmpdir(), 'agent-code-monitor-')) + roots.push(root) + const store = new MonitorHistoryStore(root, 'run-ops', () => 2000) + store.record({ ...snapshot(2000), operations: aggregator.snapshot(2000, 100).operations }, null, null, 0, 0) + await store.settled() + const saved = JSON.parse(await readFile(join(root, 'runs', 'run-ops', 'operations.json'), 'utf8')) as Array<{ name: string; outcome: string }> + expect(saved).toHaveLength(MONITOR_OPERATIONS.length * MONITOR_OUTCOMES.length) + expect(saved.map(entry => `${entry.name}:${entry.outcome}`)).toContain('heap.snapshot:timeout') + }) + it('keeps an existing destination intact when report creation fails', async () => { const root = await mkdtemp(join(tmpdir(), 'agent-code-monitor-')) roots.push(root) From 1c6a28b63beaea6a7c47fed7bba7aeeacad0bd58 Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sat, 26 Sep 2026 23:13:30 -0700 Subject: [PATCH 12/48] docs(plan): closed orchestration children follow their parent's replacement (#1283 item 1) Co-Authored-By: Claude Opus 5.5 --- ...27-orchestration-tombstone-parent-carry.md | 28 +++++++++++++++++++ 1 file changed, 28 insertions(+) create mode 100644 docs/plans/2026-09-27-orchestration-tombstone-parent-carry.md diff --git a/docs/plans/2026-09-27-orchestration-tombstone-parent-carry.md b/docs/plans/2026-09-27-orchestration-tombstone-parent-carry.md new file mode 100644 index 000000000..e46ce23fe --- /dev/null +++ b/docs/plans/2026-09-27-orchestration-tombstone-parent-carry.md @@ -0,0 +1,28 @@ +# Closed orchestration children follow their parent's replacement (#1283 item 1) + +## Problem +`OrchestrationBridge` keeps a small in-memory tombstone for every child closed through the orchestration MCP (`closedAgents`, 24 h / 500 entries). This is how a parent can still see "closed" as a state after the renderer removes the pane. A tombstone matches a parent by its frozen `orchestrationParentId` / `orchestrationRootId`. + +When the parent P is replaced by P′, the renderer remaps every LIVE child's pointers to P′ (`remapSessionsRelationships`). This covers replace, provider switch, resume, rewind, Reload Agents and rehydrate. Main is never told, so its tombstones still say P: +- P′'s `list_agents` / `read_run_outputs` omit every closed child; +- `read_agent` on a closed child returns not-found; +- the `parentSessionByChildSession` hints still point at P, so a child's prompt boundary invalidates P's status cache instead of P′'s. + +## Evidence +- `OrchestrationBridge.agentMatchesParent` compares the stored ids directly. +- No IPC or main path ever learns the renderer's old→new session id map for orchestration. `workflows:carry-session` is the only such channel, and it deliberately skips a newConversation swap, which live children do NOT skip. +- Found by the Stage 3 C2 hunt (temp/quality-loop/hunt-c2.md, row 7). + +## Decisions (defaults) +- **One new IPC, `orchestration:carry-parent {from, to}`.** It is called at exactly the sites that remap live children's pointers, and it is fire-and-forget like `carryWorkflowRuns`: a failed carry leaves the pre-fix behaviour, and never blocks a swap. It does not piggyback on `workflows:carry-session`, because that one follows the conversation (it is skipped for newConversation), while orchestration pointers follow the PANE. +- **Main rewrites at carry time.** Stored tombstones whose parent or root is `from` now say `to`, so a parent reading them sees consistent ids. The same goes for `parentSessionByChildSession` hints, and both status caches are invalidated. +- **Plus a bounded alias chain for tombstones created later.** `closeAgent` reads the child BEFORE it asks the renderer to close it. A swap in between hands `noteClosed` a record that still says P. `noteClosed` resolves the ids through the aliases (following P→P′→P″, with a hop cap against cycles). Aliases share the tombstones' TTL and size cap. +- **Trust:** the same as `workflows:carry-session` and `goal-loop:carry`. An application window can already list, read and close any orchestration child it can see. +- **Out of scope:** items 2 (Codex handoff lease, #1338) and 3 (PTY attach counts) of #1283. So this PR says `Refs #1283`, not `Fixes`. + +## Tests +- `OrchestrationBridge.test.ts` (fail-first): + - a tombstone of P is listed, read and returned by `read_run_outputs` for P′ after `carryParent(P, P′)`; + - a close whose pre-read happened before the carry is still found under P′; + - a two-hop lineage works, and a cycle terminates. +- Renderer: `carryOrchestrationParents` is called with the committed idMap on replace and reload, and not for a successor killed as an orphan. From c518c09e1d10043e4e58fc07578bfc8a32d085e6 Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sat, 26 Sep 2026 23:16:47 -0700 Subject: [PATCH 13/48] fix(orchestration): closed children follow their parent's replacement in main Tombstones of MCP-closed children named the parent's old session id, so the parent's successor could not list, read or collect them. carryParent rewrites stored tombstones and child->parent hints, and keeps a bounded, acyclic alias chain for a close whose read predates the swap. Co-Authored-By: Claude Opus 5.5 --- .../OrchestrationBridge.parentCarry.test.ts | 126 ++++++++++++++++++ src/main/orchestration/OrchestrationBridge.ts | 83 +++++++++++- 2 files changed, 208 insertions(+), 1 deletion(-) create mode 100644 src/main/orchestration/OrchestrationBridge.parentCarry.test.ts diff --git a/src/main/orchestration/OrchestrationBridge.parentCarry.test.ts b/src/main/orchestration/OrchestrationBridge.parentCarry.test.ts new file mode 100644 index 000000000..42c0fd00b --- /dev/null +++ b/src/main/orchestration/OrchestrationBridge.parentCarry.test.ts @@ -0,0 +1,126 @@ +import { expect, it, vi } from 'vitest' + +// #1283 item 1. A child closed through the orchestration MCP leaves a +// tombstone in main that names its parent P. When P is replaced by P' (reload, +// provider switch, resume, rewind, Reload Agents, rehydrate), the renderer +// remaps every LIVE child to P' but main was never told, so P' lost sight of +// every closed child: list_agents and read_run_outputs omitted them, and +// read_agent said not-found. +const sent: Array<{ requestId: string; type: string; parentSessionId?: string; sessionId?: string }> = [] + +vi.mock('@main/window/windowRegistry.js', () => ({ + sendToWindow: (_windowId: string, _channel: string, request: never) => { + sent.push(request) + return true + }, + windowForSession: () => 'test-window', +})) + +const { OrchestrationBridge } = await import('@main/orchestration/OrchestrationBridge.js') +type Bridge = InstanceType + +const agent = (sessionId: string, parent: string, root = parent) => ({ + sessionId, kind: 'claude' as const, cwd: '/tmp/project', + orchestrationParentId: parent, orchestrationRootId: root, orchestrationRunId: 'run-1', + lifecycleState: 'completed' as const, statusSummary: 'done', +}) + +async function next(type: string) { + await vi.waitFor(() => expect(sent.some(request => request.type === type)).toBe(true)) + const index = sent.findIndex(request => request.type === type) + return sent.splice(index, 1)[0]! +} + +/** The real close flow: main reads the child first, then asks the renderer to + * close it, and tombstones the read. `beforeClose` runs between the two. */ +async function closeChild(bridge: Bridge, parent: string, child: ReturnType, beforeClose?: () => void) { + const closing = bridge.closeAgent({ parentSessionId: parent, sessionId: child.sessionId }) + const read = await next('read-agent') + bridge.resolve({ requestId: read.requestId, ok: true, type: 'read-agent', output: { agent: child, messages: [], finalAssistantText: 'finished' } } as never) + const close = await next('close-agent') + beforeClose?.() + bridge.resolve({ requestId: close.requestId, ok: true, type: 'close-agent', result: { closedSessionIds: [child.sessionId] } } as never) + await closing +} + +async function listFor(bridge: Bridge, parent: string) { + const listing = bridge.listAgents({ parentSessionId: parent }) + const request = await next('list-agents') + bridge.resolve({ requestId: request.requestId, ok: true, type: 'list-agents', agents: [] } as never) + return (await listing).map(row => [row.sessionId, row.orchestrationParentId, row.lifecycleState]) +} + +async function readFor(bridge: Bridge, parent: string, sessionId: string) { + const reading = bridge.readAgent({ parentSessionId: parent, sessionId }) + const request = await next('read-agent') + // The pane is gone, so the renderer cannot find the child. + bridge.resolve({ requestId: request.requestId, ok: false, type: 'read-agent', message: 'not found' } as never) + return await reading +} + +it('shows a closed child to the parent\'s successor after the parent is replaced', async () => { + sent.length = 0 + const bridge = new OrchestrationBridge() + await closeChild(bridge, 'parent-a', agent('child-1', 'parent-a')) + + bridge.carryParent('parent-a', 'parent-b') + + expect(await listFor(bridge, 'parent-b')).toEqual([['child-1', 'parent-b', 'closed']]) + expect((await readFor(bridge, 'parent-b', 'child-1')).agent).toMatchObject({ sessionId: 'child-1', orchestrationParentId: 'parent-b', orchestrationRootId: 'parent-b' }) + const outputs = bridge.readRunOutputs({ parentSessionId: 'parent-b', runId: 'run-1' }) + const request = await next('read-run-outputs') + bridge.resolve({ requestId: request.requestId, ok: true, type: 'read-run-outputs', outputs: [] } as never) + expect((await outputs).map(output => output.agent.sessionId)).toEqual(['child-1']) + // The replaced id no longer owns it. + expect(await listFor(bridge, 'parent-a')).toEqual([]) +}) + +// closeAgent reads the child BEFORE the renderer closes it; a replacement in +// between hands the tombstone a record that still names the old parent. +it('files a close that straddles the replacement under the successor', async () => { + sent.length = 0 + const bridge = new OrchestrationBridge() + await closeChild(bridge, 'parent-a', agent('child-1', 'parent-a'), () => bridge.carryParent('parent-a', 'parent-b')) + + expect(await listFor(bridge, 'parent-b')).toEqual([['child-1', 'parent-b', 'closed']]) +}) + +it('follows a parent replaced twice, and stops on a cycle', async () => { + sent.length = 0 + const bridge = new OrchestrationBridge() + bridge.carryParent('parent-a', 'parent-b') + bridge.carryParent('parent-b', 'parent-c') + // A grandchild rooted at the replaced parent, closed after both swaps from a + // stale read. + await closeChild(bridge, 'child-1', agent('grandchild', 'child-1', 'parent-a')) + expect((await readFor(bridge, 'child-1', 'grandchild')).agent).toMatchObject({ orchestrationParentId: 'child-1', orchestrationRootId: 'parent-c' }) + + // parent-a comes back as a live pane (Undo Close can do this). Its old edge + // out (a -> b) must go: otherwise a child it closes now would be filed + // under b, and the edges would form a loop. + bridge.carryParent('parent-c', 'parent-a') + await closeChild(bridge, 'parent-c', agent('child-2', 'parent-c')) + await closeChild(bridge, 'parent-a', agent('child-3', 'parent-a')) + expect(await listFor(bridge, 'parent-a')).toEqual([ + ['grandchild', 'child-1', 'closed'], ['child-2', 'parent-a', 'closed'], ['child-3', 'parent-a', 'closed'], + ]) +}) + +// The child -> parent hint decides whose status cache a child's prompt +// boundary invalidates. After the swap it must be the successor's, or the +// successor keeps serving a status computed before the child changed. +it('invalidates the successor\'s status cache when a child of the replaced parent changes', async () => { + sent.length = 0 + const bridge = new OrchestrationBridge() + const created = bridge.createAgent({ parentSessionId: 'parent-a', kind: 'claude' }) + const create = await next('create-agent') + bridge.resolve({ requestId: create.requestId, ok: true, type: 'create-agent', agent: agent('child-1', 'parent-a') } as never) + await created + bridge.carryParent('parent-a', 'parent-b') + // Fill parent-b's cache. Another (unrelated) child's unknown hint would + // clear every cache, so only child-1's hint is exercised. + await listFor(bridge, 'parent-b') + bridge.notePromptSubmitted('child-1') + void bridge.listAgents({ parentSessionId: 'parent-b' }) + await vi.waitFor(() => expect(sent.some(request => request.type === 'list-agents' && request.parentSessionId === 'parent-b')).toBe(true)) +}) diff --git a/src/main/orchestration/OrchestrationBridge.ts b/src/main/orchestration/OrchestrationBridge.ts index 28b188b85..313c332b0 100644 --- a/src/main/orchestration/OrchestrationBridge.ts +++ b/src/main/orchestration/OrchestrationBridge.ts @@ -255,6 +255,21 @@ export class OrchestrationBridge { private readonly createCallsInFlight = new Map>() private readonly closedAgents = new Map() private readonly parentSessionByChildSession = new Map() + /** + * Replaced parent id -> its successor (#1283 item 1). The renderer remaps a + * LIVE child's orchestrationParentId/RootId when its parent pane gets a new + * session id (reload, provider switch, resume, rewind, Reload Agents, + * rehydrate), but tombstones live only here, so carryParent is how main + * hears about it. + * + * WHY an alias map on top of rewriting the tombstones in carryParent: a + * rewrite only reaches tombstones that already exist. closeAgent reads the + * child BEFORE the renderer closes it, so a replacement landing between the + * two hands noteClosed a record that still names the old parent; noteClosed + * resolves it through these edges. Bounded by the tombstones' own TTL and + * cap: an alias older than every tombstone it could still serve is useless. + */ + private readonly replacedParents = new Map() private readonly listAgentsCache = new Map>() private readonly readRunOutputsCache = new Map>() private lastPrunedAt = 0 @@ -947,10 +962,70 @@ export class OrchestrationBridge { } } + /** + * The parent pane `from` now runs as session `to` (#1283 item 1): closed + * children filed under `from` belong to `to`. Called by the renderer at the + * same commit points where it remaps live children (see replacedParents). + * + * WHY not piggyback on workflows:carry-session: workflow runs follow the + * CONVERSATION and are deliberately not carried when a different + * conversation is swapped into the pane; orchestration pointers follow the + * PANE, and the renderer remaps live children in both cases. One channel + * cannot honour both rules. + */ + carryParent(from: unknown, to: unknown): void { + // Clone-boundary input from the renderer; a malformed carry must be a + // no-op, never an alias keyed by `undefined`. + if (typeof from !== 'string' || typeof to !== 'string' || !from || !to || from === to) return + // `to` is a live pane again (it may itself have been replaced earlier and + // come back, e.g. Undo Close). An edge OUT of it would now send its new + // children elsewhere, and dropping it is also what keeps the edges acyclic: + // after this, `to` reaches nothing, so `from -> to` cannot close a loop. + this.replacedParents.delete(to) + this.replacedParents.delete(from) + this.replacedParents.set(from, { to, at: Date.now() }) + for (const record of this.closedAgents.values()) { + const agent = record.output.agent + if (agent.orchestrationParentId !== from && agent.orchestrationRootId !== from) continue + record.output = { + ...record.output, + agent: { + ...agent, + orchestrationParentId: agent.orchestrationParentId === from ? to : agent.orchestrationParentId, + orchestrationRootId: agent.orchestrationRootId === from ? to : agent.orchestrationRootId, + }, + } + } + // The prompt-boundary hint: a child of `from` changing state must + // invalidate `to`'s status cache, the one its parent now polls. + for (const [child, parent] of this.parentSessionByChildSession) { + if (parent === from) this.parentSessionByChildSession.set(child, to) + } + this.invalidateStatusCache(from) + this.invalidateStatusCache(to) + this.pruneCoordinationMetadata() + } + + /** The live successor of a possibly replaced parent id (see replacedParents). */ + private currentParentId(sessionId: string): string { + let current = sessionId + // carryParent keeps the edges acyclic; the cap is a backstop so a future + // edit that breaks that can never hang main. + for (let hops = 0; hops <= this.replacedParents.size; hops++) { + const next = this.replacedParents.get(current) + if (!next) return current + current = next.to + } + return current + } + private noteClosed(output: OrchestrationAgentOutput): void { const closedAt = Date.now() const agent: OrchestrationAgentRecord = { ...this.enrichAgent(output.agent), + // The read may predate a replacement of the parent (see replacedParents). + orchestrationParentId: this.currentParentId(output.agent.orchestrationParentId), + orchestrationRootId: this.currentParentId(output.agent.orchestrationRootId), lifecycleState: 'closed', completedAt: output.agent.completedAt ?? output.agent.lastActivityAt, lastActivityAt: closedAt, @@ -1150,7 +1225,8 @@ export class OrchestrationBridge { if ( now - this.lastPrunedAt < PRUNE_INTERVAL_MS && this.promptDeliveries.size <= MAX_PROMPT_DELIVERIES && - this.closedAgents.size <= MAX_CLOSED_AGENTS + this.closedAgents.size <= MAX_CLOSED_AGENTS && + this.replacedParents.size <= MAX_CLOSED_AGENTS ) { return } @@ -1184,6 +1260,11 @@ export class OrchestrationBridge { } } trimMapToNewest(this.closedAgents, MAX_CLOSED_AGENTS) + // Same lifetime as the tombstones they serve (see replacedParents). + for (const [from, edge] of this.replacedParents) { + if (now - edge.at > ORCHESTRATION_METADATA_TTL_MS) this.replacedParents.delete(from) + } + trimMapToNewest(this.replacedParents, MAX_CLOSED_AGENTS) } } From f5d44f36931d528ee94546e709e8d18b0ad10bdd Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sat, 26 Sep 2026 23:19:57 -0700 Subject: [PATCH 14/48] fix(orchestration): the renderer carries closed children to the parent's successor orchestration:carry-parent, called wherever live children are remapped (replace incl. newConversation, Reload Agents) and on Undo Close, only for committed successors. Co-Authored-By: Claude Opus 5.5 --- src/main/ipc/orchestration.ts | 11 +++++++ src/preload/api/orchestration.ts | 5 ++++ .../reloadAgentsOwnership.renderer.test.tsx | 7 +++-- .../replacementOwnership.renderer.test.tsx | 10 +++++-- .../src/workspace/hook/actions/session.ts | 6 +++- .../workspace/hook/actions/successorCarry.ts | 29 +++++++++++++++++++ .../src/workspace/hook/actions/undoClose.ts | 5 +++- .../undoCloseSuccessorCarry.renderer.test.tsx | 8 +++-- 8 files changed, 72 insertions(+), 9 deletions(-) diff --git a/src/main/ipc/orchestration.ts b/src/main/ipc/orchestration.ts index 6e1379c70..39af17694 100644 --- a/src/main/ipc/orchestration.ts +++ b/src/main/ipc/orchestration.ts @@ -11,4 +11,15 @@ export function registerOrchestrationIpc(bridge: OrchestrationBridge): void { return true }, ) + // #1283 item 1: the renderer commits every pane replacement and remaps its + // LIVE orchestration children itself; only it can tell main that the + // tombstones of closed children follow too. Same trust as + // workflows:carry-session and goal-loop:carry: an application window can + // already list, read and close any orchestration child it can see. + ipcMain.handle( + 'orchestration:carry-parent', + (_evt, request: { from?: unknown; to?: unknown } | undefined) => { + bridge.carryParent(request?.from, request?.to) + }, + ) } diff --git a/src/preload/api/orchestration.ts b/src/preload/api/orchestration.ts index 9c4e42ef3..92859f9cf 100644 --- a/src/preload/api/orchestration.ts +++ b/src/preload/api/orchestration.ts @@ -15,4 +15,9 @@ export const orchestrationApi = { resolveOrchestrationRequest: ( response: OrchestrationRendererResponse, ): Promise => ipcRenderer.invoke('orchestration:response', response), + + // #1283 item 1: the parent pane `from` now runs as `to`; main's tombstones + // of its MCP-closed children follow (see OrchestrationBridge.carryParent). + carryOrchestrationParent: (from: string, to: string): Promise => + ipcRenderer.invoke('orchestration:carry-parent', { from, to }), } diff --git a/src/renderer/src/workspace/hook/actions/reloadAgentsOwnership.renderer.test.tsx b/src/renderer/src/workspace/hook/actions/reloadAgentsOwnership.renderer.test.tsx index 64633a415..92f6da56d 100644 --- a/src/renderer/src/workspace/hook/actions/reloadAgentsOwnership.renderer.test.tsx +++ b/src/renderer/src/workspace/hook/actions/reloadAgentsOwnership.renderer.test.tsx @@ -67,11 +67,12 @@ function harness(heldKind: 'claude' | 'codex' = 'claude', opts: { heldSpawnRejec const killHolds = new Map>() const killOwnedSession = vi.fn(async (req: { sessionId: string }) => killHolds.get(req.sessionId) ?? true) const carryWorkflowRuns = vi.fn(async (_from: string, _to: string) => undefined) - window.api = { ...originalApi, spawnSession, killOwnedSession, controlGoalLoop: vi.fn(async () => null), carryGoalLoop: vi.fn(async () => null), carryWorkflowRuns } + const carryOrchestrationParent = vi.fn(async (_from: string, _to: string) => undefined) + window.api = { ...originalApi, spawnSession, killOwnedSession, controlGoalLoop: vi.fn(async () => null), carryGoalLoop: vi.fn(async () => null), carryWorkflowRuns, carryOrchestrationParent } const hook = renderHook(() => useSessionActions(state, writer.setState, setRuntimes, refs)) // The Claude agent is first in the snapshot, so it is the one in flight. const order = Object.keys(recorded.sessions).filter(id => id === claudeLane || id === codexLane) - return { hook, writer, refs, spawnSession, killOwnedSession, killHolds, release, claudeLane, codexLane, order, carryWorkflowRuns } + return { hook, writer, refs, spawnSession, killOwnedSession, killHolds, release, claudeLane, codexLane, order, carryWorkflowRuns, carryOrchestrationParent } } it('does not bring back an agent closed while its respawn was in flight', async () => { @@ -111,6 +112,8 @@ it('carries workflow runs only to committed successors, never to an orphan', asy await act(async () => { h.release(); await reload }) expect(h.killOwnedSession).toHaveBeenCalledWith(expect.objectContaining({ sessionId: 'claude-restarted', caller: 'reload.orphaned-successor' })) expect(h.carryWorkflowRuns.mock.calls).toEqual([[h.codexLane, 'codex-restarted']]) + // #1283 item 1: the closed children of the committed one only. + expect(h.carryOrchestrationParent.mock.calls).toEqual([[h.codexLane, 'codex-restarted']]) }) it('does not double an agent replaced while its respawn was in flight', async () => { diff --git a/src/renderer/src/workspace/hook/actions/replacementOwnership.renderer.test.tsx b/src/renderer/src/workspace/hook/actions/replacementOwnership.renderer.test.tsx index 8a6123246..087b8f38f 100644 --- a/src/renderer/src/workspace/hook/actions/replacementOwnership.renderer.test.tsx +++ b/src/renderer/src/workspace/hook/actions/replacementOwnership.renderer.test.tsx @@ -53,10 +53,11 @@ function replaceHarness(defaults: string[]) { refs.latestRuntimesRef.current = useAppStore.getState().workspaceRuntimes const carryGoalLoop = vi.fn(async (_from: string, _to: string) => null) const carryWorkflowRuns = vi.fn(async (_from: string, _to: string) => undefined) + const carryOrchestrationParent = vi.fn(async (_from: string, _to: string) => undefined) const controlGoalLoop = vi.fn(async (_request: { sessionId: string; action: string }) => null) - window.api = { ...originalApi, spawnSession: vi.fn(async () => ({ sessionId: 'successor' })), killOwnedSession: vi.fn(async () => true), carryGoalLoop, carryWorkflowRuns, controlGoalLoop } + window.api = { ...originalApi, spawnSession: vi.fn(async () => ({ sessionId: 'successor' })), killOwnedSession: vi.fn(async () => true), carryGoalLoop, carryWorkflowRuns, carryOrchestrationParent, controlGoalLoop } const mounted = renderHook(() => useSessionActions(state, useAppStore.getState().setWorkspaceState, useAppStore.getState().setWorkspaceRuntimes, refs)) - return { mounted, carryGoalLoop, carryWorkflowRuns, controlGoalLoop } + return { mounted, carryGoalLoop, carryWorkflowRuns, carryOrchestrationParent, controlGoalLoop } } it('hands the pane\'s goal loop to the successor when the replacement commits', async () => { @@ -149,9 +150,12 @@ it('hands the pane\'s workflow runs to the successor of the same conversation', }) it('does not hand workflow runs to a different conversation swapped into the pane', async () => { - const { mounted, carryWorkflowRuns } = replaceHarness(['goal_loop']) + const { mounted, carryWorkflowRuns, carryOrchestrationParent } = replaceHarness(['goal_loop']) await act(async () => { await mounted.result.current.replaceSession('/recorded/project', { targetSessionId: 'source', kind: 'claude', resumeSessionId: 'other-conversation', newConversation: true }) }) expect(carryWorkflowRuns).not.toHaveBeenCalled() + // #1283 item 1: orchestration children follow the PANE, as the live ones + // remapped in the same commit do, so the closed ones are carried even here. + expect(carryOrchestrationParent).toHaveBeenCalledWith('source', 'successor') }) diff --git a/src/renderer/src/workspace/hook/actions/session.ts b/src/renderer/src/workspace/hook/actions/session.ts index 3d68b7db0..fe7de6457 100644 --- a/src/renderer/src/workspace/hook/actions/session.ts +++ b/src/renderer/src/workspace/hook/actions/session.ts @@ -1,5 +1,5 @@ import { tldrIdentityForReplacement, tldrIdentityForSession } from '@renderer/features/tldr/identity' -import { carryGoalLoops, carryWorkflowRuns, stopGoalLoops } from '@renderer/workspace/hook/actions/successorCarry' +import { carryGoalLoops, carryOrchestrationParents, carryWorkflowRuns, stopGoalLoops } from '@renderer/workspace/hook/actions/successorCarry' import { hasReportingDomain } from '@shared/types/tldr' import { getRendererProviderCapabilities } from '@providers/registry.renderer.capabilities' import { @@ -1458,6 +1458,9 @@ export function useSessionActions( // (#1280). A different conversation swapped into the pane does not // inherit them: they belong to the conversation that started them. if (!opts?.newConversation) carryWorkflowRuns(idMap) + // Closed orchestration children follow the PANE, like the live ones + // remapped in the commit above, newConversation included (#1283). + carryOrchestrationParents(idMap) setRuntimes(prev => { // Replacement can await spawn and backend retirement while the user // keeps editing. Transfer the latest draft in the same state update @@ -1737,6 +1740,7 @@ export function useSessionActions( // reaches this line, so its predecessor's runs are never handed to a // process that is being killed. carryWorkflowRuns(new Map([[oldId, newId]])) + carryOrchestrationParents(new Map([[oldId, newId]])) if (hasDurableProviderSession(fresh)) { void loadInitialHistoryForSession({ sessionId: newId, meta: fresh, refs, setRuntimes }) } diff --git a/src/renderer/src/workspace/hook/actions/successorCarry.ts b/src/renderer/src/workspace/hook/actions/successorCarry.ts index 45be27e04..cad00b1f0 100644 --- a/src/renderer/src/workspace/hook/actions/successorCarry.ts +++ b/src/renderer/src/workspace/hook/actions/successorCarry.ts @@ -32,6 +32,35 @@ export function carryWorkflowRuns(idMap: ReadonlyMap): void { } } +/** + * Tell main that each replaced pane's MCP-closed orchestration children now + * belong to its successor (#1283 item 1), so the successor still lists, reads + * and collects them. + * + * WHY every site that remaps live children, newConversation included, unlike + * carryWorkflowRuns: runs follow the CONVERSATION, but orchestration pointers + * follow the PANE (remapSessionsRelationships rewrites every live child's + * orchestrationParentId/RootId on every committed swap). Main's tombstones + * are the closed half of the same relationship, so they take the same rule. + * + * WHY not rehydrate: local session ids are stable across a renderer reload + * (rehydrate.ts: remapping them duplicated live backends), and a full restart + * starts main with no tombstones. There is nothing to carry. + * + * Fire-and-forget: a failed carry leaves the tombstones under the old id (the + * pre-fix behaviour) and never blocks the swap. + */ +export function carryOrchestrationParents(idMap: ReadonlyMap): void { + const carry = window.api?.carryOrchestrationParent + if (!carry) return + for (const [oldId, newId] of idMap) { + if (oldId === newId) continue + void carry(oldId, newId).catch(error => { + console.warn('[orchestration] carry to the replacement session failed:', error) + }) + } +} + /** End the loop of each replaced pane that did NOT get it carried (#1287 * review A2): its old id is gone from the workspace, so no pane could ever * resume or stop it, and a successor without goal_loop could not complete diff --git a/src/renderer/src/workspace/hook/actions/undoClose.ts b/src/renderer/src/workspace/hook/actions/undoClose.ts index a6987a51c..9310e3c88 100644 --- a/src/renderer/src/workspace/hook/actions/undoClose.ts +++ b/src/renderer/src/workspace/hook/actions/undoClose.ts @@ -1,4 +1,4 @@ -import { carryWorkflowRuns, handOverGoalLoops, stopGoalLoops } from '@renderer/workspace/hook/actions/successorCarry' +import { carryOrchestrationParents, carryWorkflowRuns, handOverGoalLoops, stopGoalLoops } from '@renderer/workspace/hook/actions/successorCarry' import { carriedRelationships } from '@renderer/workspace/idRemap' import { sessionDisplayTitle } from '@renderer/workspace/sessionDisplayTitle' import { sessionMcpOverrides } from '@renderer/workspace/mcpDomains' @@ -460,6 +460,9 @@ export function useUndoCloseAction( : await restoreTabEntry(entry, recordingPublish) if (result === 'retryable-failure') return result carryWorkflowRuns(successors) + // The restored pane resumes the same conversation that closed those + // children through the MCP; it should still see them (#1283). + carryOrchestrationParents(successors) handOverGoalLoops(successors, hasGoalLoopTools) const closedIds = entry.type === 'session' ? [entry.sessionId] diff --git a/src/renderer/src/workspace/hook/actions/undoCloseSuccessorCarry.renderer.test.tsx b/src/renderer/src/workspace/hook/actions/undoCloseSuccessorCarry.renderer.test.tsx index 24e4de49e..a643f18df 100644 --- a/src/renderer/src/workspace/hook/actions/undoCloseSuccessorCarry.renderer.test.tsx +++ b/src/renderer/src/workspace/hook/actions/undoCloseSuccessorCarry.renderer.test.tsx @@ -28,9 +28,10 @@ function mount(entry: unknown, makeSpawn: (writer: Writer) => ReturnType undefined) + const carryOrchestrationParent = vi.fn(async (_from: string, _to: string) => undefined) const carryGoalLoop = vi.fn(async (_from: string, _to: string) => null) const controlGoalLoop = vi.fn(async (_request: { sessionId: string; action: string }) => null) - window.api = { ...originalApi, carryWorkflowRuns, carryGoalLoop, controlGoalLoop, killOwnedSession: vi.fn(async () => true) } + window.api = { ...originalApi, carryWorkflowRuns, carryOrchestrationParent, carryGoalLoop, controlGoalLoop, killOwnedSession: vi.fn(async () => true) } const spawn = makeSpawn(writer) let actions!: ReturnType function Harness(): React.JSX.Element { @@ -38,7 +39,7 @@ function mount(entry: unknown, makeSpawn: (writer: Writer) => ReturnType } const mounted = render() - return { carryWorkflowRuns, carryGoalLoop, controlGoalLoop, writer, refs, undo: () => actions.undoClose(), unmount: () => mounted.unmount() } + return { carryWorkflowRuns, carryOrchestrationParent, carryGoalLoop, controlGoalLoop, writer, refs, undo: () => actions.undoClose(), unmount: () => mounted.unmount() } } it('hands a restored pane the workflow runs of the pane it restores', async () => { @@ -48,6 +49,8 @@ it('hands a restored pane the workflow runs of the pane it restores', async () = }, () => vi.fn().mockResolvedValue('restored-pane')) await act(async () => { await harness.undo() }) expect(harness.carryWorkflowRuns).toHaveBeenCalledWith('closed-pane', 'restored-pane') + // #1283 item 1: and the children it closed through the orchestration MCP. + expect(harness.carryOrchestrationParent).toHaveBeenCalledWith('closed-pane', 'restored-pane') harness.unmount() }) @@ -135,6 +138,7 @@ it('ends the loop when the restore bails and consumes the entry', async () => { await act(async () => { await harness.undo() }) expect(harness.carryGoalLoop).not.toHaveBeenCalled() expect(harness.carryWorkflowRuns).not.toHaveBeenCalled() + expect(harness.carryOrchestrationParent).not.toHaveBeenCalled() expect(harness.controlGoalLoop.mock.calls).toEqual([[{ sessionId: 'closed-pane', action: 'stop' }]]) expect(harness.refs.undoStackRef.current.length).toBe(0) harness.unmount() From 770e6922edbc03d5185c2fd9b094bf50aaa09e77 Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sat, 26 Sep 2026 23:25:34 -0700 Subject: [PATCH 15/48] test(claude): pin the stranded paint window against the recorded lag #1358 verification b: the only window test reinstalled fake timers, which reset the clock and put strandedAt in the future, so a 500 ms window survived. One clock throughout; the next delivery starts 6 s after the failure and the text paints at +7 s. Red with 500 ms and 6.5 s. Co-Authored-By: Claude Opus 5.5 --- .../sessionManager.strandedDelivery.test.ts | 23 +++++++++++++++++++ 1 file changed, 23 insertions(+) diff --git a/src/main/sessionManager.strandedDelivery.test.ts b/src/main/sessionManager.strandedDelivery.test.ts index 799578ad2..4088ea3d6 100644 --- a/src/main/sessionManager.strandedDelivery.test.ts +++ b/src/main/sessionManager.strandedDelivery.test.ts @@ -180,6 +180,29 @@ it('waits for the stranded text to paint before writing, then clears it', async expect(writes.slice(1)).toEqual(['\x15', 'the next task', '\r']) }) +// #1358 verification b: the window must cover the recorded lag, measured +// from the failure. One fake clock throughout (a second useFakeTimers() resets +// the clock, which put strandedAt in the future and let any positive window +// pass). The next delivery starts 6 s after the strand and the text paints at +// +7 s, inside the 0.7-3.8 s recorded lag's 8 s window; a window shorter than +// that writes the new prompt beside the unpainted bytes. +it('still waits for the late paint when the next delivery starts seconds after the strand', async () => { + const { manager, session, writes } = claudeLike() + vi.useFakeTimers() + const first = manager.deliverPromptToAgent('s1', 'an earlier prompt that painted late') + await vi.advanceTimersByTimeAsync(8_000) + expect(await first).toMatchObject({ ok: false, code: 'absorption-timeout' }) + // The failure, not the end of the advance above, is where the lag starts. + const { at: strandedAt } = (manager as unknown as { strandedDeliveries: Map }).strandedDeliveries.get('s1')! + await vi.advanceTimersByTimeAsync(strandedAt + 6_000 - Date.now()) + const next = manager.deliverPromptToAgent('s1', 'the next task') + await vi.advanceTimersByTimeAsync(1_000) + session.paintLate() + await vi.advanceTimersByTimeAsync(10_000) + await expect(next).resolves.toMatchObject({ ok: true }) + expect(writes.slice(1)).toEqual(['\x15', 'the next task', '\r']) +}) + // #1358 reviews a and c: whether Ctrl+U removes an image pill is not // established, so an image delivery's leftovers are never reclaimed. it('does not mark an image delivery that stranded', async () => { From f98eb890cc37ecfe55e72aadc01cb8902d5f2d6c Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sat, 26 Sep 2026 23:43:43 -0700 Subject: [PATCH 16/48] fix(orchestration): a child created across its parent's replacement is filed under the successor #1369 review a: createAgent's answer and adoptLateResponse stored the requested (replaced) parent as the child's hint and invalidated only its cache, so the successor served a stale empty list. Both now resolve through the alias chain. Review b: the preload->IPC->bridge boundary, the same-conversation replace carry and alias pruning are now pinned; every single mutation is red. Co-Authored-By: Claude Opus 5.5 --- src/main/ipc/orchestration.test.ts | 42 ++++++++++ .../OrchestrationBridge.parentCarry.test.ts | 77 +++++++++++++++++++ src/main/orchestration/OrchestrationBridge.ts | 28 +++++-- .../replacementOwnership.renderer.test.tsx | 5 +- 4 files changed, 145 insertions(+), 7 deletions(-) create mode 100644 src/main/ipc/orchestration.test.ts diff --git a/src/main/ipc/orchestration.test.ts b/src/main/ipc/orchestration.test.ts new file mode 100644 index 000000000..03355f558 --- /dev/null +++ b/src/main/ipc/orchestration.test.ts @@ -0,0 +1,42 @@ +import { expect, it, vi } from 'vitest' + +// #1369 review b: the carry crosses preload -> IPC -> OrchestrationBridge, and +// every other test injected one side (the renderer mocks window.api, the main +// tests call carryParent directly), so dropping the handler or the preload +// method left them all green. Here ONE electron mock routes ipcRenderer.invoke +// into whatever ipcMain.handle registered: the real preload method reaches +// the real handler and the real bridge, as in the app. +const handlers = vi.hoisted(() => new Map unknown>()) + +vi.mock('electron', () => ({ + ipcMain: { + handle: (channel: string, handler: (event: unknown, ...args: unknown[]) => unknown) => { handlers.set(channel, handler) }, + }, + ipcRenderer: { + invoke: async (channel: string, ...args: unknown[]) => { + const handler = handlers.get(channel) + if (!handler) throw new Error(`No handler registered for '${channel}'`) + return await handler({}, ...args) + }, + }, +})) +vi.mock('@main/window/windowRegistry.js', () => ({ sendToWindow: () => true, windowForSession: () => 'test-window' })) + +const { OrchestrationBridge } = await import('@main/orchestration/OrchestrationBridge.js') +const { registerOrchestrationIpc } = await import('./orchestration.js') +const { orchestrationApi } = await import('@preload/api/orchestration.js') + +it('carries a replaced parent from the renderer API into the bridge', async () => { + const bridge = new OrchestrationBridge() + registerOrchestrationIpc(bridge) + await orchestrationApi.carryOrchestrationParent('parent-a', 'parent-b') + expect((bridge as unknown as { replacedParents: Map }).replacedParents.get('parent-a')?.to).toBe('parent-b') +}) + +it('ignores a malformed carry instead of recording an alias for undefined', async () => { + const bridge = new OrchestrationBridge() + registerOrchestrationIpc(bridge) + await handlers.get('orchestration:carry-parent')!({}, undefined) + await handlers.get('orchestration:carry-parent')!({}, { from: 'parent-a' }) + expect((bridge as unknown as { replacedParents: Map }).replacedParents.size).toBe(0) +}) diff --git a/src/main/orchestration/OrchestrationBridge.parentCarry.test.ts b/src/main/orchestration/OrchestrationBridge.parentCarry.test.ts index 42c0fd00b..4891a6c05 100644 --- a/src/main/orchestration/OrchestrationBridge.parentCarry.test.ts +++ b/src/main/orchestration/OrchestrationBridge.parentCarry.test.ts @@ -124,3 +124,80 @@ it('invalidates the successor\'s status cache when a child of the replaced paren void bridge.listAgents({ parentSessionId: 'parent-b' }) await vi.waitFor(() => expect(sent.some(request => request.type === 'list-agents' && request.parentSessionId === 'parent-b')).toBe(true)) }) + +// #1369 review a: a create requested under A and answered after A became B. +// B polls meanwhile and caches what it saw; when the create finally answers, +// the child's hint must name B and B's cache must be dropped, or wait_agents +// on B keeps reporting the cached list (done) while the child runs. The +// bridge serves one renderer request at a time, so B's poll is answered +// before the create is dispatched; fake timers freeze the 250 ms cache so +// only an invalidation, never expiry, can send the next poll to the renderer. +it('files a child whose create is answered after the parent was replaced under the successor', async () => { + sent.length = 0 + vi.useFakeTimers() + try { + const bridge = new OrchestrationBridge() + bridge.carryParent('parent-a', 'parent-b') + const listing = bridge.listAgents({ parentSessionId: 'parent-b' }) + await vi.advanceTimersByTimeAsync(0) + const list = sent.splice(sent.findIndex(request => request.type === 'list-agents'), 1)[0]! + bridge.resolve({ requestId: list.requestId, ok: true, type: 'list-agents', agents: [] } as never) + expect(await listing).toEqual([]) + + // A's in-flight create_agent call, answered after the swap. + const created = bridge.createAgent({ parentSessionId: 'parent-a', kind: 'claude' }) + await vi.advanceTimersByTimeAsync(0) + const create = sent.splice(sent.findIndex(request => request.type === 'create-agent'), 1)[0]! + bridge.resolve({ requestId: create.requestId, ok: true, type: 'create-agent', agent: agent('child-1', 'parent-a') } as never) + await created + expect((bridge as unknown as { parentSessionByChildSession: Map }).parentSessionByChildSession.get('child-1')).toBe('parent-b') + + void bridge.listAgents({ parentSessionId: 'parent-b' }) + await vi.advanceTimersByTimeAsync(0) + expect(sent.some(request => request.type === 'list-agents' && request.parentSessionId === 'parent-b')).toBe(true) + } finally { + vi.useRealTimers() + } +}) + +// The same, for a create that timed out and is adopted when the renderer +// finally answers (adoptLateResponse). +it('files a late-adopted child under the parent\'s successor', async () => { + sent.length = 0 + vi.useFakeTimers() + try { + const bridge = new OrchestrationBridge() + const created = bridge.createAgent({ parentSessionId: 'parent-a', kind: 'claude' }).catch(error => error) + await vi.advanceTimersByTimeAsync(0) + const create = sent.find(request => request.type === 'create-agent')! + await vi.advanceTimersByTimeAsync(30_000) + expect(await created).toBeInstanceOf(Error) + bridge.carryParent('parent-a', 'parent-b') + bridge.resolve({ requestId: create.requestId, ok: true, type: 'create-agent', agent: agent('child-1', 'parent-a') } as never) + expect((bridge as unknown as { parentSessionByChildSession: Map }).parentSessionByChildSession.get('child-1')).toBe('parent-b') + } finally { + vi.useRealTimers() + } +}) + +// #1369 reviews a and b (surviving mutants): the alias map is bounded like the +// tombstones it serves, by the same 24 h TTL and 500-entry cap. +it('forgets aliases after the tombstone TTL and keeps at most 500', async () => { + vi.useFakeTimers() + try { + const bridge = new OrchestrationBridge() + const aliases = (bridge as unknown as { replacedParents: Map }).replacedParents + bridge.carryParent('parent-old', 'parent-new') + vi.setSystemTime(Date.now() + 24 * 60 * 60 * 1000 + 1) + bridge.carryParent('parent-x', 'parent-y') + expect(aliases.has('parent-old')).toBe(false) + expect(aliases.has('parent-x')).toBe(true) + + for (let i = 0; i < 510; i++) bridge.carryParent(`pane-${i}`, `pane-${i}-next`) + expect(aliases.size).toBe(500) + expect(aliases.has('pane-509')).toBe(true) + expect(aliases.has('pane-0')).toBe(false) + } finally { + vi.useRealTimers() + } +}) diff --git a/src/main/orchestration/OrchestrationBridge.ts b/src/main/orchestration/OrchestrationBridge.ts index 313c332b0..5859b343f 100644 --- a/src/main/orchestration/OrchestrationBridge.ts +++ b/src/main/orchestration/OrchestrationBridge.ts @@ -314,9 +314,7 @@ export class OrchestrationBridge { createdAt: Date.now(), promptSubmissionCount: 0, }) - this.parentSessionByChildSession.set(response.agent.sessionId, params.parentSessionId) - this.closedAgents.delete(response.agent.sessionId) - this.invalidateStatusCache(params.parentSessionId) + this.noteCreatedChild(response.agent.sessionId, params.parentSessionId) return this.enrichAgent(response.agent) } @@ -605,9 +603,7 @@ export class OrchestrationBridge { createdAt: Date.now(), promptSubmissionCount: 0, }) - this.parentSessionByChildSession.set(response.agent.sessionId, parentSessionId) - this.closedAgents.delete(response.agent.sessionId) - this.invalidateStatusCache(parentSessionId) + this.noteCreatedChild(response.agent.sessionId, parentSessionId) this.journal?.recordIncident({ kind: 'orchestration.late_response_adopted', severity: 'warn', @@ -1006,6 +1002,26 @@ export class OrchestrationBridge { this.pruneCoordinationMetadata() } + /** + * Record a child that a create answered for, under its parent's LIVE id. + * + * WHY resolve here (#1369 review a): a create is requested under parent A + * and can be answered after A was replaced by B (the spawn can take tens of + * seconds; a timed-out create is adopted even later). carryParent only + * rewrites hints that already exist, so storing A made the child's prompt + * boundaries invalidate A's status cache while B kept serving the empty + * list it cached during the spawn, and `wait_agents` on B could report done + * with an active child. Both A's and B's caches are invalidated: A's may + * still hold a list computed before the swap. + */ + private noteCreatedChild(childSessionId: string, requestedParentId: string): void { + const parentSessionId = this.currentParentId(requestedParentId) + this.parentSessionByChildSession.set(childSessionId, parentSessionId) + this.closedAgents.delete(childSessionId) + this.invalidateStatusCache(requestedParentId) + if (parentSessionId !== requestedParentId) this.invalidateStatusCache(parentSessionId) + } + /** The live successor of a possibly replaced parent id (see replacedParents). */ private currentParentId(sessionId: string): string { let current = sessionId diff --git a/src/renderer/src/workspace/hook/actions/replacementOwnership.renderer.test.tsx b/src/renderer/src/workspace/hook/actions/replacementOwnership.renderer.test.tsx index 087b8f38f..11b0fe04f 100644 --- a/src/renderer/src/workspace/hook/actions/replacementOwnership.renderer.test.tsx +++ b/src/renderer/src/workspace/hook/actions/replacementOwnership.renderer.test.tsx @@ -142,11 +142,14 @@ it('Reload Agents carries only to successors that keep Goal Loop tools', async ( // same conversation continuing in a successor takes them along, with or // without Goal Loop tools; a different conversation swapped in does not. it('hands the pane\'s workflow runs to the successor of the same conversation', async () => { - const { mounted, carryWorkflowRuns } = replaceHarness([]) + const { mounted, carryWorkflowRuns, carryOrchestrationParent } = replaceHarness([]) await act(async () => { await mounted.result.current.replaceSession('/recorded/project', { targetSessionId: 'source', kind: 'claude', resumeSessionId: 'native-source' }) }) expect(carryWorkflowRuns).toHaveBeenCalledWith('source', 'successor') + // #1369 review b: the same conversation carries its closed orchestration + // children too (the newConversation case is pinned below). + expect(carryOrchestrationParent).toHaveBeenCalledWith('source', 'successor') }) it('does not hand workflow runs to a different conversation swapped into the pane', async () => { From 35aa4b8ef0626c24ba23c1449bb19f6a3f328283 Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sat, 26 Sep 2026 23:49:26 -0700 Subject: [PATCH 17/48] test(orchestration): pin the carry's cache invalidations and the renderer helper's containment #1369 review c: dropping either status-cache invalidation in carryParent, the renderer's identity-pair skip or its .catch left every test green. Also corrects the delete(to) comment: no wired caller revives an id. Co-Authored-By: Claude Opus 5.5 --- .../OrchestrationBridge.parentCarry.test.ts | 24 +++++++++++++++++ src/main/orchestration/OrchestrationBridge.ts | 10 ++++--- .../actions/successorCarry.renderer.test.ts | 27 +++++++++++++++++++ 3 files changed, 57 insertions(+), 4 deletions(-) create mode 100644 src/renderer/src/workspace/hook/actions/successorCarry.renderer.test.ts diff --git a/src/main/orchestration/OrchestrationBridge.parentCarry.test.ts b/src/main/orchestration/OrchestrationBridge.parentCarry.test.ts index 4891a6c05..738ea291f 100644 --- a/src/main/orchestration/OrchestrationBridge.parentCarry.test.ts +++ b/src/main/orchestration/OrchestrationBridge.parentCarry.test.ts @@ -201,3 +201,27 @@ it('forgets aliases after the tombstone TTL and keeps at most 500', async () => vi.useRealTimers() } }) + +// #1369 review c: the carry drops both parents' status caches. The successor +// may have cached "no closed children" in the spawn -> commit window, before +// the carry landed; the replaced id may still hold the list that owned them. +// Fake timers freeze the 250 ms cache, so only the invalidation can send the +// next poll to the renderer. +it.each(['parent-a', 'parent-b'])('drops %s\'s cached status when the parent is carried', async parent => { + sent.length = 0 + vi.useFakeTimers() + try { + const bridge = new OrchestrationBridge() + const listing = bridge.listAgents({ parentSessionId: parent }) + await vi.advanceTimersByTimeAsync(0) + const list = sent.splice(sent.findIndex(request => request.type === 'list-agents'), 1)[0]! + bridge.resolve({ requestId: list.requestId, ok: true, type: 'list-agents', agents: [] } as never) + await listing + bridge.carryParent('parent-a', 'parent-b') + void bridge.listAgents({ parentSessionId: parent }) + await vi.advanceTimersByTimeAsync(0) + expect(sent.some(request => request.type === 'list-agents' && request.parentSessionId === parent)).toBe(true) + } finally { + vi.useRealTimers() + } +}) diff --git a/src/main/orchestration/OrchestrationBridge.ts b/src/main/orchestration/OrchestrationBridge.ts index 5859b343f..a8bdc02b5 100644 --- a/src/main/orchestration/OrchestrationBridge.ts +++ b/src/main/orchestration/OrchestrationBridge.ts @@ -973,10 +973,12 @@ export class OrchestrationBridge { // Clone-boundary input from the renderer; a malformed carry must be a // no-op, never an alias keyed by `undefined`. if (typeof from !== 'string' || typeof to !== 'string' || !from || !to || from === to) return - // `to` is a live pane again (it may itself have been replaced earlier and - // come back, e.g. Undo Close). An edge OUT of it would now send its new - // children elsewhere, and dropping it is also what keeps the edges acyclic: - // after this, `to` reaches nothing, so `from -> to` cannot close a loop. + // `to` is a live pane. None of today's callers (replace, Reload Agents, + // Undo Close) passes a `to` that already has an edge out: each mints a + // fresh id (#1369 review c). The delete is what keeps the edges acyclic + // by construction anyway: after it, `to` reaches nothing, so `from -> to` + // cannot close a loop, even if a future caller revives a replaced id + // (an edge out of it would then send its new children elsewhere). this.replacedParents.delete(to) this.replacedParents.delete(from) this.replacedParents.set(from, { to, at: Date.now() }) diff --git a/src/renderer/src/workspace/hook/actions/successorCarry.renderer.test.ts b/src/renderer/src/workspace/hook/actions/successorCarry.renderer.test.ts new file mode 100644 index 000000000..16e367043 --- /dev/null +++ b/src/renderer/src/workspace/hook/actions/successorCarry.renderer.test.ts @@ -0,0 +1,27 @@ +import { afterEach, expect, it, vi } from 'vitest' + +import { carryOrchestrationParents } from './successorCarry' + +// #1369 review c: the renderer half of the orchestration carry. It runs right +// after a swap commits, so it must never throw into that commit (a failed IPC +// is logged, not rejected), and an identity pair (a pane that kept its id) +// is not a replacement. +const originalApi = window.api +afterEach(() => { window.api = originalApi; vi.restoreAllMocks() }) + +it('carries each replaced id, skips a pane that kept its id, and contains a failed IPC', async () => { + const warn = vi.spyOn(console, 'warn').mockImplementation(() => {}) + const unhandled = vi.fn() + process.on('unhandledRejection', unhandled) + const carry = vi.fn(async (from: string, _to: string) => { if (from === 'broken') throw new Error('ipc gone') }) + window.api = { ...originalApi, carryOrchestrationParent: carry } as never + try { + expect(() => carryOrchestrationParents(new Map([['old', 'new'], ['same', 'same'], ['broken', 'next']]))).not.toThrow() + await vi.waitFor(() => expect(warn).toHaveBeenCalledWith('[orchestration] carry to the replacement session failed:', expect.any(Error))) + expect(carry.mock.calls).toEqual([['old', 'new'], ['broken', 'next']]) + await new Promise(resolve => setTimeout(resolve, 0)) + expect(unhandled).not.toHaveBeenCalled() + } finally { + process.off('unhandledRejection', unhandled) + } +}) From 725a8b83408c67f00a0bd835c63314c34a58c3c8 Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sat, 26 Sep 2026 23:52:55 -0700 Subject: [PATCH 18/48] =?UTF-8?q?docs(plan):=20Undo=20Close=20=E2=80=94=20?= =?UTF-8?q?a=20restored=20parent's=20live=20children=20follow=20it=20(#137?= =?UTF-8?q?3)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 --- .../2026-09-27-undo-close-live-children.md | 21 +++++++++++++++++++ 1 file changed, 21 insertions(+) create mode 100644 docs/plans/2026-09-27-undo-close-live-children.md diff --git a/docs/plans/2026-09-27-undo-close-live-children.md b/docs/plans/2026-09-27-undo-close-live-children.md new file mode 100644 index 000000000..3e63a9c71 --- /dev/null +++ b/docs/plans/2026-09-27-undo-close-live-children.md @@ -0,0 +1,21 @@ +# Undo Close: a restored parent's live children follow it (#1373) + +## Problem +Undo Close restores a closed pane under a fresh session id. Its LIVE children (orchestration children and linked panes that stayed open) keep `orchestrationParentId` / `orchestrationRootId` / `linkedParentId` pointing at the dead id. Restore only remaps pointers inside the restored entry (`remapMetaLineage` on the carried rows) and on the undo stack (`remapSingleEntryLineage`). Replace and Reload Agents, by contrast, remap every session (`remapSessionsRelationships`). + +The renderer's orchestration visibility gate compares ids for equality (`orchestrationMcp.ts`), so the restored parent cannot list, read or prompt its own surviving children. They also render as top-level rows. + +## Evidence +- Found by review c of #1369 (`temp/review-1369/report-c.md`, Suspicions). +- `undoClose.ts`: the single-pane commit writes only `sessions[newSessionId]`; the project commit remaps only the carried rows. + +## Decisions (defaults) +- **At both restore commits, map every live session's pointers through the restore's old→new ids with `remapMetaLineage`.** + - WHY not `remapSessionsRelationships`: that one drops a pointer whose target is not in the record. A child of ANOTHER still-closed parent would lose its link, and a later undo of that parent could no longer relink it. `remapMetaLineage` keeps unmapped ids by design ("lineage describes one restore"). + - Rows that no pointer touches keep their object identity, because `remapMetaLineage` returns the same meta. +- Main's orchestration tombstones already follow Undo Close through `carryOrchestrationParents` (#1369). This is the live half. It lands independently: no shared code beyond the file. + +## Tests (fail-first) +- Undo Close of a single pane with a live orchestration child and a live linked child: both now point at the restored id. +- Undo Close of a project whose member has a live child in ANOTHER project: that child follows too. +- A child of a different, still-closed parent keeps its pointer (not dropped). From de4c6717bdea22eeecb6a04a1d2962712081b2f0 Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sat, 26 Sep 2026 23:54:30 -0700 Subject: [PATCH 19/48] fix(undo-close): a restored parent's live children follow it to its new id Undo Close remapped pointers only on the rows it restored, so a restored parent's live orchestration and linked children kept the dead id and the parent could not list, read or prompt them. Both restore commits now map every live session through the restore's lineage (remapMetaLineage keeps unmapped ids, so a child of another still-closed pane keeps its link). Co-Authored-By: Claude Opus 5.5 --- .../src/workspace/hook/actions/undoClose.ts | 36 ++++++++- .../undoCloseLiveChildren.renderer.test.tsx | 81 +++++++++++++++++++ 2 files changed, 115 insertions(+), 2 deletions(-) create mode 100644 src/renderer/src/workspace/hook/actions/undoCloseLiveChildren.renderer.test.tsx diff --git a/src/renderer/src/workspace/hook/actions/undoClose.ts b/src/renderer/src/workspace/hook/actions/undoClose.ts index a6987a51c..4392e7013 100644 --- a/src/renderer/src/workspace/hook/actions/undoClose.ts +++ b/src/renderer/src/workspace/hook/actions/undoClose.ts @@ -77,6 +77,35 @@ type PublishLineage = (lineage: UndoLineage) => void * has not landed in this snapshot; production `spawn` writes SessionMeta into * workspace state itself, so in the app it is normally present. */ + +/** + * Point every live session's cross-session pointers (orchestration parent and + * root, linked parent) at the restored ids (#1373). + * + * WHY every session and not only the restored rows: a pane's LIVE children + * (orchestration children, linked panes left open) are not part of its undo + * entry, so remapping only the carried rows left them naming the dead id. The + * renderer's orchestration visibility gate compares ids, so the restored + * parent could not list, read or prompt its own children, and they rendered + * as top-level rows. Replace and Reload Agents already remap every session; + * this is the same rule on the undo path. + * + * WHY remapMetaLineage and not remapSessionsRelationships: the latter DROPS a + * pointer whose target is not in the record. A child of a different pane that + * is still closed (still on the undo stack) would lose its link, and undoing + * that pane later could no longer relink it. remapMetaLineage keeps any id the + * map does not name, and returns the same object when nothing changed, so + * untouched rows keep their identity. + */ +function remapLiveLineage( + sessions: Record, + idMap: ReadonlyMap, +): Record { + const out: Record = {} + for (const [sessionId, meta] of Object.entries(sessions)) out[sessionId] = remapMetaLineage(meta, idMap) + return out +} + export function carryDurableMeta(spawned: SessionMeta | undefined, closed: SessionMeta): SessionMeta { // Extension panes skip spawn entirely: all of their metadata is durable UI // identity, including the view ID. The process-specific allowlist below is @@ -261,7 +290,8 @@ export function useUndoCloseAction( // files a session has always done. activeTabId: projectId, sessions: { - ...prev.sessions, + // Live children of the closed pane follow it to its new id (#1373). + ...remapLiveLineage(prev.sessions, new Map([[entry.sessionId, newSessionId]])), // carryDurableMeta restores membership too — `projectId` and, // critically, `joinedAt`, so the row returns to its old position // in the index rather than jumping to the bottom. @@ -388,7 +418,9 @@ export function useUndoCloseAction( const insertIdx = Math.min(entry.tabIndex, prev.tabs.length) const tabs = [...prev.tabs] tabs.splice(insertIdx, 0, restoredTab) - const sessions = { ...prev.sessions } + // Live children of the restored members, wherever they live, follow + // them to their new ids (#1373). + const sessions = remapLiveLineage(prev.sessions, idMap) for (const [newId, closed] of carried) { sessions[newId] = { // Relationship pointers follow the restore: a linked child diff --git a/src/renderer/src/workspace/hook/actions/undoCloseLiveChildren.renderer.test.tsx b/src/renderer/src/workspace/hook/actions/undoCloseLiveChildren.renderer.test.tsx new file mode 100644 index 000000000..8b8fa2bd3 --- /dev/null +++ b/src/renderer/src/workspace/hook/actions/undoCloseLiveChildren.renderer.test.tsx @@ -0,0 +1,81 @@ +import { act, render } from '@testing-library/react' +import { afterEach, expect, it, vi } from 'vitest' + +import { useUndoCloseAction } from '@renderer/workspace/hook/actions/undoClose' +import { makeRefs, sessionActionsWithSpawn, stateWriter } from '@renderer/workspace/hook/actions/testing/paneActionsHarness' +import { freshStage } from '@renderer/workspace/dispatch/gridShape' +import type { SessionMeta, WorkspaceState } from '@renderer/workspace/types' + +// #1373: Undo Close restores a pane under a fresh session id, and its LIVE +// children (orchestration children, linked panes that stayed open) kept +// pointing at the dead id. The orchestration visibility gate compares ids, so +// the restored parent could not list, read or prompt its own children. Replace +// and Reload Agents already remap every session; Undo Close only remapped the +// rows it restored. +const originalApi = window.api +afterEach(() => { window.api = originalApi }) + +const meta = (title: string, extra: Partial = {}): SessionMeta => + ({ cwd: '/projects/agent-code', kind: 'claude', title, projectId: 'tab-live', joinedAt: 1, ...extra }) as SessionMeta + +function mount(entry: unknown, sessions: Record, spawn: ReturnType) { + const state: WorkspaceState = { + tabs: [{ id: 'tab-live', title: 'live' }, { id: 'tab-parent', title: 'agent-code' }], activeTabId: 'tab-live', stage: freshStage(), + sessions, pinnedSessionIds: [], + } + const refs = makeRefs(state) + const writer = stateWriter(state, refs) + refs.undoStackRef.current.push(entry as never) + window.api = { + ...originalApi, carryWorkflowRuns: vi.fn(async () => undefined), carryGoalLoop: vi.fn(async () => null), + controlGoalLoop: vi.fn(async () => null), killOwnedSession: vi.fn(async () => true), + } as never + let actions!: ReturnType + function Harness(): React.JSX.Element { + actions = useUndoCloseAction(state, writer.setState, refs, sessionActionsWithSpawn(spawn), vi.fn()) + return
+ } + const mounted = render() + return { writer, undo: () => actions.undoClose(), unmount: () => mounted.unmount() } +} + +it('points a restored pane\'s live orchestration and linked children at its new id', async () => { + const unrelated = meta('Unrelated') + const harness = mount( + { type: 'session', closedAt: Date.now(), sessionId: 'closed-parent', sessionMeta: meta('Parent', { projectId: 'tab-parent' }) }, + { + worker: meta('Worker', { orchestrationParentId: 'closed-parent', orchestrationRootId: 'closed-parent' }), + linked: meta('Linked', { linkedParentId: 'closed-parent' }), + // A child of ANOTHER closed parent that is still on the stack: a later + // undo of that parent must still find it, so its pointer is kept. + orphan: meta('Orphan', { orchestrationParentId: 'other-closed-parent', orchestrationRootId: 'other-closed-parent' }), + unrelated, + }, + vi.fn().mockResolvedValue('restored-parent'), + ) + await act(async () => { await harness.undo() }) + const sessions = harness.writer.getState().sessions + expect(sessions.worker).toMatchObject({ orchestrationParentId: 'restored-parent', orchestrationRootId: 'restored-parent' }) + expect(sessions.linked).toMatchObject({ linkedParentId: 'restored-parent' }) + expect(sessions.orphan).toMatchObject({ orchestrationParentId: 'other-closed-parent', orchestrationRootId: 'other-closed-parent' }) + // Rows no pointer touches keep their identity (no needless re-render). + expect(sessions.unrelated).toBe(unrelated) + harness.unmount() +}) + +it('points the live children of a restored project\'s members at their new ids', async () => { + const harness = mount( + { + type: 'tab', closedAt: Date.now(), tab: { id: 'closed-tab', title: 'agent-code' }, tabIndex: 1, + sessions: [{ sessionId: 'lead', meta: meta('Lead', { projectId: 'closed-tab' }) }, { sessionId: 'second', meta: meta('Second', { projectId: 'closed-tab' }) }], + }, + { + // Live in ANOTHER project, so the project close did not take it. + worker: meta('Worker', { orchestrationParentId: 'second', orchestrationRootId: 'lead' }), + }, + vi.fn().mockResolvedValueOnce('new-lead').mockResolvedValueOnce('new-second'), + ) + await act(async () => { await harness.undo() }) + expect(harness.writer.getState().sessions.worker).toMatchObject({ orchestrationParentId: 'new-second', orchestrationRootId: 'new-lead' }) + harness.unmount() +}) From a3246961710f3a7da6f4df39e317447ce45f37a4 Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sun, 27 Sep 2026 00:05:56 -0700 Subject: [PATCH 20/48] fix(orchestration): the renderer files a child created across a parent swap under the successor #1369 verification a and b: createOrchestrationAgent captured the parent id, awaited the spawn, and filed the child under it; a replacement committed meanwhile had only remapped existing children, so the successor could not list, read or close the new child. The committed-swap helper now records a bounded, acyclic successor map per window, and filing resolves parent and root through it. Red without the resolution; every piece pinned. Co-Authored-By: Claude Opus 5.5 --- .../src/workspace/hook/actions/pane.ts | 14 ++++-- .../actions/successorCarry.renderer.test.ts | 19 ++++++++ .../workspace/hook/actions/successorCarry.ts | 47 +++++++++++++++++++ .../orchestrationRuntime.renderer.test.tsx | 35 ++++++++++++++ 4 files changed, 111 insertions(+), 4 deletions(-) diff --git a/src/renderer/src/workspace/hook/actions/pane.ts b/src/renderer/src/workspace/hook/actions/pane.ts index 5b6be6ec7..6a81d27d0 100644 --- a/src/renderer/src/workspace/hook/actions/pane.ts +++ b/src/renderer/src/workspace/hook/actions/pane.ts @@ -1,6 +1,7 @@ import { DEFAULT_PROVIDER, effectiveProviderRuntime } from '@shared/types/providerKind' import { SESSION_START_FAILED_MESSAGE } from '@shared/types/session' import { enabledAgentProviderChoices } from '@renderer/workspace/providerChoices' +import { currentOrchestrationParent } from '@renderer/workspace/hook/actions/successorCarry' import { expandSessionCloseTargets, expandTabCloseTargets, @@ -1383,13 +1384,18 @@ export function usePaneActions( builtInMcpDomains: params.builtInMcpDomains, }) + // The parent may have been replaced while the spawn was awaited (#1369 + // verification a/b): file the child under the live ids, as the swap's + // remap did for the children that already existed. + const parentId = currentOrchestrationParent(params.parentId) + const rootId = currentOrchestrationParent(rootParentId) const agent: OrchestrationAgentRecord = { sessionId, kind: params.kind, cwd, ...(params.title ? { title: params.title } : {}), - orchestrationParentId: params.parentId, - orchestrationRootId: rootParentId, + orchestrationParentId: parentId, + orchestrationRootId: rootId, ...(params.runId ? { orchestrationRunId: params.runId } : {}), ...(params.role ? { orchestrationRole: params.role } : {}), } @@ -1408,8 +1414,8 @@ export function usePaneActions( cwd, kind: params.kind, ...(params.title ? { title: params.title } : {}), - orchestrationParentId: params.parentId, - orchestrationRootId: rootParentId, + orchestrationParentId: parentId, + orchestrationRootId: rootId, ...(params.runId ? { orchestrationRunId: params.runId } : {}), ...(params.role ? { orchestrationRole: params.role } : {}), // Filed in the root parent's project; see createLinkedAgent. diff --git a/src/renderer/src/workspace/hook/actions/successorCarry.renderer.test.ts b/src/renderer/src/workspace/hook/actions/successorCarry.renderer.test.ts index 16e367043..42fcd85f3 100644 --- a/src/renderer/src/workspace/hook/actions/successorCarry.renderer.test.ts +++ b/src/renderer/src/workspace/hook/actions/successorCarry.renderer.test.ts @@ -25,3 +25,22 @@ it('carries each replaced id, skips a pane that kept its id, and contains a fail process.off('unhandledRejection', unhandled) } }) + +// #1369 verification a/b: the renderer's copy of the lineage. Own ids per +// case: the map is per-window module state. +it('resolves a parent through its successors, stops at a revived pane, and keeps at most 500', async () => { + const { currentOrchestrationParent } = await import('./successorCarry') + window.api = { ...originalApi, carryOrchestrationParent: vi.fn(async () => undefined) } as never + carryOrchestrationParents(new Map([['lin-a', 'lin-b']])) + carryOrchestrationParents(new Map([['lin-b', 'lin-c']])) + expect(currentOrchestrationParent('lin-a')).toBe('lin-c') + // lin-a comes back as a live pane: its old edge must go (no loop, and a + // child created under it now stays with it). + carryOrchestrationParents(new Map([['lin-c', 'lin-a']])) + expect(currentOrchestrationParent('lin-a')).toBe('lin-a') + expect(currentOrchestrationParent('lin-b')).toBe('lin-a') + + for (let i = 0; i < 510; i++) carryOrchestrationParents(new Map([[`cap-${i}`, `cap-${i}-next`]])) + expect(currentOrchestrationParent('cap-509')).toBe('cap-509-next') + expect(currentOrchestrationParent('cap-0')).toBe('cap-0') +}) diff --git a/src/renderer/src/workspace/hook/actions/successorCarry.ts b/src/renderer/src/workspace/hook/actions/successorCarry.ts index cad00b1f0..0f9f7e678 100644 --- a/src/renderer/src/workspace/hook/actions/successorCarry.ts +++ b/src/renderer/src/workspace/hook/actions/successorCarry.ts @@ -51,6 +51,9 @@ export function carryWorkflowRuns(idMap: ReadonlyMap): void { * pre-fix behaviour) and never blocks the swap. */ export function carryOrchestrationParents(idMap: ReadonlyMap): void { + for (const [oldId, newId] of idMap) { + if (oldId !== newId) recordOrchestrationSuccessor(oldId, newId) + } const carry = window.api?.carryOrchestrationParent if (!carry) return for (const [oldId, newId] of idMap) { @@ -117,3 +120,47 @@ export function handOverGoalLoops( carryGoalLoops(new Map(pairs.filter(([, newId]) => capable(newId)))) stopGoalLoops(pairs.filter(([, newId]) => !capable(newId)).map(([oldId]) => oldId)) } + +/** + * Replaced pane id -> its successor, as this window committed them (#1369 + * verification a and b). + * + * WHY the renderer needs its own copy of main's alias chain: an orchestration + * create captures its parent id, then awaits the child's spawn (tens of + * seconds under load). A replacement committed meanwhile remaps only the + * children already in the store, so the new child was filed under the + * retired id, and the successor could not list, read or close it (the + * visibility gate compares ids). The filing step resolves through this map. + * + * WHY module state and not WorkspaceRefs: it is written by the same helper + * every committed swap already calls (replace, Reload Agents, Undo Close), and + * that helper has no refs. It is per window, which matches ownership: a + * pane's swaps and its creates run in the window that owns it. A renderer + * reload clears it, which is harmless because local ids are stable across a + * reload. Bounded like main's alias map; acyclic for the same reason (a carry + * to `to` drops `to`'s own edge). + */ +const orchestrationSuccessors = new Map() +const MAX_ORCHESTRATION_SUCCESSORS = 500 + +function recordOrchestrationSuccessor(from: string, to: string): void { + orchestrationSuccessors.delete(to) + orchestrationSuccessors.delete(from) + orchestrationSuccessors.set(from, to) + while (orchestrationSuccessors.size > MAX_ORCHESTRATION_SUCCESSORS) { + const oldest = orchestrationSuccessors.keys().next() + if (oldest.done) break + orchestrationSuccessors.delete(oldest.value) + } +} + +/** The live successor of a possibly replaced orchestration parent id. */ +export function currentOrchestrationParent(sessionId: string): string { + let current = sessionId + for (let hops = 0; hops <= orchestrationSuccessors.size; hops++) { + const next = orchestrationSuccessors.get(current) + if (!next) return current + current = next + } + return current +} diff --git a/src/renderer/src/workspace/hook/orchestrationRuntime.renderer.test.tsx b/src/renderer/src/workspace/hook/orchestrationRuntime.renderer.test.tsx index 8706d1140..12c8e2b5f 100644 --- a/src/renderer/src/workspace/hook/orchestrationRuntime.renderer.test.tsx +++ b/src/renderer/src/workspace/hook/orchestrationRuntime.renderer.test.tsx @@ -109,6 +109,41 @@ describe('renderer orchestration runtime creation', () => { expect(spawnSession).toHaveBeenCalledExactlyOnceWith(expect.objectContaining({ kind: 'pi', providerRuntime: 'terminal', cwd: '/repo/child' })) }) + // #1369 verification a and b: a create captures its parent, then awaits the + // child's spawn; a replacement committed meanwhile remaps only children + // already in the store, so the child was filed under the retired id and the + // successor could not list, read or close it. Own ids: the successor map is + // per-window module state and must not leak into the other cases. + it('files a child whose parent was replaced during its spawn under the successor', async () => { + const { carryOrchestrationParents } = await import('./actions/successorCarry') + useAppStore.setState(state => ({ workspaceState: { ...state.workspaceState, sessions: { + ...state.workspaceState.sessions, + 'swap-root': { kind: 'claude', cwd: '/repo', projectId: 'project', joinedAt: 2 }, + 'swap-parent': { kind: 'claude', cwd: '/repo', orchestrationParentId: 'swap-root', orchestrationRootId: 'swap-root', projectId: 'project', joinedAt: 3 }, + } } })) + let release!: () => void + spawnSession.mockImplementationOnce(() => new Promise(resolve => { release = () => resolve({ sessionId: 'child' }) })) + renderHook(() => useWorkspace()) + const creating = dispatch({ requestId: 'swap', type: 'create-agent', parentSessionId: 'swap-parent', kind: 'claude', cwd: '/repo/child' }) + await act(async () => { await vi.advanceTimersByTimeAsync(0) }) + // Replacements commit while the spawn is pending (the parent's pane AND + // the root's, e.g. Reload Agents), and each committed swap records its + // lineage. + act(() => { + useAppStore.setState(state => { + const { 'swap-parent': retired, 'swap-root': retiredRoot, ...rest } = state.workspaceState.sessions + return { workspaceState: { ...state.workspaceState, sessions: { ...rest, 'swap-successor': retired!, 'swap-root-next': retiredRoot! } } } + }) + carryOrchestrationParents(new Map([['swap-parent', 'swap-successor'], ['swap-root', 'swap-root-next']])) + }) + release() + await creating + expect(useAppStore.getState().workspaceState.sessions.child).toMatchObject({ orchestrationParentId: 'swap-successor', orchestrationRootId: 'swap-root-next' }) + expect(resolved).toHaveBeenCalledWith(expect.objectContaining({ requestId: 'swap', ok: true, agent: expect.objectContaining({ orchestrationParentId: 'swap-successor' }) })) + await dispatch({ requestId: 'swap-list', type: 'list-agents', parentSessionId: 'swap-successor' }) + expect(resolved).toHaveBeenLastCalledWith(expect.objectContaining({ requestId: 'swap-list', ok: true, agents: [expect.objectContaining({ sessionId: 'child' })] })) + }) + it('refuses unsupported Claude terminal before spawn even without the main bridge', async () => { renderHook(() => useWorkspace()) await dispatch({ requestId: 'unsupported', type: 'create-agent', parentSessionId: 'parent', kind: 'claude', providerRuntime: 'terminal' }) From f6c66f60ba2ad8234f2335ee0372384320ddee18 Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sun, 27 Sep 2026 00:09:05 -0700 Subject: [PATCH 21/48] fix(orchestration): a bootstrap mark from a replaced parent reaches its successor q85: the mark is the last step of a delivery that can land long after the create began (a prompt waiting for the composer, or #1370's late adoption). It was addressed with the parent id the tool call captured; after a replacement that id has no window and owns no children, so the mark was lost. Resolved through the alias chain; the combined late-create + replacement path was reproduced against #1375 and passes with this. Co-Authored-By: Claude Opus 5.5 --- .../OrchestrationBridge.parentCarry.test.ts | 16 ++++++++++++++++ src/main/orchestration/OrchestrationBridge.ts | 13 +++++++++++-- 2 files changed, 27 insertions(+), 2 deletions(-) diff --git a/src/main/orchestration/OrchestrationBridge.parentCarry.test.ts b/src/main/orchestration/OrchestrationBridge.parentCarry.test.ts index 738ea291f..21b1d65e0 100644 --- a/src/main/orchestration/OrchestrationBridge.parentCarry.test.ts +++ b/src/main/orchestration/OrchestrationBridge.parentCarry.test.ts @@ -225,3 +225,19 @@ it.each(['parent-a', 'parent-b'])('drops %s\'s cached status when the parent is vi.useRealTimers() } }) + +// q85: a bootstrap mark is the last step of a delivery that can land long +// after the create began (a prompt that waited for the child's composer, or a +// create adopted late, #1370). It is addressed with the parent id the tool +// call captured; after a replacement that id has no window and owns no +// children in the renderer, so the mark must go to the successor. +it('addresses a bootstrap mark from a replaced parent to its successor', async () => { + sent.length = 0 + const bridge = new OrchestrationBridge() + bridge.carryParent('parent-a', 'parent-b') + const marking = bridge.markBootstrapPromptDelivered({ parentSessionId: 'parent-a', sessionId: 'child-1' }) + const mark = await next('mark-bootstrap-prompt-delivered') + expect(mark.parentSessionId).toBe('parent-b') + bridge.resolve({ requestId: mark.requestId, ok: true, type: 'mark-bootstrap-prompt-delivered', agent: agent('child-1', 'parent-b') } as never) + expect(await marking).toMatchObject({ sessionId: 'child-1', orchestrationParentId: 'parent-b' }) +}) diff --git a/src/main/orchestration/OrchestrationBridge.ts b/src/main/orchestration/OrchestrationBridge.ts index a8bdc02b5..c021a1a61 100644 --- a/src/main/orchestration/OrchestrationBridge.ts +++ b/src/main/orchestration/OrchestrationBridge.ts @@ -530,17 +530,26 @@ export class OrchestrationBridge { parentSessionId: string sessionId: string }): Promise { + // WHY resolve (q85): the mark is the LAST step of a bootstrap delivery, + // which can land long after the create began: when the prompt waited for + // the child's composer, or when a create outlived the 30 s deadline and was + // adopted late (#1370). The caller addresses it with the parent id its + // tool call captured. If that parent was replaced meanwhile, the retired + // id has no window (the request could not be routed) and owns no children + // in the renderer (ownership checks compare ids), so the mark was lost and + // the child looked never-bootstrapped. The live successor owns both. + const parentSessionId = this.currentParentId(params.parentSessionId) const response = await this.request({ requestId: randomUUID(), type: 'mark-bootstrap-prompt-delivered', - parentSessionId: params.parentSessionId, + parentSessionId, sessionId: params.sessionId, }) if (!response.ok) throw new Error(response.message) if (response.type !== 'mark-bootstrap-prompt-delivered') { throw new Error(`Unexpected orchestration response: ${response.type}`) } - this.invalidateStatusCache(params.parentSessionId) + this.invalidateStatusCache(parentSessionId) return this.enrichAgent(response.agent) } From 81314044aadd90b3154d25cf6cb47bf31f8f4e25 Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sun, 27 Sep 2026 00:19:18 -0700 Subject: [PATCH 22/48] fix(orchestration): a queued bootstrap mark resolves its parent at dispatch #1369 verification a, round 3: the bridge serves one renderer request at a time, so a mark resolved when queued could still dispatch to a parent replaced while it waited. Only the mark is rewritten; other requests act for the parent that asked. Co-Authored-By: Claude Opus 5.5 --- .../OrchestrationBridge.parentCarry.test.ts | 19 +++++++++++++++++++ src/main/orchestration/OrchestrationBridge.ts | 16 ++++++++++++++-- 2 files changed, 33 insertions(+), 2 deletions(-) diff --git a/src/main/orchestration/OrchestrationBridge.parentCarry.test.ts b/src/main/orchestration/OrchestrationBridge.parentCarry.test.ts index 21b1d65e0..55c35970b 100644 --- a/src/main/orchestration/OrchestrationBridge.parentCarry.test.ts +++ b/src/main/orchestration/OrchestrationBridge.parentCarry.test.ts @@ -241,3 +241,22 @@ it('addresses a bootstrap mark from a replaced parent to its successor', async ( bridge.resolve({ requestId: mark.requestId, ok: true, type: 'mark-bootstrap-prompt-delivered', agent: agent('child-1', 'parent-b') } as never) expect(await marking).toMatchObject({ sessionId: 'child-1', orchestrationParentId: 'parent-b' }) }) + +// #1369 verification a, round 3: the bridge serves ONE renderer request at a +// time, so a mark can wait in the queue while another request runs. A swap +// landing in that wait must still redirect it: resolution happens when the +// request is dispatched, not when it was queued. +it('redirects a bootstrap mark queued before the parent was replaced', async () => { + sent.length = 0 + const bridge = new OrchestrationBridge() + const listing = bridge.listAgents({ parentSessionId: 'parent-a' }) + const list = await next('list-agents') + const marking = bridge.markBootstrapPromptDelivered({ parentSessionId: 'parent-a', sessionId: 'child-1' }) + bridge.carryParent('parent-a', 'parent-b') + bridge.resolve({ requestId: list.requestId, ok: true, type: 'list-agents', agents: [] } as never) + await listing + const mark = await next('mark-bootstrap-prompt-delivered') + expect(mark.parentSessionId).toBe('parent-b') + bridge.resolve({ requestId: mark.requestId, ok: true, type: 'mark-bootstrap-prompt-delivered', agent: agent('child-1', 'parent-b') } as never) + await marking +}) diff --git a/src/main/orchestration/OrchestrationBridge.ts b/src/main/orchestration/OrchestrationBridge.ts index c021a1a61..6a8d7e65a 100644 --- a/src/main/orchestration/OrchestrationBridge.ts +++ b/src/main/orchestration/OrchestrationBridge.ts @@ -549,7 +549,8 @@ export class OrchestrationBridge { if (response.type !== 'mark-bootstrap-prompt-delivered') { throw new Error(`Unexpected orchestration response: ${response.type}`) } - this.invalidateStatusCache(parentSessionId) + // Re-resolved: the parent may have been replaced while the mark waited. + this.invalidateStatusCache(this.currentParentId(parentSessionId)) return this.enrichAgent(response.agent) } @@ -857,8 +858,19 @@ export class OrchestrationBridge { } private async dispatchRendererRequest( - request: OrchestrationRendererRequest, + queued: OrchestrationRendererRequest, ): Promise { + // WHY a bootstrap mark resolves its parent HERE, at dispatch (#1369 + // verification a, round 3): the bridge serves one renderer request at a + // time, so a mark can sit in the queue while another request runs. A + // replacement landing in that wait makes the parent it was queued with a + // retired id: no window to route to, no ownership of the child. Resolving + // when it was queued (markBootstrapPromptDelivered) is too early. Only the + // mark is rewritten: every other request is a question or an action BY + // that parent, and answering it for a different session would be wrong. + const request = queued.type === 'mark-bootstrap-prompt-delivered' + ? { ...queued, parentSessionId: this.currentParentId(queued.parentSessionId) } + : queued return await new Promise((resolve, reject) => { const TIMEOUT_MS = 30_000 const timer = setTimeout(() => { From 1b00555564e9b1da035a8461c1ca19619dd495f1 Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sun, 27 Sep 2026 01:28:39 -0700 Subject: [PATCH 23/48] docs(plan): reloading a live child keeps its pointer to a closed, restorable parent (#1379) Co-Authored-By: Claude Opus 5.5 --- ...6-09-27-replace-keeps-closed-parent-pointer.md | 15 +++++++++++++++ 1 file changed, 15 insertions(+) create mode 100644 docs/plans/2026-09-27-replace-keeps-closed-parent-pointer.md diff --git a/docs/plans/2026-09-27-replace-keeps-closed-parent-pointer.md b/docs/plans/2026-09-27-replace-keeps-closed-parent-pointer.md new file mode 100644 index 000000000..1b74edeec --- /dev/null +++ b/docs/plans/2026-09-27-replace-keeps-closed-parent-pointer.md @@ -0,0 +1,15 @@ +# Reloading a live child keeps its pointer to a closed, restorable parent (#1379) + +## Problem +Replace (reload, provider switch, resume, rewind) and Reload Agents commit through `remapSessionsRelationships(sessions, idMap)`. Its `knownSessionIds` defaults to the live record's keys, and a pointer whose target is in neither the idMap nor that set is DROPPED. A parent that is closed but still on the undo stack is in neither. So reloading its live child erases `linkedParentId` / `orchestrationParentId` / `orchestrationRootId`, and undoing the parent later (with #1374's relink) finds nothing to relink. The two stay live and permanently unlinked. Found by review c of #1374. + +## Decision +- **Both commits pass `knownSessionIds` = live ids ∪ the undo stack's restorable session ids.** `UndoCloseStack.restorableSessionIds()` collects them from session, tab and group entries, after the same expiry prune that `pop()` uses. +- **Why not "never drop":** the drop is right for a target that can never come back. A deleted pane or an expired undo entry dangling forever misleads the index and the orchestration gate. Only what undo can still restore is worth keeping. +- **An entry that later expires** leaves a pointer to an id nothing can restore. That is the same observable state as the drop (a top-level row, invisible to any parent), and it is cleaned the next time that child is remapped. + +## Tests (fail-first) +- **Replace:** a live child reloaded while its parent sits on the undo stack keeps its pointers. +- **Reload Agents:** the same. +- **The control:** a pointer to a parent that is neither live nor restorable is still dropped. +- **`UndoCloseStack.restorableSessionIds`:** session, tab and group entries, and expired entries excluded. From 39d1595609e41e4e5bdd5a9a35dc70adf4ae58a5 Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sun, 27 Sep 2026 01:33:35 -0700 Subject: [PATCH 24/48] fix(sessions): a reloaded child keeps its pointer to a closed, restorable parent #1379 (from #1374 review c): replace and Reload Agents dropped any relationship pointer whose target was neither live nor in their idMap; a parent closed but still on the undo stack is neither, so reloading its live child erased the link and undoing the parent later had nothing to relink. Both commits now treat the undo stack's restorable session ids as known (UndoCloseStack.restorableSessionIds, pruned like pop). A pointer to a parent nothing can restore is still dropped. Red before; each call site pinned. Co-Authored-By: Claude Opus 5.5 --- src/renderer/src/lib/undoClose.test.ts | 17 +++++ src/renderer/src/lib/undoClose.ts | 24 ++++++ .../closedParentPointer.renderer.test.tsx | 73 +++++++++++++++++++ .../src/workspace/hook/actions/session.ts | 29 +++++++- 4 files changed, 140 insertions(+), 3 deletions(-) create mode 100644 src/renderer/src/workspace/hook/actions/closedParentPointer.renderer.test.tsx diff --git a/src/renderer/src/lib/undoClose.test.ts b/src/renderer/src/lib/undoClose.test.ts index c0f51c424..5551cdbee 100644 --- a/src/renderer/src/lib/undoClose.test.ts +++ b/src/renderer/src/lib/undoClose.test.ts @@ -123,3 +123,20 @@ describe('UndoCloseStack', () => { expect(stack.pop()).toBeNull() }) }) + +// #1379: what a waiting entry could bring back, for replace/Reload Agents' +// relationship remap. +describe('UndoCloseStack.restorableSessionIds', () => { + it('collects session, project and group members, and skips expired entries', () => { + let now = 1_000_000 + const stack = new UndoCloseStack(() => now) + const meta = { cwd: '/p', kind: 'claude', projectId: 't', joinedAt: 0 } as never + stack.push({ type: 'session', closedAt: now - UNDO_CLOSE_RETENTION_MS - 1, sessionId: 'old' as never, sessionMeta: meta } as never) + stack.push({ type: 'session', closedAt: now, sessionId: 'single' as never, sessionMeta: meta } as never) + stack.push({ type: 'tab', closedAt: now, tab: { id: 't', title: 't' }, tabIndex: 0, sessions: [{ sessionId: 'tab-a' as never, meta }] } as never) + stack.push({ type: 'group', closedAt: now, entries: [{ type: 'session', closedAt: now, sessionId: 'grouped' as never, sessionMeta: meta }] } as never) + expect([...stack.restorableSessionIds()].sort()).toEqual(['grouped', 'single', 'tab-a']) + now += UNDO_CLOSE_RETENTION_MS + 1 + expect(stack.restorableSessionIds().size).toBe(0) + }) +}) diff --git a/src/renderer/src/lib/undoClose.ts b/src/renderer/src/lib/undoClose.ts index b5efae26f..2aa48e2b4 100644 --- a/src/renderer/src/lib/undoClose.ts +++ b/src/renderer/src/lib/undoClose.ts @@ -265,6 +265,30 @@ export class UndoCloseStack { return this.entries.length } + /** + * Every session id a waiting entry could bring back (#1379). Pruned first, + * exactly as `pop()` would be, so an expired entry never counts. + * + * WHY the stack answers this: replace and Reload Agents drop any + * relationship pointer whose target is neither live nor in their idMap. A + * parent that is closed but still restorable is neither, so reloading its + * live child erased the child's link, and undoing the parent later had + * nothing to relink. The swap passes these ids as "still known". + */ + restorableSessionIds(): Set { + this.prune() + const ids = new Set() + const add = (entry: SingleClosedEntry): void => { + if (entry.type === 'session') ids.add(entry.sessionId) + else for (const member of entry.sessions) ids.add(member.sessionId) + } + for (const entry of this.entries) { + if (entry.type === 'group') entry.entries.forEach(add) + else add(entry) + } + return ids + } + /** Rewrite every waiting entry's anchors after a successful restore. See * `UndoLineage` for why this runs on restore and never on merge. */ remapLineage(lineage: UndoLineage): void { diff --git a/src/renderer/src/workspace/hook/actions/closedParentPointer.renderer.test.tsx b/src/renderer/src/workspace/hook/actions/closedParentPointer.renderer.test.tsx new file mode 100644 index 000000000..fc51452bf --- /dev/null +++ b/src/renderer/src/workspace/hook/actions/closedParentPointer.renderer.test.tsx @@ -0,0 +1,73 @@ +import { act, cleanup, renderHook } from '@testing-library/react' +import { afterEach, expect, it, vi } from 'vitest' +import { useAppStore } from '@renderer/app-state/store' +import { emptyRuntime } from '@renderer/session-runtime/state' +import type { SessionMeta } from '@renderer/workspace/types' +import { useSessionActions } from './session' +import { makeRefs } from './testing/paneActionsHarness' +vi.mock('@renderer/workspace/hook/actions/initialHistory', () => ({ loadInitialHistoryForSession: vi.fn(async () => undefined) })) + +// #1379: replace and Reload Agents drop any relationship pointer whose target +// is neither live nor in their idMap. A parent that is CLOSED but still on the +// undo stack is neither, so reloading its live child erased the child's link +// and undoing the parent later had nothing to relink: two live panes, +// permanently unlinked. A pointer to what undo can still restore must survive; +// a pointer to what nothing can bring back is still dropped. +const original = useAppStore.getState() +const originalApi = window.api +afterEach(() => { cleanup(); useAppStore.setState(original, true); window.api = originalApi }) + +const child = (parent: string): SessionMeta => ({ + kind: 'claude', cwd: '/recorded/project', providerSessionId: 'native-child', projectId: 'project', joinedAt: 0, + linkedParentId: parent, orchestrationParentId: parent, orchestrationRootId: parent, +} as SessionMeta) + +function harness(parent: string, restorable: boolean) { + useAppStore.setState({ workspaceState: { ...original.workspaceState, activeTabId: 'project', + tabs: [{ id: 'project', title: 'Project' }], + sessions: { child: child(parent) }, + }, workspaceRuntimes: { child: { ...emptyRuntime(), processStatus: 'started' } } }) + const state = useAppStore.getState().workspaceState + const refs = makeRefs(state) + refs.latestRuntimesRef.current = useAppStore.getState().workspaceRuntimes + if (restorable) { + refs.undoStackRef.current.push({ + type: 'session', closedAt: Date.now(), sessionId: parent, + sessionMeta: { kind: 'claude', cwd: '/recorded/project', projectId: 'project', joinedAt: 1 }, + } as never) + } + window.api = { + ...originalApi, spawnSession: vi.fn(async () => ({ sessionId: 'child-next' })), killOwnedSession: vi.fn(async () => true), + carryGoalLoop: vi.fn(async () => null), carryWorkflowRuns: vi.fn(async () => undefined), controlGoalLoop: vi.fn(async () => null), + } + return renderHook(() => useSessionActions(state, useAppStore.getState().setWorkspaceState, useAppStore.getState().setWorkspaceRuntimes, refs)) +} + +const pointers = (id: string) => { + const meta = useAppStore.getState().workspaceState.sessions[id] + return { linkedParentId: meta?.linkedParentId, orchestrationParentId: meta?.orchestrationParentId, orchestrationRootId: meta?.orchestrationRootId } +} + +it('keeps a replaced child\'s pointers to a parent that is closed but restorable', async () => { + const mounted = harness('closed-parent', true) + await act(async () => { + expect(await mounted.result.current.replaceSession('/recorded/project', { targetSessionId: 'child', kind: 'claude', resumeSessionId: 'native-child' })).toBe('child-next') + }) + expect(pointers('child-next')).toEqual({ linkedParentId: 'closed-parent', orchestrationParentId: 'closed-parent', orchestrationRootId: 'closed-parent' }) +}) + +it('keeps the pointers through Reload Agents too', async () => { + const mounted = harness('closed-parent', true) + await act(async () => { await mounted.result.current.reloadAgentSessions(true) }) + expect(pointers('child-next')).toEqual({ linkedParentId: 'closed-parent', orchestrationParentId: 'closed-parent', orchestrationRootId: 'closed-parent' }) +}) + +// The control: nothing can bring this parent back, so the dangling link is +// still dropped (the pre-existing, correct behaviour). +it('still drops pointers to a parent that is neither live nor restorable', async () => { + const mounted = harness('gone-parent', false) + await act(async () => { + await mounted.result.current.replaceSession('/recorded/project', { targetSessionId: 'child', kind: 'claude', resumeSessionId: 'native-child' }) + }) + expect(pointers('child-next')).toEqual({ linkedParentId: undefined, orchestrationParentId: undefined, orchestrationRootId: undefined }) +}) diff --git a/src/renderer/src/workspace/hook/actions/session.ts b/src/renderer/src/workspace/hook/actions/session.ts index 3d68b7db0..baafff706 100644 --- a/src/renderer/src/workspace/hook/actions/session.ts +++ b/src/renderer/src/workspace/hook/actions/session.ts @@ -364,6 +364,21 @@ function waitForSessionInputReady( /** A reload successor whose agent went away while it spawned. */ type ReloadOrphan = { newId: SessionId; owner: Pick } +/** + * The ids a replacement's relationship remap may keep pointing at: every live + * session, plus every session the undo stack can still restore (#1379). See + * UndoCloseStack.restorableSessionIds for why a closed-but-restorable parent + * must not be treated as gone. + */ +function knownOrRestorable( + sessions: Record, + undoStack: { restorableSessionIds(): Set }, +): Set { + const known = new Set(Object.keys(sessions) as SessionId[]) + for (const id of undoStack.restorableSessionIds()) known.add(id) + return known +} + export function useSessionActions( state: { activeTabId: string; sessions: Record; tabs: Tab[] }, setState: WorkspaceSetState, @@ -1431,7 +1446,12 @@ export function useSessionActions( // swap has to update those too or the child renders top-level and // parent-scoped orchestration reads break. (rehydrate already does // this; reload/switch/resume/rewind funnel through here and didn't.) - sessions: remapSessionsRelationships(sessions, idMap), + // + // WHY the undo stack's ids count as "known" (#1379): a pointer + // whose target is neither live nor in idMap is dropped, and a + // parent that is closed but still restorable is neither. Dropping + // it here meant undoing that parent later had nothing to relink. + sessions: remapSessionsRelationships(sessions, idMap, knownOrRestorable(sessions, refs.undoStackRef.current)), // A pinned agent that gets a fresh id on reload/switch must follow // to the new id instead of silently dropping out of the Pinned list. pinnedSessionIds: remapPinnedSessionIds(prev.pinnedSessionIds, idMap), @@ -1491,6 +1511,7 @@ export function useSessionActions( refs.latestRuntimesRef, refs.seenUuidsRef, refs.stateRef, + refs.undoStackRef, setRuntimes, setState, spawn, @@ -1815,8 +1836,9 @@ export function useSessionActions( ...prev, // A fresh id: remap relationship pointers across all sessions // (children keep pointing at the right parent), the pinned list - // (pins follow) and the lanes. - sessions: remapSessionsRelationships(sessions, idMap), + // (pins follow) and the lanes. A pointer to a closed parent that + // undo can still restore survives, as in replace (#1379). + sessions: remapSessionsRelationships(sessions, idMap, knownOrRestorable(sessions, refs.undoStackRef.current)), pinnedSessionIds: remapPinnedSessionIds(prev.pinnedSessionIds, idMap), stage: remapTiledLanes(prev.stage, idMap), } @@ -1828,6 +1850,7 @@ export function useSessionActions( refs.latestRuntimesRef, refs.seenUuidsRef, refs.stateRef, + refs.undoStackRef, refs.useProxyStreamingRef, setRuntimes, setState, From 713408622bdebcaf49d8fe4b891298ee9756d0f3 Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sun, 27 Sep 2026 01:36:34 -0700 Subject: [PATCH 25/48] test(sessions): give the recovery test's hand-built refs an undo stack Replace now consults refs.undoStackRef for restorable parents (#1379); production refs always carry it, this test's partial mock did not. Co-Authored-By: Claude Opus 5.5 --- .../workspace/hook/actions/sessionRecovery.renderer.test.tsx | 3 +++ 1 file changed, 3 insertions(+) diff --git a/src/renderer/src/workspace/hook/actions/sessionRecovery.renderer.test.tsx b/src/renderer/src/workspace/hook/actions/sessionRecovery.renderer.test.tsx index 1847cdd3a..e5ab170cb 100644 --- a/src/renderer/src/workspace/hook/actions/sessionRecovery.renderer.test.tsx +++ b/src/renderer/src/workspace/hook/actions/sessionRecovery.renderer.test.tsx @@ -9,6 +9,7 @@ import type { SessionId, WorkspaceState } from '@renderer/workspace/types' import { killSessionBackendIfOwned, useSessionActions } from './session' import { freshStage } from '@renderer/workspace/dispatch/gridShape' +import { UndoCloseStack } from '@renderer/lib/undoClose' import { oneLaneStage } from '@renderer/workspace/testing/stageFixtures' const originalApiDescriptor = Object.getOwnPropertyDescriptor(window, 'api') @@ -426,6 +427,8 @@ describe('useSessionActions recovery retry', () => { useProxyStreamingRef: ref(false), defaultBuiltInMcpDomainsRef: ref([]), seenUuidsRef: ref({}), + // Replace consults the undo stack for restorable parents (#1379). + undoStackRef: ref(new UndoCloseStack()), } as unknown as WorkspaceRefs const setState = (next: WorkspaceState | ((prev: WorkspaceState) => WorkspaceState)) => { state = typeof next === 'function' ? next(state) : next From 82772d6b2521ede6f728763dc476fa8665e38322 Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sun, 27 Sep 2026 01:50:27 -0700 Subject: [PATCH 26/48] fix(sessions): relink end to end on #1374, and drop kept pointers once the parent can never come back #1387 review a: - (major) keeping the pointer only helps if restoring the parent relinks it: #1374 is merged in (Depends on #1374) and an end-to-end test closes P, reloads its live child, undoes P and asserts the child names P' (red without #1374's remap). - (minor) a kept pointer lingered as a ghost when P's entry left the stack without a restore. dropPointersTo drops pointers to exactly those ids (never unrelated or cross-window ones) when an entry is consumed as stale/failed, expires or is evicted; the stack notifies on a microtask because prune can run inside a setState updater. Each path is pinned. Co-Authored-By: Claude Opus 5.5 --- src/renderer/src/lib/undoClose.test.ts | 39 ++++++++++ src/renderer/src/lib/undoClose.ts | 78 +++++++++++++++++-- .../closedParentPointer.renderer.test.tsx | 44 ++++++++++- .../src/workspace/hook/actions/undoClose.ts | 27 ++++++- 4 files changed, 175 insertions(+), 13 deletions(-) diff --git a/src/renderer/src/lib/undoClose.test.ts b/src/renderer/src/lib/undoClose.test.ts index 5551cdbee..2da89623d 100644 --- a/src/renderer/src/lib/undoClose.test.ts +++ b/src/renderer/src/lib/undoClose.test.ts @@ -140,3 +140,42 @@ describe('UndoCloseStack.restorableSessionIds', () => { expect(stack.restorableSessionIds().size).toBe(0) }) }) + +// #1387 review a (minor): a pointer kept for a restorable parent must go once +// the parent's entry leaves the stack without a restore. +describe('dropPointersTo and the stack\'s dropped notifications', () => { + it('drops only pointers to gone, non-live ids, and keeps untouched rows', async () => { + const { dropPointersTo } = await import('./undoClose') + const child = { cwd: '/p', kind: 'claude', linkedParentId: 'gone', orchestrationParentId: 'gone', orchestrationRootId: 'other' } as unknown as SessionMeta + const other = { cwd: '/p', kind: 'claude', orchestrationParentId: 'cross-window' } as unknown as SessionMeta + const sessions = { child, other } as Record + const out = dropPointersTo(sessions as never, new Set(['gone']) as never) as Record + expect(out.child).toEqual({ cwd: '/p', kind: 'claude', orchestrationRootId: 'other' }) + expect(out.other).toBe(other) + expect(dropPointersTo(sessions as never, new Set(['nothing']) as never)).toBe(sessions) + }) + + it('keeps a pointer to a listed id that is live (defensive: an id the stack reports but the workspace still holds)', async () => { + const { dropPointersTo } = await import('./undoClose') + const parent = { cwd: '/p', kind: 'claude' } as unknown as SessionMeta + const child = { cwd: '/p', kind: 'claude', orchestrationParentId: 'parent' } as unknown as SessionMeta + const sessions = { parent, child } as Record + expect(dropPointersTo(sessions as never, new Set(['parent']) as never)).toBe(sessions) + }) + + it('reports expired and evicted entries on a microtask', async () => { + let now = 1_000_000 + const stack = new UndoCloseStack(() => now) + const dropped: string[][] = [] + stack.setDroppedListener(ids => { dropped.push([...ids].sort()) }) + const meta = { cwd: '/p', kind: 'claude', projectId: 't', joinedAt: 0 } as never + stack.push({ type: 'session', closedAt: now, sessionId: 'first' as never, sessionMeta: meta } as never) + for (let i = 0; i < UNDO_CLOSE_MAX_ENTRIES; i++) stack.push({ type: 'session', closedAt: now, sessionId: `s${i}` as never, sessionMeta: meta } as never) + await Promise.resolve() + expect(dropped).toEqual([['first']]) + now += UNDO_CLOSE_RETENTION_MS + 1 + expect(stack.length).toBe(0) + await Promise.resolve() + expect(dropped.at(-1)?.length).toBe(UNDO_CLOSE_MAX_ENTRIES) + }) +}) diff --git a/src/renderer/src/lib/undoClose.ts b/src/renderer/src/lib/undoClose.ts index 2aa48e2b4..9df400934 100644 --- a/src/renderer/src/lib/undoClose.ts +++ b/src/renderer/src/lib/undoClose.ts @@ -231,10 +231,75 @@ export function remapSingleEntryLineage(entry: SingleClosedEntry, lineage: UndoL } } +/** Every session id an entry could restore (session, project and group members). */ +function entrySessionIds(entry: ClosedEntry, into: Set = new Set()): Set { + if (entry.type === 'group') { + for (const member of entry.entries) entrySessionIds(member, into) + } else if (entry.type === 'session') { + into.add(entry.sessionId) + } else { + for (const member of entry.sessions) into.add(member.sessionId) + } + return into +} + +/** + * Drop relationship pointers that name one of `gone` while that id is not a + * live session (#1387 review a). Replace and Reload Agents now KEEP a pointer + * to a closed parent while undo can still restore it (#1379). When the entry + * then leaves the stack WITHOUT being restored (expired, evicted, or consumed + * as stale/failed), nothing else would ever clear that pointer: the child kept + * naming a dead id until restart, still counted as an orchestration worker, + * and was invisible to every parent. Only pointers to exactly these ids are + * touched, so unrelated or cross-window pointers are left alone. Untouched rows + * keep their identity. + */ +export function dropPointersTo( + sessions: Record, + gone: ReadonlySet, +): Record { + let changed = false + const out: Record = {} + const dead = (id: SessionId | undefined): boolean => id !== undefined && gone.has(id) && !(id in sessions) + for (const [id, meta] of Object.entries(sessions) as Array<[SessionId, SessionMeta]>) { + if (!dead(meta.linkedParentId) && !dead(meta.orchestrationParentId) && !dead(meta.orchestrationRootId)) { + out[id] = meta + continue + } + const next = { ...meta } + if (dead(meta.linkedParentId)) delete next.linkedParentId + if (dead(meta.orchestrationParentId)) delete next.orchestrationParentId + if (dead(meta.orchestrationRootId)) delete next.orchestrationRootId + out[id] = next + changed = true + } + return changed ? out : sessions +} + // ---- Stack ---- export class UndoCloseStack { private entries: ClosedEntry[] = [] + private droppedListener: ((ids: Set) => void) | null = null + + /** + * Told the session ids of entries that left the stack WITHOUT a restore + * (expired or evicted), so the workspace can drop pointers to them (see + * dropPointersTo). Delivered on a microtask, never synchronously: prune runs + * inside restorableSessionIds, which is called from within a setState + * updater, and a listener that set state there would nest updates. + */ + setDroppedListener(listener: ((ids: Set) => void) | null): void { + this.droppedListener = listener + } + + private notifyDropped(entries: ClosedEntry[]): void { + const listener = this.droppedListener + if (!listener || entries.length === 0) return + const ids = new Set() + for (const entry of entries) entrySessionIds(entry, ids) + queueMicrotask(() => listener(ids)) + } constructor(private readonly now: () => number = Date.now) {} @@ -243,6 +308,7 @@ export class UndoCloseStack { this.prune() this.entries.push(entry) if (this.entries.length > UNDO_CLOSE_MAX_ENTRIES) { + this.notifyDropped(this.entries.slice(0, this.entries.length - UNDO_CLOSE_MAX_ENTRIES)) this.entries = this.entries.slice(-UNDO_CLOSE_MAX_ENTRIES) } } @@ -278,14 +344,7 @@ export class UndoCloseStack { restorableSessionIds(): Set { this.prune() const ids = new Set() - const add = (entry: SingleClosedEntry): void => { - if (entry.type === 'session') ids.add(entry.sessionId) - else for (const member of entry.sessions) ids.add(member.sessionId) - } - for (const entry of this.entries) { - if (entry.type === 'group') entry.entries.forEach(add) - else add(entry) - } + for (const entry of this.entries) entrySessionIds(entry, ids) return ids } @@ -298,6 +357,9 @@ export class UndoCloseStack { private prune(): void { const cutoff = this.now() - UNDO_CLOSE_RETENTION_MS + const expired = this.entries.filter(e => e.closedAt <= cutoff) + if (expired.length === 0) return this.entries = this.entries.filter(e => e.closedAt > cutoff) + this.notifyDropped(expired) } } diff --git a/src/renderer/src/workspace/hook/actions/closedParentPointer.renderer.test.tsx b/src/renderer/src/workspace/hook/actions/closedParentPointer.renderer.test.tsx index fc51452bf..93bc14d54 100644 --- a/src/renderer/src/workspace/hook/actions/closedParentPointer.renderer.test.tsx +++ b/src/renderer/src/workspace/hook/actions/closedParentPointer.renderer.test.tsx @@ -4,7 +4,8 @@ import { useAppStore } from '@renderer/app-state/store' import { emptyRuntime } from '@renderer/session-runtime/state' import type { SessionMeta } from '@renderer/workspace/types' import { useSessionActions } from './session' -import { makeRefs } from './testing/paneActionsHarness' +import { makeRefs, sessionActionsWithSpawn } from './testing/paneActionsHarness' +import { useUndoCloseAction } from './undoClose' vi.mock('@renderer/workspace/hook/actions/initialHistory', () => ({ loadInitialHistoryForSession: vi.fn(async () => undefined) })) // #1379: replace and Reload Agents drop any relationship pointer whose target @@ -40,7 +41,8 @@ function harness(parent: string, restorable: boolean) { ...originalApi, spawnSession: vi.fn(async () => ({ sessionId: 'child-next' })), killOwnedSession: vi.fn(async () => true), carryGoalLoop: vi.fn(async () => null), carryWorkflowRuns: vi.fn(async () => undefined), controlGoalLoop: vi.fn(async () => null), } - return renderHook(() => useSessionActions(state, useAppStore.getState().setWorkspaceState, useAppStore.getState().setWorkspaceRuntimes, refs)) + const mounted = renderHook(() => useSessionActions(state, useAppStore.getState().setWorkspaceState, useAppStore.getState().setWorkspaceRuntimes, refs)) + return Object.assign(mounted, { refs, state }) } const pointers = (id: string) => { @@ -71,3 +73,41 @@ it('still drops pointers to a parent that is neither live nor restorable', async }) expect(pointers('child-next')).toEqual({ linkedParentId: undefined, orchestrationParentId: undefined, orchestrationRootId: undefined }) }) + +// #1387 review a (major): keeping the pointer only matters if restoring the +// parent then relinks it (#1374's remap of live rows, stacked under this PR). +// End to end: close P, reload its live child C to C', undo P to P': C' must +// point at P', or C' renders top-level and P' cannot see it. +it('relinks a reloaded child to its parent restored by Undo Close', async () => { + const mounted = harness('closed-parent', true) + await act(async () => { + expect(await mounted.result.current.replaceSession('/recorded/project', { targetSessionId: 'child', kind: 'claude', resumeSessionId: 'native-child' })).toBe('child-next') + }) + const undo = renderHook(() => useUndoCloseAction( + mounted.state, useAppStore.getState().setWorkspaceState, mounted.refs, + sessionActionsWithSpawn(vi.fn().mockResolvedValue('parent-restored')), + )) + await act(async () => { await undo.result.current.undoClose() }) + expect(useAppStore.getState().workspaceState.sessions['parent-restored']).toBeDefined() + expect(pointers('child-next')).toEqual({ linkedParentId: 'parent-restored', orchestrationParentId: 'parent-restored', orchestrationRootId: 'parent-restored' }) +}) + +// #1387 review a (minor): if P's entry is consumed WITHOUT a restore (its +// project was removed, so it is stale), nothing can bring P back; the kept +// pointer must not linger as a ghost. +it('drops the kept pointers once the parent\'s entry is consumed without a restore', async () => { + const mounted = harness('closed-parent', true) + await act(async () => { + await mounted.result.current.replaceSession('/recorded/project', { targetSessionId: 'child', kind: 'claude', resumeSessionId: 'native-child' }) + }) + expect(pointers('child-next').orchestrationParentId).toBe('closed-parent') + // P's project disappears (e.g. merged away): the undo is judged stale. + act(() => { useAppStore.getState().setWorkspaceState(prev => ({ ...prev, tabs: [{ id: 'elsewhere', title: 'Elsewhere' }] })) }) + mounted.refs.stateRef.current = useAppStore.getState().workspaceState + const undo = renderHook(() => useUndoCloseAction( + mounted.state, useAppStore.getState().setWorkspaceState, mounted.refs, + sessionActionsWithSpawn(vi.fn().mockResolvedValue('never')), + )) + await act(async () => { await undo.result.current.undoClose() }) + expect(pointers('child-next')).toEqual({ linkedParentId: undefined, orchestrationParentId: undefined, orchestrationRootId: undefined }) +}) diff --git a/src/renderer/src/workspace/hook/actions/undoClose.ts b/src/renderer/src/workspace/hook/actions/undoClose.ts index 4392e7013..080d5ee6a 100644 --- a/src/renderer/src/workspace/hook/actions/undoClose.ts +++ b/src/renderer/src/workspace/hook/actions/undoClose.ts @@ -4,7 +4,7 @@ import { sessionDisplayTitle } from '@renderer/workspace/sessionDisplayTitle' import { sessionMcpOverrides } from '@renderer/workspace/mcpDomains' import { DEFAULT_PROVIDER, isAgentSessionKind } from '@shared/types/providerKind' import { MISSING_WORKSPACE_FOLDER_PREFIX, SESSION_START_FAILED_MESSAGE } from '@shared/types/session' -import { useCallback, useState } from 'react' +import { useCallback, useEffect, useState } from 'react' import type { SessionId, @@ -14,6 +14,7 @@ import type { } from '@renderer/workspace/types' import { remapTiledLanes } from '@renderer/workspace/dispatch/tiledDispatchSelectors' import { + dropPointersTo, remapMetaLineage, remapSingleEntryLineage, } from '@renderer/lib/undoClose' @@ -480,6 +481,22 @@ export function useUndoCloseAction( // they are finished history in the store, not live state to stop. // `publish` is the source of truth for "came back": each restore calls it // after its commit and only then, with exactly the old -> new pairs. + // Pointers kept for a restorable parent (#1379) must go once nothing can + // restore it. See dropPointersTo. + const dropGhostPointers = useCallback((gone: Set) => { + if (gone.size === 0) return + setState(prev => { + const sessions = dropPointersTo(prev.sessions, gone) + return sessions === prev.sessions ? prev : { ...prev, sessions } + }) + }, [setState]) + // Entries that expire or are evicted leave the stack without a restore. + useEffect(() => { + const stack = refs.undoStackRef.current + stack.setDroppedListener(dropGhostPointers) + return () => stack.setDroppedListener(null) + }, [dropGhostPointers, refs.undoStackRef]) + const restoreSingleEntry = useCallback( async (entry: SingleClosedEntry, publish: PublishLineage): Promise => { const successors = new Map() @@ -496,10 +513,14 @@ export function useUndoCloseAction( const closedIds = entry.type === 'session' ? [entry.sessionId] : entry.sessions.map(member => member.sessionId) - stopGoalLoops(closedIds.filter(id => !successors.has(id))) + const notBack = closedIds.filter(id => !successors.has(id)) + stopGoalLoops(notBack) + // A consumed entry whose members did not come back can never be + // restored now: drop live pointers to them (#1387 review a). + dropGhostPointers(new Set(notBack)) return result }, - [hasGoalLoopTools, restoreSessionEntry, restoreTabEntry], + [dropGhostPointers, hasGoalLoopTools, restoreSessionEntry, restoreTabEntry], ) // Replay one close OPERATION's units last-first (see ClosedGroup). From 7a5c9c909ebd9f457d70f94dc779b1d92f95d415 Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sun, 27 Sep 2026 02:06:02 -0700 Subject: [PATCH 27/48] fix(undo-close): a group drops pointers to members it consumed, from members restored after them #1387 review a, round 2: undo replays a group newest-first, so a parent consumed as stale (folder gone) could precede its child in the same group; the child then came back carrying pointers to a parent that can never return. The group remembers consumed ids and drops pointers to them once it has replayed. Review b, round 2: a hook-level test now proves the stack's expiry notification reaches the workspace (red without the listener). Co-Authored-By: Claude Opus 5.5 --- .../src/workspace/hook/actions/undoClose.ts | 14 +++- .../undoCloseGhostPointers.renderer.test.tsx | 74 +++++++++++++++++++ 2 files changed, 87 insertions(+), 1 deletion(-) create mode 100644 src/renderer/src/workspace/hook/actions/undoCloseGhostPointers.renderer.test.tsx diff --git a/src/renderer/src/workspace/hook/actions/undoClose.ts b/src/renderer/src/workspace/hook/actions/undoClose.ts index 080d5ee6a..8526d3ec2 100644 --- a/src/renderer/src/workspace/hook/actions/undoClose.ts +++ b/src/renderer/src/workspace/hook/actions/undoClose.ts @@ -550,6 +550,14 @@ export function useUndoCloseAction( // review: a deleted-folder member was retried, and re-toasted, on // every later ⌘⇧T). let consumedAny = false + // WHY the group remembers consumed ids (#1387 review a, round 2): undo + // replays newest-first, so a parent can be consumed as stale BEFORE its + // child in the same group comes back, carrying its old pointers to that + // parent. The per-entry drop in restoreSingleEntry ran while the child + // was not live yet, so it missed it. Dropping once the whole group has + // replayed catches every member restored after the consumed one. + const consumedIds = new Set() + const dropConsumed = (): void => dropGhostPointers(consumedIds) while (remaining.length > 0) { const member = remaining[remaining.length - 1] remaining = remaining.slice(0, -1) @@ -561,6 +569,8 @@ export function useUndoCloseAction( restoredAny = true } else if (result === 'stale') { consumedAny = true + if (member.type === 'session') consumedIds.add(member.sessionId) + else for (const closed of member.sessions) consumedIds.add(closed.sessionId) } else if (result === 'retryable-failure') { if (!restoredAny && !consumedAny) return 'retryable-failure' const rest = [...remaining, member] @@ -568,12 +578,14 @@ export function useUndoCloseAction( refs.undoStackRef.current.push(leftover) // Part of the group came back; say what did not (#1242). showToast(restoreFailureMessage(leftover), RESTORE_FAILURE_TOAST_MS) + dropConsumed() return 'restored' } } + dropConsumed() return restoredAny ? 'restored' : 'stale' }, - [refs.undoStackRef, restoreSingleEntry, showToast], + [dropGhostPointers, refs.undoStackRef, restoreSingleEntry, showToast], ) const undoClose = useCallback(async () => { diff --git a/src/renderer/src/workspace/hook/actions/undoCloseGhostPointers.renderer.test.tsx b/src/renderer/src/workspace/hook/actions/undoCloseGhostPointers.renderer.test.tsx new file mode 100644 index 000000000..d58ccfc37 --- /dev/null +++ b/src/renderer/src/workspace/hook/actions/undoCloseGhostPointers.renderer.test.tsx @@ -0,0 +1,74 @@ +import { act, render } from '@testing-library/react' +import { expect, it, vi } from 'vitest' + +import { useUndoCloseAction } from '@renderer/workspace/hook/actions/undoClose' +import { makeRefs, sessionActionsWithSpawn, stateWriter } from '@renderer/workspace/hook/actions/testing/paneActionsHarness' +import { MissingWorkspaceDirectoryError } from '@main/workspaceDirectory' +import { UNDO_CLOSE_RETENTION_MS } from '@renderer/lib/undoClose' +import { freshStage } from '@renderer/workspace/dispatch/gridShape' +import type { SessionMeta, WorkspaceState } from '@renderer/workspace/types' + +// #1387: a pointer kept for a restorable parent (#1379) must not outlive the +// last chance to restore that parent. + +function mount(sessions: Record, spawn: ReturnType) { + const state: WorkspaceState = { + tabs: [{ id: 'tab-parent', title: 'agent-code' }], activeTabId: 'tab-parent', stage: freshStage(), + sessions, pinnedSessionIds: [], + } + const refs = makeRefs(state) + const writer = stateWriter(state, refs) + let actions!: ReturnType + function Harness(): React.JSX.Element { + actions = useUndoCloseAction(state, writer.setState, refs, sessionActionsWithSpawn(spawn), vi.fn()) + return
+ } + const mounted = render() + return { refs, writer, undo: () => actions.undoClose(), rerender: () => mounted.rerender(), unmount: () => mounted.unmount() } +} + +const meta = (extra: Partial): SessionMeta => + ({ cwd: '/projects/agent-code', kind: 'claude', projectId: 'tab-parent', joinedAt: 1, ...extra }) as SessionMeta + +// Review a, round 2: ONE close operation recorded [child C, parent P]. Undo +// replays last-first: P's folder is gone, so P is consumed as stale; C then +// comes back carrying its old pointers to P. P can never return, so C must not +// keep naming it. +it('drops pointers to a group member consumed as stale from a member restored after it', async () => { + const gone = '/projects/deleted-worktree' + const relayed = `Error invoking remote method 'session:spawn': ${String(new MissingWorkspaceDirectoryError(gone))}` + const spawn = vi.fn().mockRejectedValueOnce(new Error(relayed)).mockResolvedValueOnce('restored-child') + const harness = mount({}, spawn) + harness.refs.undoStackRef.current.push({ + type: 'group', closedAt: Date.now(), + entries: [ + { type: 'session', closedAt: Date.now(), sessionId: 'child', sessionMeta: meta({ linkedParentId: 'parent', orchestrationParentId: 'parent', orchestrationRootId: 'parent' }) }, + { type: 'session', closedAt: Date.now(), sessionId: 'parent', sessionMeta: meta({ cwd: gone }) }, + ], + } as never) + await act(async () => { await harness.undo() }) + const restored = harness.writer.getState().sessions['restored-child'] + expect(restored).toBeDefined() + expect(restored).not.toHaveProperty('linkedParentId') + expect(restored).not.toHaveProperty('orchestrationParentId') + expect(restored).not.toHaveProperty('orchestrationRootId') + harness.unmount() +}) + +// Review b, round 2: the stack's expiry notification must actually reach the +// workspace (the listener registration in useUndoCloseAction), not just fire. +it('drops a live child\'s kept pointer when its parent\'s undo entry expires', async () => { + const spy = vi.spyOn(Date, 'now') + const start = 5_000_000_000 + spy.mockReturnValue(start) + const harness = mount({ child: meta({ orchestrationParentId: 'parent', orchestrationRootId: 'parent' }) }, vi.fn()) + harness.refs.undoStackRef.current.push({ type: 'session', closedAt: start, sessionId: 'parent', sessionMeta: meta({}) } as never) + spy.mockReturnValue(start + UNDO_CLOSE_RETENTION_MS + 1) + // The next read of the stack (a render reads its length) prunes the entry. + await act(async () => { harness.rerender(); await Promise.resolve() }) + const child = harness.writer.getState().sessions.child + expect(child).not.toHaveProperty('orchestrationParentId') + expect(child).not.toHaveProperty('orchestrationRootId') + spy.mockRestore() + harness.unmount() +}) From 02e03d92002afd879c4369f59cf458bb55929067 Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sun, 27 Sep 2026 02:12:25 -0700 Subject: [PATCH 28/48] fix(undo-close): a leftover pushed back for retry no longer points at a member the group consumed #1387 review a, round 3: when a group consumed parent P as stale and then hit a transient failure on child C, C went back on the stack with its entry metadata still naming P, and the retry restored that dead pointer. The leftover's metadata now drops pointers to the replay's consumed ids (dropEntryPointersTo). Red before, on the reviewer's exact sequence. Co-Authored-By: Claude Opus 5.5 --- src/renderer/src/lib/undoClose.ts | 23 ++++++++++++++ .../src/workspace/hook/actions/undoClose.ts | 8 ++++- .../undoCloseGhostPointers.renderer.test.tsx | 30 +++++++++++++++++++ 3 files changed, 60 insertions(+), 1 deletion(-) diff --git a/src/renderer/src/lib/undoClose.ts b/src/renderer/src/lib/undoClose.ts index 9df400934..f57602016 100644 --- a/src/renderer/src/lib/undoClose.ts +++ b/src/renderer/src/lib/undoClose.ts @@ -276,6 +276,29 @@ export function dropPointersTo( return changed ? out : sessions } +/** + * The same rule as dropPointersTo, applied to a waiting entry's OWN metadata + * (#1387 review a, round 3). A group member pushed back for a retry must not + * carry pointers to a member the same group already consumed: that parent can + * never return, and the later restore would bring the dead pointer back. + */ +export function dropEntryPointersTo(entry: ClosedEntry, gone: ReadonlySet): ClosedEntry { + if (gone.size === 0) return entry + const clean = (meta: SessionMeta): SessionMeta => { + if (!(meta.linkedParentId && gone.has(meta.linkedParentId)) && + !(meta.orchestrationParentId && gone.has(meta.orchestrationParentId)) && + !(meta.orchestrationRootId && gone.has(meta.orchestrationRootId))) return meta + const next = { ...meta } + if (next.linkedParentId && gone.has(next.linkedParentId)) delete next.linkedParentId + if (next.orchestrationParentId && gone.has(next.orchestrationParentId)) delete next.orchestrationParentId + if (next.orchestrationRootId && gone.has(next.orchestrationRootId)) delete next.orchestrationRootId + return next + } + if (entry.type === 'group') return { ...entry, entries: entry.entries.map(member => dropEntryPointersTo(member, gone) as SingleClosedEntry) } + if (entry.type === 'session') return { ...entry, sessionMeta: clean(entry.sessionMeta) } + return { ...entry, sessions: entry.sessions.map(member => ({ ...member, meta: clean(member.meta) })) } +} + // ---- Stack ---- export class UndoCloseStack { diff --git a/src/renderer/src/workspace/hook/actions/undoClose.ts b/src/renderer/src/workspace/hook/actions/undoClose.ts index 8526d3ec2..87481193f 100644 --- a/src/renderer/src/workspace/hook/actions/undoClose.ts +++ b/src/renderer/src/workspace/hook/actions/undoClose.ts @@ -14,6 +14,7 @@ import type { } from '@renderer/workspace/types' import { remapTiledLanes } from '@renderer/workspace/dispatch/tiledDispatchSelectors' import { + dropEntryPointersTo, dropPointersTo, remapMetaLineage, remapSingleEntryLineage, @@ -574,7 +575,12 @@ export function useUndoCloseAction( } else if (result === 'retryable-failure') { if (!restoredAny && !consumedAny) return 'retryable-failure' const rest = [...remaining, member] - const leftover: ClosedEntry = rest.length === 1 ? rest[0] : { ...entry, entries: rest } + // Strip pointers to members this replay consumed from what goes + // back for retry, or the retry restores them (#1387 review a r3). + const leftover: ClosedEntry = dropEntryPointersTo( + rest.length === 1 ? rest[0] : { ...entry, entries: rest }, + consumedIds, + ) refs.undoStackRef.current.push(leftover) // Part of the group came back; say what did not (#1242). showToast(restoreFailureMessage(leftover), RESTORE_FAILURE_TOAST_MS) diff --git a/src/renderer/src/workspace/hook/actions/undoCloseGhostPointers.renderer.test.tsx b/src/renderer/src/workspace/hook/actions/undoCloseGhostPointers.renderer.test.tsx index d58ccfc37..1e467d47e 100644 --- a/src/renderer/src/workspace/hook/actions/undoCloseGhostPointers.renderer.test.tsx +++ b/src/renderer/src/workspace/hook/actions/undoCloseGhostPointers.renderer.test.tsx @@ -55,6 +55,36 @@ it('drops pointers to a group member consumed as stale from a member restored af harness.unmount() }) +// Review a, round 3: same group, but C's spawn fails TRANSIENTLY after P was +// consumed. C goes back on the stack as a leftover; its entry metadata must no +// longer name P, or the next Undo restores C pointing at a parent that can +// never return. +it('strips pointers to a consumed member from a leftover pushed back for retry', async () => { + const gone = '/projects/deleted-worktree' + const relayed = `Error invoking remote method 'session:spawn': ${String(new MissingWorkspaceDirectoryError(gone))}` + const spawn = vi.fn() + .mockRejectedValueOnce(new Error(relayed)) + .mockRejectedValueOnce(new Error('spawn timed out')) + .mockResolvedValueOnce('restored-child') + const harness = mount({}, spawn) + harness.refs.undoStackRef.current.push({ + type: 'group', closedAt: Date.now(), + entries: [ + { type: 'session', closedAt: Date.now(), sessionId: 'child', sessionMeta: meta({ title: 'Child', linkedParentId: 'parent', orchestrationParentId: 'parent', orchestrationRootId: 'parent' }) }, + { type: 'session', closedAt: Date.now(), sessionId: 'parent', sessionMeta: meta({ cwd: gone }) }, + ], + } as never) + await act(async () => { await harness.undo() }) + expect(harness.refs.undoStackRef.current.length).toBe(1) + await act(async () => { await harness.undo() }) + const restored = harness.writer.getState().sessions['restored-child'] + expect(restored).toBeDefined() + expect(restored).not.toHaveProperty('linkedParentId') + expect(restored).not.toHaveProperty('orchestrationParentId') + expect(restored).not.toHaveProperty('orchestrationRootId') + harness.unmount() +}) + // Review b, round 2: the stack's expiry notification must actually reach the // workspace (the listener registration in useUndoCloseAction), not just fire. it('drops a live child\'s kept pointer when its parent\'s undo entry expires', async () => { From fcadd1934aeefd30aec15d28aac938ae1f47329a Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sun, 27 Sep 2026 04:27:53 -0700 Subject: [PATCH 29/48] docs(plans): C5 fail-all batch, rows verified on main (#1251) Co-Authored-By: Claude Opus 5.5 --- docs/plans/2026-09-27-c5-fail-all-batch.md | 33 ++++++++++++++++++++++ 1 file changed, 33 insertions(+) create mode 100644 docs/plans/2026-09-27-c5-fail-all-batch.md diff --git a/docs/plans/2026-09-27-c5-fail-all-batch.md b/docs/plans/2026-09-27-c5-fail-all-batch.md new file mode 100644 index 000000000..fdf919351 --- /dev/null +++ b/docs/plans/2026-09-27-c5-fail-all-batch.md @@ -0,0 +1,33 @@ +# C5 fail-all batch (#1251) + +Source: the read-only C5 hunt in `temp/quality-loop/hunt-c5.md`, rows 8–13. Each row was verified on origin/main `5e22c7b0` before any fix. "Fail-first" means the new test was run red against main's implementation first. + +## Principle + +One bad record must cost only itself. Two constraints shape every fix: + +1. **Owner rule: do not delete stuff often (2026-09-27).** A record this build cannot read is carried verbatim whenever the file is rewritten, never dropped. +2. **Ambiguity fails closed (q40).** + - A skipped approval grants nothing. + - A skipped ledger row can only fail to protect a bundle it does not name. + - Destructive transforms keep refusing (row 10). + +## Rows + +| Row | Verified on main | Decision | Test (fail-first) | +|---|---|---|---| +| 8: Codex rollouts, `conversations/sources/codex.ts` | Yes. `readRolloutHead` streams through readline, which rethrows EACCES/EIO. `fromHead` and both discovery loops await it with no catch, so `discover()` rejected and the Codex column emptied. | Skip the unreadable rollout (it has no cwd to scope it by); cache nothing, so it is retried later. | `codex.system.test.ts`, "skips an unreadable rollout…": red with EACCES. | +| 9: monitor incidents, `performance/MonitorHistoryStore.ts` | Yes, and worse than the hunt said. One unparseable row hid the run's whole incident list. On a helper restart in that run, `persistIncidents` merged from an empty list and rewrote `incidents.json`, erasing the readable evidence AND the unknown row. Realistic: Preview and stable builds share this directory. | Parse per row. Carry unknown rows per run in `foreignIncidents`, and re-append them on every rewrite via `incidentFileBody`. Leave room under `INCIDENT_LIMIT`, keep the run from expiry-by-emptiness, and mark the store degraded. | `MonitorHistoryStore.test.ts`, "keeps readable incidents…": red (`null` for the readable incident). | +| 10: Pi JSONL, `providerSwitch/piTranscript.ts` | Real, but **strict by design**. `loadPiSnapshotAt` feeds destructive transforms (switch, duplicate, rewind). Skipping a malformed middle line would move or rewind a conversation with a silent hole. | No change. The error names the file and line and never reaches a toast raw. | none | +| 11: workflow approvals, `workflows/WorkflowSourceApprovalStore.ts` | Yes. `load()` threw on the first bad entry and never set `loaded`, so every `authorize()` rethrew: all repository workflows were blocked. | Skip the entry (it approves nothing, so its source is prompted again) and carry it verbatim through `persist()`. A wrong file version still throws, because that is not one bad row. | `WorkflowSourceApprovalStore.test.ts`, "honours valid approvals…": red. | +| 12: TLDR batch, `main/tldr/ipc.ts` | Yes. `z.array(z.string().refine(validTldrIdentity))` rejected the whole batch, and Agent Activity reads every TLDR and goal in one batch. | Keep the payload shape strict (a bounded array of bounded strings) and drop invalid identities. This is exact: the store only writes valid identities, so an invalid one has no record. | New `tldr/ipc.test.ts`: red (ZodError). A second test pins that malformed payloads are still refused. | +| 13: legacy bundle ledger, `storage/debugRetention.ts` | Yes. A JSON-valid non-entry line (`null`, or a row with a non-string `bundlePath`) threw TypeError, which rejected `collectArtifacts` and stopped every prune pass. | Extract `parseManualLegacyBundlePaths`, then shape-check each row. A non-string reason counts as manual, so retention keeps the bundle. | `debugRetention.test.ts`, "keeps every readable manual row…": red (`null.event`). | + +Rows 14–15 (key vault index, agent-name registry, tmux recovery) are strict by design per the issue and are only recorded there. + +## Residuals + +- **Row 9:** + - Carried foreign rows never expire on their own. They leave disk only with their run directory (budget pruning, clear). + - A run whose incident file holds only foreign rows is kept from expiry-by-emptiness. It is still pruned by the data budget. +- **Row 12:** a renderer that sends an invalid identity gets no record and no error for it. That is the same answer as "no TLDR yet". From 318c0d5d201e273c228c253c8dbf4a08c310a9fc Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sun, 27 Sep 2026 04:27:53 -0700 Subject: [PATCH 30/48] fix(conversations): one unreadable Codex rollout no longer empties the list (#1251 row 8) Co-Authored-By: Claude Opus 5.5 --- .../sources/codex.system.test.ts | 30 ++++++++++++++++++- src/main/conversations/sources/codex.ts | 18 ++++++++++- 2 files changed, 46 insertions(+), 2 deletions(-) diff --git a/src/main/conversations/sources/codex.system.test.ts b/src/main/conversations/sources/codex.system.test.ts index 8b6fd394c..abd80057d 100644 --- a/src/main/conversations/sources/codex.system.test.ts +++ b/src/main/conversations/sources/codex.system.test.ts @@ -1,4 +1,4 @@ -import { rename } from 'node:fs/promises' +import { chmod, readdir, rename } from 'node:fs/promises' import { join } from 'node:path' import { DatabaseSync } from 'node:sqlite' import { afterEach, describe, expect, it } from 'vitest' @@ -90,6 +90,34 @@ describe('Codex conversation source', () => { expect(everywhere.filter(r => r.origin === 'scan')).toHaveLength(counts.codex.unindexedSampled) }) + it('skips an unreadable rollout instead of failing the whole Codex list (#1251 row 8)', async () => { + // readline's async iterator rethrows a stream error (EACCES here, EIO on a + // failing disk), and nothing between readRolloutHead and discover() caught + // it, so one rollout the app cannot open emptied the Codex column. + const { corpus, source, listWorktrees } = await setup() + const counts = corpus.manifest.counts as { codex: { inFamily: number; unindexedSampled: number } } + const rollouts = (await readdir(join(corpus.codexHome, 'sessions'), { recursive: true })) + .filter(name => /rollout-.*\.jsonl$/.test(name)).map(name => join(corpus.codexHome, 'sessions', name)) + expect(rollouts.length).toBeGreaterThan(1) + for (const file of rollouts) await chmod(file, 0o000) + // unshift: permissions come back before the corpus cleanup removes the tree. + cleanups.unshift(async () => { for (const file of rollouts) await chmod(file, 0o600) }) + const family = await resolveFamily('/fixture/repo', 'everywhere', { listWorktrees }) + + // Index path: the indexed rows never open a rollout and must all survive; + // the unindexed union is what reads heads, and it now skips what it cannot. + const indexed = await source.discover({ scope: 'everywhere', family }) + expect(indexed.filter(r => r.origin === 'index').length).toBeGreaterThanOrEqual(counts.codex.inFamily) + expect(indexed.filter(r => r.origin === 'scan')).toHaveLength(0) + + // Fallback path: one readable rollout still lists beside unreadable ones. + await chmod(rollouts[0]!, 0o600) + await rename(join(corpus.codexHome, 'state_5.sqlite'), join(corpus.codexHome, 'state_5.sqlite.away')) + const fresh = new CodexConversationSource({ codexHome: corpus.codexHome }) + const scanned = await fresh.discover({ scope: 'everywhere', family }) + expect(scanned.map(r => r.file)).toEqual([rollouts[0]]) + }) + it('falls back to the rollout scan when the index is missing and reports why', async () => { const { corpus, source, listWorktrees } = await setup() await rename(join(corpus.codexHome, 'state_5.sqlite'), join(corpus.codexHome, 'state_5.sqlite.away')) diff --git a/src/main/conversations/sources/codex.ts b/src/main/conversations/sources/codex.ts index 5a25536ca..5735586fb 100644 --- a/src/main/conversations/sources/codex.ts +++ b/src/main/conversations/sources/codex.ts @@ -169,7 +169,23 @@ export class CodexConversationSource implements ConversationSource { } } const cached = this.heads.get(file) - const head = cached && cached.mtime === mtime ? cached.head : await readRolloutHead(file) + let head: RolloutHead + if (cached && cached.mtime === mtime) head = cached.head + else { + // WHY one unreadable rollout is skipped here (#1251 row 8): readline's + // async iterator rethrows a stream error (EACCES, EIO, a file removed + // between the walk and the read), and both callers await this in a plain + // loop, so a single rollout the app cannot open used to reject + // discover() and empty the whole Codex column. Skipping costs exactly + // the row that cannot be labelled anyway (without its head there is no + // cwd to scope it by). Nothing is cached, so the next discovery retries + // it once the file is readable again. + try { + head = await readRolloutHead(file) + } catch { + return null + } + } this.heads.set(file, { mtime, head }) if (scope.scope !== 'everywhere' && !scope.family.matches(head.cwd)) return null return { From e14cd3eb3ffe1742d24a07ccf26d1602860b667c Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sun, 27 Sep 2026 04:27:53 -0700 Subject: [PATCH 31/48] fix(performance): keep readable incidents and carry unknown rows through rewrites (#1251 row 9) Co-Authored-By: Claude Opus 5.5 --- .../performance/MonitorHistoryStore.test.ts | 25 ++++++++ src/main/performance/MonitorHistoryStore.ts | 57 ++++++++++++++----- 2 files changed, 69 insertions(+), 13 deletions(-) diff --git a/src/main/performance/MonitorHistoryStore.test.ts b/src/main/performance/MonitorHistoryStore.test.ts index ea27de350..80eb5df73 100644 --- a/src/main/performance/MonitorHistoryStore.test.ts +++ b/src/main/performance/MonitorHistoryStore.test.ts @@ -108,6 +108,31 @@ describe('bounded local performance history', () => { }) }) + it('keeps readable incidents and never erases a row it does not recognise (#1251 row 9)', async () => { + // A preview build can write an incident rule this build does not know, and + // the owner moves between the Preview and stable channels. One such row + // used to hide the run's whole incident list AND, on a helper restart in + // that run, persistIncidents replaced the file with only the new engine's + // rows, erasing the evidence it could not read. + const root = await mkdtemp(join(tmpdir(), 'agent-code-monitor-')) + roots.push(root) + const foreign = { ...incident, id: 9, at: 10_500, rule: 'rule-from-a-newer-build' } + await mkdir(join(root, 'runs', 'run-mixed'), { recursive: true }) + await writeFile(join(root, 'runs', 'run-mixed', 'incidents.json'), JSON.stringify([incident, foreign])) + + const restarted = new MonitorHistoryStore(root, 'run-mixed') + await restarted.settled() + restarted.record(snapshot(11_000), null, [{ ...incident, id: 2, at: 11_000 }], 0, 1) + await restarted.settled() + + expect(await restarted.readIncident(10_000, 1)).toMatchObject({ rule: 'renderer-stall' }) + expect(await restarted.readIncident(11_000, 2)).toMatchObject({ rule: 'renderer-stall' }) + expect(restarted.status().state).toBe('degraded') + const onDisk = JSON.parse(await readFile(join(root, 'runs', 'run-mixed', 'incidents.json'), 'utf8')) as Array<{ id: number; rule: string }> + expect(onDisk.map(row => row.id).sort()).toEqual([1, 2, 9]) + expect(onDisk.find(row => row.id === 9)).toEqual(foreign) + }) + it('repairs a torn append and keeps coarse tiers peak-preserving', async () => { const root = await mkdtemp(join(tmpdir(), 'agent-code-monitor-')) roots.push(root) diff --git a/src/main/performance/MonitorHistoryStore.ts b/src/main/performance/MonitorHistoryStore.ts index 2db7c5fea..4bead304d 100644 --- a/src/main/performance/MonitorHistoryStore.ts +++ b/src/main/performance/MonitorHistoryStore.ts @@ -64,6 +64,19 @@ export class MonitorHistoryStore { private exporting = false private index = new Map() private incidentRuns = new Map() + // Rows of a run's incidents.json that this build cannot parse, kept + // verbatim per run (#1251 row 9). WHY they are carried instead of dropped: + // the owner moves between the Preview and stable channels, so a newer build + // can leave an incident rule this one does not know. Hiding the run's whole + // list for one such row lost the readable evidence, and every rewrite of the + // file (a helper restart's persistIncidents, capture repair, expiry, + // eviction) then replaced it with only the rows this build understood, + // deleting the rest. Invariant: every write of a run's incident file goes + // through incidentFileBody, which re-appends these rows. They never enter + // incidentRuns, so queries, exports and the 50-incident eviction only ever + // see rows this build can vouch for; the rows leave disk only with the run + // directory itself (budget pruning, retention, clear). + private foreignIncidents = new Map() private repairedTails = new Set() // False until startup indexing completes. Retention deletes any run the // index does not know about, so a partial index (EPERM, ENOSPC or an I/O @@ -296,7 +309,7 @@ export class MonitorHistoryStore { try { await rm(join(this.root, RUNS_DIR), { recursive: true, force: true }) await mkdir(this.runDir, { recursive: true }) - this.index.clear(); this.incidentRuns.clear(); this.repairedTails.clear(); this.indexed = true + this.index.clear(); this.incidentRuns.clear(); this.foreignIncidents.clear(); this.repairedTails.clear(); this.indexed = true this.bytes = 0; this.shortened = false; this.degraded = false this.operationFingerprint = ''; this.unindexedRuns.clear() this.lastMaintenanceAt = -Infinity @@ -374,13 +387,13 @@ export class MonitorHistoryStore { } } const file = join(this.root, RUNS_DIR, run, 'incidents.json') - const stored = await this.readIncidentFile(file) + const stored = await this.readIncidentFile(file, run) // Every retained run is repaired, not only the current one. A crash or // force-quit in ANY earlier run left its last capture as "capturing" // forever, and only a helper restart within the same run fixed it. const repaired = stored.map(incident => incident.state === 'capturing' ? { ...incident, state: 'interrupted' as const } : incident) if (repaired.some((incident, position) => incident !== stored[position])) { - await this.replaceBounded(file, JSON.stringify(repaired), INCIDENT_BUDGET).catch(() => { this.degraded = true }) + await this.replaceBounded(file, this.incidentFileBody(run, repaired), INCIDENT_BUDGET).catch(() => { this.degraded = true }) } if (repaired.length) this.incidentRuns.set(run, repaired) } @@ -405,10 +418,13 @@ export class MonitorHistoryStore { const existing = this.incidentRuns.get(this.runId) ?? [] const merged = new Map(existing.map(incident => [`${incident.at}:${incident.id}`, incident])) for (const incident of current) merged.set(`${incident.at}:${incident.id}`, incident) - const rows = [...merged.values()].sort((a, b) => a.at - b.at).slice(-INCIDENT_LIMIT) + // Room is left for this run's carried foreign rows, or the rewritten file + // would exceed INCIDENT_LIMIT and the next launch would reject all of it. + const room = Math.max(0, INCIDENT_LIMIT - (this.foreignIncidents.get(this.runId)?.length ?? 0)) + const rows = room ? [...merged.values()].sort((a, b) => a.at - b.at).slice(-room) : [] // Memory mirrors disk: a capacity-shortened write keeps the previous rows, // which is what status, queries and the next launch will actually find. - if (!(await this.replaceBounded(join(this.runDir, 'incidents.json'), JSON.stringify(rows), INCIDENT_BUDGET))) return + if (!(await this.replaceBounded(join(this.runDir, 'incidents.json'), this.incidentFileBody(this.runId, rows), INCIDENT_BUDGET))) return this.incidentRuns.set(this.runId, rows) await this.enforceIncidentLimit() } @@ -545,7 +561,7 @@ export class MonitorHistoryStore { // with no remaining points or incidents holds only an unattributable // operations snapshot, so it is retention-expired, not capacity-pruned. if (this.indexed) for (const run of await this.runNames()) { - if (run === this.runId || this.incidentRuns.has(run) || this.unindexedRuns.has(run) || [...this.index.values()].some(entry => entry.run === run)) continue + if (run === this.runId || this.incidentRuns.has(run) || this.foreignIncidents.has(run) || this.unindexedRuns.has(run) || [...this.index.values()].some(entry => entry.run === run)) continue await rm(join(this.root, RUNS_DIR, run), { recursive: true, force: true }) } this.bytes = await this.diskBytes() @@ -580,11 +596,12 @@ export class MonitorHistoryStore { private async writeRunIncidents(run: string, rows: MonitorIncident[]): Promise { const file = join(this.root, RUNS_DIR, run, 'incidents.json') - if (rows.length) { - if (await this.replaceBounded(file, JSON.stringify(rows), INCIDENT_BUDGET)) this.incidentRuns.set(run, rows) - } else { + if (!rows.length && !this.foreignIncidents.has(run)) { await rm(file, { force: true }) this.incidentRuns.delete(run) + } else if (await this.replaceBounded(file, this.incidentFileBody(run, rows), INCIDENT_BUDGET)) { + if (rows.length) this.incidentRuns.set(run, rows) + else this.incidentRuns.delete(run) } } @@ -691,14 +708,27 @@ export class MonitorHistoryStore { } catch { return null } } - private async readIncidentFile(file: string): Promise { + private incidentFileBody(run: string, rows: MonitorIncident[]): string { + return JSON.stringify([...rows, ...(this.foreignIncidents.get(run) ?? [])]) + } + + private async readIncidentFile(file: string, run: string): Promise { try { + // Whole-file refusals stay whole-file: an oversized or non-array file is + // not "one row this build does not know" but a file it cannot trust at + // all, and nothing here rewrites it (the run has no incidentRuns entry). if ((await stat(file)).size > INCIDENT_BUDGET) { this.degraded = true; return [] } const value: unknown = JSON.parse(await readFile(file, 'utf8')) if (!Array.isArray(value) || value.length > INCIDENT_LIMIT) { this.degraded = true; return [] } - const parsed = value.map(parseMonitorIncident) - if (parsed.some(incident => incident === null)) { this.degraded = true; return [] } - return parsed as MonitorIncident[] + const rows: MonitorIncident[] = [] + const foreign: unknown[] = [] + for (const row of value) { + const incident = parseMonitorIncident(row) + if (incident) rows.push(incident) + else foreign.push(row) + } + if (foreign.length) { this.degraded = true; this.foreignIncidents.set(run, foreign) } + return rows } catch (error) { if ((error as NodeJS.ErrnoException).code !== 'ENOENT') this.degraded = true return [] @@ -753,6 +783,7 @@ export class MonitorHistoryStore { await rm(join(this.root, RUNS_DIR, run), { recursive: true, force: true }) for (const [file, entry] of [...this.index]) if (entry.run === run) { this.index.delete(file); this.repairedTails.delete(file) } this.incidentRuns.delete(run) + this.foreignIncidents.delete(run) this.unindexedRuns.delete(run) total = Math.max(0, total - size); this.shortened = true } From 77384ab301517a56684a44e79b82456f4811d1fb Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sun, 27 Sep 2026 04:27:53 -0700 Subject: [PATCH 32/48] fix(workflows): one invalid source approval no longer blocks every workflow (#1251 row 11) Co-Authored-By: Claude Opus 5.5 --- .../WorkflowSourceApprovalStore.test.ts | 32 ++++++++++++++++++- .../workflows/WorkflowSourceApprovalStore.ts | 27 ++++++++++++---- 2 files changed, 52 insertions(+), 7 deletions(-) diff --git a/src/main/workflows/WorkflowSourceApprovalStore.test.ts b/src/main/workflows/WorkflowSourceApprovalStore.test.ts index ab847797d..986d99daa 100644 --- a/src/main/workflows/WorkflowSourceApprovalStore.test.ts +++ b/src/main/workflows/WorkflowSourceApprovalStore.test.ts @@ -1,4 +1,4 @@ -import { mkdtemp } from 'node:fs/promises' +import { mkdtemp, readFile, writeFile } from 'node:fs/promises' import { tmpdir } from 'node:os' import { join } from 'node:path' @@ -36,4 +36,34 @@ describe('WorkflowSourceApprovalStore', () => { .resolves.toBe(false) expect(shouldNotPrompt).toHaveBeenCalledOnce() }) + + it('honours valid approvals beside an unreadable entry and never approves or drops that entry (#1251 row 11)', async () => { + // One bad entry used to throw from load() on every authorize(), so every + // repository workflow failed until the user hand-edited the file. + const root = await mkdtemp(join(tmpdir(), 'workflow-source-approval-')) + const filePath = join(root, 'approvals.json') + const unreadable = { canonicalIdentity: '/repo/.claude/workflows/other.js', sourceHash: 'not-a-sha', approvedAt: '2026-09-01T00:00:00.000Z' } + await writeFile(filePath, JSON.stringify({ + version: 1, + approvals: [ + { canonicalIdentity: request.canonicalIdentity, sourceHash: request.sourceHash, approvedAt: '2026-09-01T00:00:00.000Z' }, + unreadable, + ], + })) + const store = new WorkflowSourceApprovalStore(filePath) + const shouldNotPrompt = vi.fn(async () => false) + await expect(store.authorize(request, shouldNotPrompt)).resolves.toBe(true) + expect(shouldNotPrompt).not.toHaveBeenCalled() + + // The unreadable entry grants nothing: its source still needs a prompt. + const approve = vi.fn(async () => true) + const other = { ...request, canonicalIdentity: unreadable.canonicalIdentity, sourceHash: 'c'.repeat(64) } + await expect(store.authorize(other, approve)).resolves.toBe(true) + expect(approve).toHaveBeenCalledOnce() + + // Persisting the new grant keeps the entry this build could not read. + const written = JSON.parse(await readFile(filePath, 'utf8')) as { approvals: unknown[] } + expect(written.approvals).toContainEqual(unreadable) + expect(written.approvals).toHaveLength(3) + }) }) diff --git a/src/main/workflows/WorkflowSourceApprovalStore.ts b/src/main/workflows/WorkflowSourceApprovalStore.ts index e947baff3..93c7d097b 100644 --- a/src/main/workflows/WorkflowSourceApprovalStore.ts +++ b/src/main/workflows/WorkflowSourceApprovalStore.ts @@ -12,7 +12,7 @@ type StoredApproval = { type StoredApprovalFile = { version: 1 - approvals: StoredApproval[] + approvals: unknown[] } /** @@ -25,6 +25,17 @@ type StoredApprovalFile = { export class WorkflowSourceApprovalStore { private readonly filePath: string private readonly approvals = new Map() + // Entries this build could not read, carried verbatim (#1251 row 11). + // WHY skip-and-carry instead of the old throw: load() threw on the first bad + // entry and never set `loaded`, so every authorize() rethrew and one bad row + // blocked every repository workflow until the file was hand-edited. Skipping + // still fails closed where it matters: an unreadable entry never enters + // `approvals`, so its source is prompted for again exactly like a new one. + // WHY carry rather than drop: persist() rewrites the whole file on the next + // grant, and dropping would silently delete a record a newer build may + // understand. The whole-file version check below still throws, because a + // file of an unknown version is not one bad row. + private readonly unreadable: unknown[] = [] private readonly prompts = new Map>() private loaded = false @@ -83,7 +94,8 @@ export class WorkflowSourceApprovalStore { !/^[a-f0-9]{64}$/.test(value.sourceHash) || typeof value.approvedAt !== 'string' ) { - throw new Error(`Workflow source approval entry is invalid: ${this.filePath}`) + this.unreadable.push(value) + continue } const approval = value as StoredApproval this.approvals.set(approvalKey(approval.canonicalIdentity, approval.sourceHash), approval) @@ -97,10 +109,13 @@ export class WorkflowSourceApprovalStore { const temporary = `${this.filePath}.tmp-${process.pid}-${randomUUID()}` const document: StoredApprovalFile = { version: 1, - approvals: [...this.approvals.values()].sort((left, right) => ( - left.canonicalIdentity.localeCompare(right.canonicalIdentity) || - left.sourceHash.localeCompare(right.sourceHash) - )), + approvals: [ + ...[...this.approvals.values()].sort((left, right) => ( + left.canonicalIdentity.localeCompare(right.canonicalIdentity) || + left.sourceHash.localeCompare(right.sourceHash) + )), + ...this.unreadable, + ], } try { await writeFile(temporary, `${JSON.stringify(document)}\n`, { encoding: 'utf8', mode: 0o600 }) From bb4c15741771bec1ac3407bf106e06e1bc909a7a Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sun, 27 Sep 2026 04:27:53 -0700 Subject: [PATCH 33/48] fix(tldr): drop invalid identities from a read batch instead of failing it (#1251 row 12) Co-Authored-By: Claude Opus 5.5 --- src/main/tldr/ipc.test.ts | 71 +++++++++++++++++++++++++++++++++++++++ src/main/tldr/ipc.ts | 12 ++++++- 2 files changed, 82 insertions(+), 1 deletion(-) create mode 100644 src/main/tldr/ipc.test.ts diff --git a/src/main/tldr/ipc.test.ts b/src/main/tldr/ipc.test.ts new file mode 100644 index 000000000..79387c5d8 --- /dev/null +++ b/src/main/tldr/ipc.test.ts @@ -0,0 +1,71 @@ +import { describe, expect, it, vi } from 'vitest' + +import type { TldrStore } from './TldrStore.js' +import type { TldrEnforcement } from './enforcement.js' + +// vi.hoisted, because vi.mock factories are hoisted above ordinary consts. +const { handlers } = vi.hoisted(() => ({ + handlers: new Map unknown>(), +})) + +vi.mock('electron', () => ({ + ipcMain: { + handle: (channel: string, handler: (event: unknown, ...args: unknown[]) => unknown) => { handlers.set(channel, handler) }, + on: () => {}, + }, + systemPreferences: {}, +})) +vi.mock('@main/window/windowRegistry.js', () => ({ + windowIdFor: () => 'window-one', + getBrowserWindow: () => ({ id: 1 }), + broadcastToWindows: () => {}, +})) +vi.mock('@main/dictation/macHotkeyHelper.js', () => ({ ensureMacHotkeyHelperBinary: async () => null })) +vi.mock('./holdRelease.js', () => ({ watchMacTldrRelease: () => () => {} })) + +const { registerGoalIpc, registerTldrIpc } = await import('./ipc.js') + +function fakes() { + const read = vi.fn(async (identities: string[]) => Object.fromEntries(identities.map(id => [id, { text: `tldr of ${id}` }]))) + const status = vi.fn((identities: string[]) => Object.fromEntries(identities.map(id => [id, 'active']))) + const store = { read, on: () => {} } as unknown as TldrStore + const enforcement = { status } as unknown as Pick + return { store, enforcement, read, status } +} + +const sender = { mainFrame: {} } +const event = { sender, senderFrame: sender.mainFrame } + +describe('TLDR read IPC batches (#1251 row 12)', () => { + // One identity outside the stored-identity alphabet used to fail zod's + // array parse, so Agent Activity's single batched read rejected and every + // TLDR (and goal) on screen went blank. A TLDR can only ever be written under + // a valid identity (the MCP writer validates the same predicate), so an + // invalid one has no record to return: dropping it from the batch answers it + // exactly as the store would, "none", without taking the rest down. + it('answers the valid identities of a batch that also holds an invalid one', async () => { + const { store, enforcement, read, status } = fakes() + registerTldrIpc(store, enforcement) + registerGoalIpc(store) + const batch = ['session-1', 'not a/valid identity', 'session-2'] + + await expect(handlers.get('tldr:read')!(event, batch)).resolves.toEqual({ + 'session-1': { text: 'tldr of session-1' }, + 'session-2': { text: 'tldr of session-2' }, + }) + expect(read).toHaveBeenLastCalledWith(['session-1', 'session-2']) + await handlers.get('goal:read')!(event, batch) + expect(read).toHaveBeenLastCalledWith(['session-1', 'session-2']) + await handlers.get('tldr:enforcement')!(event, batch) + expect(status).toHaveBeenLastCalledWith(['session-1', 'session-2']) + }) + + it('still refuses a payload that is not a bounded list of strings', async () => { + const { store, enforcement } = fakes() + registerTldrIpc(store, enforcement) + expect(() => handlers.get('tldr:read')!(event, 'session-1')).toThrow() + expect(() => handlers.get('tldr:read')!(event, [42])).toThrow() + expect(() => handlers.get('tldr:read')!(event, Array.from({ length: 10_001 }, (_, i) => `s${i}`))).toThrow() + expect(() => handlers.get('tldr:read')!(event, ['x'.repeat(100_000)])).toThrow() + }) +}) diff --git a/src/main/tldr/ipc.ts b/src/main/tldr/ipc.ts index f2f9c4ed4..c2da8f4d6 100644 --- a/src/main/tldr/ipc.ts +++ b/src/main/tldr/ipc.ts @@ -16,7 +16,17 @@ function assertApplicationWindow(event: Electron.IpcMainInvokeEvent): void { } } -const identityList = z.array(z.string().refine(validTldrIdentity)).max(10_000) +// WHY invalid identities are dropped from a batch instead of failing it +// (#1251 row 12): Agent Activity reads every visible agent's TLDR and goal in +// ONE batch, and a single identity outside the alphabet used to reject the +// whole parse and blank every row. Dropping is exact, not lenient: TldrStore +// only ever writes under identities that pass validTldrIdentity, so an invalid +// one has no record, and "absent from the result" is the answer the store +// would give it anyway. The shape stays strict (a bounded array of bounded +// strings), so a malformed payload is still refused outright. The 256-char +// element cap only bounds what zod copies; the predicate's own limit is 128. +const identityList = z.array(z.string().max(256)).max(10_000) + .transform(identities => identities.filter(validTldrIdentity)) const singleIdentity = z.string().refine(validTldrIdentity) /** From 129f8c5a85b270a7286070e6d29c61f3302f43ab Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sun, 27 Sep 2026 04:27:53 -0700 Subject: [PATCH 34/48] fix(storage): one malformed legacy ledger row no longer stops debug pruning (#1251 row 13) Co-Authored-By: Claude Opus 5.5 --- src/main/storage/debugRetention.test.ts | 32 ++++++++++++++++++++++++- src/main/storage/debugRetention.ts | 23 +++++++++++++----- 2 files changed, 48 insertions(+), 7 deletions(-) diff --git a/src/main/storage/debugRetention.test.ts b/src/main/storage/debugRetention.test.ts index e4c0683f7..4dcca5395 100644 --- a/src/main/storage/debugRetention.test.ts +++ b/src/main/storage/debugRetention.test.ts @@ -4,7 +4,7 @@ import { join } from 'node:path' import { afterEach, beforeEach, describe, expect, it } from 'vitest' -import { collectSessionRecordingDirs, runPrunePasses } from './debugRetention.js' +import { collectSessionRecordingDirs, parseManualLegacyBundlePaths, runPrunePasses } from './debugRetention.js' import type { DebugStorageArtifact, DebugStorageBucket, @@ -226,3 +226,33 @@ describe('runPrunePasses', () => { expect(result).toEqual({ removed: 0, bytesFreed: 0, remainingBytes: 500 }) }) }) + +describe('parseManualLegacyBundlePaths (#1251 row 13)', () => { + // The legacy ledger is append-only JSONL written across many app versions. + // A row that parses as JSON but is not a saved-entry object (a bare `null`, + // a number, an entry without a string bundlePath) used to throw out of the + // loop (`null.event`, `resolve(undefined)`), which rejected collectArtifacts + // and so stopped EVERY prune pass, for every bucket, on every trigger. + it('keeps every readable manual row and skips rows that are not saved-entry objects', () => { + const raw = [ + JSON.stringify({ event: 'saved', reason: 'manual', bundlePath: '/bundles/2026-01-01T00-00-00' }), + 'null', + '42', + '"saved"', + JSON.stringify({ event: 'saved', reason: 'manual' }), + JSON.stringify({ event: 'saved', reason: 'manual', bundlePath: 42 }), + JSON.stringify({ event: 'saved', reason: 7, bundlePath: '/bundles/2026-01-03T00-00-00' }), + '{not json', + JSON.stringify({ event: 'saved', reason: 'autosave-crash', bundlePath: '/bundles/2026-01-02T00-00-00' }), + JSON.stringify({ event: 'saved', reason: 'manual', bundlePath: '/bundles/2026-01-04T00-00-00' }), + ].join('\n') + expect([...parseManualLegacyBundlePaths(raw)]).toEqual([ + '/bundles/2026-01-01T00-00-00', + // A non-string reason is not an autosave label, and an unlabelled save + // was user-triggered in the versions that wrote this ledger, so it stays + // protected: when in doubt, retention keeps the bundle. + '/bundles/2026-01-03T00-00-00', + '/bundles/2026-01-04T00-00-00', + ]) + }) +}) diff --git a/src/main/storage/debugRetention.ts b/src/main/storage/debugRetention.ts index 8bdaace51..6c4065811 100644 --- a/src/main/storage/debugRetention.ts +++ b/src/main/storage/debugRetention.ts @@ -625,25 +625,36 @@ function isProtectedFromDebugPrune(artifact: Artifact): boolean { } async function loadManualLegacyBundlePaths(): Promise> { - const manual = new Set() let raw: string try { raw = await readFile(DEBUG_BUNDLE_LOG_FILE, 'utf8') } catch { - return manual + return new Set() } + return parseManualLegacyBundlePaths(raw) +} +export function parseManualLegacyBundlePaths(raw: string): Set { + const manual = new Set() for (const line of raw.split('\n')) { const trimmed = line.trim() if (!trimmed) continue - let entry: DebugBundleLogEntry + let parsed: unknown try { - entry = JSON.parse(trimmed) as DebugBundleLogEntry + parsed = JSON.parse(trimmed) } catch { continue } - if (entry.event !== 'saved') continue - if (isAutosaveDebugBundleReason(entry.reason)) continue + // WHY a shape check and not only the JSON.parse guard (#1251 row 13): a + // line can be valid JSON and still not an entry (`null`, a number, a row + // from a build that wrote bundlePath differently). Such a row threw here, + // which rejected collectArtifacts and stopped every prune pass for every + // bucket. Skipping it can only fail to protect a bundle the row does not + // name, so it never exposes a manual bundle to deletion. + if (typeof parsed !== 'object' || parsed === null) continue + const entry = parsed as Partial & { bundlePath?: unknown; reason?: unknown } + if (entry.event !== 'saved' || typeof entry.bundlePath !== 'string') continue + if (isAutosaveDebugBundleReason(typeof entry.reason === 'string' ? entry.reason : null)) continue // WHY manual legacy classification comes from the old mixed ledger instead // of folder contents: every bundle contains a manifest, but reading // thousands of manifests during retention would turn a cheap directory From 56d60268ca6755f60060abc6a58b2dd759510da8 Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sun, 27 Sep 2026 05:19:32 -0700 Subject: [PATCH 35/48] revert(storage): row 13 moves to #1417 (manager q109) Row 13 of #1251 and steering q109 on #1417 both change loadManualLegacyBundlePaths, so the row ships with #1417 (commit 02df3652 there), with a test through the real loader. This PR no longer touches debugRetention. Co-Authored-By: Claude Opus 5.5 --- src/main/storage/debugRetention.test.ts | 32 +------------------------ src/main/storage/debugRetention.ts | 23 +++++------------- 2 files changed, 7 insertions(+), 48 deletions(-) diff --git a/src/main/storage/debugRetention.test.ts b/src/main/storage/debugRetention.test.ts index 4dcca5395..e4c0683f7 100644 --- a/src/main/storage/debugRetention.test.ts +++ b/src/main/storage/debugRetention.test.ts @@ -4,7 +4,7 @@ import { join } from 'node:path' import { afterEach, beforeEach, describe, expect, it } from 'vitest' -import { collectSessionRecordingDirs, parseManualLegacyBundlePaths, runPrunePasses } from './debugRetention.js' +import { collectSessionRecordingDirs, runPrunePasses } from './debugRetention.js' import type { DebugStorageArtifact, DebugStorageBucket, @@ -226,33 +226,3 @@ describe('runPrunePasses', () => { expect(result).toEqual({ removed: 0, bytesFreed: 0, remainingBytes: 500 }) }) }) - -describe('parseManualLegacyBundlePaths (#1251 row 13)', () => { - // The legacy ledger is append-only JSONL written across many app versions. - // A row that parses as JSON but is not a saved-entry object (a bare `null`, - // a number, an entry without a string bundlePath) used to throw out of the - // loop (`null.event`, `resolve(undefined)`), which rejected collectArtifacts - // and so stopped EVERY prune pass, for every bucket, on every trigger. - it('keeps every readable manual row and skips rows that are not saved-entry objects', () => { - const raw = [ - JSON.stringify({ event: 'saved', reason: 'manual', bundlePath: '/bundles/2026-01-01T00-00-00' }), - 'null', - '42', - '"saved"', - JSON.stringify({ event: 'saved', reason: 'manual' }), - JSON.stringify({ event: 'saved', reason: 'manual', bundlePath: 42 }), - JSON.stringify({ event: 'saved', reason: 7, bundlePath: '/bundles/2026-01-03T00-00-00' }), - '{not json', - JSON.stringify({ event: 'saved', reason: 'autosave-crash', bundlePath: '/bundles/2026-01-02T00-00-00' }), - JSON.stringify({ event: 'saved', reason: 'manual', bundlePath: '/bundles/2026-01-04T00-00-00' }), - ].join('\n') - expect([...parseManualLegacyBundlePaths(raw)]).toEqual([ - '/bundles/2026-01-01T00-00-00', - // A non-string reason is not an autosave label, and an unlabelled save - // was user-triggered in the versions that wrote this ledger, so it stays - // protected: when in doubt, retention keeps the bundle. - '/bundles/2026-01-03T00-00-00', - '/bundles/2026-01-04T00-00-00', - ]) - }) -}) diff --git a/src/main/storage/debugRetention.ts b/src/main/storage/debugRetention.ts index 6c4065811..8bdaace51 100644 --- a/src/main/storage/debugRetention.ts +++ b/src/main/storage/debugRetention.ts @@ -625,36 +625,25 @@ function isProtectedFromDebugPrune(artifact: Artifact): boolean { } async function loadManualLegacyBundlePaths(): Promise> { + const manual = new Set() let raw: string try { raw = await readFile(DEBUG_BUNDLE_LOG_FILE, 'utf8') } catch { - return new Set() + return manual } - return parseManualLegacyBundlePaths(raw) -} -export function parseManualLegacyBundlePaths(raw: string): Set { - const manual = new Set() for (const line of raw.split('\n')) { const trimmed = line.trim() if (!trimmed) continue - let parsed: unknown + let entry: DebugBundleLogEntry try { - parsed = JSON.parse(trimmed) + entry = JSON.parse(trimmed) as DebugBundleLogEntry } catch { continue } - // WHY a shape check and not only the JSON.parse guard (#1251 row 13): a - // line can be valid JSON and still not an entry (`null`, a number, a row - // from a build that wrote bundlePath differently). Such a row threw here, - // which rejected collectArtifacts and stopped every prune pass for every - // bucket. Skipping it can only fail to protect a bundle the row does not - // name, so it never exposes a manual bundle to deletion. - if (typeof parsed !== 'object' || parsed === null) continue - const entry = parsed as Partial & { bundlePath?: unknown; reason?: unknown } - if (entry.event !== 'saved' || typeof entry.bundlePath !== 'string') continue - if (isAutosaveDebugBundleReason(typeof entry.reason === 'string' ? entry.reason : null)) continue + if (entry.event !== 'saved') continue + if (isAutosaveDebugBundleReason(entry.reason)) continue // WHY manual legacy classification comes from the old mixed ledger instead // of folder contents: every bundle contains a manifest, but reading // thousands of manifests during retention would turn a cheap directory From 4cd7c706773054bf5cbd49b68441bc2df5797612 Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sun, 27 Sep 2026 05:24:58 -0700 Subject: [PATCH 36/48] fix(performance): a wholly refused incident file is set aside, never overwritten or expired (#1411 review) Co-Authored-By: Claude Opus 5.5 --- .../performance/MonitorHistoryStore.test.ts | 72 ++++++++++++++++++- src/main/performance/MonitorHistoryStore.ts | 57 ++++++++++++--- 2 files changed, 118 insertions(+), 11 deletions(-) diff --git a/src/main/performance/MonitorHistoryStore.test.ts b/src/main/performance/MonitorHistoryStore.test.ts index 80eb5df73..8bdd96081 100644 --- a/src/main/performance/MonitorHistoryStore.test.ts +++ b/src/main/performance/MonitorHistoryStore.test.ts @@ -1,4 +1,4 @@ -import { appendFile, mkdir, mkdtemp, readFile, rm, writeFile } from 'node:fs/promises' +import { appendFile, mkdir, mkdtemp, readFile, readdir, rm, stat, writeFile } from 'node:fs/promises' import { tmpdir } from 'node:os' import { join } from 'node:path' import { afterEach, describe, expect, it } from 'vitest' @@ -133,6 +133,76 @@ describe('bounded local performance history', () => { expect(onDisk.find(row => row.id === 9)).toEqual(foreign) }) + // Review of #1411 (a, b, c): a file refused WHOLE (not an array, a newer + // format, over the row limit, oversized) kept no marker, so a restarted + // helper's next incident replaced it with only the new row, and it was lost. + for (const [label, body] of [ + ['a newer-format object', JSON.stringify({ version: 2, incidents: [incident] })], + ['51 valid rows, one over the limit', JSON.stringify(Array.from({ length: 51 }, (_, index) => ({ ...incident, id: index + 1, at: 1000 + index })))], + ['malformed JSON', '[{"id":1,'], + ] as const) { + it(`sets a wholly refused current-run incident file aside instead of overwriting it: ${label}`, async () => { + const root = await mkdtemp(join(tmpdir(), 'agent-code-monitor-')) + roots.push(root) + const runDir = join(root, 'runs', 'run-refused') + await mkdir(runDir, { recursive: true }) + await writeFile(join(runDir, 'incidents.json'), body) + + const restarted = new MonitorHistoryStore(root, 'run-refused') + await restarted.settled() + restarted.record(snapshot(11_000), null, [{ ...incident, id: 99, at: 11_000 }], 0, 1) + await restarted.settled() + + expect(await restarted.readIncident(11_000, 99)).toMatchObject({ rule: 'renderer-stall' }) + const aside = (await readdir(runDir)).filter(name => name.startsWith('incidents.refused-')) + expect(aside).toHaveLength(1) + expect(await readFile(join(runDir, aside[0]!), 'utf8')).toBe(body) + expect(restarted.status().state).toBe('degraded') + }) + } + + // Review of #1411 (a, b, c), a surviving mutation: maintenance deletes a + // prior run it believes empty. A run whose incident file holds only rows + // this build cannot read, or that it refused whole, is not empty. + it('never expires a prior run whose incidents it could not read', async () => { + const root = await mkdtemp(join(tmpdir(), 'agent-code-monitor-')) + roots.push(root) + const foreignOnly = join(root, 'runs', 'run-foreign-only') + const refused = join(root, 'runs', 'run-refused-whole') + await mkdir(foreignOnly, { recursive: true }) + await mkdir(refused, { recursive: true }) + await writeFile(join(foreignOnly, 'incidents.json'), JSON.stringify([{ ...incident, rule: 'rule-from-a-newer-build' }])) + await writeFile(join(refused, 'incidents.json'), JSON.stringify({ version: 2 })) + + const store = new MonitorHistoryStore(root, 'run-now') + await store.settled() + store.record(snapshot(90_000), null, [], 0, 0) + await store.settled() + + await expect(stat(foreignOnly)).resolves.toBeTruthy() + await expect(stat(refused)).resolves.toBeTruthy() + }) + + // Review of #1411 (a): with the run's file full of carried rows, a new + // readable incident has no room. Keeping the carried rows is right; hiding + // that the new one was not kept is not. + it('reports coverage as shortened when carried rows leave no room for a new incident', async () => { + const root = await mkdtemp(join(tmpdir(), 'agent-code-monitor-')) + roots.push(root) + const runDir = join(root, 'runs', 'run-full') + await mkdir(runDir, { recursive: true }) + await writeFile(join(runDir, 'incidents.json'), JSON.stringify(Array.from({ length: 50 }, (_, index) => ({ ...incident, id: index + 1, rule: 'rule-from-a-newer-build' })))) + + const store = new MonitorHistoryStore(root, 'run-full') + await store.settled() + store.record(snapshot(11_000), null, [{ ...incident, id: 99, at: 11_000 }], 0, 1) + await store.settled() + + expect(store.status().shortened).toBe(true) + const onDisk = JSON.parse(await readFile(join(runDir, 'incidents.json'), 'utf8')) as unknown[] + expect(onDisk).toHaveLength(50) + }) + it('repairs a torn append and keeps coarse tiers peak-preserving', async () => { const root = await mkdtemp(join(tmpdir(), 'agent-code-monitor-')) roots.push(root) diff --git a/src/main/performance/MonitorHistoryStore.ts b/src/main/performance/MonitorHistoryStore.ts index 4bead304d..2687df7bb 100644 --- a/src/main/performance/MonitorHistoryStore.ts +++ b/src/main/performance/MonitorHistoryStore.ts @@ -1,5 +1,5 @@ import { createReadStream, createWriteStream } from 'node:fs' -import { appendFile, mkdir, open, readdir, readFile, rm, stat, writeFile } from 'node:fs/promises' +import { appendFile, mkdir, open, readdir, readFile, rename, rm, stat, writeFile } from 'node:fs/promises' import { once } from 'node:events' import { finished } from 'node:stream/promises' import { dirname, join } from 'node:path' @@ -77,6 +77,17 @@ export class MonitorHistoryStore { // see rows this build can vouch for; the rows leave disk only with the run // directory itself (budget pruning, retention, clear). private foreignIncidents = new Map() + // Runs whose incidents.json was refused WHOLE (not an array, over the row + // limit, oversized, malformed, unreadable) — review of #1411 (a, b, c). + // Unlike a foreign row, such a file cannot be carried row by row, and the + // first version kept no marker at all: the run then looked like it had no + // incidents, so the current run's next persistIncidents replaced the file + // with only the new rows and maintenance expired a prior run as empty. + // Invariant: a refused file is never written over. The current run sets it + // aside (setRefusedIncidentsAside) before its first write; a prior run is + // kept by maintenance and leaves disk only with its directory (budget + // pruning, clear), like foreign rows. + private refusedIncidentRuns = new Set() private repairedTails = new Set() // False until startup indexing completes. Retention deletes any run the // index does not know about, so a partial index (EPERM, ENOSPC or an I/O @@ -309,7 +320,7 @@ export class MonitorHistoryStore { try { await rm(join(this.root, RUNS_DIR), { recursive: true, force: true }) await mkdir(this.runDir, { recursive: true }) - this.index.clear(); this.incidentRuns.clear(); this.foreignIncidents.clear(); this.repairedTails.clear(); this.indexed = true + this.index.clear(); this.incidentRuns.clear(); this.foreignIncidents.clear(); this.refusedIncidentRuns.clear(); this.repairedTails.clear(); this.indexed = true this.bytes = 0; this.shortened = false; this.degraded = false this.operationFingerprint = ''; this.unindexedRuns.clear() this.lastMaintenanceAt = -Infinity @@ -422,9 +433,16 @@ export class MonitorHistoryStore { // would exceed INCIDENT_LIMIT and the next launch would reject all of it. const room = Math.max(0, INCIDENT_LIMIT - (this.foreignIncidents.get(this.runId)?.length ?? 0)) const rows = room ? [...merged.values()].sort((a, b) => a.at - b.at).slice(-room) : [] + // Carried rows are kept over new ones (the owner's "do not delete stuff"), + // but a readable incident that found no room is missing evidence, so say + // so (review of #1411, a). Trimming to INCIDENT_LIMIT itself is the normal + // cap and is not a shortfall. + if (rows.length < Math.min(merged.size, INCIDENT_LIMIT)) this.shortened = true + const file = join(this.runDir, 'incidents.json') + if (this.refusedIncidentRuns.has(this.runId) && !(await this.setRefusedIncidentsAside(file))) return // Memory mirrors disk: a capacity-shortened write keeps the previous rows, // which is what status, queries and the next launch will actually find. - if (!(await this.replaceBounded(join(this.runDir, 'incidents.json'), this.incidentFileBody(this.runId, rows), INCIDENT_BUDGET))) return + if (!(await this.replaceBounded(file, this.incidentFileBody(this.runId, rows), INCIDENT_BUDGET))) return this.incidentRuns.set(this.runId, rows) await this.enforceIncidentLimit() } @@ -561,7 +579,7 @@ export class MonitorHistoryStore { // with no remaining points or incidents holds only an unattributable // operations snapshot, so it is retention-expired, not capacity-pruned. if (this.indexed) for (const run of await this.runNames()) { - if (run === this.runId || this.incidentRuns.has(run) || this.foreignIncidents.has(run) || this.unindexedRuns.has(run) || [...this.index.values()].some(entry => entry.run === run)) continue + if (run === this.runId || this.incidentRuns.has(run) || this.foreignIncidents.has(run) || this.refusedIncidentRuns.has(run) || this.unindexedRuns.has(run) || [...this.index.values()].some(entry => entry.run === run)) continue await rm(join(this.root, RUNS_DIR, run), { recursive: true, force: true }) } this.bytes = await this.diskBytes() @@ -712,14 +730,32 @@ export class MonitorHistoryStore { return JSON.stringify([...rows, ...(this.foreignIncidents.get(run) ?? [])]) } + /** + * Moves a refused incidents.json to `incidents.refused-.json` in the same + * run directory, keeping its bytes (and counting them in the budget), so the + * run can record new incidents. False, and nothing is written, when the + * move fails: losing the new rows beats overwriting the refused ones. + */ + private async setRefusedIncidentsAside(file: string): Promise { + try { + await rename(file, join(dirname(file), `incidents.refused-${Date.now()}.json`)) + } catch (error) { + if ((error as NodeJS.ErrnoException).code !== 'ENOENT') { this.degraded = true; this.shortened = true; return false } + } + this.refusedIncidentRuns.delete(this.runId) + return true + } + private async readIncidentFile(file: string, run: string): Promise { + // Whole-file refusals stay whole-file: an oversized, non-array, over-limit, + // malformed or unreadable file is not "one row this build does not know" + // but a file it cannot trust at all. It is marked refused, never rewritten; + // see refusedIncidentRuns. + const refuse = (): MonitorIncident[] => { this.degraded = true; this.refusedIncidentRuns.add(run); return [] } try { - // Whole-file refusals stay whole-file: an oversized or non-array file is - // not "one row this build does not know" but a file it cannot trust at - // all, and nothing here rewrites it (the run has no incidentRuns entry). - if ((await stat(file)).size > INCIDENT_BUDGET) { this.degraded = true; return [] } + if ((await stat(file)).size > INCIDENT_BUDGET) return refuse() const value: unknown = JSON.parse(await readFile(file, 'utf8')) - if (!Array.isArray(value) || value.length > INCIDENT_LIMIT) { this.degraded = true; return [] } + if (!Array.isArray(value) || value.length > INCIDENT_LIMIT) return refuse() const rows: MonitorIncident[] = [] const foreign: unknown[] = [] for (const row of value) { @@ -730,7 +766,7 @@ export class MonitorHistoryStore { if (foreign.length) { this.degraded = true; this.foreignIncidents.set(run, foreign) } return rows } catch (error) { - if ((error as NodeJS.ErrnoException).code !== 'ENOENT') this.degraded = true + if ((error as NodeJS.ErrnoException).code !== 'ENOENT') return refuse() return [] } } @@ -784,6 +820,7 @@ export class MonitorHistoryStore { for (const [file, entry] of [...this.index]) if (entry.run === run) { this.index.delete(file); this.repairedTails.delete(file) } this.incidentRuns.delete(run) this.foreignIncidents.delete(run) + this.refusedIncidentRuns.delete(run) this.unindexedRuns.delete(run) total = Math.max(0, total - size); this.shortened = true } From 0d5de99d67402e59650381f4066fbaabb0a14bdf Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sun, 27 Sep 2026 05:24:58 -0700 Subject: [PATCH 37/48] fix(conversations): type each Codex index row by value; count and report skipped rows and rollouts (#1411 review) Co-Authored-By: Claude Opus 5.5 --- .../sources/codex.system.test.ts | 20 ++++++ src/main/conversations/sources/codex.ts | 63 +++++++++++++++++-- 2 files changed, 79 insertions(+), 4 deletions(-) diff --git a/src/main/conversations/sources/codex.system.test.ts b/src/main/conversations/sources/codex.system.test.ts index abd80057d..6927c8d8e 100644 --- a/src/main/conversations/sources/codex.system.test.ts +++ b/src/main/conversations/sources/codex.system.test.ts @@ -109,6 +109,9 @@ describe('Codex conversation source', () => { const indexed = await source.discover({ scope: 'everywhere', family }) expect(indexed.filter(r => r.origin === 'index').length).toBeGreaterThanOrEqual(counts.codex.inFamily) expect(indexed.filter(r => r.origin === 'scan')).toHaveLength(0) + // Review of #1411 (c): the skip must not look like a complete result. + expect(counts.codex.unindexedSampled).toBeGreaterThan(0) + expect(source.lastDowngradeReason()).toMatch(/skipped \d+ unreadable rollout/) // Fallback path: one readable rollout still lists beside unreadable ones. await chmod(rollouts[0]!, 0o600) @@ -116,6 +119,23 @@ describe('Codex conversation source', () => { const fresh = new CodexConversationSource({ codexHome: corpus.codexHome }) const scanned = await fresh.discover({ scope: 'everywhere', family }) expect(scanned.map(r => r.file)).toEqual([rollouts[0]]) + expect(fresh.lastDowngradeReason()).toMatch(/no state_N\.sqlite.*; skipped \d+ unreadable rollout/) + }) + + // Review of #1411 (b): SQLite keeps any value in any column, so one thread + // whose title is a BLOB made `.trim()` throw and rejected the whole index. + it('lists every indexed thread when one row holds a value of the wrong type (#1251 row 8)', async () => { + const { corpus, source, listWorktrees } = await setup() + const family = await resolveFamily('/fixture/repo', 'everywhere', { listWorktrees }) + const before = await source.discover({ scope: 'everywhere', family }) + const db = new DatabaseSync(join(corpus.codexHome, 'state_5.sqlite')) + const victim = (db.prepare('select id from threads where archived = 0 limit 1').get() as { id: string }).id + db.prepare("update threads set title = x'00', first_user_message = 'fallback label' where id = ?").run(victim) + db.close() + + const after = await new CodexConversationSource({ codexHome: corpus.codexHome }).discover({ scope: 'everywhere', family }) + expect(after).toHaveLength(before.length) + expect(after.find(r => r.nativeId === victim)?.userTexts).toEqual(['fallback label']) }) it('falls back to the rollout scan when the index is missing and reports why', async () => { diff --git a/src/main/conversations/sources/codex.ts b/src/main/conversations/sources/codex.ts index 5735586fb..14a490a3f 100644 --- a/src/main/conversations/sources/codex.ts +++ b/src/main/conversations/sources/codex.ts @@ -72,6 +72,39 @@ type RolloutHead = { lastUserAt: number | null } +/** + * One `threads` row, typed by value rather than trusted by column (review of + * #1411, b). SQLite stores any value in any column whatever its declared type, + * so a BLOB title made `(row.title ?? '').trim()` throw, and that one row + * rejected the whole Codex discovery. A field of the wrong type becomes its + * empty value (a title then falls back to the next label); only a row with no + * string id is dropped, because nothing can address it. + */ +function normalizeIndexRow(raw: Record): IndexRow | null { + const text = (value: unknown): string | null => typeof value === 'string' ? value : null + const num = (value: unknown): number | null => typeof value === 'number' && Number.isFinite(value) ? value : null + const id = text(raw.id) + if (!id) return null + return { + id, + rollout_path: text(raw.rollout_path) ?? '', + cwd: text(raw.cwd) ?? '', + title: text(raw.title), + first_user_message: text(raw.first_user_message), + preview: text(raw.preview), + name: text(raw.name), + source: text(raw.source) ?? '', + thread_source: text(raw.thread_source), + agent_role: text(raw.agent_role), + git_branch: text(raw.git_branch), + created_at_ms: num(raw.created_at_ms), + updated_at_ms: num(raw.updated_at_ms), + recency_at_ms: num(raw.recency_at_ms), + archived: num(raw.archived) ?? 0, + originator: text(raw.originator), + } +} + function isSubagentSource(row: IndexRow): boolean { if (row.thread_source === 'subagent') return true if (row.agent_role) return true @@ -118,6 +151,12 @@ async function readRolloutHead(file: string): Promise { export class CodexConversationSource implements ConversationSource { readonly provider = 'codex' as const private downgradeReason: string | null = null + // What one discovery skipped (review of #1411, c): a skipped rollout or index + // row used to leave the result looking complete. The counts reach the + // discovery span, lastDowngradeReason and one console warning (counts only, + // never paths). The picker itself has no degraded indicator for any source + // yet, including the existing no-index downgrade; that is a residual. + private skipped = { rollouts: 0, indexRows: 0 } private walk: { at: number; files: Map } | null = null private readonly heads = new Map() // Rollout paths learnt at discovery, so a search that reads prompts for a @@ -183,6 +222,7 @@ export class CodexConversationSource implements ConversationSource { try { head = await readRolloutHead(file) } catch { + this.skipped.rollouts++ return null } } @@ -211,16 +251,25 @@ export class CodexConversationSource implements ConversationSource { } } + private withSkipped(reason: string | null): string | null { + const { rollouts, indexRows } = this.skipped + if (!rollouts && !indexRows) return reason + const note = `skipped ${rollouts} unreadable rollout(s) and ${indexRows} malformed index row(s)` + console.warn(`[conversations.codex] ${note}`) + return reason ? `${reason}; ${note}` : note + } + async discover(scope: SourceScope): Promise { const span = performanceService.span('conversations.codex.discover', { scope: scope.scope }) + this.skipped = { rollouts: 0, indexRows: 0 } const dbPath = newestCodexStateDb(this.deps.codexHome) const opened = dbPath ? openReadOnlySqlite(dbPath, CODEX_INDEX_COLUMNS) : { ok: false as const, reason: `no state_N.sqlite under ${this.deps.codexHome}` } if (!opened.ok) { - this.downgradeReason = opened.reason const rows = await this.scanEverything(scope) - span.end({ mode: 'scan', rows: rows.length }) + this.downgradeReason = this.withSkipped(opened.reason) + span.end({ mode: 'scan', rows: rows.length, ...this.skipped }) return rows } this.downgradeReason = null @@ -263,7 +312,12 @@ export class CodexConversationSource implements ConversationSource { } const where = predicates.length > 0 ? `where archived = 0 and (${predicates.join(' or ')})` : 'where archived = 0' const columns = CODEX_INDEX_COLUMNS.threads.map(c => `"${c}"`).join(', ') - for (const row of opened.db.prepare(`select ${columns} from threads ${where}`).all(...args) as unknown as IndexRow[]) { + for (const raw of opened.db.prepare(`select ${columns} from threads ${where}`).all(...args) as Array>) { + const row = normalizeIndexRow(raw) + if (!row) { + this.skipped.indexRows++ + continue + } this.rolloutPaths.set(row.id, row.rollout_path) const title = (row.title ?? '').trim() || (row.first_user_message ?? '').trim() || (row.preview ?? '').trim() const name = (row.name ?? '').trim() @@ -302,7 +356,8 @@ export class CodexConversationSource implements ConversationSource { const row = await this.fromHead(file, meta.mtime, meta.id, scope) if (row) rows.push(row) } - span.end({ mode: 'index', rows: rows.length, unindexed: rows.filter(r => r.origin === 'scan').length }) + this.downgradeReason = this.withSkipped(null) + span.end({ mode: 'index', rows: rows.length, unindexed: rows.filter(r => r.origin === 'scan').length, ...this.skipped }) return rows } From 6077c0a059f0f38fc2adf7da3ea4fd67384addcc Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sun, 27 Sep 2026 05:24:58 -0700 Subject: [PATCH 38/48] fix(tldr): history answers an invalid identity with an empty list (#1411 review) Co-Authored-By: Claude Opus 5.5 --- src/main/tldr/ipc.test.ts | 17 +++++++++++++++-- src/main/tldr/ipc.ts | 15 +++++++++++---- 2 files changed, 26 insertions(+), 6 deletions(-) diff --git a/src/main/tldr/ipc.test.ts b/src/main/tldr/ipc.test.ts index 79387c5d8..ab8f685e5 100644 --- a/src/main/tldr/ipc.test.ts +++ b/src/main/tldr/ipc.test.ts @@ -28,9 +28,10 @@ const { registerGoalIpc, registerTldrIpc } = await import('./ipc.js') function fakes() { const read = vi.fn(async (identities: string[]) => Object.fromEntries(identities.map(id => [id, { text: `tldr of ${id}` }]))) const status = vi.fn((identities: string[]) => Object.fromEntries(identities.map(id => [id, 'active']))) - const store = { read, on: () => {} } as unknown as TldrStore + const history = vi.fn(async (identity: string) => [{ identity }]) + const store = { read, history, on: () => {} } as unknown as TldrStore const enforcement = { status } as unknown as Pick - return { store, enforcement, read, status } + return { store, enforcement, read, status, history } } const sender = { mainFrame: {} } @@ -68,4 +69,16 @@ describe('TLDR read IPC batches (#1251 row 12)', () => { expect(() => handlers.get('tldr:read')!(event, Array.from({ length: 10_001 }, (_, i) => `s${i}`))).toThrow() expect(() => handlers.get('tldr:read')!(event, ['x'.repeat(100_000)])).toThrow() }) + + it('answers history for an invalid identity with an empty list, as the batch reads do', async () => { + const { store, enforcement, history } = fakes() + registerTldrIpc(store, enforcement) + registerGoalIpc(store) + for (const channel of ['tldr:history', 'goal:history']) { + await expect(handlers.get(channel)!(event, 'not a/valid identity')).resolves.toEqual([]) + await expect(handlers.get(channel)!(event, 'session-1')).resolves.toEqual([{ identity: 'session-1' }]) + expect(() => handlers.get(channel)!(event, 42)).toThrow() + } + expect(history).toHaveBeenCalledTimes(2) + }) }) diff --git a/src/main/tldr/ipc.ts b/src/main/tldr/ipc.ts index c2da8f4d6..e58fd6c60 100644 --- a/src/main/tldr/ipc.ts +++ b/src/main/tldr/ipc.ts @@ -27,7 +27,15 @@ function assertApplicationWindow(event: Electron.IpcMainInvokeEvent): void { // element cap only bounds what zod copies; the predicate's own limit is 128. const identityList = z.array(z.string().max(256)).max(10_000) .transform(identities => identities.filter(validTldrIdentity)) -const singleIdentity = z.string().refine(validTldrIdentity) +const singleIdentity = z.string().max(256) +// History answers an invalid identity the way the batch reads do (review of +// #1411, b): it cannot have a record, so its history is empty, not an error +// that puts the history modal into its failure state. A non-string or +// oversized payload is still refused by the parse. +const historyFor = (store: TldrStore, raw: unknown) => { + const identity = singleIdentity.parse(raw) + return validTldrIdentity(identity) ? store.history(identity) : Promise.resolve([]) +} /** * Record an unobservable hold ONCE per app run. @@ -151,14 +159,13 @@ export function registerTldrIpc( if (hold && hold.token === token) hold.cancel() }) const identities = identityList - const identity = singleIdentity ipcMain.handle('tldr:read', (event, raw: unknown) => { assertApplicationWindow(event) return store.read(identities.parse(raw)) }) ipcMain.handle('tldr:history', (event, raw: unknown) => { assertApplicationWindow(event) - return store.history(identity.parse(raw)) + return historyFor(store, raw) }) // Read-only, like every renderer TLDR API: whether this identity's provider // hooks have reached main. The renderer uses it to say when enforcement is not @@ -184,7 +191,7 @@ export function registerGoalIpc(store: TldrStore): void { }) ipcMain.handle('goal:history', (event, raw: unknown) => { assertApplicationWindow(event) - return store.history(singleIdentity.parse(raw)) + return historyFor(store, raw) }) store.on('changed', (update: TldrUpdate) => broadcastToWindows('goal:changed', update)) } From 18d7da87552d3f73a40c83821fdef32bb0c91cb1 Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sun, 27 Sep 2026 05:24:58 -0700 Subject: [PATCH 39/48] test(workflows): an approval entry missing approvedAt prompts (#1411 review survivor) Co-Authored-By: Claude Opus 5.5 --- .../workflows/WorkflowSourceApprovalStore.test.ts | 15 +++++++++++++++ 1 file changed, 15 insertions(+) diff --git a/src/main/workflows/WorkflowSourceApprovalStore.test.ts b/src/main/workflows/WorkflowSourceApprovalStore.test.ts index 986d99daa..02f36bcfc 100644 --- a/src/main/workflows/WorkflowSourceApprovalStore.test.ts +++ b/src/main/workflows/WorkflowSourceApprovalStore.test.ts @@ -66,4 +66,19 @@ describe('WorkflowSourceApprovalStore', () => { expect(written.approvals).toContainEqual(unreadable) expect(written.approvals).toHaveLength(3) }) + + // Review of #1411 (a), a surviving mutation: dropping the approvedAt check + // still passed. An entry for the exact source that lacks approvedAt is not a + // grant this build wrote, so it must prompt, never authorize. + it('does not treat an entry missing approvedAt as an approval', async () => { + const root = await mkdtemp(join(tmpdir(), 'workflow-source-approval-')) + const filePath = join(root, 'approvals.json') + await writeFile(filePath, JSON.stringify({ + version: 1, + approvals: [{ canonicalIdentity: request.canonicalIdentity, sourceHash: request.sourceHash }], + })) + const prompt = vi.fn(async () => false) + await expect(new WorkflowSourceApprovalStore(filePath).authorize(request, prompt)).resolves.toBe(false) + expect(prompt).toHaveBeenCalledOnce() + }) }) From 7204503c1c49440357822f9db662251c31856b56 Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sun, 27 Sep 2026 05:24:58 -0700 Subject: [PATCH 40/48] docs(plans): #1411 round-1 review disposition Co-Authored-By: Claude Opus 5.5 --- docs/plans/2026-09-27-c5-fail-all-batch.md | 14 ++++++++++++++ 1 file changed, 14 insertions(+) diff --git a/docs/plans/2026-09-27-c5-fail-all-batch.md b/docs/plans/2026-09-27-c5-fail-all-batch.md index fdf919351..cebbd243b 100644 --- a/docs/plans/2026-09-27-c5-fail-all-batch.md +++ b/docs/plans/2026-09-27-c5-fail-all-batch.md @@ -31,3 +31,17 @@ Rows 14–15 (key vault index, agent-name registry, tmux recovery) are strict by - Carried foreign rows never expire on their own. They leave disk only with their run directory (budget pruning, clear). - A run whose incident file holds only foreign rows is kept from expiry-by-emptiness. It is still pruned by the data budget. - **Row 12:** a renderer that sends an invalid identity gets no record and no error for it. That is the same answer as "no TLDR yet". + +## Review round 1 (a, b, c codex at `129f8c5a`): all FIX-BEFORE-MERGE + +| Finding | Verdict | Change | +|---|---|---| +| **a1 / b1 / c1, major:** a WHOLLY refused incident file (a newer-format object, over the row limit, oversized, malformed JSON, unreadable) kept no marker. The current run's next incident replaced it with only the new row, and maintenance expired a prior run holding one as empty. | valid | The run is marked refused. The current run moves the refused file aside to `incidents.refused-.json` (same run dir, counted in the byte budget) before its first write, and writes nothing if the move fails. Maintenance keeps refused runs. Three fail-first cases (object, 51 rows, malformed JSON). | +| **a and c survivor:** the foreign-only prior-run expiry guard was unpinned. | valid | The prior-run test covers a foreign-only run and a refused run; removing either guard goes red. | +| **b2, major:** one indexed row with a wrong-typed value (a BLOB title) made `.trim()` throw and rejected the whole index. | valid | `normalizeIndexRow` types each field by value. A wrong type becomes the empty value, so the title falls back; only a row with no string id is dropped (and counted). Fail-first with a BLOB title on the recorded corpus: the row still lists under its fallback label. | +| **c2, minor:** a skipped rollout left discovery looking complete. | valid | Skips are counted into the discovery span, `lastDowngradeReason` and one console warning (counts only). **Residual:** the picker has no degraded indicator for any source yet, including the existing no-index downgrade. | +| **b3, minor:** `tldr:history` and `goal:history` threw for an invalid identity that the batch reads skip. | valid | Empty history for an invalid identity; a non-string payload is still refused. Mutant red. | +| **a3, minor:** with the file full of carried rows, a new incident was silently not kept. | valid | Carried rows still win (owner rule), but `shortened` is set. Fail-first. | +| **a survivor:** dropping the `approvedAt` check passed. | valid | An entry missing `approvedAt` prompts. Mutant red. | +| **b survivor:** the row 13 loader returning an empty set passed. | valid | Row 13 moved to #1417 (manager q109: the same loader as steering q109), with a test through the real loader there. This PR no longer touches `debugRetention`. | +| **a2, minor:** carried rows change position on rewrite. | declined | Values are all kept. Neither reader gives order any authority: incidents are sorted by `at`, and approvals are keyed by identity plus hash. Preserving the original interleaving would need positional bookkeeping for no reader. | From c38a67a64391960eb27d5727f7861af612c7d218 Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sun, 27 Sep 2026 07:02:34 -0700 Subject: [PATCH 41/48] fix(performance): a set-aside refused file keeps its run and never collides (#1411 review round 2) Co-Authored-By: Claude Opus 5.5 --- .../performance/MonitorHistoryStore.test.ts | 93 ++++++++++++++++++- src/main/performance/MonitorHistoryStore.ts | 24 ++++- 2 files changed, 111 insertions(+), 6 deletions(-) diff --git a/src/main/performance/MonitorHistoryStore.test.ts b/src/main/performance/MonitorHistoryStore.test.ts index 8bdd96081..d72a67754 100644 --- a/src/main/performance/MonitorHistoryStore.test.ts +++ b/src/main/performance/MonitorHistoryStore.test.ts @@ -1,7 +1,7 @@ -import { appendFile, mkdir, mkdtemp, readFile, readdir, rm, stat, writeFile } from 'node:fs/promises' +import { appendFile, chmod, mkdir, mkdtemp, readFile, readdir, rm, stat, writeFile } from 'node:fs/promises' import { tmpdir } from 'node:os' import { join } from 'node:path' -import { afterEach, describe, expect, it } from 'vitest' +import { afterEach, describe, expect, it, vi } from 'vitest' import type { MonitorWorkerSnapshot } from '@shared/performance/monitorSnapshot.js' import { MonitorHistoryStore } from './MonitorHistoryStore.js' @@ -161,6 +161,95 @@ describe('bounded local performance history', () => { }) } + // Review of #1411, round 2 (a, b, c): once the refused file was set aside, + // nothing marked its run, so a later run's retention expired the run's + // readable incidents, found it empty, and deleted the directory with the + // set-aside bytes in it. + it('keeps a run holding a set-aside refused file after its readable incidents expire', async () => { + const root = await mkdtemp(join(tmpdir(), 'agent-code-monitor-')) + roots.push(root) + const runA = join(root, 'runs', 'run-a') + const refused = JSON.stringify({ version: 2, incidents: [incident] }) + await mkdir(runA, { recursive: true }) + await writeFile(join(runA, 'incidents.json'), refused) + const first = new MonitorHistoryStore(root, 'run-a') + await first.settled() + first.record(snapshot(11_000), null, [{ ...incident, id: 2, at: 11_000 }], 0, 1) + await first.settled() + + const later = new MonitorHistoryStore(root, 'run-b') + await later.settled() + later.record(snapshot(11_000 + 8 * 24 * 60 * 60_000), null, [], 0, 0) + await later.settled() + + const aside = (await readdir(runA)).filter(name => name.startsWith('incidents.refused-')) + expect(aside).toHaveLength(1) + expect(await readFile(join(runA, aside[0]!), 'utf8')).toBe(refused) + }) + + // Review of #1411, round 2 (a, b, c): the set-aside name was only the time, + // and rename replaces an existing file, so a second refusal in the same + // millisecond erased the first. Also pins that recording again after a + // set-aside writes normally and does not set the new file aside (a stale + // refusal marker survived mutation in round 2). + it('keeps every refused file when two refusals set aside in the same millisecond', async () => { + vi.useFakeTimers({ toFake: ['Date'] }) + vi.setSystemTime(123_456) + try { + const root = await mkdtemp(join(tmpdir(), 'agent-code-monitor-')) + roots.push(root) + const runDir = join(root, 'runs', 'run-twice') + await mkdir(runDir, { recursive: true }) + const firstBody = JSON.stringify({ version: 2, generation: 'first' }) + const secondBody = JSON.stringify({ version: 3, generation: 'second' }) + await writeFile(join(runDir, 'incidents.json'), firstBody) + const store = new MonitorHistoryStore(root, 'run-twice') + await store.settled() + store.record(snapshot(11_000), null, [{ ...incident, id: 2, at: 11_000 }], 0, 1) + await store.settled() + store.record(snapshot(12_000), null, [{ ...incident, id: 2, at: 11_000 }, { ...incident, id: 3, at: 12_000 }], 0, 1) + await store.settled() + expect((await readdir(runDir)).filter(name => name.startsWith('incidents.refused-'))).toHaveLength(1) + const current = JSON.parse(await readFile(join(runDir, 'incidents.json'), 'utf8')) as Array<{ id: number }> + expect(current.map(row => row.id)).toEqual([2, 3]) + + await writeFile(join(runDir, 'incidents.json'), secondBody) + const restarted = new MonitorHistoryStore(root, 'run-twice') + await restarted.settled() + restarted.record(snapshot(13_000), null, [{ ...incident, id: 4, at: 13_000 }], 0, 1) + await restarted.settled() + + const aside = (await readdir(runDir)).filter(name => name.startsWith('incidents.refused-')) + const bodies = await Promise.all(aside.map(name => readFile(join(runDir, name), 'utf8'))) + expect(bodies.sort()).toEqual([firstBody, secondBody].sort()) + } finally { + vi.useRealTimers() + } + }) + + // Review of #1411, round 2 (b): the guard that writes nothing when the + // set-aside rename fails was unpinned. Losing the new rows beats + // overwriting the refused ones. + it('writes nothing over a refused file it could not set aside', async () => { + const root = await mkdtemp(join(tmpdir(), 'agent-code-monitor-')) + roots.push(root) + const runDir = join(root, 'runs', 'run-stuck') + await mkdir(runDir, { recursive: true }) + const refused = JSON.stringify({ version: 2 }) + await writeFile(join(runDir, 'incidents.json'), refused) + const store = new MonitorHistoryStore(root, 'run-stuck') + await store.settled() + await chmod(runDir, 0o500) + try { + store.record(snapshot(11_000), null, [{ ...incident, id: 2, at: 11_000 }], 0, 1) + await store.settled() + } finally { + await chmod(runDir, 0o700) + } + expect(await readFile(join(runDir, 'incidents.json'), 'utf8')).toBe(refused) + expect(store.status().shortened).toBe(true) + }) + // Review of #1411 (a, b, c), a surviving mutation: maintenance deletes a // prior run it believes empty. A run whose incident file holds only rows // this build cannot read, or that it refused whole, is not empty. diff --git a/src/main/performance/MonitorHistoryStore.ts b/src/main/performance/MonitorHistoryStore.ts index 2687df7bb..bf0b4fed5 100644 --- a/src/main/performance/MonitorHistoryStore.ts +++ b/src/main/performance/MonitorHistoryStore.ts @@ -2,6 +2,7 @@ import { createReadStream, createWriteStream } from 'node:fs' import { appendFile, mkdir, open, readdir, readFile, rename, rm, stat, writeFile } from 'node:fs/promises' import { once } from 'node:events' import { finished } from 'node:stream/promises' +import { randomUUID } from 'node:crypto' import { dirname, join } from 'node:path' import { createInterface } from 'node:readline' import type { MonitorIncident } from '@shared/performance/monitorIncidents.js' @@ -26,6 +27,7 @@ const INCIDENT_BUDGET = 8 * 1024 * 1024 const OPERATIONS_BUDGET = 1024 * 1024 const REPORT_BUDGET = 8 * 1024 * 1024 const LINE_LIMIT = 16 * 1024 +const REFUSED_INCIDENTS_PREFIX = 'incidents.refused-' const INCIDENT_LIMIT = MONITOR_POLICY.incidentCount // About seven minutes of rolled-up points (67 per minute across tiers). A disk // that stalls longer sheds new points as "shortened" coverage instead of @@ -88,6 +90,13 @@ export class MonitorHistoryStore { // kept by maintenance and leaves disk only with its directory (budget // pruning, clear), like foreign rows. private refusedIncidentRuns = new Set() + // Runs that HOLD a set-aside `incidents.refused-*.json` (review of #1411, + // round 2). Setting a refused file aside clears refusedIncidentRuns, so the + // run then looked empty again once its readable incidents expired, and + // maintenance deleted the directory with the set-aside bytes in it. Found by + // listing each run at startup and added on every set-aside; maintenance + // never expires such a run. Only budget pruning and clear() remove it. + private refusedAsideRuns = new Set() private repairedTails = new Set() // False until startup indexing completes. Retention deletes any run the // index does not know about, so a partial index (EPERM, ENOSPC or an I/O @@ -320,7 +329,7 @@ export class MonitorHistoryStore { try { await rm(join(this.root, RUNS_DIR), { recursive: true, force: true }) await mkdir(this.runDir, { recursive: true }) - this.index.clear(); this.incidentRuns.clear(); this.foreignIncidents.clear(); this.refusedIncidentRuns.clear(); this.repairedTails.clear(); this.indexed = true + this.index.clear(); this.incidentRuns.clear(); this.foreignIncidents.clear(); this.refusedIncidentRuns.clear(); this.refusedAsideRuns.clear(); this.repairedTails.clear(); this.indexed = true this.bytes = 0; this.shortened = false; this.degraded = false this.operationFingerprint = ''; this.unindexedRuns.clear() this.lastMaintenanceAt = -Infinity @@ -376,6 +385,8 @@ export class MonitorHistoryStore { try { await this.cleanupTemps() for (const run of await this.runNames()) { + const entries = await readdir(join(this.root, RUNS_DIR, run)).catch(() => [] as string[]) + if (entries.some(name => name.startsWith(REFUSED_INCIDENTS_PREFIX))) this.refusedAsideRuns.add(run) for (const resolution of TIERS) { const file = join(this.root, RUNS_DIR, run, `${resolution}.jsonl`) try { @@ -579,7 +590,7 @@ export class MonitorHistoryStore { // with no remaining points or incidents holds only an unattributable // operations snapshot, so it is retention-expired, not capacity-pruned. if (this.indexed) for (const run of await this.runNames()) { - if (run === this.runId || this.incidentRuns.has(run) || this.foreignIncidents.has(run) || this.refusedIncidentRuns.has(run) || this.unindexedRuns.has(run) || [...this.index.values()].some(entry => entry.run === run)) continue + if (run === this.runId || this.incidentRuns.has(run) || this.foreignIncidents.has(run) || this.refusedIncidentRuns.has(run) || this.refusedAsideRuns.has(run) || this.unindexedRuns.has(run) || [...this.index.values()].some(entry => entry.run === run)) continue await rm(join(this.root, RUNS_DIR, run), { recursive: true, force: true }) } this.bytes = await this.diskBytes() @@ -731,18 +742,22 @@ export class MonitorHistoryStore { } /** - * Moves a refused incidents.json to `incidents.refused-.json` in the same + * Moves a refused incidents.json to `incidents.refused--.json` in the same * run directory, keeping its bytes (and counting them in the budget), so the * run can record new incidents. False, and nothing is written, when the * move fails: losing the new rows beats overwriting the refused ones. */ private async setRefusedIncidentsAside(file: string): Promise { try { - await rename(file, join(dirname(file), `incidents.refused-${Date.now()}.json`)) + // A UUID, not only the time (review of #1411, round 2): rename replaces an + // existing destination, so a second refusal in the same millisecond, or + // after the clock stepped back, overwrote the first set-aside file. + await rename(file, join(dirname(file), `${REFUSED_INCIDENTS_PREFIX}${Date.now()}-${randomUUID()}.json`)) } catch (error) { if ((error as NodeJS.ErrnoException).code !== 'ENOENT') { this.degraded = true; this.shortened = true; return false } } this.refusedIncidentRuns.delete(this.runId) + this.refusedAsideRuns.add(this.runId) return true } @@ -821,6 +836,7 @@ export class MonitorHistoryStore { this.incidentRuns.delete(run) this.foreignIncidents.delete(run) this.refusedIncidentRuns.delete(run) + this.refusedAsideRuns.delete(run) this.unindexedRuns.delete(run) total = Math.max(0, total - size); this.shortened = true } From 7a51d5dfa0b792b693d13cedf7d79b8f799873a9 Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sun, 27 Sep 2026 07:02:34 -0700 Subject: [PATCH 42/48] test(conversations): every projected Codex column survives a wrong type (#1411 review round 2 survivors) Co-Authored-By: Claude Opus 5.5 --- .../conversations/sources/codex.system.test.ts | 17 ++++++++++++++--- 1 file changed, 14 insertions(+), 3 deletions(-) diff --git a/src/main/conversations/sources/codex.system.test.ts b/src/main/conversations/sources/codex.system.test.ts index 6927c8d8e..c37ce9bfa 100644 --- a/src/main/conversations/sources/codex.system.test.ts +++ b/src/main/conversations/sources/codex.system.test.ts @@ -129,13 +129,24 @@ describe('Codex conversation source', () => { const family = await resolveFamily('/fixture/repo', 'everywhere', { listWorktrees }) const before = await source.discover({ scope: 'everywhere', family }) const db = new DatabaseSync(join(corpus.codexHome, 'state_5.sqlite')) - const victim = (db.prepare('select id from threads where archived = 0 limit 1').get() as { id: string }).id - db.prepare("update threads set title = x'00', first_user_message = 'fallback label' where id = ?").run(victim) + const [victim, second] = (db.prepare('select id from threads where archived = 0 limit 2').all() as Array<{ id: string }>).map(row => row.id) + // Every string column the row projects, not only the title (review of + // #1411, round 2: guards on `source` and the others survived mutation). + db.prepare(`update threads set title = x'00', preview = x'00', name = x'00', source = x'00', thread_source = x'00', + agent_role = x'00', git_branch = x'00', originator = x'00', cwd = x'00', rollout_path = x'00', + created_at_ms = x'00', updated_at_ms = x'00', first_user_message = 'fallback label' where id = ?`).run(victim) + // The fallback chain itself: an empty title falls to a BLOB first message, + // which must fall through to the preview rather than throw. + db.prepare("update threads set title = '', first_user_message = x'00', preview = 'preview label' where id = ?").run(second) db.close() const after = await new CodexConversationSource({ codexHome: corpus.codexHome }).discover({ scope: 'everywhere', family }) expect(after).toHaveLength(before.length) - expect(after.find(r => r.nativeId === victim)?.userTexts).toEqual(['fallback label']) + const victimRow = after.find(r => r.nativeId === victim) + expect(victimRow?.userTexts).toEqual(['fallback label']) + expect(victimRow?.cwd).toBeNull() + expect(victimRow?.gitBranch).toBeNull() + expect(after.find(r => r.nativeId === second)?.userTexts).toEqual(['preview label']) }) it('falls back to the rollout scan when the index is missing and reports why', async () => { From 5689181222d98348171de110e08fc1927a4df1e6 Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sun, 27 Sep 2026 07:02:35 -0700 Subject: [PATCH 43/48] docs(plans): #1411 round-2 disposition; picker residual filed as #1433 Co-Authored-By: Claude Opus 5.5 --- docs/plans/2026-09-27-c5-fail-all-batch.md | 11 +++++++++++ 1 file changed, 11 insertions(+) diff --git a/docs/plans/2026-09-27-c5-fail-all-batch.md b/docs/plans/2026-09-27-c5-fail-all-batch.md index cebbd243b..586cec716 100644 --- a/docs/plans/2026-09-27-c5-fail-all-batch.md +++ b/docs/plans/2026-09-27-c5-fail-all-batch.md @@ -45,3 +45,14 @@ Rows 14–15 (key vault index, agent-name registry, tmux recovery) are strict by | **a survivor:** dropping the `approvedAt` check passed. | valid | An entry missing `approvedAt` prompts. Mutant red. | | **b survivor:** the row 13 loader returning an empty set passed. | valid | Row 13 moved to #1417 (manager q109: the same loader as steering q109), with a test through the real loader there. This PR no longer touches `debugRetention`. | | **a2, minor:** carried rows change position on rewrite. | declined | Values are all kept. Neither reader gives order any authority: incidents are sorted by `at`, and approvals are keyed by identity plus hash. Preserving the original interleaving would need positional bookkeeping for no reader. | + +## Review round 2 (a, b, c codex at `7204503c`): all FIX-BEFORE-MERGE (final round) + +| Finding | Verdict | Change | +|---|---|---| +| **a / b / c, major:** a set-aside refused file was deleted by a later run's retention. Setting it aside cleared the refusal marker, so once the run's readable incidents expired it looked empty. | valid | `refusedAsideRuns`: every run holding an `incidents.refused-*` file (listed at startup, added on set-aside) is never expired; only budget pruning and `clear()` remove it. Fail-first with a day-8 later run. | +| **a / b / c, major:** the aside name was only `Date.now()`, and `rename` replaces, so a same-millisecond second refusal overwrote the first | valid | The name adds a UUID. Fail-first with a fixed clock and two refusals: both bodies are kept. | +| **c survivor:** removing `refusedIncidentRuns.delete` after a set-aside (a stale marker) | valid | The collision test records again after a set-aside: one aside file, and the canonical file holds both incidents. Mutant red. | +| **b survivor:** the rename-failure guard | valid | A read-only run dir makes the set-aside fail; the refused file stays byte-for-byte and `shortened` is set. Mutant red. | +| **a / c survivors:** the `source`, `first_user_message`, `cwd` and `git_branch` type guards | valid | The wrong-type test now BLOBs every projected string column, plus a second row whose empty title falls to a BLOB first message; all four mutants red. | +| **c3, minor:** the picker still shows a partial Codex list as complete | residual, **filed as #1433** | Needs a per-source degraded status through `Discovery` and a curated picker line (q39); out of scope for a fail-all fix. | From 3c807a1b915814834248d74435a50bda1cfecf2b Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sun, 27 Sep 2026 07:08:22 -0700 Subject: [PATCH 44/48] fix(performance): a failed run listing is unknown, not empty; the run is kept (steering q115) Startup found set-aside refused files by listing each run with .catch(() => []). A failed listing read as 'no set-aside file', so a prior run holding only one had no marker and the next maintenance deleted it with the refused bytes (the unknown-as-empty shape q109 forbade). Now the run is marked unindexed (never expired) and the store degraded; only ENOENT means nothing to find. Real temp run + one injected per-run listing failure: red at 56891812 (ENOENT on the deleted run); removing the unindexed mark is red. Co-Authored-By: Claude Opus 5.5 --- docs/plans/2026-09-27-c5-fail-all-batch.md | 12 ++++ ...MonitorHistoryStore.listingFailure.test.ts | 60 +++++++++++++++++++ src/main/performance/MonitorHistoryStore.ts | 17 +++++- 3 files changed, 87 insertions(+), 2 deletions(-) create mode 100644 src/main/performance/MonitorHistoryStore.listingFailure.test.ts diff --git a/docs/plans/2026-09-27-c5-fail-all-batch.md b/docs/plans/2026-09-27-c5-fail-all-batch.md index 586cec716..95bd63d0b 100644 --- a/docs/plans/2026-09-27-c5-fail-all-batch.md +++ b/docs/plans/2026-09-27-c5-fail-all-batch.md @@ -56,3 +56,15 @@ Rows 14–15 (key vault index, agent-name registry, tmux recovery) are strict by | **b survivor:** the rename-failure guard | valid | A read-only run dir makes the set-aside fail; the refused file stays byte-for-byte and `shortened` is set. Mutant red. | | **a / c survivors:** the `source`, `first_user_message`, `cwd` and `git_branch` type guards | valid | The wrong-type test now BLOBs every projected string column, plus a second row whose empty title falls to a BLOB first message; all four mutants red. | | **c3, minor:** the picker still shows a partial Codex list as complete | residual, **filed as #1433** | Needs a per-source degraded status through `Discovery` and a curated picker line (q39); out of scope for a fail-all fix. | + +## Steering q115 (after round 2) + +**Finding.** Startup found set-aside refused files by listing each run with `.catch(() => [])`. A failed listing therefore read as "no set-aside file". A prior run holding ONLY a set-aside file had no marker, so the next maintenance deleted it with the refused bytes. This is the unknown-as-empty shape q109 forbade. + +**Fix.** A failed per-run listing is unknown: the run is marked unindexed, which maintenance never expires, and the store is degraded. Only ENOENT (the run is already gone) means there is nothing to find. + +**Test.** `MonitorHistoryStore.listingFailure.test.ts` uses a real temp run holding only an `incidents.refused-*` file and injects one failed plain listing of that run at startup (the parent `runNames()` listing succeeds). Maintenance then runs after reads recover, and the run and its exact bytes must survive. +- Red at `56891812`: `ENOENT` on the deleted run. +- Removing the unindexed mark: red. + +**Loss-path audit.** Whole runs leave disk only through `clear()`, budget `pruneRuns` and maintenance expiry, and expiry is guarded by the incident, foreign, refused, set-aside and unindexed markers. The other file removals are expired tier files, an empty `incidents.json` and `*.tmp` scratch. The remaining `.catch(() => [])` listings either sweep only `*.tmp` or undercount bytes, which makes budget pruning less aggressive, never more. diff --git a/src/main/performance/MonitorHistoryStore.listingFailure.test.ts b/src/main/performance/MonitorHistoryStore.listingFailure.test.ts new file mode 100644 index 000000000..8082bff9a --- /dev/null +++ b/src/main/performance/MonitorHistoryStore.listingFailure.test.ts @@ -0,0 +1,60 @@ +import { mkdir, mkdtemp, readFile, readdir, rm, writeFile } from 'node:fs/promises' +import { tmpdir } from 'node:os' +import { join } from 'node:path' +import { afterEach, expect, it, vi } from 'vitest' +import type { MonitorWorkerSnapshot } from '@shared/performance/monitorSnapshot.js' + +// One per-run listing fails, once, while the parent listing succeeds: the +// shape of a transient EIO/EACCES on a single run directory. Everything else +// is the real filesystem. Its own file because this mock covers the module. +const failOnce = vi.hoisted(() => ({ path: null as string | null })) +vi.mock('node:fs/promises', async importOriginal => { + const real = await importOriginal() + return { + ...real, + readdir: ((...args: Parameters) => { + // Only the plain name listing (no options): cleanupTemps and runBytes list + // the same directory with file types first and must not absorb the fault. + if (failOnce.path !== null && String(args[0]) === failOnce.path && args[1] === undefined) { + failOnce.path = null + return Promise.reject(Object.assign(new Error('injected EIO'), { code: 'EIO' })) + } + return (real.readdir as (...a: unknown[]) => unknown)(...args) + }) as typeof real.readdir, + } +}) +const { MonitorHistoryStore } = await import('./MonitorHistoryStore.js') + +const roots: string[] = [] +afterEach(async () => { await Promise.all(roots.splice(0).map(root => rm(root, { recursive: true, force: true }))) }) +const snapshot = (at: number): MonitorWorkerSnapshot => ({ + schemaVersion: 1, sampledAt: at, + main: { at, cpuPercent: 2, rss: 1024, heapUsed: 256, heapLimit: 2048, loopMeanMs: 20, loopP99Ms: 22, loopMaxMs: 25, sleepGap: false }, + windows: [], operations: [], recent: [], workerRss: 4096, +}) + +// Steering q115: startup found set-aside refused files by listing each run +// and read a failed listing as "no files" (`.catch(() => [])`). A prior run +// holding ONLY a set-aside file then had no marker at all, and the next +// maintenance deleted it with the refused bytes, the same unknown-as-empty +// shape q109 forbade for the debug ledger. +it('keeps a run whose listing failed at startup, with its set-aside bytes intact', async () => { + const root = await mkdtemp(join(tmpdir(), 'agent-code-monitor-')) + roots.push(root) + const runOld = join(root, 'runs', 'run-old') + await mkdir(runOld, { recursive: true }) + const aside = 'incidents.refused-1-00000000-0000-4000-8000-000000000000.json' + const refused = JSON.stringify({ version: 2, incidents: ['evidence'] }) + await writeFile(join(runOld, aside), refused) + + failOnce.path = runOld + const store = new MonitorHistoryStore(root, 'run-now') + await store.settled() + expect(failOnce.path).toBeNull() + store.record(snapshot(90_000), null, [], 0, 0) + await store.settled() + + expect(await readdir(runOld)).toContain(aside) + expect(await readFile(join(runOld, aside), 'utf8')).toBe(refused) + expect(store.status().state).toBe('degraded') +}) diff --git a/src/main/performance/MonitorHistoryStore.ts b/src/main/performance/MonitorHistoryStore.ts index bf0b4fed5..c0078dbdf 100644 --- a/src/main/performance/MonitorHistoryStore.ts +++ b/src/main/performance/MonitorHistoryStore.ts @@ -385,8 +385,21 @@ export class MonitorHistoryStore { try { await this.cleanupTemps() for (const run of await this.runNames()) { - const entries = await readdir(join(this.root, RUNS_DIR, run)).catch(() => [] as string[]) - if (entries.some(name => name.startsWith(REFUSED_INCIDENTS_PREFIX))) this.refusedAsideRuns.add(run) + // A failed listing is UNKNOWN, not empty (steering q115): read as "no + // set-aside file", a prior run holding only one had no marker left and + // the next maintenance deleted it with the refused bytes. Such a run is + // marked unindexed, which maintenance never expires, and the store is + // degraded. Only ENOENT (the run is already gone) means nothing to find. + // Budget pruning (pruneRuns) stays the only way this run leaves disk. + try { + const entries = await readdir(join(this.root, RUNS_DIR, run)) + if (entries.some(name => name.startsWith(REFUSED_INCIDENTS_PREFIX))) this.refusedAsideRuns.add(run) + } catch (error) { + if ((error as NodeJS.ErrnoException).code !== 'ENOENT') { + this.degraded = true + this.unindexedRuns.add(run) + } + } for (const resolution of TIERS) { const file = join(this.root, RUNS_DIR, run, `${resolution}.jsonl`) try { From ab1a4b43b15f484ae51a7fb2288774cd64135b32 Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sun, 27 Sep 2026 08:15:26 -0700 Subject: [PATCH 45/48] docs(plan): happy-dom 20.14.5 bump and the OffscreenCanvas shim gap (#1365) Co-Authored-By: Claude Opus 5.5 --- ...-09-27-happy-dom-20-14-offscreen-canvas.md | 63 +++++++++++++++++++ 1 file changed, 63 insertions(+) create mode 100644 docs/plans/2026-09-27-happy-dom-20-14-offscreen-canvas.md diff --git a/docs/plans/2026-09-27-happy-dom-20-14-offscreen-canvas.md b/docs/plans/2026-09-27-happy-dom-20-14-offscreen-canvas.md new file mode 100644 index 000000000..c3b53814c --- /dev/null +++ b/docs/plans/2026-09-27-happy-dom-20-14-offscreen-canvas.md @@ -0,0 +1,63 @@ +# happy-dom 20.9 → 20.14.5 (#1365) + +## Why this exists + +Dependabot #1322 moved happy-dom 20.9.0 → 20.14.5 and its quality gate (run +36168289491, head `cfc31862`, 2026-09-25T17:37Z) failed in three renderer +files. #1360 took #1322's other bumps with happy-dom held back. #1365 asks which +happy-dom change broke event dispatch, and forbids weakening the tests to get +the bump through. + +## What the evidence says + +Reproduced locally with happy-dom 20.14.5 loaded (verified via +`navigator.userAgent` → `HappyDOM/20.14.5`; a symlinked vitest silently loads +the shared tree's 20.9.0, so the version has to be checked, not assumed). + +1. **`terminalWheelBoundary.xterm.renderer.test.ts`: real, and it comes from + happy-dom.** All 5 tests fail in `term.open()` with `value must not be + falsy` from xterm's `WidthCache` (`RendererUtils.throwIfFalsy`). The cause + is not event dispatch. happy-dom **20.10.0** added a global + `OffscreenCanvas` (`lib/canvas/OffscreenCanvas.js`; 20.9.0 has no + `lib/canvas` at all). xterm's WidthCache prefers + `new OffscreenCanvas(1, 1)` whenever the global exists, so it bypasses the + test's shim, which patches only `HTMLCanvasElement.prototype.getContext`. + happy-dom's `OffscreenCanvas.getContext` returns `null` when no + `canvasAdapter` is configured, which is exactly the "no 2D context" gap the + shim already fills for ``. + xterm's CharSizeService also probes `new OffscreenCanvas(100, 100)` for its + TextMetrics strategy. It wraps that probe in try/catch and falls back to + the DOM strategy, so it degrades on its own, and shim 2 (the measure-span + size) still applies. +2. **`CommandKeybindingsRow.capture.renderer.test.tsx` (7 tests) and + `ConversationsPicker.renderer.test.tsx` (2 tests): not happy-dom.** The + failures were `Unable to find … role "button" and name "Add"` or + `"everywhere"`, while the rendered DOM in the same log shows + `"Add a shortcut to New Tab"`. #1221 (keyboard-first) renamed those + controls. The tests were updated in `e092232c` at 17:52Z, 15 minutes after + the #1322 run. On current main both files pass unchanged with 20.14.5 + (34/34 across the three files, with only the xterm file failing). + +## Plan + +1. Commit this plan. +2. Fail-first: bump `happy-dom` to `^20.14.5` in package.json and the lockfile, + touching nothing else. The existing xterm wheel test then fails on the real + path (`open()` → WidthCache). No new test is needed: the existing one is + the real-data oracle. +3. Fix: shim 1 installs the same fake 2D context on + `OffscreenCanvas.prototype.getContext` when the global exists, restored in + `afterEach` like the others. Update the WHY block: it still claims happy-dom + has no OffscreenCanvas, and that is now false. Keep the DOM char-measure + strategy: the fake context has no `fontBoundingBox*`, so xterm's + CharSizeService still falls back to the DOM strategy that shim 2 sizes. + Say so, so the next bump does not have to re-derive it. + Rejected alternative: deleting `globalThis.OffscreenCanvas` for the test. + That would hide the path Electron takes in production (Chromium has + OffscreenCanvas, so WidthCache uses it there too). Shimming it keeps the + test on production's branch. +4. Run the three files plus the renderer suite once at the end. + +## Non-goals + +No assertion changes. No other dependency moves. From 53f8b0394e71d3ddfa75d7461a004a6909501038 Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sun, 27 Sep 2026 08:15:56 -0700 Subject: [PATCH 46/48] chore(deps): bump happy-dom 20.9.0 -> 20.14.5 (fail-first, #1365) happy-dom 20.10 added a global OffscreenCanvas. xterm's WidthCache prefers it over , so the real-xterm wheel test's 2D-context shim is bypassed and term.open() throws 'value must not be falsy'. The shim fix follows. Co-Authored-By: Claude Opus 5.5 --- package-lock.json | 24 +++++++++++++++++++----- package.json | 2 +- 2 files changed, 20 insertions(+), 6 deletions(-) diff --git a/package-lock.json b/package-lock.json index 7b4df100d..d41c54701 100644 --- a/package-lock.json +++ b/package-lock.json @@ -76,7 +76,7 @@ "electron-updater": "^6.8.9", "electron-vite": "^5.0.0", "esbuild": "0.25.12", - "happy-dom": "^20.9.0", + "happy-dom": "^20.14.5", "tailwindcss": "^4.2.2", "tsx": "^4.23.15", "typescript": "^5.5.0", @@ -4778,6 +4778,19 @@ "dev": true, "license": "MIT" }, + "node_modules/buffer-image-size": { + "version": "0.6.4", + "resolved": "https://registry.npmjs.org/buffer-image-size/-/buffer-image-size-0.6.4.tgz", + "integrity": "sha512-nEh+kZOPY1w+gcCMobZ6ETUp9WfibndnosbpwB1iJk/8Gt5ZF2bhS6+B6bPYz424KtwsR6Rflc3tCz1/ghX2dQ==", + "dev": true, + "license": "MIT", + "dependencies": { + "@types/node": "*" + }, + "engines": { + "node": ">=4.0" + } + }, "node_modules/builder-util": { "version": "26.16.0", "resolved": "https://registry.npmjs.org/builder-util/-/builder-util-26.16.0.tgz", @@ -6688,18 +6701,19 @@ "license": "ISC" }, "node_modules/happy-dom": { - "version": "20.9.0", - "resolved": "https://registry.npmjs.org/happy-dom/-/happy-dom-20.9.0.tgz", - "integrity": "sha512-GZZ9mKe8r646NUAf/zemnGbjYh4Bt8/MqASJY+pSm5ZDtc3YQox+4gsLI7yi1hba6o+eCsGxpHn5+iEVn31/FQ==", + "version": "20.14.5", + "resolved": "https://registry.npmjs.org/happy-dom/-/happy-dom-20.14.5.tgz", + "integrity": "sha512-x/RzkpWO40bTjIoT30iQtt64FLLmH/iRcUCN2X//bLx7H3ifkdfPXyqsro/OYtqzIAhiLMMA7mmiOR9C3NOKjQ==", "dev": true, "license": "MIT", "dependencies": { "@types/node": ">=20.0.0", "@types/whatwg-mimetype": "^3.0.2", "@types/ws": "^8.18.1", + "buffer-image-size": "^0.6.4", "entities": "^7.0.1", "whatwg-mimetype": "^3.0.0", - "ws": "^8.18.3" + "ws": "^8.21.0" }, "engines": { "node": ">=20.0.0" diff --git a/package.json b/package.json index 36da775e9..12d89ba1c 100644 --- a/package.json +++ b/package.json @@ -144,7 +144,7 @@ "electron-updater": "^6.8.9", "electron-vite": "^5.0.0", "esbuild": "0.25.12", - "happy-dom": "^20.9.0", + "happy-dom": "^20.14.5", "tailwindcss": "^4.2.2", "tsx": "^4.23.15", "typescript": "^5.5.0", From 12e8f5feb8f758f74d4c6eca3d793f53815a9dec Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sun, 27 Sep 2026 08:16:33 -0700 Subject: [PATCH 47/48] test(terminal): shim OffscreenCanvas 2D context for the real-xterm wheel test (#1365) Co-Authored-By: Claude Opus 5.5 --- ...rminalWheelBoundary.xterm.renderer.test.ts | 29 +++++++++++++++++-- 1 file changed, 26 insertions(+), 3 deletions(-) diff --git a/src/renderer/src/workspace/terminal/terminalWheelBoundary.xterm.renderer.test.ts b/src/renderer/src/workspace/terminal/terminalWheelBoundary.xterm.renderer.test.ts index b6fa6ad30..92cf73942 100644 --- a/src/renderer/src/workspace/terminal/terminalWheelBoundary.xterm.renderer.test.ts +++ b/src/renderer/src/workspace/terminal/terminalWheelBoundary.xterm.renderer.test.ts @@ -21,10 +21,21 @@ import { attachTerminalWheelBoundary } from './terminalWheelBoundary' // and scroll dimension is 0 (PR #792 review round 1 stopped there). The shims // below supply only that missing browser environment: // 1. A 2D context exposing the two members WidthCache touches (`font` and -// `measureText().width`). It returns a constant glyph width. +// `measureText().width`). It returns a constant glyph width. It goes on +// BOTH `HTMLCanvasElement` and `OffscreenCanvas`. happy-dom 20.10 added a +// global `OffscreenCanvas` whose `getContext` returns null when no canvas +// adapter is configured, and WidthCache prefers `new OffscreenCanvas(1, 1)` +// over `` whenever that global exists. With only the `` +// shim, `open()` threw `value must not be falsy` again on 20.14.5 (#1365). +// We shim rather than delete the global: Chromium has OffscreenCanvas, so +// production xterm takes the OffscreenCanvas branch as well. // 2. A fixed size for xterm's `.xterm-char-measure-element` span, which -// CharSizeService's DOM strategy reads. happy-dom has no OffscreenCanvas, -// so the DOM strategy is the one xterm picks. +// CharSizeService's DOM strategy reads. CharSizeService first tries a +// TextMetrics strategy on `new OffscreenCanvas(100, 100)`. It needs +// `fontBoundingBoxAscent`/`Descent`, which the fake context deliberately +// lacks, so that constructor throws and xterm's own try/catch falls back to +// the DOM strategy that this shim sizes. If the fake ever gains those +// fields, xterm switches strategies and this shim becomes dead. // 3. Explicit 0px left/top padding on `.xterm-screen`, which a browser // computes from xterm.css. happy-dom's computed style returns '' for unset // padding, and MouseCoordsService parseInt()s it into NaN. @@ -76,6 +87,18 @@ function installBrowserShims(): void { } as unknown as typeof canvasPrototype.getContext restoreShims.push(() => { canvasPrototype.getContext = originalGetContext }) + // A `typeof` guard, not an assertion. Without the global, xterm falls back + // to ``, which the shim above already covers, so its absence is + // not a failure. + if (typeof OffscreenCanvas !== 'undefined') { + const offscreenPrototype = OffscreenCanvas.prototype + const originalOffscreenGetContext = offscreenPrototype.getContext + offscreenPrototype.getContext = function (this: OffscreenCanvas, contextId: string) { + return contextId === '2d' ? fakeContext : originalOffscreenGetContext.call(this, contextId as '2d') + } as unknown as typeof offscreenPrototype.getContext + restoreShims.push(() => { offscreenPrototype.getContext = originalOffscreenGetContext }) + } + const sizes = [['offsetWidth', GLYPH_WIDTH_PX * MEASURE_SPAN_REPEAT], ['offsetHeight', GLYPH_HEIGHT_PX]] as const for (const [property, size] of sizes) { const original = Object.getOwnPropertyDescriptor(HTMLElement.prototype, property) From 033a7d81b8d950df99e0ad12cedf6f44fc38ea97 Mon Sep 17 00:00:00 2001 From: Julius Olsson Date: Sun, 27 Sep 2026 08:56:58 -0700 Subject: [PATCH 48/48] test(terminal): correct the padding provenance in the wheel test WHY block (#1446 review c) Co-Authored-By: Claude Opus 5.5 --- .../terminalWheelBoundary.xterm.renderer.test.ts | 14 ++++++++------ 1 file changed, 8 insertions(+), 6 deletions(-) diff --git a/src/renderer/src/workspace/terminal/terminalWheelBoundary.xterm.renderer.test.ts b/src/renderer/src/workspace/terminal/terminalWheelBoundary.xterm.renderer.test.ts index 92cf73942..001cab9ba 100644 --- a/src/renderer/src/workspace/terminal/terminalWheelBoundary.xterm.renderer.test.ts +++ b/src/renderer/src/workspace/terminal/terminalWheelBoundary.xterm.renderer.test.ts @@ -36,17 +36,19 @@ import { attachTerminalWheelBoundary } from './terminalWheelBoundary' // lacks, so that constructor throws and xterm's own try/catch falls back to // the DOM strategy that this shim sizes. If the fake ever gains those // fields, xterm switches strategies and this shim becomes dead. -// 3. Explicit 0px left/top padding on `.xterm-screen`, which a browser -// computes from xterm.css. happy-dom's computed style returns '' for unset -// padding, and MouseCoordsService parseInt()s it into NaN. +// 3. Explicit 0px left/top padding on `.xterm-screen`. A browser computes +// 0px from the CSS initial value: xterm.css has no padding rule for +// `.xterm-screen`, and neither does our styles.css. happy-dom's computed +// style returns '' for unset padding, and MouseCoordsService parseInt()s it +// into NaN. // Separately, happy-dom's WheelEvent extends UIEvent rather than MouseEvent, // so it lacks the modifier and clientX/clientY fields every browser WheelEvent // carries; `wheel()` below sets them explicitly. Either gap yields // `ESC[<64;NaN;NaNM` instead of a real report, a happy-dom artifact rather // than an xterm bug. The first real run had both gaps, and a run with only the -// padding shim still got NaN from the missing clientX. That unset padding alone -// also yields NaN comes from reading happy-dom's computed style (it has no -// default padding) and was not observed on its own. +// padding shim still got NaN from the missing clientX. A run with clientX set +// but without the padding shim (#1446 review c) produced exactly the same bytes, +// so each gap alone is enough to break the report. // Everything asserted is computed by xterm itself from real WheelEvents: // - StandardWheelEvent delta normalization; // - the `fastScrollSensitivity` Alt multiplier;