feat(core): shared host-hook protocol (S8a) - #161
Open
BrainerVirus wants to merge 12 commits into
Open
BrainerVirus wants to merge 12 commits into
BrainerVirus wants to merge 12 commits into
Conversation
Records written by a newer runtime may carry keys an older reader does not know. Strict record schemas rejected them as recovery_required, so one additive field would brick every older host reading the same store (D17). Reads now drop exactly the keys zod reports as unrecognized and parse again; any other violation still fails, and writes stay strict. hostSchema gains claude_code so the Claude Code adapter does not need to touch the task contract. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
core/hooks is the one implementation of session start (including the compact restore), per-turn context, pre-shell branch policy, and subagent start for every host. Hosts only parse native payloads and render decisions. Parsers ignore unknown keys and check only the keys a mapping needs. A pre-tool parse failure denies only on hosts whose descriptor is fail-closed; other events fail open with a diagnostic. Each host declares a descriptor of what it supports, with citations. capabilitiesFor derives engine capabilities from it; an undocumented axis never backs an enforced capability. session-context moves to hooks/context with one unfinished-task offer and session predicate. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Codex rejected any unknown payload key and denied PreToolUse on every parse error, so a Codex release that adds a field would deny every tool call. The Codex hook is now a thin mapping onto core/hooks: unknown keys are ignored and a broken payload passes through with a stderr diagnostic. Pi and OpenCode take branch policy, session context, the unfinished-task offer, and their capabilities from core/hooks and their descriptors. The capability arrays are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
hooks-cursor.json registered beforeShellExecution, but the hook always answered allow. It now maps onto core/hooks and denies direct branch creation onto a protected or noncompliant name with exit 2, like the other hosts. subagentStart keeps its Cursor-only worker assignment path. Hook bundles now import core/hooks and direct modules instead of the core barrel, and a test pins that they load no doctor or setup code. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Substituting a raw Windows cwd into fixture JSON produced invalid escape sequences; the path is now JSON-escaped. Bundle metafile inputs are normalized to forward slashes before matching. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> # Conflicts: # packages/workit-codex/hooks/workit-hook.ts # packages/workit-cursor/hooks/workit-hook.ts
…ripped records Reader tolerance dropped every unknown key, which removed the fail-closed upgrade guard for fields an older reader must not ignore. Records may now list such paths in a top-level critical array: a reader that would strip one fails with the upgrade message, so it neither acts on nor rewrites the record. Recovery copies a snapshot back verbatim, so it refuses any record it could only read by dropping fields. Under a plain union, stripping now keeps the branch that drops the fewest keys instead of the first that parses, so a key one branch knows is never lost. The design records the rule for new fields and the Claude PowerShell matcher, and a follow-up to verify Cursor's allow live. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Codex parsing still rejected payloads over fields branch policy never reads, such as a permission mode Codex adds later, and the parse failure then let protected-branch creation through. A Codex shell gate now needs only cwd and the command; the command may be an argv array, and a shell -c script is checked as the script. Claude treats PowerShell like Bash, and shows the PreCompact notice through systemMessage. Cursor accepts multi-root workspaces and checks a shell command against its own cwd. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… sessions - commandText finds the script after the first script flag, so [bash, -e, -c, s], [powershell.exe, -Command, s] and [cmd, /c, s] are checked as s. - Numeric segments in declared critical paths match any array element, so authorityRefs.0.future fails closed like authorityRefs.*.future. - A Codex session without an id gets no writer-acquire actor, and id-less sessions no longer share one per-process offer key. - Cursor subagentStart assigns workers in the workspace root that contains the payload cwd, the same rule as other events. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Slice S8a of the workit-next design (
docs/workit-next/design.md§1, §5).What changes
packages/workit-core/src/hooks/: one host-hook implementation, exported as@brainervirus/workit-core/hooks.protocol.ts: the protocol types.handle.ts:handleHook,dispatchHook, and the fail policy.context.ts: moved fromcore/session-context.ts. It holds one unfinished-task offer and one session predicate for every host.policy.ts: branch policy for shell commands.descriptor.ts: the per-host capability descriptor andcapabilitiesFor.run.ts: the stdin→stdout hook process, which can also run asrun.ts <host>.hosts/{claude-code,codex,cursor,opencode,pi}.ts: field mappings and descriptors.PreToolUseon any parse error, so a Codex release that added one field would have denied every tool call. Now a pre-tool parse failure denies only when the host's descriptor setsfailClosed(Cursor, which is unchanged). Other hosts pass the call through and write a diagnostic to stderr.codexCapabilities,cursorCapabilities,piCapabilitiesand OpenCode's inline V2 array. The old function names are kept as thin wrappers. Anundocumentedaxis never backs anenforcedcapability, andpartialsupport lowersenforcedtoagent_guided.parseStoredRecorddrops exactly the keys zod reports as unrecognized and parses again. Any other schema violation still fails. Writes stay strict. A field added by a newer runtime therefore no longer turns older readers intorecovery_required.hostSchemagainsclaude_code.beforeShellExecutionnow enforces branch policy. It denies with exit 2 and anagent_messagethat includes the correction; before, it always returned allow.subagentStartkeeps its Cursor-only worker-assignment path.Acceptance (design §5, S8a)
main, Claude PreToolUsegit checkout -b mainpiped to the hook →permissionDecision:"deny",protected_ref; no host receivesallowtest/workit-core/hooks/claude-code.test.ts(spawnsrun.ts claude_code; checks that every Claude fixture and the Codex PreToolUse output never contain"allow")beforeShellExecution→ deny with exit 2; the parity table includes cursor and claude_codetest/workit-cursor/task-hooks.test.ts(spawned hook, exit 2);route-denial-parity.test.tsnow has cursor and claude_code rows;hooks/protocol.test.tschecks every host against the design §1.3 mapping table plus native fixtureshooks/lenient-parse.test.tsundocumentedaxis → neverenforcedhooks/descriptor.test.ts(sets each axis of each descriptor to undocumented; also snapshots the descriptors)source:"compact"→<workit-task-context>hooks/claude-code.test.ts(Claude and Codex)hooks/bundle.test.ts(metafile check; no core barrel either)Verification (measured, Node 24.20.0)
bun run lint,bun run format:check,tsc --noEmit,bun run knip(63 core files reachable from 26 entries): all clean.bun testafterbun run build, with origin/main merged (including test: cull duplicate/tautological tests, fix env-brittle checks, split unit and packaging tiers #155): 1449 pass, 0 fail. CI is green on all 15 checks, including Windows. The first CI run found a Windows-only fixture-path bug, fixed in 4526aee.codexCapabilities(both surfaces),cursorCapabilities(all 32 availability combinations),piCapabilitiesand the OpenCode V2 array were compared against the new descriptors. All 100 comparisons were identical.handleCodexHook/handleCursorHookwere compared on 54 payloads (session start with and without tasks, shell commands, writes, subagents, malformed input). There were 0 differences, apart from the intended Cursor deny.undocumentedcounted as usableChanges to existing tests (intended behavior changes)
codex/cli.test.ts: an unknown key now parsesok:true(D17).protocol-ergonomics.test.ts: a record from a newer runtime that only adds a field is now readable. The "upgrade Workit" message is still checked, on a newer record that actually violates the schema.Deviations from the design
task-enginecome to about 459 KB. Hook bundles no longer load the core barrel, doctor or setup. A regression ceiling of 620 KB is pinned until the graph is split.tool.preevent for non-shell pre-tool gates (Codex/ClaudePreToolUsewith other tools, CursorpreToolUse). It always returnsnone.label,context.task(session-bound|single-active, so Codex and Cursor keep showing the workspace's single active task),interaction.{questions,writeBoundary}(Pi and OpenCode need them forenforcedclaims), andcapabilities(the rules).effective()and theknip.jsonchange were not needed.test/fixtures/hooks/<host>were written from documented and design-verified field lists, not captured from live sessions.subagentStop/preCompactand unknown-event parse failures now return{}. Before, they returned a deny JSON with exit 0, which was not blocking.Risks
{permission:"allow"}. Cursor still answers{permission:"allow"}for compliant shell commands and allowed pre-tool calls; this contract is unchanged and existing tests pin it. Whether Cursor'sallowskips its own approval prompt has not been verified.Review fixes (016f0c7, 31214d8)
cwdand the command (D17). A CodexPreToolUsegate now requires onlycwd,tool_nameand the command. Every other field is optional, including model, session, turn, transcript, tool-use id and permission mode; an unknown permission mode or start source is accepted.tool_input.commandmay be an argv array:[sh, -c|-lc, script]checks the script, and other arrays are joined with quoting. A test coverspermission_mode:"weird", argv commands,bash -lc,unified-exec, and the bare{cwd, tool_name, tool_input}payload; all of them still denymain.critical: string[]listing dotted paths, with*for array items. A reader that would strip a critical path, or its parent, fails with "requires fields this Workit cannot read …; upgrade Workit". Mutations read first, so they refuse and the file stays untouched. Recovery parses snapshots inrewritemode and refuses any record it could only read by stripping keys. The rule is documented inparseStoredRecordand in design §0 feat: editable templates + canonical multi-platform rules (Phase 2) #7: a new field must be ignorable, or the writer must declare it critical. Tests cover both cases, the parent-path case, and recovery.z.union, the branch that drops the fewest keys now wins. The synthetic repro{a,b,future}now keepsb.PowerShell. It gets the same shell policy asBash. The design's hooks.json now has aPowerShellPreToolUsematcher.beforeShellExecutionchecks the policy against the payloadcwdwhen it is set, otherwise againstworkspace_roots[0]. A test uses a strict workspace glob that matches only a nested repository.PreCompact. The notice is now carried in the top-levelsystemMessage.unified-exec, union branch selection, thepath.isAbsolutecheck oncwd, and argv handling.allow. Unchanged, since it was already there. A design §0 feat: flow rails — issue prompt fix + subagent-driven enforcement #11 follow-up asks to verify against a live Cursor whether{permission:"allow"}skips Cursor's own approval prompt.Measured after merging #153 (Node 24.20.0).
bun run lint(type-aware),format:check,tscandknipare clean.bun testafter build: 1458 pass, 0 fail.unified-execdroppedisAbsolutecheck removedCA-06, an isolatednpm install, timed out at 122 s on the first run; it passed on rerun.Re-verification residuals (7c2f29e)
commandTextscans past leading options to the first script flag (-…c,-Command,/c), so["bash","-e","-c",s],["powershell.exe","-NoProfile","-Command",s]and["cmd","/c",s]are checked ass. Scanning stops at the first positional argument, so["bash","script.sh","-c","x"]is not unwrapped.criticalpaths are normalized to*, so a record declaringintent.data.authorityRefs.0.futurefails closed.subagentStart. Workers are assigned in the workspace root that contains the payloadcwd, falling back to the first root. This is the same rule as other events. Test: a multi-root workspace with the task in the second root.Measured.
bun testafter build: 1463 pass, 0 fail.-Command//cflags removedsubagentStartusingworkspace_roots[0]🤖 Generated with Claude Code