Skip to content

Pass every Claude goal to the CLI as a structured message - #2070

Merged
ppXD merged 1 commit into
mainfrom
fix/pass-every-claude-goal-as-a-structured-message
Oct 6, 2026
Merged

ppXD merged 1 commit into
mainfrom
fix/pass-every-claude-goal-as-a-structured-message

Conversation

@ppXD

@ppXD ppXD commented Oct 6, 2026

Copy link
Copy Markdown
Owner

Summary

  • On a text stdin the pinned Claude CLI 2.1.263 acts on the goal before the model sees it, in plan mode as in bypass. It reads every @path in the goal into the model request and the session transcript: absolute and ~ paths, directory listings, symlinks out of the workspace, and ~/.mcp.json, which under bubblewrap is the run's MCP declaration and token. It also runs a goal that starts with /word as a command: /security-review runs git, and an unknown word ends the run with exit 0 and no model turn, which the run records as Completed. ClaudeCodeHarness.BuildInvocation (backend/src/CodeSpace.Core/Services/Agents/Harnesses/Claude/ClaudeCodeHarness.cs) now passes --input-format stream-json and writes one NDJSON user message. The goal is the first text block, byte for byte; the constant Begin with the task above. is the last. The CLI parses mentions and commands only from the last block. Fresh, --resume, revise and reviewer prompts all take this path.
  • A goal the CLI would treat as blank is refused. In a structured message the CLI drops a block that JavaScript's trim() empties (U+FEFF and every Zs included), then runs the model on the trailer alone and reports success. IsBlankToTheCli treats .NET whitespace plus U+FEFF as blank. Invisible characters the CLI keeps (U+200B, U+3164) are still sent.
  • The message is serialized with relaxed escaping because the launch pipe re-encodes stdin with its own JSON encoder: non-ASCII text is escaped once, by the pipe. Characters the message must escape itself (quotes, backslashes, control characters, U+FEFF/U+2028, non-BMP characters) are escaped again on the pipe and cost at most twice the goal's pipe size. Launch-size accounting measures the encoded message.

Test plan

  • Unit: argv and message shape; 9 adversarial goals reach the first block byte for byte; trailer pinned; blank refusal, including U+FEFF, and a theory over every character ECMAScript trim() removes; U+200B/U+3164 still sent; non-ASCII escaped once and JSON-escaped characters bounded at 2x on the pipe; launch preflight measures the message. Full unit suite: 11755 passed, 1 skipped.
  • Integration: the real executor hands the goal through byte for byte, and a shared awk goal reader with a drift detector covers the 6 Claude-dialect fake CLIs. The 21 classes that exercise the Claude harness: 405 passed, 1 skipped.
  • E2E GoalChannelE2ETests (real CLI 2.1.263, stub model on loopback, fake secrets): mentions, resumed session, /fix, /security-review. 4/4 in plan mode and 4/4 in bypass on macOS. The text-channel positive control reads all 14 planted secrets. For a slash word it must reach its result with no model turn and record the word as its own command; mutations with empty control stdin or an unknown control flag turn it red.
  • Linux sandbox-isolation lanes: root (bwrap, floor 96) and non-root uid 1654 (floor 15), each with the four [goal-channel-e2e] ran markers

The Claude harness wrote the goal to the CLI's stdin as text. On that
channel the pinned 2.1.263 CLI acts on the goal before the model sees
it, in plan mode as in bypass. Every @path after start of text,
whitespace (JavaScript's \s, so a BOM, NBSP, U+3000 and U+2028 count)
or CJK punctuation is read into the model request and the session
transcript: absolute and ~ paths, directory listings, symlinks out of
the workspace. A goal that opens with /word runs as a command:
/security-review runs git without the CLI's own hardening, /heapdump
writes a heap snapshot holding the run's tokens, and an unknown word
ends the run with exit 0, is_error false and no turn, which the run
records as Completed. Goals carry text from pull requests, repositories
and other models, and plan-mode reviewers are no exception.

The CLI parses mentions and commands out of the last text block of a
stream-json user message only. BuildInvocation, the one place every
Claude prompt is built (fresh, --resume, revise, reviewer), now passes
--input-format stream-json and writes one NDJSON user message: the goal
as its first block, byte for byte, and a constant trailer as the last.
Escaping the sigils instead would change what the model reads and has
to copy the CLI's \s exactly; a .NET \s port misses the BOM and the
file is read anyway.

The message is serialized with relaxed escaping, because the launch
pipe measures and re-encodes stdin with its own JSON encoder: non-ASCII
text is escaped once, by the pipe. What the message must escape itself
(quotes, backslashes, controls, characters outside the BMP) is escaped
again on the pipe and costs at most twice the goal's pipe size. A blank
goal is refused: in a structured message the CLI drops a block its
JavaScript trim() empties, U+FEFF included, and runs the model on the
trailer alone, reporting success.

The block order is undocumented, so a real-CLI E2E pins it in plan and
bypass, fresh and resumed, against a text-channel positive control
that must read the same planted secrets, or run the same slash word as
its own command. The fake CLIs that serve the Claude dialect decode the
message with a shared awk reader, pinned against the real encoder. The
--append-system-prompt and Stop-hook reason channels were probed and
are inert to @ and /.
@ppXD
ppXD merged commit b9b14ad into main Oct 6, 2026
6 checks passed
@ppXD
ppXD deleted the fix/pass-every-claude-goal-as-a-structured-message branch October 6, 2026 11:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant