Skip to content

feat(ai): Claude Code provider — run the CMS agent on a Claude subscription - #577

Open
cakelesscoder wants to merge 4 commits into
CoreBunch:mainfrom
cakelesscoder:pr/claude-code-provider
Open

cakelesscoder wants to merge 4 commits into
CoreBunch:mainfrom
cakelesscoder:pr/claude-code-provider

Conversation

@cakelesscoder

@cakelesscoder cakelesscoder commented Sep 29, 2026 •

Copy link
Copy Markdown

Proposed in #589 — please read that first. It asks whether you want this direction at all, and I'm happy to close this PR if the answer is no.

Summary

Adds a Claude Code AI provider that runs the CMS agent on the operator's Claude subscription (Pro/Max) instead of a metered Anthropic API key.

You type in the normal AI panel. Behind the scenes the driver spawns the headless claude CLI, points it at Instatic's own MCP server, and streams its output back into the chat — so the agent is real Claude Code (its tool loop, its harness) editing the live workspace you have open. This is the inverse of connecting an external Claude Code to Instatic over MCP connectors: here Instatic drives Claude Code, but editing still flows through the same live editor bridge.

AI panel  →  POST /admin/api/ai/chat/site  →  resolveDriver('claude-code')
                                                    │
                     Bun.spawn('claude', --output-format stream-json …)
                                                    │
                        claude CLI ──MCP (bearer)──▶ /_instatic/mcp
                                                    │
                     stream-json stdout → AiStreamEvent → the open workspace

Why a subprocess, not an SDK. ai-driver-isolation.test.ts bans every provider SDK. This driver is compliant by construction — it imports nothing and spawns a binary. That is also the only thing that works: the subscription cost model exists solely in the CLI/subscription auth path, and the Agent SDK bills API credits, which would defeat the purpose.

Layout. Three modules by responsibility, each comfortably inside the 700-line budget — no GRANDFATHERED entry:

Module Lines Responsibility
claudeCode.ts 309 What to ask for — provider definition, argv/env, session-mode decision
claudeCodeProcess.ts 199 How to run it — one spawn, its failure modes, the outcome each produces
claudeCodeEvents.ts 277 How to read what came back — stream-json validation and translation

claudeCodeEvents.ts is pure and synchronous (no subprocess, no IO), which is what makes the translation layer testable from recorded CLI output.

Auth is the machine's Claude subscription. ANTHROPIC_API_KEY / ANTHROPIC_AUTH_TOKEN are stripped from the child env so the CLI can't silently fall back to metered billing, and --bare is never used (it forces API-key auth).

Capabilities. An internal MCP connector mints a per-user bearer token granted exactly the chatting user's own capabilities, so the spawned CLI can do precisely what that user could through the built-in agent, never more.

Tool surface. The CLI sees only this server's mcp__instatic__* tools — no Bash, Edit or WebFetch in the CMS chat.

Credential-less. The provider needs no secret, but a conversation must reference a credential row, so it stores an inert baseUrl sentinel (claude-code://local) the driver never dials. This satisfies the existing ai_creds_apikey_shape_check without a migration.

The part worth reviewing closely

--tools "" strips built-ins. The CLI's tool search (ENABLE_TOOL_SEARCH, default auto) defers a large MCP toolset behind the ToolSearch built-in — which --tools "" strips. With ~50 tools this server crosses the deferral threshold, so the combination leaves the model with zero callable tools.

A model with no tools doesn't error. It narrates a tool call in prose (`get_context` … calling that now) and ends the turn subtype: "success". From the composer that is indistinguishable from the agent just stopping a few seconds in.

The driver sets ENABLE_TOOL_SEARCH=false, which presents all tools directly and is also faster (~5s to first tool call vs ~13–19s through search round trips). Because that's a CLI default the integration now depends on, it is also checked: the driver reads the tool list from the CLI's system/init event and aborts the turn with a clear error if no mcp__instatic__* tool is present, rather than letting the model bluff.

The same principle covers the other ways a subprocess goes quiet — each terminal and named rather than silent:

Failure Handling
Wedged child stderr drained concurrently (an unread pipe blocks the writer once full); idle watchdog kills a CLI that stops emitting, SIGTERM→SIGKILL
Session mismatch --resume of a missing session / --session-id of an existing one each retry once as the other mode — safe because the failed attempt emitted nothing
Subscription rate limit surfaced with its reset time
CLI format drift unreadable output lines counted and logged
Abandoned stream child reaped instead of leaked

Each turn logs [ai/claude-code] session <id> started — N CMS tool(s), which is the first thing to check if the agent misbehaves.

Docker

Dockerfile.claude-code overlays the base image with the self-contained claude binary; compose.claude-code.yml supplies a subscription token via CLAUDE_CODE_OAUTH_TOKEN. No home-directory mount.

Verification

  • bun run build
  • bun test
  • bun run lint
  • Docker/deployment check — overlay image built and run; provider exercised end-to-end against a live site (real MCP tool calls, multi-turn session resume, and the failure paths above).

src/__tests__/ai/claudeCodeMapping.test.ts — 14 tests over the stream-json translation: text deltas, thinking suppression, tool-call/result pairing and prefix stripping, the init tool-count guard and MCP server status, rate-limit classification, usage/context mapping, failing-result capture, and unreadable-line counting.

Checklist

  • Tests cover behavior changes.
  • Docs were updated — docs/features/claude-code-provider.md covers the architecture, setup (bare-metal and Docker), the tool-visibility constraint, and the limitations below. It's listed in the documentation map, and agent.md no longer claims that every driver talks to a REST API over HTTP/SSE — it names the exception and lists the three new driver modules.
  • No compatibility shim was added for old pre-release behavior.
  • No secrets, local databases, uploads, or generated artifacts are included.

Scope and caveats

  • Single-operator. Driving a subscription programmatically to back an app is a grey area of Anthropic's terms. This is intended for an operator editing their own site, not multi-tenant serving. Documented as such.
  • Subscription rate limits apply.
  • Auth is machine-wide: one Claude account for the server, not per Instatic user. Per-user scoping is handled by the MCP connector's capability grant. I'm happy to add an in-app connect/re-auth flow — the CLI exposes auth status --json and an OAuth flow that proxies cleanly through a browser — if that's something you'd want in-tree.

I'm running this in production against a live site, and I'm happy to adjust anything here, including dropping it if a CLI-spawning provider isn't a direction you want to take.

cakelesscoder and others added 4 commits September 29, 2026 16:42
Runs the CMS agent on the machine's Claude subscription instead of a metered
API key. The driver spawns the headless `claude` CLI, points it at Instatic's
own MCP server, and streams its stream-json output back into the chat — so the
agent is real Claude Code editing the live workspace through the same editor
bridge the built-in agent uses.

It is the deliberate exception to the SDK ban and compliant by construction:
it imports nothing and spawns a binary. It is also the only option that works,
since the subscription cost model exists solely in the CLI auth path — the
Agent SDK bills API credits, which would defeat the purpose.

An internal MCP connector mints a per-user bearer token granted exactly the
chatting user's own capabilities, so the spawned CLI can do precisely what that
user could through the agent panel, never more.

Three modules by responsibility: what to ask for (claudeCode.ts), how to run it
(claudeCodeProcess.ts), how to read what came back (claudeCodeEvents.ts). The
last is pure and synchronous, so the translation layer is unit-testable from
recorded CLI output — which is what the new tests do.

A subprocess has many more ways to go quiet than an HTTP call, and all of them
look identical from the composer: a few seconds of output, then nothing. Each is
terminal and named — a toolless turn, a wedged child, a session mismatch, a rate
limit, CLI format drift, an abandoned stream. The toolless case is the subtle
one: `--tools ''` strips built-ins, but the CLI defers a large MCP toolset behind
the `ToolSearch` built-in, so the combination leaves the model with zero callable
tools. It then narrates a tool call in prose and ends the turn `success`.
ENABLE_TOOL_SEARCH=false fixes the cause; the init tool count is checked anyway,
so a future CLI that ignores the env var fails loudly instead of bluffing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude Code authenticates outside Instatic, so its connect form has no secret to
collect — but the flow assumed every provider hands over either an API key or an
endpoint URL.

Rather than special-case one provider id in ProvidersTab, make the idea explicit
in the provider catalogue: `credentialLess` says the provider brings its own
auth, `credentialLessHint` is shown where the inputs would have been, and
`sentinelBaseUrl` is the inert value stored on the row so it still satisfies the
API's auth-shape check. The form and the detail panel read those fields and know
nothing about Claude Code specifically, so any future provider that authenticates
elsewhere gets the same treatment.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The base image has no `claude` binary. Dockerfile.claude-code overlays it with
the self-contained installer output at /opt/claude/claude and bakes
INSTATIC_CLAUDE_BIN, kept separate from the main Dockerfile so the base image
keeps tracking upstream cleanly.

compose.claude-code.yml supplies subscription auth as CLAUDE_CODE_OAUTH_TOKEN
(from `claude setup-token`). No home-directory mount, and no API key: the driver
strips ANTHROPIC_API_KEY but preserves this var, so the CLI cannot silently fall
back to metered billing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Architecture and request flow, why a subprocess rather than the Agent SDK, setup
for both bare-metal and Docker, the tool-visibility constraint that makes
ENABLE_TOOL_SEARCH=false load-bearing, the failure modes the driver now names,
and the scope limits — single-operator, subscription rate limits, and auth being
machine-wide rather than per user.

Also wires it into the existing docs: the feature appears in the documentation
map, and agent.md no longer claims that *every* driver talks to a REST API over
HTTP/SSE — it names the one deliberate exception and lists the three new driver
modules alongside the HTTP ones.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant