Skip to content

About

Command Code API 反代代理,兼容 OpenAI 与 Anthropic 接口 | Reverse proxy exposing Command Code API as OpenAI- and Anthropic-compatible endpoints

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

 
 

Repository files navigation

Command Code Proxy

中文文档

A multi-account Command Code reverse proxy with a management console and OpenAI / Anthropic compatible endpoints.

Built by analyzing official CLI network traffic to accurately replicate the Command Code API request protocol, including device-fingerprint and lifecycle pre-requests.

Features: OpenAI Chat Completions / Responses API (/v1/responses) + Anthropic Messages API | Streaming & non-streaming | Tool calling (tool_use) | Multimodal image input | Reasoning effort | Dynamic model list | Cache hit metrics | Device fingerprint disguise (per-key, auto-refresh) | x-api-key auth (Anthropic SDK) | Client disconnect detection with upstream abort | Zero-output → 429 auto-retry | Consecutive timeout → 429 auto-retry | Privacy-aware logging

Community: Linux.do — a friendly Chinese tech community.

Current console: hybrid protocol roots

Management listens on 3050; private inference listens on 3051. Authenticate with a console-issued client API key, not an upstream user_ credential.

Account groups and outlet keys

Use 分组管理 to create or rename account pools. Select 所属分组 when creating or editing an upstream account or client API key. Each client key can list models and send requests only through accounts in its group, for both /v1 and /provider/v1. Existing account/client model allowlists, priority, weights, concurrency, cooldowns and native Responses permissions still apply within the group. An empty or unavailable group never falls back to another group.

The version 2 SQLite migration atomically assigns all existing accounts and clients to 默认分组 (default). Credentials, account settings, statistics, sessions and existing key hashes remain unchanged; clients do not need to replace their keys. The default group cannot be deleted, and other groups must be empty before deletion. Moving a client also moves the scope of every key in its rotation grace period. New accounts and clients default to this group when group_id is omitted. Authenticated group management uses /command/api/groups; account/client create and update requests accept group_id.

Back up SQLite on the server before deployment. Older versions ignore group bindings: after configuring separate groups, do not downgrade to a version without groups while isolated keys remain enabled.

  • /v1 preserves existing Chat, Messages and translated Responses via the CLI /alpha/generate backend.
  • /provider/v1/responses forwards native Responses to the official Provider API. /provider/v1/chat/completions still accepts and returns Chat format through the existing compatibility backend.
  • Enable provider_responses_enabled only on accounts whose Provider API permission has been verified, then refresh their catalogs. Go plans have no Provider API access. Native Responses additionally requires the model's reported supported_endpoints and both account/client ACLs.
  • Each root's /models describes its effective endpoints and responses_backend (provider or cli). Native failures never fall back to the CLI backend, and Chat is not forced into Responses. Replacing credentials disables native access until it is verified again.
  • Native input, images, tools, opaque reasoning, SSE deltas and usage are preserved. Requests always use store:false; server-side history, background requests and remote MCP tools are rejected. Send full history and execute MCP tools locally. Missing terminal events and failed streams are not counted as successful; Retry-After is retained and requests are not replayed.
  • Management generation tests default to Chat and can explicitly select native Responses.

The single-file, upstream-key and CLI-fingerprint instructions below describe the older standalone backend. For the console, use server.mjs, the management UI and these two inference roots.

Quick Start

npm start        # Start (the repo ships with config.json listening on http://0.0.0.0:3050)
npm run dev      # Watch mode (auto-reload on file changes)

API Key is passed via the Authorization request header (or x-api-key for Anthropic SDKs) — no need to store it in config files. Key must start with user_ (automatically matched with any prefix, e.g. Bearer token_user_xxx):

curl http://127.0.0.1:3050/v1/chat/completions \
  -H "Authorization: Bearer user_xxxxxxxxx" \
  -H "Content-Type: application/json" \
  -d '{"model":"deepseek/deepseek-v4-flash","messages":[{"role":"user","content":"hi"}]}'

File Structure

commandcode/
├── config.json           # Port / log path etc.
├── LICENSE               # MIT License
├── package.json          # npm start / npm run dev
├── proxy.mjs             # Single-file proxy core (~1900 lines)
├── Dockerfile            # Container build (node:22-alpine)
├── docker-compose.yml    # Container orchestration
├── .dockerignore         # Build context exclusions
├── .github/
│   └── workflows/
│       └── docker-publish.yml  # release branch / v* tags → GHCR multi-arch (latest + release)
├── captured-requests/    # Captured CLI traffic (protocol analysis reference)
├── README.md             # This document (English)
└── README_zh.md          # Chinese documentation

Configuration

config.json

Field Default Description
port 3000 Listen port (repo config.json ships with 3050)
host 0.0.0.0 Listen address
apiBase https://api.commandcode.ai CC API base URL
projectSlug cc-proxy x-project-slug header
apiKey "" Optional fallback API key (requests can also send it via header)
logFile "" Log file path (empty = console only)
logLevel info Log level
useProviderModels true Dynamically fetch model list from Provider API
modelRefreshIntervalMs 300000 Model list cache refresh interval (5 min)
zdr false Request ZDR-only routing from Command Code
cliMode agent Envelope mode. Upstream enum: agent / learning / custom-agent / custom-agent-create / title-gen / tool-desc / compact / vision
cliSessionMode interactive mode inside the lifecycle metadata (a different enum: interactive / non-interactive)
fingerprintSalt "" Salt for the device fingerprint — use it to rotate the whole fleet's identity (one key still always reports one device)
deviceProjectDir "" Faked project directory (empty = built-in C:\Users\dev\projects\app); changing it gives every account a different device
emptySystemPlaceholder true Send a space placeholder when there is no system prompt, preventing upstream from injecting its ~7.5K-token default (#17)

Environment Variables

Variable Default Description
PORT 3000 (shipped config.json uses 3050) Listen port → port
HOST 0.0.0.0 Listen address → host
CC_API_BASE https://api.commandcode.ai Upstream base URL → apiBase
CC_UPSTREAM_PROXY (unset) Global fallback for accounts without a dedicated upstream proxy; see "Upstream proxy" below → upstreamProxy
PROJECT_SLUG cc-proxy x-project-slug → projectSlug
LOG_FILE empty Log file → logFile (synchronous writes, see Other notes)
CC_USE_PROVIDER_MODELS true Fetch the model list dynamically → useProviderModels
CMD_ZDR off 1 enables ZDR-only routing → zdr
CC_CLI_MODE agent Envelope mode → cliMode
CC_CLI_SESSION_MODE interactive Lifecycle metadata mode → cliSessionMode
CC_FINGERPRINT_SALT empty Fingerprint salt → fingerprintSalt
CC_DEVICE_PROJECT_DIR empty Faked project directory → deviceProjectDir
CC_EMPTY_SYSTEM_PLACEHOLDER true Space placeholder for a missing system prompt; false disables → emptySystemPlaceholder
CC_MAX_BODY_MB 100 Max request body size in MB; oversized requests get 413
CC_MAX_TOOL_IMAGE_MB 6 Total per-request budget for tool screenshots (base64); older ones become a placeholder; 0 disables. See Tool screenshot budget
CC_STREAM_IDLE_MS 30000 Streaming upstream read idle timeout; see Upstream idle timeouts
CC_NONSTREAM_IDLE_MS 90000 Non-streaming upstream read idle timeout
CC_UPSTREAM_RETRY_MAX 2 Retries for upstream disconnects before any byte is written downstream; 0 disables; see Upstream transient retry
CC_UPSTREAM_RETRY_BASE_MS 400 Backoff base in ms; the actual delay is base × attempt number
CC_MAX_INFLIGHT 0 (unlimited) In-process request cap; over-limit returns 503; see In-flight cap
CC_CLIENT_DRAIN_TIMEOUT_MS unset (disabled) Drop the client once downstream backpressure blocks longer than this; see Stalled clients
CC_KEEPALIVE_TIMEOUT_MS 65000 Backend keep-alive timeout (headersTimeout is set to +1s automatically). Must be larger than the reverse proxy's keepalive_timeout — see keep-alive ordering

When enabled, the proxy sends x-cmd-zdr: 1 on Command Code generation requests and the fingerprint/lifecycle initialization requests. It does not add the header to the npm version check or the proxy's /provider/v1/models catalog request. This requests Command Code's ZDR-only routing; the upstream service remains the authority for actual retention and provider availability.

Request body limit: independent of config.json — requests larger than 100 MB are rejected with HTTP 413 (the connection is kept alive and drained, not reset). Override with CC_MAX_BODY_MB (positive integer, unit: MB).

⚠️ Memory amplification: a request body exists in several copies before it reaches upstream; measured peak ≈ body size × 5.1–7.4 (7 MB → +52 MB, 20 MB → +116 MB, while a request rejected with 413 costs only ×1.05). The default CC_MAX_BODY_MB=100 therefore implies up to ~550 MB for a single request, and that limit is per-request, not global. See Memory & Deployment.

Tool screenshot budget

Images inside tool results are sent upstream as separate image blocks (an input_image inside function_call_output.output is no longer JSON.stringify-ed into the tool-result text). Base64 is extremely expensive once upstream tokenizes it as text — a single 2.76 MB Codex Desktop screenshot measured ≈ 1.92M tokens, blowing straight through a 1M window:

400 This model's maximum context length is 1048576 tokens. However, you requested
    1986800 tokens (1922800 in the messages, 64000 in the completion)

Images are still re-sent in full every turn: one real session carried 11 screenshots ≈5.5 MB of base64, which — with the memory amplification above (×5.1–7.4) — pushed the proxy to anon-rss 532 MB and triggered a global OOM on a 1 GB box. CC_MAX_TOOL_IMAGE_MB (default 6) keeps the newest images within budget (at least one) and replaces the rest with [older tool screenshot omitted: image budget exceeded], so the model knows an image was dropped instead of assuming none ever existed. Only tool screenshots are affected; images you attach yourself are untouched.

Raise the budget if the model really needs every screenshot, but size memory as budget × in-flight × 5–7 (see in-flight cap).

Upstream proxy (upstreamProxy / CC_UPSTREAM_PROXY)

In the console's upstream account editor, assign one proxy URL per account. http://, https://, and socks5:// are supported, with optional user:pass@ authentication. A blank value on edit keeps the existing proxy; the remove checkbox clears it. URLs are encrypted at rest, and management responses expose only the scheme, host and port.

The account list's exit-IP check follows Sub2API's proxy probe: it requests fixed IP services through that account's proxy and displays the observed exit IP, region and latency. Results are saved with the account and cleared when its proxy changes. This confirms the network route; use the account generation test to verify that a model accepts the exit region.

The global setting below is used only for accounts without a dedicated proxy:

{ "upstreamProxy": "http://127.0.0.1:7890" }
CC_UPSTREAM_PROXY=http://127.0.0.1:7890 npm start
  • Applies to generation, fingerprint/lifecycle pre-requests, model catalogs and billing queries.
  • Does not touch the local listener, /health, or the npm version check.
  • HTTP/HTTPS proxies use CONNECT tunnels; SOCKS5 uses an optional username/password handshake. No new dependency is required.
  • Each upstream request opens its own tunnel connection. TLS is end-to-end: the certificate is validated against the target hostname, never against the proxy.
  • Routing the fingerprint/lifecycle pre-requests through the same proxy matters: if they went out direct while generation went through the proxy, one account would register from two different IPs — exactly the inconsistency you are trying to avoid.
  • Credentials in the proxy URL are never logged: only the scheme, host and port show up.

Node's built-in fetch does not read HTTPS_PROXY/HTTP_PROXY. The official env-var route requires Node ≥ 22.21 / 24.5 plus NODE_USE_ENV_PROXY=1; this option works without either.

API Endpoints

POST /v1/chat/completions

OpenAI Chat Completions compatible. Supports streaming, non-streaming, tool calling, multimodal image input, and reasoning effort.

Request parameters:

Parameter Required Description
model Yes Model ID (see model list)
messages Yes Conversation messages, supports system/user/assistant/tool roles
max_tokens No Max tokens to generate (default 64000)
stream No SSE streaming (default false)
temperature No Sampling temperature (0-2)
reasoning_effort No Reasoning intensity: low/medium/high/max
tools No Tool definitions (OpenAI function calling format)
tool_choice No Tool selection strategy
parallel_tool_calls No Allow parallel tool calls

Simple request:

{
  "model": "deepseek/deepseek-v4-flash",
  "messages": [{ "role": "user", "content": "hello" }],
  "stream": true
}

Multimodal image input (vision model required):

{
  "model": "xiaomi/mimo-v2.5",
  "messages": [{
    "role": "user",
    "content": [
      { "type": "text", "text": "Describe this image" },
      { "type": "image_url", "image_url": { "url": "data:image/jpeg;base64,..." } }
    ]
  }]
}

Tool calling:

{
  "model": "deepseek/deepseek-v4-flash",
  "messages": [...],
  "tools": [{
    "type": "function",
    "function": { "name": "get_weather", "description": "...", "parameters": {...} }
  }],
  "tool_choice": "auto"
}

Streaming response (SSE):

data: {"id":"chatcmpl-xxx","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"role":"assistant","reasoning_content":"thinking..."}}]}

data: {"id":"chatcmpl-xxx","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"Hello"}}]}

data: {"id":"chatcmpl-xxx","object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"stop"}],"usage":{"prompt_tokens":10,"completion_tokens":20,"total_tokens":30,"prompt_tokens_details":{"cached_tokens":8}}}

data: [DONE]

Non-streaming response (with cache hits):

{
  "id": "chatcmpl-xxx",
  "object": "chat.completion",
  "created": 1234567890,
  "model": "deepseek/deepseek-v4-flash",
  "choices": [{
    "index": 0,
    "message": {
      "role": "assistant",
      "content": "Hello!",
      "reasoning_content": "The user said hello, I should respond."
    },
    "finish_reason": "stop"
  }],
  "usage": {
    "prompt_tokens": 7558,
    "completion_tokens": 42,
    "total_tokens": 7600,
    "prompt_tokens_details": { "cached_tokens": 7552 }
  }
}

POST /v1/messages

Anthropic Messages API compatible endpoint. Supports streaming, non-streaming, and tool calling.

Request body:

{
  "model": "claude-sonnet-4-6",
  "max_tokens": 1000,
  "system": "You are a helpful assistant.",
  "messages": [
    { "role": "user", "content": "hello" }
  ],
  "stream": true
}

Anthropic protocol conversion (automatic):

Concept Anthropic Format Conversion
System prompt Top-level system field Auto-converted to OpenAI system message
Message content content array (text/tool_use/tool_result) Auto-mapped to corresponding roles
Tool results tool_result blocks in user messages Auto-converted to role: "tool"
Tool definitions input_schema Auto-mapped to parameters
tool_choice {type:"auto"/"any"/"tool"} any→required, tool→function object
Reasoning thinking.budget_tokens Auto-mapped to reasoning_effort (≥10000→high, ≥5000→medium, ≥2000→low)
Stop reason end_turn/max_tokens/tool_use Auto-mapped to stop/length/tool_calls
Token usage input_tokens/output_tokens + cache Passed through, cache fields mapped to Anthropic format

Streaming response (SSE, Anthropic format):

event: message_start
data: {"type":"message_start","message":{"id":"msg_xxx","type":"message","role":"assistant","content":[],"model":"...","usage":{"input_tokens":0,"output_tokens":0}}}

event: content_block_start
data: {"type":"content_block_start","index":0,"content_block":{"type":"text","text":""}}

event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"Hello"}}

event: content_block_stop
data: {"type":"content_block_stop","index":0}

event: message_delta
data: {"type":"message_delta","delta":{"stop_reason":"end_turn"},"usage":{"output_tokens":10,"cache_read_input_tokens":0,"input_tokens":100}}

event: message_stop
data: {"type":"message_stop"}

Non-streaming response:

{
  "id": "msg_xxx",
  "type": "message",
  "role": "assistant",
  "model": "deepseek/deepseek-v4-flash",
  "content": [{ "type": "text", "text": "Hello!" }],
  "stop_reason": "end_turn",
  "stop_sequence": null,
  "usage": {
    "input_tokens": 7558,
    "output_tokens": 42,
    "cache_read_input_tokens": 7552,
    "cache_creation_input_tokens": null
  }
}

POST /v1/responses

OpenAI Responses API (what Codex and the newer OpenAI SDKs speak).

The request side is translated: input (message array; items may omit type), instructions, max_output_tokens, temperature, top_p, reasoning, tools and tool_choice all map onto the CC envelope; the response comes back in Responses shape (object: "response", output array, usage, status). Streaming is SSE: response.created / response.in_progress / response.output_item.added|done / response.content_part.added|done / response.output_text.delta|done / response.reasoning_summary_text.delta|done / response.function_call_arguments.delta|done, terminated by response.completed (response.incomplete when truncated by max_output_tokens, response.failed on error).

  • Stateless: previous_response_id is not supported and answers 400 — send the full input every turn (the proxy stores no conversation history).
  • Errors use the Responses shape: {"error":{"message":...,"type":...}}.
  • Shares the same upstream call path, cache breakpoints and idle watchdog as /v1/chat/completions.
  • First-token silence and keep-alive: response.created / response.in_progress are emitted immediately once upstream returns 200, and a : keepalive SSE comment follows every 5 s while waiting. Time-to-first-token for a reasoning model with a large prompt measured 15–40 s; that silent window used to put zero bytes on the wire, so intermediate layers (EdgeOne's origin idle timeout measured ~15 s) or client first-byte timeouts cut the connection — visible as nginx 499 with body_bytes_sent=0 and a client retrying every 15 s. SSE comments must be ignored by clients per spec (/v1/messages uses event: ping, but Responses has no ping event and an unknown event type risks strict-parser errors).
curl http://127.0.0.1:3050/v1/responses \
  -H "Authorization: Bearer user_xxxxxxxxx" -H "Content-Type: application/json" \
  -d '{"model":"deepseek/deepseek-v4-flash","input":[{"role":"user","content":[{"type":"input_text","text":"hi"}]}]}'

GET /v1/models

Returns available model list. Fetched dynamically from Provider API (5 min cache), falls back to hardcoded list on failure.

GET /health

Health check. Returns OK.

Error Codes

Produced by the proxy itself:

HTTP When
400 Body is not valid JSON, input is empty, or an unsupported previous_response_id is used
401 API key missing / malformed (must start with user_; sent via Authorization: Bearer or x-api-key)
404 Unknown path
413 Body exceeds CC_MAX_BODY_MB (connection kept alive and drained, not reset)
429 Zero output tokens, stream idle timeout (30s streaming / 90s non-streaming), or an upstream rate-limit mapping — all carry Retry-After so SDKs back off; after 3 consecutive timeouts a "reduce context" hint is returned
502 CC upstream error (connection-level failures such as fetch failed also land here)
503 CC_MAX_INFLIGHT is set and the in-flight cap is exceeded (type: server_busy)

Upstream CC status mapping (CC_STATUS_MAP; anything unlisted becomes 502 upstream_error):

Upstream Downstream
400 → 400 invalid_request_error 401 → 401 authentication_error
402 → 429 rate_limit_error (payment failures are treated as rate limits) 403 → 401 authentication_error
404 → 404 not_found 422 → 400 invalid_request_error
429 → 429 rate_limit_error (with retry_after: 30) 500 / 502 → 502 upstream_error
503 → 503 temporarily_unavailable anything else → 502 upstream_error

The machine-readable classification from the upstream error body (error.code, e.g. BAD_REQUEST / USAGE_EXCEEDED) is passed through as error.code.

Model List

The proxy returns a live model list via GET /v1/models. Below are common models for reference; the actual list depends on the live API response — see Command Code Pricing for plan details.

Common Models

Model ID Provider
claude-sonnet-4-6 / claude-opus-4-8 / claude-opus-4-7 / claude-haiku-4-5-20251001 Anthropic
gpt-5.5 / gpt-5.4 / gpt-5.4-mini / gpt-5.3-codex OpenAI
deepseek/deepseek-v4-pro / deepseek/deepseek-v4-flash DeepSeek
moonshotai/Kimi-K2.6 / moonshotai/Kimi-K2.5 Kimi
zai-org/GLM-5.1 / zai-org/GLM-5 GLM
MiniMaxAI/MiniMax-M3 / MiniMaxAI/MiniMax-M2.7 / MiniMaxAI/MiniMax-M2.5 MiniMax
Qwen/Qwen3.7-Max / Qwen/Qwen3.6-Max-Preview / Qwen/Qwen3.6-Plus Qwen
stepfun/Step-3.7-Flash / stepfun/Step-3.5-Flash Step
xiaomi/mimo-v2.5-pro / xiaomi/mimo-v2.5 Xiaomi (image input supported)
google/gemini-3.5-flash / google/gemini-3.1-flash-lite Gemini

⚠️ Some models (e.g. deepseek-v4-flash, claude-sonnet-4-6) do not support image input. Use xiaomi/mimo-v2.5, Kimi-K2.5, or other vision models for multimodal.

Integration Examples

Python (OpenAI SDK)

from openai import OpenAI

client = OpenAI(
    api_key="user_xxxxxxxxx",
    base_url="http://127.0.0.1:3050/v1",
)

response = client.chat.completions.create(
    model="deepseek/deepseek-v4-flash",
    messages=[{"role": "user", "content": "hello"}],
    stream=True,
)
for chunk in response:
    print(chunk.choices[0].delta.content or "", end="")

cURL

curl http://127.0.0.1:3050/v1/chat/completions \
  -H "Authorization: Bearer user_xxxxxxxxx" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek/deepseek-v4-flash",
    "messages": [{"role": "user", "content": "hello"}],
    "stream": true
  }'

Cursor

Add a Custom Provider in Cursor settings:

  • API Base URL: http://127.0.0.1:3050/v1
  • API Key: user_xxxxxxxxx
  • Model: Choose from the model list

Anthropic (Python SDK)

import anthropic

client = anthropic.Anthropic(
    api_key="user_xxxxxxxxx",
    base_url="http://127.0.0.1:3050",
)
message = client.messages.create(
    model="deepseek/deepseek-v4-flash",
    max_tokens=1000,
    system="You are helpful.",
    messages=[{"role": "user", "content": "hello"}],
)
print(message.content[0].text)

The Anthropic SDK authenticates via the x-api-key header — supported by the proxy natively (no Authorization header needed).

OpenCode

{
  "provider": "openai-compatible",
  "baseUrl": "http://127.0.0.1:3050/v1",
  "apiKey": "user_xxxxxxxxx"
}

Anti-Detection

Aligned line-by-line against the official npm package source (command-code@1.53.1; dist/cli.mjs is minified but not obfuscated). Newer npm releases only raise a drift warning — the proxy never silently bumps the version it claims:

Mechanism Implementation
Device Fingerprint POST /alpha/fingerprint/record before first request per key; signal values (Windows MachineGuid shape, real-shaped MACs, DESKTOP-xxxxxx hostname) are derived deterministically from the API key and hashed exactly like the CLI, so one key always reports the same device — across restarts, memory reclamation and multiple instances (bulk reset via CC_FINGERPRINT_SALT)
Lifecycle Events POST /alpha/lifecycle-events (cli_session_exists, metadata {sessionId, cliVersion, mode, os}) sent in parallel with the fingerprint on key init
Per-Key Session One session per API key, 12h expiry + 1h random jitter
Version x-command-code-version reports the protocol version actually implemented (currently 1.53.1); newer npm releases only raise a drift warning, never a silent version bump
CLI Envelope 9 keys: config / memory / taste / skills / permissionMode / threadId / mode / promptCache / params
OpenTelemetry traceparent (W3C Trace Context)
Environment x-cli-environment: production, x-taste-learning: "false", User-Agent: cli
Project Slug x-project-slug = slugify(DEVICE_PROFILE.projectDir) — same source as config.workingDir (default C:\Users\dev\projects\app, override with CC_DEVICE_PROJECT_DIR)
Single Source of Device Truth Fingerprint / config.environment / config.workingDir / x-project-slug / lifecycle os all read one DEVICE_PROFILE (win32 / x64) — so they cannot contradict each other ("fingerprint says win32, environment says linux"), and the host's real platform, Node version and cwd are never handed upstream
Reasoning Effort reasoning_effort pass-through (low/medium/high/max)
Key Validation Regex user_[a-zA-Z0-9_-]+ on Authorization: Bearer or x-api-key, auto-cleans extra paths/prefixes, rejects sk-xxx format
Stream Timeout 30s streaming / 90s non-streaming → 429 with SDK auto-retry
Consecutive Timeout 3 consecutive timeouts before "reduce context" hint
Zero-Output Guard outputTokens=0 → 429 rate_limit_error (SDK auto-retry, anti false billing)
Upstream Abort AbortController on client disconnect + all error paths
Privacy Logging No API key fragments, no error bodies, no stack traces in logs

Protocol Details

CC API Request Structure

{
  "config": {
    "workingDir": "C:\\project",
    "date": "2026-06-07",
    "environment": "win32",
    "structure": [],
    "isGitRepo": false,
    "currentBranch": "",
    "mainBranch": "",
    "gitStatus": "",
    "recentCommits": []
  },
  "memory": null,
  "taste": null,
  "skills": null,
  "permissionMode": "standard",
  "params": {
    "model": "deepseek/deepseek-v4-flash",
    "messages": [...],
    "max_tokens": 64000,
    "stream": true,
    "reasoning_effort": "max"
  }
}

config.environment and config.workingDir come from DEVICE_PROFILE (not from the host), and skills is null (not an empty string).

Conditional fields: system (extracted from system messages), temperature, reasoning_effort, tools (mapped to CC input_schema format), tool_choice, parallel_tool_calls. When a prompt_cache_key is present (or the client already set cache_control), the breakpoint lands on the last system block — caching is prefix-based and system is that prefix.

CC API Image Message Format

The CLI sends images in this format:

{
  "role": "user",
  "content": [
    { "type": "image", "image": "data:image/jpeg;base64,..." },
    { "type": "text", "text": "What does this image say?" }
  ]
}

The proxy receives OpenAI image_url format and converts it to the above CC format transparently.

Docker Deployment

Pull from GHCR

GitHub Actions publishes multi-arch images (linux/amd64 + linux/arm64) to the GitHub Container Registry:

Tag Source Notes
:release release branch Tracks the release branch
:latest release branch or v* tag Same digest as :release. "Only updated on version tags" was the old behaviour — it left :latest users stuck on an old build where new endpoints 404'd (#28)
docker pull ghcr.io/maxeaglet/commandcode-proxy:release
docker run -d --name cc-proxy -p 3050:3050 -e PORT=3050 ghcr.io/maxeaglet/commandcode-proxy:release

The image is public — no login required to pull. After upgrading, confirm the digest actually changed (docker inspect --format '{{index .RepoDigests 0}}') rather than assuming your local cache is current.

Quick Start (docker compose)

docker compose up -d

The proxy will listen on http://0.0.0.0:3050. Set PROXY_PORT to customize the host port:

PROXY_PORT=13050 docker compose up -d

Build from Source

docker build -t commandcode-proxy:latest .
docker run -d -p 3050:3050 -e PORT=3050 commandcode-proxy:latest

Multi-Architecture Build

npm run docker:build:multi

Environment Variables

Only two are container-specific; everything else lives in the Environment Variables table above:

Variable Default Description
PORT 3050 Container listen port
PROXY_PORT 3050 Host port (compose only)

In-flight Cap (Optional)

Off by default (CC_MAX_INFLIGHT unset = no concurrency limit), so existing behaviour is unchanged.

This project is a pure proxy layer; concurrency control belongs downstream — use your reverse proxy for per-IP / per-key limits (see the limit_conn block in Memory & Deployment). This option is not a replacement for that; it only covers running without a reverse proxy (which both the Dockerfile and npm start invite) with an in-process, global-only guard:

CC_MAX_INFLIGHT=32 npm start    # at most 32 concurrent requests

Over the limit it returns 503 + Retry-After: 5 + type: server_busy — a shape the official OpenAI / Anthropic SDKs retry with backoff, instead of the client seeing a connection reset. /health and / are exempt so liveness probes and orchestrators never receive a 503 because business traffic is busy.

Why it exists: memory is in-flight × (0.13 MB + 5.5 × body_MB). CC_MAX_BODY_MB bounds only the per-request term; nothing bounds the multiplier — at the default 100 MB, N concurrent requests can cost N × 550 MB.

⚠️ Enabling this is not the same as being memory-safe: 32 × 550 MB still exceeds a small box. For a hard bound, lower CC_MAX_BODY_MB as well.

Upstream Idle Timeouts

Two upstream read idle watchdogs; on expiry the proxy returns 429 (with retry_after) so the SDK retries automatically:

Env var Default Applies to
CC_STREAM_IDLE_MS 30000 Streaming requests
CC_NONSTREAM_IDLE_MS 90000 Non-streaming requests

Semantics: they measure only the time spent waiting inside reader.read(), reset on every received chunk — not the total request duration. As long as upstream keeps emitting, the watchdog never fires, even for a request that has been running for tens of minutes.

The defaults differ from the official CLI, and that is a known trade-off (#19): the official CLI has no upstream idle timeout at all — deobfuscating command-code@1.50.0 shows every createApiClient({ baseUrl }) call site passes no timeout, and 700+ second stalls complete successfully. This proxy keeps 30 s to catch genuinely dead connections; the cost is that a reasoning model's long prefill/first-token stall can be killed.

If you see 429 Response timeout or zero output tokens where the log shows elapsedMs ≈ 30000 and bytesReceived = 0, the watchdog killed a healthy stall — raise it:

CC_STREAM_IDLE_MS=300000 npm start      # 5 minutes

⚠️ A false kill costs more than one failed request: the abort returns 429 + retry_after, the SDK retries automatically, and a retry resends the entire context — so each false kill re-pays the full prefill on long conversations.

Upstream Transient Retry

The CC upstream sometimes kills a connection mid-stream at peak hours (peer RST/FIN), which surfaces in Node as TypeError: terminated. Before this change the proxy handed that straight to the client as 502 {"error":{"message":"Upstream error: terminated","type":"proxy_error"}}.

As long as no byte has been written downstream yet, the request never started from the client's point of view — so the proxy can absorb the blip internally instead of making the client eat a 502 and resend its whole context.

  • Retry requires all four: ① the upstream did not finish — either a transport-level drop (terminated / ECONNRESET / ECONNREFUSED / EPIPE / ETIMEDOUT / UND_ERR_SOCKET / socket hang up / other side closed / fetch failed), or a clean peer FIN that carried no finish event at all (the same class of truncation as an RST); ② nothing written downstream yet (streaming: no header/event emitted; non-streaming: headersSent still false); ③ the client is still connected; ④ the retry cap is not reached.
  • Once any header or event has gone downstream, the proxy never retries — the semantics are already committed and a retry would duplicate text.
  • STREAM_IDLE_TIMEOUT (the 429 "reduce your context" signal) is deliberately passed through and is never retried.
  • If an upstream error event (429 / 503 …) has already been parsed, the proxy does not retry, and if the connection then drops it still surfaces that semantic error instead of overwriting it with a transport 502 (reporting "upstream at capacity" as "the proxy broke" would be misleading).
  • A client that disconnects during the backoff abandons the retry — the client is gone, another upstream call would only burn quota.
  • Retries are invisible to the client: it sees a single 200 whose body comes from the attempt that succeeded.
Env var Default Description
CC_UPSTREAM_RETRY_MAX 2 Max retries (3 attempts in total); 0 disables the behaviour
CC_UPSTREAM_RETRY_BASE_MS 400 Backoff base in ms; the actual delay is base × attempt number

Logs to look for: Upstream stream terminated before first byte - retrying / Upstream error before first byte - retrying / Upstream stream ended incomplete before first byte - retrying (a retry happened; attempt / maxAttempts / cause or reason are in the structured fields), Upstream retry recovered (the retry actually delivered a normal response) and Upstream retry abandoned (client disconnected during backoff). The startup banner's upstreamRetry field shows the effective values.

CC_UPSTREAM_RETRY_MAX=0 npm start        # disable retries (behaviour reverts to before this change)
CC_UPSTREAM_RETRY_BASE_MS=800 npm start  # wider backoff (default 400ms)

Scope: the retry loop covers both the streaming and non-streaming paths of /v1/chat/completions. /v1/messages and /v1/responses have a different structure and are not covered by this change (drops are still reported as before).

Memory & Deployment

Measurements reproduced from issue #20 (Node v24, loopback mock upstream).

Rule of thumb for per-request memory:

RSS ≈ 70 MB + in-flight × (0.13 MB + 5.5 × body_MB)

Streaming responses apply backpressure

When res.write() returns false (the socket write buffer passed highWaterMark), reading from upstream pauses, so the response no longer accumulates unbounded in memory:

Scenario (200 MB upstream stream, client stops reading after sending) Peak RSS delta
Before the fix +586 MB (66 → 652 MB)
After the fix +4 MB (backpressure propagates upstream, which stalls after ~8 MB)

This is not only a hostile-client problem — throttled/mobile links, a client blocked on tool execution, or a client that already gave up but whose TCP stack has not sent RST all trigger it.

Request body is ~5.5× its size

The body exists in several copies before being forwarded: chunks[] / Buffer.concat / utf8 string / JSON.parse object tree / buildCcRequest second object tree / JSON.stringify serialized body.

body cap peak delta status
7 MB 100 MB +52 MB (7.4×) 200
20 MB 100 MB +116 MB (5.8×) 200
20 MB 8 MB +21 MB (1.05×) 413

At startup a warn is logged when the implied worst case is ≥ 500 MB. The limit is per request and the proxy does no in-flight limiting of its own — a public deployment must add both at the reverse proxy.

Suggested nginx front

Rejecting in nginx means the body is never materialized in the Node process at all:

map $http_authorization $cc_key { default $http_authorization; "" $http_x_api_key; }
map "" $cc_global_key { default "global"; }

limit_conn_zone $binary_remote_addr zone=cc_ip:10m;
limit_conn_zone $cc_key             zone=cc_key:10m;
limit_conn_zone $cc_global_key      zone=cc_global:10m;

location /v1/ {
    client_max_body_size 4m;   # must be <= CC_MAX_BODY_MB
    limit_conn cc_ip     8;
    limit_conn cc_key    4;
    limit_conn cc_global 32;   # this *is* the memory ceiling
    limit_conn_status 429;
    proxy_pass http://127.0.0.1:3050;
    proxy_http_version 1.1;
    proxy_set_header Connection "";
    proxy_buffering off;
    proxy_read_timeout 300s;   # must exceed the 30s stream idle timeout
}

⚠️ keep-alive ordering: proxy_set_header Connection "" above keeps connections to the backend alive, so the reverse proxy's upstream { keepalive_timeout ...; } must be smaller than the backend's CC_KEEPALIVE_TIMEOUT_MS (default 65s). Get it backwards and nginx reuses a connection the backend has already FIN'd, then hits EPIPE while writing the POST body (visible in the nginx log as sendfile() failed (32: Broken pipe)); POST is not idempotent and nginx does not retry it by default — the client gets a bare 502.

Stalled clients (neither reading nor disconnecting)

Once backpressure is in effect, a client that neither reads nor disconnects keeps its request and upstream connection alive indefinitely. Measured residual cost:

Stalled connections RSS delta Upstream connections held
1 +5 MB 1
10 +45 MB 10
50 +248 MB 50 (held forever)

The cost is bounded, does not leak, and is reclaimed as soon as the client disconnects (the RSS curve stays flat) — but the number of connections itself is unbounded.

This is left unhandled by default, because a stalled client is indistinguishable at the protocol level from a legitimate client blocked on tool execution, and the official CLI has no upstream idle timeout at all (see #19) — adding an aggressive timeout would repeat the mistake of killing healthy requests.

To cap it, opt in:

# only drop a client that has been blocked downstream for over 60s;
# a client that keeps making drain progress never triggers this
CC_CLIENT_DRAIN_TIMEOUT_MS=60000 npm start

Measured with the timeout enabled (50 stalled connections): upstream connections held goes from 50 (forever) → 0, with no post-drop draining of upstream.

A more robust cap still belongs at the reverse proxy (limit_conn), since only it knows how much concurrency a given deployment can afford.

Other notes

  • logFile uses appendFileSync — synchronous writes on the event loop. Under public load they serialize the loop; prefer leaving it empty and collecting stdout.
  • systemd guard rails: set MemoryMax= and NODE_OPTIONS=--max-old-space-size= so an overshoot kills the proxy, not sshd/nginx.
  • Multi-account + multiple instances: sessionStore / keyStateStore are per-process Maps, so the same API key served by two instances gets two different sessions and two different device fingerprints — upstream sees one account on multiple machines. Scale with consistent hashing on the API key (hash $cc_key consistent), not round-robin.

Disclaimer

This project is for educational and research purposes only.

  • Unofficial: This project is not affiliated with Command Code in any way.
  • Personal Use: Users assume all responsibility. Please comply with the Command Code Terms of Service.
  • API Key: This project does not collect, upload, or leak your API Key. The key is sent per request via the Authorization: Bearer <key> or x-api-key header and is never logged; an optional apiKey field in config.json serves only as a local fallback and never leaves your machine.
  • Compliance: The protocol is based on passive observation of local CLI network traffic. No unauthorized access, cracking, or tampering of the server has been performed.
  • Account Risk: Keep usage frequency consistent with normal CLI usage. Extremely high concurrent calls may trigger risk controls.

Development

# Start with watch mode (auto-reload on file changes)
npm run dev

About

Command Code API 反代代理,兼容 OpenAI 与 Anthropic 接口 | Reverse proxy exposing Command Code API as OpenAI- and Anthropic-compatible endpoints

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages