Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,8 @@ All notable changes to this project. Format: [Keep a Changelog](https://keepacha

- **fix(commands): rename `/review` → `/review-branch` so it stops shadowing the bundled `/code-review`.** CC 2.1.223 made `/review` the alias of the bundled `/code-review` — Claude Code's multi-agent reviewer, including the cloud `ultra` mode. A project skill of that name wins it (verified headlessly on 2.1.241: a project skill named `review` ran for `/review`, and the same held for `plan` against the built-in `/plan`), so **every scaffolded project was silently hiding the better built-in behind this simpler single-pass skill** — overlap that turned into a real capability loss the day the alias shipped. The skill moves to `templates/commands/review-branch/` with `name: review-branch`, and its description now positions it honestly ("a quick single-pass review; Claude Code's bundled `/code-review` is the deeper multi-agent one"). Both are reachable again. Updated across `config_schema.py`, `configure.py`'s pattern-integration map, `templates/INDEX.md`, the `/investigate` and `/plan-eng-review` cross-references, docs 02/03/05/09/10/11, README, and the example project. **Migration:** the configurator has no mechanism to delete a file it previously wrote, so an upgraded project keeps the old `.claude/skills/review/` alongside the new one — and the stale copy still shadows the alias. New `/verify-setup` **check 13** detects exactly that pair and tells the user to `rm -rf .claude/skills/review`. `/plan` is left alone deliberately: it shadows a built-in *command* rather than a bundled skill, and plan mode stays reachable via Shift+Tab, so it's a name clash rather than a lost capability — README now says so and points at the rename if you'd rather keep the shortcut.

- **docs: re-baseline the MCP context claims against tool search, and scope `/infinite` against dynamic workflows.** Two of the project's headline claims had been overtaken by Claude Code and were overstating what the modules buy. **MCP.** README claimed per-task profiles "drop a bloated 4-MCP baseline from ~49% context to under 5%", and `docs/04` asserted "every MCP tool is a chunk of JSON schema loaded at session start". Tool search defers MCP schemas by default (`alwaysLoad: true` is the opt-*out*), so the premise no longer holds. Measured rather than re-guessed — four local stdio servers advertising twelve tools each (48 total) against an otherwise identical one-turn session on CC 2.1.245: **26,665 tokens with no MCP servers, 27,361 deferred (+696, ~14/tool), 40,993 with `alwaysLoad: true` (+14,328, ~298/tool)** — deferral removes ~95% of the schema cost, and the numbers reproduced exactly across runs. The claim was also embedded in four *shipped* templates, which is worse than in the docs because it lands in every user's project: `check-context/SKILL.md` (its budget guardrails and the "MCP > 10%" flag), `claude-ctx.sh`'s rationale comment, `servers-cookbook.md` (which already explained deferral correctly a few sections earlier, so it contradicted itself), and the `mcp.minimal.json` profile comment. All corrected. Profiles are now documented for what they still genuinely buy — which servers *connect*: startup time, auth prompts, cold start, and the blast radius `--strict-mcp-config` enforces — and `docs/06` picks up the same correction. The dated `experiments-memory` example keeps its original result with a superseding **addendum** rather than a rewrite, because an experiment log records what was true when it ran. **`/infinite`.** Dynamic workflows now do staged, resumable, budgeted fan-out with structured output between stages; hand-rolled wave batching is the weaker instrument for that job. The skill opens with a decision table sending staged / merge-heavy / resumable work to a workflow, and keeps the one case it is genuinely good at — N variants of a single spec into disjoint slots with no cross-iteration coordination. README's module row and a new `docs/04` section say the same.

- **fix(portability): scaffolding from Windows emitted CRLF into every generated file; `--check` crashed; two hooks assumed tools that aren't there.** Four bugs, all invisible to a Linux-only CI. **(1) CRLF output.** `Path.write_text()` opens in text mode with `newline=None`, which translates `\n` to `os.linesep` — so a Windows scaffold wrote CRLF into all 13 hooks, `claude-ctx`, `CLAUDE.md` and `settings.json` (measured: 100% CRLF, zero bare LF). Harmless *on* Windows, where Git Bash strips the CR, but the shebang carried one, and `.claude/` is meant to be committed and shared (the generated CLAUDE.md says so, and `small-team` is a shipped persona) — so a teammate's Linux/macOS clone got `bad interpreter: /usr/bin/env bash^M` on every hook. All six write sites now route through a `write_text_lf()` helper using `open(..., newline="\n")` (`Path.write_text` only gained a `newline` parameter in 3.10; the project floor is 3.8). **(2) `--check` crashed on Windows** with `UnicodeEncodeError` on the `✓` under cp1252 — the exact command CONTRIBUTING gives new contributors. Python uses UTF-8 for a real Windows console, so this bit hardest when output was piped or redirected (CI logs, `cc-configure > setup.log`, an editor terminal). Both streams are now retagged UTF-8 with `errors="replace"` where the runtime allows it. **(3) The Stop hook's report depended on an unguarded `jq -n`.** jq isn't preinstalled on macOS, Windows, or most Linux distros; without it the checks ran, the hook exited 0, and Claude never learned anything had failed — the hook's entire purpose, lost silently. It now prefers jq, falls back to python3, and as a last resort prints the report to stderr with an explicit "NOT reaching Claude" warning. **(4) The package-availability gate disabled itself on stock macOS.** It bounded each probe with GNU `timeout`, absent there; the resulting rc=127 took the "probe inconclusive" branch, so the gate stood down on the first package for every macOS user — fail-open, so no false denials, but the advertised `safety` feature simply did not work, and said nothing. It now falls back to `gtimeout`, then to an unbounded probe. **Mitigation for what the configurator can't control:** the scaffold appends a `.gitattributes` block (`*.sh text eol=lf`, `claude-ctx`, `.claude/skills/**/scripts/*`) through the same managed-block append the `.gitignore` block uses — written when absent, line-level-unioned into a stale block, never touching the user's own rules. Git cannot record the *executable* bit that way, and `core.filemode=false` (the Windows default) drops it, so the generated `CLAUDE.md` § Repo bootstrap now tells Windows committers to run `git update-index --chmod=+x claude-ctx`. New `test/portability/` fixtures cover all of it by behavior, not by grepping source: every generated file is CRLF-free and every hook shebang is clean; the `.gitattributes` block is written, idempotent and union-safe; and the two hooks are run under a PATH that hides jq, python3 and timeout. README's runtime-dependency list now names jq and the timeout fallback; the platform table reflects what CI actually exercises. The `python-uv-fastapi` example and all five persona snapshots are regenerated (`.gitattributes` is new output).

- **fix(install): the `ln -sf` shortcut silently froze on Windows; PATH-independent shim, honest prerequisites.** Under Git Bash/MSYS, `ln -s` **copies** unless `MSYS=winsymlinks:nativestrict` is set *and* the user has Developer Mode or admin — so `~/.local/bin/cc-configure` became a snapshot of configure.py taken on install day. `git -C ~/.cc-configurator pull` updated the clone while the command kept running the old copy, with nothing to indicate it. The installer now verifies the link actually resolved (`[ -L ]`) and otherwise writes a two-line `exec python3 <clone>/configure.py "$@"` shim, which tracks the clone on every platform; both checks live in the `if` condition, because under `set -e` a trailing `[ -L … ] && link_ok=1` in a then-block would abort the installer on precisely the platform the fallback exists for. Also: `git` is now checked up front (it was used before the first guard), the missing-python3 hint is OS-aware instead of always saying `apt install python3`, and the closing usage banner no longer advertises `--preset aggressive` and the legacy `commands-core` / `token-efficiency-pro` module IDs — both deprecated and slated for removal in v3.0.
Expand Down
4 changes: 2 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -131,8 +131,8 @@ There are 11 modules; legacy IDs (`commands-core`, `agents`, `token-efficiency-p
| **git-workflow** | PostToolUse formatter on Write/Edit, Stop hook running typecheck / lint / tests. |
| **token-efficiency** | Path-scoped `.claude/rules/` starters + PreCompact snapshot hook. `tier` flag: `basic` (default) ships discipline rules + snapshot only; `pro` adds bash-output truncation hook + always-loaded discipline rules. |
| **commands** | Slash commands + agents + microbits. `subset` flag (linear ordering: `curated ⊂ full ⊂ rigorous`): **`curated`** = 3 essential skills (`/plan`, `/commit`, `/verify-setup`) + the `code-reviewer` agent. **`full`** (default) = 9 workflow skills (adds `/review-branch`, `/ship`, `/sync-docs`, `/check-context`, `/session-retro`, `/retrofit`) + 4 agents (`code-reviewer`, `test-runner`, `doc-writer`, `security-auditor`) + 4 discipline microbits (`/freeze`, `/unfreeze`, `/guard`, `/careful`) + the `microbit-enforcer.sh` PreToolUse hook. **`rigorous`** = `full` + `/investigate` + `/plan-eng-review`, the rigor skills that embed `templates/commands/_patterns/` cross-cutting blocks (confidence gate, independent verification, no-fix-without-investigation, AI-slop detection). **Naming note:** a project skill wins its name against Claude Code's built-ins (verified on CC 2.1.241). `/review-branch` is deliberately *not* called `review`, because CC 2.1.223 made `/review` the alias of the bundled `/code-review` and taking that name would hide Claude Code's multi-agent reviewer. `/plan` does still shadow the built-in plan-mode shortcut — plan mode remains reachable with Shift+Tab, so this is a name clash rather than a lost capability; rename the skill if you'd rather keep the shortcut. The `security-auditor` frontmatter wires Sonatype's dependency-management MCP (`https://mcp.guide.sonatype.com/mcp`) scoped to that agent — active only when it runs, so ~0 baseline context cost. Set `SONATYPE_TOKEN` env var to enable ([generate a token](https://guide.sonatype.com/settings/tokens)). |
| **mcp** | `.mcp.json` generated from selected servers, plus **per-task profiles** (`.mcp.research.json`, `.mcp.frontend.json`, `.mcp.minimal.json`) and an executable `./claude-ctx` wrapper that launches Claude with `--mcp-config <profile> --strict-mcp-config` — drops a bloated 4-MCP baseline from ~49% context to under 5%. |
| **multi-agent** | Path-scoped `multi-agent-guardrails.md` (5-scenario "when not to parallel" list), `/merge-worktrees` skill, `/infinite` skill, `parallel-generator` subagent. |
| **mcp** | `.mcp.json` generated from selected servers, plus **per-task profiles** (`.mcp.research.json`, `.mcp.frontend.json`, `.mcp.minimal.json`) and an executable `./claude-ctx` wrapper that launches Claude with `--mcp-config <profile> --strict-mcp-config`. Since tool search defers MCP schemas by default, profiles are about **which servers connect** — startup time, auth prompts, blast radius — not about reclaiming context: measured on CC 2.1.245, four servers × 12 tools cost **+696 tokens deferred vs +14,328 eagerly loaded** (`alwaysLoad: true`). See [`docs/04`](docs/04-subagents-mcp-orchestration.md#what-mcp-actually-costs). |
| **multi-agent** | Path-scoped `multi-agent-guardrails.md` (5-scenario "when not to parallel" list plus the runtime caps to design around), `/merge-worktrees` skill, `/infinite` skill, `parallel-generator` subagent. `/infinite` is scoped to its one good case — N variants of a single spec into disjoint slots; Claude Code's own dynamic workflows are the better tool for staged or merge-heavy fan-out, and the skill says so up front. |
| **github-actions** | `.github/workflows/claude.yml` pinned to `anthropics/claude-code-action@v1`. Triggers on `@claude` mentions in issues, PR comments, and PR reviews. |
| **ui** | Custom status line (project \| branch \| model \| ctx% \| OS+tool-version chip like `deb13 · pg17 · node20 · py3.13`; optionally effort + thinking indicators when 2.1.119+ is running), `statusline-last-prompt.sh` variant, and a "plan" output style. Sub-flag: `no_version_chip` omits the chip (sets `CC_STATUSLINE_NO_VERSION_CHIP=1`) for narrow terminals or privacy. |
| **recommend-plugins** | Drops `docs/recommended-plugins.md` — a stack-aware list of official Claude Code plugins worth considering (always-recommended set + stack-specific picks computed from your form answers). Reference doc; refreshes on every `cc-configure` run. See [`docs/10-plugin-ecosystem.md`](docs/10-plugin-ecosystem.md) for how plugins relate to the configurator. |
Expand Down
56 changes: 48 additions & 8 deletions docs/04-subagents-mcp-orchestration.md
Original file line number Diff line number Diff line change
Expand Up @@ -106,20 +106,50 @@ Situationally valuable:
- **playwright** — any frontend project.
- **context7** — live library docs lookup. Stops hallucinated APIs.

Usually skip: "kitchen sink" servers exposing dozens of tools you won't use. Every tool definition costs tokens on every session start.
Usually skip: "kitchen sink" servers exposing dozens of tools you won't use — not because of tokens (see below) but because every extra server is another process, another auth prompt, and more surface for the model to reach for the wrong tool.

### Context bloat is real
### What MCP actually costs

Every MCP tool is a chunk of JSON schema loaded at session start. A heavy `.mcp.json` can burn thousands of tokens before you type anything — a fresh session with 4 MCP servers loaded typically costs ~37k tokens of tool descriptions alone, ~49% of a 100k window. Mitigations:
**Tool search changed this, and older advice (including earlier versions of this
page) is now wrong.** Claude Code defers MCP tool schemas by default and pulls
them in on demand, so a heavy `.mcp.json` no longer front-loads its schemas into
every prompt. `alwaysLoad: true` opts a server out of that deferral — it is the
setting that brings the old cost back.

1. **Only enable what you'll actually use this week.**
2. **Per-task profiles** — ship `.mcp.<profile>.json` files at repo root and run `claude --mcp-config <path> --strict-mcp-config` (or `./claude-ctx <profile>` using the wrapper the `mcp` module ships). `--strict-mcp-config` ignores the default hierarchy entirely. Real demo: context usage drops from 18.8% to 2.4% when scoped.
3. **Scope to subagents** — put MCP servers in a subagent's frontmatter (`mcpServers:`) so they only load for that agent.
4. **Wrap heavy servers in narrow slash commands** — if a server exposes 30 tools and you use 3, write a skill that calls those 3.
Measured on Claude Code 2.1.245 (Opus 5, one-turn session, four local stdio
servers advertising twelve tools each — 48 tools total):

| Session | Prompt tokens | MCP's share |
|---|---|---|
| No MCP servers | 26,665 | — |
| 4 servers, default (deferred) | 27,361 | +696 (~14 tokens/tool) |
| 4 servers with `alwaysLoad: true` | 40,993 | +14,328 (~298 tokens/tool) |

Deferral removes ~95% of the schema cost. Two consequences:

1. **Adding a server is cheap. Loading it eagerly is not.** Reach for
`alwaysLoad` only when a server's tools are used in nearly every turn.
2. **Per-task profiles are no longer primarily a token lever.** What they still
buy is real: fewer processes and auth prompts at startup, faster cold start,
and a smaller blast radius — `--strict-mcp-config` means only the listed
servers exist for that session. Choose a profile for those reasons.

Still worth doing regardless of tokens:

- **Only enable what you'll actually use this week.** Fewer servers, fewer ways
for a turn to go sideways.
- **Scope to subagents** — put MCP servers in a subagent's frontmatter
(`mcpServers:`) so they're only live while that agent runs.
- **Wrap heavy servers in narrow skills** — if a server exposes 30 tools and you
use 3, a skill that calls those 3 is easier for the model to aim.

### Checking the cost

`/context` in an active session shows what's loaded. `/cost` shows spend. Use both.
`/context` in an active session shows what's loaded, and it is the number to
trust for your own project — the table above is one synthetic shape, not a law.
The dominant line is the baseline itself (~26.7k tokens of system prompt and
built-in tools before any MCP server exists); no profile changes that. `/cost`
shows spend.

## Orchestration patterns

Expand All @@ -135,6 +165,16 @@ Main session edits; `code-reviewer` subagent runs in parallel after every commit
### Desktop + cloud
Local Claude Code for interactive work; Claude Code Web for long-running cloud jobs. Worktrees bridge the two. Useful for "run this large refactor while I go do something else." After both branches converge, use `/merge-worktrees` (ships in `multi-agent`) to integrate safely via a disposable branch.

### When to reach for a workflow instead

Dynamic workflows (`/workflows`) run a script Claude writes over many agents,
with control flow, structured output between stages, a token budget, resume
after failure, and a progress view. Prefer them for staged pipelines
(find → verify → synthesize), for fan-out whose results need merging or scoring,
and for anything that should survive an interruption. Hand-rolled batching —
including the `/infinite` skill the `multi-agent` module ships — is the right
shape only for N independent variants of one spec written to disjoint slots.

### Limits

- The main thread's context is still finite. Subagents help with verbosity but not with total information you're holding in your head.
Expand Down
Loading