diff --git a/CHANGELOG.md b/CHANGELOG.md index 1961173..12f13a3 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -6,6 +6,8 @@ All notable changes to this project. Format: [Keep a Changelog](https://keepacha - **fix(commands): rename `/review` → `/review-branch` so it stops shadowing the bundled `/code-review`.** CC 2.1.223 made `/review` the alias of the bundled `/code-review` — Claude Code's multi-agent reviewer, including the cloud `ultra` mode. A project skill of that name wins it (verified headlessly on 2.1.241: a project skill named `review` ran for `/review`, and the same held for `plan` against the built-in `/plan`), so **every scaffolded project was silently hiding the better built-in behind this simpler single-pass skill** — overlap that turned into a real capability loss the day the alias shipped. The skill moves to `templates/commands/review-branch/` with `name: review-branch`, and its description now positions it honestly ("a quick single-pass review; Claude Code's bundled `/code-review` is the deeper multi-agent one"). Both are reachable again. Updated across `config_schema.py`, `configure.py`'s pattern-integration map, `templates/INDEX.md`, the `/investigate` and `/plan-eng-review` cross-references, docs 02/03/05/09/10/11, README, and the example project. **Migration:** the configurator has no mechanism to delete a file it previously wrote, so an upgraded project keeps the old `.claude/skills/review/` alongside the new one — and the stale copy still shadows the alias. New `/verify-setup` **check 13** detects exactly that pair and tells the user to `rm -rf .claude/skills/review`. `/plan` is left alone deliberately: it shadows a built-in *command* rather than a bundled skill, and plan mode stays reachable via Shift+Tab, so it's a name clash rather than a lost capability — README now says so and points at the rename if you'd rather keep the shortcut. +- **docs: re-baseline the MCP context claims against tool search, and scope `/infinite` against dynamic workflows.** Two of the project's headline claims had been overtaken by Claude Code and were overstating what the modules buy. **MCP.** README claimed per-task profiles "drop a bloated 4-MCP baseline from ~49% context to under 5%", and `docs/04` asserted "every MCP tool is a chunk of JSON schema loaded at session start". Tool search defers MCP schemas by default (`alwaysLoad: true` is the opt-*out*), so the premise no longer holds. Measured rather than re-guessed — four local stdio servers advertising twelve tools each (48 total) against an otherwise identical one-turn session on CC 2.1.245: **26,665 tokens with no MCP servers, 27,361 deferred (+696, ~14/tool), 40,993 with `alwaysLoad: true` (+14,328, ~298/tool)** — deferral removes ~95% of the schema cost, and the numbers reproduced exactly across runs. The claim was also embedded in four *shipped* templates, which is worse than in the docs because it lands in every user's project: `check-context/SKILL.md` (its budget guardrails and the "MCP > 10%" flag), `claude-ctx.sh`'s rationale comment, `servers-cookbook.md` (which already explained deferral correctly a few sections earlier, so it contradicted itself), and the `mcp.minimal.json` profile comment. All corrected. Profiles are now documented for what they still genuinely buy — which servers *connect*: startup time, auth prompts, cold start, and the blast radius `--strict-mcp-config` enforces — and `docs/06` picks up the same correction. The dated `experiments-memory` example keeps its original result with a superseding **addendum** rather than a rewrite, because an experiment log records what was true when it ran. **`/infinite`.** Dynamic workflows now do staged, resumable, budgeted fan-out with structured output between stages; hand-rolled wave batching is the weaker instrument for that job. The skill opens with a decision table sending staged / merge-heavy / resumable work to a workflow, and keeps the one case it is genuinely good at — N variants of a single spec into disjoint slots with no cross-iteration coordination. README's module row and a new `docs/04` section say the same. + - **fix(portability): scaffolding from Windows emitted CRLF into every generated file; `--check` crashed; two hooks assumed tools that aren't there.** Four bugs, all invisible to a Linux-only CI. **(1) CRLF output.** `Path.write_text()` opens in text mode with `newline=None`, which translates `\n` to `os.linesep` — so a Windows scaffold wrote CRLF into all 13 hooks, `claude-ctx`, `CLAUDE.md` and `settings.json` (measured: 100% CRLF, zero bare LF). Harmless *on* Windows, where Git Bash strips the CR, but the shebang carried one, and `.claude/` is meant to be committed and shared (the generated CLAUDE.md says so, and `small-team` is a shipped persona) — so a teammate's Linux/macOS clone got `bad interpreter: /usr/bin/env bash^M` on every hook. All six write sites now route through a `write_text_lf()` helper using `open(..., newline="\n")` (`Path.write_text` only gained a `newline` parameter in 3.10; the project floor is 3.8). **(2) `--check` crashed on Windows** with `UnicodeEncodeError` on the `✓` under cp1252 — the exact command CONTRIBUTING gives new contributors. Python uses UTF-8 for a real Windows console, so this bit hardest when output was piped or redirected (CI logs, `cc-configure > setup.log`, an editor terminal). Both streams are now retagged UTF-8 with `errors="replace"` where the runtime allows it. **(3) The Stop hook's report depended on an unguarded `jq -n`.** jq isn't preinstalled on macOS, Windows, or most Linux distros; without it the checks ran, the hook exited 0, and Claude never learned anything had failed — the hook's entire purpose, lost silently. It now prefers jq, falls back to python3, and as a last resort prints the report to stderr with an explicit "NOT reaching Claude" warning. **(4) The package-availability gate disabled itself on stock macOS.** It bounded each probe with GNU `timeout`, absent there; the resulting rc=127 took the "probe inconclusive" branch, so the gate stood down on the first package for every macOS user — fail-open, so no false denials, but the advertised `safety` feature simply did not work, and said nothing. It now falls back to `gtimeout`, then to an unbounded probe. **Mitigation for what the configurator can't control:** the scaffold appends a `.gitattributes` block (`*.sh text eol=lf`, `claude-ctx`, `.claude/skills/**/scripts/*`) through the same managed-block append the `.gitignore` block uses — written when absent, line-level-unioned into a stale block, never touching the user's own rules. Git cannot record the *executable* bit that way, and `core.filemode=false` (the Windows default) drops it, so the generated `CLAUDE.md` § Repo bootstrap now tells Windows committers to run `git update-index --chmod=+x claude-ctx`. New `test/portability/` fixtures cover all of it by behavior, not by grepping source: every generated file is CRLF-free and every hook shebang is clean; the `.gitattributes` block is written, idempotent and union-safe; and the two hooks are run under a PATH that hides jq, python3 and timeout. README's runtime-dependency list now names jq and the timeout fallback; the platform table reflects what CI actually exercises. The `python-uv-fastapi` example and all five persona snapshots are regenerated (`.gitattributes` is new output). - **fix(install): the `ln -sf` shortcut silently froze on Windows; PATH-independent shim, honest prerequisites.** Under Git Bash/MSYS, `ln -s` **copies** unless `MSYS=winsymlinks:nativestrict` is set *and* the user has Developer Mode or admin — so `~/.local/bin/cc-configure` became a snapshot of configure.py taken on install day. `git -C ~/.cc-configurator pull` updated the clone while the command kept running the old copy, with nothing to indicate it. The installer now verifies the link actually resolved (`[ -L ]`) and otherwise writes a two-line `exec python3 /configure.py "$@"` shim, which tracks the clone on every platform; both checks live in the `if` condition, because under `set -e` a trailing `[ -L … ] && link_ok=1` in a then-block would abort the installer on precisely the platform the fallback exists for. Also: `git` is now checked up front (it was used before the first guard), the missing-python3 hint is OS-aware instead of always saying `apt install python3`, and the closing usage banner no longer advertises `--preset aggressive` and the legacy `commands-core` / `token-efficiency-pro` module IDs — both deprecated and slated for removal in v3.0. diff --git a/README.md b/README.md index 56d04c3..150fa12 100644 --- a/README.md +++ b/README.md @@ -131,8 +131,8 @@ There are 11 modules; legacy IDs (`commands-core`, `agents`, `token-efficiency-p | **git-workflow** | PostToolUse formatter on Write/Edit, Stop hook running typecheck / lint / tests. | | **token-efficiency** | Path-scoped `.claude/rules/` starters + PreCompact snapshot hook. `tier` flag: `basic` (default) ships discipline rules + snapshot only; `pro` adds bash-output truncation hook + always-loaded discipline rules. | | **commands** | Slash commands + agents + microbits. `subset` flag (linear ordering: `curated ⊂ full ⊂ rigorous`): **`curated`** = 3 essential skills (`/plan`, `/commit`, `/verify-setup`) + the `code-reviewer` agent. **`full`** (default) = 9 workflow skills (adds `/review-branch`, `/ship`, `/sync-docs`, `/check-context`, `/session-retro`, `/retrofit`) + 4 agents (`code-reviewer`, `test-runner`, `doc-writer`, `security-auditor`) + 4 discipline microbits (`/freeze`, `/unfreeze`, `/guard`, `/careful`) + the `microbit-enforcer.sh` PreToolUse hook. **`rigorous`** = `full` + `/investigate` + `/plan-eng-review`, the rigor skills that embed `templates/commands/_patterns/` cross-cutting blocks (confidence gate, independent verification, no-fix-without-investigation, AI-slop detection). **Naming note:** a project skill wins its name against Claude Code's built-ins (verified on CC 2.1.241). `/review-branch` is deliberately *not* called `review`, because CC 2.1.223 made `/review` the alias of the bundled `/code-review` and taking that name would hide Claude Code's multi-agent reviewer. `/plan` does still shadow the built-in plan-mode shortcut — plan mode remains reachable with Shift+Tab, so this is a name clash rather than a lost capability; rename the skill if you'd rather keep the shortcut. The `security-auditor` frontmatter wires Sonatype's dependency-management MCP (`https://mcp.guide.sonatype.com/mcp`) scoped to that agent — active only when it runs, so ~0 baseline context cost. Set `SONATYPE_TOKEN` env var to enable ([generate a token](https://guide.sonatype.com/settings/tokens)). | -| **mcp** | `.mcp.json` generated from selected servers, plus **per-task profiles** (`.mcp.research.json`, `.mcp.frontend.json`, `.mcp.minimal.json`) and an executable `./claude-ctx` wrapper that launches Claude with `--mcp-config --strict-mcp-config` — drops a bloated 4-MCP baseline from ~49% context to under 5%. | -| **multi-agent** | Path-scoped `multi-agent-guardrails.md` (5-scenario "when not to parallel" list), `/merge-worktrees` skill, `/infinite` skill, `parallel-generator` subagent. | +| **mcp** | `.mcp.json` generated from selected servers, plus **per-task profiles** (`.mcp.research.json`, `.mcp.frontend.json`, `.mcp.minimal.json`) and an executable `./claude-ctx` wrapper that launches Claude with `--mcp-config --strict-mcp-config`. Since tool search defers MCP schemas by default, profiles are about **which servers connect** — startup time, auth prompts, blast radius — not about reclaiming context: measured on CC 2.1.245, four servers × 12 tools cost **+696 tokens deferred vs +14,328 eagerly loaded** (`alwaysLoad: true`). See [`docs/04`](docs/04-subagents-mcp-orchestration.md#what-mcp-actually-costs). | +| **multi-agent** | Path-scoped `multi-agent-guardrails.md` (5-scenario "when not to parallel" list plus the runtime caps to design around), `/merge-worktrees` skill, `/infinite` skill, `parallel-generator` subagent. `/infinite` is scoped to its one good case — N variants of a single spec into disjoint slots; Claude Code's own dynamic workflows are the better tool for staged or merge-heavy fan-out, and the skill says so up front. | | **github-actions** | `.github/workflows/claude.yml` pinned to `anthropics/claude-code-action@v1`. Triggers on `@claude` mentions in issues, PR comments, and PR reviews. | | **ui** | Custom status line (project \| branch \| model \| ctx% \| OS+tool-version chip like `deb13 · pg17 · node20 · py3.13`; optionally effort + thinking indicators when 2.1.119+ is running), `statusline-last-prompt.sh` variant, and a "plan" output style. Sub-flag: `no_version_chip` omits the chip (sets `CC_STATUSLINE_NO_VERSION_CHIP=1`) for narrow terminals or privacy. | | **recommend-plugins** | Drops `docs/recommended-plugins.md` — a stack-aware list of official Claude Code plugins worth considering (always-recommended set + stack-specific picks computed from your form answers). Reference doc; refreshes on every `cc-configure` run. See [`docs/10-plugin-ecosystem.md`](docs/10-plugin-ecosystem.md) for how plugins relate to the configurator. | diff --git a/docs/04-subagents-mcp-orchestration.md b/docs/04-subagents-mcp-orchestration.md index 9581451..c73c8c2 100644 --- a/docs/04-subagents-mcp-orchestration.md +++ b/docs/04-subagents-mcp-orchestration.md @@ -106,20 +106,50 @@ Situationally valuable: - **playwright** — any frontend project. - **context7** — live library docs lookup. Stops hallucinated APIs. -Usually skip: "kitchen sink" servers exposing dozens of tools you won't use. Every tool definition costs tokens on every session start. +Usually skip: "kitchen sink" servers exposing dozens of tools you won't use — not because of tokens (see below) but because every extra server is another process, another auth prompt, and more surface for the model to reach for the wrong tool. -### Context bloat is real +### What MCP actually costs -Every MCP tool is a chunk of JSON schema loaded at session start. A heavy `.mcp.json` can burn thousands of tokens before you type anything — a fresh session with 4 MCP servers loaded typically costs ~37k tokens of tool descriptions alone, ~49% of a 100k window. Mitigations: +**Tool search changed this, and older advice (including earlier versions of this +page) is now wrong.** Claude Code defers MCP tool schemas by default and pulls +them in on demand, so a heavy `.mcp.json` no longer front-loads its schemas into +every prompt. `alwaysLoad: true` opts a server out of that deferral — it is the +setting that brings the old cost back. -1. **Only enable what you'll actually use this week.** -2. **Per-task profiles** — ship `.mcp..json` files at repo root and run `claude --mcp-config --strict-mcp-config` (or `./claude-ctx ` using the wrapper the `mcp` module ships). `--strict-mcp-config` ignores the default hierarchy entirely. Real demo: context usage drops from 18.8% to 2.4% when scoped. -3. **Scope to subagents** — put MCP servers in a subagent's frontmatter (`mcpServers:`) so they only load for that agent. -4. **Wrap heavy servers in narrow slash commands** — if a server exposes 30 tools and you use 3, write a skill that calls those 3. +Measured on Claude Code 2.1.245 (Opus 5, one-turn session, four local stdio +servers advertising twelve tools each — 48 tools total): + +| Session | Prompt tokens | MCP's share | +|---|---|---| +| No MCP servers | 26,665 | — | +| 4 servers, default (deferred) | 27,361 | +696 (~14 tokens/tool) | +| 4 servers with `alwaysLoad: true` | 40,993 | +14,328 (~298 tokens/tool) | + +Deferral removes ~95% of the schema cost. Two consequences: + +1. **Adding a server is cheap. Loading it eagerly is not.** Reach for + `alwaysLoad` only when a server's tools are used in nearly every turn. +2. **Per-task profiles are no longer primarily a token lever.** What they still + buy is real: fewer processes and auth prompts at startup, faster cold start, + and a smaller blast radius — `--strict-mcp-config` means only the listed + servers exist for that session. Choose a profile for those reasons. + +Still worth doing regardless of tokens: + +- **Only enable what you'll actually use this week.** Fewer servers, fewer ways + for a turn to go sideways. +- **Scope to subagents** — put MCP servers in a subagent's frontmatter + (`mcpServers:`) so they're only live while that agent runs. +- **Wrap heavy servers in narrow skills** — if a server exposes 30 tools and you + use 3, a skill that calls those 3 is easier for the model to aim. ### Checking the cost -`/context` in an active session shows what's loaded. `/cost` shows spend. Use both. +`/context` in an active session shows what's loaded, and it is the number to +trust for your own project — the table above is one synthetic shape, not a law. +The dominant line is the baseline itself (~26.7k tokens of system prompt and +built-in tools before any MCP server exists); no profile changes that. `/cost` +shows spend. ## Orchestration patterns @@ -135,6 +165,16 @@ Main session edits; `code-reviewer` subagent runs in parallel after every commit ### Desktop + cloud Local Claude Code for interactive work; Claude Code Web for long-running cloud jobs. Worktrees bridge the two. Useful for "run this large refactor while I go do something else." After both branches converge, use `/merge-worktrees` (ships in `multi-agent`) to integrate safely via a disposable branch. +### When to reach for a workflow instead + +Dynamic workflows (`/workflows`) run a script Claude writes over many agents, +with control flow, structured output between stages, a token budget, resume +after failure, and a progress view. Prefer them for staged pipelines +(find → verify → synthesize), for fan-out whose results need merging or scoring, +and for anything that should survive an interruption. Hand-rolled batching — +including the `/infinite` skill the `multi-agent` module ships — is the right +shape only for N independent variants of one spec written to disjoint slots. + ### Limits - The main thread's context is still finite. Subagents help with verbosity but not with total information you're holding in your head. diff --git a/docs/06-token-efficiency.md b/docs/06-token-efficiency.md index 0079a33..b5f29d7 100644 --- a/docs/06-token-efficiency.md +++ b/docs/06-token-efficiency.md @@ -10,7 +10,7 @@ Every session starts with: 2. **CLAUDE.md** (+ concatenated nested ones) — your doing. **Main lever.** 3. **Imported files via `@path`** — expand at launch. Same cost as if inlined. 4. **Auto-memory** — first 200 lines / 25 KB of MEMORY.md. Toggleable. -5. **MCP tool definitions** — every enabled server's tool schemas. **Hidden fat.** +5. **MCP tool definitions** — every enabled server's tool schemas. Mostly *not* fat any more: tool search defers them by default (~14 tokens/tool instead of ~298 — see [doc 04](04-subagents-mcp-orchestration.md#what-mcp-actually-costs)). It becomes fat again for any server you set `alwaysLoad: true` on. 6. **Available subagents' descriptions** — short, but they add up if you have many. 7. **Skills with `disable-model-invocation: false`** — descriptions only, not bodies. @@ -32,7 +32,7 @@ After that, each turn adds: Path-scoped rules only load when Claude reads matching files. Zero cost otherwise. Move anything domain-specific here (frontend rules, test rules, DB rules). See `templates/token-efficiency/dot-claude/rules/` for starters. ### 3. Audit MCP servers quarterly -Run `/context` in a session. Count the tokens eaten by MCP tool definitions. If a server exposes 30 tools and you use 3, either: +Run `/context` in a session. Count the tokens eaten by MCP tool definitions — with deferral on (the default) this should be small, and a large number means something is set to `alwaysLoad`. If a server exposes 30 tools and you use 3, either: - Remove it and call those 3 as Bash commands. - Scope it to a subagent that needs it. - Replace it with a targeted skill. @@ -53,7 +53,7 @@ Teach Claude (in CLAUDE.md rules) to prefer `grep` / structured search over `cat Common causes: -- **Over-general MCP config.** A broad server with many tools burns tokens every session whether you use it or not. +- **Over-general MCP config.** A broad server costs little while its tools stay deferred, but it still adds a process, an auth prompt, and more ways for the model to pick the wrong tool — and it burns tokens every session if you set `alwaysLoad: true`. - **Fat CLAUDE.md.** Every subdirectory loads its nested CLAUDE.md too when Claude reads files there. If every subdir has a 300-line CLAUDE.md, you're spending hugely. - **Eager skills with `disable-model-invocation: false`.** Claude scans skill descriptions at every turn to decide whether to invoke. Many skills with long descriptions = real token cost. - **Auto-memory gone wild.** `/memory` shows what's been saved. If MEMORY.md has years of stale notes, prune it or disable auto-memory. @@ -103,7 +103,7 @@ Since Claude Code 2.1.117, Pro/Max subscribers on Opus 4.6 and Sonnet 4.6 defaul - [ ] CLAUDE.md is under 200 lines. - [ ] Path-scoped rules handle all per-domain guidance. -- [ ] `/context` shows MCP tool definitions under ~3k tokens. +- [ ] `/context` shows MCP tool definitions under ~3k tokens (easy with deferral on; if it's far above, check for `alwaysLoad`). - [ ] No subagent descriptions over ~50 words. - [ ] `PreCompact` hook snapshots session state. - [ ] `/cost` and `/context` checked at least weekly. diff --git a/examples/python-uv-fastapi/.claude/skills/check-context/SKILL.md b/examples/python-uv-fastapi/.claude/skills/check-context/SKILL.md index fd82351..49bfe48 100644 --- a/examples/python-uv-fastapi/.claude/skills/check-context/SKILL.md +++ b/examples/python-uv-fastapi/.claude/skills/check-context/SKILL.md @@ -11,11 +11,11 @@ Run `/context` first (the built-in) and read the breakdown it prints. Then analy ## Budget guardrails -A fresh Claude Code session with 4 MCP servers loaded already burns ~49% of a 100k window before the user types anything — system prompt ~2.2k, system tool descriptions ~12k, MCP tool descriptions ~37k. That's the baseline to compare against. +Before the user types anything a session already carries the system prompt, built-in tool descriptions, CLAUDE.md, path-scoped rules and the skill listing. Measured on Claude Code 2.1.245 that floor was ~26.7k tokens with no MCP servers configured at all — that's the baseline to compare against, and no MCP profile changes it. MCP tool schemas are deferred by tool search unless a server sets `alwaysLoad: true`, so they should be a thin slice (~14 tokens/tool, versus ~298 when eagerly loaded). Flag the following: -- **MCP tool descriptions > 10%** of the window → too many MCP servers for this task. Recommend the user split per-task `.mcp.json.` files and run `claude --mcp-config --strict-mcp-config`. +- **MCP tool descriptions > 10%** of the window → with deferral on this is nearly impossible, so suspect a server with `alwaysLoad: true` (or a build predating tool search). Recommend dropping `alwaysLoad` first. Per-task `.mcp..json` files with `claude --mcp-config --strict-mcp-config` (or `./claude-ctx `) remain the right move when the goal is *fewer connected servers* — startup time, auth prompts, blast radius — rather than fewer tokens. - **Custom tools / skills > 5%** of the window → some skill descriptions are too verbose or too many user-invocable skills are loaded. Recommend auditing `allowed-tools`, `when_to_use`, and trimming descriptions. - **Memory (CLAUDE.md + @imports) > 10%** → CLAUDE.md has bloated. Propose moving path-scoped content into `.claude/rules/*.md` with `paths:` frontmatter so it only loads when relevant. - **Total before first turn > 40%** → the session will autocompact early. Combination of the above. diff --git a/templates/commands/check-context/SKILL.md b/templates/commands/check-context/SKILL.md index c4a0ac4..041c622 100644 --- a/templates/commands/check-context/SKILL.md +++ b/templates/commands/check-context/SKILL.md @@ -10,11 +10,11 @@ Run `/context` first (the built-in) and read the breakdown it prints. Then analy ## Budget guardrails -A fresh Claude Code session with 4 MCP servers loaded already burns ~49% of a 100k window before the user types anything — system prompt ~2.2k, system tool descriptions ~12k, MCP tool descriptions ~37k. That's the baseline to compare against. +Before the user types anything a session already carries the system prompt, built-in tool descriptions, CLAUDE.md, path-scoped rules and the skill listing. Measured on Claude Code 2.1.245 that floor was ~26.7k tokens with no MCP servers configured at all — that's the baseline to compare against, and no MCP profile changes it. MCP tool schemas are deferred by tool search unless a server sets `alwaysLoad: true`, so they should be a thin slice (~14 tokens/tool, versus ~298 when eagerly loaded). Flag the following: -- **MCP tool descriptions > 10%** of the window → too many MCP servers for this task. Recommend the user split per-task `.mcp.json.` files and run `claude --mcp-config --strict-mcp-config`. +- **MCP tool descriptions > 10%** of the window → with deferral on this is nearly impossible, so suspect a server with `alwaysLoad: true` (or a build predating tool search). Recommend dropping `alwaysLoad` first. Per-task `.mcp..json` files with `claude --mcp-config --strict-mcp-config` (or `./claude-ctx `) remain the right move when the goal is *fewer connected servers* — startup time, auth prompts, blast radius — rather than fewer tokens. - **Custom tools / skills > 5%** of the window → some skill descriptions are too verbose or too many user-invocable skills are loaded. Recommend auditing `allowed-tools`, `when_to_use`, and trimming descriptions. - **Memory (CLAUDE.md + @imports) > 10%** → CLAUDE.md has bloated. Propose moving path-scoped content into `.claude/rules/*.md` with `paths:` frontmatter so it only loads when relevant. - **Total before first turn > 40%** → the session will autocompact early. Combination of the above. diff --git a/templates/commands/infinite/SKILL.md b/templates/commands/infinite/SKILL.md index d26463f..4aabe56 100644 --- a/templates/commands/infinite/SKILL.md +++ b/templates/commands/infinite/SKILL.md @@ -7,7 +7,28 @@ allowed-tools: Read Grep Glob Bash(ls:*) Bash(mkdir:*) Bash(test:*) # Parallel spec expansion: `$ARGUMENTS` -Before anything touches disk, confirm **this is a parallelizable fanout task.** If the spec asks for sequential work, coupled features, or exploratory problem-solving, **refuse** and point at `.claude/rules/multi-agent-guardrails.md`. Parallel agents are a fanout multiplier, not a general speedup. +## First: is this the right tool? + +Claude Code ships **dynamic workflows** (`/workflows`, or asking for one in the +prompt), which orchestrate agents from a script Claude writes: real control +flow, structured output between stages, a token budget, resumability after a +failure, and a live progress view. For most fan-out work that is now the better +instrument, and this skill should not compete with it. + +| Situation | Use | +|---|---| +| N variants of **one spec**, each written to its own slot, no coordination | **this skill** | +| Stages that feed each other (find → verify → synthesize) | a dynamic workflow | +| Fan-out where results must be merged, deduped, or scored | a dynamic workflow | +| The run must survive an interruption and resume | a dynamic workflow | +| More agents than you want to hand-batch | a dynamic workflow | + +What this skill still does well is the narrow case it was written for: identical +task shape, disjoint output slots, deliberate diversification along one named +axis, and no dependency between iterations. If the work doesn't look like that, +stop and reach for a workflow instead. + +Then confirm **this is a parallelizable fanout task.** If the spec asks for sequential work, coupled features, or exploratory problem-solving, **refuse** and point at `.claude/rules/multi-agent-guardrails.md`. Parallel agents are a fanout multiplier, not a general speedup. ## Parse diff --git a/templates/experiments-memory/memory/experiments/2026-04-24-example-profile-budget.md b/templates/experiments-memory/memory/experiments/2026-04-24-example-profile-budget.md index a427ceb..8c630df 100644 --- a/templates/experiments-memory/memory/experiments/2026-04-24-example-profile-budget.md +++ b/templates/experiments-memory/memory/experiments/2026-04-24-example-profile-budget.md @@ -16,6 +16,17 @@ Delta: ~33.4k tokens reclaimed. Session context at first turn went from 49% → ## Conclusion Hypothesis held. The `--strict-mcp-config` flag is the key — without it the scoped file merges with the default hierarchy and the savings disappear. Codified as: default working mode for research/debugging sessions is now `./claude-ctx research`, not plain `claude`. `claude` without a profile is reserved for sessions that genuinely need filesystem+git+github simultaneously. +## Addendum (2026-08-25) — superseded by tool search + +Claude Code now defers MCP tool schemas by default (`alwaysLoad: true` opts a +server out), so the control condition this experiment measured no longer +happens on a current build. Re-measured with four servers advertising 48 tools +on CC 2.1.245: +696 tokens deferred versus +14,328 eagerly loaded, against a +~26.7k baseline with no servers at all. The conclusion below still holds for +`alwaysLoad` servers and for builds predating tool search; the headline saving +does not generalize. Left in place deliberately — an experiment log records what +was true when it ran, and the addendum is where the update belongs. + ## Follow-ups - Measure the `minimal` profile the same way — expected ~0 tokens for MCP, but is there overhead from the empty `mcpServers: {}` itself? - Try the `frontend` profile on a real UI bug to see if playwright alone is enough or if filesystem access is needed too. diff --git a/templates/mcp/claude-ctx.sh b/templates/mcp/claude-ctx.sh index ac5143c..b386893 100644 --- a/templates/mcp/claude-ctx.sh +++ b/templates/mcp/claude-ctx.sh @@ -5,11 +5,16 @@ # MCP hierarchy (user/project/local). Profiles live at .mcp..json in # the project root. # -# WHY this exists: a fresh session with 4 MCP servers loaded burns ~49% of a -# 100k context window before you type anything (system prompt ~2.2k + -# system tool descriptions ~12k + MCP tool descriptions ~37k). Scoping per -# task with --strict-mcp-config drops that cost to near-zero for tasks that -# don't need those servers. +# WHY this exists: --strict-mcp-config makes the listed servers the ONLY ones +# that exist for the session, so a profile controls which servers actually +# connect: fewer processes and auth prompts at startup, a faster cold start, +# and a smaller blast radius for a task that has no business touching them. +# +# Note on tokens: this used to be pitched as a context saving, and it no longer +# is. Tool search defers MCP tool schemas by default — measured on Claude Code +# 2.1.245, four servers advertising 48 tools between them cost +696 tokens +# deferred versus +14,328 with alwaysLoad: true. Use profiles for connection +# control; use `/context` if you want to see where the tokens really went. # # Usage: # ./claude-ctx research # loads only .mcp.research.json diff --git a/templates/mcp/profiles/mcp.minimal.json b/templates/mcp/profiles/mcp.minimal.json index fa77024..edb7372 100644 --- a/templates/mcp/profiles/mcp.minimal.json +++ b/templates/mcp/profiles/mcp.minimal.json @@ -1,4 +1,4 @@ { - "//": "Profile: minimal. No MCP servers. Use for pure writing/editing sessions where external tools are dead weight. Demonstrates the --strict-mcp-config savings cleanly (no ~37k of MCP tool descriptions loaded). Runs via: ./claude-ctx minimal", + "//": "Profile: minimal. No MCP servers. Use for pure writing/editing sessions where external tools are dead weight. The cleanest demonstration of --strict-mcp-config: no MCP servers connect at all, so nothing external can be reached from the session. Runs via: ./claude-ctx minimal", "mcpServers": {} } diff --git a/templates/mcp/servers-cookbook.md b/templates/mcp/servers-cookbook.md index faeed74..bf615ca 100644 --- a/templates/mcp/servers-cookbook.md +++ b/templates/mcp/servers-cookbook.md @@ -92,7 +92,7 @@ For everything else, leave `alwaysLoad` unset and let the model pull tools in as ## Per-task profiles with `claude-ctx` -A fresh session with 4 MCP servers loaded burns ~49% of a 100k context window before you type anything (system prompt ~2.2k + system tools ~12k + MCP tool descriptions ~37k). For tasks that don't need those servers, scope per-task instead of loading the full default set. +Profiles decide **which servers connect**, not how many tokens you spend. With tool search deferring schemas (see the `alwaysLoad` section above), four servers advertising 48 tools between them measured +696 tokens on Claude Code 2.1.245, versus +14,328 with `alwaysLoad: true`. What scoping still buys: fewer processes and auth prompts at startup, a faster cold start, and a session that physically cannot reach a server the task has no business touching — `--strict-mcp-config` ignores the default hierarchy entirely. Profiles live at `.mcp..json` in the project root. Three are shipped by default: