From a2b6c08afe55856cb3bf171b5313110b9f9a3652 Mon Sep 17 00:00:00 2001 From: tigers1997 Date: Tue, 25 Aug 2026 09:40:34 -0400 Subject: [PATCH] docs: re-baseline the MCP context claims against tool search; scope /infinite Two headline claims had been overtaken by Claude Code and were overstating what the modules buy. MCP. README claimed per-task profiles "drop a bloated 4-MCP baseline from ~49% context to under 5%", and docs/04 asserted that "every MCP tool is a chunk of JSON schema loaded at session start". Tool search defers MCP schemas by default (alwaysLoad: true is the opt-OUT), so the premise no longer holds. Measured rather than re-guessed. Four local stdio servers advertising twelve tools each (48 total), against an otherwise identical one-turn session on CC 2.1.245: no MCP servers 26,665 prompt tokens 4 servers, deferred 27,361 (+696, ~14 tokens/tool) 4 servers, alwaysLoad 40,993 (+14,328, ~298 tokens/tool) Deferral removes ~95% of the schema cost, and the numbers reproduced exactly across runs. The stale claim was also embedded in four SHIPPED templates, which is worse than in the docs because it lands in every user's project: check-context/SKILL.md (its budget guardrails and the "MCP > 10%" flag), claude-ctx.sh's rationale comment, servers-cookbook.md (which already explained deferral correctly a few sections earlier, so it contradicted itself), and the mcp.minimal.json profile comment. All corrected. Profiles are now documented for what they still genuinely buy -- which servers connect: startup time, auth prompts, cold start, and the blast radius --strict-mcp-config enforces. docs/06 picks up the same correction. The dated experiments-memory example keeps its original result with a superseding addendum rather than a rewrite, because an experiment log records what was true when it ran; that is also a better demonstration of the format. /infinite. Dynamic workflows now do staged, resumable, budgeted fan-out with structured output between stages, and hand-rolled wave batching is the weaker instrument for that job. The skill opens with a decision table routing staged, merge-heavy or resumable work to a workflow, and keeps the one case it is genuinely good at: N variants of a single spec into disjoint slots with no cross-iteration coordination. README's module row and a new docs/04 section say the same. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01JVndNviHZSnbKJnWP7jFbV --- CHANGELOG.md | 2 + README.md | 4 +- docs/04-subagents-mcp-orchestration.md | 56 ++++++++++++++++--- docs/06-token-efficiency.md | 8 +-- .../.claude/skills/check-context/SKILL.md | 4 +- templates/commands/check-context/SKILL.md | 4 +- templates/commands/infinite/SKILL.md | 23 +++++++- .../2026-04-24-example-profile-budget.md | 11 ++++ templates/mcp/claude-ctx.sh | 15 +++-- templates/mcp/profiles/mcp.minimal.json | 2 +- templates/mcp/servers-cookbook.md | 2 +- 11 files changed, 105 insertions(+), 26 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 1961173..12f13a3 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -6,6 +6,8 @@ All notable changes to this project. Format: [Keep a Changelog](https://keepacha - **fix(commands): rename `/review` → `/review-branch` so it stops shadowing the bundled `/code-review`.** CC 2.1.223 made `/review` the alias of the bundled `/code-review` — Claude Code's multi-agent reviewer, including the cloud `ultra` mode. A project skill of that name wins it (verified headlessly on 2.1.241: a project skill named `review` ran for `/review`, and the same held for `plan` against the built-in `/plan`), so **every scaffolded project was silently hiding the better built-in behind this simpler single-pass skill** — overlap that turned into a real capability loss the day the alias shipped. The skill moves to `templates/commands/review-branch/` with `name: review-branch`, and its description now positions it honestly ("a quick single-pass review; Claude Code's bundled `/code-review` is the deeper multi-agent one"). Both are reachable again. Updated across `config_schema.py`, `configure.py`'s pattern-integration map, `templates/INDEX.md`, the `/investigate` and `/plan-eng-review` cross-references, docs 02/03/05/09/10/11, README, and the example project. **Migration:** the configurator has no mechanism to delete a file it previously wrote, so an upgraded project keeps the old `.claude/skills/review/` alongside the new one — and the stale copy still shadows the alias. New `/verify-setup` **check 13** detects exactly that pair and tells the user to `rm -rf .claude/skills/review`. `/plan` is left alone deliberately: it shadows a built-in *command* rather than a bundled skill, and plan mode stays reachable via Shift+Tab, so it's a name clash rather than a lost capability — README now says so and points at the rename if you'd rather keep the shortcut. +- **docs: re-baseline the MCP context claims against tool search, and scope `/infinite` against dynamic workflows.** Two of the project's headline claims had been overtaken by Claude Code and were overstating what the modules buy. **MCP.** README claimed per-task profiles "drop a bloated 4-MCP baseline from ~49% context to under 5%", and `docs/04` asserted "every MCP tool is a chunk of JSON schema loaded at session start". Tool search defers MCP schemas by default (`alwaysLoad: true` is the opt-*out*), so the premise no longer holds. Measured rather than re-guessed — four local stdio servers advertising twelve tools each (48 total) against an otherwise identical one-turn session on CC 2.1.245: **26,665 tokens with no MCP servers, 27,361 deferred (+696, ~14/tool), 40,993 with `alwaysLoad: true` (+14,328, ~298/tool)** — deferral removes ~95% of the schema cost, and the numbers reproduced exactly across runs. The claim was also embedded in four *shipped* templates, which is worse than in the docs because it lands in every user's project: `check-context/SKILL.md` (its budget guardrails and the "MCP > 10%" flag), `claude-ctx.sh`'s rationale comment, `servers-cookbook.md` (which already explained deferral correctly a few sections earlier, so it contradicted itself), and the `mcp.minimal.json` profile comment. All corrected. Profiles are now documented for what they still genuinely buy — which servers *connect*: startup time, auth prompts, cold start, and the blast radius `--strict-mcp-config` enforces — and `docs/06` picks up the same correction. The dated `experiments-memory` example keeps its original result with a superseding **addendum** rather than a rewrite, because an experiment log records what was true when it ran. **`/infinite`.** Dynamic workflows now do staged, resumable, budgeted fan-out with structured output between stages; hand-rolled wave batching is the weaker instrument for that job. The skill opens with a decision table sending staged / merge-heavy / resumable work to a workflow, and keeps the one case it is genuinely good at — N variants of a single spec into disjoint slots with no cross-iteration coordination. README's module row and a new `docs/04` section say the same. + - **fix(portability): scaffolding from Windows emitted CRLF into every generated file; `--check` crashed; two hooks assumed tools that aren't there.** Four bugs, all invisible to a Linux-only CI. **(1) CRLF output.** `Path.write_text()` opens in text mode with `newline=None`, which translates `\n` to `os.linesep` — so a Windows scaffold wrote CRLF into all 13 hooks, `claude-ctx`, `CLAUDE.md` and `settings.json` (measured: 100% CRLF, zero bare LF). Harmless *on* Windows, where Git Bash strips the CR, but the shebang carried one, and `.claude/` is meant to be committed and shared (the generated CLAUDE.md says so, and `small-team` is a shipped persona) — so a teammate's Linux/macOS clone got `bad interpreter: /usr/bin/env bash^M` on every hook. All six write sites now route through a `write_text_lf()` helper using `open(..., newline="\n")` (`Path.write_text` only gained a `newline` parameter in 3.10; the project floor is 3.8). **(2) `--check` crashed on Windows** with `UnicodeEncodeError` on the `✓` under cp1252 — the exact command CONTRIBUTING gives new contributors. Python uses UTF-8 for a real Windows console, so this bit hardest when output was piped or redirected (CI logs, `cc-configure > setup.log`, an editor terminal). Both streams are now retagged UTF-8 with `errors="replace"` where the runtime allows it. **(3) The Stop hook's report depended on an unguarded `jq -n`.** jq isn't preinstalled on macOS, Windows, or most Linux distros; without it the checks ran, the hook exited 0, and Claude never learned anything had failed — the hook's entire purpose, lost silently. It now prefers jq, falls back to python3, and as a last resort prints the report to stderr with an explicit "NOT reaching Claude" warning. **(4) The package-availability gate disabled itself on stock macOS.** It bounded each probe with GNU `timeout`, absent there; the resulting rc=127 took the "probe inconclusive" branch, so the gate stood down on the first package for every macOS user — fail-open, so no false denials, but the advertised `safety` feature simply did not work, and said nothing. It now falls back to `gtimeout`, then to an unbounded probe. **Mitigation for what the configurator can't control:** the scaffold appends a `.gitattributes` block (`*.sh text eol=lf`, `claude-ctx`, `.claude/skills/**/scripts/*`) through the same managed-block append the `.gitignore` block uses — written when absent, line-level-unioned into a stale block, never touching the user's own rules. Git cannot record the *executable* bit that way, and `core.filemode=false` (the Windows default) drops it, so the generated `CLAUDE.md` § Repo bootstrap now tells Windows committers to run `git update-index --chmod=+x claude-ctx`. New `test/portability/` fixtures cover all of it by behavior, not by grepping source: every generated file is CRLF-free and every hook shebang is clean; the `.gitattributes` block is written, idempotent and union-safe; and the two hooks are run under a PATH that hides jq, python3 and timeout. README's runtime-dependency list now names jq and the timeout fallback; the platform table reflects what CI actually exercises. The `python-uv-fastapi` example and all five persona snapshots are regenerated (`.gitattributes` is new output). - **fix(install): the `ln -sf` shortcut silently froze on Windows; PATH-independent shim, honest prerequisites.** Under Git Bash/MSYS, `ln -s` **copies** unless `MSYS=winsymlinks:nativestrict` is set *and* the user has Developer Mode or admin — so `~/.local/bin/cc-configure` became a snapshot of configure.py taken on install day. `git -C ~/.cc-configurator pull` updated the clone while the command kept running the old copy, with nothing to indicate it. The installer now verifies the link actually resolved (`[ -L ]`) and otherwise writes a two-line `exec python3 /configure.py "$@"` shim, which tracks the clone on every platform; both checks live in the `if` condition, because under `set -e` a trailing `[ -L … ] && link_ok=1` in a then-block would abort the installer on precisely the platform the fallback exists for. Also: `git` is now checked up front (it was used before the first guard), the missing-python3 hint is OS-aware instead of always saying `apt install python3`, and the closing usage banner no longer advertises `--preset aggressive` and the legacy `commands-core` / `token-efficiency-pro` module IDs — both deprecated and slated for removal in v3.0. diff --git a/README.md b/README.md index 56d04c3..150fa12 100644 --- a/README.md +++ b/README.md @@ -131,8 +131,8 @@ There are 11 modules; legacy IDs (`commands-core`, `agents`, `token-efficiency-p | **git-workflow** | PostToolUse formatter on Write/Edit, Stop hook running typecheck / lint / tests. | | **token-efficiency** | Path-scoped `.claude/rules/` starters + PreCompact snapshot hook. `tier` flag: `basic` (default) ships discipline rules + snapshot only; `pro` adds bash-output truncation hook + always-loaded discipline rules. | | **commands** | Slash commands + agents + microbits. `subset` flag (linear ordering: `curated ⊂ full ⊂ rigorous`): **`curated`** = 3 essential skills (`/plan`, `/commit`, `/verify-setup`) + the `code-reviewer` agent. **`full`** (default) = 9 workflow skills (adds `/review-branch`, `/ship`, `/sync-docs`, `/check-context`, `/session-retro`, `/retrofit`) + 4 agents (`code-reviewer`, `test-runner`, `doc-writer`, `security-auditor`) + 4 discipline microbits (`/freeze`, `/unfreeze`, `/guard`, `/careful`) + the `microbit-enforcer.sh` PreToolUse hook. **`rigorous`** = `full` + `/investigate` + `/plan-eng-review`, the rigor skills that embed `templates/commands/_patterns/` cross-cutting blocks (confidence gate, independent verification, no-fix-without-investigation, AI-slop detection). **Naming note:** a project skill wins its name against Claude Code's built-ins (verified on CC 2.1.241). `/review-branch` is deliberately *not* called `review`, because CC 2.1.223 made `/review` the alias of the bundled `/code-review` and taking that name would hide Claude Code's multi-agent reviewer. `/plan` does still shadow the built-in plan-mode shortcut — plan mode remains reachable with Shift+Tab, so this is a name clash rather than a lost capability; rename the skill if you'd rather keep the shortcut. The `security-auditor` frontmatter wires Sonatype's dependency-management MCP (`https://mcp.guide.sonatype.com/mcp`) scoped to that agent — active only when it runs, so ~0 baseline context cost. Set `SONATYPE_TOKEN` env var to enable ([generate a token](https://guide.sonatype.com/settings/tokens)). | -| **mcp** | `.mcp.json` generated from selected servers, plus **per-task profiles** (`.mcp.research.json`, `.mcp.frontend.json`, `.mcp.minimal.json`) and an executable `./claude-ctx` wrapper that launches Claude with `--mcp-config --strict-mcp-config` — drops a bloated 4-MCP baseline from ~49% context to under 5%. | -| **multi-agent** | Path-scoped `multi-agent-guardrails.md` (5-scenario "when not to parallel" list), `/merge-worktrees` skill, `/infinite` skill, `parallel-generator` subagent. | +| **mcp** | `.mcp.json` generated from selected servers, plus **per-task profiles** (`.mcp.research.json`, `.mcp.frontend.json`, `.mcp.minimal.json`) and an executable `./claude-ctx` wrapper that launches Claude with `--mcp-config --strict-mcp-config`. Since tool search defers MCP schemas by default, profiles are about **which servers connect** — startup time, auth prompts, blast radius — not about reclaiming context: measured on CC 2.1.245, four servers × 12 tools cost **+696 tokens deferred vs +14,328 eagerly loaded** (`alwaysLoad: true`). See [`docs/04`](docs/04-subagents-mcp-orchestration.md#what-mcp-actually-costs). | +| **multi-agent** | Path-scoped `multi-agent-guardrails.md` (5-scenario "when not to parallel" list plus the runtime caps to design around), `/merge-worktrees` skill, `/infinite` skill, `parallel-generator` subagent. `/infinite` is scoped to its one good case — N variants of a single spec into disjoint slots; Claude Code's own dynamic workflows are the better tool for staged or merge-heavy fan-out, and the skill says so up front. | | **github-actions** | `.github/workflows/claude.yml` pinned to `anthropics/claude-code-action@v1`. Triggers on `@claude` mentions in issues, PR comments, and PR reviews. | | **ui** | Custom status line (project \| branch \| model \| ctx% \| OS+tool-version chip like `deb13 · pg17 · node20 · py3.13`; optionally effort + thinking indicators when 2.1.119+ is running), `statusline-last-prompt.sh` variant, and a "plan" output style. Sub-flag: `no_version_chip` omits the chip (sets `CC_STATUSLINE_NO_VERSION_CHIP=1`) for narrow terminals or privacy. | | **recommend-plugins** | Drops `docs/recommended-plugins.md` — a stack-aware list of official Claude Code plugins worth considering (always-recommended set + stack-specific picks computed from your form answers). Reference doc; refreshes on every `cc-configure` run. See [`docs/10-plugin-ecosystem.md`](docs/10-plugin-ecosystem.md) for how plugins relate to the configurator. | diff --git a/docs/04-subagents-mcp-orchestration.md b/docs/04-subagents-mcp-orchestration.md index 9581451..c73c8c2 100644 --- a/docs/04-subagents-mcp-orchestration.md +++ b/docs/04-subagents-mcp-orchestration.md @@ -106,20 +106,50 @@ Situationally valuable: - **playwright** — any frontend project. - **context7** — live library docs lookup. Stops hallucinated APIs. -Usually skip: "kitchen sink" servers exposing dozens of tools you won't use. Every tool definition costs tokens on every session start. +Usually skip: "kitchen sink" servers exposing dozens of tools you won't use — not because of tokens (see below) but because every extra server is another process, another auth prompt, and more surface for the model to reach for the wrong tool. -### Context bloat is real +### What MCP actually costs -Every MCP tool is a chunk of JSON schema loaded at session start. A heavy `.mcp.json` can burn thousands of tokens before you type anything — a fresh session with 4 MCP servers loaded typically costs ~37k tokens of tool descriptions alone, ~49% of a 100k window. Mitigations: +**Tool search changed this, and older advice (including earlier versions of this +page) is now wrong.** Claude Code defers MCP tool schemas by default and pulls +them in on demand, so a heavy `.mcp.json` no longer front-loads its schemas into +every prompt. `alwaysLoad: true` opts a server out of that deferral — it is the +setting that brings the old cost back. -1. **Only enable what you'll actually use this week.** -2. **Per-task profiles** — ship `.mcp..json` files at repo root and run `claude --mcp-config --strict-mcp-config` (or `./claude-ctx ` using the wrapper the `mcp` module ships). `--strict-mcp-config` ignores the default hierarchy entirely. Real demo: context usage drops from 18.8% to 2.4% when scoped. -3. **Scope to subagents** — put MCP servers in a subagent's frontmatter (`mcpServers:`) so they only load for that agent. -4. **Wrap heavy servers in narrow slash commands** — if a server exposes 30 tools and you use 3, write a skill that calls those 3. +Measured on Claude Code 2.1.245 (Opus 5, one-turn session, four local stdio +servers advertising twelve tools each — 48 tools total): + +| Session | Prompt tokens | MCP's share | +|---|---|---| +| No MCP servers | 26,665 | — | +| 4 servers, default (deferred) | 27,361 | +696 (~14 tokens/tool) | +| 4 servers with `alwaysLoad: true` | 40,993 | +14,328 (~298 tokens/tool) | + +Deferral removes ~95% of the schema cost. Two consequences: + +1. **Adding a server is cheap. Loading it eagerly is not.** Reach for + `alwaysLoad` only when a server's tools are used in nearly every turn. +2. **Per-task profiles are no longer primarily a token lever.** What they still + buy is real: fewer processes and auth prompts at startup, faster cold start, + and a smaller blast radius — `--strict-mcp-config` means only the listed + servers exist for that session. Choose a profile for those reasons. + +Still worth doing regardless of tokens: + +- **Only enable what you'll actually use this week.** Fewer servers, fewer ways + for a turn to go sideways. +- **Scope to subagents** — put MCP servers in a subagent's frontmatter + (`mcpServers:`) so they're only live while that agent runs. +- **Wrap heavy servers in narrow skills** — if a server exposes 30 tools and you + use 3, a skill that calls those 3 is easier for the model to aim. ### Checking the cost -`/context` in an active session shows what's loaded. `/cost` shows spend. Use both. +`/context` in an active session shows what's loaded, and it is the number to +trust for your own project — the table above is one synthetic shape, not a law. +The dominant line is the baseline itself (~26.7k tokens of system prompt and +built-in tools before any MCP server exists); no profile changes that. `/cost` +shows spend. ## Orchestration patterns @@ -135,6 +165,16 @@ Main session edits; `code-reviewer` subagent runs in parallel after every commit ### Desktop + cloud Local Claude Code for interactive work; Claude Code Web for long-running cloud jobs. Worktrees bridge the two. Useful for "run this large refactor while I go do something else." After both branches converge, use `/merge-worktrees` (ships in `multi-agent`) to integrate safely via a disposable branch. +### When to reach for a workflow instead + +Dynamic workflows (`/workflows`) run a script Claude writes over many agents, +with control flow, structured output between stages, a token budget, resume +after failure, and a progress view. Prefer them for staged pipelines +(find → verify → synthesize), for fan-out whose results need merging or scoring, +and for anything that should survive an interruption. Hand-rolled batching — +including the `/infinite` skill the `multi-agent` module ships — is the right +shape only for N independent variants of one spec written to disjoint slots. + ### Limits - The main thread's context is still finite. Subagents help with verbosity but not with total information you're holding in your head. diff --git a/docs/06-token-efficiency.md b/docs/06-token-efficiency.md index 0079a33..b5f29d7 100644 --- a/docs/06-token-efficiency.md +++ b/docs/06-token-efficiency.md @@ -10,7 +10,7 @@ Every session starts with: 2. **CLAUDE.md** (+ concatenated nested ones) — your doing. **Main lever.** 3. **Imported files via `@path`** — expand at launch. Same cost as if inlined. 4. **Auto-memory** — first 200 lines / 25 KB of MEMORY.md. Toggleable. -5. **MCP tool definitions** — every enabled server's tool schemas. **Hidden fat.** +5. **MCP tool definitions** — every enabled server's tool schemas. Mostly *not* fat any more: tool search defers them by default (~14 tokens/tool instead of ~298 — see [doc 04](04-subagents-mcp-orchestration.md#what-mcp-actually-costs)). It becomes fat again for any server you set `alwaysLoad: true` on. 6. **Available subagents' descriptions** — short, but they add up if you have many. 7. **Skills with `disable-model-invocation: false`** — descriptions only, not bodies. @@ -32,7 +32,7 @@ After that, each turn adds: Path-scoped rules only load when Claude reads matching files. Zero cost otherwise. Move anything domain-specific here (frontend rules, test rules, DB rules). See `templates/token-efficiency/dot-claude/rules/` for starters. ### 3. Audit MCP servers quarterly -Run `/context` in a session. Count the tokens eaten by MCP tool definitions. If a server exposes 30 tools and you use 3, either: +Run `/context` in a session. Count the tokens eaten by MCP tool definitions — with deferral on (the default) this should be small, and a large number means something is set to `alwaysLoad`. If a server exposes 30 tools and you use 3, either: - Remove it and call those 3 as Bash commands. - Scope it to a subagent that needs it. - Replace it with a targeted skill. @@ -53,7 +53,7 @@ Teach Claude (in CLAUDE.md rules) to prefer `grep` / structured search over `cat Common causes: -- **Over-general MCP config.** A broad server with many tools burns tokens every session whether you use it or not. +- **Over-general MCP config.** A broad server costs little while its tools stay deferred, but it still adds a process, an auth prompt, and more ways for the model to pick the wrong tool — and it burns tokens every session if you set `alwaysLoad: true`. - **Fat CLAUDE.md.** Every subdirectory loads its nested CLAUDE.md too when Claude reads files there. If every subdir has a 300-line CLAUDE.md, you're spending hugely. - **Eager skills with `disable-model-invocation: false`.** Claude scans skill descriptions at every turn to decide whether to invoke. Many skills with long descriptions = real token cost. - **Auto-memory gone wild.** `/memory` shows what's been saved. If MEMORY.md has years of stale notes, prune it or disable auto-memory. @@ -103,7 +103,7 @@ Since Claude Code 2.1.117, Pro/Max subscribers on Opus 4.6 and Sonnet 4.6 defaul - [ ] CLAUDE.md is under 200 lines. - [ ] Path-scoped rules handle all per-domain guidance. -- [ ] `/context` shows MCP tool definitions under ~3k tokens. +- [ ] `/context` shows MCP tool definitions under ~3k tokens (easy with deferral on; if it's far above, check for `alwaysLoad`). - [ ] No subagent descriptions over ~50 words. - [ ] `PreCompact` hook snapshots session state. - [ ] `/cost` and `/context` checked at least weekly. diff --git a/examples/python-uv-fastapi/.claude/skills/check-context/SKILL.md b/examples/python-uv-fastapi/.claude/skills/check-context/SKILL.md index fd82351..49bfe48 100644 --- a/examples/python-uv-fastapi/.claude/skills/check-context/SKILL.md +++ b/examples/python-uv-fastapi/.claude/skills/check-context/SKILL.md @@ -11,11 +11,11 @@ Run `/context` first (the built-in) and read the breakdown it prints. Then analy ## Budget guardrails -A fresh Claude Code session with 4 MCP servers loaded already burns ~49% of a 100k window before the user types anything — system prompt ~2.2k, system tool descriptions ~12k, MCP tool descriptions ~37k. That's the baseline to compare against. +Before the user types anything a session already carries the system prompt, built-in tool descriptions, CLAUDE.md, path-scoped rules and the skill listing. Measured on Claude Code 2.1.245 that floor was ~26.7k tokens with no MCP servers configured at all — that's the baseline to compare against, and no MCP profile changes it. MCP tool schemas are deferred by tool search unless a server sets `alwaysLoad: true`, so they should be a thin slice (~14 tokens/tool, versus ~298 when eagerly loaded). Flag the following: -- **MCP tool descriptions > 10%** of the window → too many MCP servers for this task. Recommend the user split per-task `.mcp.json.` files and run `claude --mcp-config --strict-mcp-config`. +- **MCP tool descriptions > 10%** of the window → with deferral on this is nearly impossible, so suspect a server with `alwaysLoad: true` (or a build predating tool search). Recommend dropping `alwaysLoad` first. Per-task `.mcp..json` files with `claude --mcp-config --strict-mcp-config` (or `./claude-ctx `) remain the right move when the goal is *fewer connected servers* — startup time, auth prompts, blast radius — rather than fewer tokens. - **Custom tools / skills > 5%** of the window → some skill descriptions are too verbose or too many user-invocable skills are loaded. Recommend auditing `allowed-tools`, `when_to_use`, and trimming descriptions. - **Memory (CLAUDE.md + @imports) > 10%** → CLAUDE.md has bloated. Propose moving path-scoped content into `.claude/rules/*.md` with `paths:` frontmatter so it only loads when relevant. - **Total before first turn > 40%** → the session will autocompact early. Combination of the above. diff --git a/templates/commands/check-context/SKILL.md b/templates/commands/check-context/SKILL.md index c4a0ac4..041c622 100644 --- a/templates/commands/check-context/SKILL.md +++ b/templates/commands/check-context/SKILL.md @@ -10,11 +10,11 @@ Run `/context` first (the built-in) and read the breakdown it prints. Then analy ## Budget guardrails -A fresh Claude Code session with 4 MCP servers loaded already burns ~49% of a 100k window before the user types anything — system prompt ~2.2k, system tool descriptions ~12k, MCP tool descriptions ~37k. That's the baseline to compare against. +Before the user types anything a session already carries the system prompt, built-in tool descriptions, CLAUDE.md, path-scoped rules and the skill listing. Measured on Claude Code 2.1.245 that floor was ~26.7k tokens with no MCP servers configured at all — that's the baseline to compare against, and no MCP profile changes it. MCP tool schemas are deferred by tool search unless a server sets `alwaysLoad: true`, so they should be a thin slice (~14 tokens/tool, versus ~298 when eagerly loaded). Flag the following: -- **MCP tool descriptions > 10%** of the window → too many MCP servers for this task. Recommend the user split per-task `.mcp.json.` files and run `claude --mcp-config --strict-mcp-config`. +- **MCP tool descriptions > 10%** of the window → with deferral on this is nearly impossible, so suspect a server with `alwaysLoad: true` (or a build predating tool search). Recommend dropping `alwaysLoad` first. Per-task `.mcp..json` files with `claude --mcp-config --strict-mcp-config` (or `./claude-ctx `) remain the right move when the goal is *fewer connected servers* — startup time, auth prompts, blast radius — rather than fewer tokens. - **Custom tools / skills > 5%** of the window → some skill descriptions are too verbose or too many user-invocable skills are loaded. Recommend auditing `allowed-tools`, `when_to_use`, and trimming descriptions. - **Memory (CLAUDE.md + @imports) > 10%** → CLAUDE.md has bloated. Propose moving path-scoped content into `.claude/rules/*.md` with `paths:` frontmatter so it only loads when relevant. - **Total before first turn > 40%** → the session will autocompact early. Combination of the above. diff --git a/templates/commands/infinite/SKILL.md b/templates/commands/infinite/SKILL.md index d26463f..4aabe56 100644 --- a/templates/commands/infinite/SKILL.md +++ b/templates/commands/infinite/SKILL.md @@ -7,7 +7,28 @@ allowed-tools: Read Grep Glob Bash(ls:*) Bash(mkdir:*) Bash(test:*) # Parallel spec expansion: `$ARGUMENTS` -Before anything touches disk, confirm **this is a parallelizable fanout task.** If the spec asks for sequential work, coupled features, or exploratory problem-solving, **refuse** and point at `.claude/rules/multi-agent-guardrails.md`. Parallel agents are a fanout multiplier, not a general speedup. +## First: is this the right tool? + +Claude Code ships **dynamic workflows** (`/workflows`, or asking for one in the +prompt), which orchestrate agents from a script Claude writes: real control +flow, structured output between stages, a token budget, resumability after a +failure, and a live progress view. For most fan-out work that is now the better +instrument, and this skill should not compete with it. + +| Situation | Use | +|---|---| +| N variants of **one spec**, each written to its own slot, no coordination | **this skill** | +| Stages that feed each other (find → verify → synthesize) | a dynamic workflow | +| Fan-out where results must be merged, deduped, or scored | a dynamic workflow | +| The run must survive an interruption and resume | a dynamic workflow | +| More agents than you want to hand-batch | a dynamic workflow | + +What this skill still does well is the narrow case it was written for: identical +task shape, disjoint output slots, deliberate diversification along one named +axis, and no dependency between iterations. If the work doesn't look like that, +stop and reach for a workflow instead. + +Then confirm **this is a parallelizable fanout task.** If the spec asks for sequential work, coupled features, or exploratory problem-solving, **refuse** and point at `.claude/rules/multi-agent-guardrails.md`. Parallel agents are a fanout multiplier, not a general speedup. ## Parse diff --git a/templates/experiments-memory/memory/experiments/2026-04-24-example-profile-budget.md b/templates/experiments-memory/memory/experiments/2026-04-24-example-profile-budget.md index a427ceb..8c630df 100644 --- a/templates/experiments-memory/memory/experiments/2026-04-24-example-profile-budget.md +++ b/templates/experiments-memory/memory/experiments/2026-04-24-example-profile-budget.md @@ -16,6 +16,17 @@ Delta: ~33.4k tokens reclaimed. Session context at first turn went from 49% → ## Conclusion Hypothesis held. The `--strict-mcp-config` flag is the key — without it the scoped file merges with the default hierarchy and the savings disappear. Codified as: default working mode for research/debugging sessions is now `./claude-ctx research`, not plain `claude`. `claude` without a profile is reserved for sessions that genuinely need filesystem+git+github simultaneously. +## Addendum (2026-08-25) — superseded by tool search + +Claude Code now defers MCP tool schemas by default (`alwaysLoad: true` opts a +server out), so the control condition this experiment measured no longer +happens on a current build. Re-measured with four servers advertising 48 tools +on CC 2.1.245: +696 tokens deferred versus +14,328 eagerly loaded, against a +~26.7k baseline with no servers at all. The conclusion below still holds for +`alwaysLoad` servers and for builds predating tool search; the headline saving +does not generalize. Left in place deliberately — an experiment log records what +was true when it ran, and the addendum is where the update belongs. + ## Follow-ups - Measure the `minimal` profile the same way — expected ~0 tokens for MCP, but is there overhead from the empty `mcpServers: {}` itself? - Try the `frontend` profile on a real UI bug to see if playwright alone is enough or if filesystem access is needed too. diff --git a/templates/mcp/claude-ctx.sh b/templates/mcp/claude-ctx.sh index ac5143c..b386893 100644 --- a/templates/mcp/claude-ctx.sh +++ b/templates/mcp/claude-ctx.sh @@ -5,11 +5,16 @@ # MCP hierarchy (user/project/local). Profiles live at .mcp..json in # the project root. # -# WHY this exists: a fresh session with 4 MCP servers loaded burns ~49% of a -# 100k context window before you type anything (system prompt ~2.2k + -# system tool descriptions ~12k + MCP tool descriptions ~37k). Scoping per -# task with --strict-mcp-config drops that cost to near-zero for tasks that -# don't need those servers. +# WHY this exists: --strict-mcp-config makes the listed servers the ONLY ones +# that exist for the session, so a profile controls which servers actually +# connect: fewer processes and auth prompts at startup, a faster cold start, +# and a smaller blast radius for a task that has no business touching them. +# +# Note on tokens: this used to be pitched as a context saving, and it no longer +# is. Tool search defers MCP tool schemas by default — measured on Claude Code +# 2.1.245, four servers advertising 48 tools between them cost +696 tokens +# deferred versus +14,328 with alwaysLoad: true. Use profiles for connection +# control; use `/context` if you want to see where the tokens really went. # # Usage: # ./claude-ctx research # loads only .mcp.research.json diff --git a/templates/mcp/profiles/mcp.minimal.json b/templates/mcp/profiles/mcp.minimal.json index fa77024..edb7372 100644 --- a/templates/mcp/profiles/mcp.minimal.json +++ b/templates/mcp/profiles/mcp.minimal.json @@ -1,4 +1,4 @@ { - "//": "Profile: minimal. No MCP servers. Use for pure writing/editing sessions where external tools are dead weight. Demonstrates the --strict-mcp-config savings cleanly (no ~37k of MCP tool descriptions loaded). Runs via: ./claude-ctx minimal", + "//": "Profile: minimal. No MCP servers. Use for pure writing/editing sessions where external tools are dead weight. The cleanest demonstration of --strict-mcp-config: no MCP servers connect at all, so nothing external can be reached from the session. Runs via: ./claude-ctx minimal", "mcpServers": {} } diff --git a/templates/mcp/servers-cookbook.md b/templates/mcp/servers-cookbook.md index faeed74..bf615ca 100644 --- a/templates/mcp/servers-cookbook.md +++ b/templates/mcp/servers-cookbook.md @@ -92,7 +92,7 @@ For everything else, leave `alwaysLoad` unset and let the model pull tools in as ## Per-task profiles with `claude-ctx` -A fresh session with 4 MCP servers loaded burns ~49% of a 100k context window before you type anything (system prompt ~2.2k + system tools ~12k + MCP tool descriptions ~37k). For tasks that don't need those servers, scope per-task instead of loading the full default set. +Profiles decide **which servers connect**, not how many tokens you spend. With tool search deferring schemas (see the `alwaysLoad` section above), four servers advertising 48 tools between them measured +696 tokens on Claude Code 2.1.245, versus +14,328 with `alwaysLoad: true`. What scoping still buys: fewer processes and auth prompts at startup, a faster cold start, and a session that physically cannot reach a server the task has no business touching — `--strict-mcp-config` ignores the default hierarchy entirely. Profiles live at `.mcp..json` in the project root. Three are shipped by default: