From a4bac73983a14e268a5d18b2fc425963b10af11d Mon Sep 17 00:00:00 2001 From: tigers1997 Date: Mon, 24 Aug 2026 16:12:23 -0400 Subject: [PATCH 1/2] chore(skills): sync discipline-skills v6.0.2 -> v6.3.0 (obra/superpowers) Four upstream releases since PR #85, all landing in the seven forked skills. v6.0.3 moved SDD's scratch files out of .git/ (Claude Code treats it as a protected path and denies agent writes, which blocked an implementer subagent mid-run) into a self-ignoring .superpowers/sdd/ working-tree directory, resolved by a new shared script scripts/sdd-workspace. v6.2.0 made that workspace plan-scoped (.superpowers/sdd//, so a follow-up plan can't read the previous plan's ledger as its own); review-package gained the plan file as its first argument. The review-fix loop now resumes the implementer instead of dispatching fresh, with a scoped re-review prompt (re-review-prompt.md, NEW) and a five-round circuit breaker. A library-wide compression campaign removed the Bottom Line / Key Principles / Advantages / Integration / "Why This Matters" sections, folding the load-bearing arguments into Excuse/Reality rationalization tables. finishing-a-development-branch no longer offers "Discard this work" in its menu (discard is explicit-request only), creates PRs with whichever forge tooling is present, and fixes a real bug where the worktree path was recomputed after cleanup had already changed directory. v6.3.0 teaches brainstorming to classify a request as spike / bounded / architectural and scale the ceremony to it (only the architectural path writes a spec; the approval gate never scales). SDD controllers now issue recorded rulings instead of stalling on plan conflicts, ledger the pre-flight conflict scan as a table, batch small same-shape tasks into one dispatch, and forbid implementers and reviewers from spawning their own subagents (duplicate review seats). Plans carry a Spec: pointer. finishing-a-development-branch stops and asks when `git worktree remove` refuses because of untracked files. Local edits re-applied per SYNC.md: `superpowers:` prefixes stripped (14 sites across three SKILL.md files), brainstorming's ## Visual Companion section and the visual-companion step of its *architectural* checklist removed (renumbered to 8; the spike/bounded lists never had one), executing-plans' subagents note reframed project-neutral. One NEW local edit: subagent-driven-development's final whole-branch review now points at the configurator's own code-reviewer subagent (.claude/agents/code-reviewer.md, with a task-reviewer-prompt.md fallback when the commands module isn't installed) instead of upstream's ../requesting-code-review/code-reviewer.md -- three digraph labels plus the ## Final Review paragraph. This retires the broken-link papercut carried since v5.1.0 and documented in SYNC.md. The former "remove the requesting-code-review / test-driven-development lines from ## Integration" edits are obsolete; upstream dropped that section in v6.2.0. Module paths 15 -> 17 (re-review-prompt.md, scripts/sdd-workspace); test-module-files-exist.sh and test-scaffold-installs-skills.sh updated (the scaffold test now asserts all three scripts ship executable); the three persona snapshots that include the module regenerated. SYNC.md pinned to v6.3.0 (2026-08-12) with a fresh delta paragraph, a CRLF note for Windows checkouts, and a renumbered canonical-edit list. NOTICE and docs/10-plugin-ecosystem.md still said v5.1.0; both bumped, along with the module description and README's carve-out paragraphs. The new script's index mode was set with `git add --chmod=+x` (this checkout has core.filemode=false). Claude Code compat: 2.1.116-2.1.150 (unchanged by this commit). Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01JVndNviHZSnbKJnWP7jFbV --- CHANGELOG.md | 2 + NOTICE | 2 +- README.md | 4 +- config_schema.py | 4 +- docs/10-plugin-ecosystem.md | 2 +- examples/persona-small-team/expected-tree.txt | 2 + .../expected-tree.txt | 2 + examples/persona-solo-newer/expected-tree.txt | 2 + templates/discipline-skills/SYNC.md | 127 +++- .../discipline-skills/brainstorming/SKILL.md | 127 +++- .../executing-plans/SKILL.md | 16 +- .../finishing-a-development-branch/SKILL.md | 212 +++--- .../subagent-driven-development/SKILL.md | 620 +++++++++++------- .../implementer-prompt.md | 21 +- .../re-review-prompt.md | 115 ++++ .../scripts/review-package | 25 +- .../scripts/sdd-workspace | 40 ++ .../scripts/task-brief | 9 +- .../task-reviewer-prompt.md | 29 +- .../using-git-worktrees/SKILL.md | 53 +- .../verification-before-completion/SKILL.md | 19 - .../discipline-skills/writing-plans/SKILL.md | 9 +- .../test-module-files-exist.sh | 2 + .../test-scaffold-installs-skills.sh | 9 +- 24 files changed, 945 insertions(+), 508 deletions(-) create mode 100644 templates/discipline-skills/subagent-driven-development/re-review-prompt.md create mode 100755 templates/discipline-skills/subagent-driven-development/scripts/sdd-workspace diff --git a/CHANGELOG.md b/CHANGELOG.md index 0358897..a757fbe 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -4,6 +4,8 @@ All notable changes to this project. Format: [Keep a Changelog](https://keepacha ## Unreleased +- **chore(skills): sync discipline-skills v6.0.2 → v6.3.0 (obra/superpowers).** Four upstream releases since #85, all landing in the seven forked skills. **v6.0.3** moved SDD's scratch files out of `.git/` (Claude Code denies agent writes there) into a self-ignoring `.superpowers/sdd/` working-tree directory resolved by a new shared script, `scripts/sdd-workspace`. **v6.2.0** made that workspace plan-scoped (`.superpowers/sdd//`; `scripts/review-package` now takes the plan file first: `review-package PLAN_FILE BASE HEAD`), restructured the review-fix loop to resume the implementer with a scoped re-review (`re-review-prompt.md`, NEW) and a five-round circuit breaker, ran a library-wide compression campaign (the Bottom Line / Key Principles / Advantages / Integration / "Why This Matters" sections are gone; `using-git-worktrees` and `finishing-a-development-branch` gained Excuse/Reality rationalization tables), dropped "Discard this work" from the finishing menu (discard is explicit-request-only), made PR creation forge-agnostic and fixed the worktree path being recomputed after the cleanup `cd`. **v6.3.0** teaches `brainstorming` to classify requests as spike / bounded / architectural and scale the ceremony (only the architectural path writes a spec; the approval gate never scales), has SDD controllers issue recorded rulings instead of stalling on plan conflicts (ledgered pre-flight scan table, batched same-shape tasks, a hard "implementers and reviewers never spawn subagents" contract in both prompt templates, a `**Spec:**` pointer in the `writing-plans` header, reviewers re-reading illegible evidence instead of re-running suites), and stops `finishing-a-development-branch` from `--force`-removing a worktree that holds untracked files. Local edits re-applied per SYNC.md: `superpowers:` prefixes stripped (14 sites across three SKILL.md files), the brainstorming `## Visual Companion` section and the visual-companion step of the *architectural* checklist removed (8 items; the spike/bounded lists have none), the executing-plans subagents note reframed project-neutral. **One new local edit:** SDD's final whole-branch review now points at the configurator's own `code-reviewer` subagent (`.claude/agents/code-reviewer.md`, with a `task-reviewer-prompt.md` fallback when the `commands` module isn't installed) instead of upstream's `../requesting-code-review/code-reviewer.md` — three digraph labels plus the `## Final Review` paragraph — which retires the broken-link papercut carried since v5.1.0. The former "remove the requesting-code-review / test-driven-development lines from `## Integration`" edits are obsolete (upstream dropped the section in v6.2.0). Module `paths:` 15 → 17 (`re-review-prompt.md`, `scripts/sdd-workspace`); `test-module-files-exist.sh` and `test-scaffold-installs-skills.sh` updated (the scaffold test now asserts all three scripts ship executable); the three persona snapshots that include the module regenerated; SYNC.md pinned to v6.3.0 (2026-08-12) with a fresh delta paragraph, a CRLF note for Windows checkouts, and a renumbered canonical-edit list; the module description, README carve-out paragraphs, `NOTICE` and `docs/10` (the last two still said v5.1.0) bumped. On a `core.filemode=false` checkout the new script's index mode was set with `git add --chmod=+x`. Upstream's v6.2.0 Windows fix (`"shell": "bash"` on its SessionStart hook) is adopted configurator-wide in the companion compat entry. + - **feat(hooks): `stop-run-checks` can run a check inside its docker-compose service (dogfood F3).** A containerized project's check loop no-op'd because the toolchain lives in the image, not on the host PATH (F2 warned about this). A `CHECKS` entry now takes an optional 3rd field — `label|command|service` — and when a service is named the Stop hook runs the check inside it: `docker compose exec -T ` if the service is up, else `docker compose run --rm ` (a throwaway instance of the *existing* service — no new container, no host rebuild, no path translation). The host-PATH/manifest guards are bypassed for container checks; infra-not-ready (no docker/compose, or service undefined) skips silently (fail-open), and a non-zero container exit is reported as a normal check FAIL. Two-field `label|command` entries — including host commands containing a literal `|` — are unchanged; an entry is container-bound only when it has ≥2 pipes and a bare-token final field. The F2 `[ STACK WARNINGS ]` note now points at this field. New `test/stop-run-checks/test-container-checks.sh` (run-vs-exec by service state, skip when compose/service absent, FAIL on non-zero, 2-field stays host-side) via a `docker` stub. **Deferred follow-up:** a conditional intake field that detects `docker-compose.yml` and pre-populates the service per check (this format is the substrate). No `tested_up_to` / `CC_VERSION` bump. - **fix(preflight): host-PATH gate handles path-like binaries (`./gradlew`); drop redundant `import shutil`; tighten the retrofit-migration test (review follow-ups to #88/#89).** Three non-blocking advisories from the split dogfood PRs: **(1)** `check_stack_reality`'s container-note gate used `shutil.which(b)`, which always returns `None` for a path-like binary like `./gradlew` (a project-local wrapper, never on PATH) — so a Gradle-wrapper project with a `Dockerfile`/compose and no root `build.gradle` got a spurious "toolchain lives in the container" note. The gate now file-checks path-like binaries (`(target_dir / b).exists()`) and only consults `shutil.which` for bare names. **(2)** the inline `import shutil` inside `check_stack_reality` was redundant (shutil is imported at module top) and is removed. **(3)** `test/retrofit-hooks/test-sessionstart-matcher-migration.sh` case 2 now asserts the marker-clear migrates *out* of a user's mixed matcherless group (appearing exactly once, under `startup|clear`), not merely that the user hook survives. New `./gradlew` case in `test/cc-manifest/test-stack-reality-preflight.sh`. No behavior change for non-path stacks. diff --git a/NOTICE b/NOTICE index 3f49051..6c8b7c0 100644 --- a/NOTICE +++ b/NOTICE @@ -9,7 +9,7 @@ version 3 (AGPL-3.0). See LICENSE for the full text. Bundled third-party code under different (compatible) licenses: * templates/discipline-skills/ - Forked from obra/superpowers v5.1.0 + Forked from obra/superpowers v6.3.0 (https://github.com/obra/superpowers) Copyright (c) 2025 Jesse Vincent Licensed under the MIT License diff --git a/README.md b/README.md index 6b649e9..20dfcaa 100644 --- a/README.md +++ b/README.md @@ -259,7 +259,7 @@ Additional patterns were distilled from the MIT-licensed [garrytan/gstack](https - The `security-auditor` agent's confidence gate (≥8), false-positive exclusion list, concrete-exploit requirement, and lightweight STRIDE checklist are distilled from gstack's `/cso` security-review skill. - The `slop-scan` PostToolUse hook in `templates/safety/` adapts the AI-slop pattern catalog from gstack. -The `discipline-skills` module ships a curated 7-skill subset (brainstorming, writing-plans, executing-plans, verification-before-completion, using-git-worktrees, subagent-driven-development, finishing-a-development-branch) forked from the MIT-licensed [obra/superpowers](https://github.com/obra/superpowers) v6.0.2 plugin by Jesse Vincent (© 2025 Jesse Vincent). The skill bodies are lightly edited: bare cross-references replace `superpowers:`-prefixed ones, the visual-companion section is stripped from brainstorming, the `requesting-code-review` and `test-driven-development` Integration lines are removed from `subagent-driven-development` (we don't ship those upstream skills), and a slim SessionStart bootstrap replaces the upstream `using-superpowers` injection. See `templates/discipline-skills/LICENSE` for the upstream MIT notice and `templates/discipline-skills/SYNC.md` for the maintainer-internal sync workflow. +The `discipline-skills` module ships a curated 7-skill subset (brainstorming, writing-plans, executing-plans, verification-before-completion, using-git-worktrees, subagent-driven-development, finishing-a-development-branch) forked from the MIT-licensed [obra/superpowers](https://github.com/obra/superpowers) v6.3.0 plugin by Jesse Vincent (© 2025 Jesse Vincent). The skill bodies are lightly edited: bare cross-references replace `superpowers:`-prefixed ones, the visual-companion section is stripped from brainstorming, `subagent-driven-development`'s final whole-branch review points at the configurator's `code-reviewer` subagent instead of the upstream `requesting-code-review` template (we don't ship that skill), and a slim SessionStart bootstrap replaces the upstream `using-superpowers` injection. See `templates/discipline-skills/LICENSE` for the upstream MIT notice and `templates/discipline-skills/SYNC.md` for the maintainer-internal sync workflow. The configurator, its modules, the preflight-check architecture (`--check` / `check_schema_url` / `check_hook_weight` / `check_github_remote`), the stack-preset system, and all other scaffolding are original work. @@ -269,6 +269,6 @@ The configurator, its modules, the preflight-check architecture (`--check` / `ch The configurator switched from MIT to AGPL-3.0 to close the SaaS-loophole used to fork copyleft projects into closed managed services (the "MongoDB on AWS" pattern). Self-hosting and use in open-source projects remain free; closed-source or commercial managed-service use requires a separate license. -Carve-out: `templates/discipline-skills/` remains under the **MIT License** (© 2025 Jesse Vincent), forked from the MIT-licensed [obra/superpowers](https://github.com/obra/superpowers) v6.0.2 plugin. The MIT terms travel with those files when users install the module into their own projects. See `NOTICE` for the full bundled-license breakdown and `templates/discipline-skills/LICENSE` for the upstream MIT notice. +Carve-out: `templates/discipline-skills/` remains under the **MIT License** (© 2025 Jesse Vincent), forked from the MIT-licensed [obra/superpowers](https://github.com/obra/superpowers) v6.3.0 plugin. The MIT terms travel with those files when users install the module into their own projects. See `NOTICE` for the full bundled-license breakdown and `templates/discipline-skills/LICENSE` for the upstream MIT notice. Past releases tagged before this commit remain available under the MIT license they shipped under; the AGPL-3.0 terms apply to all subsequent code. diff --git a/config_schema.py b/config_schema.py index 9214234..f54d683 100644 --- a/config_schema.py +++ b/config_schema.py @@ -236,7 +236,7 @@ { "id": "discipline-skills", "title": "Curated discipline skills (brainstorming, planning, verification, worktrees)", - "description": "Seven discipline skills forked from the MIT-licensed obra/superpowers v6.0.2 plugin: brainstorming, writing-plans, executing-plans, verification-before-completion, using-git-worktrees, subagent-driven-development, finishing-a-development-branch. Ships as project-level .claude/skills/ so it costs ~930 fewer SessionStart tokens than installing the full upstream plugin (no using-superpowers injection, no descriptions for the 7 unused upstream skills). Includes a slim SessionStart bootstrap that auto-suppresses when the upstream `superpowers` plugin is also installed. See docs/10-plugin-ecosystem.md and templates/discipline-skills/SYNC.md for the curation rationale and upstream-sync workflow.", + "description": "Seven discipline skills forked from the MIT-licensed obra/superpowers v6.3.0 plugin: brainstorming, writing-plans, executing-plans, verification-before-completion, using-git-worktrees, subagent-driven-development, finishing-a-development-branch. Ships as project-level .claude/skills/ so it costs ~930 fewer SessionStart tokens than installing the full upstream plugin (no using-superpowers injection, no descriptions for the 7 unused upstream skills). Includes a slim SessionStart bootstrap that auto-suppresses when the upstream `superpowers` plugin is also installed. See docs/10-plugin-ecosystem.md and templates/discipline-skills/SYNC.md for the curation rationale and upstream-sync workflow.", "paths": [ "discipline-skills/LICENSE", "discipline-skills/brainstorming/SKILL.md", @@ -249,8 +249,10 @@ "discipline-skills/subagent-driven-development/SKILL.md", "discipline-skills/subagent-driven-development/implementer-prompt.md", "discipline-skills/subagent-driven-development/task-reviewer-prompt.md", + "discipline-skills/subagent-driven-development/re-review-prompt.md", "discipline-skills/subagent-driven-development/scripts/review-package", "discipline-skills/subagent-driven-development/scripts/task-brief", + "discipline-skills/subagent-driven-development/scripts/sdd-workspace", "discipline-skills/finishing-a-development-branch/SKILL.md", "discipline-skills/hooks/sessionstart-discipline.sh", ], diff --git a/docs/10-plugin-ecosystem.md b/docs/10-plugin-ecosystem.md index 24a7111..6ab132e 100644 --- a/docs/10-plugin-ecosystem.md +++ b/docs/10-plugin-ecosystem.md @@ -86,7 +86,7 @@ You're trading "deterministic, fork-able templates I control" for "Anthropic-mai ## Discipline skills: bundled vs. upstream plugin -The configurator ships a `discipline-skills` module — a curated 7-skill subset forked from the MIT-licensed `obra/superpowers` v5.1.0 plugin: `brainstorming`, `writing-plans`, `executing-plans`, `verification-before-completion`, `using-git-worktrees`, `subagent-driven-development`, `finishing-a-development-branch`. They land at project-level `.claude/skills//SKILL.md`. A slim SessionStart hook (`sessionstart-discipline.sh`) primes the model with a terse seven-skill bootstrap. +The configurator ships a `discipline-skills` module — a curated 7-skill subset forked from the MIT-licensed `obra/superpowers` v6.3.0 plugin: `brainstorming`, `writing-plans`, `executing-plans`, `verification-before-completion`, `using-git-worktrees`, `subagent-driven-development`, `finishing-a-development-branch`. They land at project-level `.claude/skills//SKILL.md`. A slim SessionStart hook (`sessionstart-discipline.sh`) primes the model with a terse seven-skill bootstrap. **Why ship a fork instead of just recommending the upstream plugin:** diff --git a/examples/persona-small-team/expected-tree.txt b/examples/persona-small-team/expected-tree.txt index 28008d7..5df2fde 100644 --- a/examples/persona-small-team/expected-tree.txt +++ b/examples/persona-small-team/expected-tree.txt @@ -49,7 +49,9 @@ ./.claude/skills/ship/SKILL.md ./.claude/skills/subagent-driven-development/SKILL.md ./.claude/skills/subagent-driven-development/implementer-prompt.md +./.claude/skills/subagent-driven-development/re-review-prompt.md ./.claude/skills/subagent-driven-development/scripts/review-package +./.claude/skills/subagent-driven-development/scripts/sdd-workspace ./.claude/skills/subagent-driven-development/scripts/task-brief ./.claude/skills/subagent-driven-development/task-reviewer-prompt.md ./.claude/skills/sync-docs/SKILL.md diff --git a/examples/persona-solo-experienced/expected-tree.txt b/examples/persona-solo-experienced/expected-tree.txt index b35e78e..e1e674b 100644 --- a/examples/persona-solo-experienced/expected-tree.txt +++ b/examples/persona-solo-experienced/expected-tree.txt @@ -45,7 +45,9 @@ ./.claude/skills/ship/SKILL.md ./.claude/skills/subagent-driven-development/SKILL.md ./.claude/skills/subagent-driven-development/implementer-prompt.md +./.claude/skills/subagent-driven-development/re-review-prompt.md ./.claude/skills/subagent-driven-development/scripts/review-package +./.claude/skills/subagent-driven-development/scripts/sdd-workspace ./.claude/skills/subagent-driven-development/scripts/task-brief ./.claude/skills/subagent-driven-development/task-reviewer-prompt.md ./.claude/skills/sync-docs/SKILL.md diff --git a/examples/persona-solo-newer/expected-tree.txt b/examples/persona-solo-newer/expected-tree.txt index 46fbd09..a8a85ab 100644 --- a/examples/persona-solo-newer/expected-tree.txt +++ b/examples/persona-solo-newer/expected-tree.txt @@ -27,7 +27,9 @@ ./.claude/skills/plan/SKILL.md ./.claude/skills/subagent-driven-development/SKILL.md ./.claude/skills/subagent-driven-development/implementer-prompt.md +./.claude/skills/subagent-driven-development/re-review-prompt.md ./.claude/skills/subagent-driven-development/scripts/review-package +./.claude/skills/subagent-driven-development/scripts/sdd-workspace ./.claude/skills/subagent-driven-development/scripts/task-brief ./.claude/skills/subagent-driven-development/task-reviewer-prompt.md ./.claude/skills/using-git-worktrees/SKILL.md diff --git a/templates/discipline-skills/SYNC.md b/templates/discipline-skills/SYNC.md index 31f54e8..1dba044 100644 --- a/templates/discipline-skills/SYNC.md +++ b/templates/discipline-skills/SYNC.md @@ -6,11 +6,53 @@ upstream `obra/superpowers` plugin so we can keep them aligned over time. ## Source pin - Upstream: https://github.com/obra/superpowers -- Last synced from: `claude-plugins-official` marketplace, **v6.0.2** (released 2026-06-17) +- Last synced from: tag **v6.3.0** (released 2026-08-12; `claude-plugins-official` + marketplace entry pins the same commit, `b36e0829`) - License: MIT (see `LICENSE` in this directory) - Copyright: © 2025 Jesse Vincent -**v5.1.0 → v6.0.2 sync notes (2026-06-17).** Upstream's v6.0.0 was a major release: per-task reviewer prompts unified into one (`task-reviewer-prompt.md` replaces both `spec-reviewer-prompt.md` and `code-quality-reviewer-prompt.md`); two new bash scripts (`scripts/review-package` + `scripts/task-brief`) move diff and task text to files so they don't park permanently in the controller's context; the legacy global worktree directory (`~/.config/superpowers/worktrees/`) was dropped in favor of project-local `.worktrees/`; `writing-plans` gained Global Constraints + per-task Interfaces blocks; brainstorming's visual companion was rewritten with a real security model (we still strip it entirely); and prose across all skills moved from "Claude / Task tool" to "your agent / dispatch a subagent" so the skills work cross-harness (Claude Code, Codex, Copilot, Gemini, Pi, Antigravity). Our SYNC.md "embed code-reviewer.md inline in code-quality-reviewer-prompt.md" rewrite (v5.1.0 era) is RETIRED — the file no longer exists upstream and the unified `task-reviewer-prompt.md` has no external dependency. v6.0.1 and v6.0.2 are bug-fix releases on top of v6.0.0 (Codex bootstrap fixes, eval-submodule cleanup) with no impact on the 7 forked skills. +**v6.0.2 → v6.3.0 sync notes (2026-08-24).** Four upstream releases since the last +sync, all landing in the seven forked skills: + +- **v6.0.3 (2026-06-18)** — SDD scratch files moved out of `.git/` (Claude Code + denies agent writes there) into a self-ignoring `.superpowers/sdd/` working-tree + directory resolved by a new shared script, `scripts/sdd-workspace`; `task-brief` + and `review-package` now call it instead of `git rev-parse --git-path sdd`. +- **v6.1.0 / v6.1.1 (2026-06-30 / 07-02)** — `using-superpowers` bootstrap compressed, + per-harness tool references pruned, Codex packaging. No content change in the + seven forked skills; the bootstrap is a skill we replace with our own hook anyway. +- **v6.2.0 (2026-07-24)** — the big one. SDD's workspace is now **plan-scoped** + (`.superpowers/sdd//`; `review-package` gained the plan file as its + first argument: `review-package PLAN_FILE BASE HEAD`); the review-fix loop resumes + the implementer instead of dispatching fresh, gets a scoped re-review prompt + (`re-review-prompt.md`, NEW) and a five-round circuit breaker; SKILL.md is + reorganized by lifecycle. A library-wide **compression campaign** removed the + Bottom Line / Key Principles / Advantages / Integration / "Why This Matters" + sections from `brainstorming`, `verification-before-completion`, `executing-plans`, + `subagent-driven-development`, `using-git-worktrees` and `writing-plans`, folding + the load-bearing arguments into Excuse/Reality rationalization tables. + `finishing-a-development-branch` no longer offers "Discard this work" in its menu + (discard survives only as an explicit-request path), creates PRs with whichever + forge tooling is present, and fixes a real bug (worktree path captured before the + cleanup `cd`). Upstream's SessionStart hook also gained `"shell": "bash"` for + Windows (CC ≥ 2.1.81 resolves Git Bash directly instead of falling back to + PowerShell) — the configurator adopted that key on **every** shipped command hook + in the same release as this sync. +- **v6.3.0 (2026-08-12)** — `brainstorming` now classifies each request as + **spike / bounded / architectural** and scales the ceremony (only the + architectural path writes a spec and hands off to `writing-plans`; the approval + gate never scales). SDD controllers issue recorded **rulings** instead of stalling + on plan conflicts, the pre-flight conflict scan is written to the ledger as a + table, small same-shape tasks batch into one dispatch, implementers and reviewers + are forbidden from spawning their own subagents (duplicate review seats), plans + carry a `**Spec:**` pointer (`writing-plans` header), and reviewers re-read + illegible evidence instead of re-running suites. `finishing-a-development-branch` + stops and asks when `git worktree remove` refuses because of untracked files + (never `--force` on its own). + +Skill inventory unchanged (14/14 upstream; we still fork the same 7). Two files +added to the module's `paths:` (`re-review-prompt.md`, `scripts/sdd-workspace`), +15 → 17 paths. ## What we fork @@ -18,25 +60,38 @@ Seven of the upstream's fourteen skills: | Skill | File | Notes | |---|---|---| -| brainstorming | `brainstorming/SKILL.md` | Visual companion checklist item dropped + `## Visual Companion` section stripped + `visual-companion.md` and `scripts/` excluded from the module's `paths:` (we don't ship the browser server) | -| writing-plans | `writing-plans/SKILL.md` | `superpowers:` prefix stripped from cross-references | -| executing-plans | `executing-plans/SKILL.md` | `superpowers:` prefix stripped + the "Superpowers works much better with access to subagents" line reframed as project-neutral capability note; the parenthetical pointer to `../using-superpowers/references/` was dropped (we don't ship `using-superpowers`) | +| brainstorming | `brainstorming/SKILL.md` | `## Visual Companion` section stripped + the "Offer the visual companion just-in-time" step dropped from the **Architectural** checklist (renumbered to 8 items; the Spike and Bounded lists are untouched). `visual-companion.md` and `scripts/` excluded from the module's `paths:` (we don't ship the browser server) | +| writing-plans | `writing-plans/SKILL.md` | `superpowers:` prefix stripped from cross-references (5 sites in v6.3.0) | +| executing-plans | `executing-plans/SKILL.md` | `superpowers:` prefix stripped + the "Tell your human partner that Superpowers works much better…" note reframed as a project-neutral capability note without the harness list or the `../using-superpowers/references/` pointer (we don't ship `using-superpowers`). v6.2.0 dropped the upstream `## Integration` section, so nothing else is local | | verification-before-completion | `verification-before-completion/SKILL.md` | Verbatim | -| using-git-worktrees | `using-git-worktrees/SKILL.md` | Verbatim (v6.0.0 dropped the global `~/.config/superpowers/worktrees/` path; our fork inherits this stricter project-local convention cleanly) | -| subagent-driven-development | `subagent-driven-development/SKILL.md` | `superpowers:` prefix stripped (including the graphviz `Use superpowers:finishing-a-development-branch` green-fill node); the Integration section's `**requesting-code-review** - Code review template for the final whole-branch review` line REMOVED (we don't ship that skill); the entire `**Subagents should use:** test-driven-development` block REMOVED | -| finishing-a-development-branch | `finishing-a-development-branch/SKILL.md` | Verbatim (v6.0.0 also drops the global worktree path here; the v6.0.0 "forge-neutral" rewrite that drops hardcoded `gh pr create` is inherited cleanly) | +| using-git-worktrees | `using-git-worktrees/SKILL.md` | Verbatim | +| subagent-driven-development | `subagent-driven-development/SKILL.md` | `superpowers:` prefix stripped (6 sites, including the graphviz green-fill node `"Use finishing-a-development-branch"`). The **final whole-branch review** is pointed at the configurator's `code-reviewer` subagent instead of upstream's `../requesting-code-review/code-reviewer.md` (a skill we don't ship): three digraph node labels become `"Dispatch final code reviewer (code-reviewer subagent)"` and the `## Final Review` paragraph names `.claude/agents/code-reviewer.md` with a `task-reviewer-prompt.md` fallback when the `commands` module isn't installed. This resolves the papercut carried since v5.1.0. v6.2.0 dropped the upstream `## Integration` section, so the former "remove the requesting-code-review / test-driven-development lines" edits are obsolete | +| finishing-a-development-branch | `finishing-a-development-branch/SKILL.md` | Verbatim | Supporting prompt templates carried over: - `brainstorming/spec-document-reviewer-prompt.md` (verbatim) - `writing-plans/plan-document-reviewer-prompt.md` (verbatim) -- `subagent-driven-development/implementer-prompt.md` (verbatim — v6.0.0 rewrote it to read task-brief from file, require explicit `model:` field, and add RED/GREEN TDD evidence format; the rewrite is inherited as-is, no local edits) -- `subagent-driven-development/task-reviewer-prompt.md` (verbatim — NEW in v6.0.0, replaces the v5 split spec-reviewer + code-quality-reviewer files; no external dependency on `requesting-code-review`) - -Supporting scripts carried over (NEW in v6.0.0): -- `subagent-driven-development/scripts/review-package` (bash, executable; writes review diff package to a file the reviewer reads in one call) -- `subagent-driven-development/scripts/task-brief` (bash, executable; extracts one task's full text from a plan into a file the implementer reads in one call) - -Both scripts ship without a `.sh` extension (matching upstream); the configurator's exec-bit heuristic was widened in this sync to also honor the source file's user-execute bit, so the scripts arrive at user installs with `+x` set. +- `subagent-driven-development/implementer-prompt.md` (verbatim — v6.3.0 adds the + "You Do Not Dispatch Subagents" contract and the resume-on-findings fix flow) +- `subagent-driven-development/task-reviewer-prompt.md` (verbatim — v6.3.0 adds the + no-subagents contract, batched-brief file-by-file checking, and the + "evidence you cannot see is not evidence that doesn't exist" rule) +- `subagent-driven-development/re-review-prompt.md` (verbatim — NEW in v6.2.0; the + scoped re-review used by every fix round) + +Supporting scripts carried over (all bash, executable, no `.sh` extension — +matching upstream): +- `subagent-driven-development/scripts/sdd-workspace` (NEW in v6.0.3; resolves and + creates `/.superpowers/sdd//`, writing the self-ignoring + `.gitignore`) +- `subagent-driven-development/scripts/review-package` (signature changed in + v6.2.0: `PLAN_FILE BASE HEAD [OUTFILE]`) +- `subagent-driven-development/scripts/task-brief` + +The configurator's exec-bit heuristic honors the source file's user-execute bit, +so the scripts arrive at user installs with `+x`. On a Windows checkout +(`core.filemode=false`) a new script's index mode must be set explicitly — +`git add --chmod=+x ` — or it lands in the repo as `100644`. ## What we don't fork @@ -44,10 +99,10 @@ Seven upstream skills are intentionally excluded: - `using-superpowers` — replaced by our slimmer `hooks/sessionstart-discipline.sh` bootstrap - `dispatching-parallel-agents` — overlaps `multi-agent/dot-claude/rules/multi-agent-guardrails.md` -- `requesting-code-review` — overlaps configurator's `/review` skill + `code-reviewer` agent. **Known papercut**: `subagent-driven-development/SKILL.md` still has one "final whole-branch review" line (around line 268 post-v6 port) that references `../requesting-code-review/code-reviewer.md`; that link is broken in our fork (always was). Not addressed in this sync to keep the port scope tight; revisit if it surfaces in user reports. +- `requesting-code-review` — overlaps the configurator's `/review` skill + `code-reviewer` agent. The one upstream reference to it (SDD's final whole-branch review) is rewritten to use that agent — see the table above; no broken link remains - `receiving-code-review` — not surfaced today; reconsider in a later sync -- `systematic-debugging` — overlaps configurator's `/investigate` skill (and v6.0.0's "no longer trips extended-thinking" fix doesn't change the overlap) -- `test-driven-development` — TDD discipline is implicit in `writing-plans` task structure and `implementer-prompt.md`'s RED/GREEN evidence format (the latter is now v6-explicit) +- `systematic-debugging` — overlaps the configurator's `/investigate` skill +- `test-driven-development` — TDD discipline is implicit in `writing-plans` task structure and `implementer-prompt.md`'s RED/GREEN evidence format (v6.2.0 renamed its reference doc to `writing-good-tests.md`; still not shipped) - `writing-skills` — too meta for the configurator's default kit; skill authors install full superpowers Also explicitly excluded at the file level (subdirectories of forked skills we don't carry): @@ -57,31 +112,43 @@ Also explicitly excluded at the file level (subdirectories of forked skills we d ## Sync workflow -1. Watch `obra/superpowers` releases. Latest known: **v6.0.2 (2026-06-17)**. -2. When a new release ships, diff each of the seven forked skills: +1. Watch `obra/superpowers` releases. Latest known: **v6.3.0 (2026-08-12)**. +2. When a new release ships, diff each of the seven forked skills. Either point at + the plugin cache (`~/.claude/plugins/cache//superpowers//skills`) + or clone the tag: ```bash - SP=/home/bob/.claude/plugins/cache/claude-plugins-official/superpowers//skills + git clone --depth 1 --branch v https://github.com/obra/superpowers.git /tmp/sp + SP=/tmp/sp/skills for s in brainstorming writing-plans executing-plans verification-before-completion using-git-worktrees subagent-driven-development finishing-a-development-branch; do diff -ur templates/discipline-skills/$s $SP/$s done ``` + On a Windows checkout with `core.autocrlf=true`, textually identical files diff + as whole-file changes (CRLF vs LF) — compare with `diff --strip-trailing-cr` + before treating a file as changed. 3. For each meaningful upstream change, port the substance and re-apply the local edits documented below. Cosmetic edits (whitespace, anchor tweaks) skip. -4. **File-list change check.** If an upstream skill gained/lost files (e.g. v5→v6 dropped the two reviewer prompts and added `task-reviewer-prompt.md` + `scripts/`), update **both** the `paths:` list in `config_schema.py` (discipline-skills module) and the table above. -5. **Exec-bit check.** If new scripts ship without a `.sh` extension (like v6's `review-package` / `task-brief`), confirm they have exec bits in the upstream cache (`ls -l`), and verify the configurator's `executable` heuristic (`configure.py` around lines 1779 / 1799) still picks them up — the current rule honors `.sh`-suffix OR source-file user-execute bit. +4. **File-list change check.** If an upstream skill gained/lost files (v5→v6 dropped two reviewer prompts and added `task-reviewer-prompt.md` + `scripts/`; v6.0.3 added `scripts/sdd-workspace`; v6.2.0 added `re-review-prompt.md`), update **both** the `paths:` list in `config_schema.py` (discipline-skills module) and the table above, plus `test/discipline-skills/test-module-files-exist.sh` and `test-scaffold-installs-skills.sh`, then regenerate the persona `expected-tree.txt` fixtures. +5. **Exec-bit check.** If new scripts ship without a `.sh` extension, confirm they have exec bits in the upstream tree (`ls -l`), that the configurator's `executable` heuristic (`configure.py`, the two `collect_files` call sites) still picks them up, and that the git index mode is `100755` (`git ls-files -s`). 6. Update the "Source pin" line at the top of this file with the new version + date, and add a brief delta paragraph. 7. Mention the bump in the release CHANGELOG entry. 8. Run `test/discipline-skills/` to confirm no regressions (especially `test-no-superpowers-prefix.sh` and `test-module-files-exist.sh`). ## Local edits — the canonical list -When porting a new upstream release, re-apply ALL of these. The fast path is `sed -i 's/superpowers://g' …` over the three SKILL.md files that carry the prefix, plus the targeted Edits in items 1, 2, and 5 below. +When porting a new upstream release, re-apply ALL of these. The fast path is +`sed -i 's/superpowers://g'` over the three SKILL.md files that carry the prefix, +plus the targeted edits in items 1, 3, and 4 below. Verify with a grep over every +shipped file for `superpowers:`, `requesting-code-review`, `using-superpowers`, +`test-driven-development`, and `visual-companion` / `Visual Companion` — all five +must come back empty. -1. **brainstorming/SKILL.md** — remove the entire `## Visual Companion` section (last section of the file). Drop the `Offer the visual companion just-in-time` step from the numbered checklist (item 2 in upstream's list) and renumber the rest. The process-flow digraph in v6.0.2 no longer contains visual-companion nodes (upstream removed them), so the v5-era digraph edits are obsolete. +1. **brainstorming/SKILL.md** — remove the entire `## Visual Companion` section (last section of the file). In the `## Checklist`, drop the `Offer the visual companion just-in-time` step from the **Architectural** list (item 2 in upstream's v6.3.0 list) and renumber the rest (9 → 8 items). The Spike and Bounded lists have no visual-companion step. The process-flow digraph carries no visual-companion nodes. 2. **writing-plans/SKILL.md** — strip the `superpowers:` prefix from every cross-skill reference. -3. **executing-plans/SKILL.md** — strip the `superpowers:` prefix from every cross-skill reference. Reframe the "Tell your human partner that Superpowers works much better" line as a project-neutral capability note (upstream says "Superpowers works much better"; we say "This skill works much better") and drop the parenthetical pointer to `../using-superpowers/references/` (we don't ship `using-superpowers`). -4. **subagent-driven-development/SKILL.md** — strip `superpowers:` prefix everywhere (this also handles the graphviz green-fill node update from `"Use superpowers:finishing-a-development-branch"` to `"Use finishing-a-development-branch"`). In `## Integration → Required workflow skills`, REMOVE the `**requesting-code-review** - Code review template for the final whole-branch review` line (we don't ship that skill). REMOVE the entire `**Subagents should use:** test-driven-development` block (4 lines). -5. **No more reviewer-prompt embedding.** Pre-v6 the canonical edit list had a fifth item: rewrite `code-quality-reviewer-prompt.md` to embed `requesting-code-review/code-reviewer.md` inline. That edit is RETIRED — v6.0.0 unified the two reviewer prompts into `task-reviewer-prompt.md` (which has no external dependency), and the original `code-quality-reviewer-prompt.md` file is gone. If a future upstream re-introduces split reviewers, revisit. +3. **executing-plans/SKILL.md** — strip the `superpowers:` prefix from every cross-skill reference. Replace the whole `**Note:** Tell your human partner that Superpowers works much better…` sentence with: *"**Note:** This skill works much better with access to subagents. The quality of its work will be significantly higher when run on Claude Code (the platform this configurator targets). If subagents are available, use subagent-driven-development instead of this skill."* (drops the harness list and the `../using-superpowers/references/` pointer). +4. **subagent-driven-development/SKILL.md** — strip `superpowers:` everywhere (this also rewrites the graphviz green-fill node to `"Use finishing-a-development-branch"`). Then point the final whole-branch review at the configurator's agent: replace every `"Dispatch final code reviewer (../requesting-code-review/code-reviewer.md)"` node label (3 sites: declaration + two edges) with `"Dispatch final code reviewer (code-reviewer subagent)"`, and in `## Final Review` replace *"using requesting-code-review's [code-reviewer.md](../requesting-code-review/code-reviewer.md)."* with *"using this project's `code-reviewer` subagent (`.claude/agents/code-reviewer.md`, shipped by the configurator's `commands` module; if it isn't installed, dispatch a general-purpose subagent with [task-reviewer-prompt.md](task-reviewer-prompt.md) scoped to the whole branch)."* + +Retired edits (kept for archaeology): the v5-era "embed `code-reviewer.md` inline in `code-quality-reviewer-prompt.md`" rewrite (file gone since v6.0.0) and the v6.0.x "remove the `requesting-code-review` / `Subagents should use: test-driven-development` lines from `## Integration`" edits (section gone since v6.2.0). If a future upstream re-introduces either, revisit. ## Why we forked rather than depending on the upstream plugin -See `docs/10-plugin-ecosystem.md` § "Discipline skills: bundled vs. upstream plugin" for the full rationale. Short version: ~930 tokens saved per session, curation control (we pick what ships), rugpull immunity (upstream removed three slash commands in v5.1.0 — bit us once already), and the configurator's existing module pipeline gives us a clean distribution path. The v6.0.0 multi-harness expansion (Codex, Copilot, Pi, Antigravity) widens the gap between "what we ship" and "what the full plugin offers" — users wanting cross-harness support should install the upstream plugin; users on Claude Code who want a smaller, curated kit use ours. +See `docs/10-plugin-ecosystem.md` § "Discipline skills: bundled vs. upstream plugin" for the full rationale. Short version: ~930 tokens saved per session, curation control (we pick what ships), rugpull immunity (upstream removed three slash commands in v5.1.0 — bit us once already), and the configurator's existing module pipeline gives us a clean distribution path. The v6.x multi-harness expansion (Codex, Copilot, Cursor, Pi, Hermes, Devin, Antigravity) widens the gap between "what we ship" and "what the full plugin offers" — users wanting cross-harness support should install the upstream plugin; users on Claude Code who want a smaller, curated kit use ours. diff --git a/templates/discipline-skills/brainstorming/SKILL.md b/templates/discipline-skills/brainstorming/SKILL.md index f32ff73..9893f93 100644 --- a/templates/discipline-skills/brainstorming/SKILL.md +++ b/templates/discipline-skills/brainstorming/SKILL.md @@ -7,20 +7,91 @@ description: "You MUST use this before any creative work - creating features, bu Help turn ideas into fully formed designs and specs through natural collaborative dialogue. -Start by understanding the current project context, then ask questions one at a time to refine the idea. Once you understand what you're building, present the design and get user approval. +Start by classifying how much process the request needs, then work +through your path: understand the context, refine the idea, present a +design, and get your human partner's approval. -Do NOT invoke any implementation skill, write any code, scaffold any project, or take any implementation action until you have presented a design and the user has approved it. This applies to EVERY project regardless of perceived simplicity. +Do NOT invoke any implementation skill, write any code, scaffold any +project, or take any implementation action until you have told your +human partner what you intend and they have approved it. This applies +to EVERY task on EVERY path below — the ceremony scales with the task; +the approval gate never does. -## Anti-Pattern: "This Is Too Simple To Need A Design" - -Every project goes through this process. A todo list, a single-function utility, a config change — all of them. "Simple" projects are where unexamined assumptions cause the most wasted work. The design can be short (a few sentences for truly simple projects), but you MUST present it and get approval. +## Three Paths + +Before your first question, classify the request and say the +classification out loud — "this looks bounded, so I'll present a short +design here rather than write a spec" — so your human partner can +override it: + +- **Spike** — a feasibility question ("can we...", "is it possible...", + "quick and dirty is fine") whose output is an answer, not code you + keep. Present the question and what you'll try in 2-3 sentences, get + a nod, then find out as cheaply as correctness allows. No design + doc, no spec file. Report findings as a recommendation; anything you + built stays labeled throwaway. +- **Bounded** — a well-scoped change to code that already exists in + this repo: a new flag, a small endpoint, a one-file fix. + Understanding the kind of app is not enough — bounded means the flow + you are changing is already here to read. If there is no existing + flow to change, the task is not bounded. Ask the clarifying + questions that matter, present a short design IN CHAT (a few + sentences to a few short paragraphs), and STOP. Implementation + starts only after your human partner says yes to that design — a + bounded task's approval is as hard a gate as an architectural + one. No spec file, no implementation plan document. +- **Architectural** — new projects, new subsystems, changes that + restructure how components fit together or alter interfaces others + depend on. Follow the full process: questions, approaches, sectioned + design, written spec, then the writing-plans skill. + +When in doubt between two paths, take the heavier one. The ratchet is +one-way: hidden complexity discovered mid-task upgrades the path — +stop, say so, and step up. Nothing downgrades mid-task. + +## Anti-Pattern: "Too Simple To Need Approval" + +Every path ends with your human partner approving your intent before +implementation. A todo list, a single-function utility, a config +change — the design may be two sentences in chat, but you MUST present +it and get approval. "Simple" tasks are where unexamined assumptions +cause the most wasted work. What scales with simplicity is the +artifact, never the approval. + +## Red Flags + +| Thought | Reality | +|---------|---------| +| "This is too simple to need a design" | Simple means a short design, not no design. Two sentences in chat, then approval. | +| "I'll call it bounded and skip the spec" | Reaching for a label to skip work IS the doubt — take the heavier path. | +| "It's bounded and the design is obvious — I'll start while they read it" | The gate is the approval, not the design's length. Present, then stop until you hear yes. | +| "I understand this kind of app, so it's bounded" | Bounded measures the repo, not your familiarity. A new project has no existing flow — it is architectural. | +| "The spike works, so I'll keep the code" | A spike's output is an answer. Keeping the code is a new request — classify it. | +| "It grew, but I'm almost done — no need to re-classify" | Hidden complexity upgrades the path mid-task. Stop and say so. | +| "They approved the spike, so the follow-up change is approved too" | Each task gets its own classification and its own approval. | ## Checklist -You MUST create a task for each of these items and complete them in order: +Classify first, announce the path, then create a task for each item on +your path and complete them in order. + +**Spike:** +1. **Explore project context** — enough to frame the probe +2. **Present question + probe plan** — 2-3 sentences +3. **Get approval** — a nod is enough +4. **Investigate** — as cheaply as correctness allows +5. **Report findings** — a recommendation; label anything built as throwaway + +**Bounded:** +1. **Explore project context** — check files, docs, recent commits +2. **Ask clarifying questions** — one at a time, the ones that matter +3. **Present short design in chat** — approach, files touched, testing +4. **Get approval** — STOP and wait for an explicit yes; presenting the design and starting in the same breath is skipping the gate +5. **Implement** — proceed with the normal development workflow (TDD applies); no plan document +**Architectural:** 1. **Explore project context** — check files, docs, recent commits 2. **Ask clarifying questions** — one at a time, understand purpose/constraints/success criteria 3. **Propose 2-3 approaches** — with trade-offs and your recommendation @@ -34,6 +105,13 @@ You MUST create a task for each of these items and complete them in order: ```dot digraph brainstorming { + "Classify: spike / bounded / architectural" [shape=diamond]; + "Present question + probe (2-3 sentences)" [shape=box]; + "Ask clarifying questions (bounded)" [shape=box]; + "Present short design in chat" [shape=box]; + "Human approves?" [shape=diamond]; + "Investigate; report recommendation" [shape=doublecircle]; + "Implement via normal workflow (no plan doc)" [shape=doublecircle]; "Explore project context" [shape=box]; "Ask clarifying questions" [shape=box]; "Propose 2-3 approaches" [shape=box]; @@ -43,7 +121,17 @@ digraph brainstorming { "Spec self-review\n(fix inline)" [shape=box]; "User reviews spec?" [shape=diamond]; "Invoke writing-plans skill" [shape=doublecircle]; - + "Hidden complexity? Upgrade path" [shape=box]; + + "Classify: spike / bounded / architectural" -> "Present question + probe (2-3 sentences)" [label="spike"]; + "Classify: spike / bounded / architectural" -> "Ask clarifying questions (bounded)" [label="bounded"]; + "Classify: spike / bounded / architectural" -> "Explore project context" [label="architectural"]; + "Present question + probe (2-3 sentences)" -> "Human approves?"; + "Ask clarifying questions (bounded)" -> "Present short design in chat"; + "Present short design in chat" -> "Human approves?"; + "Human approves?" -> "Investigate; report recommendation" [label="spike: yes"]; + "Human approves?" -> "Implement via normal workflow (no plan doc)" [label="bounded: yes"]; + "Hidden complexity? Upgrade path" -> "Classify: spike / bounded / architectural"; "Explore project context" -> "Ask clarifying questions"; "Ask clarifying questions" -> "Propose 2-3 approaches"; "Propose 2-3 approaches" -> "Present design sections"; @@ -57,10 +145,21 @@ digraph brainstorming { } ``` -**The terminal state is invoking writing-plans.** Do NOT invoke frontend-design, mcp-builder, or any other implementation skill. The ONLY skill you invoke after brainstorming is writing-plans. +**Terminal states are path-bound.** Architectural: the ONLY skill you +invoke after brainstorming is writing-plans — never frontend-design, +mcp-builder, or any other implementation skill. Bounded: after +approval, implementation proceeds directly through the normal +development workflow; no plan document. Spike: the terminal state is a +reported recommendation. ## The Process +The subsections below serve the bounded and architectural paths (a +spike stops at "present the probe, get a nod"). Sections from +**Exploring approaches** onward are architectural-path depth — for +bounded work, context plus a few questions plus a short in-chat design +is the whole process. + **Understanding the idea:** - Check out the current project state first (files, docs, recent commits) @@ -76,6 +175,7 @@ digraph brainstorming { - Propose 2-3 different approaches with trade-offs - Present options conversationally with your recommendation and reasoning - Lead with your recommended option and explain why +- YAGNI ruthlessly - remove unnecessary features from every approach and design **Presenting the design:** @@ -98,7 +198,7 @@ digraph brainstorming { - Where existing code has problems that affect the work (e.g., a file that's grown too large, unclear boundaries, tangled responsibilities), include targeted improvements as part of the design - the way a good developer improves code they're working in. - Don't propose unrelated refactoring. Stay focused on what serves the current goal. -## After the Design +## After the Design (architectural path) **Documentation:** @@ -128,12 +228,3 @@ Wait for the user's response. If they request changes, make them and re-run the - Invoke the writing-plans skill to create a detailed implementation plan - Do NOT invoke any other skill. writing-plans is the next step. - -## Key Principles - -- **One question at a time** - Don't overwhelm with multiple questions -- **Multiple choice preferred** - Easier to answer than open-ended when possible -- **YAGNI ruthlessly** - Remove unnecessary features from all designs -- **Explore alternatives** - Always propose 2-3 approaches before settling -- **Incremental validation** - Present design, get approval before moving on -- **Be flexible** - Go back and clarify when something doesn't make sense diff --git a/templates/discipline-skills/executing-plans/SKILL.md b/templates/discipline-skills/executing-plans/SKILL.md index 0033699..31889b3 100644 --- a/templates/discipline-skills/executing-plans/SKILL.md +++ b/templates/discipline-skills/executing-plans/SKILL.md @@ -16,10 +16,11 @@ Load plan, review critically, execute all tasks, report when complete. ## The Process ### Step 1: Load and Review Plan -1. Read plan file -2. Review critically - identify any questions or concerns about the plan -3. If concerns: Raise them with your human partner before starting -4. If no concerns: Create todos for the plan items and proceed +1. Ensure an isolated workspace: use using-git-worktrees to create one or verify the existing one +2. Read plan file +3. Review critically - identify any questions or concerns about the plan +4. If concerns: Raise them with your human partner before starting +5. If no concerns: Create todos for the plan items and proceed ### Step 2: Execute Tasks @@ -61,10 +62,3 @@ After all tasks complete and verified: - Reference skills when plan says to - Stop when blocked, don't guess - Never start implementation on main/master branch without explicit user consent - -## Integration - -**Required workflow skills:** -- **using-git-worktrees** - Ensures isolated workspace (creates one or verifies existing) -- **writing-plans** - Creates the plan this skill executes -- **finishing-a-development-branch** - Complete development after all tasks diff --git a/templates/discipline-skills/finishing-a-development-branch/SKILL.md b/templates/discipline-skills/finishing-a-development-branch/SKILL.md index 7f5337a..fa8aeca 100644 --- a/templates/discipline-skills/finishing-a-development-branch/SKILL.md +++ b/templates/discipline-skills/finishing-a-development-branch/SKILL.md @@ -1,71 +1,58 @@ --- name: finishing-a-development-branch -description: Use when implementation is complete, all tests pass, and you need to decide how to integrate the work - guides completion of development work by presenting structured options for merge, PR, or cleanup +description: Use when implementation is complete, all tests pass, and you need to decide how to integrate the work --- # Finishing a Development Branch ## Overview -Guide completion of development work by presenting clear options and handling chosen workflow. - **Core principle:** Verify tests → Detect environment → Present options → Execute choice → Clean up. **Announce at start:** "I'm using the finishing-a-development-branch skill to complete this work." -## The Process +## Step 1: Verify Tests -### Step 1: Verify Tests +Run the project's full test suite (`npm test` / `cargo test` / `pytest` / `go test ./...`). -**Before presenting options, verify tests pass:** +**If tests fail**, report the failures and stop — the menu comes after a green suite: -```bash -# Run project's test suite -npm test / cargo test / pytest / go test ./... -``` - -**If tests fail:** ``` Tests failing ( failures). Must fix before completing: [Show failures] - -Cannot proceed with merge/PR until tests pass. ``` -Stop. Don't proceed to Step 2. +**If tests pass:** continue to Step 2. -**If tests pass:** Continue to Step 2. - -### Step 2: Detect Environment - -**Determine workspace state before presenting options:** +## Step 2: Detect Environment ```bash GIT_DIR=$(cd "$(git rev-parse --git-dir)" 2>/dev/null && pwd -P) GIT_COMMON=$(cd "$(git rev-parse --git-common-dir)" 2>/dev/null && pwd -P) +# Capture now, while still inside the workspace — Step 5 changes directory +# before cleanup (Step 6) needs this value +WORKTREE_PATH=$(git rev-parse --show-toplevel) ``` This determines which menu to show and how cleanup works: | State | Menu | Cleanup | |-------|------|---------| -| `GIT_DIR == GIT_COMMON` (normal repo) | Standard 4 options | No worktree to clean up | -| `GIT_DIR != GIT_COMMON`, named branch | Standard 4 options | Provenance-based (see Step 6) | -| `GIT_DIR != GIT_COMMON`, detached HEAD | Reduced 3 options (no merge) | No cleanup (externally managed) | +| `GIT_DIR == GIT_COMMON` (normal repo) | Standard 3 options | No worktree to clean up | +| `GIT_DIR != GIT_COMMON`, named branch | Standard 3 options | Provenance-based (see Step 6) | +| `GIT_DIR != GIT_COMMON`, detached HEAD | Reduced 2 options (no merge) | Externally managed — leave in place | -### Step 3: Determine Base Branch +## Step 3: Determine Base Branch -```bash -# Try common base branches -git merge-base HEAD main 2>/dev/null || git merge-base HEAD master 2>/dev/null -``` - -Or ask: "This branch split from main - is that correct?" +The base branch is whatever this work forked from — usually named in the +plan, the conversation, or the branch's upstream. If it is not already +known, ask: "This branch split from - is that correct?" +Confirm before merging: merging into the wrong base is expensive to undo. -### Step 4: Present Options +## Step 4: Present Options -**Normal repo and named-branch worktree — present exactly these 4 options:** +**Normal repo and named-branch worktree — present exactly these 3 options:** ``` Implementation complete. What would you like to do? @@ -73,28 +60,30 @@ Implementation complete. What would you like to do? 1. Merge back to locally 2. Push and create a Pull Request 3. Keep the branch as-is (I'll handle it later) -4. Discard this work Which option? ``` -**Detached HEAD — present exactly these 3 options:** +**Detached HEAD — present exactly these 2 options:** ``` Implementation complete. You're on a detached HEAD (externally managed workspace). 1. Push as new branch and create a Pull Request 2. Keep as-is (I'll handle it later) -3. Discard this work Which option? ``` -**Don't add explanation** - keep options concise. +Present the menu exactly as written — concise, with every option coming +from the list above. Discarding the work happens only in response to your +human partner explicitly asking for it (see "If your human partner asks to +discard the work" below). Wait for their answer; the integration decision +is theirs. -### Step 5: Execute Choice +## Step 5: Execute Choice -#### Option 1: Merge Locally +### Option 1: Merge Locally ```bash # Get main repo root for CWD safety @@ -108,34 +97,43 @@ git merge # Verify tests on merged result - -# Only after merge succeeds: cleanup worktree (Step 6), then delete branch ``` -Then: Cleanup worktree (Step 6), then delete branch: +If tests fail on the merged result: stop, leave the worktree and branch in +place, and investigate — nothing has been pushed, so the merge is local +and recoverable. + +Once the merged result is green: clean up the worktree (Step 6), then +delete the branch: ```bash git branch -d ``` -#### Option 2: Push and Create PR +### Option 2: Push and Create PR ```bash -# Push branch git push -u origin +# From a detached HEAD, name the new branch on the remote: +# git push origin HEAD:refs/heads/ ``` -**Do NOT clean up worktree** — user needs it alive to iterate on PR feedback. +Then create the pull/merge request against with the forge's +tooling — its CLI if one is available, or the creation URL most forges +print when you push — following the repo's PR template and conventions if +present, and report the URL to your human partner. -#### Option 3: Keep As-Is +Keep the worktree — your human partner iterates on PR feedback there. + +### Option 3: Keep As-Is Report: "Keeping branch . Worktree preserved at ." -**Don't cleanup worktree.** +### If your human partner asks to discard the work -#### Option 4: Discard +This path exists only as a response to an explicit request to throw the +work away. Confirm first: -**Confirm first:** ``` This will permanently delete: - Branch @@ -145,41 +143,62 @@ This will permanently delete: Type 'discard' to confirm. ``` -Wait for exact confirmation. +Wait for that exact confirmation. When it arrives: -If confirmed: ```bash MAIN_ROOT=$(git -C "$(git rev-parse --git-common-dir)/.." rev-parse --show-toplevel) cd "$MAIN_ROOT" ``` -Then: Cleanup worktree (Step 6), then force-delete branch: +Then clean up the worktree (Step 6) and force-delete the branch: + ```bash git branch -D ``` -### Step 6: Cleanup Workspace +## Step 6: Cleanup Workspace -**Only runs for Options 1 and 4.** Options 2 and 3 always preserve the worktree. - -```bash -GIT_DIR=$(cd "$(git rev-parse --git-dir)" 2>/dev/null && pwd -P) -GIT_COMMON=$(cd "$(git rev-parse --git-common-dir)" 2>/dev/null && pwd -P) -WORKTREE_PATH=$(git rev-parse --show-toplevel) -``` +**Runs for Option 1 and confirmed discards.** Options 2 and 3 always +preserve the worktree. Both callers have already changed directory to the +main repo root — worktree removal must run from outside the worktree — +and use the `GIT_DIR`/`GIT_COMMON`/`WORKTREE_PATH` values captured in +Step 2, from before that directory change. **If `GIT_DIR == GIT_COMMON`:** Normal repo, no worktree to clean up. Done. -**If worktree path is under `.worktrees/` or `worktrees/`:** Superpowers created this worktree — we own cleanup. +**If `WORKTREE_PATH` is under `.worktrees/` or `worktrees/`:** Superpowers +created this worktree — we own cleanup: ```bash -MAIN_ROOT=$(git -C "$(git rev-parse --git-common-dir)/.." rev-parse --show-toplevel) -cd "$MAIN_ROOT" git worktree remove "$WORKTREE_PATH" git worktree prune # Self-healing: clean up any stale registrations ``` -**Otherwise:** The host environment (harness) owns this workspace. Do NOT remove it. If your platform provides a workspace-exit tool, use it. Otherwise, leave the workspace in place. +**If removal is refused** (`contains modified or untracked files`): the +worktree holds files that exist nowhere else — uncommitted plans, notes, +or scratch work. Never `--force` on your own initiative. Show your human +partner what is at stake and ask: + +```bash +git -C "$WORKTREE_PATH" status --porcelain -uall +``` + +``` +Worktree removal refused — these files were never committed: + + + +1. Commit them to before cleanup +2. Move them into
+3. Delete them (unrecoverable) + +Which? +``` + +Carry out the choice, then remove the worktree. + +**Otherwise:** The host environment owns this workspace — leave it in +place. If your platform provides a workspace-exit tool, use it. ## Quick Reference @@ -188,54 +207,19 @@ git worktree prune # Self-healing: clean up any stale registrations | 1. Merge locally | yes | - | - | yes | | 2. Create PR | - | yes | yes | - | | 3. Keep as-is | - | - | yes | - | -| 4. Discard | - | - | - | yes (force) | - -## Common Mistakes - -**Skipping test verification** -- **Problem:** Merge broken code, create failing PR -- **Fix:** Always verify tests before offering options - -**Open-ended questions** -- **Problem:** "What should I do next?" is ambiguous -- **Fix:** Present exactly 4 structured options (or 3 for detached HEAD) - -**Cleaning up worktree for Option 2** -- **Problem:** Remove worktree user needs for PR iteration -- **Fix:** Only cleanup for Options 1 and 4 - -**Deleting branch before removing worktree** -- **Problem:** `git branch -d` fails because worktree still references the branch -- **Fix:** Merge first, remove worktree, then delete branch - -**Running git worktree remove from inside the worktree** -- **Problem:** Command fails silently when CWD is inside the worktree being removed -- **Fix:** Always `cd` to main repo root before `git worktree remove` - -**Cleaning up harness-owned worktrees** -- **Problem:** Removing a worktree the harness created causes phantom state -- **Fix:** Only clean up worktrees under `.worktrees/` or `worktrees/` - -**No confirmation for discard** -- **Problem:** Accidentally delete work -- **Fix:** Require typed "discard" confirmation - -## Red Flags - -**Never:** -- Proceed with failing tests -- Merge without verifying tests on result -- Delete work without confirmation -- Force-push without explicit request -- Remove a worktree before confirming merge success -- Clean up worktrees you didn't create (provenance check) -- Run `git worktree remove` from inside the worktree - -**Always:** -- Verify tests before offering options -- Detect environment before presenting menu -- Present exactly 4 options (or 3 for detached HEAD) -- Get typed confirmation for Option 4 -- Clean up worktree for Options 1 & 4 only -- `cd` to main repo root before worktree removal -- Run `git worktree prune` after removal +| Discard (explicit request only) | - | - | - | yes (force) | + +## Common Rationalizations + +| Excuse | Reality | +|--------|---------| +| "Tests passed earlier this session" | Run the suite on the tree you are about to integrate. A green run only proves the tree it ran on. | +| "They obviously want it merged" | Integration is your human partner's decision. Present the menu and wait. | +| "They seem done with this feature — I'll offer to discard it" | The menu is complete as written. Discard happens only when your human partner asks for it in so many words. | +| "'Yeah, get rid of it' counts as confirmation" | Only the typed word `discard` authorizes deletion. | +| "The PR is up, so the worktree is clutter now" | PR feedback gets fixed in that worktree. It stays until the work lands. | +| "This other worktree looks stale — I'll clean it too" | Clean up only worktrees under `.worktrees/` or `worktrees/`. Everything else belongs to the host. | +| "Removal refused — `--force` is just finishing the cleanup" | The refusal means files exist only in that worktree. `--force` destroys them permanently. Show your human partner and ask. | +| "The merged-result failure is probably flaky" | A failing merged result stops everything. Branch and worktree stay put while you investigate. | +| "The base branch is obviously main" | Confirm the fork point or ask. Merging into the wrong base is expensive to undo. | +| "The push was rejected — force-push will fix it" | A rejected push means the remote moved. Investigate; force-push only on your human partner's explicit request. | diff --git a/templates/discipline-skills/subagent-driven-development/SKILL.md b/templates/discipline-skills/subagent-driven-development/SKILL.md index bf7eaaa..f484d36 100644 --- a/templates/discipline-skills/subagent-driven-development/SKILL.md +++ b/templates/discipline-skills/subagent-driven-development/SKILL.md @@ -14,7 +14,21 @@ Execute plan by dispatching a fresh implementer subagent per task, a task review **Narration:** between tool calls, narrate at most one short line — the ledger and the tool results carry the record. -**Continuous execution:** Do not pause to check in with your human partner between tasks. Execute all tasks from the plan without stopping. The only reasons to stop are: BLOCKED status you cannot resolve, ambiguity that genuinely prevents progress, or all tasks complete. "Should I continue?" prompts and progress summaries waste their time — they asked you to execute the plan, so execute it. +**Continuous execution:** Do not pause to check in with your human partner between tasks. Execute all tasks from the plan without stopping. The only reasons to stop are the four named below, or all tasks complete. "Should I continue?" prompts and progress summaries waste their time — they asked you to execute the plan, so execute it. + +**Rulings, not stalls.** A running plan does not wait on a human. Conflicts, +ambiguities, plan defects, a cap you would have asked to exceed — decide +them. The spec is the binding authority, the plan is its argument, and your +judgment settles what neither answers. Record every decision in the ledger as +`Ruling: — — `, and keep +going. A wrong ruling costs rework your human partner can see and undo; a +session parked on a question costs their whole day and buys nothing. + +Four things stop you, and only these: an irreversible or destructive +operation; a security-sensitive action; a side effect outside this worktree +that norms say you ask about first (a merge, a push to a shared branch, a +publish); and a plan so broken that every path forward is a guess. For those, +stop and ask. ## When to Use @@ -51,50 +65,121 @@ digraph process { subgraph cluster_per_task { label="Per Task"; "Dispatch implementer subagent (./implementer-prompt.md)" [shape=box]; - "Implementer subagent asks questions?" [shape=diamond]; + "Implementer asks questions?" [shape=diamond]; "Answer questions, provide context" [shape=box]; - "Implementer subagent implements, tests, commits, self-reviews" [shape=box]; - "Write diff file, dispatch task reviewer subagent (./task-reviewer-prompt.md)" [shape=box]; - "Task reviewer reports spec ✅ and quality approved?" [shape=diamond]; - "Dispatch fix subagent for Critical/Important findings" [shape=box]; - "Mark task complete in todo list and progress ledger" [shape=box]; + "Implementer implements, tests, commits, self-reviews" [shape=box]; + "Generate review package, dispatch task reviewer (./task-reviewer-prompt.md)" [shape=box]; + "Spec ✅ and quality approved?" [shape=diamond]; + "Finding conflicts with plan text?" [shape=diamond]; + "Rule on the conflict, ledger the ruling" [shape=box]; + "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model" [shape=box]; + "Dispatch scoped re-review (./re-review-prompt.md)" [shape=box]; + "All findings addressed?" [shape=diamond]; + "R = 5?" [shape=diamond]; + "Adjudicate each open finding" [shape=box]; + "Any load-bearing finding?" [shape=diamond]; + "Rule and continue; stop only if every path forward is a guess" [shape=box]; + "Park findings in ledger with rulings" [shape=box]; + "Append completion to ledger, mark todo complete" [shape=box]; } - "Read plan, note context and global constraints, create todos" [shape=box]; + "Setup: worktree, ledger check, read plan, pre-flight review" [shape=box]; "More tasks remain?" [shape=diamond]; - "Dispatch final code reviewer subagent (../requesting-code-review/code-reviewer.md)" [shape=box]; + "Dispatch final code reviewer (code-reviewer subagent)" [shape=box]; + "Final findings? ONE fix dispatch, one scoped re-review, adjudicate residuals" [shape=box]; + "Final review clean: delete this plan's workspace" [shape=box]; "Use finishing-a-development-branch" [shape=box style=filled fillcolor=lightgreen]; - "Read plan, note context and global constraints, create todos" -> "Dispatch implementer subagent (./implementer-prompt.md)"; - "Dispatch implementer subagent (./implementer-prompt.md)" -> "Implementer subagent asks questions?"; - "Implementer subagent asks questions?" -> "Answer questions, provide context" [label="yes"]; - "Answer questions, provide context" -> "Dispatch implementer subagent (./implementer-prompt.md)"; - "Implementer subagent asks questions?" -> "Implementer subagent implements, tests, commits, self-reviews" [label="no"]; - "Implementer subagent implements, tests, commits, self-reviews" -> "Write diff file, dispatch task reviewer subagent (./task-reviewer-prompt.md)"; - "Write diff file, dispatch task reviewer subagent (./task-reviewer-prompt.md)" -> "Task reviewer reports spec ✅ and quality approved?"; - "Task reviewer reports spec ✅ and quality approved?" -> "Dispatch fix subagent for Critical/Important findings" [label="no"]; - "Dispatch fix subagent for Critical/Important findings" -> "Write diff file, dispatch task reviewer subagent (./task-reviewer-prompt.md)" [label="re-review"]; - "Task reviewer reports spec ✅ and quality approved?" -> "Mark task complete in todo list and progress ledger" [label="yes"]; - "Mark task complete in todo list and progress ledger" -> "More tasks remain?"; + "Setup: worktree, ledger check, read plan, pre-flight review" -> "Dispatch implementer subagent (./implementer-prompt.md)"; + "Dispatch implementer subagent (./implementer-prompt.md)" -> "Implementer asks questions?"; + "Implementer asks questions?" -> "Answer questions, provide context" [label="yes"]; + "Answer questions, provide context" -> "Implementer implements, tests, commits, self-reviews"; + "Implementer asks questions?" -> "Implementer implements, tests, commits, self-reviews" [label="no"]; + "Implementer implements, tests, commits, self-reviews" -> "Generate review package, dispatch task reviewer (./task-reviewer-prompt.md)"; + "Generate review package, dispatch task reviewer (./task-reviewer-prompt.md)" -> "Spec ✅ and quality approved?"; + "Spec ✅ and quality approved?" -> "Append completion to ledger, mark todo complete" [label="yes"]; + "Spec ✅ and quality approved?" -> "Finding conflicts with plan text?" [label="no"]; + "Finding conflicts with plan text?" -> "Rule on the conflict, ledger the ruling" [label="yes"]; + "Rule on the conflict, ledger the ruling" -> "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model"; + "Finding conflicts with plan text?" -> "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model" [label="no"]; + "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model" -> "Dispatch scoped re-review (./re-review-prompt.md)"; + "Dispatch scoped re-review (./re-review-prompt.md)" -> "All findings addressed?"; + "All findings addressed?" -> "Append completion to ledger, mark todo complete" [label="yes"]; + "All findings addressed?" -> "R = 5?" [label="no"]; + "R = 5?" -> "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model" [label="no - next round"]; + "R = 5?" -> "Adjudicate each open finding" [label="yes - breaker trips"]; + "Adjudicate each open finding" -> "Any load-bearing finding?"; + "Any load-bearing finding?" -> "Rule and continue; stop only if every path forward is a guess" [label="yes"]; + "Any load-bearing finding?" -> "Park findings in ledger with rulings" [label="no"]; + "Park findings in ledger with rulings" -> "Append completion to ledger, mark todo complete"; + "Append completion to ledger, mark todo complete" -> "More tasks remain?"; "More tasks remain?" -> "Dispatch implementer subagent (./implementer-prompt.md)" [label="yes"]; - "More tasks remain?" -> "Dispatch final code reviewer subagent (../requesting-code-review/code-reviewer.md)" [label="no"]; - "Dispatch final code reviewer subagent (../requesting-code-review/code-reviewer.md)" -> "Use finishing-a-development-branch"; + "More tasks remain?" -> "Dispatch final code reviewer (code-reviewer subagent)" [label="no"]; + "Dispatch final code reviewer (code-reviewer subagent)" -> "Final findings? ONE fix dispatch, one scoped re-review, adjudicate residuals"; + "Final findings? ONE fix dispatch, one scoped re-review, adjudicate residuals" -> "Final review clean: delete this plan's workspace"; + "Final review clean: delete this plan's workspace" -> "Use finishing-a-development-branch"; } ``` -## Pre-Flight Plan Review +## Setup -Before dispatching Task 1, scan the plan once for conflicts: +Ensure the work happens in an isolated workspace: use +using-git-worktrees to create one or verify the existing one. +Never start implementation on a main/master branch without your human +partner's explicit consent. + +Conversation memory does not survive compaction. In real sessions, +controllers that lost their place have re-dispatched entire completed task +sequences — the single most expensive failure observed. Track progress in +a ledger file, not only in todos. + +- Each plan owns a workspace: at skill start, run this skill's + `scripts/sdd-workspace PLAN_FILE` — it prints the plan's git-ignored + directory (`/.superpowers/sdd//`), home to + every artifact for THIS plan: ledger, briefs, reports, review packages. + Another plan's directory is never yours to read or write. +- Check for this plan's ledger at `/progress.md`. If its first + line names your plan file, tasks with a `Task : complete` line are DONE + — do not re-dispatch them; resume at the first task without one. A task + whose last line is a fix round is mid-loop: resume the loop at the next + round. A ledger whose first line names a different plan file — or a stray + ledger at the old flat path `.superpowers/sdd/progress.md` — is another + plan's progress: leave it in place and start your own, fresh. +- Create the ledger with its identity as the first line: + `# SDD ledger — plan: `. +- The ledger is your recovery map: the commits it names exist in git even + when your context no longer remembers creating them. After compaction, + trust the ledger and `git log` over your own recollection. +- `git clean -fdx` will destroy the workspace (it's git-ignored scratch); if + that happens, recover from `git log`. + +Read the plan once, note its context and Global Constraints, and create a +todo per task. If the plan names a Spec, read that too: the spec is the +authority the plan argues from, and conflicts inside the plan resolve +against it. A plan with no reachable spec gets a ledger note saying so — +rulings made without one are provisional. + +Before dispatching Task 1, scan the plan once for conflicts, writing down +what you checked as you check it: - tasks that contradict each other or the plan's Global Constraints - anything the plan explicitly mandates that the review rubric treats as a defect (a test that asserts nothing, verbatim duplication of a logic block) -Present everything you find to your human partner as one batched question — -each finding beside the plan text that mandates it, asking which governs — -before execution begins, not one interrupt per discovery mid-plan. If the -scan is clean, proceed without comment. The review loop remains the net for -conflicts that only emerge from implementation. +The scan's output is a table, not a verdict. One row for every pair of tasks +that share a file or an interface: the two tasks, what one produces against +what the other consumes, and what you found. One row for every task: whether +its own text agrees with itself — the tests it specifies against the code it +specifies, the files it creates against the files it later touches. "The scan +is clean" without those rows is not a scan you ran. + +Write the table to the ledger. Rule on everything you find before execution +begins — each finding against the plan text that mandates it — and record +each ruling in the ledger. If the scan is clean, proceed without comment. +Rule on each conflict it surfaces — the spec is the binding authority, the +plan is its argument — record the ruling beside its row, and dispatch +Task 1. The review loop remains the net for conflicts that only emerge from +implementation. ## Model Selection @@ -110,7 +195,11 @@ capable available model, not the session default. **Review tasks**: choose the model with the same judgment, scaled to the diff's size, complexity, and risk. A small mechanical diff does not need the -most capable model; a subtle concurrency change does. +most capable model; a subtle concurrency change does. Scoped re-reviews of +small fix diffs take a cheap-to-mid tier. + +**Fix-loop escalation (rounds 4-5)**: use a model at least one tier above +the implementer that got stuck. **Always specify the model explicitly when dispatching a subagent.** An omitted model inherits your session's model — often the most capable and @@ -129,11 +218,76 @@ that implementer. Single-file mechanical fixes also take the cheapest tier. - Touches multiple files with integration concerns → standard model - Requires design judgment or broad codebase understanding → most capable model -## Handling Implementer Status +## The Task Loop + +**Batch small same-shape work.** When the plan lists several tasks that are +each a small, independent edit of the same kind — the same one-line fix, +constant change, or field addition repeated across files — do not dispatch +one subagent per task. Compose ONE dispatch brief listing every file and +its change, send the whole batch to a single subagent, and review its diff +as one unit. Reserve one-dispatch-per-task for work that needs its own +judgment, its own tests, or its own review surface. + +Everything you paste into a dispatch prompt — and everything a subagent +prints back — stays resident in your context for the rest of the session +and is re-read on every later turn. Hand artifacts over as files. + +**Waiting on dispatched subagents:** never poll a wait interface with +short timeouts, and never sit in one silent, open-ended wait either. +While you have local work — ledger updates, packaging the next review, +reading reports — keep working; child results arrive on their own. +When you are genuinely idle, wait in bounded stretches (five to ten +minutes, where your platform allows), and between stretches post one +line of status and reconcile your live children: list them, and chase +any that finished without reporting. A bounded stretch keeps nearly +all of a long wait's efficiency while guaranteeing a stuck or lost +child is noticed within minutes, not at the end of the session. + +### 1. Dispatch the implementer + +Record BASE (`git rev-parse HEAD`) before dispatching — the review package +and fix-round diffs need it. + +- **Task brief:** before dispatching an implementer, run this skill's + `scripts/task-brief PLAN_FILE N` — it extracts the task's full text to a + uniquely named file and prints the path. Compose the dispatch so the + brief stays the single source of + requirements. Your dispatch should contain: (1) one line on where this + task fits in the project; (2) the brief path, introduced as "read this + first — it is your requirements, with the exact values to use verbatim"; + (3) interfaces and decisions from earlier tasks that the brief cannot + know; (4) your resolution of any ambiguity you noticed in the brief; + (5) the report-file path and report contract. Exact values (numbers, + magic strings, signatures, test cases) appear only in the brief. Never + make a subagent read the whole plan file. +- **Report file:** name the implementer's report file after the brief + (brief `…/task-N-brief.md` → report `…/task-N-report.md`) and put it in + the dispatch prompt. The implementer writes the full report there and + returns only status, commits, a one-line test summary, and concerns. +- A dispatch prompt describes one task, not the session's history. Do not + paste accumulated prior-task summaries ("state after Tasks 1-3") into + later dispatches — a real session's dispatch hit 42k chars of which 99% + was pasted history. A fresh subagent needs its task, the interfaces it + touches, and the global constraints. Nothing else. +- The dispatch carries the no-subagents contract (it is in the + implementer template): the implementer never dispatches subagents — + not helpers, and never a reviewer. Review arrives from you, after the + report. In real sessions, every reviewer a worker spawned duplicated + the task review the controller dispatched anyway — a full extra + review seat per task. +- If an earlier task parked a finding in the area this task touches, carry + a pointer to that ledger entry in the dispatch. +- Record the implementer's agent identity from the dispatch result — + fix-loop rounds 1-3 resume this agent. +- Never dispatch multiple implementation subagents in parallel (conflicts). + +Template: [implementer-prompt.md](implementer-prompt.md) + +### 2. Handle the report Implementer subagents report one of four statuses. Handle each appropriately: -**DONE:** Generate the review package (`scripts/review-package BASE HEAD`, from this skill's directory — it prints the unique file path it wrote; BASE is the commit you recorded before dispatching the implementer — never `HEAD~1`, which silently drops all but the last commit of a multi-commit task), then dispatch the task reviewer with the printed path. +**DONE:** Generate the review package (`scripts/review-package PLAN_FILE BASE HEAD`, from this skill's directory — it prints the unique file path it wrote; BASE is the commit you recorded before dispatching the implementer — never `HEAD~1`, which silently drops all but the last commit of a multi-commit task), then dispatch the task reviewer with the printed path. **DONE_WITH_CONCERNS:** The implementer completed the work but flagged doubts. Read the concerns before proceeding. If the concerns are about correctness or scope, address them before review. If they're observations (e.g., "this file is getting large"), note them and proceed to review. @@ -143,24 +297,41 @@ Implementer subagents report one of four statuses. Handle each appropriately: 1. If it's a context problem, provide more context and re-dispatch with the same model 2. If the task requires more reasoning, re-dispatch with a more capable model 3. If the task is too large, break it into smaller pieces -4. If the plan itself is wrong, escalate to the human +4. If the plan itself is wrong, rule on the correction, ledger it, and re-dispatch with the ruling carried in the dispatch **Never** ignore an escalation or force the same model to retry without changes. If the implementer said it's stuck, something needs to change. -## Handling Reviewer ⚠️ Items +If the implementer asks questions — before starting or mid-task — answer +clearly and completely, provide additional context if needed, and don't +rush it into implementation. -The task reviewer may report "⚠️ Cannot verify from diff" items — requirements -that live in unchanged code or span tasks. These do not block the rest of the -review, but you must resolve each one yourself before marking the task -complete: you hold the plan and cross-task context the reviewer -lacks. If you confirm an item is a real gap, treat it as a failed spec -review — send it back to the implementer and re-review. - -## Constructing Reviewer Prompts +### 3. Review the task Per-task reviews are task-scoped gates. The broad review happens once, at the -final whole-branch review. When you fill a reviewer template: +final whole-branch review. Never skip the task review, and never accept a +report missing either verdict — spec compliance AND task quality are both +required. Implementer self-review never replaces the task review; both are +needed. +- Hand the reviewer its diff as a file: run this skill's + `scripts/review-package PLAN_FILE BASE HEAD` and pass the reviewer the file path + it prints (or, without bash: `git log --oneline`, `git diff --stat`, + and `git diff -U10` for the range, redirected to one uniquely named + file). The output never enters your own context, and the reviewer sees + the commit list, stat summary, and full diff with context in one Read + call. Use the BASE you recorded before dispatching the implementer — + never `HEAD~1`, which silently truncates multi-commit tasks. Never + dispatch a task reviewer without a diff file. +- **Reviewer inputs:** the task reviewer gets three paths — the same brief + file, the report file, and the review package — plus the global + constraints that bind the task. +- The global-constraints block you hand the reviewer is its attention + lens. Copy the binding requirements verbatim from the plan's Global + Constraints section or the spec: exact values, exact formats, and the + stated relationships between components ("same layout as X", "matches + Y"). The reviewer's template already carries the process rules (YAGNI, + test hygiene, review method) — the constraints block is for what THIS + project's spec demands. - Do not add open-ended directives like "check all uses" or "run race tests if useful" without a concrete, task-specific reason - Do not ask a reviewer to re-run tests the implementer already ran on the @@ -171,108 +342,174 @@ final whole-branch review. When you fill a reviewer template: loop. If the prompt you are writing contains "do not flag," "don't treat X as a defect," "at most Minor," or "the plan chose" — stop: you are pre-judging, usually to spare yourself a review loop. -- The global-constraints block you hand the reviewer is its attention - lens. Copy the binding requirements verbatim from the plan's Global - Constraints section or the spec: exact values, exact formats, and the - stated relationships between components ("same layout as X", "matches - Y"). The reviewer's template already carries the process rules (YAGNI, - test hygiene, review method) — the constraints block is for what THIS - project's spec demands. -- Hand the reviewer its diff as a file: run this skill's - `scripts/review-package BASE HEAD` and pass the reviewer the file path - it prints (or, without bash: `git log --oneline`, `git diff --stat`, - and `git diff -U10` for the range, redirected to one uniquely named - file). The output never enters your own context, and the reviewer sees - the commit list, stat summary, and full diff with context in one Read - call. Use the BASE you recorded before dispatching the implementer — - never `HEAD~1`, which silently truncates multi-commit tasks. -- A dispatch prompt describes one task, not the session's history. Do not - paste accumulated prior-task summaries ("state after Tasks 1-3") into - later dispatches — a real session's dispatch hit 42k chars of which 99% - was pasted history. A fresh subagent needs its task, the interfaces it - touches, and the global constraints. Nothing else. -- Dispatch fix subagents for Critical and Important findings. Record Minor - findings in the progress ledger as you go, and point the final - whole-branch review at that list so it can triage which must be fixed - before merge. A roll-up nobody reads is a silent discard. -- A finding labeled plan-mandated — or any finding that conflicts with - what the plan's text requires — is the human's decision, like any plan - contradiction: present the finding and the plan text, ask which governs. - Do not dismiss the finding because the plan mandates it, and do not - dispatch a fix that contradicts the plan without asking. -- The final whole-branch review gets a package too: run - `scripts/review-package MERGE_BASE HEAD` (MERGE_BASE = the commit the - branch started from, e.g. `git merge-base main HEAD`) and include the - printed path in the final review dispatch, so the final reviewer reads - one file instead of re-deriving the branch diff with git commands. -- Every fix dispatch carries the implementer contract: the fix subagent - re-runs the tests covering its change and reports the results. Name the - covering test files in the dispatch — a one-line fix does not need the - whole suite. Before re-dispatching the reviewer, confirm the fix report - contains the covering tests, the command run, and the output; dispatch - the re-review once all three are present. -- If the final whole-branch review returns findings, dispatch ONE fix - subagent with the complete findings list — not one fixer per finding. - Per-finding fixers each rebuild context and re-run suites; a real - session's final-review fix wave cost more than all its tasks combined. - -## File Handoffs - -Everything you paste into a dispatch prompt — and everything a subagent -prints back — stays resident in your context for the rest of the session -and is re-read on every later turn. Hand artifacts over as files: - -- **Task brief:** before dispatching an implementer, run this skill's - `scripts/task-brief PLAN_FILE N` — it extracts the task's full text to a - uniquely named file and prints the path. Compose the dispatch so the - brief stays the single source of requirements. Your dispatch should - contain: (1) one line on where this task fits in the project; (2) the - brief path, introduced as "read this first — it is your requirements, - with the exact values to use verbatim"; (3) interfaces and decisions - from earlier tasks that the brief cannot know; (4) your resolution of - any ambiguity you noticed in the brief; (5) the report-file path and - report contract. Exact values (numbers, magic strings, signatures, test - cases) appear only in the brief. -- **Report file:** name the implementer's report file after the brief - (brief `…/task-N-brief.md` → report `…/task-N-report.md`) and put it in - the dispatch prompt. The implementer writes the full report there and - returns only status, commits, a one-line test summary, and concerns. -- **Reviewer inputs:** the task reviewer gets three paths — the same brief - file, the report file, and the review package — plus the global - constraints that bind the task. -- Fix dispatches append their fix report (with test results) to the same - report file and return a short summary; re-reviews read the updated file. +The task reviewer may report "⚠️ Cannot verify from diff" items — requirements +that live in unchanged code or span tasks. These do not block the rest of the +review, but you must resolve each one yourself before marking the task +complete: you hold the plan and cross-task context the reviewer +lacks. If you confirm an item is a real gap, treat it as a failed spec +review — it enters the fix loop with the other findings. -## Durable Progress +Template: [task-reviewer-prompt.md](task-reviewer-prompt.md) -Conversation memory does not survive compaction. In real sessions, -controllers that lost their place have re-dispatched entire completed task -sequences — the single most expensive failure observed. Track progress in -a ledger file, not only in todos. +### 4. The fix loop -- At skill start, check for a ledger: - `cat "$(git rev-parse --git-path sdd)/progress.md"`. Tasks listed there - as complete are DONE — do not re-dispatch them; resume at the first task - not marked complete. -- When a task's review comes back clean, append one line to the ledger in - the same message as your other bookkeeping: - `Task N: complete (commits .., review clean)`. -- The ledger is your recovery map: the commits it names exist in git even - when your context no longer remembers creating them. After compaction, - trust the ledger and `git log` over your own recollection. +The loop triggers when the review reports spec ❌, any Critical or Important +finding, or a ⚠️ item you confirmed as a real gap. -## Prompt Templates +Before the loop starts, two routes leave it immediately: -- [implementer-prompt.md](implementer-prompt.md) - Dispatch implementer subagent -- [task-reviewer-prompt.md](task-reviewer-prompt.md) - Dispatch task reviewer subagent (spec compliance + code quality) -- Final whole-branch review: use requesting-code-review's [code-reviewer.md](../requesting-code-review/code-reviewer.md) +- Record Minor findings in the progress ledger as you go + (`Task : minor (deferred): `), and point the final + whole-branch review at that list so it can triage which must be fixed + before merge. A roll-up nobody reads is a silent discard. Minor findings + never enter the loop. +- A finding labeled plan-mandated — or any finding that conflicts with + what the plan's text requires — is yours to rule on: weigh the finding + against the plan text, decide with the spec as the binding authority, and + ledger the ruling before you act on it. Do not dismiss the finding because + the plan mandates it, and do not dispatch a fix that contradicts the plan + without a recorded ruling. +Everything else enters the loop. A fix round is one fix dispatch plus one +scoped re-review. Five rounds maximum per task: + +**Rounds 1-3 — resume the original implementer.** Send it the open findings +verbatim. Its context is intact: it knows the task, the code, and its own +choices. If your harness cannot send another message to a live subagent, +dispatch a fresh implementer carrying the brief path, the report-file path, +and the findings — the report file is the persistent memory either way. + +**Rounds 4-5 — dispatch a fresh implementer on a more capable model** (per +Model Selection), with the brief path, the report-file path, the open +findings, and this framing: "A prior implementer attempted this task +[N] times; you own it now. Read the report file for what was tried." A loop +that survives three resumes usually means the implementer cannot see its +own problem — fresh eyes and a capability bump in one move. + +**Every round, either way:** the implementer fixes, re-runs the tests +covering the amended code, appends its fix report to the same report file, +and returns the short contract. Before re-dispatching the reviewer, confirm +the fix report contains the covering tests, the command run, and the +output; dispatch the re-review once all three are present. Name the +covering test files in the fix message — a one-line fix does not need the +whole suite. + +**The re-review is scoped.** Run `scripts/review-package PLAN_FILE FIX_BASE HEAD` +where FIX_BASE is the head the previous review saw, and dispatch +[re-review-prompt.md](re-review-prompt.md) with the findings list, the +brief, the report file, and the printed diff path. The re-reviewer verdicts +each finding ADDRESSED or NOT ADDRESSED and flags new breakage in the fix +diff only. New Critical/Important breakage in the fix diff joins the open +findings list. Out-of-scope observations go to the ledger as deferred +minors — they never extend the loop. + +**After each round,** append to the ledger: +`Task : fix round /5 ( addressed, open — ; commits ..)` + +Never fix findings yourself in the controller session — your context stays +clean for coordination, and controller fixes skip review. + +**The breaker.** When round 5's re-review still leaves findings open, stop +dispatching. Adjudicate each open finding yourself — you hold the plan and +the cross-task context the reviewer lacks: + +- **The reviewer is wrong, or the point is contestable:** park it — + `Task : parked — — Ruling: `. The final + review sees both sides. +- **Real, but nothing downstream builds on it:** park it the same way, with + a ruling that says it's real and deferred. +- **Real and load-bearing** — a later task builds on it, or it reveals a + plan defect: rule on the smallest change that unblocks the dependent work, + ledger it as `Task : Ruling: — `, + and carry it into the next task's dispatch. Parking a structural failure + silently lets every dependent task build on it. Stop only when the defect + leaves every path forward a guess. + +Adjudicate only at the cap. Adjudicating earlier to end a loop is +pre-judging with a different name. Every adjudication is a ledger entry — +a silent discard is forbidden. + +### 5. Complete the task + +When the review comes back clean — or every open finding is parked with a +ruling at the cap — append the completion line to the ledger in the same +message as your other bookkeeping: + +- `Task : complete (commits .., review clean)` +- `Task : complete (commits .., parked)` after a + tripped breaker + +Then mark the todo complete and move on. Never move to the next task while +the review has open Critical/Important issues that are neither fixed nor +parked-with-ruling at the cap. + +## Final Review + +The final whole-branch review gets a package too: run +`scripts/review-package PLAN_FILE MERGE_BASE HEAD` (MERGE_BASE = the commit the +branch started from, e.g. `git merge-base main HEAD`) and include the +printed path in the final review dispatch, so the final reviewer reads +one file instead of re-deriving the branch diff with git commands. Dispatch +on the most capable available model (see Model Selection), using this +project's `code-reviewer` subagent (`.claude/agents/code-reviewer.md`, shipped by +the configurator's `commands` module; if it isn't installed, dispatch a +general-purpose subagent with [task-reviewer-prompt.md](task-reviewer-prompt.md) +scoped to the whole branch). Point it at +the ledger's deferred-minor and parked lines so it can triage which must be +fixed before merge. + +If the final whole-branch review returns findings, dispatch ONE fix subagent +with the complete findings list — not one fixer per finding. +Per-finding fixers each rebuild context and re-run suites; a real +session's final-review fix wave cost more than all its tasks combined. +Then run exactly one scoped re-review of the fix wave +(`scripts/review-package PLAN_FILE FIX_BASE HEAD` over the fix range, +[re-review-prompt.md](re-review-prompt.md)). +Adjudicate any residual findings as in the task loop's breaker: park with +rulings, or rule on the load-bearing ones and ledger what you decided. Only +the four classes above stop you here. There is no second fix wave — +residual load-bearing findings surface to your human partner when +finishing-a-development-branch presents the options. + +## Finish + +Before you delete anything, collect every ledger line containing `Ruling:` — +preflight rulings, parked findings, breaker adjudications, all of them — into +your final message under "Rulings I made", in the order you made them, each +with what it costs if wrong. The list is exhaustive: if the ledger holds a +ruling, the list holds it. That list is the only place the decisions you +took on your human partner's behalf reach them — they read it and rework +whatever you got wrong. A ruling that dies with the workspace was a decision +made in secret. + +When the final whole-branch review is clean and its fixes are merged, +delete this plan's workspace (`rm -rf `) — the git history is +the record now. Sibling directories belong to other plans; leave them +alone. + +Use finishing-a-development-branch. + +## Common Rationalizations + +| Excuse | Reality | +|--------|---------| +| "Close enough on spec compliance" | Reviewer found spec gaps = not done. Fix or hit the cap and adjudicate — those are the only exits. | +| "I'll fix it myself, dispatching is overhead" | Controller fixes pollute your context and skip review. Resume the implementer. | +| "One more round will converge" | Past the cap, rounds don't converge — the failure is structural. Adjudicate and route. | +| "The reviewer will just find something new anyway" | Scoped re-reviews verify fixes; they cannot wander. New findings on untouched code go to the ledger, not the loop. | +| "This finding is obviously wrong, I'll drop it" | You adjudicate only at the cap, and every ruling is a ledger entry. Silent discards are forbidden. | +| "The fix was small, skip the re-review" | Unreviewed fixes are how regressions land. Every round ends with a scoped re-review. | +| "Reviews slow the loop down" | The loop without reviews is just unverified churn. Reviews are the loop's brakes and steering. | +| "Ledger bookkeeping is overhead" | The ledger is what survives compaction. Controllers without one have re-dispatched entire completed task sequences. | +| "The implementer spawned its own reviewer — free extra assurance" | It's a duplicate seat reviewing the same diff; the task review is the gate. A worker-spawned reviewer is a defect to flag, not rigor. | ## Example Workflow ``` You: I'm using Subagent-Driven Development to execute this plan. +[Setup: worktree verified] [Read plan file once: docs/superpowers/plans/feature-plan.md] +[Resolve workspace: scripts/sdd-workspace docs/superpowers/plans/feature-plan.md — no ledger inside, fresh start] [Create todos for all tasks] Task 1: Hook installation script @@ -283,130 +520,51 @@ Implementer: "Before I begin - should the hook be installed at user or system le You: "User level (~/.config/superpowers/hooks/)" -Implementer: "Got it. Implementing now..." -[Later] Implementer: +Implementer: [Later] - Implemented install-hook command - Added tests, 5/5 passing - Self-review: Found I missed --force flag, added it - Committed -[Run review-package, dispatch task reviewer with the printed path] +[Run review-package PLAN_FILE BASE HEAD; dispatch task reviewer with the printed path] Task reviewer: Spec ✅ - all requirements met, nothing extra. Strengths: Good test coverage, clean. Issues: None. Task quality: Approved. -[Mark Task 1 complete] +[Ledger: Task 1: complete (commits a1b2c3d..d4e5f6a, review clean)] Task 2: Recovery modes [Run task-brief for Task 2; dispatch implementer with brief + report paths + context] -Implementer: [No questions, proceeds] -Implementer: +Implementer: [No questions] - Added verify/repair modes - 8/8 tests passing - - Self-review: All good - Committed -[Run review-package, dispatch task reviewer with the printed path] +[Run review-package PLAN_FILE BASE HEAD; dispatch task reviewer with the printed path] Task reviewer: Spec ❌: - Missing: Progress reporting (spec says "report every 100 items") - - Extra: Added --json flag (not requested) Issues (Important): Magic number (100) -[Dispatch fix subagent with all findings] -Fixer: Removed --json flag, added progress reporting, extracted PROGRESS_INTERVAL constant +[Fix round 1: resume the implementer with both findings] +Implementer: Added progress reporting, extracted PROGRESS_INTERVAL constant. + Re-ran test/recovery.test.js — 10/10 passing. Fix report appended. -[Task reviewer reviews again] -Task reviewer: Spec ✅. Task quality: Approved. +[Run review-package PLAN_FILE FIX_BASE HEAD; dispatch scoped re-review] +Re-reviewer: Missing progress reporting — ADDRESSED (src/recovery.js:41). + Magic number — ADDRESSED (src/recovery.js:7). New breakage: none. + Verdict: all findings addressed. -[Mark Task 2 complete] +[Ledger: Task 2: fix round 1/5 (2 addressed, 0 open; commits d4e5f6a..b7c8d9e)] +[Ledger: Task 2: complete (commits d4e5f6a..b7c8d9e, review clean)] ... [After all tasks] -[Dispatch final code-reviewer] -Final reviewer: All requirements met, ready to merge +[Run review-package PLAN_FILE MERGE_BASE HEAD; dispatch final code-reviewer, most capable model] +Final reviewer: All requirements met. Deferred minors triaged: none block merge. -Done! -``` +[Delete this plan's workspace — the record now lives in git] -## Advantages - -**vs. Manual execution:** -- Subagents follow TDD naturally -- Fresh context per task (no confusion) -- Parallel-safe (subagents don't interfere) -- Subagent can ask questions (before AND during work) - -**vs. Executing Plans:** -- Same session (no handoff) -- Continuous progress (no waiting) -- Review checkpoints automatic - -**Efficiency gains:** -- Controller curates exactly what context is needed; bulk artifacts move - as files, not pasted text -- Subagent gets complete information upfront -- Questions surfaced before work begins (not after) - -**Quality gates:** -- Self-review catches issues before handoff -- Task review carries two verdicts: spec compliance and code quality -- Review loops ensure fixes actually work -- Spec compliance prevents over/under-building -- Code quality ensures implementation is well-built - -**Cost:** -- More subagent invocations (implementer + reviewer per task) -- Controller does more prep work (extracting all tasks upfront) -- Review loops add iterations -- But catches issues early (cheaper than debugging later) - -## Red Flags - -**Never:** -- Start implementation on main/master branch without explicit user consent -- Skip task review, or accept a report missing either verdict (spec compliance AND task quality are both required) -- Proceed with unfixed issues -- Dispatch multiple implementation subagents in parallel (conflicts) -- Make a subagent read the whole plan file (hand it its task brief — - `scripts/task-brief` — instead) -- Skip scene-setting context (subagent needs to understand where task fits) -- Ignore subagent questions (answer before letting them proceed) -- Accept "close enough" on spec compliance (reviewer found spec issues = not done) -- Skip review loops (reviewer found issues = implementer fixes = review again) -- Let implementer self-review replace actual review (both are needed) -- Tell a reviewer what not to flag, or pre-rate a finding's severity in the - dispatch prompt ("treat it as Minor at most") — the plan's example code is - a starting point, not evidence that its weaknesses were chosen -- Dispatch a task reviewer without a diff file — generate it first - (`scripts/review-package BASE HEAD`) and name the printed path in the - prompt -- Move to next task while the review has open Critical/Important issues -- Re-dispatch a task the progress ledger already marks complete — check - the ledger (and `git log`) after any compaction or resume - -**If subagent asks questions:** -- Answer clearly and completely -- Provide additional context if needed -- Don't rush them into implementation - -**If reviewer finds issues:** -- Implementer (same subagent) fixes them -- Reviewer reviews again -- Repeat until approved -- Don't skip the re-review - -**If subagent fails task:** -- Dispatch fix subagent with specific instructions -- Don't try to fix manually (context pollution) - -## Integration - -**Required workflow skills:** -- **using-git-worktrees** - Ensures isolated workspace (creates one or verifies existing) -- **writing-plans** - Creates the plan this skill executes -- **finishing-a-development-branch** - Complete development after all tasks - -**Alternative workflow:** -- **executing-plans** - Use for parallel session instead of same-session execution +Done! Using finishing-a-development-branch. +``` diff --git a/templates/discipline-skills/subagent-driven-development/implementer-prompt.md b/templates/discipline-skills/subagent-driven-development/implementer-prompt.md index 218fcfe..5c8ecd6 100644 --- a/templates/discipline-skills/subagent-driven-development/implementer-prompt.md +++ b/templates/discipline-skills/subagent-driven-development/implementer-prompt.md @@ -47,6 +47,18 @@ Subagent (general-purpose): While iterating, run the focused test for what you're changing; run the full suite once before committing, not after every edit. + ## You Do Not Dispatch Subagents + + Do all of this task's work yourself. Never spawn a subagent to + implement part of the task, and above all never spawn a reviewer to + check your work. Self-review (below) means reading your own diff. + Review is the controller's job: after you report, it dispatches a + fresh reviewer against your diff. A reviewer you spawn duplicates + that review at full cost, and its approval counts for nothing in + the process. If you catch yourself thinking "an independent review + would strengthen my report" — that review is already scheduled. + Report instead. + ## Code Organization You reason best about code you can hold in context at once, and your edits are more @@ -106,9 +118,12 @@ Subagent (general-purpose): ## After Review Findings - If a reviewer finds issues and you fix them, re-run the tests that cover - the amended code and append the results to your report file. Reviewers - will not re-run tests for you — your report is the test evidence. + If the task review finds issues, you will be resumed with the findings. + Fix them, re-run the tests that cover the amended code, and append a fix + report to your report file: what you changed, the covering tests you + ran, the command, and the output. Reviewers will not re-run tests for + you — your report is the test evidence. Then reply with the same short + status contract as your first report. ## Report Format diff --git a/templates/discipline-skills/subagent-driven-development/re-review-prompt.md b/templates/discipline-skills/subagent-driven-development/re-review-prompt.md new file mode 100644 index 0000000..ad74b10 --- /dev/null +++ b/templates/discipline-skills/subagent-driven-development/re-review-prompt.md @@ -0,0 +1,115 @@ +# Scoped Re-Review Prompt Template + +Use this template when dispatching a re-review after a fix round. The +re-reviewer verifies the findings were addressed and checks the fix diff for +new breakage. It is not a fresh review — the full review already happened. + +**Purpose:** Verify each finding from the previous review was addressed, and +that the fix itself broke nothing. + +``` +Subagent (general-purpose): + description: "Re-review Task N fix round R" + model: [MODEL — REQUIRED: choose per SKILL.md Model Selection; an omitted + model silently inherits the session's most expensive one] + prompt: | + You are re-reviewing one task's fix round. A previous review produced + findings; an implementer has attempted to fix them. Your job is to + verdict each finding and inspect the fix diff — nothing else. + + ## The Task + + Read the task brief: [BRIEF_FILE] + + ## The Findings Under Verification + + [FINDINGS] + + ## The Fix + + Read the implementer's report (fix reports are appended at the end): + [REPORT_FILE] + + **Fix base:** [FIX_BASE_SHA] (the head the previous review saw) + **Head:** [HEAD_SHA] + **Diff file:** [DIFF_FILE] + + Read the diff file once — it contains the fix commits, a stat summary, + and the fix diff with surrounding context. Do not re-run git commands. + If the diff file is missing, fetch the diff yourself: + `git diff --stat [FIX_BASE_SHA]..[HEAD_SHA]` and + `git diff [FIX_BASE_SHA]..[HEAD_SHA]`. + + Your review is read-only on this checkout. Do not mutate the working + tree, the index, HEAD, or branch state in any way. + + ## You Do Not Dispatch Subagents + + Do all of this review yourself. Never spawn a subagent to review part + of the diff, and never spawn another reviewer for a second opinion. + This process already provides every review seat the work gets; a + reviewer you spawn duplicates one of them at full cost, and its + verdict counts for nothing. If the diff feels too large for one + pass, review it in passes yourself and say so in your report. + + ## Scope + + Your scope is the findings list and the fix diff. Verdict every finding. + Inspect the fix diff for new problems the fix itself introduced. Do NOT + re-review code the fix did not touch: if you notice an issue entirely + outside the fix diff, report it under Out-of-Scope Observations — it + does not block this task and does not extend the loop. A broad + whole-branch review happens after all tasks are complete. + + ## Tests + + The implementer re-ran the tests covering the amended code and appended + the results to the report file. Treat the report as unverified claims: + confirm the fix report names the covering tests and shows their output, + and verify the claims against the diff. Do not re-run the suite to + confirm their report. Run a test only when reading the code raises a + specific doubt that no existing run answers — and then a focused test, + never a package-wide suite. + + ## Output Format + + Your final message is the report itself: begin directly with the first + finding's verdict. Every line is a verdict, a finding with file:line, + or a check you ran — no preamble, no process narration. + + ### Finding Verdicts + + For each finding in The Findings Under Verification, in order: + - **[finding one-liner]** — ADDRESSED | NOT ADDRESSED, with file:line + evidence. "Attempted" is not addressed: the specific defect must no + longer exist. + + ### New Breakage in the Fix Diff + + Anything the fix itself broke or introduced, with severity + (Critical/Important/Minor) and file:line. "None" if clean. + + ### Out-of-Scope Observations + + Issues you noticed entirely outside the fix diff. Non-blocking; the + controller ledgers these for the final review. "None" if none. + + ### Verdict + + **Fix round:** [All findings addressed, no new Critical/Important + breakage | Findings remain open] — list the open ones. +``` + +**Placeholders:** +- `[MODEL]` — REQUIRED: reviewer model per SKILL.md Model Selection; scoped + re-reviews of small fix diffs take a cheap-to-mid tier +- `[BRIEF_FILE]` — the task brief file (same file the implementer worked from) +- `[FINDINGS]` — the Critical/Important findings and spec gaps from the + previous review, copied verbatim, one per bullet +- `[REPORT_FILE]` — the implementer's report file (fix reports appended) +- `[FIX_BASE_SHA]` — the head the previous review saw +- `[HEAD_SHA]` — current commit +- `[DIFF_FILE]` — the path `scripts/review-package PLAN_FILE FIX_BASE HEAD` printed + +**Re-reviewer returns:** per-finding verdicts (ADDRESSED / NOT ADDRESSED), +new breakage in the fix diff, out-of-scope observations, and a round verdict. diff --git a/templates/discipline-skills/subagent-driven-development/scripts/review-package b/templates/discipline-skills/subagent-driven-development/scripts/review-package index 88a0022..31852e2 100755 --- a/templates/discipline-skills/subagent-driven-development/scripts/review-package +++ b/templates/discipline-skills/subagent-driven-development/scripts/review-package @@ -4,29 +4,28 @@ # call. Using the recorded per-task BASE (not HEAD~1) keeps multi-commit # tasks intact. # -# Usage: review-package BASE HEAD [OUTFILE] -# Default OUTFILE: /sdd/review-...diff — unique per -# repo instance and per range, so concurrent sessions cannot collide and a -# re-review after fixes always gets a distinctly named fresh file. +# Usage: review-package PLAN_FILE BASE HEAD [OUTFILE] +# Default OUTFILE: /.superpowers/sdd//review-...diff +# (named per range, so a re-review after fixes gets a distinct fresh file). set -euo pipefail -if [ $# -lt 2 ] || [ $# -gt 3 ]; then - echo "usage: review-package BASE HEAD [OUTFILE]" >&2 +if [ $# -lt 3 ] || [ $# -gt 4 ]; then + echo "usage: review-package PLAN_FILE BASE HEAD [OUTFILE]" >&2 exit 2 fi -base=$1 -head=$2 +plan=$1 +base=$2 +head=$3 +[ -f "$plan" ] || { echo "no such plan file: $plan" >&2; exit 2; } git rev-parse --verify --quiet "$base" >/dev/null || { echo "bad BASE: $base" >&2; exit 2; } git rev-parse --verify --quiet "$head" >/dev/null || { echo "bad HEAD: $head" >&2; exit 2; } -if [ $# -eq 3 ]; then - out=$3 +if [ $# -eq 4 ]; then + out=$4 else - dir=$(git rev-parse --git-path sdd) - mkdir -p "$dir" - dir=$(cd "$dir" && pwd) + dir=$("$(cd "$(dirname "$0")" && pwd)/sdd-workspace" "$plan") out="$dir/review-$(git rev-parse --short "$base")..$(git rev-parse --short "$head").diff" fi diff --git a/templates/discipline-skills/subagent-driven-development/scripts/sdd-workspace b/templates/discipline-skills/subagent-driven-development/scripts/sdd-workspace new file mode 100755 index 0000000..4e2d168 --- /dev/null +++ b/templates/discipline-skills/subagent-driven-development/scripts/sdd-workspace @@ -0,0 +1,40 @@ +#!/usr/bin/env bash +# Resolve and ensure the working-tree directory SDD uses for one plan's +# short-lived artifacts: task briefs, implementer reports, review packages, +# and the progress ledger. Print the plan directory's absolute path. +# +# One directory per plan (.superpowers/sdd//) so a follow-up +# plan in the same working tree can never read or overwrite another plan's +# artifacts. A stale ledger misread as current progress makes controllers +# skip whole task sequences — plan-scoping removes that failure structurally. +# +# The workspace lives in the working tree (not under .git/) because Claude Code +# treats .git/ as a protected path and denies agent writes there — which blocks +# an implementer subagent from writing its report file. A self-ignoring +# .gitignore at .superpowers/sdd/ keeps every plan's workspace out of +# `git status` and out of accidental commits without modifying any tracked file. +# +# Single source of truth for the workspace location, so task-brief and +# review-package cannot drift to different directories. +# +# Usage: sdd-workspace PLAN_FILE +set -euo pipefail + +if [ $# -ne 1 ]; then + echo "usage: sdd-workspace PLAN_FILE" >&2 + exit 2 +fi + +plan=$1 +[ -f "$plan" ] || { echo "no such plan file: $plan" >&2; exit 2; } + +slug=$(basename "$plan" .md) +[ -n "$slug" ] && [ "$slug" != "." ] && [ "$slug" != ".." ] \ + || { echo "cannot derive a workspace name from: $plan" >&2; exit 2; } + +root=$(git rev-parse --show-toplevel) +base="$root/.superpowers/sdd" +dir="$base/$slug" +mkdir -p "$dir" +printf '*\n' > "$base/.gitignore" +cd "$dir" && pwd diff --git a/templates/discipline-skills/subagent-driven-development/scripts/task-brief b/templates/discipline-skills/subagent-driven-development/scripts/task-brief index b046a2b..612e14a 100755 --- a/templates/discipline-skills/subagent-driven-development/scripts/task-brief +++ b/templates/discipline-skills/subagent-driven-development/scripts/task-brief @@ -4,8 +4,9 @@ # through the controller's context. # # Usage: task-brief PLAN_FILE TASK_NUMBER [OUTFILE] -# Default OUTFILE: /sdd/task--brief.md — unique per repo -# instance, so concurrent sessions cannot collide. +# Default OUTFILE: /.superpowers/sdd//task--brief.md +# (per plan and per worktree; concurrent runs of the SAME plan in the same +# working tree share it). set -euo pipefail if [ $# -lt 2 ] || [ $# -gt 3 ]; then @@ -20,9 +21,7 @@ n=$2 if [ $# -eq 3 ]; then out=$3 else - dir=$(git rev-parse --git-path sdd) - mkdir -p "$dir" - dir=$(cd "$dir" && pwd) + dir=$("$(cd "$(dirname "$0")" && pwd)/sdd-workspace" "$plan") out="$dir/task-${n}-brief.md" fi diff --git a/templates/discipline-skills/subagent-driven-development/task-reviewer-prompt.md b/templates/discipline-skills/subagent-driven-development/task-reviewer-prompt.md index 588a402..ce79694 100644 --- a/templates/discipline-skills/subagent-driven-development/task-reviewer-prompt.md +++ b/templates/discipline-skills/subagent-driven-development/task-reviewer-prompt.md @@ -52,6 +52,15 @@ Subagent (general-purpose): Your review is read-only on this checkout. Do not mutate the working tree, the index, HEAD, or branch state in any way. + ## You Do Not Dispatch Subagents + + Do all of this review yourself. Never spawn a subagent to review part + of the diff, and never spawn another reviewer for a second opinion. + This process already provides every review seat the work gets; a + reviewer you spawn duplicates one of them at full cost, and its + verdict counts for nothing. If the diff feels too large for one + pass, review it in passes yourself and say so in your report. + ## Do Not Trust the Report Treat the implementer's report as unverified claims about the code. It @@ -75,6 +84,13 @@ Subagent (general-purpose): Warnings or other noise in the implementer's reported test output are findings — test output should be pristine. + Evidence you cannot see is not evidence that doesn't exist. If the + report or its test evidence looks truncated, or you cannot locate the + results it claims, re-read the file at its stated path — and if it is + genuinely missing or garbled, report that as a gap for the controller. + Re-running the suite to regenerate what you failed to read is not + verification; illegibility of the evidence is not invalidation of it. + ## Part 1: Spec Compliance Compare the diff against What Was Requested: @@ -86,6 +102,12 @@ Subagent (general-purpose): - **Misunderstood:** right feature built the wrong way, wrong problem solved + If the brief lists several files each with its own change (a batched + dispatch), check the diff against that list file by file: every listed + file must have its corresponding hunk. A listed file the diff never + touches is a Missing finding, no matter how clean the rest of the + batch looks. + If a requirement cannot be verified from this diff alone (it lives in unchanged code or spans tasks), report it as a ⚠️ item instead of broadening your search. @@ -178,11 +200,8 @@ Subagent (general-purpose): - `[BASE_SHA]` — commit before this task - `[HEAD_SHA]` — current commit - `[DIFF_FILE]` — REQUIRED: the path the controller wrote the review - package to (`scripts/review-package BASE HEAD` prints the unique path it - wrote; the package never enters the controller's context) + package to (`scripts/review-package PLAN_FILE BASE HEAD` prints the unique + path it wrote; the package never enters the controller's context) **Reviewer returns:** Spec Compliance verdict (✅/❌/⚠️), Strengths, Issues (Critical/Important/Minor), Task quality verdict - -A fix dispatch can address spec gaps and quality findings together; -re-review after fixes covers both verdicts. diff --git a/templates/discipline-skills/using-git-worktrees/SKILL.md b/templates/discipline-skills/using-git-worktrees/SKILL.md index 212c569..1381dac 100644 --- a/templates/discipline-skills/using-git-worktrees/SKILL.md +++ b/templates/discipline-skills/using-git-worktrees/SKILL.md @@ -156,47 +156,12 @@ Ready to implement | Tests fail during baseline | Report failures + ask | | No package.json/Cargo.toml | Skip dependency install | -## Common Mistakes - -### Fighting the harness - -- **Problem:** Using `git worktree add` when the platform already provides isolation -- **Fix:** Step 0 detects existing isolation. Step 1a defers to native tools. - -### Skipping detection - -- **Problem:** Creating a nested worktree inside an existing one -- **Fix:** Always run Step 0 before creating anything - -### Skipping ignore verification - -- **Problem:** Worktree contents get tracked, pollute git status -- **Fix:** Always use `git check-ignore` before creating project-local worktree - -### Assuming directory location - -- **Problem:** Creates inconsistency, violates project conventions -- **Fix:** Follow priority: explicit instructions > existing project-local directory > default - -### Proceeding with failing tests - -- **Problem:** Can't distinguish new bugs from pre-existing issues -- **Fix:** Report failures, get explicit permission to proceed - -## Red Flags - -**Never:** -- Create a worktree when Step 0 detects existing isolation -- Use `git worktree add` when you have a native worktree tool (e.g., `EnterWorktree`). This is the #1 mistake — if you have it, use it. -- Skip Step 1a by jumping straight to Step 1b's git commands -- Create worktree without verifying it's ignored (project-local) -- Skip baseline test verification -- Proceed with failing tests without asking - -**Always:** -- Run Step 0 detection first -- Prefer native tools over git fallback -- Follow directory priority: explicit instructions > existing project-local directory > default -- Verify directory is ignored for project-local -- Auto-detect and run project setup -- Verify clean test baseline +## Common Rationalizations + +| Excuse | Reality | +|--------|---------| +| "I'm obviously not in a worktree — no need to check" | Run Step 0. Harness-created isolation and submodules both fool eyeballing; the detection commands settle it. | +| "`git worktree add` is quicker than hunting for a native tool" | A native tool (e.g. `EnterWorktree`) owns placement, branching, and cleanup. Bypassing it is the #1 mistake — it creates phantom state your harness can't see or manage. | +| "The worktree directory is surely ignored already" | Run `git check-ignore`. An unignored worktree directory commits the whole tree into the repo. | +| "Any directory name works" | Explicit instructions beat an existing project-local directory, which beats the `.worktrees/` default. | +| "The workspace is fresh — baseline tests can wait" | A dirty baseline makes every later failure ambiguous. Run the tests now; proceeding past failures is your human partner's call. | diff --git a/templates/discipline-skills/verification-before-completion/SKILL.md b/templates/discipline-skills/verification-before-completion/SKILL.md index 2f14076..7d45333 100644 --- a/templates/discipline-skills/verification-before-completion/SKILL.md +++ b/templates/discipline-skills/verification-before-completion/SKILL.md @@ -7,8 +7,6 @@ description: Use when about to claim work is complete, fixed, or passing, before ## Overview -Claiming work is complete without verification is dishonesty, not efficiency. - **Core principle:** Evidence before claims, always. **Violating the letter of this rule is violating the spirit of this rule.** @@ -105,15 +103,6 @@ Skip any step = lying, not verifying ❌ Trust agent report ``` -## Why This Matters - -From 24 failure memories: -- your human partner said "I don't believe you" - trust broken -- Undefined functions shipped - would crash -- Missing requirements shipped - incomplete features -- Time wasted on false completion → redirect → rework -- Violates: "Honesty is a core value. If you lie, you'll be replaced." - ## When To Apply **ALWAYS before:** @@ -129,11 +118,3 @@ From 24 failure memories: - Paraphrases and synonyms - Implications of success - ANY communication suggesting completion/correctness - -## The Bottom Line - -**No shortcuts for verification.** - -Run the command. Read the output. THEN claim the result. - -This is non-negotiable. diff --git a/templates/discipline-skills/writing-plans/SKILL.md b/templates/discipline-skills/writing-plans/SKILL.md index f6debd2..8bbe6b2 100644 --- a/templates/discipline-skills/writing-plans/SKILL.md +++ b/templates/discipline-skills/writing-plans/SKILL.md @@ -66,6 +66,9 @@ independently testable deliverable. **Tech Stack:** [Key technologies/libraries] +**Spec:** [path to the spec/design doc this plan implements — the plan +argues from the spec, so the spec travels with it; executors read both] + ## Global Constraints [The spec's project-wide requirements — version floors, dependency limits, @@ -135,12 +138,6 @@ Every step must contain the actual content an engineer needs. These are **plan f - Steps that describe what to do without showing how (code blocks required for code steps) - References to types, functions, or methods not defined in any task -## Remember -- Exact file paths always -- Complete code in every step — if a step changes code, show the code -- Exact commands with expected output -- DRY, YAGNI, TDD, frequent commits - ## Self-Review After writing the complete plan, look at the spec with fresh eyes and check the plan against it. This is a checklist you run yourself — not a subagent dispatch. diff --git a/test/discipline-skills/test-module-files-exist.sh b/test/discipline-skills/test-module-files-exist.sh index 24983c1..1fd68fd 100755 --- a/test/discipline-skills/test-module-files-exist.sh +++ b/test/discipline-skills/test-module-files-exist.sh @@ -15,8 +15,10 @@ paths=( "templates/discipline-skills/subagent-driven-development/SKILL.md" "templates/discipline-skills/subagent-driven-development/implementer-prompt.md" "templates/discipline-skills/subagent-driven-development/task-reviewer-prompt.md" + "templates/discipline-skills/subagent-driven-development/re-review-prompt.md" "templates/discipline-skills/subagent-driven-development/scripts/review-package" "templates/discipline-skills/subagent-driven-development/scripts/task-brief" + "templates/discipline-skills/subagent-driven-development/scripts/sdd-workspace" "templates/discipline-skills/finishing-a-development-branch/SKILL.md" "templates/discipline-skills/hooks/sessionstart-discipline.sh" ) diff --git a/test/discipline-skills/test-scaffold-installs-skills.sh b/test/discipline-skills/test-scaffold-installs-skills.sh index 3a436d3..1218158 100755 --- a/test/discipline-skills/test-scaffold-installs-skills.sh +++ b/test/discipline-skills/test-scaffold-installs-skills.sh @@ -38,14 +38,15 @@ jq -e '.hooks.SessionStart[] | .hooks[] | select(.command | contains("sessionsta || { echo "FAIL: missing brainstorming/spec-document-reviewer-prompt.md"; exit 1; } [ -f "$tmp/.claude/skills/writing-plans/plan-document-reviewer-prompt.md" ] \ || { echo "FAIL: missing writing-plans/plan-document-reviewer-prompt.md"; exit 1; } -for sub in implementer-prompt.md task-reviewer-prompt.md; do +for sub in implementer-prompt.md task-reviewer-prompt.md re-review-prompt.md; do [ -f "$tmp/.claude/skills/subagent-driven-development/$sub" ] \ || { echo "FAIL: missing subagent-driven-development/$sub"; exit 1; } done -# v6.0.0 added bash scripts (no .sh extension) used by SDD's file-handoff flow. -# Both must ship executable so the implementer/reviewer can invoke them directly. -for script in review-package task-brief; do +# v6.0.0 added bash scripts (no .sh extension) used by SDD's file-handoff flow; +# v6.0.3 added sdd-workspace (the shared workspace resolver both call). +# All must ship executable so the controller can invoke them directly. +for script in review-package task-brief sdd-workspace; do [ -f "$tmp/.claude/skills/subagent-driven-development/scripts/$script" ] \ || { echo "FAIL: missing subagent-driven-development/scripts/$script"; exit 1; } [ -x "$tmp/.claude/skills/subagent-driven-development/scripts/$script" ] \ From 63b98b57b812aec78afa596f76cc21255e3a6158 Mon Sep 17 00:00:00 2001 From: tigers1997 Date: Mon, 24 Aug 2026 16:13:04 -0400 Subject: [PATCH 2/2] chore(compat): CC 2.1.183-2.1.241 survey; retire dead autoMode + Write() defaults; shell:bash on every hook tested_up_to 2.1.150 -> 2.1.220. The no-lone-bumps gate finally cleared: SchemaStore #5723 (-> 2.1.150) was closed unmerged, but #5867 (-> 2.1.195, merged 2026-07-03) and #6131 (-> 2.1.220, merged 2026-07-27) landed. The keys held since v2.6.0 (fallbackModel, disableBundledSkills) are schema-validated, as are agent, claudeMdExcludes, skillListingBudgetFraction/MaxDescChars, sandbox.credentials (deny form) and worktree.symlinkDirectories/sparsePaths -- all promoted to doc-only opt-in stubs in settings.local.json.example. Every release 2.1.183 -> 2.1.241 was read verbatim; the full survey is recorded in the CLAUDE_CODE_COMPAT comment block. Four shipped defaults were dead or wrong, not merely stale: 1. autoMode.hard_deny in project settings has been ignored since CC 2.1.207 -- the classifier reads autoMode only from ~/.claude/settings.json or managed settings, because .claude/settings.json and .claude/settings.local.json both live in the repo and could inject allow rules. The block v2.6.0 promoted to an active default did nothing. Removed from the safety patch; documented as a user-scope block (with "$defaults") in docs/05-safety-permissions.md. A retrofit strips the exact shipped block -- a user-edited variant survives -- and check_settings_validates now flags any project-scope autoMode. 2. Write(.env) / Write(.env.*) deny rules were never consulted. Claude Code checks file permissions against Edit(path) and Read(path) only, and 2.1.210 added a startup warning for exactly these spellings. Dropped from the core deny list (the Edit(.env*) rules already covered the Write tool); a retrofit strips those two shipped strings; the preflight flags any others via the new _find_unconsulted_path_rules. 3. Every shipped command hook now declares "shell": "bash" (13 entries: 9 in the settings patches, 4 embedded in config_schema). Without it a Windows session with no Git Bash sends .sh hooks to PowerShell, where they die on a parser error; with it Claude Code resolves Git for Windows directly and prompts to install it when missing. Older CC ignores the key. Adopted after upstream superpowers v6.2.0 fixed its own SessionStart hook the same way. _merge_hook_groups backfills the key onto configurator-owned entries on retrofit, never onto a user's own command or an explicit shell choice. 4. CC 2.1.223 made /review the bundled alias of /code-review -- the collision the 2.1.146 survey note was watching for. Rather than assume, this was tested headlessly on 2.1.241: a project skill named `review` wins over the bundled alias, and `plan` wins over the built-in /plan shortcut. No rename; the shadowing is documented in README, docs/03 and the skill's own description. Wording brought in line with the runtime: subagent nesting (5 -> off -> 3 by default at 2.1.219, CLAUDE_CODE_MAX_SUBAGENT_SPAWN_DEPTH), the 20-concurrent cap (2.1.217), background-by-default subagents (2.1.198), the fork subagent type (2.1.232), Explore inheriting the session model instead of Haiku (2.1.198), the new SessionStart `fork` source (2.1.214), hook-matcher semantics (hyphen exact- match 2.1.195, comma separators 2.1.191, cwd-anchored single-segment if: 2.1.214), DirectoryAdded (2.1.219), Todo-tool withdrawal on Fable 5 / Opus 4.8+ / Sonnet 5 (2.1.233), and the Sonnet 5 / Opus 5 model landscape. The shipped matchers are pipe-joined tool names, so none of the matcher changes affect behavior. --check now rejects a UTF-8 BOM at the top of any shipped SKILL.md or agent file (CC before 2.1.239 silently ignored BOM-prefixed files). A retrofit also re-runs check_settings_validates against the merged result and surfaces findings under [ MERGED ] -- the one place a user's own settings.json enters the pipeline. Held / out of territory: sandbox.network.strictAllowlist, sandbox.filesystem. disabled and dialogExpiry are user/managed-only per the settings reference; crossSessionInbound is project-honored but not yet in SchemaStore (the sync stops at 2.1.220) and is held for the next survey; UI prefs, marketplace/plugin-source/ self-hosted-runner/gateway keys and the removed /agents wizard need no action. New test/retrofit-hooks/test-retired-defaults-migration.sh and two new cases in test-preflight-detects-violations.sh. The python-uv-fastapi example is regenerated, which also catches it up to the docker-compose check field from c29786d (it was last written by cc-configure 2.6.0). Claude Code compat: 2.1.116-2.1.220. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01JVndNviHZSnbKJnWP7jFbV --- CHANGELOG.md | 2 + README.md | 6 +- config_schema.py | 121 +++++++++++++++++- configure.py | 114 ++++++++++++++++- docs/03-commands-and-hooks.md | 18 ++- docs/04-subagents-mcp-orchestration.md | 8 +- docs/05-safety-permissions.md | 18 ++- docs/06-token-efficiency.md | 7 +- .../.claude/.cc-manifest.json | 6 +- .../.claude/hooks/microbit-enforcer.sh | 2 +- .../.claude/hooks/stop-run-checks.sh | 71 ++++++++-- .../python-uv-fastapi/.claude/settings.json | 19 +-- .../.claude/settings.local.json.example | 43 ++++++- .../.claude/skills/review/SKILL.md | 2 +- templates/commands/infinite/SKILL.md | 6 +- .../microbit-enforcer/microbit-enforcer.sh | 2 +- .../microbit-enforcer/settings-patch.json | 4 +- templates/commands/review/SKILL.md | 2 +- templates/core/dot-claude/settings.json | 2 - .../dot-claude/settings.local.json.example | 43 ++++++- templates/git-workflow/settings-patch.json | 2 + .../rules/multi-agent-guardrails.md | 7 + templates/safety/settings-patch.json | 13 +- .../safety/settings-patch.slop-scan.json | 1 + .../settings-patch.tier-pro.json | 1 + .../test-retired-defaults-migration.sh | 72 +++++++++++ .../test-preflight-detects-violations.sh | 21 ++- 27 files changed, 532 insertions(+), 81 deletions(-) create mode 100644 test/retrofit-hooks/test-retired-defaults-migration.sh diff --git a/CHANGELOG.md b/CHANGELOG.md index a757fbe..0bbca69 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -4,6 +4,8 @@ All notable changes to this project. Format: [Keep a Changelog](https://keepacha ## Unreleased +- **chore(compat): CC 2.1.183–2.1.241 survey — `tested_up_to` 2.1.150 → 2.1.220; retire the dead project-scope `autoMode` default and the `Write(.env*)` deny rules; `shell: "bash"` on every shipped hook; `/review` shadowing documented.** The no-lone-bumps gate finally cleared: SchemaStore's #5723 (→ 2.1.150) was closed unmerged, but #5867 (→ 2.1.195, merged 2026-07-03) and #6131 (→ 2.1.220, merged 2026-07-27) landed, so `tested_up_to` moves to **2.1.220** and the keys held since v2.6.0 (`fallbackModel`, `disableBundledSkills`) plus `agent`, `claudeMdExcludes`, `skillListingBudgetFraction` / `skillListingMaxDescChars`, `sandbox.credentials` (deny form) and `worktree.symlinkDirectories` / `sparsePaths` are schema-validated and promoted to doc-only opt-in stubs in `settings.local.json.example` (validation stamp 2026-07-27; the file also gains an `ANTHROPIC_DEFAULT_MODEL` note and an explicit "autoMode is NOT read from this file" warning). Every release 2.1.183 → 2.1.241 was read verbatim (2.1.182/184/188/189/192/194/213/230 are absent upstream); the installed CC is 2.1.241. **Template-affecting findings, all acted on:** **(1)** Since CC 2.1.207 the auto-mode classifier reads `autoMode` only from `~/.claude/settings.json` or managed settings — the live `auto-mode-config` doc states neither `.claude/settings.json` nor `.claude/settings.local.json` is read (both live in the repo) — so the `autoMode.hard_deny` block the safety module has shipped as an active default since v2.6.0 was dead config. It is removed from `templates/safety/settings-patch.json`; the old two rules are documented as a user-scope copy-paste block in `docs/05-safety-permissions.md` (with `"$defaults"`), a retrofit strips the *exact* shipped block from project settings (a user-edited variant is preserved), and `check_settings_validates` now flags any project-scope `autoMode` under `[ SETTINGS WARNINGS ]`. **(2)** CC 2.1.210 warns at startup about `Write(path)` / `NotebookEdit(path)` / `Glob(path)` rules, which it never consulted (`Edit(path)` covers Write and NotebookEdit; `Read(path)` covers Glob/Grep) — `Write(.env)` / `Write(.env.*)` are dropped from the core deny list (`Edit(.env)` / `Edit(.env.*)` stay), a retrofit strips those two shipped strings, and the preflight flags any remaining never-consulted path rule. New `RETIRED_PROJECT_AUTOMODE` / `RETIRED_DENY_RULES` constants in `configure.py`; the `[ MERGED ]` summary reports "removed N retired configurator default(s)". **(3)** Every shipped command hook (9 entries across the safety / slop-scan / git-workflow / token-efficiency-pro / microbit patches + 4 schema-embedded entries: PreCompact snapshot, the two mcp drift-check groups, the discipline bootstrap) now declares `"shell": "bash"` — a schema-validated key (CC 2.1.81+) that makes Windows sessions resolve Git for Windows directly instead of falling back to PowerShell and dying on the `.sh` entrypoint, and prompts to install Git Bash when it's missing; older CC ignores the key. Adopted after upstream superpowers v6.2.0 fixed its own SessionStart hook the same way. `_merge_hook_groups` backfills the key onto configurator-owned entries on retrofit (never onto a user's own commands or an explicit `shell` choice); README's Windows row rewritten. **(4)** CC 2.1.223 made `/review` the bundled alias of `/code-review` — the collision the 2.1.146 survey note watched for. The skills doc says a same-named project skill overrides a bundled one, and a headless check on 2.1.241 (a project skill named `review`, then `claude -p /review`) ran the project skill — likewise a project `plan` skill over the built-in `/plan` plan-mode shortcut — so no rename: the shadowing is documented in README's `commands` row, `docs/03`'s starter kit, and `review/SKILL.md`'s description ("Distinct from Claude Code's built-in /code-review"). **(5)** Subagent runtime changed under the docs: nesting default 5 → off (2.1.217) → 3 (2.1.219, `CLAUDE_CODE_MAX_SUBAGENT_SPAWN_DEPTH`), 20-concurrent cap (`CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS`, 2.1.217), background-by-default subagents with a narrower tool set (2.1.198), `fork` subagent type on by default (2.1.232), `Explore` inheriting the session model capped at Opus instead of Haiku (2.1.198), the Task tool's `mode` parameter deprecated (2.1.212) — `docs/04`, `docs/06`, `templates/commands/infinite/SKILL.md` and `multi-agent-guardrails.md` (new "Runtime caps to design around" section) reworded; the leaf-design guidance stands. **(6)** SessionStart gained the `fork` source (2.1.214; forks used to report `resume`): the microbit marker-clear stays `startup|clear` (markers now also persist across `/fork`, which is right) and the mcp drift-check stays `startup|resume`; comments and `docs/03` list the new source. Hook-matcher semantics (hyphenated names exact-match since 2.1.195, comma separators since 2.1.191, cwd-anchored single-segment `if:` patterns since 2.1.214) leave the pipe-joined shipped matchers unaffected and are documented in `docs/03`, whose event table also gains `DirectoryAdded` (2.1.219) and the settings example the `shell` key. **(7)** TodoWrite/Task* tools are withdrawn on Fable 5 / Opus 4.8+ / Sonnet 5 (2.1.233; `CLAUDE_CODE_ENABLE_TODO_TOOLS=1` restores) — no shipped template named them; noted in `docs/06`, which also refreshes the model lever for Sonnet 5 (2.1.197, Claude Code's default) and Opus 5 (2.1.219) and the `xhigh` effort default. **(8)** `--check` now rejects a UTF-8 BOM at the top of any shipped SKILL.md / agent file (CC before 2.1.239 silently ignored BOM-prefixed files). **Held / out of territory:** `sandbox.network.strictAllowlist` (2.1.219), `sandbox.filesystem.disabled` (2.1.216) and `dialogExpiry` (2.1.224) are user/managed-only per the settings reference; `crossSessionInbound` (2.1.224) is project-honored but not yet in SchemaStore (sync stops at 2.1.220) — held for the next sync; UI prefs (`keybindingFlavor`, `spellcheck`, `emojiCompletionEnabled`, `axScreenReader`, `vimInsertModeRemaps`, `respondToBashCommands`, `awaySummaryEnabled`), marketplace/plugin-source/self-hosted-runner/gateway keys and the removed `/agents` wizard need no configurator action. Full survey recorded in `CLAUDE_CODE_COMPAT`'s comment block. New tests: `test/retrofit-hooks/test-retired-defaults-migration.sh` (shipped autoMode + Write rules migrate out, `shell` backfilled on every hook, a user-edited autoMode block survives) and two new cases in `test/schema-hygiene/test-preflight-detects-violations.sh`; the `python-uv-fastapi` example regenerated. Claude Code compat: **2.1.116–2.1.220**. + - **chore(skills): sync discipline-skills v6.0.2 → v6.3.0 (obra/superpowers).** Four upstream releases since #85, all landing in the seven forked skills. **v6.0.3** moved SDD's scratch files out of `.git/` (Claude Code denies agent writes there) into a self-ignoring `.superpowers/sdd/` working-tree directory resolved by a new shared script, `scripts/sdd-workspace`. **v6.2.0** made that workspace plan-scoped (`.superpowers/sdd//`; `scripts/review-package` now takes the plan file first: `review-package PLAN_FILE BASE HEAD`), restructured the review-fix loop to resume the implementer with a scoped re-review (`re-review-prompt.md`, NEW) and a five-round circuit breaker, ran a library-wide compression campaign (the Bottom Line / Key Principles / Advantages / Integration / "Why This Matters" sections are gone; `using-git-worktrees` and `finishing-a-development-branch` gained Excuse/Reality rationalization tables), dropped "Discard this work" from the finishing menu (discard is explicit-request-only), made PR creation forge-agnostic and fixed the worktree path being recomputed after the cleanup `cd`. **v6.3.0** teaches `brainstorming` to classify requests as spike / bounded / architectural and scale the ceremony (only the architectural path writes a spec; the approval gate never scales), has SDD controllers issue recorded rulings instead of stalling on plan conflicts (ledgered pre-flight scan table, batched same-shape tasks, a hard "implementers and reviewers never spawn subagents" contract in both prompt templates, a `**Spec:**` pointer in the `writing-plans` header, reviewers re-reading illegible evidence instead of re-running suites), and stops `finishing-a-development-branch` from `--force`-removing a worktree that holds untracked files. Local edits re-applied per SYNC.md: `superpowers:` prefixes stripped (14 sites across three SKILL.md files), the brainstorming `## Visual Companion` section and the visual-companion step of the *architectural* checklist removed (8 items; the spike/bounded lists have none), the executing-plans subagents note reframed project-neutral. **One new local edit:** SDD's final whole-branch review now points at the configurator's own `code-reviewer` subagent (`.claude/agents/code-reviewer.md`, with a `task-reviewer-prompt.md` fallback when the `commands` module isn't installed) instead of upstream's `../requesting-code-review/code-reviewer.md` — three digraph labels plus the `## Final Review` paragraph — which retires the broken-link papercut carried since v5.1.0. The former "remove the requesting-code-review / test-driven-development lines from `## Integration`" edits are obsolete (upstream dropped the section in v6.2.0). Module `paths:` 15 → 17 (`re-review-prompt.md`, `scripts/sdd-workspace`); `test-module-files-exist.sh` and `test-scaffold-installs-skills.sh` updated (the scaffold test now asserts all three scripts ship executable); the three persona snapshots that include the module regenerated; SYNC.md pinned to v6.3.0 (2026-08-12) with a fresh delta paragraph, a CRLF note for Windows checkouts, and a renumbered canonical-edit list; the module description, README carve-out paragraphs, `NOTICE` and `docs/10` (the last two still said v5.1.0) bumped. On a `core.filemode=false` checkout the new script's index mode was set with `git add --chmod=+x`. Upstream's v6.2.0 Windows fix (`"shell": "bash"` on its SessionStart hook) is adopted configurator-wide in the companion compat entry. - **feat(hooks): `stop-run-checks` can run a check inside its docker-compose service (dogfood F3).** A containerized project's check loop no-op'd because the toolchain lives in the image, not on the host PATH (F2 warned about this). A `CHECKS` entry now takes an optional 3rd field — `label|command|service` — and when a service is named the Stop hook runs the check inside it: `docker compose exec -T ` if the service is up, else `docker compose run --rm ` (a throwaway instance of the *existing* service — no new container, no host rebuild, no path translation). The host-PATH/manifest guards are bypassed for container checks; infra-not-ready (no docker/compose, or service undefined) skips silently (fail-open), and a non-zero container exit is reported as a normal check FAIL. Two-field `label|command` entries — including host commands containing a literal `|` — are unchanged; an entry is container-bound only when it has ≥2 pipes and a bare-token final field. The F2 `[ STACK WARNINGS ]` note now points at this field. New `test/stop-run-checks/test-container-checks.sh` (run-vs-exec by service state, skip when compose/service absent, FAIL on non-zero, 2-field stays host-side) via a `docker` stub. **Deferred follow-up:** a conditional intake field that detects `docker-compose.yml` and pre-populates the service per check (this format is the substrate). No `tested_up_to` / `CC_VERSION` bump. diff --git a/README.md b/README.md index 20dfcaa..6e9221b 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ Headless CLI that generates Claude Code project scaffolding — `CLAUDE.md`, `.c - **git** — used by the installer and (later) by generated hooks / statusline - **bash** — the shipped hook scripts and `./claude-ctx` wrapper all use `#!/usr/bin/env bash` - **curl** — only for the one-shot install command -- **Claude Code 2.1.116–2.1.132** — the range the current templates are tested against (see `CLAUDE_CODE_COMPAT` in `config_schema.py`). The configurator runs a preflight check and prints a `[ VERSION WARNINGS ]` block if the installed `claude` is outside this range. Older CC silently drops features (agent-frontmatter `mcpServers: http`; `DISABLE_UPDATES`; `permissions.disableBypassPermissionsMode`). Each release states its compat range in `CHANGELOG.md`. +- **Claude Code 2.1.116–2.1.220** — the range the current templates are tested against (see `CLAUDE_CODE_COMPAT` in `config_schema.py`; the upper bound tracks the last SchemaStore settings-schema sync, and the shipped templates were surveyed against every upstream release through 2.1.241). The configurator runs a preflight check and prints a `[ VERSION WARNINGS ]` block if the installed `claude` is outside this range. Older CC silently drops features (agent-frontmatter `mcpServers: http`; `DISABLE_UPDATES`; `permissions.disableBypassPermissionsMode`). Each release states its compat range in `CHANGELOG.md`. **Generated projects need, depending on which modules you enable:** - **`ui`** — `python3` + `git` (used by `statusline.sh` and the last-prompt variant) @@ -26,7 +26,7 @@ Headless CLI that generates Claude Code project scaffolding — `CLAUDE.md`, `.c | --- | --- | | **Linux** (Debian/Ubuntu/Arch/Fedora/…) | Primary target. Everything works out of the box. | | **macOS** (12+) | Works. Bash 3.2 from the system is sufficient; scripts use POSIX-safe flags. | -| **Windows** | Claude Code 2.1.120+ runs natively on Windows — when Git Bash is absent, Claude Code falls back to PowerShell as its shell tool. However, the `.sh` hook scripts this project ships still need a bash interpreter to execute. Three options: install Git Bash (via Git for Windows), use WSL, or translate each hook to PowerShell and set `"shell": "powershell"` on the hook entry (per the Claude Code docs). The template directory uses `dot-claude/` (rewritten to `.claude/` at install) so the templates browse and sync cleanly on filesystems and tools that special-case dotfiles. | +| **Windows** | Claude Code 2.1.120+ runs natively on Windows — when Git Bash is absent, Claude Code falls back to PowerShell as its shell tool. The `.sh` hook scripts this project ships still need a bash interpreter, so every shipped hook entry declares `"shell": "bash"` (honored by Claude Code 2.1.81+): with Git for Windows installed the hooks run under Git Bash even when Claude Code's own shell tool fell back to PowerShell, and without it Claude Code prompts to install Git Bash instead of failing silently. Alternatives: use WSL, or translate a hook to PowerShell and set `"shell": "powershell"` on its entry. The template directory uses `dot-claude/` (rewritten to `.claude/` at install) so the templates browse and sync cleanly on filesystems and tools that special-case dotfiles. | ## Install @@ -130,7 +130,7 @@ There are 11 modules; legacy IDs (`commands-core`, `agents`, `token-efficiency-p | **safety** | PreToolUse hooks (block dangerous bash + scan Write/Edit for secrets + gate `apt\|brew\|dnf\|yum\|pacman\|apk install` for packages not in any configured repo, with structured denial listing available siblings + detected installed version); `permissions.disableBypassPermissionsMode: "disable"` hard-blocks `--dangerously-skip-permissions`. Sub-flags: `lockdown` (sets `DISABLE_UPDATES=1` — blocks autoupdates AND manual `claude update`; for air-gapped / enterprise environments); `slop_scan` (PostToolUse hook on Write/Edit/NotebookEdit flagging filler / marketing-voice / hedging / em-dash patterns; `slop_scan_action=warn\|block`, `slop_scan_density` and `slop_scan_imports` opt-in). All non-`custom` personas pre-set `slop_scan=true` action=warn. | | **git-workflow** | PostToolUse formatter on Write/Edit, Stop hook running typecheck / lint / tests. | | **token-efficiency** | Path-scoped `.claude/rules/` starters + PreCompact snapshot hook. `tier` flag: `basic` (default) ships discipline rules + snapshot only; `pro` adds bash-output truncation hook + always-loaded discipline rules. | -| **commands** | Slash commands + agents + microbits. `subset` flag (linear ordering: `curated ⊂ full ⊂ rigorous`): **`curated`** = 3 essential skills (`/plan`, `/commit`, `/verify-setup`) + the `code-reviewer` agent. **`full`** (default) = 9 workflow skills (adds `/review`, `/ship`, `/sync-docs`, `/check-context`, `/session-retro`, `/retrofit`) + 4 agents (`code-reviewer`, `test-runner`, `doc-writer`, `security-auditor`) + 4 discipline microbits (`/freeze`, `/unfreeze`, `/guard`, `/careful`) + the `microbit-enforcer.sh` PreToolUse hook. **`rigorous`** = `full` + `/investigate` + `/plan-eng-review`, the rigor skills that embed `templates/commands/_patterns/` cross-cutting blocks (confidence gate, independent verification, no-fix-without-investigation, AI-slop detection). The `security-auditor` frontmatter wires Sonatype's dependency-management MCP (`https://mcp.guide.sonatype.com/mcp`) scoped to that agent — active only when it runs, so ~0 baseline context cost. Set `SONATYPE_TOKEN` env var to enable ([generate a token](https://guide.sonatype.com/settings/tokens)). | +| **commands** | Slash commands + agents + microbits. `subset` flag (linear ordering: `curated ⊂ full ⊂ rigorous`): **`curated`** = 3 essential skills (`/plan`, `/commit`, `/verify-setup`) + the `code-reviewer` agent. **`full`** (default) = 9 workflow skills (adds `/review`, `/ship`, `/sync-docs`, `/check-context`, `/session-retro`, `/retrofit`) + 4 agents (`code-reviewer`, `test-runner`, `doc-writer`, `security-auditor`) + 4 discipline microbits (`/freeze`, `/unfreeze`, `/guard`, `/careful`) + the `microbit-enforcer.sh` PreToolUse hook. **`rigorous`** = `full` + `/investigate` + `/plan-eng-review`, the rigor skills that embed `templates/commands/_patterns/` cross-cutting blocks (confidence gate, independent verification, no-fix-without-investigation, AI-slop detection). `/review` and `/plan` share their names with Claude Code built-ins (the `/review` alias of the bundled `/code-review` since CC 2.1.223, and the `/plan` plan-mode shortcut); a project skill wins by name (verified on CC 2.1.241), so reach the built-in multi-agent review with `/code-review` and plan mode with Shift+Tab. The `security-auditor` frontmatter wires Sonatype's dependency-management MCP (`https://mcp.guide.sonatype.com/mcp`) scoped to that agent — active only when it runs, so ~0 baseline context cost. Set `SONATYPE_TOKEN` env var to enable ([generate a token](https://guide.sonatype.com/settings/tokens)). | | **mcp** | `.mcp.json` generated from selected servers, plus **per-task profiles** (`.mcp.research.json`, `.mcp.frontend.json`, `.mcp.minimal.json`) and an executable `./claude-ctx` wrapper that launches Claude with `--mcp-config --strict-mcp-config` — drops a bloated 4-MCP baseline from ~49% context to under 5%. | | **multi-agent** | Path-scoped `multi-agent-guardrails.md` (5-scenario "when not to parallel" list), `/merge-worktrees` skill, `/infinite` skill, `parallel-generator` subagent. | | **github-actions** | `.github/workflows/claude.yml` pinned to `anthropics/claude-code-action@v1`. Triggers on `@claude` mentions in issues, PR comments, and PR reviews. | diff --git a/config_schema.py b/config_schema.py index f54d683..acea365 100644 --- a/config_schema.py +++ b/config_schema.py @@ -89,6 +89,7 @@ "hooks": [ { "type": "command", + "shell": "bash", "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/pre-compact-snapshot.sh", "timeout": 15, } @@ -215,6 +216,7 @@ "hooks": [ { "type": "command", + "shell": "bash", "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/sessionstart-drift-check.sh", "timeout": 5, } @@ -225,6 +227,7 @@ "hooks": [ { "type": "command", + "shell": "bash", "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/sessionstart-drift-check.sh", "timeout": 5, } @@ -263,6 +266,7 @@ "hooks": [ { "type": "command", + "shell": "bash", "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/sessionstart-discipline.sh", "timeout": 5, } @@ -404,7 +408,11 @@ # configurator against. Newer is likely fine but unverified. CLAUDE_CODE_COMPAT = { "min_version": "2.1.116", # agent mcpServers http (2.1.116/117) - "tested_up_to": "2.1.150", # SchemaStore PR #5706 merged 2026-05-23 (sync to + "tested_up_to": "2.1.220", # SchemaStore sync PR #6131 merged 2026-07-27 + # (v2.1.220). Surveyed through 2.1.241 on + # 2026-08-24 - see the LAST paragraph of this + # block. History follows, oldest first: + # 2.1.150 bump - SchemaStore PR #5706 merged 2026-05-23 (sync to # v2.1.143), unblocking nine settings.json keys for # opt-in templating: skillOverrides (token- # efficiency-pro tier-pro patch), worktree.baseRef @@ -570,6 +578,117 @@ # survey - still open/draft, untouched since # 2026-05-24 despite dep #5728 merged # 2026-06-01. Issue #14920 unchanged. + # 2.1.183-2.1.241 survey (2026-08-24; installed CC + # 2.1.241): tested_up_to 2.1.150 -> 2.1.220. The + # no-lone-bumps gate cleared: #5723 (->2.1.150) was + # closed unmerged, but #5867 (->2.1.195, merged + # 2026-07-03) and #6131 (->2.1.220, merged + # 2026-07-27; #6127 then dropped a stray + # leftArrowOpensAgents key, 2026-08-03) landed, so the + # held keys fallbackModel + disableBundledSkills are + # schema-validated - as are agent, claudeMdExcludes, + # skillListingBudgetFraction/MaxDescChars, + # sandbox.credentials (deny form; mask is user/ + # managed-only) and worktree.symlinkDirectories - + # all promoted to doc-only opt-in stubs in + # settings.local.json.example (2026-07-27 stamp). + # Versions absent from the upstream CHANGELOG + # (skipped/internal): 2.1.182, 184, 188, 189, 192, + # 194, 213, 230. + # TEMPLATE-AFFECTING (acted on): + # (1) 2.1.207 - the classifier no longer reads + # autoMode from repo-resident settings; the live + # auto-mode-config doc says neither + # .claude/settings.json nor settings.local.json is + # read -> the safety patch's active + # autoMode.hard_deny default (v2.6.0 promotion) was + # dead config. Removed; documented as a + # ~/.claude/settings.json opt-in (docs/05); a + # retrofit strips the exact shipped block and + # [ SETTINGS WARNINGS ] flags any other project- + # scope autoMode. + # (2) 2.1.210 - startup warning for Write(path)/ + # NotebookEdit(path)/Glob(path) rules (never + # consulted; Edit()/Read() cover them) -> dropped + # Write(.env) + Write(.env.*) from the core deny + # list (Edit(.env*) stays); retrofit strips those + # two shipped strings; preflight flags others. + # (3) 2.1.223 - /review is now the bundled alias + # of /code-review (the collision the 2.1.146 note + # above watched for). Project skills win by name: + # the skills doc says a same-named skill overrides + # a bundled skill, and a headless check on 2.1.241 + # (project skill `review` -> `claude -p /review`) + # ran the project skill; same result for `plan` + # vs the built-in /plan shortcut. No rename; + # documented as shadowing (README, docs/03, + # review/SKILL.md description). + # (4) 2.1.217/2.1.219 - nesting default 1 then 3 + # (was 5 since 2.1.172), env + # CLAUDE_CODE_MAX_SUBAGENT_SPAWN_DEPTH; 20- + # concurrent cap (CLAUDE_CODE_MAX_CONCURRENT_ + # SUBAGENTS, 2.1.217); a 200/session spawn cap + # came (2.1.212) and went (2.1.224) -> docs/04, + # infinite/SKILL.md and multi-agent-guardrails.md + # reworded. + # (5) 2.1.198/2.1.232 - subagents background by + # default, `fork` subagent type on by default, + # Explore inherits the session model (capped at + # Opus, was Haiku); Task tool `mode` param + # deprecated (2.1.212) -> docs/04 + docs/06. + # (6) 2.1.214 - SessionStart source `fork` + # (forks used to report `resume`) -> microbit + # marker-clear stays `startup|clear` (markers + # persist across forks - correct); the mcp + # drift-check stays `startup|resume` (a fork is + # a live copy, closer to compact than to a cold + # resume). Comments/docs list the new source. + # (7) hook-matcher semantics: hyphenated names + # exact-match (2.1.195), comma separators work + # (2.1.191), single-segment `if:` dir patterns + # cwd-anchored (2.1.214) - shipped matchers are + # pipe-joined tool names, unaffected; docs/03. + # (8) 2.1.233 - TodoWrite/Task* tools withdrawn + # on Fable 5 / Opus 4.8+ / Sonnet 5 + # (CLAUDE_CODE_ENABLE_TODO_TOOLS=1 restores); + # shipped templates never named them; docs/06. + # (9) hookCommand `shell` (schema enum + # bash|powershell, CC 2.1.81+): adopted on every + # shipped command hook after upstream superpowers + # v6.2.0 fixed its Windows SessionStart hook the + # same way; retrofit backfills it on configurator- + # owned entries; README Windows row rewritten. + # (10) 2.1.239 - BOM-prefixed skill/agent .md + # files were silently ignored before -> --check + # now rejects a BOM in any shipped frontmatter + # file. + # Model landscape: Sonnet 5 (2.1.197, CC's default + # model) and Opus 5 (2.1.219, default Opus, 1M); + # scaffold default stays `fable`. + # OUT OF TERRITORY / HELD: sandbox.network. + # strictAllowlist (2.1.219), sandbox.filesystem. + # disabled (2.1.216), dialogExpiry (2.1.224) - + # user/managed-only per settings-reference; + # crossSessionInbound (2.1.224; any file, stricter + # project value honored) - not in SchemaStore yet + # (sync stops at 2.1.220), HELD; keybindingFlavor, + # spellcheck, emojiCompletionEnabled, + # axScreenReader, vimInsertModeRemaps, + # respondToBashCommands, awaySummaryEnabled - UI + # prefs; marketplace aliases / archive + command + # plugin sources / self-hosted runners / gateway + # keys - enterprise or plugin territory. Env + # mention-only: ANTHROPIC_DEFAULT_MODEL (2.1.236), + # CLAUDE_CODE_MAX_WEB_SEARCHES_PER_SESSION + + # CLAUDE_CODE_MCP_AUTO_BACKGROUND_MS (2.1.212), + # CLAUDE_CODE_TOOL_MEMORY_LIMIT + + # CLAUDE_CODE_WEBFETCH_CACHE_TTL_MS (2.1.233), + # CLAUDE_CODE_GOAL_CHECKIN_MINUTES + + # CLAUDE_CODE_PROJECT_DIR_NAME (2.1.234). New hook + # event in window: DirectoryAdded (2.1.219). + # `/agents` wizard removed (2.1.198; no template + # referenced it). Issue #14920 (disable individual + # plugin skills) still open. } diff --git a/configure.py b/configure.py index dab3eb1..aec158f 100755 --- a/configure.py +++ b/configure.py @@ -733,7 +733,11 @@ def err(source, msg): and rel.parts[2] == "agents" and f.suffix == ".md" ): - fm = _frontmatter_block(f.read_text(encoding="utf-8")) + text = f.read_text(encoding="utf-8") + if text.startswith("\ufeff"): + err(src, "file starts with a UTF-8 BOM — Claude Code before 2.1.239 " + "silently ignores BOM-prefixed skills/agents; save without BOM") + fm = _frontmatter_block(text) if not fm: err(src, "missing YAML frontmatter (--- block at top of file)") continue @@ -890,6 +894,12 @@ def check_settings_validates(settings: dict) -> list: 3. Duplicate hook commands — the same hook `command` wired into more than one group under one event, so it fires multiple times per call (the F1 dogfood symptom). See _find_duplicate_hook_commands. + 4. `autoMode` in project settings — Claude Code >= 2.1.207 reads the + key only from ~/.claude/settings.json or managed settings, so a + project-scope block is silently ignored (v2.8.0). + 5. `Write(path)` / `NotebookEdit(path)` / `Glob(path)` / `MultiEdit(path)` + permission rules — never consulted; CC >= 2.1.210 warns at startup. + Only Edit(path) and Read(path) rules govern file access (v2.8.0). """ warnings = [] valid_overrides = {"on", "name-only", "user-invocable-only", "off"} @@ -930,9 +940,47 @@ def check_settings_validates(settings: dict) -> list: "group from a prior scaffold; run cc-configure --retrofit (the " "merge now collapses these) or remove the duplicate group." ) + + if "autoMode" in settings: + warnings.append( + "settings.json carries an `autoMode` block, but Claude Code >= 2.1.207 " + "reads autoMode only from ~/.claude/settings.json or managed settings " + "— project and local settings are ignored (code.claude.com/docs/en/" + "auto-mode-config). Move the block to ~/.claude/settings.json (see " + "docs/05-safety-permissions.md) and delete it here." + ) + + stale = _find_unconsulted_path_rules(settings) + if stale: + warnings.append( + "settings.json permission rules use path forms Claude Code never " + f"consults and warns about at startup (>= 2.1.210): {stale[:4]}" + + (f" (+{len(stale) - 4} more)" if len(stale) > 4 else "") + + ". Only Edit(path) and Read(path) govern file access — rewrite them " + "(Edit covers Write/NotebookEdit; Read covers Glob/Grep)." + ) return warnings +def _find_unconsulted_path_rules(settings: dict) -> list: + """Return permission rules of the form Write(...) / NotebookEdit(...) / + Glob(...) / MultiEdit(...) found in permissions.allow/ask/deny. Claude Code + checks file permissions against Edit(path) and Read(path) rules only; the + other spellings are accepted but never consulted (startup warning since + 2.1.210). Bare tool names (`Write`) are fine — only the parenthesised + path form is flagged.""" + import re + perms = settings.get("permissions") + if not isinstance(perms, dict): + return [] + found = [] + for key in ("allow", "ask", "deny"): + for rule in perms.get(key, []) or []: + if isinstance(rule, str) and re.match(r"^(Write|NotebookEdit|Glob|MultiEdit)\(", rule): + found.append(f"{key}: {rule}") + return found + + KNOWN_STACK_MANIFESTS = ( "package.json", "pyproject.toml", @@ -1368,9 +1416,13 @@ def _clone(x): # matchers pattern (the mcp drift-check ships under both `startup` and # `resume`) and never touches a user's own command (absent from new_groups). new_matchers = {} + new_shell = {} for ng in new_groups: for c in _hook_commands(ng): new_matchers.setdefault(c, set()).add(ng.get("matcher")) + for h in ng.get("hooks", []) or []: + if isinstance(h, dict) and h.get("command") and h.get("shell"): + new_shell[h["command"]] = h["shell"] for g in out: gm = g.get("matcher") hooks = g.get("hooks") @@ -1383,6 +1435,17 @@ def _clone(x): ] out = [g for g in out if g.get("hooks")] + # Field backfill (v2.8.0): the configurator started declaring `shell` on + # its command hooks (Windows: resolve Git Bash instead of PowerShell). The + # command-keyed union above keeps the user's existing entry, so copy the + # key onto configurator-owned entries that predate it. Never touches a + # user's own commands (absent from new_groups) or an explicit shell choice. + for g in out: + for h in g.get("hooks", []) or []: + if (isinstance(h, dict) and h.get("command") in new_shell + and "shell" not in h): + h["shell"] = new_shell[h["command"]] + groups_added = 0 commands_added = 0 for ng in new_groups: @@ -1413,6 +1476,23 @@ def _clone(x): return out, groups_added, commands_added +# Configurator defaults retired in v2.8.0 (CC 2.1.183-2.1.241 survey). A +# retrofit removes ONLY these exact shipped values from a project's existing +# settings.json; anything the user edited is left alone (and +# check_settings_validates flags it instead). +# - autoMode: Claude Code >= 2.1.207 reads the key only from +# ~/.claude/settings.json or managed settings, never from repo-resident +# files (code.claude.com/docs/en/auto-mode-config) - the block shipped by +# v2.6.0-v2.7.x safety patches was dead config. +# - Write(.env) / Write(.env.*): Edit(path) rules already cover the Write +# tool, and CC >= 2.1.210 warns at startup about Write(path) rules, +# which it never consults. +RETIRED_PROJECT_AUTOMODE = { + "hard_deny": ["Running executable files", "Writing to system directories"], +} +RETIRED_DENY_RULES = ("Write(.env)", "Write(.env.*)") + + def deep_merge_settings(existing: dict, new: dict): """Merge a user's existing .claude/settings.json with the configurator's new version. Returns (merged_dict, summary_str). @@ -1439,9 +1519,17 @@ def deep_merge_settings(existing: dict, new: dict): user's deliberate overrides). - statusLine, model: preserve existing if set; otherwise use new. - Unknown top-level keys: pass through verbatim from existing. + - Retired configurator defaults (RETIRED_PROJECT_AUTOMODE, the + RETIRED_DENY_RULES strings): removed when they match exactly what a + prior release shipped; user-edited variants survive. """ out = dict(existing) - counts = {"perms_added": 0, "hook_groups_added": 0, "hook_cmds_added": 0, "env_added": 0} + counts = {"perms_added": 0, "hook_groups_added": 0, "hook_cmds_added": 0, "env_added": 0, + "retired": 0} + + if out.get("autoMode") == RETIRED_PROJECT_AUTOMODE: + del out["autoMode"] + counts["retired"] += 1 if "$schema" in new: out["$schema"] = new["$schema"] @@ -1452,6 +1540,10 @@ def deep_merge_settings(existing: dict, new: dict): for key in ("allow", "ask", "deny", "additionalDirectories"): if key in new_perms: existing_list = out_perms.get(key, []) + if key == "deny": + kept = [r for r in existing_list if r not in RETIRED_DENY_RULES] + counts["retired"] += len(existing_list) - len(kept) + existing_list = kept merged = _merge_unique_list(existing_list, new_perms[key]) counts["perms_added"] += len(merged) - len(existing_list) out_perms[key] = merged @@ -1489,6 +1581,9 @@ def deep_merge_settings(existing: dict, new: dict): msg = (f"preserved existing config; added {counts['perms_added']} permission rule(s), " f"{counts['hook_groups_added']} hook group(s), {counts['hook_cmds_added']} hook command(s), " f"{counts['env_added']} env var(s)") + if counts["retired"]: + msg += (f"; removed {counts['retired']} retired configurator default(s) " + "(project-scope autoMode / Write(.env*) deny rules)") return out, msg @@ -1533,6 +1628,13 @@ def apply_structured_merges(files, target_dir): new_data = json.loads(f["content"]) if target == ".claude/settings.json": merged_data, msg = deep_merge_settings(existing_data, new_data) + # The preflight only sees the template merge; a retrofit is the + # one place a user's own settings.json enters the pipeline, so + # re-validate the merged result (v2.8.0: project-scope autoMode, + # never-consulted Write()/Glob() rules, stale duplicate groups). + merge_warnings = check_settings_validates(merged_data) + if merge_warnings: + f["merge_warnings"] = merge_warnings else: merged_data, msg = deep_merge_mcp(existing_data, new_data) f["content"] = json.dumps(merged_data, indent=2) + "\n" @@ -2728,8 +2830,9 @@ def main(): print(bold(yellow("[ SETTINGS WARNINGS ]"))) for w in settings_warnings: print(f" {yellow('!')} {w}") - print(dim(" These trigger Claude Code's settings-validator complaints in the user's editor.")) - print(dim(" File a configurator issue with the offending key + module — this is a template bug.")) + print(dim(" These trigger Claude Code's settings-validator complaints in the user's editor,")) + print(dim(" or name keys/rules current Claude Code silently ignores. If the offending entry")) + print(dim(" came from a template (not your own settings.json), file a configurator issue.")) # Surface module-level prerequisites that can't be fixed by the configurator. module_warnings = [] @@ -2905,6 +3008,9 @@ def main(): print(bold(blue("[ MERGED ]"))) for target_rel, msg in merge_messages: print(f" {tick} {target_rel} — {msg}") + for f in files: + for w in f.get("merge_warnings", []): + print(f" {yellow('!')} {f['target']}: {w}") if collision_report: print() diff --git a/docs/03-commands-and-hooks.md b/docs/03-commands-and-hooks.md index 3d026ff..41e5ffe 100644 --- a/docs/03-commands-and-hooks.md +++ b/docs/03-commands-and-hooks.md @@ -49,8 +49,8 @@ Great for injecting live repo state into the prompt. The `/review` and `/commit` ### Starter kit -- `/plan` — forces a structured plan before edits. -- `/review` — code review against main. +- `/plan` — forces a structured plan before edits. Shadows Claude Code's built-in `/plan` (the plan-mode shortcut) — a project skill wins by name; Shift+Tab still enters plan mode. +- `/review` — code review against main via the `code-reviewer` agent. Shadows the bundled `/review` alias that CC 2.1.223 added for `/code-review` (verified on 2.1.241) — type `/code-review` for the built-in multi-agent review. - `/commit` — Conventional Commits from staged diff. - `/ship` — full pre-push gauntlet. - `/sync-docs` — update `CLAUDE.md` / rules from recent work. @@ -63,11 +63,11 @@ Hooks are scripts that run on Claude Code lifecycle events. They're deterministi ### The event surface -Claude Code 2026 fires 27+ events. The ones that matter most day-to-day: +Claude Code fires 31 documented events (`code.claude.com/docs/en/hooks`). The ones that matter most day-to-day: | Event | When it fires | Common use | |---|---|---| -| `SessionStart` | Session begins or resumes | Inject current git status, prune logs | +| `SessionStart` | Session starts; `matcher` filters on the source: `startup`, `resume`, `clear`, `compact`, `fork` (2.1.214+ — forks used to report `resume`) | Inject current git status, prune logs | | `UserPromptSubmit` | User presses Enter | Log prompts, reject disallowed patterns | | `PreToolUse` | Before any tool call | Block dangerous bash, scan for secrets | | `PostToolUse` | After a tool call succeeds | Format files, run fast checks | @@ -76,6 +76,7 @@ Claude Code 2026 fires 27+ events. The ones that matter most day-to-day: | `SessionEnd` | Session terminates | Flush logs, send a summary | | `InstructionsLoaded` | When CLAUDE.md loads | Debug which files are in context | | `WorktreeCreate` | Before worktree creation | Abort if conditions wrong (only event where nonzero exit code blocks) | +| `DirectoryAdded` | After `/add-dir` registers another working directory (2.1.219+) | Re-run drift/stack checks for the new tree | ### Handler types @@ -134,6 +135,7 @@ From `templates/`: "hooks": [ { "type": "command", + "shell": "bash", "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/block-dangerous-bash.sh", "timeout": 10 } @@ -146,12 +148,18 @@ From `templates/`: The `$CLAUDE_PROJECT_DIR` env var is always set to the project root — use it instead of relative paths. Quote it for Windows compatibility. +### Matcher semantics (changed in 2026) + +- Matchers are exact-match sets separated by `|`. Hyphenated names (`code-reviewer`, `mcp__brave-search`) exact-match since 2.1.195; before that they were unanchored regexes (`code-reviewer` also fired for `senior-code-reviewer` — anchor as `^code-reviewer$` on older versions). +- Comma-separated matchers (`"Bash,PowerShell"`) work since 2.1.191; on earlier versions they silently never fired. The shipped templates use `|`. +- A hook `if:` condition with a single-segment directory pattern (`Edit(src/**)`) matches only `/src` since 2.1.214; write `Edit(**/src/**)` for any depth. `deny`/`ask` permission rules keep their any-depth match. + ### Defaults and caps - Default timeouts: 600s command, 30s prompt, 60s agent. - Injected context (additionalContext, systemMessage, stdout) capped at 10,000 chars. - Multiple hooks per event run in parallel. Identical commands are deduplicated. -- Windows: set `"shell": "powershell"` on command hooks. +- `shell`: every shipped command hook declares `"shell": "bash"` (CC 2.1.81+; the key is in the settings schema). Without it, a Windows session with no Git Bash defaults hooks to PowerShell and the `.sh` entrypoint dies on a parser error; with it, Claude Code resolves Git for Windows directly and prompts to install it when missing. Translate a hook to PowerShell and set `"shell": "powershell"` only when you want to drop the bash dependency. ### Debugging hooks diff --git a/docs/04-subagents-mcp-orchestration.md b/docs/04-subagents-mcp-orchestration.md index d144737..9581451 100644 --- a/docs/04-subagents-mcp-orchestration.md +++ b/docs/04-subagents-mcp-orchestration.md @@ -56,7 +56,9 @@ Only `name` and `description` are required. See `templates/agents/` for four rea - Each subagent starts with its own system prompt + minimal env info. **No inherited conversation**. - Only the subagent's final response returns to the parent. Verbose tool output stays in the subagent's transcript. -- **Subagents can spawn other subagents (CC ≥ 2.1.172, capped at 5 nesting levels).** Older CC versions can't. Prefer leaf designs anyway — use skills or chain through the main thread — so the architecture survives older CC and doesn't bury context behind deep subagent transcripts. +- **Subagents can spawn nested subagents — 3 levels below the main thread by default (CC ≥ 2.1.219; 2.1.172–2.1.216 allowed 5, 2.1.217–218 shipped nesting off).** Tune with `CLAUDE_CODE_MAX_SUBAGENT_SPAWN_DEPTH`; at the cap Claude Code withholds the Agent tool from the subagent. Prefer leaf designs anyway — use skills or chain through the main thread — so the architecture survives older CC and doesn't bury context behind deep subagent transcripts. +- **Subagents run in the background by default (CC ≥ 2.1.198)** with a narrower built-in tool set; their permission prompts surface in the main session. At most 20 run concurrently (`CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS`, 2.1.217+); the 21st spawn fails with an error that tells Claude not to retry. +- **Forks are the opposite of isolation.** `subagent_type: "fork"` (on by default since 2.1.232; `/subtask` from the prompt) inherits the whole conversation and prompt cache. Use a fork to continue *this* context elsewhere; use a named subagent for a specialist with its own prompt. - Parent `bypassPermissions` or `acceptEdits` **overrides** any `permissionMode` set on the subagent. ### Parallel subagents @@ -69,7 +71,7 @@ For sustained parallelism that exceeds context, use **agent teams** (separate se ### Built-in subagents you get for free -- `Explore` (Haiku, read-only) — fast codebase exploration. +- `Explore` (read-only) — fast codebase exploration. Since CC 2.1.198 it inherits the session model (capped at Opus) instead of always running on Haiku; define a project agent named `Explore` with `model: haiku` to pin it cheap. - `Plan` (inherits, read-only) — structured planning. - `general-purpose` — default catch-all. - `statusline-setup`. @@ -137,7 +139,7 @@ Local Claude Code for interactive work; Claude Code Web for long-running cloud j - The main thread's context is still finite. Subagents help with verbosity but not with total information you're holding in your head. - Parallel subagents that return detailed results still fill the parent. Prefer "ship a summary, not a transcript." -- Subagent nesting caps at 5 levels (CC ≥ 2.1.172; earlier versions: 0). If you need 3+ levels, the design is usually still wrong — re-shape into a fanout or pipeline through the main thread. +- Subagent nesting is capped — 3 levels below the main thread by default since CC 2.1.219 (`CLAUDE_CODE_MAX_SUBAGENT_SPAWN_DEPTH`). If you need 3+ levels, the design is usually still wrong — re-shape into a fanout or pipeline through the main thread, or a dynamic Workflow for fan-outs beyond a handful of workers. ## Recommendations diff --git a/docs/05-safety-permissions.md b/docs/05-safety-permissions.md index 394a769..0711ddb 100644 --- a/docs/05-safety-permissions.md +++ b/docs/05-safety-permissions.md @@ -36,7 +36,7 @@ In `.claude/settings.json`: "permissions": { "allow": ["Read", "Grep", "Bash(git status)", "Bash(npm test:*)"], "ask": ["Bash(git push:*)", "Bash(rm:*)"], - "deny": ["Write(.env)", "Bash(sudo:*)", "Bash(curl * | sh:*)"] + "deny": ["Edit(.env)", "Bash(sudo:*)", "Bash(curl * | sh:*)"] } } ``` @@ -49,11 +49,25 @@ Rules of thumb: ### Patterns that matter - `Bash(git push:*)` — matches any `git push …`. The colon-star is the prefix match. -- `Write(.env*)` — any Write whose path starts with `.env`. +- `Edit(.env.*)` — any edit whose path starts with `.env.`. `Edit(path)` rules cover the Write and NotebookEdit tools too; `Write(path)`, `NotebookEdit(path)` and `Glob(path)` rules are never consulted and Claude Code 2.1.210+ warns about them at startup — write `Edit()` / `Read()` rules only (a `Read` deny also blocks edits on the same path). - `Bash(rm -rf /:*)` — specific dangerous commands. The starter settings in `templates/core/dot-claude/settings.json` are a balanced baseline. +### Auto-mode rules live in user scope + +Auto mode (the classifier that reviews actions instead of prompting; the starting mode on Pro/Max/Team since CC 2.1.207) reads its `autoMode` customizations only from `~/.claude/settings.json` or managed settings. It deliberately ignores `.claude/settings.json` and `.claude/settings.local.json` — both live in the repo, so a checked-in file could inject its own allow rules (`code.claude.com/docs/en/auto-mode-config`). Configurator releases before v2.8.0 shipped an `autoMode.hard_deny` block in project settings; that block has been dead since 2.1.207, so the safety module no longer writes it, a retrofit removes the exact shipped block, and `[ SETTINGS WARNINGS ]` flags any other project-scope `autoMode`. To keep the old defaults, add them to `~/.claude/settings.json`: + +```json +{ + "autoMode": { + "hard_deny": ["$defaults", "Running executable files", "Writing to system directories"] + } +} +``` + +`"$defaults"` keeps the built-in rules; `soft_deny`, `allow` and `environment` follow the same shape. Two related keys *are* honored from project settings: `useAutoModeDuringPlan: false` (prompt instead of classify during plan mode) and `disableAutoMode: "disable"` (remove auto mode from the Shift+Tab cycle). + ## Layer 2: Hooks Permissions are checked against tool-call metadata. Hooks run actual code. Use them for things that need inspection, not just pattern-matching. diff --git a/docs/06-token-efficiency.md b/docs/06-token-efficiency.md index 8aa14d1..0079a33 100644 --- a/docs/06-token-efficiency.md +++ b/docs/06-token-efficiency.md @@ -81,15 +81,16 @@ Do a "context budget review" monthly: ## The model-choice lever -- Haiku for cheap, fast, read-only work — great for `Explore`, grep-style research, the `doc-writer` subagent. +- Haiku for cheap, fast, read-only work — grep-style research, the `doc-writer` subagent. Since CC 2.1.198 the built-in `Explore` agent inherits the session model (capped at Opus) instead of running on Haiku; define a project `Explore` agent with `model: haiku` to keep exploration cheap. - Fable 5 (`fable`, CC 2.1.170+) for day-to-day edits and reviews — the scaffolded default. It is the most capable model and premium-priced (check the platform pricing page for current rates); the rest of this doc's levers (Haiku subagents, narrowing flags, reset rhythm) are what keep that affordable. -- Opus for high-stakes review and deep refactors on Claude Code older than 2.1.170, or where your org's plan doesn't include Fable 5. +- Opus (Claude Opus 5 on CC 2.1.219+) for high-stakes review and deep refactors on Claude Code older than 2.1.170, or where your org's plan doesn't include Fable 5. Sonnet (Claude Sonnet 5 on CC 2.1.197+, Claude Code's own default model) is the cost-balanced middle for day-to-day edits when Fable's pricing doesn't fit. +- Task-tracking tools (`TodoWrite`, `TaskCreate`/`TaskUpdate`/…) are withdrawn on Fable 5, Opus 4.8+/5 and Sonnet 5 as of CC 2.1.233 — the shipped skills never named them, so nothing breaks, but set `CLAUDE_CODE_ENABLE_TODO_TOOLS=1` if a custom skill of yours depends on them. Most sessions can run primarily on Fable 5 with Haiku subagents for reads. On older Claude Code the `fable` alias isn't selectable — drop the project back to `sonnet` via `.claude/settings.local.json` (the example file ships the stub). ### Effort level (Pro/Max only) -Since Claude Code 2.1.117, Pro/Max subscribers on Opus 4.6 and Sonnet 4.6 default to `effort: high` (was `medium`). The default is already tuned — do **not** manually downgrade to `medium` thinking you're saving tokens. You're not; you're getting a less capable response for the same budget. If you want cheaper, drop to Haiku or use a lighter model. Reserve `effort: minimal` for skills that are mechanical (e.g. `/sync-docs`, `/check-context`, `/session-retro` — the `eff_effort_minimal` toggle stamps this automatically on those skills' frontmatter). +Since Claude Code 2.1.117, Pro/Max subscribers on Opus 4.6 and Sonnet 4.6 default to `effort: high` (was `medium`); on Fable 5, Opus 4.7+ and Sonnet 5 Claude Code's own default is `xhigh`. The default is already tuned — do **not** manually downgrade to `medium` thinking you're saving tokens. You're not; you're getting a less capable response for the same budget. If you want cheaper, drop to Haiku or use a lighter model. Reserve `effort: minimal` for skills that are mechanical (e.g. `/sync-docs`, `/check-context`, `/session-retro` — the `eff_effort_minimal` toggle stamps this automatically on those skills' frontmatter). ## When you hit the context wall mid-session diff --git a/examples/python-uv-fastapi/.claude/.cc-manifest.json b/examples/python-uv-fastapi/.claude/.cc-manifest.json index ce6b71c..30bd711 100644 --- a/examples/python-uv-fastapi/.claude/.cc-manifest.json +++ b/examples/python-uv-fastapi/.claude/.cc-manifest.json @@ -1,8 +1,8 @@ { "manifest_version": 2, - "written_at": "2026-06-10T03:58:53Z", - "written_by": "cc-configure 2.6.0", - "written_by_sha": "1ec8c2c", + "written_at": "2026-08-24T17:56:11Z", + "written_by": "cc-configure 2.7.0", + "written_by_sha": "c29786d", "mcp_servers": [], "stack_manifests": [], "check_commands": { diff --git a/examples/python-uv-fastapi/.claude/hooks/microbit-enforcer.sh b/examples/python-uv-fastapi/.claude/hooks/microbit-enforcer.sh index 111bc1f..e967564 100755 --- a/examples/python-uv-fastapi/.claude/hooks/microbit-enforcer.sh +++ b/examples/python-uv-fastapi/.claude/hooks/microbit-enforcer.sh @@ -14,7 +14,7 @@ # Lifecycle: a SessionStart hook (registered alongside this hook by the # configurator's settings-patch, matcher startup|clear) clears all three # files on a fresh session or /clear. Markers are session-scoped but -# survive --resume and compaction, so a long session keeps its markers. +# survive --resume, compaction and /fork, so a long session keeps its markers. set -euo pipefail diff --git a/examples/python-uv-fastapi/.claude/hooks/stop-run-checks.sh b/examples/python-uv-fastapi/.claude/hooks/stop-run-checks.sh index b38b35c..23d3f08 100755 --- a/examples/python-uv-fastapi/.claude/hooks/stop-run-checks.sh +++ b/examples/python-uv-fastapi/.claude/hooks/stop-run-checks.sh @@ -46,8 +46,15 @@ except Exception: ' 2>/dev/null || echo 0)" if [ "$BG_COUNT" -gt 0 ]; then exit 0; fi -# label|command — labels are display-only, commands come from cc-configure. -# Leave a value empty after the `|` to skip that check entirely. +# label|command or label|command|service +# - labels are display-only; commands come from cc-configure. +# - leave the command empty (label|) to skip a check entirely. +# - add a docker-compose service as a 3rd field to run that check INSIDE the +# container (reuses your existing service): runs `docker compose exec -T +# ` if it's up, else `docker compose run --rm +# `; skips silently if docker/compose or the service is absent. +# Service names are single tokens; a containerized command can't contain `|`. +# Example: "test|pytest|backend" "typecheck|pnpm typecheck|frontend" CHECKS=( "typecheck|uv run mypy ." "lint|uv run ruff check" @@ -71,24 +78,64 @@ manifest_for() { esac } +# Echo a working `docker compose` invocation, or return non-zero when the +# docker compose v2 CLI isn't usable here (so container checks skip, fail-open). +compose() { + command -v docker >/dev/null 2>&1 || return 1 + docker compose version >/dev/null 2>&1 || return 1 + echo "docker compose" +} + +# Probe the compose CLI once (it's a daemon round-trip); container-bound checks +# below reuse $CC. Empty string ⇒ compose unusable ⇒ those checks skip. +CC="$(compose)" || CC="" + REPORT="" for entry in "${CHECKS[@]}"; do label="${entry%%|*}" - cmd="${entry#*|}" + rest="${entry#*|}" + + # Optional 3rd field = compose service. An entry is container-bound iff it + # has >=2 pipes AND its final field is a bare token (compose service names + # match [A-Za-z0-9._-]+). Otherwise it's a host command — `cmd` is everything + # after the first pipe, so host commands may still contain a literal `|`. + pipes="$(printf '%s' "$entry" | tr -cd '|' | wc -c | tr -d ' ')" + last="${entry##*|}" + if [ "$pipes" -ge 2 ] && printf '%s' "$last" | grep -qxE '[A-Za-z0-9._-]+'; then + service="$last" + cmd="${rest%|*}" + else + service="" + cmd="$rest" + fi # Skip empty (user opted out of this check during intake). [ -z "$cmd" ] && continue - # Skip if the first binary in the command isn't on PATH. - first="$(printf '%s' "$cmd" | awk '{print $1}')" - if ! command -v "$first" >/dev/null 2>&1; then continue; fi - - # Skip if the stack manifest hasn't landed yet. Project-relative path — - # the cd above pins us to CLAUDE_PROJECT_DIR. - manifest="$(manifest_for "$first")" - if [ -n "$manifest" ] && [ ! -f "$manifest" ]; then continue; fi + if [ -z "$service" ]; then + # --- HOST branch (unchanged) --- + first="$(printf '%s' "$cmd" | awk '{print $1}')" + if ! command -v "$first" >/dev/null 2>&1; then continue; fi + manifest="$(manifest_for "$first")" + if [ -n "$manifest" ] && [ ! -f "$manifest" ]; then continue; fi + run_cmd="$cmd" + else + # --- CONTAINER branch --- + [ -z "$CC" ] && { echo "[stop-check] ${label}: docker compose unavailable; skipping container check" >&2; continue; } + if ! $CC config --services 2>/dev/null | grep -qxF "$service"; then + echo "[stop-check] ${label}: compose service '${service}' not defined; skipping" >&2 + continue + fi + # -T on both: hooks run non-interactively, so disable PTY allocation (else + # compose < v2.2.0 prints "input device is not a TTY" into the report). + if $CC ps --status running --services 2>/dev/null | grep -qxF "$service"; then + run_cmd="$CC exec -T ${service} ${cmd}" + else + run_cmd="$CC run --rm -T ${service} ${cmd}" + fi + fi - out=$(eval "$cmd" 2>&1) && status=0 || status=$? + out=$(eval "$run_cmd" 2>&1) && status=0 || status=$? if [ $status -eq 0 ]; then REPORT="${REPORT}[stop-check] ${label}: OK"$'\n' else diff --git a/examples/python-uv-fastapi/.claude/settings.json b/examples/python-uv-fastapi/.claude/settings.json index 04d7b9c..3ab025e 100644 --- a/examples/python-uv-fastapi/.claude/settings.json +++ b/examples/python-uv-fastapi/.claude/settings.json @@ -39,8 +39,6 @@ "Bash(curl * | sh:*)", "Bash(curl * | bash:*)", "Bash(wget * | sh:*)", - "Write(.env)", - "Write(.env.*)", "Edit(.env)", "Edit(.env.*)" ], @@ -53,11 +51,13 @@ "hooks": [ { "type": "command", + "shell": "bash", "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/block-dangerous-bash.sh", "timeout": 10 }, { "type": "command", + "shell": "bash", "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/check-package-availability.sh", "timeout": 10 } @@ -68,6 +68,7 @@ "hooks": [ { "type": "command", + "shell": "bash", "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/scan-secrets.sh", "timeout": 10 } @@ -78,6 +79,7 @@ "hooks": [ { "type": "command", + "shell": "bash", "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/microbit-enforcer.sh", "timeout": 5 } @@ -90,6 +92,7 @@ "hooks": [ { "type": "command", + "shell": "bash", "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/format-on-write.sh", "timeout": 30 } @@ -100,6 +103,7 @@ "hooks": [ { "type": "command", + "shell": "bash", "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/truncate-bash-output.sh", "timeout": 10 } @@ -111,6 +115,7 @@ "hooks": [ { "type": "command", + "shell": "bash", "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/stop-run-checks.sh", "timeout": 120 } @@ -122,6 +127,7 @@ "hooks": [ { "type": "command", + "shell": "bash", "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/pre-compact-snapshot.sh", "timeout": 15 } @@ -134,6 +140,7 @@ "hooks": [ { "type": "command", + "shell": "bash", "command": "rm -f \"$CLAUDE_PROJECT_DIR\"/.claude/.frozen \"$CLAUDE_PROJECT_DIR\"/.claude/.guarded \"$CLAUDE_PROJECT_DIR\"/.claude/.careful || true", "timeout": 5 } @@ -145,11 +152,5 @@ "SLOP_SCAN_ACTION": "warn", "CLAUDE_BASH_MAX_LINES": "80" }, - "model": "fable", - "autoMode": { - "hard_deny": [ - "Running executable files", - "Writing to system directories" - ] - } + "model": "fable" } diff --git a/examples/python-uv-fastapi/.claude/settings.local.json.example b/examples/python-uv-fastapi/.claude/settings.local.json.example index 5e72b93..22d53a6 100644 --- a/examples/python-uv-fastapi/.claude/settings.local.json.example +++ b/examples/python-uv-fastapi/.claude/settings.local.json.example @@ -2,7 +2,7 @@ "$schema": "https://json.schemastore.org/claude-code-settings.json", "//": "Personal settings — gitignored. Override anything from .claude/settings.json here. Copy this file to .claude/settings.local.json to activate; the .example suffix means Claude Code does not parse this file directly.", - "// env": "Durable home for env vars including MCP auth tokens — e.g. SONATYPE_TOKEN (security-auditor agent) or GITHUB_TOKEN (mcp_github). Persists across shells, so MCPs keep working even when Claude is launched from a shell that doesn't inherit your tokens. Add entries below alongside EDITOR. CLAUDE_CODE_STOP_HOOK_BLOCK_CAP (CC 2.1.143+, overrides default 8-block cap for consecutive Stop hook blocks) lives here too — add the env var if a safety Stop hook intentionally blocks repeatedly.", + "// env": "Durable home for env vars including MCP auth tokens — e.g. SONATYPE_TOKEN (security-auditor agent) or GITHUB_TOKEN (mcp_github). Persists across shells, so MCPs keep working even when Claude is launched from a shell that doesn't inherit your tokens. Add entries below alongside EDITOR. CLAUDE_CODE_STOP_HOOK_BLOCK_CAP (CC 2.1.143+, overrides default 8-block cap for consecutive Stop hook blocks) lives here too — add the env var if a safety Stop hook intentionally blocks repeatedly. ANTHROPIC_DEFAULT_MODEL (CC 2.1.236+) is the env-var alternative to the `model` key below: it sets the model new sessions start on while a /model pick still overrides it and persists.", "env": { "EDITOR": "code" }, @@ -14,11 +14,15 @@ ] }, - "//opt-ins": "===== OPT-IN STUBS — uncomment by removing the leading `// ` from any block, then drop this `//opt-ins` line. Each block is a documented Claude Code setting the configurator does not ship as an active default but supports users opting into per-project. All schemastore-validated as of 2026-05-23. Inline `//N` keys are explainer notes; the values shown are reasonable defaults you'll likely want to tune.", + "//opt-ins": "===== OPT-IN STUBS — uncomment by removing the leading `// ` from any block, then drop this `//opt-ins` line. Each block is a documented Claude Code setting the configurator does not ship as an active default but supports users opting into per-project. All schemastore-validated as of 2026-07-27 (SchemaStore sync to Claude Code 2.1.220, PR #6131). Inline `//N` keys are explainer notes; the values shown are reasonable defaults you'll likely want to tune.", + + "//autoMode-notes": "NOT an opt-in here: the auto-mode classifier does not read `autoMode` (hard_deny / soft_deny / allow / environment / classifyAllShell) from this file or from .claude/settings.json — both live in the repo, so a checked-in file could inject allow rules (CC 2.1.207+; code.claude.com/docs/en/auto-mode-config). Put autoMode in ~/.claude/settings.json instead; docs/05-safety-permissions.md has the copy-paste block that earlier configurator releases shipped here.", "// sandbox": { "//1": "sandbox.network.deniedDomains (CC 2.1.113+) — data-exfiltration-resistant baseline. Takes effect only when the sandbox is otherwise active for the command. Supports wildcards (*.example.com).", "//2": "sandbox.failIfUnavailable (CC 2.1.143+) — when true, sandbox startup is a hard failure if dependencies are missing (no fall-back to non-sandboxed execution). Fail-closed posture for safety-sensitive projects.", + "//3": "sandbox.credentials (CC 2.1.187+) — hide credential files and secret env vars from sandboxed Bash commands. `deny` entries are honored from any settings scope (this file included) and merge across scopes; `mask` entries, allowPlaintextInject, awsPairs and sigv4 are honored only from ~/.claude/settings.json, managed settings, or --settings. Env-var denial still applies when filesystem isolation is off; file denial doesn't. MCP servers are not sandboxed Bash, so an MCP token exported in the `env` block above keeps working for its server while Bash can't read it.", + "//4": "Not stubbed because project scope can't set them: sandbox.network.strictAllowlist (CC 2.1.219+) and sandbox.filesystem.disabled (CC 2.1.216+) are read only from user, managed, or --settings sources.", "network": { "deniedDomains": [ "pastebin.com", @@ -33,21 +37,37 @@ "uguu.se" ] }, - "failIfUnavailable": true + "failIfUnavailable": true, + "credentials": { + "files": [ + { "path": "~/.aws/credentials", "mode": "deny" }, + { "path": "~/.ssh", "mode": "deny" } + ], + "envVars": [ + { "name": "GITHUB_TOKEN", "mode": "deny" }, + { "name": "SONATYPE_TOKEN", "mode": "deny" } + ] + } }, "// worktree": { "//1": "worktree.baseRef (CC 2.1.133+) — values: 'fresh' (default; new worktrees from origin/HEAD) or 'head' (preserves unpushed local commits).", "//2": "worktree.bgIsolation (CC 2.1.143+) — values: 'worktree' (default; background sessions get their own worktree) or 'none' (background sessions edit the live working copy directly).", + "//3": "worktree.symlinkDirectories — directories symlinked from the main checkout into each worktree instead of duplicated on disk (nothing is symlinked by default). node_modules is the usual candidate; drop the entry if your build writes into it per-branch.", + "//4": "worktree.sparsePaths (CC 2.1.76+) — array of directories to check out via git sparse-checkout (cone mode) in each worktree; only those paths hit the disk. Monorepo-only; omit it entirely (an empty array is not a no-op) unless you know which paths a session needs.", "baseRef": "head", - "bgIsolation": "worktree" + "bgIsolation": "worktree", + "symlinkDirectories": ["node_modules"] }, + "// agent": "code-reviewer", + "//agent-notes": "agent — name of a subagent (built-in or .claude/agents/.md) whose system prompt, tool restrictions and model the MAIN thread takes on for every session in this project (same as `claude --agent `; the choice persists on resume). Never a configurator default: it replaces the default Claude Code system prompt entirely. Useful for a review-only checkout or a docs-only worktree.", + "// prUrlTemplate": "https://gitlab.com/{owner}/{repo}/-/merge_requests/{number}", - "//prUrlTemplate-notes": "prUrlTemplate (CC 2.1.119+) — footer PR badge + tool-result summaries follow this URL template instead of github.com. Placeholders: {host}, {owner}, {repo}, {number}. Examples: GitLab as shown above; Bitbucket: https://bitbucket.org/{owner}/{repo}/pull-requests/{number}; GHE: https://gh.example.com/{owner}/{repo}/pull/{number}.", + "//prUrlTemplate-notes": "prUrlTemplate (CC 2.1.119+) — footer PR badge + tool-result summaries follow this URL template instead of github.com. Placeholders: {host}, {owner}, {repo}, {number}. Examples: GitLab as shown above; Bitbucket: https://bitbucket.org/{owner}/{repo}/pull-requests/{number}; GHE: https://gh.example.com/{owner}/{repo}/pull/{number}. Since CC 2.1.234 GitLab remotes with an authenticated `glab` CLI get an MR badge without this setting.", "// subagentStatusLine": { - "//1": "subagentStatusLine (CC 2.1.143+) — distinct statusline for subagent runs so they're visually separable from the parent session. Requires statusline.sh to handle the --subagent flag (the shipped script does not yet — wrap or fork before uncommenting).", + "//1": "subagentStatusLine (CC 2.1.143+) — distinct statusline for subagent runs so they're visually separable from the parent session. Requires statusline.sh to handle the --subagent flag (the shipped script does not yet — wrap or fork before uncommenting). The payload carries the subagent's model and, since CC 2.1.214, its reasoning effort.", "//2": "Note on the sibling statusLine.hideVimModeIndicator opt-in (CC 2.1.143+): only useful if you wrote a custom statusline.sh that renders its own vim mode display. To enable, copy your full statusLine block from .claude/settings.json (settings.local.json's statusLine REPLACES rather than merges with the base config) and add `\"hideVimModeIndicator\": true` to it.", "type": "command", "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/statusline.sh --subagent" @@ -59,6 +79,15 @@ "//3": "Add your own skill-name → enum-value entries below. The configurator does not ship default entries because guessing which skills to compress is project-specific." }, + "//token-efficiency-notes": "Four context-budget keys, all project-scope-honored: claudeMdExcludes — glob patterns (matched against absolute paths, merged across settings layers) for CLAUDE.md files to skip when memory loads, e.g. another team's tree in a monorepo or vendored checkouts; skillListingBudgetFraction — share of the context window reserved for the per-turn skill listing (default 0.01 = 1%; descriptions are truncated to fit, so raise it when your skill descriptions get cut short, lower it to fit more skills); skillListingMaxDescChars — per-skill cap on the description + when-to-use text in that listing; disableBundledSkills (CC 2.1.169+) — removes every bundled skill and workflow except /doctor (`/code-review`, `/loop`, `/claude-api`, `/dataviz`, `/deep-research`, `/batch`, `/debug` …) and hides built-in slash commands from the model; project, user and plugin skills are unaffected. Weigh it against losing `/code-review` — the built-in multi-agent review — before turning it on.", + "// claudeMdExcludes": ["**/vendor/**/CLAUDE.md", "**/third_party/**/CLAUDE.md"], + "// skillListingBudgetFraction": 0.02, + "// skillListingMaxDescChars": 2048, + "// disableBundledSkills": true, + + "// fallbackModel": ["opus", "sonnet"], + "//fallbackModel-notes": "fallbackModel (CC 2.1.166+) — ordered chain (up to three entries; model names or aliases; 'default' expands to the account default) tried when the primary model is overloaded, unavailable, or returns a non-retryable server error. Position carries meaning, so Claude Code takes the whole array from the highest-precedence file that sets it. The --fallback-model CLI flag overrides it.", + "// model": "sonnet", - "//model-notes": "Per-machine model override. The scaffolded project default is 'fable' (Claude Fable 5 — requires CC 2.1.170+; premium pricing; safety-classifier fallback in cybersecurity/biology domains; unavailable under zero-data-retention). Uncomment to drop THIS machine back to 'sonnet' when its CC predates 2.1.170, Fable 5 isn't in your org's plan, or the cost posture doesn't fit. 'best' is the middle path: resolves to Fable 5 where available, latest Opus otherwise." + "//model-notes": "Per-machine model override. The scaffolded project default is 'fable' (Claude Fable 5 — requires CC 2.1.170+; premium pricing; safety-classifier fallback in cybersecurity/biology domains; unavailable under zero-data-retention). Uncomment to drop THIS machine back to 'sonnet' (Claude Sonnet 5 on CC 2.1.197+, Claude Code's own default model) when its CC predates 2.1.170, Fable 5 isn't in your org's plan, or the cost posture doesn't fit. 'opus' resolves to Claude Opus 5 on CC 2.1.219+. 'best' is the middle path: resolves to Fable 5 where available, latest Opus otherwise." } diff --git a/examples/python-uv-fastapi/.claude/skills/review/SKILL.md b/examples/python-uv-fastapi/.claude/skills/review/SKILL.md index 856cf10..ea97059 100644 --- a/examples/python-uv-fastapi/.claude/skills/review/SKILL.md +++ b/examples/python-uv-fastapi/.claude/skills/review/SKILL.md @@ -1,6 +1,6 @@ --- name: review -description: Code-review the current branch's changes against main. Use before committing or pushing. +description: Code-review the current branch's changes against main with the code-reviewer agent. Use before committing or pushing. Distinct from Claude Code's built-in /code-review. argument-hint: "[optional focus area]" allowed-tools: Read Grep Glob Bash(git diff:*) Bash(git log:*) Bash(git status) context: fork diff --git a/templates/commands/infinite/SKILL.md b/templates/commands/infinite/SKILL.md index 8cab1b1..d26463f 100644 --- a/templates/commands/infinite/SKILL.md +++ b/templates/commands/infinite/SKILL.md @@ -48,13 +48,13 @@ Pick a batching plan based on `count`: | count | strategy | |-------------|----------------------------------------------------------------------| -| 1-5 | Launch all subagents simultaneously in a single parallel Task call. | +| 1-5 | Launch all subagents simultaneously in a single parallel Agent-tool call (the tool was named Task before CC 2.1.63; the alias still works). | | 6-20 | Batches of 5. Launch batch, wait for all, then launch next. | | infinite | Waves of 3-5. After each wave, check context usage; if > 50% start a fresh session and hand off the directory snapshot. Stop if user says stop. | ## Phase 4 — Parallel agent coordination -Dispatch subagents via the Task tool using the `parallel-generator` subagent. Each call's prompt must contain exactly these five sections (keep them labeled): +Dispatch subagents via the Agent tool (formerly Task) using the `parallel-generator` subagent. Claude Code caps concurrent subagents at 20 (`CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS`, CC ≥ 2.1.217) — the wave sizes above stay well under it. Each call's prompt must contain exactly these five sections (keep them labeled): 1. **spec_context** — the full contents of `spec_file` (or a summary if > 2k chars; include the path). 2. **claimed_slots_manifest** — the snapshot from Phase 2, updated with any slots claimed by earlier waves. @@ -62,7 +62,7 @@ Dispatch subagents via the Task tool using the `parallel-generator` subagent. Ea 4. **diversification_axis** — the axis this iteration should differ on, named explicitly. e.g. "this iteration must use a dark color palette"; "this iteration must favor imperative style over declarative." 5. **quality_standards** — the must-be-true bullets from Phase 1, verbatim. -**Do not ask subagents to coordinate with each other.** Their contexts don't share — even with nested-subagent support (CC ≥ 2.1.172, ≤ 5 levels), a child can't reach a sibling's transcript. The claimed_slots_manifest + diversification_axis does the coordination for them. +**Do not ask subagents to coordinate with each other.** Their contexts don't share — even with nested-subagent support (3 levels by default since CC 2.1.219, `CLAUDE_CODE_MAX_SUBAGENT_SPAWN_DEPTH`), a child can't reach a sibling's transcript. The claimed_slots_manifest + diversification_axis does the coordination for them. ## After each wave diff --git a/templates/commands/microbit-enforcer/microbit-enforcer.sh b/templates/commands/microbit-enforcer/microbit-enforcer.sh index 111bc1f..e967564 100755 --- a/templates/commands/microbit-enforcer/microbit-enforcer.sh +++ b/templates/commands/microbit-enforcer/microbit-enforcer.sh @@ -14,7 +14,7 @@ # Lifecycle: a SessionStart hook (registered alongside this hook by the # configurator's settings-patch, matcher startup|clear) clears all three # files on a fresh session or /clear. Markers are session-scoped but -# survive --resume and compaction, so a long session keeps its markers. +# survive --resume, compaction and /fork, so a long session keeps its markers. set -euo pipefail diff --git a/templates/commands/microbit-enforcer/settings-patch.json b/templates/commands/microbit-enforcer/settings-patch.json index e99277d..d217cf6 100644 --- a/templates/commands/microbit-enforcer/settings-patch.json +++ b/templates/commands/microbit-enforcer/settings-patch.json @@ -1,5 +1,5 @@ { - "//": "Auto-installed alongside the freeze/unfreeze/guard/careful microbits when commands.subset is 'full' or 'rigorous'. PreToolUse rejects Write/Edit/NotebookEdit when a marker is present; SessionStart (matcher startup|clear) clears markers only on a fresh session or /clear — resume and compaction preserve them so a long session keeps its freezes.", + "//": "Auto-installed alongside the freeze/unfreeze/guard/careful microbits when commands.subset is 'full' or 'rigorous'. PreToolUse rejects Write/Edit/NotebookEdit when a marker is present; SessionStart (matcher startup|clear) clears markers only on a fresh session or /clear — resume, compaction and fork (`source: fork`, CC 2.1.214+) preserve them so a long session keeps its freezes.", "hooks": { "PreToolUse": [ { @@ -7,6 +7,7 @@ "hooks": [ { "type": "command", + "shell": "bash", "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/microbit-enforcer.sh", "timeout": 5 } @@ -19,6 +20,7 @@ "hooks": [ { "type": "command", + "shell": "bash", "command": "rm -f \"$CLAUDE_PROJECT_DIR\"/.claude/.frozen \"$CLAUDE_PROJECT_DIR\"/.claude/.guarded \"$CLAUDE_PROJECT_DIR\"/.claude/.careful || true", "timeout": 5 } diff --git a/templates/commands/review/SKILL.md b/templates/commands/review/SKILL.md index 856cf10..ea97059 100644 --- a/templates/commands/review/SKILL.md +++ b/templates/commands/review/SKILL.md @@ -1,6 +1,6 @@ --- name: review -description: Code-review the current branch's changes against main. Use before committing or pushing. +description: Code-review the current branch's changes against main with the code-reviewer agent. Use before committing or pushing. Distinct from Claude Code's built-in /code-review. argument-hint: "[optional focus area]" allowed-tools: Read Grep Glob Bash(git diff:*) Bash(git log:*) Bash(git status) context: fork diff --git a/templates/core/dot-claude/settings.json b/templates/core/dot-claude/settings.json index 0b9636a..d2a2a59 100644 --- a/templates/core/dot-claude/settings.json +++ b/templates/core/dot-claude/settings.json @@ -39,8 +39,6 @@ "Bash(curl * | sh:*)", "Bash(curl * | bash:*)", "Bash(wget * | sh:*)", - "Write(.env)", - "Write(.env.*)", "Edit(.env)", "Edit(.env.*)" ] diff --git a/templates/core/dot-claude/settings.local.json.example b/templates/core/dot-claude/settings.local.json.example index 5e72b93..22d53a6 100644 --- a/templates/core/dot-claude/settings.local.json.example +++ b/templates/core/dot-claude/settings.local.json.example @@ -2,7 +2,7 @@ "$schema": "https://json.schemastore.org/claude-code-settings.json", "//": "Personal settings — gitignored. Override anything from .claude/settings.json here. Copy this file to .claude/settings.local.json to activate; the .example suffix means Claude Code does not parse this file directly.", - "// env": "Durable home for env vars including MCP auth tokens — e.g. SONATYPE_TOKEN (security-auditor agent) or GITHUB_TOKEN (mcp_github). Persists across shells, so MCPs keep working even when Claude is launched from a shell that doesn't inherit your tokens. Add entries below alongside EDITOR. CLAUDE_CODE_STOP_HOOK_BLOCK_CAP (CC 2.1.143+, overrides default 8-block cap for consecutive Stop hook blocks) lives here too — add the env var if a safety Stop hook intentionally blocks repeatedly.", + "// env": "Durable home for env vars including MCP auth tokens — e.g. SONATYPE_TOKEN (security-auditor agent) or GITHUB_TOKEN (mcp_github). Persists across shells, so MCPs keep working even when Claude is launched from a shell that doesn't inherit your tokens. Add entries below alongside EDITOR. CLAUDE_CODE_STOP_HOOK_BLOCK_CAP (CC 2.1.143+, overrides default 8-block cap for consecutive Stop hook blocks) lives here too — add the env var if a safety Stop hook intentionally blocks repeatedly. ANTHROPIC_DEFAULT_MODEL (CC 2.1.236+) is the env-var alternative to the `model` key below: it sets the model new sessions start on while a /model pick still overrides it and persists.", "env": { "EDITOR": "code" }, @@ -14,11 +14,15 @@ ] }, - "//opt-ins": "===== OPT-IN STUBS — uncomment by removing the leading `// ` from any block, then drop this `//opt-ins` line. Each block is a documented Claude Code setting the configurator does not ship as an active default but supports users opting into per-project. All schemastore-validated as of 2026-05-23. Inline `//N` keys are explainer notes; the values shown are reasonable defaults you'll likely want to tune.", + "//opt-ins": "===== OPT-IN STUBS — uncomment by removing the leading `// ` from any block, then drop this `//opt-ins` line. Each block is a documented Claude Code setting the configurator does not ship as an active default but supports users opting into per-project. All schemastore-validated as of 2026-07-27 (SchemaStore sync to Claude Code 2.1.220, PR #6131). Inline `//N` keys are explainer notes; the values shown are reasonable defaults you'll likely want to tune.", + + "//autoMode-notes": "NOT an opt-in here: the auto-mode classifier does not read `autoMode` (hard_deny / soft_deny / allow / environment / classifyAllShell) from this file or from .claude/settings.json — both live in the repo, so a checked-in file could inject allow rules (CC 2.1.207+; code.claude.com/docs/en/auto-mode-config). Put autoMode in ~/.claude/settings.json instead; docs/05-safety-permissions.md has the copy-paste block that earlier configurator releases shipped here.", "// sandbox": { "//1": "sandbox.network.deniedDomains (CC 2.1.113+) — data-exfiltration-resistant baseline. Takes effect only when the sandbox is otherwise active for the command. Supports wildcards (*.example.com).", "//2": "sandbox.failIfUnavailable (CC 2.1.143+) — when true, sandbox startup is a hard failure if dependencies are missing (no fall-back to non-sandboxed execution). Fail-closed posture for safety-sensitive projects.", + "//3": "sandbox.credentials (CC 2.1.187+) — hide credential files and secret env vars from sandboxed Bash commands. `deny` entries are honored from any settings scope (this file included) and merge across scopes; `mask` entries, allowPlaintextInject, awsPairs and sigv4 are honored only from ~/.claude/settings.json, managed settings, or --settings. Env-var denial still applies when filesystem isolation is off; file denial doesn't. MCP servers are not sandboxed Bash, so an MCP token exported in the `env` block above keeps working for its server while Bash can't read it.", + "//4": "Not stubbed because project scope can't set them: sandbox.network.strictAllowlist (CC 2.1.219+) and sandbox.filesystem.disabled (CC 2.1.216+) are read only from user, managed, or --settings sources.", "network": { "deniedDomains": [ "pastebin.com", @@ -33,21 +37,37 @@ "uguu.se" ] }, - "failIfUnavailable": true + "failIfUnavailable": true, + "credentials": { + "files": [ + { "path": "~/.aws/credentials", "mode": "deny" }, + { "path": "~/.ssh", "mode": "deny" } + ], + "envVars": [ + { "name": "GITHUB_TOKEN", "mode": "deny" }, + { "name": "SONATYPE_TOKEN", "mode": "deny" } + ] + } }, "// worktree": { "//1": "worktree.baseRef (CC 2.1.133+) — values: 'fresh' (default; new worktrees from origin/HEAD) or 'head' (preserves unpushed local commits).", "//2": "worktree.bgIsolation (CC 2.1.143+) — values: 'worktree' (default; background sessions get their own worktree) or 'none' (background sessions edit the live working copy directly).", + "//3": "worktree.symlinkDirectories — directories symlinked from the main checkout into each worktree instead of duplicated on disk (nothing is symlinked by default). node_modules is the usual candidate; drop the entry if your build writes into it per-branch.", + "//4": "worktree.sparsePaths (CC 2.1.76+) — array of directories to check out via git sparse-checkout (cone mode) in each worktree; only those paths hit the disk. Monorepo-only; omit it entirely (an empty array is not a no-op) unless you know which paths a session needs.", "baseRef": "head", - "bgIsolation": "worktree" + "bgIsolation": "worktree", + "symlinkDirectories": ["node_modules"] }, + "// agent": "code-reviewer", + "//agent-notes": "agent — name of a subagent (built-in or .claude/agents/.md) whose system prompt, tool restrictions and model the MAIN thread takes on for every session in this project (same as `claude --agent `; the choice persists on resume). Never a configurator default: it replaces the default Claude Code system prompt entirely. Useful for a review-only checkout or a docs-only worktree.", + "// prUrlTemplate": "https://gitlab.com/{owner}/{repo}/-/merge_requests/{number}", - "//prUrlTemplate-notes": "prUrlTemplate (CC 2.1.119+) — footer PR badge + tool-result summaries follow this URL template instead of github.com. Placeholders: {host}, {owner}, {repo}, {number}. Examples: GitLab as shown above; Bitbucket: https://bitbucket.org/{owner}/{repo}/pull-requests/{number}; GHE: https://gh.example.com/{owner}/{repo}/pull/{number}.", + "//prUrlTemplate-notes": "prUrlTemplate (CC 2.1.119+) — footer PR badge + tool-result summaries follow this URL template instead of github.com. Placeholders: {host}, {owner}, {repo}, {number}. Examples: GitLab as shown above; Bitbucket: https://bitbucket.org/{owner}/{repo}/pull-requests/{number}; GHE: https://gh.example.com/{owner}/{repo}/pull/{number}. Since CC 2.1.234 GitLab remotes with an authenticated `glab` CLI get an MR badge without this setting.", "// subagentStatusLine": { - "//1": "subagentStatusLine (CC 2.1.143+) — distinct statusline for subagent runs so they're visually separable from the parent session. Requires statusline.sh to handle the --subagent flag (the shipped script does not yet — wrap or fork before uncommenting).", + "//1": "subagentStatusLine (CC 2.1.143+) — distinct statusline for subagent runs so they're visually separable from the parent session. Requires statusline.sh to handle the --subagent flag (the shipped script does not yet — wrap or fork before uncommenting). The payload carries the subagent's model and, since CC 2.1.214, its reasoning effort.", "//2": "Note on the sibling statusLine.hideVimModeIndicator opt-in (CC 2.1.143+): only useful if you wrote a custom statusline.sh that renders its own vim mode display. To enable, copy your full statusLine block from .claude/settings.json (settings.local.json's statusLine REPLACES rather than merges with the base config) and add `\"hideVimModeIndicator\": true` to it.", "type": "command", "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/statusline.sh --subagent" @@ -59,6 +79,15 @@ "//3": "Add your own skill-name → enum-value entries below. The configurator does not ship default entries because guessing which skills to compress is project-specific." }, + "//token-efficiency-notes": "Four context-budget keys, all project-scope-honored: claudeMdExcludes — glob patterns (matched against absolute paths, merged across settings layers) for CLAUDE.md files to skip when memory loads, e.g. another team's tree in a monorepo or vendored checkouts; skillListingBudgetFraction — share of the context window reserved for the per-turn skill listing (default 0.01 = 1%; descriptions are truncated to fit, so raise it when your skill descriptions get cut short, lower it to fit more skills); skillListingMaxDescChars — per-skill cap on the description + when-to-use text in that listing; disableBundledSkills (CC 2.1.169+) — removes every bundled skill and workflow except /doctor (`/code-review`, `/loop`, `/claude-api`, `/dataviz`, `/deep-research`, `/batch`, `/debug` …) and hides built-in slash commands from the model; project, user and plugin skills are unaffected. Weigh it against losing `/code-review` — the built-in multi-agent review — before turning it on.", + "// claudeMdExcludes": ["**/vendor/**/CLAUDE.md", "**/third_party/**/CLAUDE.md"], + "// skillListingBudgetFraction": 0.02, + "// skillListingMaxDescChars": 2048, + "// disableBundledSkills": true, + + "// fallbackModel": ["opus", "sonnet"], + "//fallbackModel-notes": "fallbackModel (CC 2.1.166+) — ordered chain (up to three entries; model names or aliases; 'default' expands to the account default) tried when the primary model is overloaded, unavailable, or returns a non-retryable server error. Position carries meaning, so Claude Code takes the whole array from the highest-precedence file that sets it. The --fallback-model CLI flag overrides it.", + "// model": "sonnet", - "//model-notes": "Per-machine model override. The scaffolded project default is 'fable' (Claude Fable 5 — requires CC 2.1.170+; premium pricing; safety-classifier fallback in cybersecurity/biology domains; unavailable under zero-data-retention). Uncomment to drop THIS machine back to 'sonnet' when its CC predates 2.1.170, Fable 5 isn't in your org's plan, or the cost posture doesn't fit. 'best' is the middle path: resolves to Fable 5 where available, latest Opus otherwise." + "//model-notes": "Per-machine model override. The scaffolded project default is 'fable' (Claude Fable 5 — requires CC 2.1.170+; premium pricing; safety-classifier fallback in cybersecurity/biology domains; unavailable under zero-data-retention). Uncomment to drop THIS machine back to 'sonnet' (Claude Sonnet 5 on CC 2.1.197+, Claude Code's own default model) when its CC predates 2.1.170, Fable 5 isn't in your org's plan, or the cost posture doesn't fit. 'opus' resolves to Claude Opus 5 on CC 2.1.219+. 'best' is the middle path: resolves to Fable 5 where available, latest Opus otherwise." } diff --git a/templates/git-workflow/settings-patch.json b/templates/git-workflow/settings-patch.json index 734abb3..3022e7f 100644 --- a/templates/git-workflow/settings-patch.json +++ b/templates/git-workflow/settings-patch.json @@ -9,6 +9,7 @@ "hooks": [ { "type": "command", + "shell": "bash", "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/format-on-write.sh", "timeout": 30 } @@ -20,6 +21,7 @@ "hooks": [ { "type": "command", + "shell": "bash", "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/stop-run-checks.sh", "timeout": 120 } diff --git a/templates/multi-agent/dot-claude/rules/multi-agent-guardrails.md b/templates/multi-agent/dot-claude/rules/multi-agent-guardrails.md index 8725218..d56f78b 100644 --- a/templates/multi-agent/dot-claude/rules/multi-agent-guardrails.md +++ b/templates/multi-agent/dot-claude/rules/multi-agent-guardrails.md @@ -34,6 +34,13 @@ Ask all five; a "no" on any of them means stop and rethink. - [ ] **Cleanup plan?** Worktrees and merged branches get removed after a successful integration. - [ ] **Escape hatch?** A clear path to abort and reset if one agent goes sideways — without losing the work of the others. +## Runtime caps to design around (Claude Code ≥ 2.1.217) + +- At most 20 subagents run concurrently per session (`CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS`); the 21st spawn fails and the error tells Claude not to retry — size waves accordingly. +- Nesting is capped at 3 levels below the main thread by default (`CLAUDE_CODE_MAX_SUBAGENT_SPAWN_DEPTH`, since 2.1.219); at the cap the Agent tool is withheld, so a worker that "needs" helpers is a design smell, not a config problem. +- Subagents run in the background by default with a narrower tool set; a fork (`subagent_type: "fork"`) inherits the whole conversation — the opposite of isolation. +- For fan-outs beyond a handful of workers, Claude Code's dynamic Workflows (`ultracode`, or "use a workflow") give you a scripted, resumable pipeline; `/infinite` stays the right tool for N variants of one spec. + ## Rule of thumb **Solo Claude is the default.** Reach for multi-agent when the work is truly parallelizable: N variants of the same spec, N independent audits, N isolated files to process. Everything else, do sequential. diff --git a/templates/safety/settings-patch.json b/templates/safety/settings-patch.json index 86444e1..1e1e869 100644 --- a/templates/safety/settings-patch.json +++ b/templates/safety/settings-patch.json @@ -1,13 +1,7 @@ { "//": "Merged into .claude/settings.json. PreToolUse hooks guard dangerous Bash and scan Write/Edit for secrets. permissions.disableBypassPermissionsMode='disable' hard-disables --dangerously-skip-permissions for this project, even if the user passes the flag.", - "//2": "autoMode.hard_deny (CC 2.1.136+, schemastore-validated 2026-05-23) is on by default with two reasonable category blocks. The classifier only fires under `claude --auto-mode`, so this is a no-op for standard manual sessions — pure upside for auto-mode users. Tune the strings or remove entries to match your project's threat model.", - "//3": "Optional opt-ins (sandbox.network.deniedDomains, sandbox.failIfUnavailable, env CLAUDE_CODE_STOP_HOOK_BLOCK_CAP) are documented in templates/core/dot-claude/settings.local.json.example with copy-paste-ready stubs. Hook authoring note (CC 2.1.139+): hookCommand entries accept 'args: string[]' for direct-exec form and 'continueOnBlock: boolean' for prompt hooks.", - "autoMode": { - "hard_deny": [ - "Running executable files", - "Writing to system directories" - ] - }, + "//2": "autoMode is deliberately NOT shipped here: since CC 2.1.207 the auto-mode classifier reads `autoMode` only from ~/.claude/settings.json or managed settings — never from .claude/settings.json or .claude/settings.local.json (code.claude.com/docs/en/auto-mode-config: both live in the repo, so a checked-in file could inject its own allow rules). The pre-v2.8 default (`hard_deny: [\"Running executable files\", \"Writing to system directories\"]`) is documented as a user-scope opt-in in docs/05-safety-permissions.md; a retrofit removes that exact shipped block from project settings and [ SETTINGS WARNINGS ] flags any other project-scope autoMode.", + "//3": "Optional opt-ins (sandbox.network.deniedDomains, sandbox.failIfUnavailable, sandbox.credentials, env CLAUDE_CODE_STOP_HOOK_BLOCK_CAP) are documented in templates/core/dot-claude/settings.local.json.example with copy-paste-ready stubs. Hook authoring notes: hookCommand entries accept 'args: string[]' for direct-exec form and 'continueOnBlock: boolean' for prompt hooks (CC 2.1.139+); every shipped command hook declares 'shell: \"bash\"' (CC 2.1.81+, schema-validated) so Windows sessions resolve Git Bash directly instead of falling back to PowerShell and failing on the .sh entrypoint — older CC ignores the key.", "hooks": { "PreToolUse": [ { @@ -15,11 +9,13 @@ "hooks": [ { "type": "command", + "shell": "bash", "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/block-dangerous-bash.sh", "timeout": 10 }, { "type": "command", + "shell": "bash", "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/check-package-availability.sh", "timeout": 10 } @@ -30,6 +26,7 @@ "hooks": [ { "type": "command", + "shell": "bash", "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/scan-secrets.sh", "timeout": 10 } diff --git a/templates/safety/settings-patch.slop-scan.json b/templates/safety/settings-patch.slop-scan.json index c288dac..9244b29 100644 --- a/templates/safety/settings-patch.slop-scan.json +++ b/templates/safety/settings-patch.slop-scan.json @@ -8,6 +8,7 @@ "hooks": [ { "type": "command", + "shell": "bash", "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/slop-scan.sh", "timeout": 10 } diff --git a/templates/token-efficiency/settings-patch.tier-pro.json b/templates/token-efficiency/settings-patch.tier-pro.json index 6b71677..daa0d2e 100644 --- a/templates/token-efficiency/settings-patch.tier-pro.json +++ b/templates/token-efficiency/settings-patch.tier-pro.json @@ -10,6 +10,7 @@ "hooks": [ { "type": "command", + "shell": "bash", "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/truncate-bash-output.sh", "timeout": 10 } diff --git a/test/retrofit-hooks/test-retired-defaults-migration.sh b/test/retrofit-hooks/test-retired-defaults-migration.sh new file mode 100644 index 0000000..e273674 --- /dev/null +++ b/test/retrofit-hooks/test-retired-defaults-migration.sh @@ -0,0 +1,72 @@ +#!/usr/bin/env bash +# v2.8.0 retired two configurator defaults that current Claude Code ignores: +# - a project-scope `autoMode` block (CC >= 2.1.207 reads autoMode only from +# ~/.claude/settings.json / managed settings) +# - `Write(.env)` / `Write(.env.*)` deny rules (Edit(path) already covers the +# Write tool; CC >= 2.1.210 warns at startup about Write(path) rules) +# and started declaring `shell: "bash"` on every shipped command hook. +# A retrofit must (1) remove exactly the shipped autoMode block and the two +# Write rules, (2) backfill `shell` onto configurator-owned hook entries that +# predate it, and (3) leave a user-edited autoMode block alone. +set -euo pipefail + +tmp=$(mktemp -d) +trap "rm -rf $tmp" EXIT + +python3 configure.py --persona solo-experienced --yes --dir "$tmp" >/dev/null + +# Simulate a pre-v2.8.0 install: shipped autoMode block, Write(.env*) deny +# rules, and hook entries without the `shell` key. +python3 - "$tmp/.claude/settings.json" <<'PY' +import json, sys +p = sys.argv[1] +d = json.load(open(p)) +d["autoMode"] = {"hard_deny": ["Running executable files", "Writing to system directories"]} +deny = d["permissions"]["deny"] +deny[:0] = ["Write(.env)", "Write(.env.*)"] +for groups in d["hooks"].values(): + for g in groups: + for h in g.get("hooks", []): + h.pop("shell", None) +json.dump(d, open(p, "w"), indent=2) +PY + +out=$(python3 configure.py --persona solo-experienced --yes --dir "$tmp") + +python3 - "$tmp/.claude/settings.json" <<'PY' +import json, sys +d = json.load(open(sys.argv[1])) +assert "autoMode" not in d, f"shipped autoMode block survived retrofit: {d.get('autoMode')}" +deny = d["permissions"]["deny"] +assert "Write(.env)" not in deny and "Write(.env.*)" not in deny, f"Write(.env*) rules survived: {deny}" +assert "Edit(.env)" in deny and "Edit(.env.*)" in deny, f"Edit(.env*) rules missing: {deny}" +missing = [h["command"] for groups in d["hooks"].values() for g in groups + for h in g.get("hooks", []) if h.get("type") == "command" and "shell" not in h] +assert not missing, f"shell not backfilled on: {missing}" +print(" migrated: autoMode removed, Write(.env*) removed, shell backfilled on all hooks") +PY + +printf '%s\n' "$out" | grep -q "retired configurator default" \ + || { echo "FAIL: [ MERGED ] summary does not mention the retired defaults"; printf '%s\n' "$out" | grep -i merged; exit 1; } + +# --- A user-edited autoMode block must survive (and be flagged, not deleted). +tmp2=$(mktemp -d) +trap "rm -rf $tmp $tmp2" EXIT +python3 configure.py --persona solo-experienced --yes --dir "$tmp2" >/dev/null +python3 - "$tmp2/.claude/settings.json" <<'PY' +import json, sys +p = sys.argv[1] +d = json.load(open(p)) +d["autoMode"] = {"hard_deny": ["Running executable files", "Never run terraform apply"]} +json.dump(d, open(p, "w"), indent=2) +PY +python3 configure.py --persona solo-experienced --yes --dir "$tmp2" >/dev/null +python3 - "$tmp2/.claude/settings.json" <<'PY' +import json, sys +d = json.load(open(sys.argv[1])) +assert d.get("autoMode") == {"hard_deny": ["Running executable files", "Never run terraform apply"]}, \ + f"user-edited autoMode was altered: {d.get('autoMode')}" +print(" preserved: user-edited autoMode block untouched") +PY + +echo "PASS: retired defaults migrate out on retrofit; user-edited autoMode preserved" diff --git a/test/schema-hygiene/test-preflight-detects-violations.sh b/test/schema-hygiene/test-preflight-detects-violations.sh index 2ed92a1..7162911 100755 --- a/test/schema-hygiene/test-preflight-detects-violations.sh +++ b/test/schema-hygiene/test-preflight-detects-violations.sh @@ -1,6 +1,7 @@ #!/usr/bin/env bash -# Verify check_settings_validates() actually catches the two regression -# classes: top-level //-prefixed keys and a non-object skillOverrides value. +# Verify check_settings_validates() actually catches its regression classes: +# top-level //-prefixed keys, a non-object skillOverrides value, project-scope +# autoMode (ignored since CC 2.1.207) and never-consulted Write()/Glob() rules. # Constructs synthetic settings dicts and asserts the check fires. set -euo pipefail @@ -29,10 +30,22 @@ bad4 = {"skillOverrides": {"my-skill": "unknown-value"}} w = check_settings_validates(bad4) assert any("skillOverrides" in msg for msg in w), f"expected skillOverrides-value warning, got: {w}" +# Case 6: autoMode in project settings (ignored by CC >= 2.1.207) +bad6 = {"autoMode": {"hard_deny": ["Running executable files"]}} +w = check_settings_validates(bad6) +assert any("autoMode" in msg and "~/.claude/settings.json" in msg for msg in w), f"expected autoMode-scope warning, got: {w}" + +# Case 7: never-consulted path rules (Write/NotebookEdit/Glob/MultiEdit) +bad7 = {"permissions": {"deny": ["Write(.env)", "Edit(.env)"], "allow": ["Glob(src/**)", "Read"]}} +w = check_settings_validates(bad7) +assert any("Write(.env)" in msg and "Glob(src/**)" in msg for msg in w), f"expected path-rule warning, got: {w}" +assert not any("Edit(.env)" in msg for msg in w), f"Edit() rule must not be flagged: {w}" + # Case 5: clean settings — no warnings -clean = {"$schema": "x", "skillOverrides": {"my-skill": "name-only"}, "env": {"FOO": "bar"}} +clean = {"$schema": "x", "skillOverrides": {"my-skill": "name-only"}, "env": {"FOO": "bar"}, + "permissions": {"deny": ["Edit(.env)", "Read(secrets/**)", "Write"]}} w = check_settings_validates(clean) assert w == [], f"expected clean output, got: {w}" -print("PASS: check_settings_validates catches all 4 violation classes + clean case") +print("PASS: check_settings_validates catches all 6 violation classes + clean case") EOF