From 7a0f15eca20c34cea7810585bebdbfb3d024ca7e Mon Sep 17 00:00:00 2001 From: Siddharth Kapoor Date: Mon, 31 Aug 2026 05:27:40 -0400 Subject: [PATCH] refactor(spec): lean phase files and task-skill check discovery MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit **Lean phase-file template** - Drop Requirements, Task Runner Commands, Design, and References boilerplate from `implementation/phase-N.md`; keep the lite Contract block and the doc-aware `## Doc Scope` section - Move process rules (brevity, build-system preservation, terse prose) into the executor agent definition so phase files stay task documents - Tighten the executor prompt: terse prose rule plus no direct build/test tool calls — discover recipes with the task skill instead - Restructure phase files to Introduction / Implementation / Validation / Deviations sections, with a validation line per task and a closing full-suite Validation + Deviations pair **Task-skill check wiring** - Replace recorded task-runner commands across planner, executor, updater, and agent-prompt wiring with `/task detect` + `/task list` discovery at execution time — never raw npm/pytest/go-test invocations - Update planning (new.md, guide.md, update.md), review, and validation (validation-loop.md, validation-prompt.md) checks to validate the lean shape and task-skill references instead of runner-command metadata - Keep runner commands out of bash fences in guide.md and out of validation prompt code blocks - Skip the code-completeness gate for lite plans (no `**Code:**` blocks) in new.md, review.md, and validation-loop.md --- skills/spec/agents/spec-executor.md | 8 ++- skills/spec/agents/spec-planner.md | 6 +- skills/spec/agents/spec-updater.md | 2 +- skills/spec/references/agent-prompt.md | 4 +- .../references/implementation-template.md | 72 ++++++------------- skills/spec/references/validation-loop.md | 8 ++- skills/spec/references/validation-prompt.md | 15 ++-- skills/spec/references/workflows/guide.md | 6 +- skills/spec/references/workflows/new.md | 8 ++- skills/spec/references/workflows/review.md | 6 +- skills/spec/references/workflows/update.md | 2 +- 11 files changed, 56 insertions(+), 81 deletions(-) diff --git a/skills/spec/agents/spec-executor.md b/skills/spec/agents/spec-executor.md index 96253ab..e79845d 100644 --- a/skills/spec/agents/spec-executor.md +++ b/skills/spec/agents/spec-executor.md @@ -55,12 +55,16 @@ You are precise, minimal, and disciplined. You follow implementation specs exact - If a task requires an unplanned build system change, STOP and flag it **Hygiene after every task (non-negotiable):** -- Run format → lint → typecheck → tests using the project's task runner commands from plan metadata +- Run format → lint → typecheck → tests via the task skill — phase files carry no runner commands - Fix all failures before marking the task complete - Never leave a task in a state where any of these fail -- Never invent shell commands — use only task runner recipes discovered from the project's `mise.toml`, `justfile`, `Makefile`, or `package.json` scripts +- Never invent shell commands and never call build/test tools directly — discover recipes with the task skill and run those - At phase start, call `/task detect` once to identify the runner and `/task list` to enumerate recipes; reuse those names for every hygiene/build/test command +**Terse prose:** +- Everything you write — execution-log entries, reports, commit messages, markdown — is brief and terse +- One line where one line suffices; no preamble, no restating the spec, no summaries of what the checklist already tracks + **Progress tracking (non-negotiable order):** 1. Complete the task implementation 2. Run hygiene checks (format → lint → typecheck → tests) diff --git a/skills/spec/agents/spec-planner.md b/skills/spec/agents/spec-planner.md index f1979e5..36d88d0 100644 --- a/skills/spec/agents/spec-planner.md +++ b/skills/spec/agents/spec-planner.md @@ -49,7 +49,6 @@ You are thorough and opinionated. You write plans that are detailed enough to be - High-level objective (2-4 bullets) - Design overview (key decisions, not implementation detail) - Phase/task checklist (one-liner per task) -- Task runner metadata **Implementation files** — detailed, and **always show the complete code**: - Step-by-step instructions per task @@ -57,8 +56,7 @@ You are thorough and opinionated. You write plans that are detailed enough to be - **Code-first gate (non-negotiable):** Before you write ANY task into an implementation file, you must already have the exact code that task will produce. For every task, ask: *"Can I paste the literal code — every new/changed line — right now?"* If **no**, do NOT write the task yet: resolve the unknown first (read the codebase with Glob/Grep/Read, search the web, or ask the user once), then write it. Never emit a task whose code you would leave for the execution agent to figure out. - **Show the whole thing, not a sketch — and for edits, a diff, not the whole file:** For a new file, include its entire contents. For an edit to an existing file, show a **diff** — the exact old→new lines or a unified diff with context lines above and below the change — never the whole file. A whole-file replacement is reserved for tasks where the user explicitly asked for a full-file replacement; a whole-file dump where a diff is required is a plan defect. Pseudocode, ellipses (`...`), "e.g.", and "something like" are forbidden inside a task's code block — they are the exact failure this rule exists to prevent. - **Self-audit before finishing each phase file (mandatory):** re-read every task and confirm each has a non-empty, complete code block. Reject and rewrite any task whose code block: is missing or empty; contains a placeholder/stub marker from the blocklist in `references/code-completeness-blocklist.md` (the canonical list — judge by intent, not blind substring matching); shows only a signature or a comment where a body belongs; or describes the code in prose instead of showing it. Every line the execution agent will write must appear verbatim in the block. A task that fails this audit is incomplete — resolve the unknown now (read the codebase, search the web, or ask the user once) and paste the real code, or delete the task. Do not finish the phase file until every task passes. -- Task runner commands for validation after every task (format → lint → typecheck → test) discovered by reading `mise.toml`, `justfile`, `Makefile`, or `package.json` — never invent shell commands -- Every build/test/lint command MUST come from the project's task runner +- A validation line per task: "project checks green via the task skill (`/task detect`, `/task list`)" — no recorded runner commands, no raw npm/pytest/go-test invocations **User guide** — usage-focused: - What was built and how to use it @@ -71,7 +69,7 @@ Every plan you create must include these as explicit constraints in implementati 1. **Minimal changes**: Execution agent makes the smallest change that achieves the goal. No drive-by fixes. 2. **Build system preservation**: Do not modify the build system unless the plan explicitly requires it. The project must build after every task. -3. **Hygiene**: Run format → lint → typecheck → tests after every task using the project's task runners. Fix failures before moving on. +3. **Hygiene**: Executors run format → lint → typecheck → tests after every task via the task skill. Fix failures before moving on. 4. **Validation phase**: Every plan must end with a phase that confirms the full suite passes and the user guide is complete. ## Doc-Aware Planning (active only when `PERSISTENT_MODE` is set) diff --git a/skills/spec/agents/spec-updater.md b/skills/spec/agents/spec-updater.md index d52720b..58624b6 100644 --- a/skills/spec/agents/spec-updater.md +++ b/skills/spec/agents/spec-updater.md @@ -42,7 +42,7 @@ plan.md implementation/phase-N.md ─────────────────────────────── ────────────────────────────────────── High-level task checklist Step-by-step execution detail One line per task Full context, code examples, commands -Phase headers with ✅ markers Task runner commands and validation steps +Phase headers with ✅ markers Task-skill validation steps and deviations ``` **These two views must always agree.** When you modify one, ask: does the other need to change too? diff --git a/skills/spec/references/agent-prompt.md b/skills/spec/references/agent-prompt.md index e8024a1..d0e5c15 100644 --- a/skills/spec/references/agent-prompt.md +++ b/skills/spec/references/agent-prompt.md @@ -12,8 +12,8 @@ Branch: {PLAN_BRANCH} Worktree: {PLAN_WORKTREE or "(none)"} Git commits allowed: {ALLOW_COMMITS} -Task runner commands: -{Discover by reading mise.toml/justfile/Makefile/package.json directly} +Project checks: +Discover and run every build/test/lint/format command via the task skill (`/task detect`, then `/task list`) — the phase file carries no runner commands. Never call build/test tools directly and never invent shell equivalents. ## Completed phases (summary only) {One line per completed phase: "Phase 1 (Setup) ✅ — 4/4 tasks, tests passing"} diff --git a/skills/spec/references/implementation-template.md b/skills/spec/references/implementation-template.md index 99600b4..d002f67 100644 --- a/skills/spec/references/implementation-template.md +++ b/skills/spec/references/implementation-template.md @@ -1,36 +1,26 @@ # Implementation Phase File Template -Use this structure for each `implementation/phase-N.md` file created in Step 5.5. +Use this structure for each `implementation/phase-N.md` file. Phase files are task documents: a one-line introduction, the tasks with their complete code, then validation and deviations. Process rules (brevity, build-system preservation, markdown, terse prose) live in the executor agent definition — never repeat them here. Design context lives in plan.md and the user guide. Executors discover build/test/lint/format commands via the task skill, so no runner commands are recorded. ```markdown # Phase {N} - {Phase Name} ## Introduction -{Brief description of what this phase accomplishes and its role in the overall plan.} +{One or two sentences — what this phase accomplishes.} -## Requirements +## Doc Scope +{ONLY for doc-aware (`--persistent`) plans — omit this section otherwise. Per `references/doc-aware.md` Rule 3: the `**Write globs:**` list must contain at least one glob; cross-module interaction uses documented public interfaces only (Rule 4); each crossing gets a boundary callout (Rules 6–7). Verify every write target with `$SPEC_SKILL/scripts/scope.py`.} -**Brevity:** Make the smallest change that achieves the task. No drive-by refactors or unrelated fixes. +**Write globs:** {`glob1`} {`glob2`} -**Build system preservation:** Do NOT modify the build system, CI config, or dependencies unless this phase is explicitly about them. If the project built before you started, it must build after every task. If a change would require an unplanned build system modification, stop and flag it. +**Public interfaces only:** {cross-module interaction uses only the target module's documented public API — never its internals.} -**Markdown output:** Soft-wrap prose — never hard-wrap. Write each paragraph as one continuous line; do not insert manual newlines to wrap prose at a fixed column width. Newlines still separate paragraphs, list items, headings, and code fences. - -### Task Runner Commands -{List the relevant task runner commands for this phase. ALWAYS use these — never invent equivalent shell commands. Discover them by reading `mise.toml`, `justfile`, `Makefile`, or `package.json` scripts directly.} -- Build: `{e.g. just build | make build | task build | mise run build}` -- Test: `{e.g. just test | make test | task test | mise run test}` -- Lint: `{e.g. just lint | make lint | mise run lint}` -- Format: `{e.g. just fmt | make format | task fmt | mise run fmt}` - -If no task runner covers a needed operation, note: "Gap: no recipe for X — suggest adding one." - -## Design -{Describe the approach taken in this phase — what pattern, what structure, what architectural decision was made before coding begins.} +**Boundary callouts:** +- {task} — {what crosses the boundary, why it is required, and the alternative considered} ## Implementation -> **Gate (machine-checked in validation):** Every task below MUST contain a `**Code:**` block holding the **complete, literal code** it will produce — full contents for new files; for edits to existing files, a **diff** (exact old→new lines or a unified diff) with context lines above and below the change, never a whole-file dump. A whole-file replacement appears only when the task explicitly states that the user asked for a full-file replacement. The block is REJECTED if it is missing or empty, contains a placeholder/stub marker from the blocklist (see `references/code-completeness-blocklist.md` — the canonical list), shows a bare signature/comment where a body belongs, describes the code in prose instead of showing it, or replaces an entire existing file with a whole-file dump where a diff was required. The blocklist is judged by intent, not blind substring matching, so a marker used as a legitimate token (not a stand-in for missing code) does not fail the block. If you cannot show the complete code, resolve the unknown now during planning (read the codebase, search the web, or ask the user) — never pass research, open design choices, or code authoring to the execution agent. A dedicated validation agent scans for exactly these placeholders and will fail the plan until every code block is complete. +> **Gate (machine-checked in validation):** Every task below MUST contain a `**Code:**` block holding the **complete, literal code** it will produce — full contents for new files; for edits to existing files, a **diff** (exact old→new lines or a unified diff) with context lines above and below the change, never a whole-file dump. A whole-file replacement appears only when the task explicitly states that the user asked for a full-file replacement. The block is REJECTED if it is missing or empty, contains a placeholder/stub marker from the blocklist (see `references/code-completeness-blocklist.md` — the canonical list), shows a bare signature/comment where a body belongs, describes the code in prose instead of showing it, or replaces an entire existing file with a whole-file dump where a diff was required. The blocklist is judged by intent, not blind substring matching, so a marker used as a legitimate token (not a stand-in for missing code) does not fail the block. If you cannot show the complete code, resolve the unknown now during planning (read the codebase, search the web, or ask the user) — never pass research, open design choices, or code authoring to the execution agent. The executor may still apply minor mechanical fixes to the specified code in-flight — obvious typos, missing imports, or trivial type corrections that do not change a module's contract or violate an invariant — per the permitted-minor-deviations policy in `agents/spec-executor.md`; these are logged as `[MINOR-DEVIATION]` entries, not treated as failures. {For each task in this phase:} @@ -38,18 +28,14 @@ The executor may still apply minor mechanical fixes to the specified code in-fli ### Task {X}: {Task Description} **Steps:** -1. {Detailed step-by-step instructions} -2. {Include exact commands, file paths, code patterns} +1. {Concise step-by-step instructions} **Code (required — complete, never omit, never abbreviate):** ```{lang} {The COMPLETE code this task produces. New file: its entire contents. Edit to an existing file: a diff — the exact old→new lines or a unified diff with the changed lines plus context -lines above and below, so the executor sees what changes and where. Never paste the whole -file for an edit; a whole-file replacement is shown only when this task explicitly declares -a user-requested full-file replacement. Every line the execution agent will write appears -here verbatim. No ellipses, no pseudocode, no "e.g." A task with a partial, prose-only, or -whole-file-dump code block is incomplete and must not be emitted.} +lines above and below. Never paste the whole file for an edit; a whole-file replacement is +shown only when this task explicitly declares a user-requested full-file replacement.} ``` **Contract (lite mode only — replaces `**Code:**`):** @@ -65,37 +51,23 @@ Acceptance: - ``` -**Files to modify / create:** -- `path/to/file.ext` — {specific changes} - **Table rows (only when the task consumes a `tables/*.md` table):** - `tables/{set-slug}.md` — rows {#–#}: mark each row's Status `[x]` as it is completed; a row is done only when its change is applied and validated -**Validation (run after every task):** -- [ ] `{fmt command}` — no formatting changes outstanding -- [ ] `{lint command}` — zero warnings/errors -- [ ] `{test command}` — all tests pass -- [ ] *(only under `spec go --commit`, and only when CI exists)* CI is green for this committed phase — the executor invokes `/git commit --yes --fix` and the git skill blocks in its bounded fix-until-green loop. This is a HARD gate: the phase is not complete while CI is red, and if the git skill stops with CI still failing, the phase is reported FAILED. Skip silently when any of: no `--commit`, no remote, no configured CI, no `gh`/`glab` CLI. A phase never fails for the absence of CI — only for a CI that exists and stays red. - ---- +**Files to modify / create:** +- `path/to/file.ext` — {specific changes} -## Future Work / Validation +**Validation (run after every task):** +- [ ] Project checks green via the task skill (`/task detect`, `/task list` — then the repo's test/lint/typecheck recipes); fix failures before marking the task complete +- [ ] *(only under `spec go --commit`, and only when CI exists)* CI is green for this committed phase — the executor invokes `/git commit --yes --fix` and the git skill blocks in its bounded fix-until-green loop. Skip silently when any of: no `--commit`, no remote, no configured CI, no `gh`/`glab` CLI. -After all tasks complete, run the full suite: +## Validation -```bash -{fmt command} -{lint command} -{typecheck command} -{test command} -{build command} -``` +After all tasks complete, run the repo's full check set via the task skill and record the outcome: -Expected: {describe expected output} +- [ ] Full suite green (tests + validate/lint as the repo defines them) — note the exact recipes run -Note any follow-on tasks or deferred items surfaced during this phase: -- {deferred item} +## Deviations -## References -- {Related proposals, ADRs, or external references consulted for this phase} +- {Deviations, minor deviations, or deferred items surfaced during this phase — or "none"} ``` diff --git a/skills/spec/references/validation-loop.md b/skills/spec/references/validation-loop.md index 652fe67..71e5e35 100644 --- a/skills/spec/references/validation-loop.md +++ b/skills/spec/references/validation-loop.md @@ -23,7 +23,7 @@ prompt: [contents of references/validation-prompt.md with SCOPE=plan-level] ``` Validate: -- plan.md metadata (Task Runners field present, branch/worktree filled) +- plan.md metadata (no runner-command metadata, branch/worktree filled) - user-guide.md exists and has non-TODO overview - Phase names and task counts consistent across plan.md and implementation files - Inter-phase dependencies identified @@ -38,11 +38,13 @@ prompt: [contents of references/validation-prompt.md with SCOPE=phase, PHASE_N={ Validate only `implementation/phase-{N}.md` against the plan.md tasks for that phase: - Task specificity and actionability - Implementation completeness (file paths, code examples, no ambiguity) -- Task runner commands listed and used (not raw npm/pytest/go test) -- fmt/lint/typecheck/test validation block present +- Lean shape: only Introduction / Implementation / Validation / Deviations sections, plus `## Doc Scope` for doc-aware plans (no Requirements process block, no Task Runner Commands, no Design) +- Validation steps reference the task skill (no recorded runner commands, no raw npm/pytest/go test invocations) - Test coverage and success criteria - user-guide.md update instructions per task +**Lite plans skip this agent.** If the plan's metadata carries `Lite: true`, there are no `**Code:**` blocks to check — do not launch the code-completeness agent and treat `VALIDATION_CODE_COMPLETENESS_STATUS` as PASS. + **Code-completeness agent** — launch one agent (`subagent_type: general-purpose`, `model-tier: light`, `run_in_background: true`): ``` diff --git a/skills/spec/references/validation-prompt.md b/skills/spec/references/validation-prompt.md index 7ebd4e1..6d94584 100644 --- a/skills/spec/references/validation-prompt.md +++ b/skills/spec/references/validation-prompt.md @@ -19,7 +19,7 @@ Read these files: Validate the following quality criteria: **Metadata** -- Does plan.md have a "Task Runners" field with actual commands? +- Does plan.md carry no runner-command metadata? (Checks are discovered at execution time via the task skill — a "Task Runners" field is stale boilerplate.) - Are branch/worktree fields filled (or "(none)" explicitly)? **User Guide** @@ -80,12 +80,11 @@ Validate the following quality criteria for Phase {N} only: - Are code examples present for non-trivial logic? - Is there enough detail for an autonomous agent to execute without asking clarifying questions? -**Task Runner Usage** -- Does phase-{N}.md list applicable task runner commands in a "Task Runner Commands" section covering build/test/lint/format/typecheck? -- Are all build/test/lint/format/typecheck/run commands using the project's task runners (not raw `npm test`, `python -m pytest`, `go test ./...` when a task runner wraps them)? -- Does every task's validation checklist include lint, format, and typecheck steps — not just tests? -- Is there a "Phase Validation" block at the end with all five checks (fmt, lint, typecheck, test, build)? -- Is there a note that lint/format/typecheck must run after every task, not only at phase end? +**Lean shape and check discovery** +- Does phase-{N}.md contain ONLY the sections Introduction, Implementation, Validation, and Deviations — plus a `## Doc Scope` section when the plan is doc-aware (`--persistent`)? A "## Requirements" process block, a "Task Runner Commands" section, or a "## Design" section is a failure — those rules live in the executor agent definition and the design lives in plan.md. +- Does every task's validation checklist run the project's checks via the task skill (references `/task detect` / `/task list` or equivalent) — not recorded runner commands, and never raw `npm test` / `python -m pytest` / `go test` style invocations? +- Does the phase close with a Validation section (full-suite run via the task skill) and a Deviations section? +- Is the Introduction one or two sentences, with no policy prose? **User Guide** - Does each task in phase-{N}.md specify what to update in user-guide.md once complete? @@ -95,7 +94,7 @@ Validate the following quality criteria for Phase {N} only: - Are task descriptions consistent between plan.md and phase-{N}.md? **Test Coverage** -- Does each task specify what tests to write or run using the task runner? +- Does each task specify what tests to write or run via the task skill (never raw `npm test` / `python -m pytest` / `go test` style invocations)? - Are acceptance/success criteria testable? Respond ONLY in this exact format: diff --git a/skills/spec/references/workflows/guide.md b/skills/spec/references/workflows/guide.md index c19cae1..b3d9c7e 100644 --- a/skills/spec/references/workflows/guide.md +++ b/skills/spec/references/workflows/guide.md @@ -25,7 +25,6 @@ Read `.codevoyant/spec/{plan-name}/plan.md` in full. Parse: - All phases (`### Phase N - Name`) - All tasks per phase (`N. [ ] task` and `N. [x] task`) -- Task runner commands from the Metadata section - Any `## Insights` section from previous sessions Determine starting position: @@ -199,10 +198,7 @@ Then continue the guide loop with the next task. When all tasks in a phase are either done `[x]` or skipped, run the phase boundary: -1. Run tests if task runner commands are available: - ```bash - {test command from plan metadata} - ``` +1. Run the project's checks via the task skill: `/task detect` then `/task list` — run the repo's test recipe. If tests fail, report the failure and use **AskUserQuestion** (Fix before continuing / Continue anyway). 2. If all tasks were done (none skipped), mark the phase header `✅` in plan.md. diff --git a/skills/spec/references/workflows/new.md b/skills/spec/references/workflows/new.md index 2a5307f..6912880 100644 --- a/skills/spec/references/workflows/new.md +++ b/skills/spec/references/workflows/new.md @@ -12,7 +12,7 @@ This workflow's **only output** is plan files. If you are about to do anything e |---|---| | Edit a source file | Stop. Add a task to the plan instead. | | Write application code | Stop. Describe it in `implementation/phase-N.md`. | -| Run build / test / lint | Stop. Record the command in plan metadata. | +| Run build / test / lint | Stop. Executors discover checks at run time via the task skill — record nothing. | | Fix a bug you noticed | Stop. Add it as a task in the appropriate phase. | | Write a task that says "research / investigate / explore / decide / figure out X" | Stop. Resolve it **now**, during planning — read the codebase (Glob/Grep/Read) and use WebSearch/WebFetch — then write the concrete answer and code. The written plan must be delta-free; the execution agent never researches or makes open design decisions. | | Keep going after "looks good" | Stop. Your job is done. Tell the user to run `/spec go`. | @@ -345,9 +345,9 @@ Use `references/implementation-template.md`. Move ALL detailed specs here: - **The complete code for every task** — full file contents for new files; for edits, a diff (exact old→new lines or a unified diff) with context lines above and below the change, never a whole-file dump. A whole-file replacement appears only when the user explicitly asked for a full-file replacement. Not "code for non-trivial logic": all of it. No ellipses, pseudocode, or prose-only descriptions. If you can't show the code, resolve the unknown during planning rather than deferring it. - Testing and validation steps — including the `spec go --commit` CI-green check per phase. Under `--commit` with a remote, CI, and matching CI CLI, this is a HARD per-phase gate (the executor's `/git commit --yes --fix` blocks until green or stops); it must never gate a non-`--commit` run or a repo with no CI. -**Task runner constraint (CRITICAL):** Every build, test, lint, and run command MUST use the project's task runner (mise/just/Makefile/package.json scripts). Before recording any such command, call `/task detect` to identify the runner and `/task list` to see available tasks — use those names verbatim. Never invent custom shell commands when a task runner recipe exists. +**Check discovery (CRITICAL):** Phase files carry NO task runner commands — executors discover build/test/lint/format at run time via the task skill. Write each task's validation step as "project checks green via the task skill (`/task detect`, `/task list`)", never as recorded commands, and never as raw `npm test` / `pytest` style invocations. Process rules (brevity, build-system preservation, markdown, terse prose) are never written into phase files — they live in the executor agent definition. -**Doc-aware scoping (only when `PERSISTENT_MODE=true`):** every phase-N.md MUST begin its Design section with a `## Doc Scope` block (per `references/doc-aware.md` Rules 3, 4, 6). The `**Write globs:**` list MUST contain at least one glob — a phase with no write globs is a planning error (per `references/doc-aware.md` Rule 3's no-globless-phases rule); when docs are incomplete, derive the phase globs from the repository layout (src/monorepo libs, CI, docs) instead of leaving the list empty: +**Doc-aware scoping (only when `PERSISTENT_MODE=true`):** every phase-N.md MUST carry a `## Doc Scope` section directly after the Introduction (per `references/doc-aware.md` Rules 3, 4, 6) — the lean template provides this section for doc-aware plans only. The `**Write globs:**` list MUST contain at least one glob — a phase with no write globs is a planning error (per `references/doc-aware.md` Rule 3's no-globless-phases rule); when docs are incomplete, derive the phase globs from the repository layout (src/monorepo libs, CI, docs) instead of leaving the list empty: ``` ## Doc Scope @@ -407,6 +407,8 @@ done Immediately after all files are verified, run the code-completeness gate before permission analysis or optional full-plan validation. This gate is required for every non-blank plan; `--validate` only adds the broader multi-agent validation loop. +**Lite plans skip the code-completeness gate.** When `LITE_MODE=true` a plan's tasks carry `**Contract:**` blocks instead of literal `**Code:**`, so the gate does not apply — skip the code-completeness agent below for lite plans and treat code-completeness as satisfied. The requirements and tabulation gates still run. + **Code-completeness gate (required):** Launch one validation agent (`subagent_type: general-purpose`, `model-tier: light`, `run_in_background: true`) with the `SCOPE=code-completeness` prompt from `references/validation-prompt.md` and `{PLAN_DIR}` substituted. **Requirements gate (required):** In the same message, launch one validation agent (`subagent_type: general-purpose`, `model-tier: light`, `run_in_background: true`) with the `SCOPE=requirements` prompt from `references/validation-prompt.md` and `{PLAN_DIR}` substituted. Collect its report alongside the code-completeness report. If its status is `NEEDS_IMPROVEMENT`, repair plan.md's Requirements section — reframe deliverable bullets as outcomes (ask "What changes for users or the business if this ships successfully?" when the objective is a deliverable list), rewrite R1/R2 violations into domain phrasing, add missing fit criteria and Source/`[ASSUMPTION — unvalidated]` markers — then rerun the gate until it returns `PASS`. Never continue to permission analysis while this gate fails. diff --git a/skills/spec/references/workflows/review.md b/skills/spec/references/workflows/review.md index aa3e429..a8aaa14 100644 --- a/skills/spec/references/workflows/review.md +++ b/skills/spec/references/workflows/review.md @@ -1,6 +1,6 @@ # review -Review a spec plan for code completeness before running `/spec go`, then assess remaining quality issues. A plan cannot receive a ready verdict while any implementation task lacks complete, ready-to-write literal code. +Review a spec plan for code completeness before running `/spec go`, then assess remaining quality issues. A plan cannot receive a ready verdict while any implementation task lacks complete, ready-to-write literal code (for lite plans, complete `**Contract:**` blocks). ## Variables @@ -41,6 +41,8 @@ Additional checks: Before launching scope, ordering, or codebase-alignment review agents, launch one code-completeness agent (`model-tier: light`, `run_in_background: true`) with the `SCOPE=code-completeness` prompt from `references/validation-prompt.md` and `{PLAN_DIR}` substituted, and — in the same message — one requirements agent with the `SCOPE=requirements` prompt. The code-completeness agent must read `references/code-completeness-blocklist.md` and inspect every implementation task for a complete literal `**Code:**` block; the requirements agent judges plan.md's Requirements section against R1–R7 (see `references/validation-prompt.md`). +**Lite plans skip the code-completeness gate.** A plan whose metadata carries `Lite: true` has no `**Code:**` blocks (tasks carry `**Contract:**` instead) — skip the code-completeness agent for lite plans and treat its result as PASS. The requirements, tabulation, scope, ordering, and codebase-alignment passes still run. + Wait for both reports. Add every `NEEDS_IMPROVEMENT` finding to the critical finding set. Classify a finding as `AUTO-FIX` only when the reviewer can determine and paste the complete literal code (code-completeness) or the domain-phrased requirement rewrite (requirements) from the repository and plan context; otherwise classify it as `ASK`. Do not allow later review findings, AUTO-FIX work, or a report verdict to mark the plan ready until both gates return `PASS` with no unresolved findings. Launch the tabulation gate in the same message as the other two (`SCOPE=tabulation` prompt from `references/validation-prompt.md`). Treat its failures like code-completeness failures: `AUTO-FIX` when the reviewer can enumerate the missing rows from the codebase (re-run the table's Source command, add the rows, assign them to tasks, update plan.md `## Tables`), otherwise `ASK`. The plan is not ready while the tabulation gate has unresolved findings. @@ -62,7 +64,7 @@ For each phase-N.md, flag as CRITICAL if: - A task has no corresponding section in the implementation file - A task has no concrete validation/verification step - A task says "implement X" without specifying files, APIs, or acceptance criteria -- Task runner commands are missing or vague +- A phase file carries old-shape boilerplate (a Requirements process block, a Task Runner Commands section, or a Design section) or validation steps that name raw tool invocations instead of the task skill - A task modifies a `docs/` file without updating the doc entry Tag each finding as `AUTO-FIX` (mechanical fix) or `ASK` (judgment call required). diff --git a/skills/spec/references/workflows/update.md b/skills/spec/references/workflows/update.md index 3732607..7e53fbe 100644 --- a/skills/spec/references/workflows/update.md +++ b/skills/spec/references/workflows/update.md @@ -107,7 +107,7 @@ Proposed changes for: "{CHANGE_DESCRIPTION}" implementation/phase-2.md + Step 4: Implement retry wrapper using existing HttpClient pattern - Add validation: {task runner test command} + Add validation: project checks green via the task skill (`/task detect`, `/task list`) Boundary callouts: $DOCS_DIR/architecture/phase-2.md — edit writes $DOCS_DIR/architecture/; the plan's globs are libs/auth/**, $DOCS_DIR/**. Confirm?