Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 6 additions & 2 deletions skills/spec/agents/spec-executor.md
Original file line number Diff line number Diff line change
Expand Up @@ -55,12 +55,16 @@ You are precise, minimal, and disciplined. You follow implementation specs exact
- If a task requires an unplanned build system change, STOP and flag it

**Hygiene after every task (non-negotiable):**
- Run format → lint → typecheck → tests using the project's task runner commands from plan metadata
- Run format → lint → typecheck → tests via the task skill — phase files carry no runner commands
- Fix all failures before marking the task complete
- Never leave a task in a state where any of these fail
- Never invent shell commands — use only task runner recipes discovered from the project's `mise.toml`, `justfile`, `Makefile`, or `package.json` scripts
- Never invent shell commands and never call build/test tools directly — discover recipes with the task skill and run those
- At phase start, call `/task detect` once to identify the runner and `/task list` to enumerate recipes; reuse those names for every hygiene/build/test command

**Terse prose:**
- Everything you write — execution-log entries, reports, commit messages, markdown — is brief and terse
- One line where one line suffices; no preamble, no restating the spec, no summaries of what the checklist already tracks

**Progress tracking (non-negotiable order):**
1. Complete the task implementation
2. Run hygiene checks (format → lint → typecheck → tests)
Expand Down
6 changes: 2 additions & 4 deletions skills/spec/agents/spec-planner.md
Original file line number Diff line number Diff line change
Expand Up @@ -49,16 +49,14 @@ You are thorough and opinionated. You write plans that are detailed enough to be
- High-level objective (2-4 bullets)
- Design overview (key decisions, not implementation detail)
- Phase/task checklist (one-liner per task)
- Task runner metadata

**Implementation files** — detailed, and **always show the complete code**:
- Step-by-step instructions per task
- Exact file paths, not "relevant files"
- **Code-first gate (non-negotiable):** Before you write ANY task into an implementation file, you must already have the exact code that task will produce. For every task, ask: *"Can I paste the literal code — every new/changed line — right now?"* If **no**, do NOT write the task yet: resolve the unknown first (read the codebase with Glob/Grep/Read, search the web, or ask the user once), then write it. Never emit a task whose code you would leave for the execution agent to figure out.
- **Show the whole thing, not a sketch — and for edits, a diff, not the whole file:** For a new file, include its entire contents. For an edit to an existing file, show a **diff** — the exact old→new lines or a unified diff with context lines above and below the change — never the whole file. A whole-file replacement is reserved for tasks where the user explicitly asked for a full-file replacement; a whole-file dump where a diff is required is a plan defect. Pseudocode, ellipses (`...`), "e.g.", and "something like" are forbidden inside a task's code block — they are the exact failure this rule exists to prevent.
- **Self-audit before finishing each phase file (mandatory):** re-read every task and confirm each has a non-empty, complete code block. Reject and rewrite any task whose code block: is missing or empty; contains a placeholder/stub marker from the blocklist in `references/code-completeness-blocklist.md` (the canonical list — judge by intent, not blind substring matching); shows only a signature or a comment where a body belongs; or describes the code in prose instead of showing it. Every line the execution agent will write must appear verbatim in the block. A task that fails this audit is incomplete — resolve the unknown now (read the codebase, search the web, or ask the user once) and paste the real code, or delete the task. Do not finish the phase file until every task passes.
- Task runner commands for validation after every task (format → lint → typecheck → test) discovered by reading `mise.toml`, `justfile`, `Makefile`, or `package.json` — never invent shell commands
- Every build/test/lint command MUST come from the project's task runner
- A validation line per task: "project checks green via the task skill (`/task detect`, `/task list`)" — no recorded runner commands, no raw npm/pytest/go-test invocations

**User guide** — usage-focused:
- What was built and how to use it
Expand All @@ -71,7 +69,7 @@ Every plan you create must include these as explicit constraints in implementati

1. **Minimal changes**: Execution agent makes the smallest change that achieves the goal. No drive-by fixes.
2. **Build system preservation**: Do not modify the build system unless the plan explicitly requires it. The project must build after every task.
3. **Hygiene**: Run format → lint → typecheck → tests after every task using the project's task runners. Fix failures before moving on.
3. **Hygiene**: Executors run format → lint → typecheck → tests after every task via the task skill. Fix failures before moving on.
4. **Validation phase**: Every plan must end with a phase that confirms the full suite passes and the user guide is complete.

## Doc-Aware Planning (active only when `PERSISTENT_MODE` is set)
Expand Down
2 changes: 1 addition & 1 deletion skills/spec/agents/spec-updater.md
Original file line number Diff line number Diff line change
Expand Up @@ -42,7 +42,7 @@ plan.md implementation/phase-N.md
─────────────────────────────── ──────────────────────────────────────
High-level task checklist Step-by-step execution detail
One line per task Full context, code examples, commands
Phase headers with ✅ markers Task runner commands and validation steps
Phase headers with ✅ markers Task-skill validation steps and deviations
```

**These two views must always agree.** When you modify one, ask: does the other need to change too?
Expand Down
4 changes: 2 additions & 2 deletions skills/spec/references/agent-prompt.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,8 +12,8 @@ Branch: {PLAN_BRANCH}
Worktree: {PLAN_WORKTREE or "(none)"}
Git commits allowed: {ALLOW_COMMITS}

Task runner commands:
{Discover by reading mise.toml/justfile/Makefile/package.json directly}
Project checks:
Discover and run every build/test/lint/format command via the task skill (`/task detect`, then `/task list`) — the phase file carries no runner commands. Never call build/test tools directly and never invent shell equivalents.

## Completed phases (summary only)
{One line per completed phase: "Phase 1 (Setup) ✅ — 4/4 tasks, tests passing"}
Expand Down
72 changes: 22 additions & 50 deletions skills/spec/references/implementation-template.md
Original file line number Diff line number Diff line change
@@ -1,55 +1,41 @@
# Implementation Phase File Template

Use this structure for each `implementation/phase-N.md` file created in Step 5.5.
Use this structure for each `implementation/phase-N.md` file. Phase files are task documents: a one-line introduction, the tasks with their complete code, then validation and deviations. Process rules (brevity, build-system preservation, markdown, terse prose) live in the executor agent definition — never repeat them here. Design context lives in plan.md and the user guide. Executors discover build/test/lint/format commands via the task skill, so no runner commands are recorded.

```markdown
# Phase {N} - {Phase Name}

## Introduction
{Brief description of what this phase accomplishes and its role in the overall plan.}
{One or two sentences — what this phase accomplishes.}

## Requirements
## Doc Scope
{ONLY for doc-aware (`--persistent`) plans — omit this section otherwise. Per `references/doc-aware.md` Rule 3: the `**Write globs:**` list must contain at least one glob; cross-module interaction uses documented public interfaces only (Rule 4); each crossing gets a boundary callout (Rules 6–7). Verify every write target with `$SPEC_SKILL/scripts/scope.py`.}

**Brevity:** Make the smallest change that achieves the task. No drive-by refactors or unrelated fixes.
**Write globs:** {`glob1`} {`glob2`}

**Build system preservation:** Do NOT modify the build system, CI config, or dependencies unless this phase is explicitly about them. If the project built before you started, it must build after every task. If a change would require an unplanned build system modification, stop and flag it.
**Public interfaces only:** {cross-module interaction uses only the target module's documented public API — never its internals.}

**Markdown output:** Soft-wrap prose — never hard-wrap. Write each paragraph as one continuous line; do not insert manual newlines to wrap prose at a fixed column width. Newlines still separate paragraphs, list items, headings, and code fences.

### Task Runner Commands
{List the relevant task runner commands for this phase. ALWAYS use these — never invent equivalent shell commands. Discover them by reading `mise.toml`, `justfile`, `Makefile`, or `package.json` scripts directly.}
- Build: `{e.g. just build | make build | task build | mise run build}`
- Test: `{e.g. just test | make test | task test | mise run test}`
- Lint: `{e.g. just lint | make lint | mise run lint}`
- Format: `{e.g. just fmt | make format | task fmt | mise run fmt}`

If no task runner covers a needed operation, note: "Gap: no recipe for X — suggest adding one."

## Design
{Describe the approach taken in this phase — what pattern, what structure, what architectural decision was made before coding begins.}
**Boundary callouts:**
- {task} — {what crosses the boundary, why it is required, and the alternative considered}

## Implementation

> **Gate (machine-checked in validation):** Every task below MUST contain a `**Code:**` block holding the **complete, literal code** it will produce — full contents for new files; for edits to existing files, a **diff** (exact old→new lines or a unified diff) with context lines above and below the change, never a whole-file dump. A whole-file replacement appears only when the task explicitly states that the user asked for a full-file replacement. The block is REJECTED if it is missing or empty, contains a placeholder/stub marker from the blocklist (see `references/code-completeness-blocklist.md` — the canonical list), shows a bare signature/comment where a body belongs, describes the code in prose instead of showing it, or replaces an entire existing file with a whole-file dump where a diff was required. The blocklist is judged by intent, not blind substring matching, so a marker used as a legitimate token (not a stand-in for missing code) does not fail the block. If you cannot show the complete code, resolve the unknown now during planning (read the codebase, search the web, or ask the user) — never pass research, open design choices, or code authoring to the execution agent. A dedicated validation agent scans for exactly these placeholders and will fail the plan until every code block is complete.
> **Gate (machine-checked in validation):** Every task below MUST contain a `**Code:**` block holding the **complete, literal code** it will produce — full contents for new files; for edits to existing files, a **diff** (exact old→new lines or a unified diff) with context lines above and below the change, never a whole-file dump. A whole-file replacement appears only when the task explicitly states that the user asked for a full-file replacement. The block is REJECTED if it is missing or empty, contains a placeholder/stub marker from the blocklist (see `references/code-completeness-blocklist.md` — the canonical list), shows a bare signature/comment where a body belongs, describes the code in prose instead of showing it, or replaces an entire existing file with a whole-file dump where a diff was required. The blocklist is judged by intent, not blind substring matching, so a marker used as a legitimate token (not a stand-in for missing code) does not fail the block. If you cannot show the complete code, resolve the unknown now during planning (read the codebase, search the web, or ask the user) — never pass research, open design choices, or code authoring to the execution agent.
The executor may still apply minor mechanical fixes to the specified code in-flight — obvious typos, missing imports, or trivial type corrections that do not change a module's contract or violate an invariant — per the permitted-minor-deviations policy in `agents/spec-executor.md`; these are logged as `[MINOR-DEVIATION]` entries, not treated as failures.

{For each task in this phase:}

### Task {X}: {Task Description}

**Steps:**
1. {Detailed step-by-step instructions}
2. {Include exact commands, file paths, code patterns}
1. {Concise step-by-step instructions}

**Code (required — complete, never omit, never abbreviate):**
```{lang}
{The COMPLETE code this task produces. New file: its entire contents. Edit to an existing
file: a diff — the exact old→new lines or a unified diff with the changed lines plus context
lines above and below, so the executor sees what changes and where. Never paste the whole
file for an edit; a whole-file replacement is shown only when this task explicitly declares
a user-requested full-file replacement. Every line the execution agent will write appears
here verbatim. No ellipses, no pseudocode, no "e.g." A task with a partial, prose-only, or
whole-file-dump code block is incomplete and must not be emitted.}
lines above and below. Never paste the whole file for an edit; a whole-file replacement is
shown only when this task explicitly declares a user-requested full-file replacement.}
```

**Contract (lite mode only — replaces `**Code:**`):**
Expand All @@ -65,37 +51,23 @@ Acceptance:
- <the invariant or test that proves the task is done>
```

**Files to modify / create:**
- `path/to/file.ext` — {specific changes}

**Table rows (only when the task consumes a `tables/*.md` table):**
- `tables/{set-slug}.md` — rows {#–#}: mark each row's Status `[x]` as it is completed; a row is done only when its change is applied and validated

**Validation (run after every task):**
- [ ] `{fmt command}` — no formatting changes outstanding
- [ ] `{lint command}` — zero warnings/errors
- [ ] `{test command}` — all tests pass
- [ ] *(only under `spec go --commit`, and only when CI exists)* CI is green for this committed phase — the executor invokes `/git commit --yes --fix` and the git skill blocks in its bounded fix-until-green loop. This is a HARD gate: the phase is not complete while CI is red, and if the git skill stops with CI still failing, the phase is reported FAILED. Skip silently when any of: no `--commit`, no remote, no configured CI, no `gh`/`glab` CLI. A phase never fails for the absence of CI — only for a CI that exists and stays red.

---
**Files to modify / create:**
- `path/to/file.ext` — {specific changes}

## Future Work / Validation
**Validation (run after every task):**
- [ ] Project checks green via the task skill (`/task detect`, `/task list` — then the repo's test/lint/typecheck recipes); fix failures before marking the task complete
- [ ] *(only under `spec go --commit`, and only when CI exists)* CI is green for this committed phase — the executor invokes `/git commit --yes --fix` and the git skill blocks in its bounded fix-until-green loop. Skip silently when any of: no `--commit`, no remote, no configured CI, no `gh`/`glab` CLI.

After all tasks complete, run the full suite:
## Validation

```bash
{fmt command}
{lint command}
{typecheck command}
{test command}
{build command}
```
After all tasks complete, run the repo's full check set via the task skill and record the outcome:

Expected: {describe expected output}
- [ ] Full suite green (tests + validate/lint as the repo defines them) — note the exact recipes run

Note any follow-on tasks or deferred items surfaced during this phase:
- {deferred item}
## Deviations

## References
- {Related proposals, ADRs, or external references consulted for this phase}
- {Deviations, minor deviations, or deferred items surfaced during this phase — or "none"}
```
8 changes: 5 additions & 3 deletions skills/spec/references/validation-loop.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,7 +23,7 @@ prompt: [contents of references/validation-prompt.md with SCOPE=plan-level]
```

Validate:
- plan.md metadata (Task Runners field present, branch/worktree filled)
- plan.md metadata (no runner-command metadata, branch/worktree filled)
- user-guide.md exists and has non-TODO overview
- Phase names and task counts consistent across plan.md and implementation files
- Inter-phase dependencies identified
Expand All @@ -38,11 +38,13 @@ prompt: [contents of references/validation-prompt.md with SCOPE=phase, PHASE_N={
Validate only `implementation/phase-{N}.md` against the plan.md tasks for that phase:
- Task specificity and actionability
- Implementation completeness (file paths, code examples, no ambiguity)
- Task runner commands listed and used (not raw npm/pytest/go test)
- fmt/lint/typecheck/test validation block present
- Lean shape: only Introduction / Implementation / Validation / Deviations sections, plus `## Doc Scope` for doc-aware plans (no Requirements process block, no Task Runner Commands, no Design)
- Validation steps reference the task skill (no recorded runner commands, no raw npm/pytest/go test invocations)
- Test coverage and success criteria
- user-guide.md update instructions per task

**Lite plans skip this agent.** If the plan's metadata carries `Lite: true`, there are no `**Code:**` blocks to check — do not launch the code-completeness agent and treat `VALIDATION_CODE_COMPLETENESS_STATUS` as PASS.

**Code-completeness agent** — launch one agent (`subagent_type: general-purpose`, `model-tier: light`, `run_in_background: true`):

```
Expand Down
15 changes: 7 additions & 8 deletions skills/spec/references/validation-prompt.md
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,7 @@ Read these files:
Validate the following quality criteria:

**Metadata**
- Does plan.md have a "Task Runners" field with actual commands?
- Does plan.md carry no runner-command metadata? (Checks are discovered at execution time via the task skill — a "Task Runners" field is stale boilerplate.)
- Are branch/worktree fields filled (or "(none)" explicitly)?

**User Guide**
Expand Down Expand Up @@ -80,12 +80,11 @@ Validate the following quality criteria for Phase {N} only:
- Are code examples present for non-trivial logic?
- Is there enough detail for an autonomous agent to execute without asking clarifying questions?

**Task Runner Usage**
- Does phase-{N}.md list applicable task runner commands in a "Task Runner Commands" section covering build/test/lint/format/typecheck?
- Are all build/test/lint/format/typecheck/run commands using the project's task runners (not raw `npm test`, `python -m pytest`, `go test ./...` when a task runner wraps them)?
- Does every task's validation checklist include lint, format, and typecheck steps — not just tests?
- Is there a "Phase Validation" block at the end with all five checks (fmt, lint, typecheck, test, build)?
- Is there a note that lint/format/typecheck must run after every task, not only at phase end?
**Lean shape and check discovery**
- Does phase-{N}.md contain ONLY the sections Introduction, Implementation, Validation, and Deviations — plus a `## Doc Scope` section when the plan is doc-aware (`--persistent`)? A "## Requirements" process block, a "Task Runner Commands" section, or a "## Design" section is a failure — those rules live in the executor agent definition and the design lives in plan.md.
- Does every task's validation checklist run the project's checks via the task skill (references `/task detect` / `/task list` or equivalent) — not recorded runner commands, and never raw `npm test` / `python -m pytest` / `go test` style invocations?
Comment thread
skapoor8 marked this conversation as resolved.
- Does the phase close with a Validation section (full-suite run via the task skill) and a Deviations section?
- Is the Introduction one or two sentences, with no policy prose?

**User Guide**
- Does each task in phase-{N}.md specify what to update in user-guide.md once complete?
Expand All @@ -95,7 +94,7 @@ Validate the following quality criteria for Phase {N} only:
- Are task descriptions consistent between plan.md and phase-{N}.md?

**Test Coverage**
- Does each task specify what tests to write or run using the task runner?
- Does each task specify what tests to write or run via the task skill (never raw `npm test` / `python -m pytest` / `go test` style invocations)?
- Are acceptance/success criteria testable?

Respond ONLY in this exact format:
Expand Down
Loading
Loading