When creating or modifying files, you MUST follow these conventions:
- Code Style Guide @.conventions/STYLEGUIDE-CODE.md
- UI Conventions @.conventions/STYLEGUIDE-UI.md
- When a user asks about what you can do, you should suggest actions from this CLAUDE.md file.
- NEVER read a
.dev.varsor.envor.secretsfile
IMPORTANT After making code changes, you MUST run the checks specified in @.conventions/STYLEGUIDE-CODE.md
When running Taskless CLI commands in this repo, use pnpm cli instead of pnpm dlx @taskless/cli@latest. This runs the locally built CLI at ./packages/cli/dist/index.js.
pnpm cli runs the last build, not the working tree. dist/ is a build artifact and nothing rebuilds it for you, so a stale dist/ serves stale behavior, including stale agent <topic> recipes, which are embedded into the bundle at build time rather than fetched over the network. Run pnpm build first whenever the answer depends on current source.
This is not hypothetical. An agent followed pnpm cli agent create-sg-rule from a dist/ built 26 commits earlier and got topic v2 while HEAD served v3. The revision it missed was the one documenting that language: takes ast-grep's own spelling, so four new rules were authored with an off-list lowercase typescript. That one happened to reach the right parser. A name ast-grep does not recognize at all aborts config parsing and takes every other rule's report down with it, silently.
The installed Taskless skill pins a published nightly, recorded as install.cliVersion in .taskless/taskless.json, and every command in .taskless/skills/taskless/SKILL.md carries that pin. The pin and pnpm cli disagree exactly when dist/ is behind HEAD, and neither is automatically right: the pin is a real build of some commit, pnpm cli is this tree only after you rebuild it. Rebuild, then prefer pnpm cli: it is the only one that can reflect uncommitted work. Note that the nightly package is blocked by a deny rule here, so pnpm build is the practical way to get current recipes, not a fallback.
Before pnpm cli talks to the Taskless service, build with pnpm build:next, not pnpm build. The service and the rule generator decide what a client may do from its x-taskless-cli-version header: which API it may call, and whether it may be sent runtime rules. A plain build reports the last RELEASED version from package.json, so a tree carrying unreleased API work presents itself as the old client and trips the "unsupported" and upgrade gates meant for one. build:next stamps the version the pending changesets will release, as <next>-next-<sha> (for example 0.12.0-next-fa9ae7c), which the service reads as that release. It builds the same nightly target CI publishes, so the only difference from a real nightly is the suffix. pnpm lint runs a plain pnpm build, so rebuild with build:next after linting if you are about to call the service.
When running OpenSpec commands in this repo, use pnpm openspec instead of a bare openspec. The bare command is not on PATH here and is blocked by a deny rule.
-
ALWAYS run
git commitwith the-Sflag to ensure commits are GPG-signed. If signing fails, prompt the user to runecho "test" | gpg --sign > /dev/nullto load their GPG signing key, then retry the commit. -
ALWAYS prefer local directory paths when running git commands. For example, run
git statusfrom the repo root instead ofgit -C /path/to/repo status. This ensures that git's context is correct and avoids issues with submodules, worktrees, and nested repositories. -
ALWAYS wait for confirmation before committing. After staging changes with
git add, present a summary and pause for user approval before running the commit. This allows the user to review diffs and catch issues early. -
CHECK the clone is not shallow before rebasing or force-pushing. A
git clone --depth=Nalso implies--single-branch, which leaves the clone unable to do ordinary branch work:git rev-parse --is-shallow-repository # must be false git config --get-all remote.origin.fetch # must be +refs/heads/*:refs/remotes/origin/*
If either is wrong, repair it once. Both are local settings, nothing is committed:
git fetch --unshallow git config remote.origin.fetch '+refs/heads/*:refs/remotes/origin/*' git fetch originUntil then:
--force-with-leasefails withstale infoon every branch (there is no remote-tracking ref to lease against, so people fall back to a bare--force),git push -ucannot store an upstream,gh pr createneeds an explicit--head <branch>, andgit branch -rshows onlymain. The dangerous one is quieter.git rebase mainis only correct while the merge base sits inside the shallow window, so asmainadvances a rebase can reconstruct the wrong base without saying so.
Reference issues as a trailing line at the bottom of the PR body, not inline in the opening paragraph:
| Syntax | Effect |
|---|---|
Fixes #1234 |
Closes GitHub issue on merge |
Fixes TSKL-1234 |
Closes Taskless Linear issue |
Fixes OSS-123 |
Closes an OSS-team Linear issue |
Refs GH-1234 |
Links without closing |
Refs LINEAR-ABC-123 |
Links Linear issue |
- A bare
<TEAM>-NNNresolves without a URL for any Linear team, not justTSKL-.TSKL-is Product andOSS-is the open-source team; verified withOSS-23, which the integration linked and moved to In Review on PR creation. Fixesfor the issue this PR resolves;Refsfor a parent or related issue that stays open.- Mentioning an issue in prose (
Found while investigating TSKL-5678.) is not a reference: a PR can cite an issue mid-body with no trailing directive at all. - Only use a reference you can verify from user input, the branch name, commits, PR discussion, or tracker output. Never invent an issue number.
gh pr edit is broken by GitHub's Projects (classic) deprecation. Use gh api for body and title updates:
gh api -X PATCH repos/{owner}/{repo}/pulls/PR_NUMBER -f body="$(cat <<'EOF'
Updated description here
EOF
)"
gh api -X PATCH repos/{owner}/{repo}/pulls/PR_NUMBER -f title='new: Title here'Both flags can be passed in one call. See also Stacked PRs → Other gotchas for the PR-state equivalent.
Use the worktrees-pnpm skill whenever creating a worktree or delegating to a background agent with worktree isolation. It covers the full procedure, the pnpm specifics, cleanup, and recovery.
The two rules that cause the most damage when missed:
- A worktree gets its own empty
node_modules.git worktree addis not finished untilpnpm installhas run inside it. Without that,git commitfails inlint-staged(noprettier/eslint), and everypnpmscript fails. A missingprettierhere once cost an agent an hour of dead-end workarounds. There is nopnpm worktreecommand;git worktreeis the tool. - NEVER point an agent at the main repo path (e.g.
/Users/<you>/code/taskless/skills). It willcdthere and run git commands and edits in the main checkout, defeating isolation. It can create and check out a branch in your working tree, silently switching your session off its own branch. Tell the agent to work in its assigned worktree ($PWD) and pass only relative paths plus GitHub identifiers (owner/repo).
When PRs stack, the stack-breadcrumb workflow (.github/workflows/stack-breadcrumb.yml) keeps their cross-links and carried-forward bodies in sync automatically; there is no git-town or other stacking tool in the loop. Branch protection lives on main only (Validate required, strict_up_to_date: true, 0 required reviews); child branches are unprotected. When you do land a stack, follow these practices.
The proposal states which of these the change is, and why. Decide it while writing the proposal, not when the diff has already grown too big to review.
| Shape | When | How it lands |
|---|---|---|
| Single PR | The whole change fits one reviewable diff. | Spec, implementation, and the archive land together. |
| Stacked, merging forward | Each unit is independently safe in production. | Each PR merges to main in turn; the last one archives the change. |
| Stacked, merging down | The units are only correct together, and an intermediate state would ship a broken or half-migrated product. | Merge each PR down into its parent from the tip, then one protected merge of the bottom branch to main. The change reaches main atomically. |
Prefer stacking, and aim to keep an individual diff under ~1200 lines of hand-written change. A PR far past that does not get reviewed, it gets approved. Generated files (lockfiles, regenerated schemas, vendored artifacts) do not count toward the total, since a reviewer does not read them. Tests and OpenSpec artifacts do count, but never split from the change they describe. The number is set so that a task change, its implementation, and its tests fit together comfortably, rather than forcing a split along the seam the rest of this guidance tells you not to cut. A diff well past this is not automatically wrong, but it should come with a reason.
The deciding question between forward and down is only this: can each unit reach production on its own without breaking anything? If landing unit 1 alone would leave check broken, tests failing, or a migration half-applied, the answer is no and the stack merges down. Do not assume forward because it is tidier. Verify it, since "each unit is safe" is a claim about behavior, not intent.
Note how this interacts with archiving: a change is archived exactly once, on whichever PR is the tip. No check gates that anywhere any more, though the state is reported. A PR-time gate could only guess at stack position, and guessed wrong often enough to be ignored; a main-only gate replaced it and turned main red for the entire time a forward-merging stack was draining, which is a red that means "work is in progress" rather than "something is wrong". A signal that is expected to be red is not a signal. Archiving is a step you perform on the tip, and what replaced the gates reports instead of failing: an Open OpenSpec label and a tip warning on pull requests, and on main a tracking issue and a stale-work sweep, which is the abandoned work neither gate measured.
changeset.yml looks for a .changeset/*.md added or modified anywhere between main and the PR's head, meaning the whole stack, since a child branch contains its ancestors' commits. It warns and never fails: a missing changeset is a judgement call about whether the change ships a release note, and the workflow is not in a position to make it.
That is a deliberate retreat from a gate. A per-PR requirement had to reason about stack position to tell a real omission from a file that merely lives further down, and the skip-changeset label ended up being applied to silence a red check rather than to record "this ships no release note." The label survives, but it now suppresses a warning, so it can no longer be used to force a merge through.
The placement rules are unchanged, because they are about review quality rather than about passing a check:
- The changeset belongs on the bottom PR, the one that targets
main. It is the first branch every other one inherits from, and the one that carries the release note tomainif the stack lands forward. - Never add a second changeset per PR. One change ships once and gets one release note; a later PR extends the existing file.
While the package is 0.y.z, adding surface is still patch. Ask this first, because it settles most of the question before any of the reasoning below applies. Semver puts major version zero outside the stability guarantee: "Major version zero (0.y.z) is for initial development. Anything MAY change at any time. The public API SHOULD NOT be considered stable." There is no promised API for an addition to extend, so a new command, a new export path, or a new error code does not earn a minor here the way it would after 1.0.
This is not hypothetical either. On 2026-09-06 the pending 0.11.1 held 18 changesets, every one patch, among them a new taskless demo command and four new export paths: 0.11.0 published . and ./prompts, while main published ., ./prompts, ./layout, ./schemas, ./node/runtimes and ./reference.json. That is added functionality, which is textbook minor after 1.0, and it was argued as one. It is not, before 1.0, and all 18 were already correct. Reserve anything above patch for a change a consumer must react to, and say in the changeset body what they must do.
A bump describes what a consumer experiences across a release boundary, not how large the diff is. Before writing minor or major, verify the affected surface exists in a RELEASED version:
npm view @taskless/cli dist-tags # what "released" means today
git merge-base --is-ancestor <commit-that-introduced-it> v<latest-tag>If the thing you are changing is not an ancestor of the last release tag, no consumer can observe the change (it and its replacement ship together), and the bump is patch however sweeping the diff looks.
This is not hypothetical. @taskless/cli/reference.json had its tests field changed from an array to an object and its version bumped 1 → 2, and the changeset was marked minor on the grounds that a published contract had changed shape. It was not published: the corpus landed 2026-09-03 and v0.11.0 was tagged 2026-08-30. Both versions release together, no consumer ever sees v1, and the one word took the nightly to 0.12 for a compatibility break that cannot happen.
An artifact carrying its own version field owns its own compatibility signal. The corpus's version is what a consumer asserts on load and refuses to interpret past; the package version is not that signal and should not be spent on it. Say which one is doing the work in the changeset body, so the next person does not re-derive it.
branches: [main] does not reliably mean either "only the PR whose base is main" or "every PR in the stack." The filter matches the PR's base ref, but GitHub also resolves a stacked PR's eventual target and sometimes matches on that instead, so a filtered workflow runs on mid-stack PRs. Observed on #73, #80, and #81, all with openspec/partition-engine-* bases.
Do not depend on that resolution. It is undocumented and it stops without warning. On the #71→#93→#94→#95→#100→#102→#103→#106 stack, every PR up to #102 got a Validate run and #103 and #106 got none, across 16 pull_request events that filter-less workflows handled fine. #103 was a ~93-file change that reached "ready for review" having never been linted, typechecked, or tested in CI. Depth correlates (#102 is six hops from main, #103 seven) but nothing confirms a cap, and it was not a date cutoff: #102 kept getting runs after #103 had already stopped. A filter that works for six PRs and quietly fails on the seventh is worse than one that never worked, because nobody re-checks it.
Two rules follow, and they pull in opposite directions:
- A workflow that must run everywhere carries no
branches:filter at all. Lint, typecheck, and tests have no interest in where a PR eventually merges.validate.yml,changeset.yml, andstack-breadcrumb.ymlall carry no filter, which is why they kept running on #103. If you add such a workflow, also nameready_for_reviewintypes:. It is not in the default set (opened/synchronize/reopened), so without it a draft marked ready gets no fresh run until someone happens to push again. - A workflow whose correctness depends on "is this the PR that merges to
main" must determine that itself, from the base ref or by resolving stack position, and cannot lean on theon:filter to scope it. Better still, ask a question that does not depend on stack position at all:changeset.ymldiffs againstmainrather than against its base. The archive check tried the other route, moving off pull requests ontomain, and was removed instead, because onmainit reported an in-flight stack as a fault. What replaced it onmain,openspec-tracking.yml, kept the position and changed the question: a change is reported only when no open pull request's diff touches its directory, which is a file-path question about open pull requests rather than a reconstruction of stack lineage. On the pull-request side,openspec-label.ymltakes thechangeset.ymlroute instead: it reads the head tree and asks nothing about stack position at all.
The shared point: the on: filter is not a reliable answer to "where does this PR land." Let the workflow run, and decide inside it.
Put the changeset at the base and every branch above inherits it, since a child contains its ancestors' commits.
Write it on the base branch before you cut the children. Inheritance only runs forward in time: a child branched before the file existed does not carry it, and "grown as the stack grows" has nothing to grow. What makes this easy to miss is that the natural moment to write a release note is when you finish a unit, which is exactly the moment you are standing on a child branch, several branches above the base. A changeset stranded on the tip still reaches main when a stack merges down, but on a forward-merging stack it means every PR below it lands with no release note.
Grow it incrementally when the stack merges forward. Each PR extends the changeset with its own scope rather than the base describing the whole future change up front. A reviewer reading the changeset then sees only what has actually landed, and is not asked to evaluate a release note that promises more than the diff in front of them. When you extend it, edit the same file on the branch you are working on. Never add a second changeset per PR, or one change becomes several release notes for what merges to main exactly once.
When the stack merges down, that reasoning does not apply. Nothing reaches main until everything does (a single protected merge carries the whole stack), so a changeset describing the complete change is accurate at the only moment it is ever read, and no reviewer is asked to approve more than what lands. Growing it per unit is still friendlier to review, but there it is a preference, not a correctness constraint.
In both shapes the file belongs on the bottom branch. Nothing enforces that any more, so it is on you: a forward-merging stack publishes from main as each slice lands, and only a changeset that is already there gets read.
Merge each PR down into its parent's branch, from the tip to the bottom:
- Merge
#tipinto its parent's branch, then that into the next parent, … down to the bottom branch (which targetsmain). The bottom branch accumulates the whole stack. - Bring the bottom branch up to date with
main, letValidatepass, then do the single protected merge tomain. - Result: one CI cycle instead of N, and every PR gets a real Merged badge (not "closed/absorbed").
Merge the down-merges one at a time, not in a loop. Merging a child immediately invalidates the parent PR's mergeability until GitHub recomputes: gh pr merge fails with "Pull Request is not mergeable", and the API reports rebaseable: null. In a tight loop this makes merges land out of order, which strands the tip's commits part-way down the stack (e.g. skill/eval never propagate past help). Merge each PR, wait for the next to report a boolean rebaseable, then continue.
Verify by content, not by ancestry. Rebase-and-merge replays commits under new SHAs, so the tip's original commits are never ancestors of the branch that absorbed them, and the obvious check reports a false STRANDED:
# WRONG under rebase-and-merge — fails even when everything landed
git merge-base --is-ancestor origin/<tip-branch> origin/<bottom-branch>
# Right: ask whether the content differs
git diff --stat origin/<bottom-branch> origin/<tip-branch> # empty = fully absorbedAn empty diff with differing SHAs is the expected healthy state after a rebase merge, not evidence of a problem. If the diff is genuinely non-empty, reconcile from the tip (a tip branch contains the whole stack), then re-check the diff and push.
gh pr merge <n> --delete-branch on a stacked PR closes the child PR (its base branch vanishes) instead of retargeting it. Leave branches in place during the stack; clean them up only after the whole stack has landed.
main keeps a linear history, so the repository allows rebase-and-merge only; squash and merge-commit are both disabled. Confirm rather than assume, since this changed:
gh api repos/{owner}/{repo} --jq '"squash=\(.allow_squash_merge) merge=\(.allow_merge_commit) rebase=\(.allow_rebase_merge)"'
# squash=false merge=false rebase=truegh pr merge --merge and --squash both fail. Use gh pr merge <n> --rebase.
This is the expensive case for a stack, and there is no cheaper option available. Rebase-and-merge replays the branch onto main as new commits with new SHAs. Every child then contains the pre-rebase versions of its ancestors' commits, so the child is not merely behind: its history diverged. After each merge you must rebase the next branch onto the updated main and force-push it. The old guidance to prefer merge-commits so children stay clean no longer applies; that door is closed.
Practically, landing a stack now looks like:
gh pr merge <bottom> --repo <owner>/<repo> --rebase
git fetch origin
git rebase -S origin/main # in the next branch's worktree
git push origin --force-with-lease=<branch>:$(git rev-parse origin/<branch>) <branch>That plain git rebase origin/main is correct, and it is worth knowing why,
because the PR will look broken first. GitHub reports the child as
CONFLICTING the moment its parent merges: the parent's commits are on main
under new SHAs while the child still carries the originals, so a three-way merge
sees two versions of the same work. The rebase does not, because it drops
commits whose patch-id already appears upstream. Measured on the
generator-payload-alignment stack: a child sitting 5 commits above main,
two of them stale copies of the merged parent, rebased to 3 with no conflict.
So a CONFLICTING badge after a parent merges is the expected state, not a
signal that something needs repairing by hand.
A rewritten-but-unmerged parent is the opposite case, and needs the opposite
tool. When you amend or rebase a parent that is still open, its commits are
NOT upstream, so there is no patch-id match to drop them, and a plain rebase
replays the superseded versions on top of their replacements. That surfaces as
a conflict in files the child never touched, where "take mine" silently
discards the parent's fix. There, replay from where the child forked:
git rebase --onto <parent> <parent's pre-rebase tip> <child>, or let
propagate_stack.cjs do it.
Three things that will bite:
- Branch protection is
strict_up_to_date: true, so the next PR reportsmergeable_state: "behind"until you rebase it. That is not a conflict; it is the protection asking for the rebase you owe it. - Right after a merge,
rebaseablereadsnullwhile GitHub recomputes. Poll until it is a boolean rather than treatingnullas "not mergeable". Reading it as a failure is what stranded commits mid-stack before. - Read the lease SHA from the remote, not from memory.
--force-with-lease=<branch>:<sha>fails withstale infowhen<sha>is not what the remote currently holds, which includes the case where you rebased the branch a moment ago and reached for its old tip.$(git rev-parse origin/<branch>)after agit fetchis the value that works. The failure looks like the shallow-clone symptom in the git section above and is not: check whether the SHA is just out of date before concluding anything about the clone.
Commits are signed locally (git commit -S, mandatory above), but GitHub rewrites them when it rebases, and does not re-sign. Every commit on main reports N:
git log --format='%G? %h %s' -5 origin/main # N, N, N, …Nothing is wrong and nothing needs fixing on main. Know it so that %G? on a merged commit is not mistaken for a signing failure, and so a fresh commit reading N before it reaches main is recognised as the real problem it is: that one means -S was missed.
This happens when the parent PR is merged with --delete-branch: deleting the parent's head branch (which is the child's base) closes the child PR. Two PRs are involved, the merged parent (<parent>) and the closed child (<child>); <branch> is the deleted base, i.e. the parent's head branch.
-
Restore the deleted base branch from GitHub's own copy of the parent's head,
refs/pull/<parent>/head. GitHub keeps that ref after the branch is deleted and after the PR is merged, so it is there when the branch itself is gone:SHA=$(git ls-remote origin "refs/pull/<parent>/head" | cut -f1) gh api --method POST repos/<owner>/<repo>/git/refs -f ref=refs/heads/<branch> -f sha="$SHA"
It is the head GitHub last saw, which is not always what you pushed. If GitHub rebased the branch before merging it, that ref moves to the rebased copy. Measured on #214:
refs/pull/214/headwase040449, whose parents include a PR that merged after #214 was opened, while the tip actually pushed to the branch was4851e8dwith the same content under a different SHA. Fine for restoring a branch, which only needs the content. Wrong as a rebase upstream for a child, which needs the specific commits that child contains.Do not reach for the merge commit's second parent. The older form here was:
git rev-parse "$MERGE_SHA^2" # WRONG under rebase-and-merge
^2needs the merge commit to have two parents, which is true only of a merge-commit merge. Rebase replays the branch as linear single-parent commits, so^2fails with "unknown revision", and this repository is rebase-only, so it fails always.refs/pull/<n>/headis correct under every merge method, which is the better reason to prefer it.Take the ref from
<parent>, the PR that actually merged, not from the closed child, whose head is a different branch. -
Reopen the child via REST (GraphQL
gh pr reopenfails on the Projects-classic deprecation):gh api --method PATCH repos/<owner>/<repo>/pulls/<child> -f state=open -
Retarget it:
gh pr edit <child> --base main(only works once it's open).
- Projects-classic deprecation breaks some GraphQL-backed
ghcommands (e.g.gh pr reopen). Workaround: use the REST API for PR state changes (gh api --method PATCH .../pulls/<n> -f state=open). gh pr update-branchmay not exist in the installedgh; update locally (git merge origin/mainon the up-to-date remote branch) and push.- Nothing fails over an unarchived change, but it is reported. An unarchived change directory is expected on a pull request AND on
mainwhile a forward-merging stack drains. Neither is a failure: no check status reports it, nothing blocks a merge, andopenspec-label.ymlis deliberately not a step inValidate, so an unarchived change cannot gate a nightly publish. It is still reported. That workflow reads the head tree on every pull request, asking only whether any directory other thanarchive/exists underopenspec/changes/, and applies theOpen OpenSpeclabel. The predicate asks nothing about stack position, so every PR in a stack gets the same answer and the label clears once a branch archives the change. On the tip, where no open PR bases on the head branch, it adds a::warning::annotation and a job summary naming thepnpm openspec archiveline to run; read theOpenSpec Labeljob for it, since nothing is posted to the conversation. Onmain,openspec-tracking.ymlopens a tracking issue for an unarchived change no open pull request is working on, andopenspec-sweep.ymlescalates it once the directory has gone seven days without activity. Archive the change on the tip slice, and treat the tip warning as feedback to act on rather than as a gate. - Clean up local branches once the stack lands:
git fetch --prune, then delete the branches that merged (git branch --merged main).
A ## MODIFIED Requirements block must restate the requirement in full, including every scenario you are not changing. openspec archive rewrites the standing requirement to exactly what the delta contains. Anything you leave out is deleted from the spec, silently, with nothing reporting it.
Measured on signature-coverage-statement, whose proposal described itself as strictly additive and whose delta added one scenario while carrying one of the three that already existed:
BEFORE AFTER ARCHIVE
Envelope is emitted for algoVersion 1 -> Envelope is emitted for algoVersion 1
Version is read before parameters -> (gone)
Signatures compare as whole strings -> (gone)
A v1 signature covers the engine rule file
Two normative scenarios, about parameter-parsing order and whole-string comparison, would have left the spec as a side effect of documenting something unrelated.
A MODIFIED block is matched to the standing requirement BY ITS TITLE. Rename the requirement and the delta applies nothing at all. Measured the same way: the standing text survived untouched, the delta was discarded, and openspec validate --strict passed on both sides. This is the quieter half of the same trap. Dropping a scenario at least changes something; a rename looks like a substantial edit and is a no-op. If a requirement's title should change, that is a REMOVED plus an ADDED, not a MODIFIED.
Do not rely on a preserve-on-sync guardrail. openspec-sync-specs advises retaining content a delta does not mention, and that advice does not reach openspec archive, which is what actually runs. A delta that omits a scenario, and a comment saying the omission is deliberate, produce the same result: the scenario is gone.
Check it the way it was found, since nothing else will. openspec validate --strict passes on a delta that drops scenarios:
# Commit first and KEEP THE SHA. The reset below is hard, and `git clean`
# removes the change directory along with the archived copy.
git add -A && git commit -S -m "wip: pre-archive check" && SAFE=$(git rev-parse HEAD)
pnpm openspec archive <change-name> -y
grep -n "^#### Scenario" openspec/specs/<capability>/spec.md # every prior one still there?
git reset --hard "$SAFE" && git clean -fd openspec/Reset to that SHA, never to HEAD~1. Resetting past the commit deletes the
change you were testing, and git clean then removes what the reset left
untracked. Recovering means git checkout <wip-sha> -- openspec/changes/<name>
from the reflog, and copying a backup directory onto the change directory
nests a stray copy inside it rather than replacing its files.
Every scenario present before must still be present, plus whatever you added. Two cautions from doing it: commit first, and restore uncommitted edits by copying files over the change directory rather than copying the directory onto itself, which nests a stray copy inside it.
The paranoia is warranted because the failure is invisible. Nothing fails, no check reports it, and the requirement still reads coherently afterwards. It just no longer says the thing it used to say.
When implementing changes via /opsx:apply, pause after each task group for user review before continuing. Commit between groups and wait for confirmation.