You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Write docs/loopwright/principles.md: a descriptive statement of what
loopwright is opinionated about and what it treats as a swappable detail. It
changes no engine code and no workflow, and it gates nothing. Its job is to give
a name to the boundary that three open RFCs already cite, so that boundary stops
being re-argued in every design session.
The core opinions grow from six to twelve, each carrying the concrete change it
rejects. The detail territory grows from four entries to eight, each with a named
seam and its current default.
Context
Loopwright exists because of an opinion about how work flows and how "done" is
judged, not because of a stack. Today its identity is written through the stack
anyway — README.md:3-6 leads with "vendor it into any JS/TS repo", package.json:4 says "AI-assisted JS/TS repos", CLAUDE.md:1-6 describes the
project by the directory its engine lives in and the command that runs its tests,
and .loopwright/claude-md-section.md — the block install.sh injects into every
host repo — leads with "vendored under .loopwright/" before naming anything the
layer is opinionated about.
The destination is a near-autonomous loopwright, done properly rather than for
hype. That reframes every opinion below: they are not brakes on autonomy, they
are what makes autonomy not be theatre. An agent that merges on its own is only
acceptable if "done" is a mechanical verdict that cannot be bought — so the gate
as sole authority, the integrity signals and the ratchet are preconditions of
autonomy, not constraints on it.
Three downstream RFCs already depend on this document textually:
RFC: Stack connectors, decouple the quality gate from JS/TS #6: carries the defaults rule nearly verbatim — "mise is the default
provider, the way eslint is the default linter that framework starters ship
with: recommended and wired out of the box, never required."
They do not need the document to enforce anything. They need it to exist and to
name things, so they can cite it. That is why this RFC is descriptive.
Decision
The document is descriptive, not normative
docs/loopwright/principles.md. No skill consults it, no gate checks it, no RFC
is required to carry a classification line. Rejected alternative: a normative
version where /grill-rfc demands a classification line from every future RFC.
It would have meant editing a skill, and the value it adds is small — what the
downstream RFCs need is the vocabulary, which prose delivers.
"Core" means the rule has no knob, not that the rule is frozen and not that
the action it governs has no knob. An action may be configurable when the rule
itself provides for it (see P8). Changing a core opinion is an explicit amendment: a dated entry naming the RFC that changed it, in an Amendments
section. Ids are stable and amendments never renumber; a withdrawn opinion stays
in place as P4 (withdrawn, RFC #N).
Form. Each opinion carries three things: the assertion, a mandatory rejects: line naming a concrete change it refuses, and a pointer to the file
that enforces it — file name only, never a line number, and an empty pointer
where nothing enforces it. An empty pointer is itself a finding. Rationale goes
in prose below the list.
Core opinions
id
Opinion
P1
Work flows RFC issue → task sub-issues → PRs. State lives in artifacts, never in session memory. The trigger is detail; the flow is core.
P2
"Done" is determined by configured checks, never by the judgement of whoever merges. Merge — autonomous or not — requires every configured check green for the current diff and review findings resolved.
P3
Quality is a ratchet against a committed baseline, with per-file debt grandfathered. The baseline is never regenerated to hide a regression.
P4
Green earned vs bought — the integrity signals — is a first-class concern, not an add-on.
P5
Infrastructure failure blocks. Disabling a configured check is itself a violation.
P6
Agents communicate and coordinate through a mediator, never peer to peer.
P7
All loopwright policy lives in one versioned file under .loopwright/ in the host repo, and the engine's defaults are frozen. Changing what the gate demands is a PR in the host repo.
P8
Irreversible actions require a human by default. Automating one is an explicit opt-in in the host's config file.
P9
Strict TDD: the test comes before the implementation.
P10
Fresh-context review: every diff is reviewed by an agent that did not write it.
P11
A unit of work is a vertical slice, never a horizontal layer.
P12
The metric vocabulary and its semantics are core and closed. Connectors measure; they never invent a key.
Why the contested branches settled where they did
P2 absorbed review, and P8 took merge out of P2. The original text said
"merging is always a human decision" as a core opinion. It is not core — it is a
default, and a configurable one, in the spirit of the Claude Code permission
modes. What is core is the rule above it: nothing may remove the human from an
irreversible action silently. So the authority claim moved to P8, and P2 kept the
part about what "done" is. "Done" includes resolved review findings, because
that is what /babysit-pr already enforces — a doc that said otherwise would be
describing a harness that does not exist, and would authorise autonomous merge
over an unresolved review, which is the most elegant way to buy green.
Note the division of labour: "only quality-gate.mjs can fail the build" is CI
mechanics and stays in docs/loopwright/quality-gate.md. P2 is about who decides
"done", which is a different claim.
P7 merged two decisions and excludes the baseline. Engine defaults are frozen: a loopwright update never moves a host's verdict. This closes #8's
first risk ("silent gate tightening through defaults") and settles its open
question 1 — overlay is safe because the defaults are frozen; it is the freeze,
not a full copy, that removes the risk. The model is tsconfig.json: the host
writes only its overrides, the defaults live in the tool, and they do not move
without a breaking release. baseline.json is not policy and is not the file
P7 names: it is measured state, machine-written, engine-schema'd, host-committed,
never hand-edited. P3 owns it.
P8 is the one place this document changes a rule.docs/loopwright/loop-harness.md:66
states that merging is always the human's decision, absolutely. P8 makes it
opt-in-automatable. Enumerated irreversibles today: merge, publishing a tag or
release, deploy — the list is not exhaustive. Everything else the harness
automates is reversible and therefore needs no opt-in: closing an issue reopens, gh pr ready despromotes, comments edit. The one irreversible thing the harness
touches is merge, and that is the one it does not do.
A consequence, recorded because it contradicts a published RFC: the mediator
may decide the version bump, since a bump lives in a file inside the PR and the
human sees it in the same act as the merge. #9 states "The bump type is a human
decision... Agents never choose it"; it adjusts when it is grilled.
P12 was already settled by #6, which lists "the metric vocabulary and its
semantics" among what core keeps. A key a connector invents is persisted into
the baseline (quality-gate.mjs writes metrics: current wholesale) but never
scored, because evaluation iterates over the config, not over the reports — a
metric that looks measured and is not, which is exactly what P4 exists to
prevent. The connector contract is three-valued: a value, unconfigured
(warns forever, never blocks), or failed (always blocks).
The limit of autonomy is deliberately not declared. Work intake moves from
hand-written RFCs to GitHub labels to an alert webhook, and a Datadog alert
driving investigate → fix → deploy must stay possible without this document
passing judgement on it. What is not deferred: an issue and a PR exist before
anything reaches production, because P1 requires the artifact — without it there
is nowhere for the diagnosis to live, and a fix deployed without a trail is green
without evidence, in the most expensive place possible.
Detail territory
Every entry is a territory, a named seam, and the current default. This RFC does not define any of these contracts — each belongs to its own RFC.
Territory
Seam
Default today
T1
Host language/ecosystem and the tools that measure it
Defaults are opinions too. Every territory ships with one recommended
default, wired out of the box, the way framework starters ship eslint. A default
is never a requirement: there is always a documented way to swap it.
T5 is a contract of exactly four capabilities the loop requires of its agent
host: dispatch a subagent with fresh context and an explicit payload; restrict
tools per role; choose a model tier per task; an external reviewer that can write
on the PR. Anything the loop needs beyond those four is leakage — it either
enters the contract as a core change, or it is a bug. This answers the original
open question 1 with a measurable test instead of a judgement call. The engine
itself has zero coupling here: nothing under .loopwright/scripts/ references
Claude, agents, skills or sessions. Capability 2 is load-bearing for P10, because
a reviewer is read-only by being handed no write tools, not by being told to
behave.
T6 is independent of T1. CI written in a different language than the project
is ordinary, so a Go host needing Node in CI is not coupling. #6's open question
4 ("should the engine ship as a single binary?") is therefore a deferred
detail, not a question of identity.
T7 works on a non-JS host because of T6: the engine is already Node and
already has .loopwright/node_modules/, so changesets can be internal to
loopwright and still serve a Go host. If the host already uses changesets,
loopwright uses the host's.
Classification rule
If a change alters what counts as done, or how work flows, it is core and
must be argued against the opinions above. If it alters how a fact is measured
or where a step runs, it belongs behind a contract, with a default
implementation.
Scope
New docs/loopwright/principles.md.
docs/loopwright/loop-harness.md:44-77 loses the "Rules of the loop" that are
opinions — they are deleted there and replaced by a reference, since the new
doc is their only author. What stays is harness mechanics: base branch,
closing keywords, draft-until-ready, and the two grill-phase notes at :67-72
("architecture is scaffolding, not a spec") and :73-77 ("boundary-reviewer
is the first thing to cut"). Line 66 is reconciled with P8.
Identity rewritten principles-first in four places: README.md, CLAUDE.md, .loopwright/claude-md-section.md (the most important of the four — it is what
every host repo reads), and package.json:4. The "How it's wired" table at README.md:158-169 gains a row. README.md:142-143 is the second site of the
absolute merge rule — "Two rules hold the whole thing together" — and is
reconciled with P8 here, so that README.md has a single owner.
A "Known violations" section, and any issue to fix the engine's current
divergence from P5 (see Risks).
Risks
Principles written too broadly stop deciding anything. Mitigated by the
mandatory rejects: line: each opinion has to name a concrete change it
refuses, which turns a slogan into a test.
P8 is the only non-descriptive line in a descriptive document. It describes
the decided state, not the current one, and the knob it assumes does not exist
yet.
An engine divergence from P5 is known and not being fixed here. quality-gate.mjs:97-99 uses COLLECTOR_METRICS[collector] ?? [], so a
collector outside the hardcoded map that fails maps to zero metric ids,
contributes nothing to the worst status, and the gate passes. Recording it in
the doc was considered and rejected; fixing it would be a code change, which
this RFC excludes.
Twelve opinions is near the limit of what stays memorable. If the list
keeps growing, it stops being read, and an unread principle decides nothing.
Task breakdown (link sub-issues)
To be reconciled by /grill-rfc once the sub-issues are created.
Open questions
More flows for the autonomous pipeline. A one-line fix originating from an
alert does not sit comfortably inside an RFC. The flow is core for now; a
second flow for incident-shaped work is deferred, not rejected.
The limit of autonomy is intentionally undeclared. Nothing in this
document says where a human must remain. That is a decision to revisit only if
an actual problem appears, not a gap to fill pre-emptively.
Summary
Write
docs/loopwright/principles.md: a descriptive statement of whatloopwright is opinionated about and what it treats as a swappable detail. It
changes no engine code and no workflow, and it gates nothing. Its job is to give
a name to the boundary that three open RFCs already cite, so that boundary stops
being re-argued in every design session.
The core opinions grow from six to twelve, each carrying the concrete change it
rejects. The detail territory grows from four entries to eight, each with a named
seam and its current default.
Context
Loopwright exists because of an opinion about how work flows and how "done" is
judged, not because of a stack. Today its identity is written through the stack
anyway —
README.md:3-6leads with "vendor it into any JS/TS repo",package.json:4says "AI-assisted JS/TS repos",CLAUDE.md:1-6describes theproject by the directory its engine lives in and the command that runs its tests,
and
.loopwright/claude-md-section.md— the blockinstall.shinjects into everyhost repo — leads with "vendored under
.loopwright/" before naming anything thelayer is opinionated about.
The destination is a near-autonomous loopwright, done properly rather than for
hype. That reframes every opinion below: they are not brakes on autonomy, they
are what makes autonomy not be theatre. An agent that merges on its own is only
acceptable if "done" is a mechanical verdict that cannot be bought — so the gate
as sole authority, the integrity signals and the ratchet are preconditions of
autonomy, not constraints on it.
Three downstream RFCs already depend on this document textually:
and it rejects peer-to-peer messaging because it "conflicts with the mediated
principle".
detail with a default, and its section is where you swap it."
provider, the way eslint is the default linter that framework starters ship
with: recommended and wired out of the box, never required."
They do not need the document to enforce anything. They need it to exist and to
name things, so they can cite it. That is why this RFC is descriptive.
Decision
The document is descriptive, not normative
docs/loopwright/principles.md. No skill consults it, no gate checks it, no RFCis required to carry a classification line. Rejected alternative: a normative
version where
/grill-rfcdemands a classification line from every future RFC.It would have meant editing a skill, and the value it adds is small — what the
downstream RFCs need is the vocabulary, which prose delivers.
"Core" means the rule has no knob, not that the rule is frozen and not that
the action it governs has no knob. An action may be configurable when the rule
itself provides for it (see P8). Changing a core opinion is an explicit
amendment: a dated entry naming the RFC that changed it, in an
Amendmentssection. Ids are stable and amendments never renumber; a withdrawn opinion stays
in place as
P4 (withdrawn, RFC #N).Form. Each opinion carries three things: the assertion, a mandatory
rejects:line naming a concrete change it refuses, and a pointer to the filethat enforces it — file name only, never a line number, and an empty pointer
where nothing enforces it. An empty pointer is itself a finding. Rationale goes
in prose below the list.
Core opinions
.loopwright/in the host repo, and the engine's defaults are frozen. Changing what the gate demands is a PR in the host repo.Why the contested branches settled where they did
P2 absorbed review, and P8 took merge out of P2. The original text said
"merging is always a human decision" as a core opinion. It is not core — it is a
default, and a configurable one, in the spirit of the Claude Code permission
modes. What is core is the rule above it: nothing may remove the human from an
irreversible action silently. So the authority claim moved to P8, and P2 kept the
part about what "done" is. "Done" includes resolved review findings, because
that is what
/babysit-pralready enforces — a doc that said otherwise would bedescribing a harness that does not exist, and would authorise autonomous merge
over an unresolved review, which is the most elegant way to buy green.
Note the division of labour: "only
quality-gate.mjscan fail the build" is CImechanics and stays in
docs/loopwright/quality-gate.md. P2 is about who decides"done", which is a different claim.
P7 merged two decisions and excludes the baseline. Engine defaults are
frozen: a loopwright update never moves a host's verdict. This closes #8's
first risk ("silent gate tightening through defaults") and settles its open
question 1 — overlay is safe because the defaults are frozen; it is the freeze,
not a full copy, that removes the risk. The model is
tsconfig.json: the hostwrites only its overrides, the defaults live in the tool, and they do not move
without a breaking release.
baseline.jsonis not policy and is not the fileP7 names: it is measured state, machine-written, engine-schema'd, host-committed,
never hand-edited. P3 owns it.
P8 is the one place this document changes a rule.
docs/loopwright/loop-harness.md:66states that merging is always the human's decision, absolutely. P8 makes it
opt-in-automatable. Enumerated irreversibles today: merge, publishing a tag or
release, deploy — the list is not exhaustive. Everything else the harness
automates is reversible and therefore needs no opt-in: closing an issue reopens,
gh pr readydespromotes, comments edit. The one irreversible thing the harnesstouches is merge, and that is the one it does not do.
A consequence, recorded because it contradicts a published RFC: the mediator
may decide the version bump, since a bump lives in a file inside the PR and the
human sees it in the same act as the merge. #9 states "The bump type is a human
decision... Agents never choose it"; it adjusts when it is grilled.
P12 was already settled by #6, which lists "the metric vocabulary and its
semantics" among what core keeps. A key a connector invents is persisted into
the baseline (
quality-gate.mjswritesmetrics: currentwholesale) but neverscored, because evaluation iterates over the config, not over the reports — a
metric that looks measured and is not, which is exactly what P4 exists to
prevent. The connector contract is three-valued: a value,
unconfigured(warns forever, never blocks), or
failed(always blocks).The limit of autonomy is deliberately not declared. Work intake moves from
hand-written RFCs to GitHub labels to an alert webhook, and a Datadog alert
driving investigate → fix → deploy must stay possible without this document
passing judgement on it. What is not deferred: an issue and a PR exist before
anything reaches production, because P1 requires the artifact — without it there
is nowhere for the diagnosis to live, and a fix deployed without a trail is green
without evidence, in the most expensive place possible.
Detail territory
Every entry is a territory, a named seam, and the current default. This RFC does
not define any of these contracts — each belongs to its own RFC.
connectorjs(#6)toolchain providermise, never required (#6)runtimeagent-hostreleaseintakeDefaults are opinions too. Every territory ships with one recommended
default, wired out of the box, the way framework starters ship eslint. A default
is never a requirement: there is always a documented way to swap it.
T5 is a contract of exactly four capabilities the loop requires of its agent
host: dispatch a subagent with fresh context and an explicit payload; restrict
tools per role; choose a model tier per task; an external reviewer that can write
on the PR. Anything the loop needs beyond those four is leakage — it either
enters the contract as a core change, or it is a bug. This answers the original
open question 1 with a measurable test instead of a judgement call. The engine
itself has zero coupling here: nothing under
.loopwright/scripts/referencesClaude, agents, skills or sessions. Capability 2 is load-bearing for P10, because
a reviewer is read-only by being handed no write tools, not by being told to
behave.
T6 is independent of T1. CI written in a different language than the project
is ordinary, so a Go host needing Node in CI is not coupling. #6's open question
4 ("should the engine ship as a single binary?") is therefore a deferred
detail, not a question of identity.
T7 works on a non-JS host because of T6: the engine is already Node and
already has
.loopwright/node_modules/, so changesets can be internal toloopwright and still serve a Go host. If the host already uses changesets,
loopwright uses the host's.
Classification rule
If a change alters what counts as done, or how work flows, it is core and
must be argued against the opinions above. If it alters how a fact is measured
or where a step runs, it belongs behind a contract, with a default
implementation.
Scope
docs/loopwright/principles.md.docs/loopwright/loop-harness.md:44-77loses the "Rules of the loop" that areopinions — they are deleted there and replaced by a reference, since the new
doc is their only author. What stays is harness mechanics: base branch,
closing keywords, draft-until-ready, and the two grill-phase notes at
:67-72("architecture is scaffolding, not a spec") and
:73-77("boundary-revieweris the first thing to cut"). Line 66 is reconciled with P8.
README.md,CLAUDE.md,.loopwright/claude-md-section.md(the most important of the four — it is whatevery host repo reads), and
package.json:4. The "How it's wired" table atREADME.md:158-169gains a row.README.md:142-143is the second site of theabsolute merge rule — "Two rules hold the whole thing together" — and is
reconciled with P8 here, so that
README.mdhas a single owner.are the facts, the doc adjusts to them. No classification line is added to any
RFC.
gh issue viewdoes not return comments and a comment would be invisible tothe future
/grill-rfcof those RFCs. This is a close-out act on this RFC, nota task: it produces no repository diff, so paying a branch, a spec commit, a
draft PR and a gate run for it would be overhead for nothing.
Non-goals
.loopwright/scripts/or.claude/behaviour changes.RFC: Stack connectors, decouple the quality gate from JS/TS #6, T3 to RFC: Isolated agent runtime with a mediated, async channel #7, T7 to RFC: Versioned releases with changesets, and automatic changesets in the loop #9, T8 to a later RFC.
what it breaks — RFC: One config file, YAML, sectioned by feature #8 already inventories its own 44 references to
config.json.divergence from P5 (see Risks).
Risks
mandatory
rejects:line: each opinion has to name a concrete change itrefuses, which turns a slogan into a test.
deliberately: the doc is the newer decision and RFC: Versioned releases with changesets, and automatic changesets in the loop #9 adjusts when grilled. The
body pointer is what makes the contradiction visible rather than silent.
the decided state, not the current one, and the knob it assumes does not exist
yet.
quality-gate.mjs:97-99usesCOLLECTOR_METRICS[collector] ?? [], so acollector outside the hardcoded map that fails maps to zero metric ids,
contributes nothing to the worst status, and the gate passes. Recording it in
the doc was considered and rejected; fixing it would be a code change, which
this RFC excludes.
keeps growing, it stops being read, and an unread principle decides nothing.
Task breakdown (link sub-issues)
To be reconciled by
/grill-rfconce the sub-issues are created.Open questions
alert does not sit comfortably inside an RFC. The flow is core for now; a
second flow for incident-shaped work is deferred, not rejected.
document says where a human must remain. That is a decision to revisit only if
an actual problem appears, not a gap to fill pre-emptively.
config.ymlvssettings.yml) belongs to RFC: One config file, YAML, sectioned by feature #8.This doc only asserts "one file, under
.loopwright/".Refined by
/grill-rfc