Skip to content

RFC: Core principles, opinionated about process, agnostic about stack #5

Description

@SamuelDenani

Summary

Write docs/loopwright/principles.md: a descriptive statement of what
loopwright is opinionated about and what it treats as a swappable detail. It
changes no engine code and no workflow, and it gates nothing. Its job is to give
a name to the boundary that three open RFCs already cite, so that boundary stops
being re-argued in every design session.

The core opinions grow from six to twelve, each carrying the concrete change it
rejects. The detail territory grows from four entries to eight, each with a named
seam and its current default.

Context

Loopwright exists because of an opinion about how work flows and how "done" is
judged
, not because of a stack. Today its identity is written through the stack
anyway — README.md:3-6 leads with "vendor it into any JS/TS repo",
package.json:4 says "AI-assisted JS/TS repos", CLAUDE.md:1-6 describes the
project by the directory its engine lives in and the command that runs its tests,
and .loopwright/claude-md-section.md — the block install.sh injects into every
host repo — leads with "vendored under .loopwright/" before naming anything the
layer is opinionated about.

The destination is a near-autonomous loopwright, done properly rather than for
hype
. That reframes every opinion below: they are not brakes on autonomy, they
are what makes autonomy not be theatre. An agent that merges on its own is only
acceptable if "done" is a mechanical verdict that cannot be bought — so the gate
as sole authority, the integrity signals and the ratchet are preconditions of
autonomy, not constraints on it.

Three downstream RFCs already depend on this document textually:

They do not need the document to enforce anything. They need it to exist and to
name things, so they can cite it. That is why this RFC is descriptive.

Decision

The document is descriptive, not normative

docs/loopwright/principles.md. No skill consults it, no gate checks it, no RFC
is required to carry a classification line. Rejected alternative: a normative
version where /grill-rfc demands a classification line from every future RFC.
It would have meant editing a skill, and the value it adds is small — what the
downstream RFCs need is the vocabulary, which prose delivers.

"Core" means the rule has no knob, not that the rule is frozen and not that
the action it governs has no knob. An action may be configurable when the rule
itself provides for it (see P8). Changing a core opinion is an explicit
amendment: a dated entry naming the RFC that changed it, in an Amendments
section. Ids are stable and amendments never renumber; a withdrawn opinion stays
in place as P4 (withdrawn, RFC #N).

Form. Each opinion carries three things: the assertion, a mandatory
rejects: line naming a concrete change it refuses, and a pointer to the file
that enforces it — file name only, never a line number, and an empty pointer
where nothing enforces it. An empty pointer is itself a finding. Rationale goes
in prose below the list.

Core opinions

id Opinion
P1 Work flows RFC issue → task sub-issues → PRs. State lives in artifacts, never in session memory. The trigger is detail; the flow is core.
P2 "Done" is determined by configured checks, never by the judgement of whoever merges. Merge — autonomous or not — requires every configured check green for the current diff and review findings resolved.
P3 Quality is a ratchet against a committed baseline, with per-file debt grandfathered. The baseline is never regenerated to hide a regression.
P4 Green earned vs bought — the integrity signals — is a first-class concern, not an add-on.
P5 Infrastructure failure blocks. Disabling a configured check is itself a violation.
P6 Agents communicate and coordinate through a mediator, never peer to peer.
P7 All loopwright policy lives in one versioned file under .loopwright/ in the host repo, and the engine's defaults are frozen. Changing what the gate demands is a PR in the host repo.
P8 Irreversible actions require a human by default. Automating one is an explicit opt-in in the host's config file.
P9 Strict TDD: the test comes before the implementation.
P10 Fresh-context review: every diff is reviewed by an agent that did not write it.
P11 A unit of work is a vertical slice, never a horizontal layer.
P12 The metric vocabulary and its semantics are core and closed. Connectors measure; they never invent a key.

Why the contested branches settled where they did

P2 absorbed review, and P8 took merge out of P2. The original text said
"merging is always a human decision" as a core opinion. It is not core — it is a
default, and a configurable one, in the spirit of the Claude Code permission
modes. What is core is the rule above it: nothing may remove the human from an
irreversible action silently. So the authority claim moved to P8, and P2 kept the
part about what "done" is. "Done" includes resolved review findings, because
that is what /babysit-pr already enforces — a doc that said otherwise would be
describing a harness that does not exist, and would authorise autonomous merge
over an unresolved review, which is the most elegant way to buy green.

Note the division of labour: "only quality-gate.mjs can fail the build" is CI
mechanics and stays in docs/loopwright/quality-gate.md. P2 is about who decides
"done", which is a different claim.

P7 merged two decisions and excludes the baseline. Engine defaults are
frozen: a loopwright update never moves a host's verdict. This closes #8's
first risk ("silent gate tightening through defaults") and settles its open
question 1 — overlay is safe because the defaults are frozen; it is the freeze,
not a full copy, that removes the risk. The model is tsconfig.json: the host
writes only its overrides, the defaults live in the tool, and they do not move
without a breaking release. baseline.json is not policy and is not the file
P7 names: it is measured state, machine-written, engine-schema'd, host-committed,
never hand-edited. P3 owns it.

P8 is the one place this document changes a rule. docs/loopwright/loop-harness.md:66
states that merging is always the human's decision, absolutely. P8 makes it
opt-in-automatable. Enumerated irreversibles today: merge, publishing a tag or
release, deploy
— the list is not exhaustive. Everything else the harness
automates is reversible and therefore needs no opt-in: closing an issue reopens,
gh pr ready despromotes, comments edit. The one irreversible thing the harness
touches is merge, and that is the one it does not do.

A consequence, recorded because it contradicts a published RFC: the mediator
may decide the version bump
, since a bump lives in a file inside the PR and the
human sees it in the same act as the merge. #9 states "The bump type is a human
decision... Agents never choose it"
; it adjusts when it is grilled.

P12 was already settled by #6, which lists "the metric vocabulary and its
semantics"
among what core keeps. A key a connector invents is persisted into
the baseline (quality-gate.mjs writes metrics: current wholesale) but never
scored, because evaluation iterates over the config, not over the reports — a
metric that looks measured and is not, which is exactly what P4 exists to
prevent. The connector contract is three-valued: a value, unconfigured
(warns forever, never blocks), or failed (always blocks).

The limit of autonomy is deliberately not declared. Work intake moves from
hand-written RFCs to GitHub labels to an alert webhook, and a Datadog alert
driving investigate → fix → deploy must stay possible without this document
passing judgement on it. What is not deferred: an issue and a PR exist before
anything reaches production, because P1 requires the artifact — without it there
is nowhere for the diagnosis to live, and a fix deployed without a trail is green
without evidence, in the most expensive place possible.

Detail territory

Every entry is a territory, a named seam, and the current default. This RFC does
not define any of these contracts — each belongs to its own RFC.

Territory Seam Default today
T1 Host language/ecosystem and the tools that measure it connector js (#6)
T2 Stack detection and toolchain installation toolchain provider mise, never required (#6)
T3 Where and how agents execute runtime local session; container per #7
T4 Thresholds and tolerances — already config-driven
T5 Agent host agent-host Claude Code
T6 Engine implementation language — Node
T7 Release tool release changesets, internal, host-first
T8 How work enters the pipeline intake hand-written RFC

Defaults are opinions too. Every territory ships with one recommended
default, wired out of the box, the way framework starters ship eslint. A default
is never a requirement: there is always a documented way to swap it.

T5 is a contract of exactly four capabilities the loop requires of its agent
host: dispatch a subagent with fresh context and an explicit payload; restrict
tools per role; choose a model tier per task; an external reviewer that can write
on the PR. Anything the loop needs beyond those four is leakage — it either
enters the contract as a core change, or it is a bug. This answers the original
open question 1 with a measurable test instead of a judgement call. The engine
itself has zero coupling here: nothing under .loopwright/scripts/ references
Claude, agents, skills or sessions. Capability 2 is load-bearing for P10, because
a reviewer is read-only by being handed no write tools, not by being told to
behave.

T6 is independent of T1. CI written in a different language than the project
is ordinary, so a Go host needing Node in CI is not coupling. #6's open question
4 ("should the engine ship as a single binary?") is therefore a deferred
detail, not a question of identity.

T7 works on a non-JS host because of T6: the engine is already Node and
already has .loopwright/node_modules/, so changesets can be internal to
loopwright and still serve a Go host. If the host already uses changesets,
loopwright uses the host's.

Classification rule

If a change alters what counts as done, or how work flows, it is core and
must be argued against the opinions above. If it alters how a fact is measured
or where a step runs
, it belongs behind a contract, with a default
implementation.

Scope

  • New docs/loopwright/principles.md.
  • docs/loopwright/loop-harness.md:44-77 loses the "Rules of the loop" that are
    opinions — they are deleted there and replaced by a reference, since the new
    doc is their only author. What stays is harness mechanics: base branch,
    closing keywords, draft-until-ready, and the two grill-phase notes at :67-72
    ("architecture is scaffolding, not a spec") and :73-77 ("boundary-reviewer
    is the first thing to cut"). Line 66 is reconciled with P8.
  • Identity rewritten principles-first in four places: README.md, CLAUDE.md,
    .loopwright/claude-md-section.md (the most important of the four — it is what
    every host repo reads), and package.json:4. The "How it's wired" table at
    README.md:158-169 gains a row. README.md:142-143 is the second site of the
    absolute merge rule — "Two rules hold the whole thing together" — and is
    reconciled with P8 here, so that README.md has a single owner.
  • A one-way consistency test against RFC: Stack connectors, decouple the quality gate from JS/TS #6–RFC (draft): Interactive installer for first-time setup #10 while the doc is written: the RFCs
    are the facts, the doc adjusts to them. No classification line is added to any
    RFC.
  • A pointer to the doc appended to the body of RFC: Stack connectors, decouple the quality gate from JS/TS #6–RFC (draft): Interactive installer for first-time setup #10 — not a comment, since
    gh issue view does not return comments and a comment would be invisible to
    the future /grill-rfc of those RFCs. This is a close-out act on this RFC, not
    a task: it produces no repository diff, so paying a branch, a spec commit, a
    draft PR and a gate run for it would be overhead for nothing.

Non-goals

Risks

  • Principles written too broadly stop deciding anything. Mitigated by the
    mandatory rejects: line: each opinion has to name a concrete change it
    refuses, which turns a slogan into a test.
  • The doc contradicts RFC: Versioned releases with changesets, and automatic changesets in the loop #9 on the version bump the day it lands. Accepted
    deliberately: the doc is the newer decision and RFC: Versioned releases with changesets, and automatic changesets in the loop #9 adjusts when grilled. The
    body pointer is what makes the contradiction visible rather than silent.
  • P8 is the only non-descriptive line in a descriptive document. It describes
    the decided state, not the current one, and the knob it assumes does not exist
    yet.
  • An engine divergence from P5 is known and not being fixed here.
    quality-gate.mjs:97-99 uses COLLECTOR_METRICS[collector] ?? [], so a
    collector outside the hardcoded map that fails maps to zero metric ids,
    contributes nothing to the worst status, and the gate passes. Recording it in
    the doc was considered and rejected; fixing it would be a code change, which
    this RFC excludes.
  • Twelve opinions is near the limit of what stays memorable. If the list
    keeps growing, it stops being read, and an unread principle decides nothing.

Task breakdown (link sub-issues)

To be reconciled by /grill-rfc once the sub-issues are created.

Open questions

  1. More flows for the autonomous pipeline. A one-line fix originating from an
    alert does not sit comfortably inside an RFC. The flow is core for now; a
    second flow for incident-shaped work is deferred, not rejected.
  2. The limit of autonomy is intentionally undeclared. Nothing in this
    document says where a human must remain. That is a decision to revisit only if
    an actual problem appears, not a gap to fill pre-emptively.
  3. The config file's name (config.yml vs settings.yml) belongs to RFC: One config file, YAML, sectioned by feature #8.
    This doc only asserts "one file, under .loopwright/".

Refined by /grill-rfc

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    rfcRFC: top-level design and intent

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions