Skip to content

RFC: Isolated agent runtime with a mediated, async channel #7

Description

@SamuelDenani

Summary

Give each task orchestrator a fully isolated sandbox (its own clone of the repo inside its own container) and a single, mediated, asynchronous channel to coordinate with the others. Isolation is the wall; the channel is the only door, and loopwright is the one holding it. Git worktrees are dropped as an isolation mechanism. Remote execution (GitHub triggering work on my machine) is a later phase, not this RFC.

Motivation

/execute-issue already parallelizes unblocked tasks as background orchestrators, each in its own git worktree. In practice worktrees have been a recurring source of bugs when agents work in them, and even when they behave, they only isolate files and branches. My worst problems running several scopes at once came exactly from not having time to isolate each one properly.

What the current worktree setup does not give:

  • Independence: worktrees share one .git (refs, config, hooks, locks, branch checkout rules), so one sandbox's git state can leak into or block another.
  • Environment: ports, local databases, caches, dev servers, global tool state.
  • Secrets and credentials available to the session.
  • Network and filesystem outside the worktree.
  • Awareness: parallel orchestrators cannot tell each other "I changed this contract" except by colliding later in review or at merge.

Goals / Non-goals

Goals:

  • A runtime abstraction for "where a task executes", replacing direct worktree creation in /execute-issue.
  • Each sandbox is a separate clone of the repo, with no git state shared with the host checkout or other sandboxes.
  • Sandboxes run in a container with their own env, ports, network and secrets scope.
  • A mediated, async message channel between orchestrators, with every message logged.
  • Messages that live in artifacts, consistent with the principle that state never lives in session memory.

Non-goals:

  • Keeping git worktrees as a supported isolation level.
  • Remote execution triggered by GitHub (issue movement driving runs on my machine). Different way of working; separate RFC once this lands.
  • Peer-to-peer agent chat.
  • Synchronous request/response between agents.

Proposed approach

The wall (runtime): execute-issue asks a runtime for a sandbox instead of creating a worktree. A sandbox is a container holding its own clone of the repo on the task's branch. The only way work leaves it is by pushing that branch; the host checkout is never touched. The runtime is a detail behind a contract, per the principles RFC, so the container technology can change without touching the loop.

Definition format: plain Docker. The sandbox environment is loopwright's layered Dockerfile plus an entry script, driven directly through the Docker CLI. Inspecting a sandbox is docker exec -it <sandbox> bash (any editor, including Neovim, from there); previewing an app an agent started is a published port. No editor-specific tooling is assumed.

Startup and caching: build once, rebuild only when an input changes. The environment definition lives in .loopwright/runtime/ (versioned); its artifacts do not (images live in Docker, the mirror in a gitignored cache).

  • Image layers: loopwright base, then toolchain, then project dependencies. Each layer is tagged by the hash of its inputs, so a change to the lockfile rebuilds only the dependency layer, and a toolchain version bump rebuilds only from the toolchain layer down.
  • The toolchain layer is installed by the toolchain provider from the connectors RFC. With the default provider it is a mise install keyed by the hash of the tool declaration, using the same mise version loopwright pins for the host (the image fetches that version's Linux build, since the host's bundled binary may be for another OS); with native or custom the same layer runs that provider's install instead.
  • A bare mirror of the repo in the local cache, fetched before each run. Each sandbox clones from it with --reference, so a per-task clone is near instant and costs little disk.
  • Per task: start a container from the ready image, clone from the mirror onto the task's branch, run. If the task branch changes the lockfile, the sandbox runs an incremental install instead of rebuilding the image.

The door (channel): the orchestrator of the RFC run is the mediator. Task orchestrators post messages to it, never to each other. The mediator decides what to forward, to whom, and whether something needs a human.

The log: GitHub is already the state store, so the channel's durable form is issue/PR comments with a machine-readable header. A local mailbox (SQLite or a directory) may sit in front for speed, but GitHub stays the source of truth.

Alternatives considered

  • Keep worktrees as the default and add containers on top. Rejected. Worktrees have been the least reliable part of the parallel flow, and sharing one .git across sandboxes is the opposite of isolation.
  • Separate clones without containers. Fixes the git sharing, but leaves env/port/secret collisions. Could be a lightweight fallback level if containers prove too slow.
  • Peer-to-peer messaging between agents. Rejected. No single place to audit, and it conflicts with the mediated principle.
  • devcontainer as the definition format. Its main gains (attaching a configured editor, auto-forwarded ports) are VS Code features, so they add nothing to a terminal/Neovim workflow, while it adds a CLI layer and is built for one long-lived environment per folder rather than many parallel short-lived ones. Deferred, not rejected: since the runtime sits behind a contract, devcontainer can arrive later as a runtime adapter, when either (a) the remote RFC chooses GitHub Codespaces (which runs devcontainer definitions), or (b) a host asks to base sandboxes on the devcontainer.json it already has.
  • For the later remote phase: two options to start from. A GitHub self-hosted runner on my machine instead of an inbound tunnel (it polls GitHub, so no port is exposed), or Codespaces through a devcontainer adapter, with no machine of mine involved.

Risks

  • Containers add startup time and toolchain duplication per task.
  • Cache invalidation: a missed input in a layer's hash means a stale image that silently runs old tools. Layer hashes must be computed from an explicit input list, not guessed.
  • The runtime depends on the toolchain provider contract from the connectors RFC. Until it lands, the runtime can ship with a fixed mise layer.
  • A chatty channel turns into noise; the mediator needs a clear policy for what gets forwarded.
  • Shared resources that genuinely must be shared (one database, one external API sandbox) need an explicit answer, not an accidental one.

Task breakdown (link sub-issues)

To be produced by /grill-rfc.

Open questions

  1. What do agents actually need to say to each other? Not decided yet. Candidates: contract-change notices ("schema X changed"), requests to another scope, handoffs. This shapes the message schema and should be settled first in the grill.
  2. How the agent tool runs inside the container (credentials, session, permissions).
  3. Where the image cache lives when the runner is not my machine (CI, a remote runner later).
  4. Port and resource allocation policy across sandboxes.
  5. Mediator policy: when does a message get forwarded, held, or escalated to the human?
  6. Does the local mailbox exist in phase 1, or is GitHub alone enough?

Generated by Claude Code

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    rfcRFC: top-level design and intent

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions