Skip to content

Repository files navigation

clean-refactor-harness

Read this in: English · 한국어

A miner picks apart a pile of legacy spaghetti code (auth.py, db.py, main.py…) and turns it into a crystalline clean / hexagonal architecture — domain core wrapped in ports and adapters — stamped with a 'behavior preserved' machine-checked certificate.

A standalone, general-purpose AI harness that converts a spaghetti legacy codebase into clean / hexagonal architecture while preserving its captured behavior, and proves it with a machine-checked certificate. It runs largely unattended.

Honest claim model. The harness never claims a blanket "behavior preserved." It claims "captured behavior preserved + residual = X", where X is an explicit list of the effects it could not put under an oracle. What it could not observe is reported up front, alongside what it preserved.

This repo is the durable product. Any target repository it converts is a disposable test subject, used in an isolated git worktree to check that the harness itself behaves, never the other way around.


The core idea

Restructuring code is easy; doing it safely is not. The risk is silent meaning loss: a refactor that keeps the return values identical but drops a retry, reorders a transaction, or loosens an authorization check. This harness preserves observable behavior. It captures that behavior as an executable oracle before touching structure, then gates every transformation against the oracle until the code reaches the target architecture.

"Observable behavior" is defined precisely (charter §1) as six categories:

  1. Response contract — return values, status, body, serialization, field order.
  2. Error contract — exception types, error status/messages, validation order, partial failure, retryable vs terminal, rollback.
  3. Side effects — DB writes, emitted events, queue enqueues/order, email/SMS, cache mutations, file writes, external API calls (args/count/order), audit logs, metrics.
  4. Transaction / concurrency — transaction boundaries, locks, isolation level, idempotency keys, retry/backoff, ordering / race windows, timeouts.
  5. Authz / security — placement and order of permission checks, tenant isolation, PII exposure boundaries, auth flow, 404-vs-403 information leakage.
  6. Time / nondeterminism — clock, random seed, schedule/cron, debounce/throttle.

A conversion is only "done" when every unit reaches its target ring, every captured surface passes its oracle on an environment-equivalent green (not local-only), and a final A/B of baseline-vs-result is identical (or its diff is an accepted, recorded residual).


Architecture: two layers + one interface

[1] SPEC  (markdown under skills/clean-refactor/ — the driver follows it literally, the "brain")
      charter.md ....... the constitution / single source of truth for definitions
                         (§0 product+arch principles, §1 behavior taxonomy, §2 the three
                          gates + machine gate, §6a four rings + §6b DDD mapping, §0.4
                          autonomy boundary)
      00-bootstrap ..... Phase A-0: runnable-baseline gate (env-equivalent green required)
      01-discover ...... Phase A-1: find effect seams
      02-classify ...... Phase A-2: DDD classify + 4-ring placement   <-- human approval gate
      03-capture ....... Phase B-0: capture behavior goldens (6 categories + joint invariants)
      04-transform ..... Phase B-1: extract per ring (strangler-fig order)
      05-drive ......... Phase B-2: the loop, machine gates, A/B, certificate
                         (also: the state.json schema SSOT)
      06-supervise ..... the supervisor procedure (doubt + fix the harness itself)
      07-decompose ..... opt-in post-certificate pass: split oversized same-ring modules
                         (move-only + shim, gate-green per cycle, certified_head re-bound)
      08-install-guardrails
                         opt-in post-certificate pass: install .claude/ AI-session guardrails
                         (edit-time lint hook + landing-gate wiring + pointer rules; merge-only)

           driver: reads & follows          supervisor: doubts & fixes
                              |
                              v
[*] state.json  ........ the single durable interface both layers read/write.
                         Enables crash-resume and park-and-resume. Holds: baseline_pin,
                         seams, contexts, roadmap (per-unit ring + gates), dod, blockers,
                         accepted_risks, joint_invariants.
                              ^
                              |
[2] ORCHESTRATOR  (python — the deterministic unattended loop, the "body")
      run.py ........... CLI entry point
      supervisor.py .... the loop: run_until_done — respawn driver, route outcomes, and
                         validate the terminal (certificate + reconciles + terminal rescan)
      agent_steps.py ... the cognition boundary (Protocol: run_driver_cycle / classify /
                         propose_fix) + the honest "not wired" default (ManualAgentSteps)
      llm_agent_steps .. the model-backed implementation (shells out to the `claude` CLI)
      guards.py ........ deterministic safety: Tier-A/B classification, oscillation guard,
                         convergence tracker — the part a wrong LLM judgment cannot bypass
      gates.py ......... external gates: codex adversarial review, env-equivalent machine
                         gate, §S3 mutation-proof, dependency gate, terminal rescan
      dependency_gate .. in-process scanner: dependency-rule edges + placement hygiene
                         (the machine teeth behind the four-ring rule — no self-report)
      certificate.py ... §D10 certificate status + disclosure/placement/A-B reconciles
      verify_certif. ... the LANDING gate: blocks governed changes whose HEAD the
                         certificate does not bind (certified_head)
      ops.py ........... git / filesystem plumbing (forensic bundles, worktree reset)

Two design choices

  • Driver / supervisor role separation (charter §0.4). The driver may only follow the spec literally; it has no power to edit the harness or weaken a gate. The supervisor (06-supervise) is the only role that may doubt and fix the harness. Separating the follower from the doubter stops the follower from rationalizing its way around an inconvenient gate (self-confirmation bias).

  • Guards vs gates. Guards are in-process and deterministic; they cannot be argued with (e.g. a self-edit that would weaken a gate's detection power is auto-classified Tier-B and blocked). Gates are external and adversarial (codex review; an environment-equivalent machine gate enforcing effect-diff = 0). All cognition flows through the agent_steps boundary, so even a wrong model judgment is wrapped by the deterministic guard layer.


The target shape: four rings

Clean / hexagonal architecture drawn as concentric rings. The dependency rule is the whole point — every dependency points inward, and an inner ring may never import an outer one:

+-----------------------------------------------------------------------------+
|  framework (0)  --  HTTP, CLI, cron -- the outside world                    |
|                                                                             |
|    +-------------------------------------------------------------------+    |
|    |  adapter (1)  --  DB, queues, vendor SDKs, port implementations   |    |
|    |                                                                   |    |
|    |    +---------------------------------------------------------+    |    |
|    |    |  usecase (2)  --  application workflows, orchestration  |    |    |
|    |    |                                                         |    |    |
|    |    |       +-----------------------------------------+       |    |    |
|    |    |       |  domain (3)  --  entities, invariants,  |       |    |    |
|    |    |       |  pure business rules (no I/O)           |       |    |    |
|    |    |       +-----------------------------------------+       |    |    |
|    |    +---------------------------------------------------------+    |    |
|    +-------------------------------------------------------------------+    |
+-----------------------------------------------------------------------------+

Read it as framework(0) -> adapter(1) -> usecase(2) -> domain(3), dependencies inward only. DDD classification maps to rings: Core → domain, Supporting → usecase, Generic → adapter (or buy, which is out of scope).


Flow

Phase A (understand)              [human]                Phase B (transform)
bootstrap -> discover -> classify -> approve -> capture -> transform -> drive -> CERTIFICATE
                          (+env attest)  (placement)        (per ring)  (gates·A/B)
   \________________ 06-supervise wraps the whole thing as the unattended loop _____________/

Human touchpoints are exactly the legitimate ones, and every halt is a recoverable park-and-resume checkpoint (charter §0.4):

  • Placement approval (§2.2) — the one designed human-judgment gate: classify + 4-ring placement. Proposal → discussion → explicit approval.
  • Environment attestation (00-bootstrap S3) — only a human touch when there is no CI: the green source must be CI / env-equivalent sandbox / human-attested-with-evidence.
  • Hard blockers (§0.4) — missing secret/DB, env-parity loss, retry-budget exhaustion, oscillation: park, surface to a human, resume from the recorded step once provisioned.

Install as a Claude Code plugin

The harness ships as a Claude Code plugin so you can invoke it in-session instead of running the orchestrator by hand. Install it from this repo:

/plugin marketplace add tryumanshow/clean-refactor-harness
/plugin install clean-refactor-harness@tryumanshow

Those two commands are one-time setup. After that you don't type a command to use the harness. Just describe the task in natural language and Claude auto-invokes the clean-refactor skill by matching your request to its description.

For example, once installed you just type a request like this to Claude:

Convert ./backend to clean / hexagonal architecture without changing its behavior.
Run the clean-refactor harness: work in a disposable git worktree off the current
commit, capture characterization (golden) tests before any structural change, and
pause at the placement-approval gate for my review.

Claude matches that to the skill's description and runs the harness. The /clean-refactor slash command is also available as an optional explicit trigger (then describe the target), but it is not required.

The plugin is a thin wrapper. It changes how you invoke the harness, while what the harness does stays identical. The skill (skills/clean-refactor/SKILL.md) only routes to the existing charter.md + 00–06 phase specs and the Python orchestrator; it does not restate the methodology. The behavior-preservation guarantees still live entirely in those specs and in the deterministic guards/gates, so a plugin run and a direct python run.py run execute the same procedure. The only thing the plugin adds is discoverability and the /clean-refactor entry point.

Running it (the orchestrator directly)

Dependencies: the claude CLI (drives the cognitive steps headless) and the codex CLI (adversarial gate). The deterministic core (guards.py, gates.py, supervisor.py) runs on the standard library alone.

cd orchestrator/harness_supervisor

# Unattended (model-backed) — converts a target worktree until terminal:
python run.py \
  --harness  /path/to/this/repo/skills/clean-refactor \
  --worktree /path/to/disposable/target-worktree \
  --baseline <immutable_baseline_git_sha> \
  --max-cycles 30 --driver-timeout 600

# Wire-up smoke (no model calls — pauses at the cognition boundary to a human):
python supervisor.py

Swap ManualAgentSteps (the honest "not yet wired" default, which parks every cognitive step to a human) for LlmAgentSteps to run unattended; run.py does this for you.

Tests

cd orchestrator/harness_supervisor
python -m pytest -q     # 222 passed
ruff check .            # clean

The suite covers the deterministic guards (Tier classification incl. constitution-edit protection, oscillation fingerprinting, convergence), the supervisor control flow (fail-safe routing, single-classify, apply-failure handling, env-blocker parking, placement-gate routing), the run-ordering terminal (certificate status, disclosure/placement/A-B reconciles, and the terminal rescan that re-measures the actual tree instead of trusting driver-recorded numbers), the dependency/placement scanner, the landing gate, and the model-backed step parsing / fail-safe.


Status

Experimental, with one real-scale conversion behind it. The pipeline was first demonstrated end-to-end on a small target (a single-file stdlib CLI): spaghetti → clean four-ring architecture across eight strangler-fig commits, with behavior preserved (final A/B AB_IDENTICAL, golden + fitness suites green) and a cross-unit joint invariant captured and verified.

It has since driven a production FastAPI backend (~110k LOC Python) through the same procedure: full four-ring conversion with the dependency rule CI-enforced by import-linter contracts, a later package-by-feature re-layout, and a 9-cycle 07-decompose pass that split a 5,567-line usecase module down to 1,398 lines with every cycle gate-green. Cost/time per cycle and multi-context generality beyond that one codebase remain not yet systematically measured.

The spec files (charter.md, 00–07) are written in Korean (the author's working language); the orchestrator code and this README are in English. The deterministic guarantees live in the code and are language-invariant; only the cognitive layer reads the prose spec.

Direction (where this is headed)

The end state this is built toward. Each item is the trajectory implied by the gaps the Status section names — not a wishlist of unrelated features:

  • Drive the residual to zero. Today the honest claim is "captured behavior preserved + residual = X". The work is to shrink X by widening oracle coverage across all six behavior categories — especially the hardest to capture (concurrency / ordering, time / nondeterminism, authz boundaries) — so fewer effects ever fall outside the oracle.
  • Prove it beyond one codebase. One production backend has been converted end-to-end; the target is generality — more codebases, more bounded contexts, real external effects, and measured cost / time per cycle — the numbers Status calls "not yet systematically measured".
  • More unattended, the same single human gate. Keep placement approval (§2.2) as the one designed human-judgment touchpoint and push everything else hands-off — stronger machine gates, crash-resume, park-and-resume — so a run needs a human only where judgment is genuinely required.
  • Honest by construction. The certificate never upgrades to a blanket "behavior preserved"; it always ships the enumerated residual. The direction strengthens the gates, never loosens the claim.

Layout

.claude-plugin/              plugin + marketplace manifests (Claude Code packaging)
skills/clean-refactor/       the plugin skill — the repo's spine:
    SKILL.md                   entry point (routes to the spec + orchestrator)
    charter.md, 00–08.md       the spec (what to do; 07 decompose + 08 install-guardrails are opt-in)
    docs/                      design provenance: dry-run audits + the dogfooding gap log
orchestrator/harness_supervisor/   companion orchestrator (unattended runs) + tests
README.md, README.ko.md      this

About

AI harness that refactors legacy spaghetti code into clean/hexagonal architecture while preserving observable behavior — proven by a machine-checked certificate. Runs largely unattended.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages