Read this in: English · 한국어
A standalone, general-purpose AI harness that converts a spaghetti legacy codebase into clean / hexagonal architecture while preserving its captured behavior, and proves it with a machine-checked certificate. It runs largely unattended.
Honest claim model. The harness never claims a blanket "behavior preserved." It claims "captured behavior preserved + residual = X", where X is an explicit list of the effects it could not put under an oracle. What it could not observe is reported up front, alongside what it preserved.
This repo is the durable product. Any target repository it converts is a disposable test subject, used in an isolated git worktree to check that the harness itself behaves, never the other way around.
Restructuring code is easy; doing it safely is not. The risk is silent meaning loss: a refactor that keeps the return values identical but drops a retry, reorders a transaction, or loosens an authorization check. This harness preserves observable behavior. It captures that behavior as an executable oracle before touching structure, then gates every transformation against the oracle until the code reaches the target architecture.
"Observable behavior" is defined precisely (charter §1) as six categories:
- Response contract — return values, status, body, serialization, field order.
- Error contract — exception types, error status/messages, validation order, partial failure, retryable vs terminal, rollback.
- Side effects — DB writes, emitted events, queue enqueues/order, email/SMS, cache mutations, file writes, external API calls (args/count/order), audit logs, metrics.
- Transaction / concurrency — transaction boundaries, locks, isolation level, idempotency keys, retry/backoff, ordering / race windows, timeouts.
- Authz / security — placement and order of permission checks, tenant isolation, PII exposure boundaries, auth flow, 404-vs-403 information leakage.
- Time / nondeterminism — clock, random seed, schedule/cron, debounce/throttle.
A conversion is only "done" when every unit reaches its target ring, every captured surface passes its oracle on an environment-equivalent green (not local-only), and a final A/B of baseline-vs-result is identical (or its diff is an accepted, recorded residual).
[1] SPEC (markdown under skills/clean-refactor/ — the driver follows it literally, the "brain")
charter.md ....... the constitution / single source of truth for definitions
(§0 product+arch principles, §1 behavior taxonomy, §2 the three
gates + machine gate, §6a four rings + §6b DDD mapping, §0.4
autonomy boundary)
00-bootstrap ..... Phase A-0: runnable-baseline gate (env-equivalent green required)
01-discover ...... Phase A-1: find effect seams
02-classify ...... Phase A-2: DDD classify + 4-ring placement <-- human approval gate
03-capture ....... Phase B-0: capture behavior goldens (6 categories + joint invariants)
04-transform ..... Phase B-1: extract per ring (strangler-fig order)
05-drive ......... Phase B-2: the loop, machine gates, A/B, certificate
(also: the state.json schema SSOT)
06-supervise ..... the supervisor procedure (doubt + fix the harness itself)
07-decompose ..... opt-in post-certificate pass: split oversized same-ring modules
(move-only + shim, gate-green per cycle, certified_head re-bound)
08-install-guardrails
opt-in post-certificate pass: install .claude/ AI-session guardrails
(edit-time lint hook + landing-gate wiring + pointer rules; merge-only)
driver: reads & follows supervisor: doubts & fixes
|
v
[*] state.json ........ the single durable interface both layers read/write.
Enables crash-resume and park-and-resume. Holds: baseline_pin,
seams, contexts, roadmap (per-unit ring + gates), dod, blockers,
accepted_risks, joint_invariants.
^
|
[2] ORCHESTRATOR (python — the deterministic unattended loop, the "body")
run.py ........... CLI entry point
supervisor.py .... the loop: run_until_done — respawn driver, route outcomes, and
validate the terminal (certificate + reconciles + terminal rescan)
agent_steps.py ... the cognition boundary (Protocol: run_driver_cycle / classify /
propose_fix) + the honest "not wired" default (ManualAgentSteps)
llm_agent_steps .. the model-backed implementation (shells out to the `claude` CLI)
guards.py ........ deterministic safety: Tier-A/B classification, oscillation guard,
convergence tracker — the part a wrong LLM judgment cannot bypass
gates.py ......... external gates: codex adversarial review, env-equivalent machine
gate, §S3 mutation-proof, dependency gate, terminal rescan
dependency_gate .. in-process scanner: dependency-rule edges + placement hygiene
(the machine teeth behind the four-ring rule — no self-report)
certificate.py ... §D10 certificate status + disclosure/placement/A-B reconciles
verify_certif. ... the LANDING gate: blocks governed changes whose HEAD the
certificate does not bind (certified_head)
ops.py ........... git / filesystem plumbing (forensic bundles, worktree reset)
-
Driver / supervisor role separation (charter §0.4). The driver may only follow the spec literally; it has no power to edit the harness or weaken a gate. The supervisor (06-supervise) is the only role that may doubt and fix the harness. Separating the follower from the doubter stops the follower from rationalizing its way around an inconvenient gate (self-confirmation bias).
-
Guards vs gates. Guards are in-process and deterministic; they cannot be argued with (e.g. a self-edit that would weaken a gate's detection power is auto-classified Tier-B and blocked). Gates are external and adversarial (codex review; an environment-equivalent machine gate enforcing effect-diff = 0). All cognition flows through the
agent_stepsboundary, so even a wrong model judgment is wrapped by the deterministic guard layer.
Clean / hexagonal architecture drawn as concentric rings. The dependency rule is the whole point — every dependency points inward, and an inner ring may never import an outer one:
+-----------------------------------------------------------------------------+
| framework (0) -- HTTP, CLI, cron -- the outside world |
| |
| +-------------------------------------------------------------------+ |
| | adapter (1) -- DB, queues, vendor SDKs, port implementations | |
| | | |
| | +---------------------------------------------------------+ | |
| | | usecase (2) -- application workflows, orchestration | | |
| | | | | |
| | | +-----------------------------------------+ | | |
| | | | domain (3) -- entities, invariants, | | | |
| | | | pure business rules (no I/O) | | | |
| | | +-----------------------------------------+ | | |
| | +---------------------------------------------------------+ | |
| +-------------------------------------------------------------------+ |
+-----------------------------------------------------------------------------+
Read it as framework(0) -> adapter(1) -> usecase(2) -> domain(3), dependencies inward only.
DDD classification maps to rings: Core → domain, Supporting → usecase,
Generic → adapter (or buy, which is out of scope).
Phase A (understand) [human] Phase B (transform)
bootstrap -> discover -> classify -> approve -> capture -> transform -> drive -> CERTIFICATE
(+env attest) (placement) (per ring) (gates·A/B)
\________________ 06-supervise wraps the whole thing as the unattended loop _____________/
Human touchpoints are exactly the legitimate ones, and every halt is a recoverable park-and-resume checkpoint (charter §0.4):
- Placement approval (§2.2) — the one designed human-judgment gate: classify + 4-ring placement. Proposal → discussion → explicit approval.
- Environment attestation (00-bootstrap S3) — only a human touch when there is no CI: the green source must be CI / env-equivalent sandbox / human-attested-with-evidence.
- Hard blockers (§0.4) — missing secret/DB, env-parity loss, retry-budget exhaustion, oscillation: park, surface to a human, resume from the recorded step once provisioned.
The harness ships as a Claude Code plugin so you can invoke it in-session instead of running the orchestrator by hand. Install it from this repo:
/plugin marketplace add tryumanshow/clean-refactor-harness
/plugin install clean-refactor-harness@tryumanshow
Those two commands are one-time setup. After that you don't type a command to use the
harness. Just describe the task in natural language and Claude auto-invokes the
clean-refactor skill by matching your request to its description.
For example, once installed you just type a request like this to Claude:
Convert ./backend to clean / hexagonal architecture without changing its behavior.
Run the clean-refactor harness: work in a disposable git worktree off the current
commit, capture characterization (golden) tests before any structural change, and
pause at the placement-approval gate for my review.
Claude matches that to the skill's description and runs the harness. The
/clean-refactor slash command is also available as an optional explicit trigger
(then describe the target), but it is not required.
The plugin is a thin wrapper. It changes how you invoke the harness, while what the
harness does stays identical. The skill (skills/clean-refactor/SKILL.md) only routes to the existing
charter.md + 00–06 phase specs and the Python orchestrator; it does not restate the
methodology. The behavior-preservation guarantees still live entirely in those specs and in
the deterministic guards/gates, so a plugin run and a direct python run.py run execute the
same procedure. The only thing the plugin adds is discoverability and the /clean-refactor
entry point.
Dependencies: the claude CLI (drives the
cognitive steps headless) and the codex CLI (adversarial
gate). The deterministic core (guards.py, gates.py, supervisor.py) runs on the standard
library alone.
cd orchestrator/harness_supervisor
# Unattended (model-backed) — converts a target worktree until terminal:
python run.py \
--harness /path/to/this/repo/skills/clean-refactor \
--worktree /path/to/disposable/target-worktree \
--baseline <immutable_baseline_git_sha> \
--max-cycles 30 --driver-timeout 600
# Wire-up smoke (no model calls — pauses at the cognition boundary to a human):
python supervisor.pySwap ManualAgentSteps (the honest "not yet wired" default, which parks every cognitive step
to a human) for LlmAgentSteps to run unattended; run.py does this for you.
cd orchestrator/harness_supervisor
python -m pytest -q # 222 passed
ruff check . # cleanThe suite covers the deterministic guards (Tier classification incl. constitution-edit protection, oscillation fingerprinting, convergence), the supervisor control flow (fail-safe routing, single-classify, apply-failure handling, env-blocker parking, placement-gate routing), the run-ordering terminal (certificate status, disclosure/placement/A-B reconciles, and the terminal rescan that re-measures the actual tree instead of trusting driver-recorded numbers), the dependency/placement scanner, the landing gate, and the model-backed step parsing / fail-safe.
Experimental, with one real-scale conversion behind it. The pipeline was first demonstrated
end-to-end on a small target (a single-file stdlib CLI): spaghetti → clean four-ring architecture
across eight strangler-fig commits, with behavior preserved (final A/B AB_IDENTICAL, golden +
fitness suites green) and a cross-unit joint invariant captured and verified.
It has since driven a production FastAPI backend (~110k LOC Python) through the same procedure:
full four-ring conversion with the dependency rule CI-enforced by import-linter contracts,
a later package-by-feature re-layout, and a 9-cycle 07-decompose pass that split a 5,567-line
usecase module down to 1,398 lines with every cycle gate-green. Cost/time per cycle and
multi-context generality beyond that one codebase remain not yet systematically measured.
The spec files (charter.md, 00–07) are written in Korean (the author's working language);
the orchestrator code and this README are in English. The deterministic guarantees live in the
code and are language-invariant; only the cognitive layer reads the prose spec.
The end state this is built toward. Each item is the trajectory implied by the gaps the Status section names — not a wishlist of unrelated features:
- Drive the residual to zero. Today the honest claim is "captured behavior preserved +
residual = X". The work is to shrink
Xby widening oracle coverage across all six behavior categories — especially the hardest to capture (concurrency / ordering, time / nondeterminism, authz boundaries) — so fewer effects ever fall outside the oracle. - Prove it beyond one codebase. One production backend has been converted end-to-end; the target is generality — more codebases, more bounded contexts, real external effects, and measured cost / time per cycle — the numbers Status calls "not yet systematically measured".
- More unattended, the same single human gate. Keep placement approval (§2.2) as the one designed human-judgment touchpoint and push everything else hands-off — stronger machine gates, crash-resume, park-and-resume — so a run needs a human only where judgment is genuinely required.
- Honest by construction. The certificate never upgrades to a blanket "behavior preserved"; it always ships the enumerated residual. The direction strengthens the gates, never loosens the claim.
.claude-plugin/ plugin + marketplace manifests (Claude Code packaging)
skills/clean-refactor/ the plugin skill — the repo's spine:
SKILL.md entry point (routes to the spec + orchestrator)
charter.md, 00–08.md the spec (what to do; 07 decompose + 08 install-guardrails are opt-in)
docs/ design provenance: dry-run audits + the dogfooding gap log
orchestrator/harness_supervisor/ companion orchestrator (unattended runs) + tests
README.md, README.ko.md this
