Skip to content

ADD 2.0 — skill-led method, thin kernel, living specs, measured 1.7× spec-kit#163

Merged
TinDang97 merged 77 commits into
mainfrom
feat/adaptive-flow
Jul 20, 2026
Merged

ADD 2.0 — skill-led method, thin kernel, living specs, measured 1.7× spec-kit#163
TinDang97 merged 77 commits into
mainfrom
feat/adaptive-flow

Conversation

@pilotspacex-byte

@pilotspacex-byte pilotspacex-byte commented Jul 18, 2026

Copy link
Copy Markdown
Contributor

ADD 2.0 — skill-led method on a thin state kernel

The 2.0 major, M1–M8 (design ratified 2026-07-17): the skill drives the loop, add.py shrinks to a state kernel, personas become the method's adaptive brain, specs become living 5-DD documents, and the book stops shipping with the package.

Subsumes feat/strategy-intake (that branch never got a PR; its 14 unmerged commits — persona-owns-gates, thin-lane, engine-only teardown, md-cut sweep — are all ancestors of this branch, so this single PR lands both lines).

The 2.0 line (M1–M8)

  • M1 — personas measurable (61fa230a, 638cc033, f2a3e455, 99a3a5e1): persona task-kinds + route-outcome traces; 12 capability seed templates; 5 phase agents collapse into ONE add agent.
  • M2/M3 — PLAN core + 5-DD specs (b82fcbbd, ffc47d4a): §3 PLAN carries a measurable Target the gate judges; five living specs + the delta-append kernel verb.
  • M4 — inline lane + skill folds (e2b5b4b6, a68da012, d74a58bd, d593e03e): AI-routed below-scope route; SKILL.md IS the loop (phases 7→3 on-demand refs); persona proposes the lane, the freeze ratifies it.
  • M5 — kernel-trim (a7ef15e6, 9afa0efb, 91230a3a, e2d89a93, de70180a): 54→31 verbs, add.py 9,558→6,596 lines; 6 phases → direction·build·verify; ONE task template; test corpus −180 duplicated pin tests.
  • M6 — PLAN.md + book-stops-shipping (c36bc484, 3532a03f, f8c2a58e): task doc TASK.md→PLAN.md with one-shot add.py migrate (480 board docs dogfooded); book publishes at https://pilotspace.github.io/ADD/ only; 31 legacy milestones archived.
  • M7 — route scoreboard (GEPA) (ea3b5bf2): deltas reads the route traces into a per-lane scoreboard + the loop's GEPA reflection beat.
  • M8 — 2.0.0 cut + bench (f3f3dd6b, 015048d4, 52412c36): version lockstep across all 5 sources + CHANGELOG with Breaking block + release suite; --session-mode continue context-rot benchmark; results report.

Since opening — honesty wave, atomic node, task graph

  • Bench honesty wave (f4ffbe6c): 6 meter defects found and fixed (tempdir store pollution, foreign-app import, conftest collision, …); the first-edition "spec-kit collapses under evolution" claim was traced to our own meter and publicly retracted — reports revised in place, READMEs re-receipted. Corrected fresh-mode matrix: add 6/6 all-floors $17.51 · spec-kit 6/6 $10.05.

  • Atomic node (0b0192a7, 121c2feb): ONE atomic PLAN.md per task (interface = frozen contract · red suite · scope · verdict); every lane scaffold removed; §3 gains a Regression floor: binding the host repo's own suite at the gate.

  • SWE smoke, re-run under the official swebench docker eval (93360765, a4e7d677): haiku+ADD 2/3 → 3/3, cheaper ($2.15 vs $2.52; the miss's patch shrank 1,626B→559B) — appendix-H prediction 1 flips back at the tier where it failed; sonnet parity 3/3 at nearly half the wall-clock (50.9→28.1 min). Harness defect fixed: SWE workspaces now seed their own .add state (ancestor-root leak).

  • Context rot, remeasured (a4e7d677, 1d153795): with the atomic template, ONE continuing conversation over six milestones holds fidelity flat 1.0×6 with zero regressions, 22% cheaper — the mode that previously decayed 0.92→0.75 in BOTH flows. Spec-kit's recorded board in the same mode still decays with regression rates 0.33/0.29. Raw price gap (~1.7×) unchanged; appendix-H updated with both data points.

  • Task-graph-native (8d4f573b, 0dadb172, 49e588d2, 75bb1289, d5d9f0fd): milestone = scope ROOT over the task DAG. milestone-confirm compiles ## Tasks into edges new-task inherits; relate ratifies edges post-creation (freeze prints edge-hints); locate maps a failing test → owning node → in-node vs interface-regression → dependent closure (the minimal repair subgraph), and path::test resolves through §4 covers: to the exact frozen §3 clause; graph renders the DAG as mermaid with dashed planned-never-created nodes; check warns on planned drift. Verbs 31→34.

  • Negative result, published (d5a28791 → reverted f77f684e): a neighborhood-status card (ambient prints of parent contract heads) failed its gate bench — the rot curve returned at +33% cost because the card externalizes precedent, not spec. Recorded in the results doc as a design floor: localized context must be pull-based, spec-side, and complete.

  • Method refinements since the graph waves (8feeec4d, 57132feb, aa9f1f14, 3cd8792a, 0a56dede, a91f7d04): the skill now seeds a domain-fit persona at every new task when none fits; its front door analyzes intent into a task before sizing (analyst-first, not just orchestrator); run mode decouples to the pure autonomy dial (the streams: posture subsystem removed — concurrency is "a subagent per task"). Then an atomic-mode alignment sweep of the phase guides: direction.md's Grounding folds from a 7-field written block to "persist the interface, reason the rest in-context" and the retired lane/route: framing is corrected; SKILL.md reframes from orchestrator to atomic-graph navigator (node · edges · graph/locate up front) and drops the one non-ADD-specific constraint; verify.md's sensitivity vocab is compressed (the rest is load-bearing); and a drift sweep retires the last stale §0 reference (build.md re-aims its facet anchor to CONVENTIONS.md Honors). The guides now match the 2.0 atomic template; SKILL.md holds under its 9,500 B ceiling; all three skill trees stay byte-identical.

Measured (benchmark/results/…)

  • Cost target MET: ~1.7× spec-kit per milestone at equal fresh-mode trust (1.x was ~2.7×); both flows hold every trust floor in fresh mode.
  • Where ADD separates: the continuing-conversation mode (flat 1.0 vs decay + regressions) and the small-model tier (haiku 3/3 with a minimal patch vs 2/3 over-build) — both published with raw data, alongside our own retraction.

Verification

  • Full corpus 2,332 passed / 3 env skips (29 new red→green tests for the graph waves + 3 for the dependency allowlist); add.py check clean; wording-lint clean.
  • Dependency allowlist now enforced (e0980435): .add/dependencies.allowlist claimed CI rejection nothing implemented — a new corpus suite rejects any unlisted runtime dep; @clack/prompts recorded as the one approved Node dependency.
  • Updater 2.0 gaps closed (50ee1cce): global home now mirrors agents/ (the fix(installer): .claude/agents is a shared namespace — stop wiping user subagents on init/update #151 roster-drift residue — update --global propagation finally refreshes rosters + applies retired tombstones); update prints crossing nudges (sync-guidelines on any version cross, migrate on a TASK.md-era board); stale "docs refreshed" prose corrected. Corpus 2,339.
  • ENGINE_MD5 bfd47201… / PKG e0ff925d… @ run-mode-decouple (the last engine change; the method-refinement sweep edits skill guides only — engine untouched); full corpus 2,346 passed at HEAD; twins synced (tooling ×4, skill ×3, engine_pin ×4).

Remaining after merge (human-gated): tag v2.0.0 + npm/PyPI publish (in-band publish.yml on tag push).

author: Tin Dang

TinDang97 added 30 commits July 16, 2026 09:56
…session before a milestone is born

Dogfoods ADD's own intake→scope flow to spec a new-major method milestone: the
skill should shape a milestone through a persona-led, risk-proportional strategy
discussion that optimizes its task plan with the user BEFORE the milestone is
committed. Closes the gap that today intake.md only classifies into a bucket and
scope.md runs a generic co-specify — no persona lens, no task-DAG optimization
(personas apply only at design/build/advisor/verify; strategy is per-task §5 only).

5 breadth-first tasks (freeze the ## Strategy schema first):
- strategy-section — ## Strategy slot in MILESTONE.md.tmpl + engine handling
- persona-at-intake — extend add-persona selection/routing to the intake surface
- strategy-guide — new strategy.md: persona-framed discuss→optimize→converge to ~95%
- advisor-strategy-trigger — advisor refutes a high-uncertainty milestone's strategy
- risk-proportional-skip — micro/--fast bypass; zero added per-turn cost

Guardrail: ## Strategy is SOFT/advisory (like §5) — no new hard gate, no re-added
per-turn ceremony/cost (risk-proportional-skip is a first-class task), security still
HARD-STOPs. Deferred (out of scope): the AI proactively INVENTING milestones
unprompted — this shapes a request the human raises, it does not originate scope.

Created via add.py new-milestone strategy-intake --await-confirm + milestone-confirm
(by tindang); check 880 passed / 0 failed (no dangling criteria). extends
persona-learning-loop, dynamic-personas, scope-loop; relates-to
risk-proportional-ceremony, expectations-first.

author: Tin Dang <tindang.ht97@gmail.com>
…t (WIP, tests phase)

First task of the strategy-intake milestone, driven through specify → plan → freeze.
§1 rules (M1-M4 template slot + renders + placement-outside-parsed-spans + 3-twin
parity; R1 strategy_must_stay_soft, R2 tasks_parse_corrupted), §2 four gherkin
scenarios, §3 frozen contract: MILESTONE.md.tmpl gains one drafted-blank `## Strategy`
section between `## Exit criteria` and `## Close`, engine renders it verbatim — no
add.py parse/gate, no ENGINE_MD5 repin, 3 twins byte-identical.

Least-sure flag surfaced at freeze: whether placement after `## Exit criteria` leaves
the Tasks-DAG parse + pre-confirm section scan byte-behaviour-unchanged (a red test
asserts the DAG reads exactly N task rows and confirm/check stay clean) [contract/test].

Crossed into tests — red suite next. Frozen @ v1, approved by tindang.

author: Tin Dang <tindang.ht97@gmail.com>
…M brain

Re-scope the milestone from "a strategy session before a milestone is born" into the
bigger vision the user set: remove ceremony by making a fitting persona the
project-management brain — it shapes each milestone's strategy AND owns how every human
gate is communicated and paced, retiring the fixed report-template.md section list.
Persona-driven, adapts per project. Marked risk: high (method-defining).

The four report floors: show-before-ask, one-approval-at-freeze and never-pre-stamp move
from a fixed template into the persona CONTRACT (met in the persona's own voice, no fixed
section list). The fourth — security is ALWAYS HARD-STOP — is kept HARD as a STRIKEABLE
carve-out against the "floors become persona judgment" choice, because it is written into
the method constitution (release readiness floor is un-forceable) and the operating rules
forbid authoring a security auto-pass. The human may strike that one line to dissolve it.

New freeze-first task persona-owns-gates carries the persona-owned-gate contract every
other task cites. strategy-section kept, frozen @ v1. check: 884 passed, 0 failed.

author: Tin Dang <tindang.ht97@gmail.com>
…tract (WIP, tests phase)

The strategy-intake freeze-first task. Retires the fixed report-template section
list in favour of persona-owned principles: a gate report must CONVEY the decision +
its arc, the shape/plan, flags lowest-confidence-first, evidence, and a guided ask —
the persona owns structure, order, emphasis, length and cadence, adapted per project.

The four floors survive as persona-contract obligations, not layout: show-before-ask,
one-approval-at-freeze, never-pre-stamp, and security = HARD-STOP (the one hard,
un-persona-negotiable floor — the strikeable carve-out per the milestone). The engine
is untouched: the `Reported:` trace + contract/verify_report_unrecorded audit codes
stay verbatim; no ENGINE_MD5 repin.

risk: high, autonomy: conservative (method-defining trust-layer change, human gate at
verify). §1-§3 filled, frozen @ v1 approved by tindang, crossed into tests. Red suite
next: test_persona_owned_gates.py + migrate test_report_arc (SUMMARY/ordering now
persona-owned) and test_xml_convention (the report-blocks heading).

author: Tin Dang <tindang.ht97@gmail.com>
…or persona-owned gates (gate PASS)

The strategy-intake freeze-first task, complete. report-template.md is reframed
from a MANDATED ordered section list into persona-owned PRINCIPLES: a gate report
must CONVEY its content (decision + ARC · shape/plan · flags lowest-confidence-first ·
evidence · a guided APPROVE ask), but the fitting persona owns structure, order,
emphasis, length and cadence — adapted per project. A sensible default layout remains
as the persona's baseline; the mandate does not.

The four floors survive as persona-contract obligations, not layout — show-before-ask ·
one-approval-at-the-freeze · never-pre-stamp — with security = HARD-STOP marked the one
un-persona-negotiable floor (the strikeable carve-out per the milestone). SKILL.md's
report line now names the principles, not the banner→…→NEXT sequence.

Engine untouched: the `Reported:` trace + contract/verify_report_unrecorded audit codes
stay verbatim; ENGINE_MD5 unchanged (4e65596…). The mandate→principle reframing retired
now-optional prescriptive prose, freeing ~1050 B that offset the added principle +
four-floors block — report-template.md landed at 9514 B (net −109 B), so the razor-thin
reference pool stayed under budget with no rebaseline. 3 skill trees byte-identical,
SKILL.md 9489 B (< 9500 ceiling).

Tests: test_persona_owned_gates.py (10) green; migrated test_report_shape_scan_audit.py
byte-ledger (9623→9514) + test_xml_convention report-blocks heading. 114 affected
report/skill/parity tests green · check 888/0. risk: high, autonomy: conservative,
verify gate approved by tindang. Seeds the next task: redefine UDD as experience-driven
development, hosting the persona-owned gate as a text-mode UX artifact.

author: Tin Dang <tindang.ht97@gmail.com>
…PM AND UX

Extend the strategy-intake milestone per the user's direction ("move report template
to UDD as part of user experience"; chose to redefine UDD, folded into this milestone).
The milestone now spans personas owning both the project-management surface (strategy,
gates) and the user-experience surface (gates as UDD artifacts).

Broadened goal + scope: UDD is redefined from UI-design into experience-driven
development — the pillar (design.md + UDD chapter + SKILL.md trigger + the four design
axes) covers UI AND interaction/gate UX, and the persona-owned gate report is hosted as
a UDD text-mode UX artifact designed through the UDD lens.

Two new breadth-first tasks: udd-experience-pillar (redefine the pillar) and
gate-experience-udd (host the persona-owned gate as a UDD artifact; depends-on
udd-experience-pillar + persona-owns-gates). persona-owns-gates marked done; two exit
criteria added. check: 889 passed, 0 failed.

author: Tin Dang <tindang.ht97@gmail.com>
…en development + a fifth INTERACTION axis (gate PASS)

The strategy-intake milestone's UDD-redefine pillar. design.md is reframed
from UI-design into EXPERIENCE-DRIVEN development: the loop now triggers on a
UI feature OR any human-facing experience surface (a screen · an interactive
flow · a human gate), stating "UDD is experience-driven development, not
UI-only". The human gate is named an in-scope UDD surface — setting up
gate-experience-udd — without yet folding report-template.md.

The design-intake beat gains a FIFTH axis, INTERACTION (cadence · when/how to
seek the human · turn-rhythm), alongside the four originals (FIDELITY · CONCEPT
· LAYOUT · VISUAL DESIGN), which stay with frozen names. Every "four axes"
reference (design-intake beat + hard-rules) becomes "five axes" naming
INTERACTION. SKILL.md's UDD trigger broadens to "UI/experience surface → UDD
loop"; the +11 B is offset by a same-guide "the default mode" trim, landing
SKILL.md at 9490 B (< 9500 ceiling) — compress-not-rebaseline held. The three
skill trees stay byte-identical (design.md 4b340bb… · SKILL.md 756950f…).

The loop machinery is UNCHANGED (5 beats · capture/confirm · read-only binds).
Engine untouched: no add.py edit, ENGINE_MD5 unchanged (4e65596…). The
DESIGN.md.tmpl + appendix-c-glossary.md INTERACTION field is deferred (§7
SPEC·open); the report-template fold + lightweight gate loop are seeded for
gate-experience-udd.

Tests: new test_udd_experience_pillar.py (10) green — experience-driven framing
· five axes incl. INTERACTION · SKILL trigger + ceiling + 3-tree parity;
migrated test_design_intake_beat.py (15) stays green (four axis names still
asserted). report/xml/parity/lean guards (47) green · check 893/0. risk: high,
autonomy: conservative, verify gate approved by tindang. Refute-read EARNED
(self); advisor 3-lens PASS (security/concurrency/architecture CLEAR).

author: Tin Dang <tindang.ht97@gmail.com>
…te-experience-udd rename + clear dedup debt

thin-engine-loop W1 (phase-collapse-6-to-3): demote add.py from loop-driver to
helper by collapsing the two pure-bookkeeping `advance` calls. A new opt-in
`--thin` lane freezes the whole Direction bundle (spec+plan+§4 tests) in ONE
`freeze --cross` that crosses straight to build via _build_entry's SAME floor
machinery — a 3-call flow (new-task · freeze · gate) down from the 5 every
existing lane prescribes. oneshot/fast/default lanes are byte-unchanged; the
mechanical floor is intact — tamper tripwire + flag-verified + §5 scope
snapshots are all still captured through _build_entry.

  - add.py: is_thin state marker; freeze guard min-phase = specify for thin;
    `freeze --cross` thin branch (specify/plan → build in one call); new-task
    --thin flag + 2-call thin recipe
  - test_thin_engine_call_floor.py: RED→GREEN target — new-task --thin prescribes
    ≤3 engine calls
  - engine_pin.py: ENGINE_MD5 re-aimed 4e655960→ee9631f9 (add.py changed),
    synced across all 3 tooling trees
  - .add/SEAMS.md: _declared_scope anchor re-pinned 5901→5934

fold gate-experience-udd (strategy-intake): physically rename report-template.md
→ gate-udd.md across all 3 skill trees (the text-mode UDD gate surface, a member
of the UDD doc family); repoint every reference (9 skill guides · 2 book docs ·
tests); design.md gains the lightweight text-mode gate variant. The four gate
floors (show-before-ask · one-approval · never-pre-stamp · security-HARD-STOP)
are preserved verbatim in substance. Task stays at phase=build (autonomy:
conservative) — the PR review IS its human verify gate; not engine-gated here.

clear the dedup debt: the orchestration pool was 349 B over its 41300 floor at
branch HEAD (pre-existing) + 189 B from gate-udd's design.md = 538 B over.
Compressed −538 B of prose across streams/advisor/loop/design (in-file, no
rebaseline; every pinned phrase, reject code, CLI verb, and the byte-identical
<strategy> block preserved). Pool now 41278 (22 under) — greens the byte-budget
family + the fresh-checkout rollup.

Suite: only pre-existing environmental reds remain (pty_clack ×4, installer
npm-cancel ×1 — no npm/@Clack in the bare unittest env).

author: Tin Dang <tindang.ht97@gmail.com>
… docs

The tooling suite pins doc PROSE in ~278 files: each past method task was built
red/green against the exact wording it added. That ossification reddens a guard
on every doc improvement (it bit the dedup compression 3× today) and blocks the
ceremony-cut. Per direction: remove the pure prose-guard files, keeping every
test that guards a CORE VALUE (behavior · machine tokens · structure/parity).

Deleted 15 files whose SOLE assertions are "this doc/guide/render-card contains
phrase X" — no engine execution, no token/reject-code/CLI-verb guard, no md5
parity, imported by no other test:

  test_worktree_default_wording · test_two_surface · test_v11_docs ·
  test_docs_align · test_scope_level_enum · test_supported_agents_docs ·
  test_skill_todo_flag · test_arc_gate_wiring · test_report_arc ·
  test_docs_accord · test_gate_read_diet · test_risk_report_render ·
  test_template_dedup · test_question_summary_layer · test_rewrite_core

Verified before removal: each read to confirm pure-prose. Three false-positive
classes were caught and KEPT — md5 parity guards (test_*_parity), function
unit-tests (test_md_section slices markdown), and behavior-via-import
(test_installer_soul_seed runs the pip installer). After removal: 0 import
failures across the remaining 378 modules; inventory/hygiene guards
(semantic_inventory · temp_hygiene · plugin_manifest) green.

The behavioral/token coverage these prose-echoes shadowed lives on their own
modules (freeze/gate/audit behavior, verb + reject-code census, 3-tree parity) —
untouched. Docs on these surfaces are now freely editable.

author: Tin Dang <tindang.ht97@gmail.com>
…cise add.py

Per direction: the tooling suite must test the add.py engine, nothing else. The
doc/prose/wording/parity/rubric guards were ceremony — they ossified the docs
(every doc edit reddened a phrase-pin) without testing behavior. Removed all 67
tests that only inspect text: no add.py execution, no engine-package import, no
installer/real-code, no subprocess. 378 → 311 test files (−8165 lines).

Deleted class: doc-echo (test_v6_run · test_v8_docs · test_foundations_chapter ·
test_references_appendix …), wording/slang lints (test_wording_lint ·
test_ubiquitous_language), render/rubric/guide guards (test_confidence_rubric ·
test_intake_rubric · test_design_loop_guide …), byte-budget pools (test_skill_lean ·
test_skill_dedup · test_streams_persona_flow), and doc-structure guards
(test_xml_convention tag-census · test_book_parity · test_version_sync ·
test_semantic_inventory · test_worker_contract_sync). Kept test_md_section — it
unit-tests the md_section slicer function (real code).

Decoupled the survivors that leaned on deleted hubs:
  - 10 engine tests imported test_skill_lean.POOLS for a phases byte-budget
    side-check — stripped those non-engine `test_phases_pool_*` methods.
  - test_roster_portable (tests init -> AGENTS.md roster sync, real add.py) imported
    roster data from test_agent_roster; inlined AGENTS/AGENT_PHASES/AGENT_TREES/
    _agent_path/_frontmatter so it stands alone.
  - .add/SEAMS.md three-tree-parity seam cited the retired test_book_parity —
    dropped it; the engine/skill/bundle parity guards (test_engine_repin_parity ·
    test_tree_parity · test_bundle_parity) survive.

Suite: 3121 tests, only the pre-existing environmental reds remain (pty_clack ×4,
installer npm-cancel ×1 — no npm/@Clack in the bare unittest env). 0 import failures.
Docs on every surface are now freely editable — the ceremony-cut is unblocked.

NOTE: book-tree doc parity is no longer test-enforced (its guard was pure text
inspection). Engine + skill + bundle trees stay guarded; book/doc sync now relies
on discipline.

author: Tin Dang <tindang.ht97@gmail.com>
First increment of the md ceremony-cut (docs now freely editable after the
engine-only test teardown removed their prose pins). Kept every core value —
machine tokens, reject codes, functional templates, the pin-locked worker XML,
3-tree byte-identity; cut only teaching redundancy.

streams.md — gutted StreamsSafetyClausesTest from test_streams.py: 8 pure-prose
asserts that pinned exact safety wording ("name the task every time", "lease",
"circuit-breaker", "never write shared state" …). The safety CONCEPTS stay in
streams.md; their wording is now free to evolve. The real engine guards survive
untouched — slug-routing precedence (SlugRoutingPrecedenceTest), the ×3 md5
parity guard, the reject-code pins (unverified_fork_base), and the pin-locked
<strategy> block (byte-identical to advisor.md). test_streams: 13 green.

run.md (×3 trees, −383B) — compressed the "specification bundle (v7)" section,
which itself points to phases/3-plan.md as its "one home", down to that pointer +
the one load-bearing fact (one approval at the frozen contract); dropped the
historical "v7 reversal (recorded)" aside that restated the already-stated
auto default. All 12 machine tokens preserved (add.py audit/check/gate/heal,
refute_unrecorded, unguarded_high_risk_auto, goal_not_auto_ready, …); 3 copies
byte-identical.

Verified: 105 targeted tests green (test_tree_parity, test_bundle_parity, and
the run.md-referencing engine tests: setup_run_mode, verify_deepen,
earned_green_rubric, refute_record_required, high_risk_signal, goal_auto_ready_gate,
plugin_manifest) + test_streams 13. Full suite left to CI (discover run exceeds
the 2-min bound — the standing "don't block on the full suite" rule).

author: Tin Dang <tindang.ht97@gmail.com>
…/gate guides

Second increment of the md ceremony-cut. Removed pure illustration that repeats
the rule it sits under — the decision tables, reject codes, and format blocks
above each example ARE the core value; the worked example is a teaching aid that
costs tokens every load without adding a rule.

- intake.md — dropped "## Worked examples (from this project's own history)" (the
  4-row request→bucket table); the four-buckets table + tie-break test + fast-fit
  rule already carry the classification logic.
- scope.md — dropped "## Worked example" (the v4-1 walkthrough); the per-outcome
  table + drafting sections + well-formedness gate are the method.
- gate-udd.md — dropped "### Per-gate examples" (two ARC illustrations); the ARC
  format block + its field bullets already specify it.

All reject codes / machine tokens preserved (not_classified · dangling_criterion ·
duplicate_goal · ask_human · frozen_scope · split_required · missing_capture …);
3 trees byte-identical.

Also gut test_review_checklist.test_run_md_accord — a brittle wording-pin that
required run.md to contain the literal phrase "freeze review checklist". The
prior commit (d264c38) compressed that sentence to point at phases/3-plan.md as
the checklist's "one home" (navigation preserved), which tripped this cross-ref
pin — a latent red that batch didn't cover. The structural guards in that file
stay: the md_section 7-item count, the `risk: high · autonomy: conservative`
token pin, the 3-tree parity, and the ENGINE_MD5 untouched-guard.

Verified: 125 targeted tests green (review_checklist, tree/bundle parity,
intake_interview, gate_owner_marker, report_rendered_trace, udd_persona_checklist,
project_scope_lock, design_intake_beat, plugin_manifest, streams) +
add.py check 898 passed / 0 failed.

author: Tin Dang <tindang.ht97@gmail.com>
Per direction (cut ceremony deeper: strip explanatory rationale to rule +
machine tokens, keep the enforceable rule blocks). design.md 7049 -> 5811 B.

Compressed the five-beat body prose and the tool-agnostic-capture section to
the rule + tokens; dropped teaching rationale ("so the approved screen is
traceable", "the cheapest artifact that shows...", "it keeps screens
consistent"). Kept intact: the <constraints> Hard rules block, every machine
token (missing_capture, tokens.json, catalog.json, prototypes/<name>.json,
DESIGN.md, flow: design), the five axis names, and the concept-pins the
design tests assert (never-lowers-a-gate, never-an-auto-pass, engine-never-
renders, the UI-Designer/UX-Researcher dimensions).

Fixed a line-wrap that split "never lowers a gate" across a newline (broke the
phrase-pin — same class as the "default for any spawn" wrap bug).

Verified: 30 tests green (test_design_intake_beat, test_udd_persona_checklist,
tree/bundle parity); 3 trees byte-identical.

author: Tin Dang <tindang.ht97@gmail.com>
Continuing the aggressive rationale-cut. These two guides are already dense
(procedural steps, tables, output_format/exit_gate blocks, machine tokens), so
the trims are modest — the scenario-writing rationale in 1-specify and the
foundation-interview preamble in 0-setup, compressed to rule + tokens.

Preserved every command, table, machine token (amount_invalid, already_locked,
setup_unlocked, --await-lock, add.py init/lock/advance) and the exit_gate /
output_format blocks. 3 trees byte-identical.

Verified: 75 targeted tests green (tree/bundle parity, phase_bundles,
progressive_context, setup_run_mode, stale_guide_sync, plugin_manifest).

author: Tin Dang <tindang.ht97@gmail.com>
Fill the blank thin-engine-loop MILESTONE.md against its human-set goal
(≤3 add.py calls per task · 6→3 phases · loop-in-SKILL): scope, shared
ground (engine/template/skill anchors incl. the 9497/9500 SKILL ceiling,
phase-name pin census, twin-tree parity), shared decisions (route glossary,
persona-proposed/human-ratified ceremony, the four floors bind every route,
≤3 calls is happy-path not a cap), freeze-first contracts, six breadth-first
tasks (phase-collapse-3 · template-unify · skill-loop-fold ·
persona-routes-depth · engine-kernel-trim · call-census-proof) and
verifier-cited exit criteria.

Queue persona-gepa-loop as the follow-on: route-outcome traces reflected
into persona routing-rule deltas at fold-time (GEPA-style, human-folded,
never-clobber rails) — depends-on thin-engine-loop, extends strategy-intake.

Scope drafts only; task decomposition awaits the human confirm
(engine check 898 passed / 0 failed).

author: Tin Dang
…rafted + red suite (8 red / 3 floor pins)

Six tasks created with dependency DAG; phase-collapse-3 §1–§4 drafted to the
freeze: universal direction·build·verify enum, freeze --cross crosses the
whole front via _build_entry, gate compound kept, legacy read map, 3-call
recipe every lane. Wordy-test removal authorized in milestone shared
decisions; template family lean-pass added to template-unify.

author: Tin Dang
…·verify(·done)

The thin lane's 3-call walk becomes THE walk. PHASES collapses to
("direction", "build", "verify", "done"): the old front (specify · scenarios ·
plan · tests) is ONE direction span ending at the single freeze approval, and
observe folds into verify. Every task now runs new-task → freeze --by <name>
--cross → gate PASS — three engine calls, one human seam.

Engine (add.py + add_engine/constants.py, ~360 lines):
- LEGACY_PHASES: old phase names normalize at the ONE read accessor — 473
  legacy task records render/load unchanged, zero task-file rewrites; the CLI
  accepts legacy names with a mapped-to note; check's marker-parity check
  normalizes both sides (a hand-written `phase: specify` marker never reds).
- The two retired crossings' floors relocate into _build_entry's shared stack:
  cross-component hold + producer snapshot / consumer pin now fire at the one
  direction→build crossing alongside freeze gate · build-expectations · flag
  check · tamper tripwire · §5 scope snapshot. advance, freeze --cross, and
  the phase-build override run the IDENTICAL stack; --thin is a no-op.
- guide never re-teaches a passed gate: post-freeze direction steers to
  `add.py advance` on both the plain and --json surfaces.
- advance --to maps legacy tokens, stops at direction, and no-ops friendly on
  an already-reached target; init/new-task teach the 3-call recipe.

Suite migration (49 test files, red 372 → green):
- Walk helpers collapsed (a task is BORN at the freeze seam); stall phases
  move to `direction`, landings to `build`; FROZEN fixtures carry the
  least-sure flag the unified crossing now checks.
- 34 old-walk value pins deleted under the recorded wordy-test authorization
  (thin-engine-loop MILESTONE.md, 2026-07-16) — none exercise a floor; every
  freeze/gate/tamper/audit/scope/cross-component floor test MIGRATED, and
  test_phase_collapse.py pins the new 3-call walk end to end.
- Two vestigial pins retired whose targets d264c38 (human-approved md-cut)
  deleted: wave_ledger's dangling StreamsSafetyClausesTest ref, run.md's
  "seven lines" prose pin.

Book + docs ripple: ch02 flow mermaid now draws the 3 phases as subgraphs
wrapping the 7 beats (synced ×4 trees incl. bundle), diagrams/CHECKLIST.md
names the phase bands, appendix-c dogfood copy re-synced. SEAMS.md
_declared_scope anchor re-pinned 5934→5915.

Pins re-aimed once: ENGINE_MD5 6f688c41…, ENGINE_PKG_MD5 89a75e5d…; add.py /
add_engine / engine_pin byte-identical across canonical, bundle, and both
dogfood twins.

task: phase-collapse-3 (thin-engine-loop W2) — §3 FROZEN @ v1, build phase
author: Tin Dang
Claude-Session: https://claude.ai/code/session_018xZT2V9tTvjuGmPFjZw4MG
…cross · gate PASS

Human-led verify gate (risk: high · autonomy: conservative), decided by Tin Dang
2026-07-16 ("Ratify v2 + gate PASS"): §3 re-frozen @ v2 with the widened scope
(the suite-wide recipe-pin migration + the ch02 docs ripple — contract shape
unchanged), the post-freeze test changes sanctioned via the recorded re-cross,
§6 evidence filled (suite 3114 OK · live 3-call transcript · legacy render
zero-rewrite proof · 4-way engine md5 parity · EARNED refute-read · 3-lens
PASS), outcome PASS recorded — task done. ADR block harvested at done.

task: phase-collapse-3 (thin-engine-loop W2)
author: Tin Dang
Claude-Session: https://claude.ai/code/session_018xZT2V9tTvjuGmPFjZw4MG
TASK.fast.md.tmpl is gone. The fast lane is now a DERIVED render of the one
template: _strip_fast_sections(full) drops exactly the _FAST_SECTIONS heading
blocks (## 7 · OBSERVE + §6's Deep-checks / Live-verify / Refute-read /
Advisor-3-lens) and cmd_new_task splices a bare `fast: true` header — a strict
line-subset by construction, machine-checked. --oneshot adds its two headers
plus the §3 "### AI-verify record" block as splices of the same render.

Engine (add.py + add_engine/constants.py):
- _FAST_SECTIONS (constants) replaces _FALLBACK_TASK_FAST; the shared
  _FALLBACK_TASK gains the Ground SHA field the fast fallback carried.
- _render_template's fallback table shrinks to TASK.md; _strip_fast_sections
  + _AI_VERIFY_RECORD_BLOCK own the derivation mechanics.
- Pins re-aimed: ENGINE_MD5 8eaca350…, ENGINE_PKG_MD5 ed7bf3e1…; add.py /
  constants / templates byte-identical across canonical, bundle, and both
  dogfood twins; SEAMS _declared_scope re-pinned 5915→5972.

Template family lean-pass (machine-read lines pinned by the new suite):
- TASK.md.tmpl 12209→12015 B while GAINING the §1 Boundary: line — the wm2
  input-dialect floor (boundary_unfilled) now fires on BOTH lanes at freeze —
  and natively carrying `phase: direction` (the render rewrite is a no-op).
- MILESTONE.md.tmpl 4211→3729 · PROMPT.persona.md.tmpl 3225→3046 ·
  personas/_template.md.tmpl 6922→5411.
- Skill trees (×3): fast-lane.md + SKILL.md now describe the derived render
  (§3 scope v2, re-crossed — the human ratified widening before the gate).

Suite: test_template_unify.py (7 red flipped + 6 floor pins) is the task's
red suite; test_fast_lane_template.py deleted (pinned the two-template
design end-to-end, superseded); ~12 files migrated — fast section-set
{1,3,4,5,6}→{1,2,3,4,5,6}, fast-template file pins → derived-render pins,
6 freeze fixtures fill the now-universal Boundary line. Full suite green:
3092 tests OK ×3 runs; add.py check 930/0.

task: template-unify (thin-engine-loop W3) — §3 FROZEN @ v2, gate PASS
author: Tin Dang
…demand references

An ordinary task now reads ZERO phase-guide files: SKILL.md (9,324B ≤ 9,500)
narrates the 3-beat loop INLINE — DIRECTION (draft §1–§4 → the ONE freeze,
`freeze --by <name> --cross`) · BUILD · VERIFY (`gate PASS`) — and the guides
became on-demand references, never a mandated per-phase load.

Skill trees (×3, byte-identical):
- phases/{0-setup,1-specify,3-plan,4-tests}.md fold into phases/direction.md
  (13,868B): setup span (init/4-lens/run-mode/lock) · rules+scenarios
  co-specification · plan (grounding/contract/build-strategy + the seven-item
  freeze review checklist) · tests span (the declare-where-tests-live grammar);
  every pinned anchor moved VERBATIM, one <exit_gate> per beat.
- phases/5-build.md → build.md · phases/6-verify.md → verify.md (retitled,
  most-pinned content intact); phases/fast-lane.md deleted — routing lives in
  SKILL.md flag mode. Pool 33,496 → 23,586B (−30%).

Engine (§3 v2, ratified by Tin Dang mid-build + re-crossed): _PHASE_GUIDE_FILES
re-aims to the merged 3-file shape — the re-aim phase-collapse-3's own comment
reserved for this task; `add.py guide` resolves playbooks again instead of
returning null on the new tree. ENGINE_MD5 349901707a…; 4-way twin parity.

Ripple (97 reds → 0): 35 test files re-aim paths; retired structure pins
(## Exit gate → <exit_gate> census, ## Run mode → **Run mode**, the SKILL.md
phase table → inline-beat asserts, §0 → §3 Grounding) re-pinned; dropped
teachings restored in the fold (teacher library pointer, invariants-in-§3,
data fixtures, subagent ground sweep, "you run" lock wording, a wrapped
run/entry-contract phrase). Roster agents ×3 trees, beyond/design/run/scope/
streams ×3, both glossaries (TASK.fast.md wording retired), and
SEMANTIC_INVENTORY re-aimed; wording-rubric keep-terms (Objective: · living
documentation) re-homed; SEAMS _declared_scope 5972→5971.

Suite: test_skill_loop_fold.py 8/8 (red-first); full suite 3100 OK;
add.py check 929/0; wording_lint 0 findings.

task: skill-loop-fold (thin-engine-loop W4) — §3 FROZEN @ v2, gate PASS
author: Tin Dang
…reeze ratifies it

The fitting persona now proposes a task's ceremony lane in the TASK header —
`route: <full|fast|oneshot> · routed-by: <persona:<slug> | human> — <why>` —
and the freeze IS the ratify: the direction→build cross records
state.tasks[slug].route = {lane, by}. Measure-not-block throughout: a missing
or unknown lane records "unrouted" and NEVER refuses the freeze.

Engine (add.py, ENGINE_MD5 ec9a5730…, 4-way twin parity):
- _ROUTE_LINE_RE/_ROUTE_LANES/_route_record beside the header-regex cluster;
  the record write sits at the flag_verified seam in _build_entry
  (UNCONDITIONAL overwrite — a re-cross re-records).
- audit gains route_unrecorded (unrouted record, or a ratified record whose
  header line was deleted post-freeze — state is the witness, mirroring
  unflagged_freeze) + route_lane_mismatch (ratified lane contradicts the
  task's actual fast/oneshot flags). Grandfathered by key ABSENCE — no
  pre-feature record is ever retro-redded.

Doctrine (SKILL.md 9,480B ≤ 9,500, ×3 trees): flag mode shifts from
"two human-owned settings (never auto-picked)" to propose-then-ratify —
a route bullet teaches the header line; the duplicate todo jotting line
funded the bytes. direction.md's freeze paragraph names the route ratify.
No render change: a per-lane spliced line would break template-unify's
strict fast-subset pin, so the persona writes the line, never new-task.

Suite: test_persona_routes_depth.py — 8 red flipped + 3 floor pins
(unrouted-freeze-crosses · grandfather · no-new-freeze-refusal); full suite
3111 OK; wording_lint 0; SEAMS _declared_scope re-pinned 5971→5990.
This task's own header carries the dogfood route line, recorded at its
own freeze.

task: persona-routes-depth (thin-engine-loop W5) — §3 FROZEN @ v1, gate PASS
author: Tin Dang
…plicated parity/engine-pin tests

test_tree_parity.py becomes THE canonical home for tree parity + the engine
pin: skill ×3 (incl _bundled) · agents add-*.md ×3 · tooling 4-way (add.py ·
engine_pin.py · add_engine/*.py · templates/**, exists-skip ≥2) · docs 3-tree
chapter mirrors · the ONE ENGINE_MD5/ENGINE_PKG_MD5 assert. Under it, the
duplicated per-task copies are deleted: 146 files changed, +251/−1713 lines,
suite 3111→2941 tests, full run OK ×2.

Census (pinned by the new test_corpus_slim.py red suite, all green):
- ENGINE_MD5-referencing test files 112→3 (sweep · security floor · census)
- parity-named fns 180→55 (≤55) · their files 126→43 (≤45)
- dead PHASES_POOL constants 8→0
- floor def-counts hold: freeze 9 · gate_audit 18 · scope_lock 31 ·
  security 5 · advisor_relax 29 · ai_plan_verify 44 · unflagged 13
- engine untouched: ENGINE_MD5/ENGINE_PKG_MD5 byte-identical (no repin)

The AST classifier's "0 behavioral suspects" was refuted by a killed-fn
audit: 16 guards bundled inside parity-named fns were restored (NOTICES
attribution, engine hands-off scans, v16 tag census, cross-tool roster sync,
plugin-version lockstep, template field pins, cospecify anchors, no-render
invariant, wave-verify census, io_state manifest membership). Collateral
fixed: release-1.11.0 CHANGELOG anchor un-reworded; untrack-add-tooling
meta-tests re-aimed at the sweep; SEAMS.md engine-md5-repin +
three-tree-parity anchors re-aimed (scope amended, re-crossed by Tin Dang);
ci_tooling_mirror_gap's vacuous-at-HEAD "untouched by this build" git-diff
assert dropped (CI-shape guards kept). test_shared_engine_pin.py deleted —
its scans fell as duplicates and the survivor was a meta-runner; the
pin-literal scan retirement is ledgered as a §7 open delta.

Probe: one corrupted byte in _bundled SKILL.md reds exactly the sweep while
the former duplicate suites stay green — detection consolidated, still firing.

Task: .add/tasks/test-corpus-slim (frozen @ v1 by Tin Dang, gate PASS,
route: full · routed-by: persona:tdd-verifier · sensitivity: mechanical).

author: Tin Dang
…come measurable (ADD 2.0 M1)

ADD 2.0 makes personas the method's core value; this lands the measurement
substrate. A closed task-kind taxonomy (constants.TASK_KINDS, 10 kinds) is the
join key between a persona's routing claim (`task-kinds:` frontmatter, new
template slot) and a task's declared kind (`kind:` header line, same anchored
grammar family as route:/sensitivity:, read by the new PURE _task_kind).

Every recorded gate outcome now appends ONE JSON line to
.add/traces/route-outcomes.jsonl (_append_route_trace): ts · task · milestone ·
kind · lane (route record, oneshot-marker fallback) · routed_by · persona
(parsed from a persona:<slug> routed-by) · outcome · heals · recross ·
age_hours · actor. Engine-derivable fields only; degrade-safe — the trace is
telemetry, state stays the source of truth, a failed write never blocks the
verdict. This is the evidence stream the persona scoreboard and the GEPA fold
(M7) will read.

_persona_quality_warnings gains Finding C: a task-kinds value outside the
taxonomy is a named WARN (measure-not-block, mirrors Finding A's flow check).

Red/green: test_persona_task_kinds.py (12) + test_route_trace.py (8) written
red-first; full suite green after twin sync + repin (ENGINE_MD5 26d1db26,
ENGINE_PKG_MD5 991ce131) + SEAMS re-aim (_declared_scope 5990→6045,
_section_unfilled 100→101). Retired test_corpus_slim's task-scoped
ENGINE_MD5_AT_FREEZE guard — its purpose expired when that task gated PASS;
kept alive it would re-break on every legitimate engine change (this one was
the first). The durable pin lives in engine_pin.py + test_tree_parity.

refs: ADD 2.0 M1 persona-core (T1+T2 of 5) — next: roster distill 5→1,
capability seed templates, dogfood persona migration
author: Tin Dang
… loop

Evidence-driven cleanup of the test-run workflow (pytest --durations sweep,
2026-07-17): the corpus' wall-clock floor was ONE meta-test —
test_ci_tooling_mirror_gap's fresh-checkout test-job sequence (~145s, a suite
inside the suite) — plus a family of npm/node/pty/sleep-heavy IO suites that
serialize workers.

tooling/t (dev-only, NOT shipped — package.json files allowlist untouched):
  ./t            fast lane — parallel via ephemeral uv env (pytest + xdist,
                 worksteal), slow-IO + CI-mirror suites excluded: 2684 tests
                 in ~19s (was 287s serial full)
  ./t --full     whole corpus parallel (~168s) — the pre-commit/pre-PR bar
  ./t --serial   unittest discover fallback — CI parity, no uv needed

The fast lane is an ITERATION loop, never the merge bar: the excluded
installer/global/packaging/pty/lock suites still bind in --full and CI.

author: Tin Dang
…ary distilled to its runtime core (ADD 2.0 M1)

Personas are ADD 2.0's core value; this ships the seed library the per-project
seeding draws from. Twelve v2-format capability templates land under
tooling/templates/personas/ (53.6KB total vs the 256-file/12MB teacher lib they
distill): product-lead · software-architect · build-engineer ·
evidence-verifier · security-gatekeeper · ux-experience-lead · data-steward ·
release-manager · stream-orchestrator · quality-auditor · platform-engineer ·
technical-writer.

Three are MIGRATION templates — the workflow knowledge of engine verbs slated
for the 2.0 kernel trim survives as markdown playbooks with explicit HUMAN
seams: release-manager (release cut arc + graduation arc, 4 human seams),
stream-orchestrator (8-step wave loop, review-queue seam), quality-auditor
(fold ritual + spec compaction, 2 confirmation seams). A security HARD-STOP
stays unstrippable in every template.

Every template: v2 schema (task-kinds from the closed taxonomy, flow routing,
not-when sibling boundaries), verified against the live engine predicates
(_persona_missing + _persona_quality_warnings incl. Finding C) — zero bare
placeholders, zero Abilities sections, zero dying-verb pins. Synced across the
4 tooling trees; no ENGINE repin (package digest covers add_engine/*.py only).

Full fast lane + packaging/parity/persona suites green.

refs: ADD 2.0 M1 persona-core (T4 of 5)
author: Tin Dang
… agent (ADD 2.0 M1)

Personas carry the expertise; the agent carries the discipline. The 5-agent
roster (add-design/build/verify/persona/advisor, 25KB) collapses into one
agents/add.md (~5KB): the spawn prompt names the MODE (direction · build ·
verify · advise · persona), the agent loads that beat's phase guide plus the
best-fit persona by frontmatter (flow + the new task-kinds), and one shared
boundary floor binds every mode (freeze/gate/lock stay human seams; security
always HARD-STOP; verify mode still routes flow:verify first, advisor
fallback).

Ripples closed in the same change:
- constants.py PHASE_AGENT -> all phases prefer "add"; PERSONA_HINT/
  PERSONA_FIT_HINT reworded to "the add agent in persona mode".
- guidelines.py block roster -> 1-agent + modes (sync-guidelines re-ran:
  CLAUDE.md/AGENTS.md/.clinerules regenerated).
- Installer tombstones: _installer._RETIRED_AGENTS + cli.js RETIRED_AGENTS
  name the 5 retired files so `update` removes them from the SHARED
  .claude/agents namespace (tombstone-only — user agents never swept).
- PROMPT.persona.md.tmpl: the 10-row per-runner adapter table is RETIRED,
  replaced by a runner-agnostic spawn contract of four general slots —
  prompt template · persona · model · isolation (user-directed); Claude Code
  stays the one verified reference.
- Skill guides (advisor/streams/beyond/phases-build) + TASK.md.tmpl route to
  the one agent's modes.
- Test migrations: roster_portable rewritten to the 1-agent bidirectional
  contract (block modes == agent-file mode bullets, retired names rejected);
  roster_shipped/tree_parity/phase_bundles/verify_flow_value/strategy_facets/
  fold_persona_sections/installer_shared_namespace/packaging pins migrated;
  subagent_prompt pins the four contract slots.
- Repins: ENGINE_MD5 9433f6e3 · ENGINE_PKG_MD5 d82eeae0; SEAMS
  _declared_scope 6045->6048.

FULL corpus green: 2956 passed, 3 env skips.

refs: ADD 2.0 M1 persona-core (T3 of 5)
author: Tin Dang
…0 M1)

The 6 project personas each declare their slots in the closed TASK_KINDS
taxonomy (the persona scoreboard join key): book-technical-writer=docs ·
method-product-owner=feature,integration,docs · methodology-engine-dev=
feature,refactor,infra · security-gatekeeper=security · tdd-verifier=
test,feature · terminal-ux-accessibility=ui. All 6 schema-conformant +
quality-clean (Finding C validates the vocabulary). Abilities engine-verb
scrub deferred to the M5 kernel trim, when the verbs actually retire.

refs: ADD 2.0 M1 persona-core (T5 of 5 — milestone COMPLETE)
author: Tin Dang
…get, the gate judges it, shards are free (ADD 2.0 M2)

The plan becomes ADD 2.0's core artifact. Three moves, red-first
(test_plan_target.py, 7 tests; route-trace schema contract extended):

1. Measurable Target in the plan. TASK.md.tmpl §3 Contract gains a
   `Target (measurable):` line — the success bar the §6 verify evidence must
   hit, numbers not adjectives. Measure-not-block: a §3 without it still
   freezes. direction.md drafts it; the exit gate lists it.

2. target-hit at the gate. `gate <outcome> --target-hit yes|partial|no`
   records the Target judgment in state (tasks[slug].target_hit) and in the
   route-outcome trace (target_hit key — completing the persona scoreboard
   schema begun in M1). An invalid value refuses BEFORE any write
   (target_hit_invalid); absence stays null, never inferred.

3. Shard-tolerant task folder, pinned as a 2.0 contract. AI-architected
   shard files inside .add/tasks/<slug>/ (notes, evidence, sub-plans) never
   trip the §5 scope guard — the .add tree is outside the scope walk by
   construction; ShardToleranceTest makes that a contract, not an accident.
   SKILL.md: "One file = one task" -> "One plan = one task" (TASK.md is the
   engine-known spine; the AI owns the shard architecture beside it).

Budgets held by compression, never bumped: SKILL.md 9494B (<=9500; funded by
trimming the duplicated graduation cue + a book-pointer clause), TASK.md.tmpl
12196B (<=12209 family ledger). Repin ENGINE_MD5 9cc73f6e; SEAMS
_declared_scope 6048->6056. The physical TASK.md->PLAN.md rename is deferred
to M6 migrate (renaming twice would churn hundreds of tests for zero
behavior).

FULL corpus green: 2963 passed, 3 env skips.

refs: ADD 2.0 M2 plan-core-shards (complete) — next: M3 specs-5dd
author: Tin Dang
…ernel verb (ADD 2.0 M3)

The foundation becomes five LIVING spec files and lessons land in-flight,
the moment they are learned — not batched at milestone close. Red-first
(test_specs_5dd.py, 9 tests):

1. Living specs. `init` seeds `.add/specs/{domain,system,experience,quality,
   method}.md` (DDD · SDD · UDD · TDD · ADD) from ONE template —
   templates/specs/SPEC.md.tmpl rendered five ways (template-unify
   discipline) — never-clobber, never blank (the SETUP_FILES survivor
   idiom, one shared _seed_spec_file truth). Each spec: a CURRENT "Now" +
   "Decisions that bind" picture above a "## Deltas (newest first)" inbox.

2. delta-append — the last unbuilt verb of the ratified 2.0 eight-verb
   kernel. `add.py delta-append <dd> "<lesson>"` routes via the closed
   constants.SPEC_DDS map, prepends one `[open · <date>]` line directly
   under the Deltas heading (newest first), stamps the active task (or
   --task; none -> no stamp, never inferred). Unknown dd refuses BEFORE
   any write (delta_dd_unknown). Legacy tolerance: a pre-2.0 project with
   no .add/specs/ gets the target file seeded on demand — the verb never
   dies on a missing dir (dogfooded on this very repo).

3. Wiring. SKILL.md observe beat points lessons at the verb (9497B, under
   the 9500 ceiling — funded by compressing the same paragraph, never
   bumped); deltas.md documents the in-flight channel beside the frozen
   grammar; min-pillar read-spy census covers the new verb.

Twin trees synced (tooling x4, skill x3); repin ENGINE_MD5 11fe18db +
ENGINE_PKG_MD5 cd2d7e81; SEAMS _declared_scope 6056->6085.

refs: ADD 2.0 M3 specs-5dd (complete) — next: M4 skill-unify
author: Tin Dang
…f is the receipt (ADD 2.0 M4a)

Intake gains the inline lane — the route for a change too small to deserve
versioned scope. Red-first (test_inline_lane.py, 8 tests):

- intake.md: `## The inline lane` sits between the interview and the frozen
  `## The four buckets` — the lane is judged BEFORE bucketing, because
  buckets create scope and the lane exists precisely so none is. Fit rubric
  (one file / covered behavior / no new contract surface / mechanical
  sensitivity) -> no task, no milestone: make the edit; the receipt is the
  git diff + `add.py delta-append <dd>` into the living 5-DD spec (specs-5dd
  M3 verb — the spec diff IS the approval artifact).
- The floor is closed: security · data · architecture ALWAYS escalates to a
  real task (security stays HARD-STOP); the human's "make it a task"
  overrides the route, always. When in doubt, bucket.
- SKILL.md routes to the lane at the intake beat — 9493B, under the 9500
  ceiling, funded by four pin-checked compressions (each candidate phrase
  grepped against the test corpus before cutting; the one hit bound engine
  output, not SKILL.md). Line-wrap kept "inline lane" unsplit (phrase-pin
  hazard).

Frozen intake pins survive (test_intake_interview stays green). Skill twins
synced x3; no engine change — no repin.

refs: ADD 2.0 M4 skill-unify commit A — next: M4b guide-fold (8 zero-engine
guides fold into the 3 beat references)
author: Tin Dang
Save the measured-campaign verdict as the canonical user-facing explainer
for when (and why) ADD beats spec-kit: enforced vs advisory discipline,
the context-rot finding, honest ties and spec-kit wins, the enforced-
guarantees table, and testable predictions (weak models, hostile changes,
autonomous runs, teams). Linked from both READMEs and the book nav
(Appendix H, after Appendix G).

author: Tin Dang
@TinDang97
TinDang97 force-pushed the feat/adaptive-flow branch from 25e8635 to 317b65f Compare July 18, 2026 18:48
Leaderboard-class evidence pipeline: benchmark/swe/runner.py runs the
SAME pinned agent (claude -p, sonnet-5, stream-json meter) on SWE-bench
Lite instances in two arms — vanilla issue prompt vs ADD 2.0 installed
into the checkout with the loop-driving prompt. Per instance: clone
repo@base_commit, agent run, tracked-file `git diff <base>` filtered of
method artifacts (.add/, .claude/, CLAUDE.md, ...) so the prediction is
the FIX only, appended to predictions_<arm>.jsonl (official shape,
resumable). Evaluation stays official: swebench docker harness (recipe
in the module docstring). Smoke slice = three psf/requests instances
(small repo, ids validated against the HF datasets-server).

15 offline guards pin the patch filter, meter argv, arm prompts, and
smoke config; bench suite 257 green. runs-swe/ gitignored like the
other runs roots.

author: Tin Dang
@TinDang97
TinDang97 force-pushed the feat/adaptive-flow branch from 317b65f to 245c005 Compare July 18, 2026 18:51
TinDang97 added 25 commits July 19, 2026 05:10
… honestly

Full matrix, all officially evaluated (swebench 4.1.0 docker harness):
sonnet-5 vanilla 3/3 $0.95 vs add 3/3 $5.28 (every ADD patch ships a
regression test; friendly ground ties, as the wm campaign predicted).
haiku-4.5 mini probe: vanilla 3/3 $0.90 vs add 2/3 $2.52 — appendix-h
prediction 1 NOT supported at n=3, published anyway: the ADD miss had a
correct fix (F2P green) but over-built and broke two PASS_TO_PASS tests
the gate never saw, because the loop bound only its own task tests as
the floor. Diagnosed method-integration gap: in a foreign repo the host
suite IS the regression floor and must be declared. Appendix H carries
the against-us data point with the diagnosis, same as the first
retraction.

author: Tin Dang
…etired

The SWE-smoke diagnosis (cost = authoring a 161-line PLAN.md for an 18-line
fix; --oneshot never stripped the heavy body) lands as the ATG-informed cut:
one file = one atomic node — persist the interface (contract · red suite ·
scope · verdict), reason everything else in-context.

Template (191→131 lines, authored surface ~60→~22 fields), every engine
anchor preserved verbatim:
- §3 gains `Regression floor:` — the HOST repo's own suite is ALWAYS an
  inherited edge (the haiku psf__requests-863 over-build fix, now method)
- AI-verify record ships IN the template (was an --oneshot splice); an
  agent-crossed freeze is declared via `gate_mode: ai-plan-verify` in the
  header — better audit than a flag
- §5 gains `Spawn (multi-agent):` — build/verify spawns default worktree
  isolation; freeze --cross + refute-read via a cross-agent advise-mode spawn
- removed: Grounding block, gherkin scaffold (§4 test_plan is the canonical
  encoding), assumptions ladder (ONE ⚠ line stays), 9 Build-strategy facets,
  Deep checks / Live-verify / Advisor 3-lens blocks, §7 Watch

Engine: --fast/--oneshot/--thin/--full argparse + _strip_fast_sections +
_FAST_SECTIONS + _fastlane_nudge/RISK_KEYWORDS + lane state markers deleted
(creation side); READ side stays migration-tolerant (legacy fast/oneshot
state keys still honored for skip-eligibility, route records measure-not-
block). ENGINE_MD5 3e7e03c0 · ENGINE_PKG_MD5 fa82d37a @ atomic-node; 4-way
tooling twins + 3-way skill/template trees byte-synced.

Teaching + bench surfaces follow: SKILL.md flag mode now teaches the single-
template doctrine + gate_mode (9,167B ≤ 9,500B ceiling); beyond/intake/verify
guides de-laned; wm-bench + SWE prompts drop --oneshot and instruct the
gate_mode declaration; the SWE ADD prompt names the host suite as the §3
Regression floor before the gate.

Tests: 13 lane/grounding suites retired, ~40 re-aimed, test_template_atomic
added (frozen 6-tag census · comment balance · 21 engine anchors · retired-
surfaces stay retired · 3-tree parity). Corpus 2,290 passed / 3 env skips;
benchmark suite 258 passed; add.py check 412/0.

refs: #163 (feat/adaptive-flow) · benchmark/results/2026-07-swe-smoke.md
author: Tin Dang
… into §3 Target

The atomic-node follow-through: the pre-registration job the block carried is
triple-covered — §3 `Target (measurable)` (judged at the gate via --target-hit),
§3 `Regression floor`, and the §6 refute-read. In SWE transcripts agents filled
it with "tests pass — confirmed by pytest": zero information past what the gate
already checks. Its one unique remainder — pre-declaring outcomes tests can't
show — folds into the Target guidance ("name any outcome tests can't show
(boots · renders) + how it's confirmed").

Template (4 twins): `### Build expectations` block removed; Target line extended.
Engine (4 twins): the opt-in `build_expectations_unfilled` gate retired — with
the template no longer scaffolding the block the gate is unfireable for new
plans; legacy plans with filled blocks pass unchanged, unfilled ones now cross
(measure-not-block direction). `_section_unfilled` survives serving the
contract-fill gate only. ENGINE_MD5 8d44e6ed · ENGINE_PKG_MD5 ec7f8093.

Teaching (3 skill trees): SKILL.md beat-1 drops the §6 mention (9,141B);
direction.md tests-production list trimmed; verify.md fill-before-build section
removed + stale Part-one drift cleaned (checkboxes still citing the retired
Build-expectations / Ground SHA / Live-verify surfaces re-aimed to §3 Target).

Tests: test_build_expectations_gate retired (predicate stays covered by
test_contract_fill_gate + test_engine_extract_predicates); freeze-precedence,
refute-ordering, form-tags, verify-rollup, template-atomic suites re-aimed
(block joins the retired-surfaces guard). SEAMS _declared_scope pin 4434→4427.

Corpus 2,281 passed / 3 env skips · benchmark 258 · add.py check 416/0.

refs: #163 (feat/adaptive-flow)
author: Tin Dang
…state.json

The SWE installer dropped .add/tooling but never the engine state, so an agent
that skips `add.py init` (haiku does) makes root discovery walk UP and find the
HOST checkout's .add five levels above — the whole loop then runs against this
repo (observed live: haiku wrote its psf__requests-863 task, src and tests into
the host .add as fix-hooks-issue; yesterday's fix-issue/fix-hooks-lists residue
was the same leak, not a smoke walk). install_add now ends with the engine init
(`add.py init --name swe-fix --stage mvp`), anchoring discovery in the
workspace; a WorkspaceIsolationTest guard pins the step.

refs: #163 · sibling of harness-workspace-isolation a87ed1e (wm bench)
author: Tin Dang
…on-1 resolution

benchmark/results/2026-07-atomic-remeasure.md: the atomic-template re-run vs
recorded baselines, officially evaluated —
- SWE smoke: haiku 2/3 $2.52 -> 3/3 $2.15 (863 resolves; patch 1626B->559B;
  the host-suite Regression floor did it); sonnet 3/3 both templates at ~half
  the wall-clock (50.9 -> 28.1 min); first-eval 2317 miss shown to be a
  live-httpbin flake by a single-instance re-eval (zero P2P on re-run).
- Cross-milestone session bench (ONE continuing conversation, 6 WMs): fidelity
  flat 1.0×6 vs old conv-carry .92->.80->.75->.17->.75->1.0 — rot eliminated
  in this sample; $17.75 vs $22.67; tests_weakened flags audited benign.
- Harness defect record: the workspace-seed leak (fixed 9336076) and which
  runs are contaminated vs canonical.

appendix-h: prediction-1 against-us note gains its resolution paragraph (the
diagnosis became method — Regression floor line — and the re-run flipped the
score back); swe-smoke report gains a superseded-for-ADD pointer, vanilla
numbers + diagnosis stand.

refs: #163
author: Tin Dang
…ta point

atomic-remeasure gains the continue-mode spec-kit table (recorded arm, not
re-run): spec-kit rots .92->.75 by wm3 AND ships regression rates .33/.29 in
the mode where atomic ADD holds flat 1.0 with zero regressions — the honesty
wave's "both arms decay identically" verdict no longer holds. Honest bounds
kept: ~1.7x price gap survives; fresh-mode matrix un-rechallenged; appendix-H
bottom line does not flip.

appendix-h: second data point paragraph (continuing-conversation mode) beside
the prediction-1 resolution.

refs: #163
author: Tin Dang
…interfaces

Milestone step 3 (ATG localized context). Measured driver: the session bench's
late-WM cost growth (wm1 $1.87 -> wm4 $3.67, turns 74 -> 158) is interface
RE-DISCOVERY — the agent re-reads the grown app to recall what prior tasks
shipped. The engine now prints a `neighborhood` card at new-task and full
status (never --brief): one line per inherited interface — parent slug ·
phase · the head of its frozen §3 contract fence · where its code lives.

Parents = declared edges (depends_on ∪ extends); a board with NO edges falls
back to the 2 most recently updated DONE tasks (the temporal neighborhood —
bench agents declare no edges, measured across every board). Degrade-safe:
unreadable or still-placeholder parent contracts are skipped; empty
neighborhood prints nothing; card caps at 3 lines. SKILL.md beat 1 teaches
"ground §3 in the card, not code re-reads" (9,252B ≤ 9,500).

test_neighborhood_status.py: 8 tests red->green (declared-edge card, recency
fallback, empty-board silence, --brief stays lean, unreadable/placeholder
skip, 3-line cap). ENGINE_MD5 0d98f693 re-pinned; add.py x4 + skill x3
synced; SEAMS _declared_scope 4427->4480.

Gate next: re-run the 6-WM session bench targeting <=$2.00/WM avg.

refs: #163 · milestone step 3 of the ratified atomic plan
author: Tin Dang
…ome real

task-graph-native W1. The milestone is the scope ROOT of the task graph
(depth = edges, not nesting) — but measured across every bench board the
graph was EMPTY (deps=[] everywhere): waves had nothing to schedule, repair
had no dependent closure, the neighborhood card fell back to recency. Three
deterministic, propose-not-block mechanisms make it real:

- compile — `milestone-confirm` reads MILESTONE.md's `## Tasks` list
  (`- [ ] <slug>   depends-on: <none|slugs>   — <line>`) into
  state.milestones[m].planned (the figure's "Compilation of T0"); prints
  `compiled task graph: N nodes · M edges`. Re-confirm RECOMPILES — the plan
  is living. Placeholder/malformed lines skip silently; the bare scaffold
  compiles to nothing.
- inherit — `new-task <slug>` with no explicit --depends-on inherits the
  planned deps VERBATIM (creation-order-proof; a dangling forward edge is
  check's existing warn, never a lost edge). Explicit --depends-on wins.
- hint — `freeze` prints `edge-hint:` when the just-declared §3 scope
  overlaps a DONE task's scope with no edge between them (the _declared_scope
  grammar both sides, containment, cap 2, silent on UNDECLARED) — a proposal
  the human/agent ratifies, never a refusal.

SKILL.md intake line teaches the compile (9,360B ≤ 9,500). test_edge_truth.py
12 tests red->green. ENGINE_MD5 re-pinned; add.py x4 + skill x3 synced;
SEAMS _declared_scope -> 4560.

Next: W2 graph-repair (locate + minimal dependent closure on contract change).

refs: #163 · task-graph-native W1
author: Tin Dang
…fy path

task-graph-native W1.5, closing the hole W1 exposed: the freeze edge-hint
proposed edges with no clean way to declare them post-creation (new-task
flags are too late; MILESTONE.md re-confirm only updates the planned map,
never existing records).

- `add.py relate <slug> --depends-on/--extends/--relates-to <slugs>` (verb
  32): ADDITIVE (append + dedup, never drops); validate-then-write on the
  source; dangling targets legal (forward edges — check's warn owns them);
  self-edge refused. The edge-hint now names it verbatim as its ratify step.
- `milestone-confirm`'s compile warns on a depends-on cycle (would deadlock
  the wave schedule) — measure-not-block: the confirm stands, the fix is an
  edit + re-confirm.

test_edge_truth.py 12->19 tests red->green. ENGINE_MD5 79b12cd4; twins x4
synced; SEAMS _declared_scope -> 4615.

refs: #163 · task-graph-native W1.5
author: Tin Dang
…corrected

test_min_pillar's self-maintaining census caught what the chained commit
missed: verb 32 (`relate`) never ran under the read-spy. It now does —
`new-task t2` + detach (`set-milestone t2 none`, so milestone-done stays
green) + `relate t2 --relates-to t`. The CI fresh-checkout mirror suite
reds until this lands (its clone runs HEAD's census against HEAD's parser).

Also corrects the W1.5 docstrings: a dangling edge target is NOT merely
"check's warn" — `add.py check` REDS a dangling ref until it resolves
(no gate ever refuses; the diagnostic is honest, the flow unblocked).
ENGINE_MD5 2d28aebe; SEAMS re-pinned.

lesson: run the FULL corpus BEFORE `git commit` in the same chain — a
census suite red after push costs a fix-forward commit.

refs: #163 · task-graph-native W1.5 follow-up
author: Tin Dang
… returned

Reverts d5a2879. The card was gated on a re-run of the 6-WM continue-mode
session bench (runs-nbr-session); it failed every axis vs the card-free
atomic baseline (runs-atomic-session): fidelity 0.92->0.80->0.75->0.33->
0.75->1.0 (vs flat 1.0x6), regression rate 0.20-0.38 at every WM (vs 0),
total $23.57 (vs $17.75). The decay is a near-exact replay of the
pre-atomic rot trajectory the atomic template had eliminated.

Mechanism (from the transcripts, recorded in
benchmark/results/2026-07-atomic-remeasure.md): the card externalizes
PRECEDENT, not spec — wm2's card quoted wm1's shipped contract, deviation
included — and its 90-char fence-head quote degenerates to auth
boilerplate when a parent's S3 fence opens with the auth line (wm4 got
zero interface signal and cratered to 0.33). Disclosed: edge-truth
commits landed mid-bench so resolved_pin drifted across WMs; additive,
unlikely causal, but the run is not pin-clean. Isolation was clean.

Kept: edge-truth graph compile + relate verb (not part of the revert).
Localized context returns as graph-native work (locate + dependent
closure over declared edges), gated on its own bench.

ENGINE_MD5 3a191a55 re-pinned; add.py x4 + engine_pin x2 + skill x3
synced; SEAMS _declared_scope 4616->4563; test_neighborhood_status.py
removed; corpus 2300 passed.

refs: #163 · gate: runs-nbr-session vs runs-atomic-session
author: Tin Dang
… repair closure

Task-graph-native wave 2 (the ATG figure's Failure Location + Minimal
Necessary Subgraph Repair, deterministic — no LLM, read-only, verb 33).

`add.py locate <ref>` in two modes:
- test PATH -> the OWNING node, via each task's §4 `Tests live in:`
  declarations (reuses _declared_test_files) with the frozen §5 scope
  snapshot as fallback (provenance named), plus the failure class:
  `in-node` (owner still live — fix inside it; its frozen suite is the
  floor) vs `interface-regression` (owner DONE — a live change broke a
  settled contract; the owner's dependent closure prints as the repair
  set). Unowned paths report cleanly (exit 0 — a finding, not an error:
  treat as host/foreign regression-floor surface).
- task SLUG -> the dependent closure directly: BFS reverse reachability
  over depends_on ∪ extends (relates_to is context, not interface — never
  enters), per-ring sorted for deterministic output, depth = shortest
  interface distance, DONE dependents kept (settled work re-verifies when
  its foundation moves).

test_graph_repair.py: 11 tests red->green (owner mapping · both classes ·
closure-on-done-owner · scope fallback · unowned floor · transitive +
extends closure · relates_to exclusion · leaf message · done-dependent
kept). test_min_pillar LIFECYCLE gains `locate t` under the read-spy.
SKILL.md build beat teaches the verb (9,405B <= 9,500, x3 trees).

ENGINE_MD5 b2869fdd re-pinned; add.py x4 + engine_pin x4 synced; SEAMS
_declared_scope 4563->4658; corpus 2,311 passed pre-commit.

refs: #163 · task-graph-native W2 (W1 edge-truth 8d4f573 · W1.5 relate
0dadb17 · card revert f77f684)
author: Tin Dang
…s frozen §3 clause

Task-graph-native wave 3 (the ATG red-dot at clause depth) + the skill
taught to drive the atomic graph loop.

Engine: `locate` grows the pytest node-id form — `locate path::test_name`
resolves the test through §4's covers map to the frozen §3 clause LINE.
The map's grammar is the template's OWN `<test_plan>` dialect (a §4 bullet
naming a test bare `test_…` or backticked, with a `covers: key[, key]`
tail) — frozen WITH the bundle, so it is tamper-guarded like the suite it
describes. Deterministic literal key match inside the §3 body; a key §3
doesn't carry is reported honestly ("not literal — the clause lives in
§1/§2 prose"), never guessed. The scaffold's unfilled placeholder bullet
parses to nothing (angle-bracket keys filtered, fail-safe); an unmapped
test is a nudge, never a gate. Plain-path and slug modes unchanged.

Skill (x3 trees) follows the new atomic graph flow:
- SKILL.md: direction beat fills covers keys; build beat teaches
  locate path::test (9,486B <= 9,500).
- direction.md: "Clause map + edges" — covers tails + declare edges at
  creation + ground §3 on parent edges' frozen PLAN.md (the spec-side,
  pull-based grounding the reverted card got wrong).
- build.md: "A red outside your suite — locate first" (in-node ·
  interface-regression · unowned=Regression-floor; repair the clause,
  not the symptom).
- Template §4 comment documents the machine-read covers tail.
- Pool fence 33,496 held by compression, not a bump: connective fat cut
  in direction.md/verify.md (all cut phrases pin-checked unpinned;
  wording-lint 0 findings): pool 33,489.

test_clause_repair.py: 9 tests red->green (covers map · clause quote ·
template-native + backticked grammar · placeholder-safe · honest-missing
· advisory-unmapped · both mode floors). ENGINE_MD5 3a19e9dd re-pinned;
add.py x4 + templates x4 + engine_pin x4 + skill x3 synced; SEAMS
_declared_scope 4658->4710; corpus 2,320 passed pre-commit.

refs: #163 · task-graph-native W3 (W2 locate 49e588d)
author: Tin Dang
…rns on planned drift

Task-graph-native wave 4, closing the ratified milestone's engine work.

`add.py graph` (verb 34, read-only, print-only) renders the live board as a
mermaid flowchart: each milestone is a subgraph wrapping its tasks (milestone
= the scope ROOT; depth lives in edges, never nesting); edge style carries
the edge type (depends-on solid --> · extends dashed -.-> · relates-to open
-.-); node class carries phase (done green · live amber). The COMPILED plan
renders too: a planned-but-never-created node appears dashed, and an edge
target that resolves to an archived record is annotated instead of dangling.
--milestone limits to one subgraph. Deterministic (sorted everywhere) —
paste straight into a GitHub mermaid fence.

The same drift is measured: `add.py check` gains a planned-drift WARN (never
red — mid-milestone a planned-not-yet-created node is normal flow) naming
each compiled `## Tasks` node with no live task and no archived record, with
the ratify path (`new-task <slug>` inherits its planned depends-on, or
re-confirm without it).

test_graph_views.py: 9 tests red->green (flowchart render · milestone
subgraph · three edge styles · phase classes · archived-target annotation ·
--milestone filter · dashed planned node · check warn + exit 0 · silent when
all created). LIFECYCLE census gains `graph`. Dogfooded on this repo's own
board (renders clean). No skill-pool spend: the verb is --help-discoverable;
SKILL.md (14B slack) and the phases pool (7B slack) stay untouched.

ENGINE_MD5 427a2501 re-pinned; add.py x4 + engine_pin x4 synced; SEAMS
_declared_scope 4710->4796; corpus 2,329 passed pre-commit.

refs: #163 · task-graph-native W4 (W1 8d4f573 · W2 49e588d · W3 75bb128)
author: Tin Dang
Review finding pre-2.0.0 tag: `.add/dependencies.allowlist` claimed "CI
rejects anything not listed" while NOTHING read the file (every corpus
"allowlist" reference is npm's `files` tarball allowlist — a different
thing), and its "Node installer uses built-in modules only" prose was
false — package.json ships `@clack/prompts` as a real runtime dependency.
The build exit gate "no dependency outside the allow-list" rested on an
honor-system doc.

Fix, red/green:
- test_dependency_allowlist.py (3 tests) IS the missing rejection — the
  corpus runs in CI, so a declared runtime dep absent from the allowlist
  now reds the build. Scope = RUNTIME deps of the shipped package (npm
  `dependencies`, pyproject `[project] dependencies`); dev/bench tooling
  never ships and stays out of scope. Third test pins doc-truth: the
  zero-dep-installer claim may not coexist with declared npm deps.
- .add/dependencies.allowlist: @clack/prompts recorded as the ONE
  approved Node runtime dependency (shipping since the installer-UX
  milestones — recording the approval is the honest state); prose
  corrected; header now cites the enforcing suite.

No engine change (no ENGINE_MD5 repin). Corpus 2,332 passed.

refs: #163 · user review ask pre-release
author: Tin Dang
…ges + doc-truth

Review of the updater (both twins) surfaced three gaps; fixed red/green
(test_updater_2_0_gaps.py, 7 tests, node twin via subprocess):

1. GLOBAL ROSTER DRIFT (open since installer-shared-namespace #151): the
   global home mirror (GLOBAL_TREES / _GLOBAL_TREES) never carried
   `agents/`, and `update --global` propagation sources registered
   projects FROM the home — so the roster soft-skipped forever: no
   refresh, no retired-agent tombstone removal. Both mirrors now carry
   agents/; propagation reuses the existing SHARED per-file lander
   (user files in .claude/agents still survive; only the five explicit
   RETIRED_AGENTS tombstones are ever removed).

2. 2.0 CROSSING WAS SILENT: a 1.x project updating into 2.0 kept a
   TASK.md board and a stale CLAUDE.md guidance block with no signpost.
   `update` now prints crossing nudges — `add.py sync-guidelines` on any
   version cross, `add.py migrate` when a tasks/*/TASK.md without a
   PLAN.md sibling is detected. NAMED, never run: python3 may be absent
   on the updater's PATH; both commands are idempotent. Twin-identical
   wording npm/pip.

3. DOC-TRUTH: pip updater log claimed "docs refreshed" (docs stopped
   shipping at book-stops-shipping) -> "managed layer reconciled",
   matching the js twin; stale "Empty today." comment on the populated
   RETIRED_AGENTS list removed.

No engine change (add.py untouched — no ENGINE_MD5 repin). Corpus 2,339
passed pre-commit.

refs: #163 · user review ask pre-release · closes the #151 OPEN residue
author: Tin Dang
…one fits

User ask: spawn/seed a persona for each new task when missing. Implemented as
domain-fit (NOT one-persona-per-task, which would sprawl near-duplicates):
each task's DIRECTION beat establishes a persona that fits its domain — seeding
one via the add agent's persona mode ONLY when none fits, reusing the existing
persona across tasks otherwise.

- SKILL.md beat 1: "load the domain-fit persona (seed via add persona-mode if
  none), then draft the bundle" — the always-read file now leads each new task
  with persona fit.
- phases/direction.md persona callout: if none fits, spawn add persona-mode to
  seed from PROJECT.md + `.add/personas-teacher/`, then load it — seed per
  DOMAIN, REUSE across tasks, never one per task (the anti-sprawl floor).

SKILL-only — no engine change, no ENGINE_MD5 repin. Both byte fences held by
compression (SKILL 9,496 <= 9,500; phases pool 33,494 <= 33,496): connective
fat cut in SKILL.md (--todo/flag-mode/lessons lines) and direction.md (milestone
ground + relate-to-map items) — every cut phrase pin-checked unpinned;
wording-lint 0 findings.

test_persona_seed_on_task.py: 3 tests red->green across all 3 skill trees
(beat-1 fit persona · direction seeds-when-none-fits via persona mode · anti-
sprawl reuse). LESSON re-applied: the persona callout line-WRAPS "none fits"
across a blockquote break — the assertions normalize whitespace (drop `\n>` +
collapse) so a wrapped phrase-pin still matches (line-wrap-splits-phrase-pin).

skill x3 synced; corpus 2,342 passed pre-commit.

refs: #163 · user ask · builds on the persona-seed-nudge engine surface
author: Tin Dang
…st orchestrating

Review the /add SKILL.md against the skill-creator standard and apply the
findings as byte-safe edits across all three byte-identical skill trees
(canonical · _bundled · .claude mirror).

Front-door capability (the headline): ADD now reads raw intent into a task
shape BEFORE sizing it. New intake.md beat "Analyze the request before you
size it" — restate the intent, extract the latent requirements, name the
unstated, surface the hidden work — placed before Interview. The SKILL.md
opening reframes the agent from "You are the orchestrator" to "You turn intent
into the right task, then drive it": analyst first, orchestrator second.

Review findings applied:
- #1 design.md: udd-tokens.md / udd-catalog.md pointers were dangling siblings;
  qualified to templates/… so they resolve.
- #2 SKILL.md setup branch now names the brownfield adopt.md path (one hop).
- #3 new terms.md decodes the loop's coined vocabulary (compound-cross,
  co-specify, earned-green refute-read, re-cross, auto-resolved PASS, the ARC);
  linked from SKILL.md.
- #5 7↔3 bridge: "three beats (seven steps, folded)" maps the description's
  seven named steps to the body's three beats.

Constraints held: SKILL.md 9,496 → 9,497 B (under the test-enforced 9,500
ceiling — every addition funded by compressing unpinned prose, no pillar
removed); all phrase-pins preserved; the three SKILL.md trees stay md5-identical.

New guard: test_intake_analyze.py locks the Analyze section, its order before
Interview, the reframed identity, the terms.md decoder link, and the qualified
design.md refs. Full tooling suite green (2350 passed); wording-lint and
semantic-inventory clean.

author: Tin Dang
…autonomy dial

Run mode was two coupled halves: the autonomy dial AND an engine-managed
`streams:` posture (parallel/sequential) that `--run-mode` forced in lockstep
(auto→parallel, conservative→sequential). Real usage rarely runs parallel
agents, and the simplified Run-mode guide already frames concurrency as
"spawn a subagent per task" — not an engine posture. So the engine no longer
manages streams: run mode IS the autonomy dial; concurrency is doc-level.

Engine (removed the run-mode streams posture only — the separate multi-milestone
`streams : N active milestones` status view is untouched):
- constants.py: drop _STREAMS_POSTURES.
- autonomy.py: drop _streams_posture / _project_streams_token / _project_streams
  and the streams regex.
- add.py: `--run-mode {auto,conservative}` now writes ONLY `autonomy:` (no
  streams line); remove the _streams_decl_line writer and the status `run mode:`
  row (redundant with `project autonomy:`); tidy the stale mirror comments.

Doc: phases/direction.md Run-mode block simplified — the vacuous Concurrency
column dropped (both rows were "one task"), the table axis renamed Autonomy,
and the tangled parallel-streams prose replaced by one line: spawn a subagent
per task; the one-approval-per-contract floor never moves. Pool 33479B (<33496).

Tests (red/green): test_setup_run_mode rewritten — --run-mode sets autonomy and
writes NO streams line; test_streams_posture.py deleted (tested only the removed
posture). Repin: ENGINE_MD5 + ENGINE_PKG_MD5 re-aimed @ run-mode-decouple,
synced across all engine trees. SEAMS.md scope-token-grammar anchor re-aimed
4796->4776 for the add.py line shift. Full tooling suite green (2346 passed).

author: Tin Dang
…e/graph model, trim dead rule

The 2.0 atomic template (`PLAN.md.tmpl`) shifted ADD to "persist the interface,
reason everything else in-context, don't write essays" and retired lane modes —
but the phase guides and the skill still taught the pre-atomic, essay-heavy,
lane-aware model. This aligns them and trims the dead weight the shift exposed.

direction.md — atomic-mode alignment (no split; the fold stays):
- Grounding section reframed from a 7-field written block ("gather BEFORE you
  freeze") to the atomic rule: persist only the Contract's Anchors + an optional
  Ground SHA; REASON Touches / Honors / Issues / Related-intent in-context and let
  the frozen Contract encode them — don't transcribe an essay into the file.
- Dead lane framing corrected: the freeze no longer "ratifies the header route:
  line — the persona's lane proposal" (lanes are retired, 0b0192a). route: is now
  an audit-only optional line the freeze records; route_unrecorded is still measured.
- Grounding exit-gate line realigned to "reasoned in-context; the interface is
  persisted, not the bullets."

SKILL.md — orchestrator -> atomic-graph navigator, and a dead rule dropped:
- The node paragraph now leads with "One task = one atomic node" and states the
  graph model up front: the frozen §3 is the interface neighbor nodes depend on;
  edges compile from the milestone; `graph` renders the DAG; `locate` walks a
  failure to its node. Graph navigation is first-class, not a "stuck?" footnote.
- Removed constraint 5 "Ask, don't guess": unpinned, not ADD-specific (generic LLM
  hygiene), already the always-loaded CLAUDE.md floor, and counter-directional to
  the repo's own proactive-lead posture. Rules 1-4 (the freeze gate, evidence-over-
  inspection, the tamper floor, the one recorded outcome incl. security HARD-STOP)
  stay.
- `## The method rationale` collapsed from a 5-line section to one line.

Net: SKILL.md 9,497 -> 9,403 B (97 B of ceiling headroom reclaimed under the
9,500 test-enforced ceiling); the three SKILL.md / direction.md trees (canonical ·
_bundled · .claude mirror) stay byte-identical.

Verification: full tooling suite green; wording-lint 0 findings; semantic-inventory
unchanged (17 pre-existing findings on HEAD, 0 introduced). No test weakened; the
security-HARD-STOP teaching and every phrase-pin anchor survive.

author: Tin Dang
…oad-bearing

Phase 3 of the atomic-mode trim. Assessed verify.md (the second-largest phase
guide, 162 lines) for dead weight and found it is largely load-bearing: the
evidence checklist, the three lenses, the deep-check, the earned-green refute-read,
the gate-outcome table, the observe tail, and the ~55-line advisor spawn XML (the
canonical runner-agnostic worker contract, pinned by test_persona_subagent_prompt:
the four contract slots, the persona-section mapping, the {{PERSONA_SLUG}} slot, the
no-persona degrade path) all earn their bytes.

The one compressible section was the `## Sensitivity` risk-class vocabulary — dense
connective prose around a load-bearing core. Tightened it without dropping a single
fact or pinned string: the `sensitivity:` declaration, the base-four definitions
(security HARD-STOP · data · architecture · mechanical), the datetime/money/timezone
=> `data` guidance with the bench wm2 evidence, the `## Sensitivity classes` EXTEND
grammar, `sensitivity_invalid`, and `advisor-gate-relax` all survive verbatim.

verify.md 162 -> 160 lines; the three trees (canonical · _bundled · .claude mirror)
stay byte-identical.

Finding (the "what to optimize" answer): verify.md's complexity is intrinsic, not
sprawl — further shrink would remove a pillar, not ceremony. The atomic-mode wins
were in direction.md and SKILL.md (prior commit 3cd8792).

Verification: full tooling suite green; wording-lint 0 findings; no test weakened;
security-HARD-STOP teaching preserved.

author: Tin Dang
…(atomic drift sweep)

A drift sweep across every skill guide for retired-concept references (lanes,
route-as-lane, streams:, TASK.md, §0 Grounding, §6 Build-expectations, phase-count
framing) turned up exactly one genuine stale reference the earlier commits missed:

build.md's strategy-facets line anchored the domain facets upstream at
"§1 Framings · §3 Schema · §0 Honors". The 2.0 atomic template numbers §1–§7 — there
is no §0; the old §0 Grounding/Honors block folded into §3 (the engine still reads
`Honors`/`Ground SHA` by line-regex, section-agnostic, so nothing breaks). Re-aimed to
"§3 Schema · CONVENTIONS.md Honors" — pointing at the real source artifact rather than
a section that no longer exists. Both pinned build.md anchors on that line
("Approach (domain strategy", "Optimization stance") preserved verbatim.

Everything else the sweep flagged is live, not drift: the "inline lane" (a current
below-scope feature), the route scoreboard's "per-lane" evidence (the shipped M7
feature; route: is now an audit-only line), and "three beats (seven steps, folded)"
(the deliberate 7-to-3 bridge). No skill guide now references §0.

verify: full tooling suite green; wording-lint 0; three skill trees byte-identical.

author: Tin Dang
…es (verified pass)

A per-passage verification of the book's pre-atomic drift against current engine
behavior. Fixed only the genuinely-stale references; left every passage that touches
a still-live nuance alone.

Fixed (verified stale):
- 07-step-5-build.md — the Pattern facet cited "(§0 Honors / CONVENTIONS.md)"; the
  2.0 atomic template numbers §1–§7 with no §0, so the facet now cites
  "(CONVENTIONS.md Honors)" — the real source, mirroring the build.md guide fix.
- 08-step-6-verify.md — the live-anchor check was framed around "§0's Ground SHA …
  record it in §6's Live-verify evidence block". test_template_atomic pins that the
  `### Live-verify evidence` and `### Grounding` blocks must NOT exist in the atomic
  template — the block is retired. Rewrote the paragraph to describe the surviving
  check (re-resolve every §3-cited symbol against the current tree, vs the shape it
  had at freeze) without the retired §0/Live-verify machinery. The concept still
  lives in the verify guide's Part-one checklist.
- appendix-c-glossary.md — the phase-name table called grounding "the §0 grounding
  map"; dropped the stale §0 → "the grounding map".

Root mirror synced: these chapters are byte-mirrored at the repo root (guarded by
test_earned_green_rubric::test_root_book_matches_canonical for ch08); all three root
copies re-synced to canonical.

Deliberately NOT touched (live nuance / maintainer's call, flagged separately):
- The components pillar (ch17 + the appendix-c Component entry): add.py comments say
  the component:/produces:/consumes: header grammar and components.toml schema-lint
  "died with the components pillar" (kernel-trim M5), and `component_green_bar_uncited`
  is gone from the engine — yet the PLAN.md.tmpl header still advertises `component:`.
  That contradiction is the engine's to resolve, not a drift-fix; ch17's fate is an
  editorial call.
- "fast lane" glossary/component references: the `--fast`/`--oneshot` template
  scaffolds are retired, but `fast-lane-skips` (benchmark mode) and a `§0 GROUND`
  skip-rationale section are still live in the engine — so a blanket rewrite would be
  wrong.

verify: full tooling suite green; wording-lint 0; canonical↔root book parity restored.

author: Tin Dang
…task complexity

Simplify the setup Run-mode step's concurrency line and fold in model-tier selection:
a spawned subagent per task now picks its model by task complexity (mid ordinary ·
top complex), matching the roster tier contract in agents/add.md. The pinned run-mode
content stays intact — `autonomy set --project`, `init --run-mode`, the sequential/auto
table, parallel as the opt-in path, one-approval-per-contract floor.

All three skill trees stay byte-identical.

verify: full tooling suite green; test_setup_run_mode + parity guards pass.

author: Tin Dang
… distil the loop

The UDD design loop referenced its personas only at beat 4 (the confirm checklist),
and restated all five beats a second time in a `## The hard rules` constraints block.
This puts the personas first and cuts the redundancy — keeping the design-quality
performance while making the guide leaner.

Integrate (personas carry the expertise, the loop the discipline):
- New "Personas carry the design — the loop carries the discipline" preamble loads the
  design-fit persona FIRST — the frontend-designer performance — before beat 0. Names
  both flow:design dimensions (UI-Designer: visual systems · component libraries ·
  pixel-craft · WCAG-AA accessibility; UX-Researcher: evidence-validated, never
  assumed), and the seed path when none fits: .add/personas-teacher/design/ (ui-designer
  · ux-researcher · ux-architect) + engineering-frontend-developer via the add agent in
  persona mode, seeded per DOMAIN and reused across screens. The persona's Critical
  Rules shape every beat; its Success Metrics become the beat-4 checklist.
- Beat 4 now references the already-loaded personas instead of re-loading them.

Distil:
- Dropped the `## The hard rules` constraints block — every rule in it (intake-before-
  domain, domain-first, reuse-before-invent, confirm-before-build, engine-never-renders,
  bind-don't-break, confirm-against-personas) was a verbatim restatement of the beat it
  came from; each survives inline in its beat, the capture section, or the new preamble.

design.md 5,875 -> 5,416 B; the run+loop+design reclaim pool holds at 18,449 <= 20,138.
All three skill trees stay byte-identical.

verify: full tooling suite green; wording-lint 0; every design.md phrase-pin
(five beats + axes/options · the UI-Designer/UX-Researcher checklist · never-an-auto-pass
· never-lowers-a-gate · no-ui-personas degrade · the engine-never-renders invariant ·
the capture convention) preserved.

author: Tin Dang
@TinDang97
TinDang97 merged commit 5275f3f into main Jul 20, 2026
9 of 10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants