ADD 2.0 — skill-led method, thin kernel, living specs, measured 1.7× spec-kit#163
Merged
Conversation
…session before a milestone is born Dogfoods ADD's own intake→scope flow to spec a new-major method milestone: the skill should shape a milestone through a persona-led, risk-proportional strategy discussion that optimizes its task plan with the user BEFORE the milestone is committed. Closes the gap that today intake.md only classifies into a bucket and scope.md runs a generic co-specify — no persona lens, no task-DAG optimization (personas apply only at design/build/advisor/verify; strategy is per-task §5 only). 5 breadth-first tasks (freeze the ## Strategy schema first): - strategy-section — ## Strategy slot in MILESTONE.md.tmpl + engine handling - persona-at-intake — extend add-persona selection/routing to the intake surface - strategy-guide — new strategy.md: persona-framed discuss→optimize→converge to ~95% - advisor-strategy-trigger — advisor refutes a high-uncertainty milestone's strategy - risk-proportional-skip — micro/--fast bypass; zero added per-turn cost Guardrail: ## Strategy is SOFT/advisory (like §5) — no new hard gate, no re-added per-turn ceremony/cost (risk-proportional-skip is a first-class task), security still HARD-STOPs. Deferred (out of scope): the AI proactively INVENTING milestones unprompted — this shapes a request the human raises, it does not originate scope. Created via add.py new-milestone strategy-intake --await-confirm + milestone-confirm (by tindang); check 880 passed / 0 failed (no dangling criteria). extends persona-learning-loop, dynamic-personas, scope-loop; relates-to risk-proportional-ceremony, expectations-first. author: Tin Dang <tindang.ht97@gmail.com>
…t (WIP, tests phase) First task of the strategy-intake milestone, driven through specify → plan → freeze. §1 rules (M1-M4 template slot + renders + placement-outside-parsed-spans + 3-twin parity; R1 strategy_must_stay_soft, R2 tasks_parse_corrupted), §2 four gherkin scenarios, §3 frozen contract: MILESTONE.md.tmpl gains one drafted-blank `## Strategy` section between `## Exit criteria` and `## Close`, engine renders it verbatim — no add.py parse/gate, no ENGINE_MD5 repin, 3 twins byte-identical. Least-sure flag surfaced at freeze: whether placement after `## Exit criteria` leaves the Tasks-DAG parse + pre-confirm section scan byte-behaviour-unchanged (a red test asserts the DAG reads exactly N task rows and confirm/check stay clean) [contract/test]. Crossed into tests — red suite next. Frozen @ v1, approved by tindang. author: Tin Dang <tindang.ht97@gmail.com>
…M brain Re-scope the milestone from "a strategy session before a milestone is born" into the bigger vision the user set: remove ceremony by making a fitting persona the project-management brain — it shapes each milestone's strategy AND owns how every human gate is communicated and paced, retiring the fixed report-template.md section list. Persona-driven, adapts per project. Marked risk: high (method-defining). The four report floors: show-before-ask, one-approval-at-freeze and never-pre-stamp move from a fixed template into the persona CONTRACT (met in the persona's own voice, no fixed section list). The fourth — security is ALWAYS HARD-STOP — is kept HARD as a STRIKEABLE carve-out against the "floors become persona judgment" choice, because it is written into the method constitution (release readiness floor is un-forceable) and the operating rules forbid authoring a security auto-pass. The human may strike that one line to dissolve it. New freeze-first task persona-owns-gates carries the persona-owned-gate contract every other task cites. strategy-section kept, frozen @ v1. check: 884 passed, 0 failed. author: Tin Dang <tindang.ht97@gmail.com>
…tract (WIP, tests phase) The strategy-intake freeze-first task. Retires the fixed report-template section list in favour of persona-owned principles: a gate report must CONVEY the decision + its arc, the shape/plan, flags lowest-confidence-first, evidence, and a guided ask — the persona owns structure, order, emphasis, length and cadence, adapted per project. The four floors survive as persona-contract obligations, not layout: show-before-ask, one-approval-at-freeze, never-pre-stamp, and security = HARD-STOP (the one hard, un-persona-negotiable floor — the strikeable carve-out per the milestone). The engine is untouched: the `Reported:` trace + contract/verify_report_unrecorded audit codes stay verbatim; no ENGINE_MD5 repin. risk: high, autonomy: conservative (method-defining trust-layer change, human gate at verify). §1-§3 filled, frozen @ v1 approved by tindang, crossed into tests. Red suite next: test_persona_owned_gates.py + migrate test_report_arc (SUMMARY/ordering now persona-owned) and test_xml_convention (the report-blocks heading). author: Tin Dang <tindang.ht97@gmail.com>
…or persona-owned gates (gate PASS) The strategy-intake freeze-first task, complete. report-template.md is reframed from a MANDATED ordered section list into persona-owned PRINCIPLES: a gate report must CONVEY its content (decision + ARC · shape/plan · flags lowest-confidence-first · evidence · a guided APPROVE ask), but the fitting persona owns structure, order, emphasis, length and cadence — adapted per project. A sensible default layout remains as the persona's baseline; the mandate does not. The four floors survive as persona-contract obligations, not layout — show-before-ask · one-approval-at-the-freeze · never-pre-stamp — with security = HARD-STOP marked the one un-persona-negotiable floor (the strikeable carve-out per the milestone). SKILL.md's report line now names the principles, not the banner→…→NEXT sequence. Engine untouched: the `Reported:` trace + contract/verify_report_unrecorded audit codes stay verbatim; ENGINE_MD5 unchanged (4e65596…). The mandate→principle reframing retired now-optional prescriptive prose, freeing ~1050 B that offset the added principle + four-floors block — report-template.md landed at 9514 B (net −109 B), so the razor-thin reference pool stayed under budget with no rebaseline. 3 skill trees byte-identical, SKILL.md 9489 B (< 9500 ceiling). Tests: test_persona_owned_gates.py (10) green; migrated test_report_shape_scan_audit.py byte-ledger (9623→9514) + test_xml_convention report-blocks heading. 114 affected report/skill/parity tests green · check 888/0. risk: high, autonomy: conservative, verify gate approved by tindang. Seeds the next task: redefine UDD as experience-driven development, hosting the persona-owned gate as a text-mode UX artifact. author: Tin Dang <tindang.ht97@gmail.com>
…PM AND UX
Extend the strategy-intake milestone per the user's direction ("move report template
to UDD as part of user experience"; chose to redefine UDD, folded into this milestone).
The milestone now spans personas owning both the project-management surface (strategy,
gates) and the user-experience surface (gates as UDD artifacts).
Broadened goal + scope: UDD is redefined from UI-design into experience-driven
development — the pillar (design.md + UDD chapter + SKILL.md trigger + the four design
axes) covers UI AND interaction/gate UX, and the persona-owned gate report is hosted as
a UDD text-mode UX artifact designed through the UDD lens.
Two new breadth-first tasks: udd-experience-pillar (redefine the pillar) and
gate-experience-udd (host the persona-owned gate as a UDD artifact; depends-on
udd-experience-pillar + persona-owns-gates). persona-owns-gates marked done; two exit
criteria added. check: 889 passed, 0 failed.
author: Tin Dang <tindang.ht97@gmail.com>
…en development + a fifth INTERACTION axis (gate PASS) The strategy-intake milestone's UDD-redefine pillar. design.md is reframed from UI-design into EXPERIENCE-DRIVEN development: the loop now triggers on a UI feature OR any human-facing experience surface (a screen · an interactive flow · a human gate), stating "UDD is experience-driven development, not UI-only". The human gate is named an in-scope UDD surface — setting up gate-experience-udd — without yet folding report-template.md. The design-intake beat gains a FIFTH axis, INTERACTION (cadence · when/how to seek the human · turn-rhythm), alongside the four originals (FIDELITY · CONCEPT · LAYOUT · VISUAL DESIGN), which stay with frozen names. Every "four axes" reference (design-intake beat + hard-rules) becomes "five axes" naming INTERACTION. SKILL.md's UDD trigger broadens to "UI/experience surface → UDD loop"; the +11 B is offset by a same-guide "the default mode" trim, landing SKILL.md at 9490 B (< 9500 ceiling) — compress-not-rebaseline held. The three skill trees stay byte-identical (design.md 4b340bb… · SKILL.md 756950f…). The loop machinery is UNCHANGED (5 beats · capture/confirm · read-only binds). Engine untouched: no add.py edit, ENGINE_MD5 unchanged (4e65596…). The DESIGN.md.tmpl + appendix-c-glossary.md INTERACTION field is deferred (§7 SPEC·open); the report-template fold + lightweight gate loop are seeded for gate-experience-udd. Tests: new test_udd_experience_pillar.py (10) green — experience-driven framing · five axes incl. INTERACTION · SKILL trigger + ceiling + 3-tree parity; migrated test_design_intake_beat.py (15) stays green (four axis names still asserted). report/xml/parity/lean guards (47) green · check 893/0. risk: high, autonomy: conservative, verify gate approved by tindang. Refute-read EARNED (self); advisor 3-lens PASS (security/concurrency/architecture CLEAR). author: Tin Dang <tindang.ht97@gmail.com>
…te-experience-udd rename + clear dedup debt
thin-engine-loop W1 (phase-collapse-6-to-3): demote add.py from loop-driver to
helper by collapsing the two pure-bookkeeping `advance` calls. A new opt-in
`--thin` lane freezes the whole Direction bundle (spec+plan+§4 tests) in ONE
`freeze --cross` that crosses straight to build via _build_entry's SAME floor
machinery — a 3-call flow (new-task · freeze · gate) down from the 5 every
existing lane prescribes. oneshot/fast/default lanes are byte-unchanged; the
mechanical floor is intact — tamper tripwire + flag-verified + §5 scope
snapshots are all still captured through _build_entry.
- add.py: is_thin state marker; freeze guard min-phase = specify for thin;
`freeze --cross` thin branch (specify/plan → build in one call); new-task
--thin flag + 2-call thin recipe
- test_thin_engine_call_floor.py: RED→GREEN target — new-task --thin prescribes
≤3 engine calls
- engine_pin.py: ENGINE_MD5 re-aimed 4e655960→ee9631f9 (add.py changed),
synced across all 3 tooling trees
- .add/SEAMS.md: _declared_scope anchor re-pinned 5901→5934
fold gate-experience-udd (strategy-intake): physically rename report-template.md
→ gate-udd.md across all 3 skill trees (the text-mode UDD gate surface, a member
of the UDD doc family); repoint every reference (9 skill guides · 2 book docs ·
tests); design.md gains the lightweight text-mode gate variant. The four gate
floors (show-before-ask · one-approval · never-pre-stamp · security-HARD-STOP)
are preserved verbatim in substance. Task stays at phase=build (autonomy:
conservative) — the PR review IS its human verify gate; not engine-gated here.
clear the dedup debt: the orchestration pool was 349 B over its 41300 floor at
branch HEAD (pre-existing) + 189 B from gate-udd's design.md = 538 B over.
Compressed −538 B of prose across streams/advisor/loop/design (in-file, no
rebaseline; every pinned phrase, reject code, CLI verb, and the byte-identical
<strategy> block preserved). Pool now 41278 (22 under) — greens the byte-budget
family + the fresh-checkout rollup.
Suite: only pre-existing environmental reds remain (pty_clack ×4, installer
npm-cancel ×1 — no npm/@Clack in the bare unittest env).
author: Tin Dang <tindang.ht97@gmail.com>
… docs The tooling suite pins doc PROSE in ~278 files: each past method task was built red/green against the exact wording it added. That ossification reddens a guard on every doc improvement (it bit the dedup compression 3× today) and blocks the ceremony-cut. Per direction: remove the pure prose-guard files, keeping every test that guards a CORE VALUE (behavior · machine tokens · structure/parity). Deleted 15 files whose SOLE assertions are "this doc/guide/render-card contains phrase X" — no engine execution, no token/reject-code/CLI-verb guard, no md5 parity, imported by no other test: test_worktree_default_wording · test_two_surface · test_v11_docs · test_docs_align · test_scope_level_enum · test_supported_agents_docs · test_skill_todo_flag · test_arc_gate_wiring · test_report_arc · test_docs_accord · test_gate_read_diet · test_risk_report_render · test_template_dedup · test_question_summary_layer · test_rewrite_core Verified before removal: each read to confirm pure-prose. Three false-positive classes were caught and KEPT — md5 parity guards (test_*_parity), function unit-tests (test_md_section slices markdown), and behavior-via-import (test_installer_soul_seed runs the pip installer). After removal: 0 import failures across the remaining 378 modules; inventory/hygiene guards (semantic_inventory · temp_hygiene · plugin_manifest) green. The behavioral/token coverage these prose-echoes shadowed lives on their own modules (freeze/gate/audit behavior, verb + reject-code census, 3-tree parity) — untouched. Docs on these surfaces are now freely editable. author: Tin Dang <tindang.ht97@gmail.com>
…cise add.py
Per direction: the tooling suite must test the add.py engine, nothing else. The
doc/prose/wording/parity/rubric guards were ceremony — they ossified the docs
(every doc edit reddened a phrase-pin) without testing behavior. Removed all 67
tests that only inspect text: no add.py execution, no engine-package import, no
installer/real-code, no subprocess. 378 → 311 test files (−8165 lines).
Deleted class: doc-echo (test_v6_run · test_v8_docs · test_foundations_chapter ·
test_references_appendix …), wording/slang lints (test_wording_lint ·
test_ubiquitous_language), render/rubric/guide guards (test_confidence_rubric ·
test_intake_rubric · test_design_loop_guide …), byte-budget pools (test_skill_lean ·
test_skill_dedup · test_streams_persona_flow), and doc-structure guards
(test_xml_convention tag-census · test_book_parity · test_version_sync ·
test_semantic_inventory · test_worker_contract_sync). Kept test_md_section — it
unit-tests the md_section slicer function (real code).
Decoupled the survivors that leaned on deleted hubs:
- 10 engine tests imported test_skill_lean.POOLS for a phases byte-budget
side-check — stripped those non-engine `test_phases_pool_*` methods.
- test_roster_portable (tests init -> AGENTS.md roster sync, real add.py) imported
roster data from test_agent_roster; inlined AGENTS/AGENT_PHASES/AGENT_TREES/
_agent_path/_frontmatter so it stands alone.
- .add/SEAMS.md three-tree-parity seam cited the retired test_book_parity —
dropped it; the engine/skill/bundle parity guards (test_engine_repin_parity ·
test_tree_parity · test_bundle_parity) survive.
Suite: 3121 tests, only the pre-existing environmental reds remain (pty_clack ×4,
installer npm-cancel ×1 — no npm/@Clack in the bare unittest env). 0 import failures.
Docs on every surface are now freely editable — the ceremony-cut is unblocked.
NOTE: book-tree doc parity is no longer test-enforced (its guard was pure text
inspection). Engine + skill + bundle trees stay guarded; book/doc sync now relies
on discipline.
author: Tin Dang <tindang.ht97@gmail.com>
First increment of the md ceremony-cut (docs now freely editable after the
engine-only test teardown removed their prose pins). Kept every core value —
machine tokens, reject codes, functional templates, the pin-locked worker XML,
3-tree byte-identity; cut only teaching redundancy.
streams.md — gutted StreamsSafetyClausesTest from test_streams.py: 8 pure-prose
asserts that pinned exact safety wording ("name the task every time", "lease",
"circuit-breaker", "never write shared state" …). The safety CONCEPTS stay in
streams.md; their wording is now free to evolve. The real engine guards survive
untouched — slug-routing precedence (SlugRoutingPrecedenceTest), the ×3 md5
parity guard, the reject-code pins (unverified_fork_base), and the pin-locked
<strategy> block (byte-identical to advisor.md). test_streams: 13 green.
run.md (×3 trees, −383B) — compressed the "specification bundle (v7)" section,
which itself points to phases/3-plan.md as its "one home", down to that pointer +
the one load-bearing fact (one approval at the frozen contract); dropped the
historical "v7 reversal (recorded)" aside that restated the already-stated
auto default. All 12 machine tokens preserved (add.py audit/check/gate/heal,
refute_unrecorded, unguarded_high_risk_auto, goal_not_auto_ready, …); 3 copies
byte-identical.
Verified: 105 targeted tests green (test_tree_parity, test_bundle_parity, and
the run.md-referencing engine tests: setup_run_mode, verify_deepen,
earned_green_rubric, refute_record_required, high_risk_signal, goal_auto_ready_gate,
plugin_manifest) + test_streams 13. Full suite left to CI (discover run exceeds
the 2-min bound — the standing "don't block on the full suite" rule).
author: Tin Dang <tindang.ht97@gmail.com>
…/gate guides Second increment of the md ceremony-cut. Removed pure illustration that repeats the rule it sits under — the decision tables, reject codes, and format blocks above each example ARE the core value; the worked example is a teaching aid that costs tokens every load without adding a rule. - intake.md — dropped "## Worked examples (from this project's own history)" (the 4-row request→bucket table); the four-buckets table + tie-break test + fast-fit rule already carry the classification logic. - scope.md — dropped "## Worked example" (the v4-1 walkthrough); the per-outcome table + drafting sections + well-formedness gate are the method. - gate-udd.md — dropped "### Per-gate examples" (two ARC illustrations); the ARC format block + its field bullets already specify it. All reject codes / machine tokens preserved (not_classified · dangling_criterion · duplicate_goal · ask_human · frozen_scope · split_required · missing_capture …); 3 trees byte-identical. Also gut test_review_checklist.test_run_md_accord — a brittle wording-pin that required run.md to contain the literal phrase "freeze review checklist". The prior commit (d264c38) compressed that sentence to point at phases/3-plan.md as the checklist's "one home" (navigation preserved), which tripped this cross-ref pin — a latent red that batch didn't cover. The structural guards in that file stay: the md_section 7-item count, the `risk: high · autonomy: conservative` token pin, the 3-tree parity, and the ENGINE_MD5 untouched-guard. Verified: 125 targeted tests green (review_checklist, tree/bundle parity, intake_interview, gate_owner_marker, report_rendered_trace, udd_persona_checklist, project_scope_lock, design_intake_beat, plugin_manifest, streams) + add.py check 898 passed / 0 failed. author: Tin Dang <tindang.ht97@gmail.com>
Per direction (cut ceremony deeper: strip explanatory rationale to rule +
machine tokens, keep the enforceable rule blocks). design.md 7049 -> 5811 B.
Compressed the five-beat body prose and the tool-agnostic-capture section to
the rule + tokens; dropped teaching rationale ("so the approved screen is
traceable", "the cheapest artifact that shows...", "it keeps screens
consistent"). Kept intact: the <constraints> Hard rules block, every machine
token (missing_capture, tokens.json, catalog.json, prototypes/<name>.json,
DESIGN.md, flow: design), the five axis names, and the concept-pins the
design tests assert (never-lowers-a-gate, never-an-auto-pass, engine-never-
renders, the UI-Designer/UX-Researcher dimensions).
Fixed a line-wrap that split "never lowers a gate" across a newline (broke the
phrase-pin — same class as the "default for any spawn" wrap bug).
Verified: 30 tests green (test_design_intake_beat, test_udd_persona_checklist,
tree/bundle parity); 3 trees byte-identical.
author: Tin Dang <tindang.ht97@gmail.com>
Continuing the aggressive rationale-cut. These two guides are already dense (procedural steps, tables, output_format/exit_gate blocks, machine tokens), so the trims are modest — the scenario-writing rationale in 1-specify and the foundation-interview preamble in 0-setup, compressed to rule + tokens. Preserved every command, table, machine token (amount_invalid, already_locked, setup_unlocked, --await-lock, add.py init/lock/advance) and the exit_gate / output_format blocks. 3 trees byte-identical. Verified: 75 targeted tests green (tree/bundle parity, phase_bundles, progressive_context, setup_run_mode, stale_guide_sync, plugin_manifest). author: Tin Dang <tindang.ht97@gmail.com>
Fill the blank thin-engine-loop MILESTONE.md against its human-set goal (≤3 add.py calls per task · 6→3 phases · loop-in-SKILL): scope, shared ground (engine/template/skill anchors incl. the 9497/9500 SKILL ceiling, phase-name pin census, twin-tree parity), shared decisions (route glossary, persona-proposed/human-ratified ceremony, the four floors bind every route, ≤3 calls is happy-path not a cap), freeze-first contracts, six breadth-first tasks (phase-collapse-3 · template-unify · skill-loop-fold · persona-routes-depth · engine-kernel-trim · call-census-proof) and verifier-cited exit criteria. Queue persona-gepa-loop as the follow-on: route-outcome traces reflected into persona routing-rule deltas at fold-time (GEPA-style, human-folded, never-clobber rails) — depends-on thin-engine-loop, extends strategy-intake. Scope drafts only; task decomposition awaits the human confirm (engine check 898 passed / 0 failed). author: Tin Dang
…rafted + red suite (8 red / 3 floor pins) Six tasks created with dependency DAG; phase-collapse-3 §1–§4 drafted to the freeze: universal direction·build·verify enum, freeze --cross crosses the whole front via _build_entry, gate compound kept, legacy read map, 3-call recipe every lane. Wordy-test removal authorized in milestone shared decisions; template family lean-pass added to template-unify. author: Tin Dang
…·verify(·done)
The thin lane's 3-call walk becomes THE walk. PHASES collapses to
("direction", "build", "verify", "done"): the old front (specify · scenarios ·
plan · tests) is ONE direction span ending at the single freeze approval, and
observe folds into verify. Every task now runs new-task → freeze --by <name>
--cross → gate PASS — three engine calls, one human seam.
Engine (add.py + add_engine/constants.py, ~360 lines):
- LEGACY_PHASES: old phase names normalize at the ONE read accessor — 473
legacy task records render/load unchanged, zero task-file rewrites; the CLI
accepts legacy names with a mapped-to note; check's marker-parity check
normalizes both sides (a hand-written `phase: specify` marker never reds).
- The two retired crossings' floors relocate into _build_entry's shared stack:
cross-component hold + producer snapshot / consumer pin now fire at the one
direction→build crossing alongside freeze gate · build-expectations · flag
check · tamper tripwire · §5 scope snapshot. advance, freeze --cross, and
the phase-build override run the IDENTICAL stack; --thin is a no-op.
- guide never re-teaches a passed gate: post-freeze direction steers to
`add.py advance` on both the plain and --json surfaces.
- advance --to maps legacy tokens, stops at direction, and no-ops friendly on
an already-reached target; init/new-task teach the 3-call recipe.
Suite migration (49 test files, red 372 → green):
- Walk helpers collapsed (a task is BORN at the freeze seam); stall phases
move to `direction`, landings to `build`; FROZEN fixtures carry the
least-sure flag the unified crossing now checks.
- 34 old-walk value pins deleted under the recorded wordy-test authorization
(thin-engine-loop MILESTONE.md, 2026-07-16) — none exercise a floor; every
freeze/gate/tamper/audit/scope/cross-component floor test MIGRATED, and
test_phase_collapse.py pins the new 3-call walk end to end.
- Two vestigial pins retired whose targets d264c38 (human-approved md-cut)
deleted: wave_ledger's dangling StreamsSafetyClausesTest ref, run.md's
"seven lines" prose pin.
Book + docs ripple: ch02 flow mermaid now draws the 3 phases as subgraphs
wrapping the 7 beats (synced ×4 trees incl. bundle), diagrams/CHECKLIST.md
names the phase bands, appendix-c dogfood copy re-synced. SEAMS.md
_declared_scope anchor re-pinned 5934→5915.
Pins re-aimed once: ENGINE_MD5 6f688c41…, ENGINE_PKG_MD5 89a75e5d…; add.py /
add_engine / engine_pin byte-identical across canonical, bundle, and both
dogfood twins.
task: phase-collapse-3 (thin-engine-loop W2) — §3 FROZEN @ v1, build phase
author: Tin Dang
Claude-Session: https://claude.ai/code/session_018xZT2V9tTvjuGmPFjZw4MG
…cross · gate PASS
Human-led verify gate (risk: high · autonomy: conservative), decided by Tin Dang
2026-07-16 ("Ratify v2 + gate PASS"): §3 re-frozen @ v2 with the widened scope
(the suite-wide recipe-pin migration + the ch02 docs ripple — contract shape
unchanged), the post-freeze test changes sanctioned via the recorded re-cross,
§6 evidence filled (suite 3114 OK · live 3-call transcript · legacy render
zero-rewrite proof · 4-way engine md5 parity · EARNED refute-read · 3-lens
PASS), outcome PASS recorded — task done. ADR block harvested at done.
task: phase-collapse-3 (thin-engine-loop W2)
author: Tin Dang
Claude-Session: https://claude.ai/code/session_018xZT2V9tTvjuGmPFjZw4MG
TASK.fast.md.tmpl is gone. The fast lane is now a DERIVED render of the one
template: _strip_fast_sections(full) drops exactly the _FAST_SECTIONS heading
blocks (## 7 · OBSERVE + §6's Deep-checks / Live-verify / Refute-read /
Advisor-3-lens) and cmd_new_task splices a bare `fast: true` header — a strict
line-subset by construction, machine-checked. --oneshot adds its two headers
plus the §3 "### AI-verify record" block as splices of the same render.
Engine (add.py + add_engine/constants.py):
- _FAST_SECTIONS (constants) replaces _FALLBACK_TASK_FAST; the shared
_FALLBACK_TASK gains the Ground SHA field the fast fallback carried.
- _render_template's fallback table shrinks to TASK.md; _strip_fast_sections
+ _AI_VERIFY_RECORD_BLOCK own the derivation mechanics.
- Pins re-aimed: ENGINE_MD5 8eaca350…, ENGINE_PKG_MD5 ed7bf3e1…; add.py /
constants / templates byte-identical across canonical, bundle, and both
dogfood twins; SEAMS _declared_scope re-pinned 5915→5972.
Template family lean-pass (machine-read lines pinned by the new suite):
- TASK.md.tmpl 12209→12015 B while GAINING the §1 Boundary: line — the wm2
input-dialect floor (boundary_unfilled) now fires on BOTH lanes at freeze —
and natively carrying `phase: direction` (the render rewrite is a no-op).
- MILESTONE.md.tmpl 4211→3729 · PROMPT.persona.md.tmpl 3225→3046 ·
personas/_template.md.tmpl 6922→5411.
- Skill trees (×3): fast-lane.md + SKILL.md now describe the derived render
(§3 scope v2, re-crossed — the human ratified widening before the gate).
Suite: test_template_unify.py (7 red flipped + 6 floor pins) is the task's
red suite; test_fast_lane_template.py deleted (pinned the two-template
design end-to-end, superseded); ~12 files migrated — fast section-set
{1,3,4,5,6}→{1,2,3,4,5,6}, fast-template file pins → derived-render pins,
6 freeze fixtures fill the now-universal Boundary line. Full suite green:
3092 tests OK ×3 runs; add.py check 930/0.
task: template-unify (thin-engine-loop W3) — §3 FROZEN @ v2, gate PASS
author: Tin Dang
…demand references
An ordinary task now reads ZERO phase-guide files: SKILL.md (9,324B ≤ 9,500)
narrates the 3-beat loop INLINE — DIRECTION (draft §1–§4 → the ONE freeze,
`freeze --by <name> --cross`) · BUILD · VERIFY (`gate PASS`) — and the guides
became on-demand references, never a mandated per-phase load.
Skill trees (×3, byte-identical):
- phases/{0-setup,1-specify,3-plan,4-tests}.md fold into phases/direction.md
(13,868B): setup span (init/4-lens/run-mode/lock) · rules+scenarios
co-specification · plan (grounding/contract/build-strategy + the seven-item
freeze review checklist) · tests span (the declare-where-tests-live grammar);
every pinned anchor moved VERBATIM, one <exit_gate> per beat.
- phases/5-build.md → build.md · phases/6-verify.md → verify.md (retitled,
most-pinned content intact); phases/fast-lane.md deleted — routing lives in
SKILL.md flag mode. Pool 33,496 → 23,586B (−30%).
Engine (§3 v2, ratified by Tin Dang mid-build + re-crossed): _PHASE_GUIDE_FILES
re-aims to the merged 3-file shape — the re-aim phase-collapse-3's own comment
reserved for this task; `add.py guide` resolves playbooks again instead of
returning null on the new tree. ENGINE_MD5 349901707a…; 4-way twin parity.
Ripple (97 reds → 0): 35 test files re-aim paths; retired structure pins
(## Exit gate → <exit_gate> census, ## Run mode → **Run mode**, the SKILL.md
phase table → inline-beat asserts, §0 → §3 Grounding) re-pinned; dropped
teachings restored in the fold (teacher library pointer, invariants-in-§3,
data fixtures, subagent ground sweep, "you run" lock wording, a wrapped
run/entry-contract phrase). Roster agents ×3 trees, beyond/design/run/scope/
streams ×3, both glossaries (TASK.fast.md wording retired), and
SEMANTIC_INVENTORY re-aimed; wording-rubric keep-terms (Objective: · living
documentation) re-homed; SEAMS _declared_scope 5972→5971.
Suite: test_skill_loop_fold.py 8/8 (red-first); full suite 3100 OK;
add.py check 929/0; wording_lint 0 findings.
task: skill-loop-fold (thin-engine-loop W4) — §3 FROZEN @ v2, gate PASS
author: Tin Dang
…reeze ratifies it
The fitting persona now proposes a task's ceremony lane in the TASK header —
`route: <full|fast|oneshot> · routed-by: <persona:<slug> | human> — <why>` —
and the freeze IS the ratify: the direction→build cross records
state.tasks[slug].route = {lane, by}. Measure-not-block throughout: a missing
or unknown lane records "unrouted" and NEVER refuses the freeze.
Engine (add.py, ENGINE_MD5 ec9a5730…, 4-way twin parity):
- _ROUTE_LINE_RE/_ROUTE_LANES/_route_record beside the header-regex cluster;
the record write sits at the flag_verified seam in _build_entry
(UNCONDITIONAL overwrite — a re-cross re-records).
- audit gains route_unrecorded (unrouted record, or a ratified record whose
header line was deleted post-freeze — state is the witness, mirroring
unflagged_freeze) + route_lane_mismatch (ratified lane contradicts the
task's actual fast/oneshot flags). Grandfathered by key ABSENCE — no
pre-feature record is ever retro-redded.
Doctrine (SKILL.md 9,480B ≤ 9,500, ×3 trees): flag mode shifts from
"two human-owned settings (never auto-picked)" to propose-then-ratify —
a route bullet teaches the header line; the duplicate todo jotting line
funded the bytes. direction.md's freeze paragraph names the route ratify.
No render change: a per-lane spliced line would break template-unify's
strict fast-subset pin, so the persona writes the line, never new-task.
Suite: test_persona_routes_depth.py — 8 red flipped + 3 floor pins
(unrouted-freeze-crosses · grandfather · no-new-freeze-refusal); full suite
3111 OK; wording_lint 0; SEAMS _declared_scope re-pinned 5971→5990.
This task's own header carries the dogfood route line, recorded at its
own freeze.
task: persona-routes-depth (thin-engine-loop W5) — §3 FROZEN @ v1, gate PASS
author: Tin Dang
…plicated parity/engine-pin tests test_tree_parity.py becomes THE canonical home for tree parity + the engine pin: skill ×3 (incl _bundled) · agents add-*.md ×3 · tooling 4-way (add.py · engine_pin.py · add_engine/*.py · templates/**, exists-skip ≥2) · docs 3-tree chapter mirrors · the ONE ENGINE_MD5/ENGINE_PKG_MD5 assert. Under it, the duplicated per-task copies are deleted: 146 files changed, +251/−1713 lines, suite 3111→2941 tests, full run OK ×2. Census (pinned by the new test_corpus_slim.py red suite, all green): - ENGINE_MD5-referencing test files 112→3 (sweep · security floor · census) - parity-named fns 180→55 (≤55) · their files 126→43 (≤45) - dead PHASES_POOL constants 8→0 - floor def-counts hold: freeze 9 · gate_audit 18 · scope_lock 31 · security 5 · advisor_relax 29 · ai_plan_verify 44 · unflagged 13 - engine untouched: ENGINE_MD5/ENGINE_PKG_MD5 byte-identical (no repin) The AST classifier's "0 behavioral suspects" was refuted by a killed-fn audit: 16 guards bundled inside parity-named fns were restored (NOTICES attribution, engine hands-off scans, v16 tag census, cross-tool roster sync, plugin-version lockstep, template field pins, cospecify anchors, no-render invariant, wave-verify census, io_state manifest membership). Collateral fixed: release-1.11.0 CHANGELOG anchor un-reworded; untrack-add-tooling meta-tests re-aimed at the sweep; SEAMS.md engine-md5-repin + three-tree-parity anchors re-aimed (scope amended, re-crossed by Tin Dang); ci_tooling_mirror_gap's vacuous-at-HEAD "untouched by this build" git-diff assert dropped (CI-shape guards kept). test_shared_engine_pin.py deleted — its scans fell as duplicates and the survivor was a meta-runner; the pin-literal scan retirement is ledgered as a §7 open delta. Probe: one corrupted byte in _bundled SKILL.md reds exactly the sweep while the former duplicate suites stay green — detection consolidated, still firing. Task: .add/tasks/test-corpus-slim (frozen @ v1 by Tin Dang, gate PASS, route: full · routed-by: persona:tdd-verifier · sensitivity: mechanical). author: Tin Dang
…come measurable (ADD 2.0 M1) ADD 2.0 makes personas the method's core value; this lands the measurement substrate. A closed task-kind taxonomy (constants.TASK_KINDS, 10 kinds) is the join key between a persona's routing claim (`task-kinds:` frontmatter, new template slot) and a task's declared kind (`kind:` header line, same anchored grammar family as route:/sensitivity:, read by the new PURE _task_kind). Every recorded gate outcome now appends ONE JSON line to .add/traces/route-outcomes.jsonl (_append_route_trace): ts · task · milestone · kind · lane (route record, oneshot-marker fallback) · routed_by · persona (parsed from a persona:<slug> routed-by) · outcome · heals · recross · age_hours · actor. Engine-derivable fields only; degrade-safe — the trace is telemetry, state stays the source of truth, a failed write never blocks the verdict. This is the evidence stream the persona scoreboard and the GEPA fold (M7) will read. _persona_quality_warnings gains Finding C: a task-kinds value outside the taxonomy is a named WARN (measure-not-block, mirrors Finding A's flow check). Red/green: test_persona_task_kinds.py (12) + test_route_trace.py (8) written red-first; full suite green after twin sync + repin (ENGINE_MD5 26d1db26, ENGINE_PKG_MD5 991ce131) + SEAMS re-aim (_declared_scope 5990→6045, _section_unfilled 100→101). Retired test_corpus_slim's task-scoped ENGINE_MD5_AT_FREEZE guard — its purpose expired when that task gated PASS; kept alive it would re-break on every legitimate engine change (this one was the first). The durable pin lives in engine_pin.py + test_tree_parity. refs: ADD 2.0 M1 persona-core (T1+T2 of 5) — next: roster distill 5→1, capability seed templates, dogfood persona migration author: Tin Dang
… loop
Evidence-driven cleanup of the test-run workflow (pytest --durations sweep,
2026-07-17): the corpus' wall-clock floor was ONE meta-test —
test_ci_tooling_mirror_gap's fresh-checkout test-job sequence (~145s, a suite
inside the suite) — plus a family of npm/node/pty/sleep-heavy IO suites that
serialize workers.
tooling/t (dev-only, NOT shipped — package.json files allowlist untouched):
./t fast lane — parallel via ephemeral uv env (pytest + xdist,
worksteal), slow-IO + CI-mirror suites excluded: 2684 tests
in ~19s (was 287s serial full)
./t --full whole corpus parallel (~168s) — the pre-commit/pre-PR bar
./t --serial unittest discover fallback — CI parity, no uv needed
The fast lane is an ITERATION loop, never the merge bar: the excluded
installer/global/packaging/pty/lock suites still bind in --full and CI.
author: Tin Dang
…ary distilled to its runtime core (ADD 2.0 M1) Personas are ADD 2.0's core value; this ships the seed library the per-project seeding draws from. Twelve v2-format capability templates land under tooling/templates/personas/ (53.6KB total vs the 256-file/12MB teacher lib they distill): product-lead · software-architect · build-engineer · evidence-verifier · security-gatekeeper · ux-experience-lead · data-steward · release-manager · stream-orchestrator · quality-auditor · platform-engineer · technical-writer. Three are MIGRATION templates — the workflow knowledge of engine verbs slated for the 2.0 kernel trim survives as markdown playbooks with explicit HUMAN seams: release-manager (release cut arc + graduation arc, 4 human seams), stream-orchestrator (8-step wave loop, review-queue seam), quality-auditor (fold ritual + spec compaction, 2 confirmation seams). A security HARD-STOP stays unstrippable in every template. Every template: v2 schema (task-kinds from the closed taxonomy, flow routing, not-when sibling boundaries), verified against the live engine predicates (_persona_missing + _persona_quality_warnings incl. Finding C) — zero bare placeholders, zero Abilities sections, zero dying-verb pins. Synced across the 4 tooling trees; no ENGINE repin (package digest covers add_engine/*.py only). Full fast lane + packaging/parity/persona suites green. refs: ADD 2.0 M1 persona-core (T4 of 5) author: Tin Dang
… agent (ADD 2.0 M1) Personas carry the expertise; the agent carries the discipline. The 5-agent roster (add-design/build/verify/persona/advisor, 25KB) collapses into one agents/add.md (~5KB): the spawn prompt names the MODE (direction · build · verify · advise · persona), the agent loads that beat's phase guide plus the best-fit persona by frontmatter (flow + the new task-kinds), and one shared boundary floor binds every mode (freeze/gate/lock stay human seams; security always HARD-STOP; verify mode still routes flow:verify first, advisor fallback). Ripples closed in the same change: - constants.py PHASE_AGENT -> all phases prefer "add"; PERSONA_HINT/ PERSONA_FIT_HINT reworded to "the add agent in persona mode". - guidelines.py block roster -> 1-agent + modes (sync-guidelines re-ran: CLAUDE.md/AGENTS.md/.clinerules regenerated). - Installer tombstones: _installer._RETIRED_AGENTS + cli.js RETIRED_AGENTS name the 5 retired files so `update` removes them from the SHARED .claude/agents namespace (tombstone-only — user agents never swept). - PROMPT.persona.md.tmpl: the 10-row per-runner adapter table is RETIRED, replaced by a runner-agnostic spawn contract of four general slots — prompt template · persona · model · isolation (user-directed); Claude Code stays the one verified reference. - Skill guides (advisor/streams/beyond/phases-build) + TASK.md.tmpl route to the one agent's modes. - Test migrations: roster_portable rewritten to the 1-agent bidirectional contract (block modes == agent-file mode bullets, retired names rejected); roster_shipped/tree_parity/phase_bundles/verify_flow_value/strategy_facets/ fold_persona_sections/installer_shared_namespace/packaging pins migrated; subagent_prompt pins the four contract slots. - Repins: ENGINE_MD5 9433f6e3 · ENGINE_PKG_MD5 d82eeae0; SEAMS _declared_scope 6045->6048. FULL corpus green: 2956 passed, 3 env skips. refs: ADD 2.0 M1 persona-core (T3 of 5) author: Tin Dang
…0 M1) The 6 project personas each declare their slots in the closed TASK_KINDS taxonomy (the persona scoreboard join key): book-technical-writer=docs · method-product-owner=feature,integration,docs · methodology-engine-dev= feature,refactor,infra · security-gatekeeper=security · tdd-verifier= test,feature · terminal-ux-accessibility=ui. All 6 schema-conformant + quality-clean (Finding C validates the vocabulary). Abilities engine-verb scrub deferred to the M5 kernel trim, when the verbs actually retire. refs: ADD 2.0 M1 persona-core (T5 of 5 — milestone COMPLETE) author: Tin Dang
…get, the gate judges it, shards are free (ADD 2.0 M2) The plan becomes ADD 2.0's core artifact. Three moves, red-first (test_plan_target.py, 7 tests; route-trace schema contract extended): 1. Measurable Target in the plan. TASK.md.tmpl §3 Contract gains a `Target (measurable):` line — the success bar the §6 verify evidence must hit, numbers not adjectives. Measure-not-block: a §3 without it still freezes. direction.md drafts it; the exit gate lists it. 2. target-hit at the gate. `gate <outcome> --target-hit yes|partial|no` records the Target judgment in state (tasks[slug].target_hit) and in the route-outcome trace (target_hit key — completing the persona scoreboard schema begun in M1). An invalid value refuses BEFORE any write (target_hit_invalid); absence stays null, never inferred. 3. Shard-tolerant task folder, pinned as a 2.0 contract. AI-architected shard files inside .add/tasks/<slug>/ (notes, evidence, sub-plans) never trip the §5 scope guard — the .add tree is outside the scope walk by construction; ShardToleranceTest makes that a contract, not an accident. SKILL.md: "One file = one task" -> "One plan = one task" (TASK.md is the engine-known spine; the AI owns the shard architecture beside it). Budgets held by compression, never bumped: SKILL.md 9494B (<=9500; funded by trimming the duplicated graduation cue + a book-pointer clause), TASK.md.tmpl 12196B (<=12209 family ledger). Repin ENGINE_MD5 9cc73f6e; SEAMS _declared_scope 6048->6056. The physical TASK.md->PLAN.md rename is deferred to M6 migrate (renaming twice would churn hundreds of tests for zero behavior). FULL corpus green: 2963 passed, 3 env skips. refs: ADD 2.0 M2 plan-core-shards (complete) — next: M3 specs-5dd author: Tin Dang
…ernel verb (ADD 2.0 M3)
The foundation becomes five LIVING spec files and lessons land in-flight,
the moment they are learned — not batched at milestone close. Red-first
(test_specs_5dd.py, 9 tests):
1. Living specs. `init` seeds `.add/specs/{domain,system,experience,quality,
method}.md` (DDD · SDD · UDD · TDD · ADD) from ONE template —
templates/specs/SPEC.md.tmpl rendered five ways (template-unify
discipline) — never-clobber, never blank (the SETUP_FILES survivor
idiom, one shared _seed_spec_file truth). Each spec: a CURRENT "Now" +
"Decisions that bind" picture above a "## Deltas (newest first)" inbox.
2. delta-append — the last unbuilt verb of the ratified 2.0 eight-verb
kernel. `add.py delta-append <dd> "<lesson>"` routes via the closed
constants.SPEC_DDS map, prepends one `[open · <date>]` line directly
under the Deltas heading (newest first), stamps the active task (or
--task; none -> no stamp, never inferred). Unknown dd refuses BEFORE
any write (delta_dd_unknown). Legacy tolerance: a pre-2.0 project with
no .add/specs/ gets the target file seeded on demand — the verb never
dies on a missing dir (dogfooded on this very repo).
3. Wiring. SKILL.md observe beat points lessons at the verb (9497B, under
the 9500 ceiling — funded by compressing the same paragraph, never
bumped); deltas.md documents the in-flight channel beside the frozen
grammar; min-pillar read-spy census covers the new verb.
Twin trees synced (tooling x4, skill x3); repin ENGINE_MD5 11fe18db +
ENGINE_PKG_MD5 cd2d7e81; SEAMS _declared_scope 6056->6085.
refs: ADD 2.0 M3 specs-5dd (complete) — next: M4 skill-unify
author: Tin Dang
…f is the receipt (ADD 2.0 M4a) Intake gains the inline lane — the route for a change too small to deserve versioned scope. Red-first (test_inline_lane.py, 8 tests): - intake.md: `## The inline lane` sits between the interview and the frozen `## The four buckets` — the lane is judged BEFORE bucketing, because buckets create scope and the lane exists precisely so none is. Fit rubric (one file / covered behavior / no new contract surface / mechanical sensitivity) -> no task, no milestone: make the edit; the receipt is the git diff + `add.py delta-append <dd>` into the living 5-DD spec (specs-5dd M3 verb — the spec diff IS the approval artifact). - The floor is closed: security · data · architecture ALWAYS escalates to a real task (security stays HARD-STOP); the human's "make it a task" overrides the route, always. When in doubt, bucket. - SKILL.md routes to the lane at the intake beat — 9493B, under the 9500 ceiling, funded by four pin-checked compressions (each candidate phrase grepped against the test corpus before cutting; the one hit bound engine output, not SKILL.md). Line-wrap kept "inline lane" unsplit (phrase-pin hazard). Frozen intake pins survive (test_intake_interview stays green). Skill twins synced x3; no engine change — no repin. refs: ADD 2.0 M4 skill-unify commit A — next: M4b guide-fold (8 zero-engine guides fold into the 3 beat references) author: Tin Dang
Save the measured-campaign verdict as the canonical user-facing explainer for when (and why) ADD beats spec-kit: enforced vs advisory discipline, the context-rot finding, honest ties and spec-kit wins, the enforced- guarantees table, and testable predictions (weak models, hostile changes, autonomous runs, teams). Linked from both READMEs and the book nav (Appendix H, after Appendix G). author: Tin Dang
TinDang97
force-pushed
the
feat/adaptive-flow
branch
from
July 18, 2026 18:48
25e8635 to
317b65f
Compare
Leaderboard-class evidence pipeline: benchmark/swe/runner.py runs the SAME pinned agent (claude -p, sonnet-5, stream-json meter) on SWE-bench Lite instances in two arms — vanilla issue prompt vs ADD 2.0 installed into the checkout with the loop-driving prompt. Per instance: clone repo@base_commit, agent run, tracked-file `git diff <base>` filtered of method artifacts (.add/, .claude/, CLAUDE.md, ...) so the prediction is the FIX only, appended to predictions_<arm>.jsonl (official shape, resumable). Evaluation stays official: swebench docker harness (recipe in the module docstring). Smoke slice = three psf/requests instances (small repo, ids validated against the HF datasets-server). 15 offline guards pin the patch filter, meter argv, arm prompts, and smoke config; bench suite 257 green. runs-swe/ gitignored like the other runs roots. author: Tin Dang
TinDang97
force-pushed
the
feat/adaptive-flow
branch
from
July 18, 2026 18:51
317b65f to
245c005
Compare
… honestly Full matrix, all officially evaluated (swebench 4.1.0 docker harness): sonnet-5 vanilla 3/3 $0.95 vs add 3/3 $5.28 (every ADD patch ships a regression test; friendly ground ties, as the wm campaign predicted). haiku-4.5 mini probe: vanilla 3/3 $0.90 vs add 2/3 $2.52 — appendix-h prediction 1 NOT supported at n=3, published anyway: the ADD miss had a correct fix (F2P green) but over-built and broke two PASS_TO_PASS tests the gate never saw, because the loop bound only its own task tests as the floor. Diagnosed method-integration gap: in a foreign repo the host suite IS the regression floor and must be declared. Appendix H carries the against-us data point with the diagnosis, same as the first retraction. author: Tin Dang
…etired The SWE-smoke diagnosis (cost = authoring a 161-line PLAN.md for an 18-line fix; --oneshot never stripped the heavy body) lands as the ATG-informed cut: one file = one atomic node — persist the interface (contract · red suite · scope · verdict), reason everything else in-context. Template (191→131 lines, authored surface ~60→~22 fields), every engine anchor preserved verbatim: - §3 gains `Regression floor:` — the HOST repo's own suite is ALWAYS an inherited edge (the haiku psf__requests-863 over-build fix, now method) - AI-verify record ships IN the template (was an --oneshot splice); an agent-crossed freeze is declared via `gate_mode: ai-plan-verify` in the header — better audit than a flag - §5 gains `Spawn (multi-agent):` — build/verify spawns default worktree isolation; freeze --cross + refute-read via a cross-agent advise-mode spawn - removed: Grounding block, gherkin scaffold (§4 test_plan is the canonical encoding), assumptions ladder (ONE ⚠ line stays), 9 Build-strategy facets, Deep checks / Live-verify / Advisor 3-lens blocks, §7 Watch Engine: --fast/--oneshot/--thin/--full argparse + _strip_fast_sections + _FAST_SECTIONS + _fastlane_nudge/RISK_KEYWORDS + lane state markers deleted (creation side); READ side stays migration-tolerant (legacy fast/oneshot state keys still honored for skip-eligibility, route records measure-not- block). ENGINE_MD5 3e7e03c0 · ENGINE_PKG_MD5 fa82d37a @ atomic-node; 4-way tooling twins + 3-way skill/template trees byte-synced. Teaching + bench surfaces follow: SKILL.md flag mode now teaches the single- template doctrine + gate_mode (9,167B ≤ 9,500B ceiling); beyond/intake/verify guides de-laned; wm-bench + SWE prompts drop --oneshot and instruct the gate_mode declaration; the SWE ADD prompt names the host suite as the §3 Regression floor before the gate. Tests: 13 lane/grounding suites retired, ~40 re-aimed, test_template_atomic added (frozen 6-tag census · comment balance · 21 engine anchors · retired- surfaces stay retired · 3-tree parity). Corpus 2,290 passed / 3 env skips; benchmark suite 258 passed; add.py check 412/0. refs: #163 (feat/adaptive-flow) · benchmark/results/2026-07-swe-smoke.md author: Tin Dang
… into §3 Target
The atomic-node follow-through: the pre-registration job the block carried is
triple-covered — §3 `Target (measurable)` (judged at the gate via --target-hit),
§3 `Regression floor`, and the §6 refute-read. In SWE transcripts agents filled
it with "tests pass — confirmed by pytest": zero information past what the gate
already checks. Its one unique remainder — pre-declaring outcomes tests can't
show — folds into the Target guidance ("name any outcome tests can't show
(boots · renders) + how it's confirmed").
Template (4 twins): `### Build expectations` block removed; Target line extended.
Engine (4 twins): the opt-in `build_expectations_unfilled` gate retired — with
the template no longer scaffolding the block the gate is unfireable for new
plans; legacy plans with filled blocks pass unchanged, unfilled ones now cross
(measure-not-block direction). `_section_unfilled` survives serving the
contract-fill gate only. ENGINE_MD5 8d44e6ed · ENGINE_PKG_MD5 ec7f8093.
Teaching (3 skill trees): SKILL.md beat-1 drops the §6 mention (9,141B);
direction.md tests-production list trimmed; verify.md fill-before-build section
removed + stale Part-one drift cleaned (checkboxes still citing the retired
Build-expectations / Ground SHA / Live-verify surfaces re-aimed to §3 Target).
Tests: test_build_expectations_gate retired (predicate stays covered by
test_contract_fill_gate + test_engine_extract_predicates); freeze-precedence,
refute-ordering, form-tags, verify-rollup, template-atomic suites re-aimed
(block joins the retired-surfaces guard). SEAMS _declared_scope pin 4434→4427.
Corpus 2,281 passed / 3 env skips · benchmark 258 · add.py check 416/0.
refs: #163 (feat/adaptive-flow)
author: Tin Dang
…state.json The SWE installer dropped .add/tooling but never the engine state, so an agent that skips `add.py init` (haiku does) makes root discovery walk UP and find the HOST checkout's .add five levels above — the whole loop then runs against this repo (observed live: haiku wrote its psf__requests-863 task, src and tests into the host .add as fix-hooks-issue; yesterday's fix-issue/fix-hooks-lists residue was the same leak, not a smoke walk). install_add now ends with the engine init (`add.py init --name swe-fix --stage mvp`), anchoring discovery in the workspace; a WorkspaceIsolationTest guard pins the step. refs: #163 · sibling of harness-workspace-isolation a87ed1e (wm bench) author: Tin Dang
…on-1 resolution benchmark/results/2026-07-atomic-remeasure.md: the atomic-template re-run vs recorded baselines, officially evaluated — - SWE smoke: haiku 2/3 $2.52 -> 3/3 $2.15 (863 resolves; patch 1626B->559B; the host-suite Regression floor did it); sonnet 3/3 both templates at ~half the wall-clock (50.9 -> 28.1 min); first-eval 2317 miss shown to be a live-httpbin flake by a single-instance re-eval (zero P2P on re-run). - Cross-milestone session bench (ONE continuing conversation, 6 WMs): fidelity flat 1.0×6 vs old conv-carry .92->.80->.75->.17->.75->1.0 — rot eliminated in this sample; $17.75 vs $22.67; tests_weakened flags audited benign. - Harness defect record: the workspace-seed leak (fixed 9336076) and which runs are contaminated vs canonical. appendix-h: prediction-1 against-us note gains its resolution paragraph (the diagnosis became method — Regression floor line — and the re-run flipped the score back); swe-smoke report gains a superseded-for-ADD pointer, vanilla numbers + diagnosis stand. refs: #163 author: Tin Dang
…ta point atomic-remeasure gains the continue-mode spec-kit table (recorded arm, not re-run): spec-kit rots .92->.75 by wm3 AND ships regression rates .33/.29 in the mode where atomic ADD holds flat 1.0 with zero regressions — the honesty wave's "both arms decay identically" verdict no longer holds. Honest bounds kept: ~1.7x price gap survives; fresh-mode matrix un-rechallenged; appendix-H bottom line does not flip. appendix-h: second data point paragraph (continuing-conversation mode) beside the prediction-1 resolution. refs: #163 author: Tin Dang
…interfaces Milestone step 3 (ATG localized context). Measured driver: the session bench's late-WM cost growth (wm1 $1.87 -> wm4 $3.67, turns 74 -> 158) is interface RE-DISCOVERY — the agent re-reads the grown app to recall what prior tasks shipped. The engine now prints a `neighborhood` card at new-task and full status (never --brief): one line per inherited interface — parent slug · phase · the head of its frozen §3 contract fence · where its code lives. Parents = declared edges (depends_on ∪ extends); a board with NO edges falls back to the 2 most recently updated DONE tasks (the temporal neighborhood — bench agents declare no edges, measured across every board). Degrade-safe: unreadable or still-placeholder parent contracts are skipped; empty neighborhood prints nothing; card caps at 3 lines. SKILL.md beat 1 teaches "ground §3 in the card, not code re-reads" (9,252B ≤ 9,500). test_neighborhood_status.py: 8 tests red->green (declared-edge card, recency fallback, empty-board silence, --brief stays lean, unreadable/placeholder skip, 3-line cap). ENGINE_MD5 0d98f693 re-pinned; add.py x4 + skill x3 synced; SEAMS _declared_scope 4427->4480. Gate next: re-run the 6-WM session bench targeting <=$2.00/WM avg. refs: #163 · milestone step 3 of the ratified atomic plan author: Tin Dang
…ome real task-graph-native W1. The milestone is the scope ROOT of the task graph (depth = edges, not nesting) — but measured across every bench board the graph was EMPTY (deps=[] everywhere): waves had nothing to schedule, repair had no dependent closure, the neighborhood card fell back to recency. Three deterministic, propose-not-block mechanisms make it real: - compile — `milestone-confirm` reads MILESTONE.md's `## Tasks` list (`- [ ] <slug> depends-on: <none|slugs> — <line>`) into state.milestones[m].planned (the figure's "Compilation of T0"); prints `compiled task graph: N nodes · M edges`. Re-confirm RECOMPILES — the plan is living. Placeholder/malformed lines skip silently; the bare scaffold compiles to nothing. - inherit — `new-task <slug>` with no explicit --depends-on inherits the planned deps VERBATIM (creation-order-proof; a dangling forward edge is check's existing warn, never a lost edge). Explicit --depends-on wins. - hint — `freeze` prints `edge-hint:` when the just-declared §3 scope overlaps a DONE task's scope with no edge between them (the _declared_scope grammar both sides, containment, cap 2, silent on UNDECLARED) — a proposal the human/agent ratifies, never a refusal. SKILL.md intake line teaches the compile (9,360B ≤ 9,500). test_edge_truth.py 12 tests red->green. ENGINE_MD5 re-pinned; add.py x4 + skill x3 synced; SEAMS _declared_scope -> 4560. Next: W2 graph-repair (locate + minimal dependent closure on contract change). refs: #163 · task-graph-native W1 author: Tin Dang
…fy path task-graph-native W1.5, closing the hole W1 exposed: the freeze edge-hint proposed edges with no clean way to declare them post-creation (new-task flags are too late; MILESTONE.md re-confirm only updates the planned map, never existing records). - `add.py relate <slug> --depends-on/--extends/--relates-to <slugs>` (verb 32): ADDITIVE (append + dedup, never drops); validate-then-write on the source; dangling targets legal (forward edges — check's warn owns them); self-edge refused. The edge-hint now names it verbatim as its ratify step. - `milestone-confirm`'s compile warns on a depends-on cycle (would deadlock the wave schedule) — measure-not-block: the confirm stands, the fix is an edit + re-confirm. test_edge_truth.py 12->19 tests red->green. ENGINE_MD5 79b12cd4; twins x4 synced; SEAMS _declared_scope -> 4615. refs: #163 · task-graph-native W1.5 author: Tin Dang
…corrected test_min_pillar's self-maintaining census caught what the chained commit missed: verb 32 (`relate`) never ran under the read-spy. It now does — `new-task t2` + detach (`set-milestone t2 none`, so milestone-done stays green) + `relate t2 --relates-to t`. The CI fresh-checkout mirror suite reds until this lands (its clone runs HEAD's census against HEAD's parser). Also corrects the W1.5 docstrings: a dangling edge target is NOT merely "check's warn" — `add.py check` REDS a dangling ref until it resolves (no gate ever refuses; the diagnostic is honest, the flow unblocked). ENGINE_MD5 2d28aebe; SEAMS re-pinned. lesson: run the FULL corpus BEFORE `git commit` in the same chain — a census suite red after push costs a fix-forward commit. refs: #163 · task-graph-native W1.5 follow-up author: Tin Dang
… returned Reverts d5a2879. The card was gated on a re-run of the 6-WM continue-mode session bench (runs-nbr-session); it failed every axis vs the card-free atomic baseline (runs-atomic-session): fidelity 0.92->0.80->0.75->0.33-> 0.75->1.0 (vs flat 1.0x6), regression rate 0.20-0.38 at every WM (vs 0), total $23.57 (vs $17.75). The decay is a near-exact replay of the pre-atomic rot trajectory the atomic template had eliminated. Mechanism (from the transcripts, recorded in benchmark/results/2026-07-atomic-remeasure.md): the card externalizes PRECEDENT, not spec — wm2's card quoted wm1's shipped contract, deviation included — and its 90-char fence-head quote degenerates to auth boilerplate when a parent's S3 fence opens with the auth line (wm4 got zero interface signal and cratered to 0.33). Disclosed: edge-truth commits landed mid-bench so resolved_pin drifted across WMs; additive, unlikely causal, but the run is not pin-clean. Isolation was clean. Kept: edge-truth graph compile + relate verb (not part of the revert). Localized context returns as graph-native work (locate + dependent closure over declared edges), gated on its own bench. ENGINE_MD5 3a191a55 re-pinned; add.py x4 + engine_pin x2 + skill x3 synced; SEAMS _declared_scope 4616->4563; test_neighborhood_status.py removed; corpus 2300 passed. refs: #163 · gate: runs-nbr-session vs runs-atomic-session author: Tin Dang
… repair closure Task-graph-native wave 2 (the ATG figure's Failure Location + Minimal Necessary Subgraph Repair, deterministic — no LLM, read-only, verb 33). `add.py locate <ref>` in two modes: - test PATH -> the OWNING node, via each task's §4 `Tests live in:` declarations (reuses _declared_test_files) with the frozen §5 scope snapshot as fallback (provenance named), plus the failure class: `in-node` (owner still live — fix inside it; its frozen suite is the floor) vs `interface-regression` (owner DONE — a live change broke a settled contract; the owner's dependent closure prints as the repair set). Unowned paths report cleanly (exit 0 — a finding, not an error: treat as host/foreign regression-floor surface). - task SLUG -> the dependent closure directly: BFS reverse reachability over depends_on ∪ extends (relates_to is context, not interface — never enters), per-ring sorted for deterministic output, depth = shortest interface distance, DONE dependents kept (settled work re-verifies when its foundation moves). test_graph_repair.py: 11 tests red->green (owner mapping · both classes · closure-on-done-owner · scope fallback · unowned floor · transitive + extends closure · relates_to exclusion · leaf message · done-dependent kept). test_min_pillar LIFECYCLE gains `locate t` under the read-spy. SKILL.md build beat teaches the verb (9,405B <= 9,500, x3 trees). ENGINE_MD5 b2869fdd re-pinned; add.py x4 + engine_pin x4 synced; SEAMS _declared_scope 4563->4658; corpus 2,311 passed pre-commit. refs: #163 · task-graph-native W2 (W1 edge-truth 8d4f573 · W1.5 relate 0dadb17 · card revert f77f684) author: Tin Dang
…s frozen §3 clause
Task-graph-native wave 3 (the ATG red-dot at clause depth) + the skill
taught to drive the atomic graph loop.
Engine: `locate` grows the pytest node-id form — `locate path::test_name`
resolves the test through §4's covers map to the frozen §3 clause LINE.
The map's grammar is the template's OWN `<test_plan>` dialect (a §4 bullet
naming a test bare `test_…` or backticked, with a `covers: key[, key]`
tail) — frozen WITH the bundle, so it is tamper-guarded like the suite it
describes. Deterministic literal key match inside the §3 body; a key §3
doesn't carry is reported honestly ("not literal — the clause lives in
§1/§2 prose"), never guessed. The scaffold's unfilled placeholder bullet
parses to nothing (angle-bracket keys filtered, fail-safe); an unmapped
test is a nudge, never a gate. Plain-path and slug modes unchanged.
Skill (x3 trees) follows the new atomic graph flow:
- SKILL.md: direction beat fills covers keys; build beat teaches
locate path::test (9,486B <= 9,500).
- direction.md: "Clause map + edges" — covers tails + declare edges at
creation + ground §3 on parent edges' frozen PLAN.md (the spec-side,
pull-based grounding the reverted card got wrong).
- build.md: "A red outside your suite — locate first" (in-node ·
interface-regression · unowned=Regression-floor; repair the clause,
not the symptom).
- Template §4 comment documents the machine-read covers tail.
- Pool fence 33,496 held by compression, not a bump: connective fat cut
in direction.md/verify.md (all cut phrases pin-checked unpinned;
wording-lint 0 findings): pool 33,489.
test_clause_repair.py: 9 tests red->green (covers map · clause quote ·
template-native + backticked grammar · placeholder-safe · honest-missing
· advisory-unmapped · both mode floors). ENGINE_MD5 3a19e9dd re-pinned;
add.py x4 + templates x4 + engine_pin x4 + skill x3 synced; SEAMS
_declared_scope 4658->4710; corpus 2,320 passed pre-commit.
refs: #163 · task-graph-native W3 (W2 locate 49e588d)
author: Tin Dang
…rns on planned drift Task-graph-native wave 4, closing the ratified milestone's engine work. `add.py graph` (verb 34, read-only, print-only) renders the live board as a mermaid flowchart: each milestone is a subgraph wrapping its tasks (milestone = the scope ROOT; depth lives in edges, never nesting); edge style carries the edge type (depends-on solid --> · extends dashed -.-> · relates-to open -.-); node class carries phase (done green · live amber). The COMPILED plan renders too: a planned-but-never-created node appears dashed, and an edge target that resolves to an archived record is annotated instead of dangling. --milestone limits to one subgraph. Deterministic (sorted everywhere) — paste straight into a GitHub mermaid fence. The same drift is measured: `add.py check` gains a planned-drift WARN (never red — mid-milestone a planned-not-yet-created node is normal flow) naming each compiled `## Tasks` node with no live task and no archived record, with the ratify path (`new-task <slug>` inherits its planned depends-on, or re-confirm without it). test_graph_views.py: 9 tests red->green (flowchart render · milestone subgraph · three edge styles · phase classes · archived-target annotation · --milestone filter · dashed planned node · check warn + exit 0 · silent when all created). LIFECYCLE census gains `graph`. Dogfooded on this repo's own board (renders clean). No skill-pool spend: the verb is --help-discoverable; SKILL.md (14B slack) and the phases pool (7B slack) stay untouched. ENGINE_MD5 427a2501 re-pinned; add.py x4 + engine_pin x4 synced; SEAMS _declared_scope 4710->4796; corpus 2,329 passed pre-commit. refs: #163 · task-graph-native W4 (W1 8d4f573 · W2 49e588d · W3 75bb128) author: Tin Dang
Review finding pre-2.0.0 tag: `.add/dependencies.allowlist` claimed "CI rejects anything not listed" while NOTHING read the file (every corpus "allowlist" reference is npm's `files` tarball allowlist — a different thing), and its "Node installer uses built-in modules only" prose was false — package.json ships `@clack/prompts` as a real runtime dependency. The build exit gate "no dependency outside the allow-list" rested on an honor-system doc. Fix, red/green: - test_dependency_allowlist.py (3 tests) IS the missing rejection — the corpus runs in CI, so a declared runtime dep absent from the allowlist now reds the build. Scope = RUNTIME deps of the shipped package (npm `dependencies`, pyproject `[project] dependencies`); dev/bench tooling never ships and stays out of scope. Third test pins doc-truth: the zero-dep-installer claim may not coexist with declared npm deps. - .add/dependencies.allowlist: @clack/prompts recorded as the ONE approved Node runtime dependency (shipping since the installer-UX milestones — recording the approval is the honest state); prose corrected; header now cites the enforcing suite. No engine change (no ENGINE_MD5 repin). Corpus 2,332 passed. refs: #163 · user review ask pre-release author: Tin Dang
…ges + doc-truth Review of the updater (both twins) surfaced three gaps; fixed red/green (test_updater_2_0_gaps.py, 7 tests, node twin via subprocess): 1. GLOBAL ROSTER DRIFT (open since installer-shared-namespace #151): the global home mirror (GLOBAL_TREES / _GLOBAL_TREES) never carried `agents/`, and `update --global` propagation sources registered projects FROM the home — so the roster soft-skipped forever: no refresh, no retired-agent tombstone removal. Both mirrors now carry agents/; propagation reuses the existing SHARED per-file lander (user files in .claude/agents still survive; only the five explicit RETIRED_AGENTS tombstones are ever removed). 2. 2.0 CROSSING WAS SILENT: a 1.x project updating into 2.0 kept a TASK.md board and a stale CLAUDE.md guidance block with no signpost. `update` now prints crossing nudges — `add.py sync-guidelines` on any version cross, `add.py migrate` when a tasks/*/TASK.md without a PLAN.md sibling is detected. NAMED, never run: python3 may be absent on the updater's PATH; both commands are idempotent. Twin-identical wording npm/pip. 3. DOC-TRUTH: pip updater log claimed "docs refreshed" (docs stopped shipping at book-stops-shipping) -> "managed layer reconciled", matching the js twin; stale "Empty today." comment on the populated RETIRED_AGENTS list removed. No engine change (add.py untouched — no ENGINE_MD5 repin). Corpus 2,339 passed pre-commit. refs: #163 · user review ask pre-release · closes the #151 OPEN residue author: Tin Dang
…one fits User ask: spawn/seed a persona for each new task when missing. Implemented as domain-fit (NOT one-persona-per-task, which would sprawl near-duplicates): each task's DIRECTION beat establishes a persona that fits its domain — seeding one via the add agent's persona mode ONLY when none fits, reusing the existing persona across tasks otherwise. - SKILL.md beat 1: "load the domain-fit persona (seed via add persona-mode if none), then draft the bundle" — the always-read file now leads each new task with persona fit. - phases/direction.md persona callout: if none fits, spawn add persona-mode to seed from PROJECT.md + `.add/personas-teacher/`, then load it — seed per DOMAIN, REUSE across tasks, never one per task (the anti-sprawl floor). SKILL-only — no engine change, no ENGINE_MD5 repin. Both byte fences held by compression (SKILL 9,496 <= 9,500; phases pool 33,494 <= 33,496): connective fat cut in SKILL.md (--todo/flag-mode/lessons lines) and direction.md (milestone ground + relate-to-map items) — every cut phrase pin-checked unpinned; wording-lint 0 findings. test_persona_seed_on_task.py: 3 tests red->green across all 3 skill trees (beat-1 fit persona · direction seeds-when-none-fits via persona mode · anti- sprawl reuse). LESSON re-applied: the persona callout line-WRAPS "none fits" across a blockquote break — the assertions normalize whitespace (drop `\n>` + collapse) so a wrapped phrase-pin still matches (line-wrap-splits-phrase-pin). skill x3 synced; corpus 2,342 passed pre-commit. refs: #163 · user ask · builds on the persona-seed-nudge engine surface author: Tin Dang
…st orchestrating Review the /add SKILL.md against the skill-creator standard and apply the findings as byte-safe edits across all three byte-identical skill trees (canonical · _bundled · .claude mirror). Front-door capability (the headline): ADD now reads raw intent into a task shape BEFORE sizing it. New intake.md beat "Analyze the request before you size it" — restate the intent, extract the latent requirements, name the unstated, surface the hidden work — placed before Interview. The SKILL.md opening reframes the agent from "You are the orchestrator" to "You turn intent into the right task, then drive it": analyst first, orchestrator second. Review findings applied: - #1 design.md: udd-tokens.md / udd-catalog.md pointers were dangling siblings; qualified to templates/… so they resolve. - #2 SKILL.md setup branch now names the brownfield adopt.md path (one hop). - #3 new terms.md decodes the loop's coined vocabulary (compound-cross, co-specify, earned-green refute-read, re-cross, auto-resolved PASS, the ARC); linked from SKILL.md. - #5 7↔3 bridge: "three beats (seven steps, folded)" maps the description's seven named steps to the body's three beats. Constraints held: SKILL.md 9,496 → 9,497 B (under the test-enforced 9,500 ceiling — every addition funded by compressing unpinned prose, no pillar removed); all phrase-pins preserved; the three SKILL.md trees stay md5-identical. New guard: test_intake_analyze.py locks the Analyze section, its order before Interview, the reframed identity, the terms.md decoder link, and the qualified design.md refs. Full tooling suite green (2350 passed); wording-lint and semantic-inventory clean. author: Tin Dang
…autonomy dial
Run mode was two coupled halves: the autonomy dial AND an engine-managed
`streams:` posture (parallel/sequential) that `--run-mode` forced in lockstep
(auto→parallel, conservative→sequential). Real usage rarely runs parallel
agents, and the simplified Run-mode guide already frames concurrency as
"spawn a subagent per task" — not an engine posture. So the engine no longer
manages streams: run mode IS the autonomy dial; concurrency is doc-level.
Engine (removed the run-mode streams posture only — the separate multi-milestone
`streams : N active milestones` status view is untouched):
- constants.py: drop _STREAMS_POSTURES.
- autonomy.py: drop _streams_posture / _project_streams_token / _project_streams
and the streams regex.
- add.py: `--run-mode {auto,conservative}` now writes ONLY `autonomy:` (no
streams line); remove the _streams_decl_line writer and the status `run mode:`
row (redundant with `project autonomy:`); tidy the stale mirror comments.
Doc: phases/direction.md Run-mode block simplified — the vacuous Concurrency
column dropped (both rows were "one task"), the table axis renamed Autonomy,
and the tangled parallel-streams prose replaced by one line: spawn a subagent
per task; the one-approval-per-contract floor never moves. Pool 33479B (<33496).
Tests (red/green): test_setup_run_mode rewritten — --run-mode sets autonomy and
writes NO streams line; test_streams_posture.py deleted (tested only the removed
posture). Repin: ENGINE_MD5 + ENGINE_PKG_MD5 re-aimed @ run-mode-decouple,
synced across all engine trees. SEAMS.md scope-token-grammar anchor re-aimed
4796->4776 for the add.py line shift. Full tooling suite green (2346 passed).
author: Tin Dang
…e/graph model, trim dead rule
The 2.0 atomic template (`PLAN.md.tmpl`) shifted ADD to "persist the interface,
reason everything else in-context, don't write essays" and retired lane modes —
but the phase guides and the skill still taught the pre-atomic, essay-heavy,
lane-aware model. This aligns them and trims the dead weight the shift exposed.
direction.md — atomic-mode alignment (no split; the fold stays):
- Grounding section reframed from a 7-field written block ("gather BEFORE you
freeze") to the atomic rule: persist only the Contract's Anchors + an optional
Ground SHA; REASON Touches / Honors / Issues / Related-intent in-context and let
the frozen Contract encode them — don't transcribe an essay into the file.
- Dead lane framing corrected: the freeze no longer "ratifies the header route:
line — the persona's lane proposal" (lanes are retired, 0b0192a). route: is now
an audit-only optional line the freeze records; route_unrecorded is still measured.
- Grounding exit-gate line realigned to "reasoned in-context; the interface is
persisted, not the bullets."
SKILL.md — orchestrator -> atomic-graph navigator, and a dead rule dropped:
- The node paragraph now leads with "One task = one atomic node" and states the
graph model up front: the frozen §3 is the interface neighbor nodes depend on;
edges compile from the milestone; `graph` renders the DAG; `locate` walks a
failure to its node. Graph navigation is first-class, not a "stuck?" footnote.
- Removed constraint 5 "Ask, don't guess": unpinned, not ADD-specific (generic LLM
hygiene), already the always-loaded CLAUDE.md floor, and counter-directional to
the repo's own proactive-lead posture. Rules 1-4 (the freeze gate, evidence-over-
inspection, the tamper floor, the one recorded outcome incl. security HARD-STOP)
stay.
- `## The method rationale` collapsed from a 5-line section to one line.
Net: SKILL.md 9,497 -> 9,403 B (97 B of ceiling headroom reclaimed under the
9,500 test-enforced ceiling); the three SKILL.md / direction.md trees (canonical ·
_bundled · .claude mirror) stay byte-identical.
Verification: full tooling suite green; wording-lint 0 findings; semantic-inventory
unchanged (17 pre-existing findings on HEAD, 0 introduced). No test weakened; the
security-HARD-STOP teaching and every phrase-pin anchor survive.
author: Tin Dang
…oad-bearing
Phase 3 of the atomic-mode trim. Assessed verify.md (the second-largest phase
guide, 162 lines) for dead weight and found it is largely load-bearing: the
evidence checklist, the three lenses, the deep-check, the earned-green refute-read,
the gate-outcome table, the observe tail, and the ~55-line advisor spawn XML (the
canonical runner-agnostic worker contract, pinned by test_persona_subagent_prompt:
the four contract slots, the persona-section mapping, the {{PERSONA_SLUG}} slot, the
no-persona degrade path) all earn their bytes.
The one compressible section was the `## Sensitivity` risk-class vocabulary — dense
connective prose around a load-bearing core. Tightened it without dropping a single
fact or pinned string: the `sensitivity:` declaration, the base-four definitions
(security HARD-STOP · data · architecture · mechanical), the datetime/money/timezone
=> `data` guidance with the bench wm2 evidence, the `## Sensitivity classes` EXTEND
grammar, `sensitivity_invalid`, and `advisor-gate-relax` all survive verbatim.
verify.md 162 -> 160 lines; the three trees (canonical · _bundled · .claude mirror)
stay byte-identical.
Finding (the "what to optimize" answer): verify.md's complexity is intrinsic, not
sprawl — further shrink would remove a pillar, not ceremony. The atomic-mode wins
were in direction.md and SKILL.md (prior commit 3cd8792).
Verification: full tooling suite green; wording-lint 0 findings; no test weakened;
security-HARD-STOP teaching preserved.
author: Tin Dang
…(atomic drift sweep)
A drift sweep across every skill guide for retired-concept references (lanes,
route-as-lane, streams:, TASK.md, §0 Grounding, §6 Build-expectations, phase-count
framing) turned up exactly one genuine stale reference the earlier commits missed:
build.md's strategy-facets line anchored the domain facets upstream at
"§1 Framings · §3 Schema · §0 Honors". The 2.0 atomic template numbers §1–§7 — there
is no §0; the old §0 Grounding/Honors block folded into §3 (the engine still reads
`Honors`/`Ground SHA` by line-regex, section-agnostic, so nothing breaks). Re-aimed to
"§3 Schema · CONVENTIONS.md Honors" — pointing at the real source artifact rather than
a section that no longer exists. Both pinned build.md anchors on that line
("Approach (domain strategy", "Optimization stance") preserved verbatim.
Everything else the sweep flagged is live, not drift: the "inline lane" (a current
below-scope feature), the route scoreboard's "per-lane" evidence (the shipped M7
feature; route: is now an audit-only line), and "three beats (seven steps, folded)"
(the deliberate 7-to-3 bridge). No skill guide now references §0.
verify: full tooling suite green; wording-lint 0; three skill trees byte-identical.
author: Tin Dang
…es (verified pass) A per-passage verification of the book's pre-atomic drift against current engine behavior. Fixed only the genuinely-stale references; left every passage that touches a still-live nuance alone. Fixed (verified stale): - 07-step-5-build.md — the Pattern facet cited "(§0 Honors / CONVENTIONS.md)"; the 2.0 atomic template numbers §1–§7 with no §0, so the facet now cites "(CONVENTIONS.md Honors)" — the real source, mirroring the build.md guide fix. - 08-step-6-verify.md — the live-anchor check was framed around "§0's Ground SHA … record it in §6's Live-verify evidence block". test_template_atomic pins that the `### Live-verify evidence` and `### Grounding` blocks must NOT exist in the atomic template — the block is retired. Rewrote the paragraph to describe the surviving check (re-resolve every §3-cited symbol against the current tree, vs the shape it had at freeze) without the retired §0/Live-verify machinery. The concept still lives in the verify guide's Part-one checklist. - appendix-c-glossary.md — the phase-name table called grounding "the §0 grounding map"; dropped the stale §0 → "the grounding map". Root mirror synced: these chapters are byte-mirrored at the repo root (guarded by test_earned_green_rubric::test_root_book_matches_canonical for ch08); all three root copies re-synced to canonical. Deliberately NOT touched (live nuance / maintainer's call, flagged separately): - The components pillar (ch17 + the appendix-c Component entry): add.py comments say the component:/produces:/consumes: header grammar and components.toml schema-lint "died with the components pillar" (kernel-trim M5), and `component_green_bar_uncited` is gone from the engine — yet the PLAN.md.tmpl header still advertises `component:`. That contradiction is the engine's to resolve, not a drift-fix; ch17's fate is an editorial call. - "fast lane" glossary/component references: the `--fast`/`--oneshot` template scaffolds are retired, but `fast-lane-skips` (benchmark mode) and a `§0 GROUND` skip-rationale section are still live in the engine — so a blanket rewrite would be wrong. verify: full tooling suite green; wording-lint 0; canonical↔root book parity restored. author: Tin Dang
…task complexity Simplify the setup Run-mode step's concurrency line and fold in model-tier selection: a spawned subagent per task now picks its model by task complexity (mid ordinary · top complex), matching the roster tier contract in agents/add.md. The pinned run-mode content stays intact — `autonomy set --project`, `init --run-mode`, the sequential/auto table, parallel as the opt-in path, one-approval-per-contract floor. All three skill trees stay byte-identical. verify: full tooling suite green; test_setup_run_mode + parity guards pass. author: Tin Dang
… distil the loop The UDD design loop referenced its personas only at beat 4 (the confirm checklist), and restated all five beats a second time in a `## The hard rules` constraints block. This puts the personas first and cuts the redundancy — keeping the design-quality performance while making the guide leaner. Integrate (personas carry the expertise, the loop the discipline): - New "Personas carry the design — the loop carries the discipline" preamble loads the design-fit persona FIRST — the frontend-designer performance — before beat 0. Names both flow:design dimensions (UI-Designer: visual systems · component libraries · pixel-craft · WCAG-AA accessibility; UX-Researcher: evidence-validated, never assumed), and the seed path when none fits: .add/personas-teacher/design/ (ui-designer · ux-researcher · ux-architect) + engineering-frontend-developer via the add agent in persona mode, seeded per DOMAIN and reused across screens. The persona's Critical Rules shape every beat; its Success Metrics become the beat-4 checklist. - Beat 4 now references the already-loaded personas instead of re-loading them. Distil: - Dropped the `## The hard rules` constraints block — every rule in it (intake-before- domain, domain-first, reuse-before-invent, confirm-before-build, engine-never-renders, bind-don't-break, confirm-against-personas) was a verbatim restatement of the beat it came from; each survives inline in its beat, the capture section, or the new preamble. design.md 5,875 -> 5,416 B; the run+loop+design reclaim pool holds at 18,449 <= 20,138. All three skill trees stay byte-identical. verify: full tooling suite green; wording-lint 0; every design.md phrase-pin (five beats + axes/options · the UI-Designer/UX-Researcher checklist · never-an-auto-pass · never-lowers-a-gate · no-ui-personas degrade · the engine-never-renders invariant · the capture convention) preserved. author: Tin Dang
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
ADD 2.0 — skill-led method on a thin state kernel
The 2.0 major, M1–M8 (design ratified 2026-07-17): the skill drives the loop,
add.pyshrinks to a state kernel, personas become the method's adaptive brain, specs become living 5-DD documents, and the book stops shipping with the package.Subsumes
feat/strategy-intake(that branch never got a PR; its 14 unmerged commits — persona-owns-gates, thin-lane, engine-only teardown, md-cut sweep — are all ancestors of this branch, so this single PR lands both lines).The 2.0 line (M1–M8)
61fa230a,638cc033,f2a3e455,99a3a5e1): persona task-kinds + route-outcome traces; 12 capability seed templates; 5 phase agents collapse into ONEaddagent.b82fcbbd,ffc47d4a): §3 PLAN carries a measurable Target the gate judges; five living specs + thedelta-appendkernel verb.e2b5b4b6,a68da012,d74a58bd,d593e03e): AI-routed below-scope route; SKILL.md IS the loop (phases 7→3 on-demand refs); persona proposes the lane, the freeze ratifies it.a7ef15e6,9afa0efb,91230a3a,e2d89a93,de70180a): 54→31 verbs, add.py 9,558→6,596 lines; 6 phases → direction·build·verify; ONE task template; test corpus −180 duplicated pin tests.c36bc484,3532a03f,f8c2a58e): task doc TASK.md→PLAN.md with one-shotadd.py migrate(480 board docs dogfooded); book publishes at https://pilotspace.github.io/ADD/ only; 31 legacy milestones archived.ea3b5bf2):deltasreads the route traces into a per-lane scoreboard + the loop's GEPA reflection beat.f3f3dd6b,015048d4,52412c36): version lockstep across all 5 sources + CHANGELOG with Breaking block + release suite;--session-mode continuecontext-rot benchmark; results report.Since opening — honesty wave, atomic node, task graph
Bench honesty wave (
f4ffbe6c): 6 meter defects found and fixed (tempdir store pollution, foreign-app import, conftest collision, …); the first-edition "spec-kit collapses under evolution" claim was traced to our own meter and publicly retracted — reports revised in place, READMEs re-receipted. Corrected fresh-mode matrix: add 6/6 all-floors $17.51 · spec-kit 6/6 $10.05.Atomic node (
0b0192a7,121c2feb): ONE atomic PLAN.md per task (interface = frozen contract · red suite · scope · verdict); every lane scaffold removed; §3 gains aRegression floor:binding the host repo's own suite at the gate.SWE smoke, re-run under the official swebench docker eval (
93360765,a4e7d677): haiku+ADD 2/3 → 3/3, cheaper ($2.15 vs $2.52; the miss's patch shrank 1,626B→559B) — appendix-H prediction 1 flips back at the tier where it failed; sonnet parity 3/3 at nearly half the wall-clock (50.9→28.1 min). Harness defect fixed: SWE workspaces now seed their own.addstate (ancestor-root leak).Context rot, remeasured (
a4e7d677,1d153795): with the atomic template, ONE continuing conversation over six milestones holds fidelity flat 1.0×6 with zero regressions, 22% cheaper — the mode that previously decayed 0.92→0.75 in BOTH flows. Spec-kit's recorded board in the same mode still decays with regression rates 0.33/0.29. Raw price gap (~1.7×) unchanged; appendix-H updated with both data points.Task-graph-native (
8d4f573b,0dadb172,49e588d2,75bb1289,d5d9f0fd): milestone = scope ROOT over the task DAG.milestone-confirmcompiles## Tasksinto edges new-task inherits;relateratifies edges post-creation (freeze prints edge-hints);locatemaps a failing test → owning node → in-node vs interface-regression → dependent closure (the minimal repair subgraph), andpath::testresolves through §4covers:to the exact frozen §3 clause;graphrenders the DAG as mermaid with dashed planned-never-created nodes;checkwarns on planned drift. Verbs 31→34.Negative result, published (
d5a28791→ revertedf77f684e): a neighborhood-status card (ambient prints of parent contract heads) failed its gate bench — the rot curve returned at +33% cost because the card externalizes precedent, not spec. Recorded in the results doc as a design floor: localized context must be pull-based, spec-side, and complete.Method refinements since the graph waves (
8feeec4d,57132feb,aa9f1f14,3cd8792a,0a56dede,a91f7d04): the skill now seeds a domain-fit persona at every new task when none fits; its front door analyzes intent into a task before sizing (analyst-first, not just orchestrator); run mode decouples to the pure autonomy dial (thestreams:posture subsystem removed — concurrency is "a subagent per task"). Then an atomic-mode alignment sweep of the phase guides: direction.md's Grounding folds from a 7-field written block to "persist the interface, reason the rest in-context" and the retired lane/route:framing is corrected; SKILL.md reframes from orchestrator to atomic-graph navigator (node · edges ·graph/locateup front) and drops the one non-ADD-specific constraint; verify.md's sensitivity vocab is compressed (the rest is load-bearing); and a drift sweep retires the last stale§0reference (build.md re-aims its facet anchor toCONVENTIONS.md Honors). The guides now match the 2.0 atomic template; SKILL.md holds under its 9,500 B ceiling; all three skill trees stay byte-identical.Measured (benchmark/results/…)
Verification
add.py checkclean; wording-lint clean.e0980435):.add/dependencies.allowlistclaimed CI rejection nothing implemented — a new corpus suite rejects any unlisted runtime dep;@clack/promptsrecorded as the one approved Node dependency.50ee1cce): global home now mirrorsagents/(the fix(installer): .claude/agents is a shared namespace — stop wiping user subagents on init/update #151 roster-drift residue —update --globalpropagation finally refreshes rosters + applies retired tombstones);updateprints crossing nudges (sync-guidelineson any version cross,migrateon a TASK.md-era board); stale "docs refreshed" prose corrected. Corpus 2,339.bfd47201…/ PKGe0ff925d…@ run-mode-decouple (the last engine change; the method-refinement sweep edits skill guides only — engine untouched); full corpus 2,346 passed at HEAD; twins synced (tooling ×4, skill ×3, engine_pin ×4).Remaining after merge (human-gated): tag
v2.0.0+ npm/PyPI publish (in-bandpublish.ymlon tag push).author: Tin Dang