diff --git a/.agents/skills/tldrgraph-init/SKILL.md b/.agents/skills/tldrgraph-init/SKILL.md index e7e58cf..8b62d20 100644 --- a/.agents/skills/tldrgraph-init/SKILL.md +++ b/.agents/skills/tldrgraph-init/SKILL.md @@ -1,178 +1,24 @@ --- name: tldrgraph-init -description: Build or continue this repository's TLDRGraph architecture graph (layers, extraction, enrichment) +description: Build or refresh source-backed feature workflows --- -# TLDRGraph: build this repository's architecture graph +# TLDRGraph: build the source-backed workflow catalog -In Claude Code or Cursor, invoke `/tldrgraph-init`. In Codex CLI, type `/skills` -and select `tldrgraph-init`, or mention `$tldrgraph-init` directly. - -One command handles layer design, extraction, enrichment, LLM route links, and embeddings: +For a first catalog, run: ```bash tldrgraph init ``` -By default TLDRGraph detects `claude`, `cursor-agent`, or `gemini`, processes every -eligible node in batches of 200, and downloads/builds the local embedding model. -In a detected coding-agent session, plain `tldrgraph init` auto-approves the full -enrichment campaign; normal terminal users and non-agent automation still get the -confirmation gate. - -Use exactly `tldrgraph init` for this workflow. Add `--yes` only if a non-agent -`needs_confirmation` response explicitly asks for approval. `--batch 200` means -all nodes in 200-node batches; `--limit 200` means stop after only 200 nodes. -Never add `--limit` or `--embeddings off` unless the user explicitly requests a -partial or no-embedding run. - -## Completion contract for agents - -Once a full enrichment campaign is approved or auto-approved, you are not done until one of -these terminal states occurs: - -- `tldrgraph init` reports `status: done`. -- Dense embeddings finish, or `init` explicitly reports that embeddings are - unavailable while preserving a queryable graph. -- A real blocking error occurs that requires user action, and you report the - exact command/output that blocked continuation. - -These are **not** terminal states: - -- `status: needs_enrichment` -- `status: needs_llm_links` -- `NEXT ACTION` -- "nodes remaining" -- "batches remaining" -- "applied N enrichment entries" - -For every non-terminal enrichment state, immediately continue the loop: - -1. Read `.tldrgraph/enrichment_request.yaml`. -2. Open the source file for every requested node. -3. Write `.tldrgraph/enrichment_response.yaml`. -4. Run `tldrgraph init` again. -5. Repeat until a terminal state occurs. - -Do not end the task with a progress-only summary such as "I applied 400 entries -and 1,438 remain." That is an incomplete run, not a final answer. If nested-agent -protection prevents a CLI from launching another agent, you are the enrichment -agent and must process the batch yourself. - -If no supported agent is available or dense embeddings cannot be built, `init` -preserves the graph and prints a resumable status. It never guesses source intent, -route links, or architectural layers. - -## `status: needs_layers` - -TLDRGraph ships **no layer templates** and will not invent an architecture. -Design one from this repository. - -1. Read `.tldrgraph/propose_layers_request.json`. It carries the symbols and - files extraction already found -- a starting point, not a substitute for - opening the code. -2. **Open real source files**: entry points first, then a representative file - from each cluster in the evidence. Work out what this codebase actually does - and where responsibility changes hands. -3. Write `.tldrgraph/propose_layers_response.json`: - -```json -{ - "utility_id": "", - "layers": [ - { - "id": "short_machine_id", - "name": "Layer 1: Human Friendly Name", - "order": 1, - "description": "One sentence on what lives here", - "rules": [ - {"file_contains": ["substring"], "exclude_file": ["optional"]}, - {"label_contains": ["SymbolNamePart"]} - ] - } - ] -} -``` - -4. Run `tldrgraph init` again; it continues with enrichment and embeddings. - -### Rules that hold for any answer - -- 3 to 6 layers, plus exactly one catch-all whose `id` equals `utility_id` and - whose `rules` are `[]`. -- Unique `id` and `name` per layer; sequential integer `order` from 1. -- Rule keys: `file_contains`, `exclude_file`, `path_regex`, `label_contains`, - `exclude_label`, `label_ends_with`, `type_in`, `id_prefix`. Values are lists of - strings. Rules are evaluated in `order` and the first match wins. -- Derive rules from paths and symbol names you actually saw. A rule matching - nothing is worse than no rule; a rule matching everything collapses the map. - -## `status: needs_confirmation` - -Detected coding-agent sessions should not reach this state for a full run. If a -non-agent run does, the output shows how many nodes need enrichment and how many -agent round-trips that implies; ask the user whether to proceed. - -- They agree: `tldrgraph init --yes` saves approval for the full campaign -- Smaller first pass: `tldrgraph init --yes --limit 100` -- They decline: stop. The graph is already built and queryable. - -## `status: needs_enrichment` - -1. Read `.tldrgraph/enrichment_request.yaml`. -2. **Open the source file of every node in it.** This is the entire point: an - intent paraphrased from a symbol name poisons semantic search with - confident-sounding noise. -3. Write `.tldrgraph/enrichment_response.yaml` -- a *different* file from the - request, which is regenerated on every run: - -```yaml -- id: "" - intent: | - What this symbol does, why it exists, and its execution logic. - input_fields: [caseId, remarks] - output_fields: [status, disposition] - calls: [ApplicationsService, pension_cases] -``` - -Every `intent` must contain **2-3 complete sentences** covering what the symbol does, -why it exists, and its source-backed behavior. Markdown headings and list markers do not -count as sentences. +When the status is `needs_feature_workflows`: -4. Run `tldrgraph init` again. Approval is saved; process any next - `needs_enrichment` batch immediately without asking the user again until - init either reports `status: done` or advances to the required `needs_llm_links` - phase. - -Inside an existing Codex/Claude/Cursor session, nested-agent protection may stop -the CLI from launching a second agent. In that case **you are the enrichment -agent**: process every 200-node batch yourself. A `needs_enrichment` status is a -continuation instruction, not a reason to stop or request confirmation. - -**Copy every `id` verbatim.** A constructed id matches nothing, is dropped, and -gets reported back to you -- but the work is wasted. - -**Never invent `fields` or `calls`.** Omit what you cannot verify in the code: an -empty list is a correct answer, a wrong `calls` entry becomes a real wrong edge. - -## `status: needs_llm_links` - -1. Read `.tldrgraph/llm_links_request.yaml`. -2. Open the referenced frontend and backend source files. -3. Write `.tldrgraph/llm_links_response.yaml` as a YAML list of - `{source, target, confidence, frontend_evidence, backend_evidence, explanation}`. - Only include source-backed links with file and line evidence. -4. Run `tldrgraph init` again. This is a required continuation state for a - complete init run unless the user explicitly requested `--no-llm-links`. - -## Once it says DONE - -```bash -tldrgraph query "" -tldrgraph trace "" "" -tldrgraph layers -tldrgraph ui --serve -``` +1. Read the `source_hash` and validation status returned by the command. +2. Identify feature outcomes, define areas, and write the v4 + `.tldrgraph/features.yaml` index before delegating workflow research. +3. Spawn one fresh source-reading subagent for each indexed feature. Never assign the entire catalog to one subagent. +4. Have each worker write only its evidence-backed v4 workflow to its assigned + `.tldrgraph/workflows/.yaml` file; do not collect or rewrite it. +5. Run the same command again. Use `tldrgraph refresh` when updating an existing catalog after source changes. -Read-only, and they never trigger enrichment. Full schema: -`.tldrgraph/AGENT_CONTRACT.md`. +Continue until the status is `done`. diff --git a/.agents/skills/tldrgraph-refresh/SKILL.md b/.agents/skills/tldrgraph-refresh/SKILL.md new file mode 100644 index 0000000..2614234 --- /dev/null +++ b/.agents/skills/tldrgraph-refresh/SKILL.md @@ -0,0 +1,24 @@ +--- +name: tldrgraph-refresh +description: Refresh source-backed feature workflows after repository source changes +--- + +# TLDRGraph: refresh the source-backed workflow catalog + +Run exactly: + +```bash +tldrgraph refresh +``` + +When the status is `needs_feature_workflows`: + +1. Read the `source_hash` and validation status returned by `tldrgraph refresh`. +2. Identify feature outcomes, define areas, and write the v4 + `.tldrgraph/features.yaml` index before delegating workflow research. +3. Spawn one fresh source-reading subagent for each indexed feature. Never assign the entire catalog to one subagent. +4. Have each worker write only its evidence-backed v4 workflow to its assigned + `.tldrgraph/workflows/.yaml` file; do not collect or rewrite it. +5. Run `tldrgraph refresh` again. + +Continue until the status is `done`. diff --git a/.claude/commands/tldrgraph-init.md b/.claude/commands/tldrgraph-init.md index e7e58cf..8b62d20 100644 --- a/.claude/commands/tldrgraph-init.md +++ b/.claude/commands/tldrgraph-init.md @@ -1,178 +1,24 @@ --- name: tldrgraph-init -description: Build or continue this repository's TLDRGraph architecture graph (layers, extraction, enrichment) +description: Build or refresh source-backed feature workflows --- -# TLDRGraph: build this repository's architecture graph +# TLDRGraph: build the source-backed workflow catalog -In Claude Code or Cursor, invoke `/tldrgraph-init`. In Codex CLI, type `/skills` -and select `tldrgraph-init`, or mention `$tldrgraph-init` directly. - -One command handles layer design, extraction, enrichment, LLM route links, and embeddings: +For a first catalog, run: ```bash tldrgraph init ``` -By default TLDRGraph detects `claude`, `cursor-agent`, or `gemini`, processes every -eligible node in batches of 200, and downloads/builds the local embedding model. -In a detected coding-agent session, plain `tldrgraph init` auto-approves the full -enrichment campaign; normal terminal users and non-agent automation still get the -confirmation gate. - -Use exactly `tldrgraph init` for this workflow. Add `--yes` only if a non-agent -`needs_confirmation` response explicitly asks for approval. `--batch 200` means -all nodes in 200-node batches; `--limit 200` means stop after only 200 nodes. -Never add `--limit` or `--embeddings off` unless the user explicitly requests a -partial or no-embedding run. - -## Completion contract for agents - -Once a full enrichment campaign is approved or auto-approved, you are not done until one of -these terminal states occurs: - -- `tldrgraph init` reports `status: done`. -- Dense embeddings finish, or `init` explicitly reports that embeddings are - unavailable while preserving a queryable graph. -- A real blocking error occurs that requires user action, and you report the - exact command/output that blocked continuation. - -These are **not** terminal states: - -- `status: needs_enrichment` -- `status: needs_llm_links` -- `NEXT ACTION` -- "nodes remaining" -- "batches remaining" -- "applied N enrichment entries" - -For every non-terminal enrichment state, immediately continue the loop: - -1. Read `.tldrgraph/enrichment_request.yaml`. -2. Open the source file for every requested node. -3. Write `.tldrgraph/enrichment_response.yaml`. -4. Run `tldrgraph init` again. -5. Repeat until a terminal state occurs. - -Do not end the task with a progress-only summary such as "I applied 400 entries -and 1,438 remain." That is an incomplete run, not a final answer. If nested-agent -protection prevents a CLI from launching another agent, you are the enrichment -agent and must process the batch yourself. - -If no supported agent is available or dense embeddings cannot be built, `init` -preserves the graph and prints a resumable status. It never guesses source intent, -route links, or architectural layers. - -## `status: needs_layers` - -TLDRGraph ships **no layer templates** and will not invent an architecture. -Design one from this repository. - -1. Read `.tldrgraph/propose_layers_request.json`. It carries the symbols and - files extraction already found -- a starting point, not a substitute for - opening the code. -2. **Open real source files**: entry points first, then a representative file - from each cluster in the evidence. Work out what this codebase actually does - and where responsibility changes hands. -3. Write `.tldrgraph/propose_layers_response.json`: - -```json -{ - "utility_id": "", - "layers": [ - { - "id": "short_machine_id", - "name": "Layer 1: Human Friendly Name", - "order": 1, - "description": "One sentence on what lives here", - "rules": [ - {"file_contains": ["substring"], "exclude_file": ["optional"]}, - {"label_contains": ["SymbolNamePart"]} - ] - } - ] -} -``` - -4. Run `tldrgraph init` again; it continues with enrichment and embeddings. - -### Rules that hold for any answer - -- 3 to 6 layers, plus exactly one catch-all whose `id` equals `utility_id` and - whose `rules` are `[]`. -- Unique `id` and `name` per layer; sequential integer `order` from 1. -- Rule keys: `file_contains`, `exclude_file`, `path_regex`, `label_contains`, - `exclude_label`, `label_ends_with`, `type_in`, `id_prefix`. Values are lists of - strings. Rules are evaluated in `order` and the first match wins. -- Derive rules from paths and symbol names you actually saw. A rule matching - nothing is worse than no rule; a rule matching everything collapses the map. - -## `status: needs_confirmation` - -Detected coding-agent sessions should not reach this state for a full run. If a -non-agent run does, the output shows how many nodes need enrichment and how many -agent round-trips that implies; ask the user whether to proceed. - -- They agree: `tldrgraph init --yes` saves approval for the full campaign -- Smaller first pass: `tldrgraph init --yes --limit 100` -- They decline: stop. The graph is already built and queryable. - -## `status: needs_enrichment` - -1. Read `.tldrgraph/enrichment_request.yaml`. -2. **Open the source file of every node in it.** This is the entire point: an - intent paraphrased from a symbol name poisons semantic search with - confident-sounding noise. -3. Write `.tldrgraph/enrichment_response.yaml` -- a *different* file from the - request, which is regenerated on every run: - -```yaml -- id: "" - intent: | - What this symbol does, why it exists, and its execution logic. - input_fields: [caseId, remarks] - output_fields: [status, disposition] - calls: [ApplicationsService, pension_cases] -``` - -Every `intent` must contain **2-3 complete sentences** covering what the symbol does, -why it exists, and its source-backed behavior. Markdown headings and list markers do not -count as sentences. +When the status is `needs_feature_workflows`: -4. Run `tldrgraph init` again. Approval is saved; process any next - `needs_enrichment` batch immediately without asking the user again until - init either reports `status: done` or advances to the required `needs_llm_links` - phase. - -Inside an existing Codex/Claude/Cursor session, nested-agent protection may stop -the CLI from launching a second agent. In that case **you are the enrichment -agent**: process every 200-node batch yourself. A `needs_enrichment` status is a -continuation instruction, not a reason to stop or request confirmation. - -**Copy every `id` verbatim.** A constructed id matches nothing, is dropped, and -gets reported back to you -- but the work is wasted. - -**Never invent `fields` or `calls`.** Omit what you cannot verify in the code: an -empty list is a correct answer, a wrong `calls` entry becomes a real wrong edge. - -## `status: needs_llm_links` - -1. Read `.tldrgraph/llm_links_request.yaml`. -2. Open the referenced frontend and backend source files. -3. Write `.tldrgraph/llm_links_response.yaml` as a YAML list of - `{source, target, confidence, frontend_evidence, backend_evidence, explanation}`. - Only include source-backed links with file and line evidence. -4. Run `tldrgraph init` again. This is a required continuation state for a - complete init run unless the user explicitly requested `--no-llm-links`. - -## Once it says DONE - -```bash -tldrgraph query "" -tldrgraph trace "" "" -tldrgraph layers -tldrgraph ui --serve -``` +1. Read the `source_hash` and validation status returned by the command. +2. Identify feature outcomes, define areas, and write the v4 + `.tldrgraph/features.yaml` index before delegating workflow research. +3. Spawn one fresh source-reading subagent for each indexed feature. Never assign the entire catalog to one subagent. +4. Have each worker write only its evidence-backed v4 workflow to its assigned + `.tldrgraph/workflows/.yaml` file; do not collect or rewrite it. +5. Run the same command again. Use `tldrgraph refresh` when updating an existing catalog after source changes. -Read-only, and they never trigger enrichment. Full schema: -`.tldrgraph/AGENT_CONTRACT.md`. +Continue until the status is `done`. diff --git a/.cursor/commands/tldrgraph-init.md b/.cursor/commands/tldrgraph-init.md index 9fa8c56..8b62d20 100644 --- a/.cursor/commands/tldrgraph-init.md +++ b/.cursor/commands/tldrgraph-init.md @@ -1,134 +1,24 @@ --- name: tldrgraph-init -description: Build or continue this repository's TLDRGraph architecture graph (layers, extraction, enrichment) +description: Build or refresh source-backed feature workflows --- -# TLDRGraph: build this repository's architecture graph +# TLDRGraph: build the source-backed workflow catalog -In Claude Code or Cursor, invoke `/tldrgraph-init`. In Codex CLI, type `/skills` -and select `tldrgraph-init`, or mention `$tldrgraph-init` directly. - -One command handles layer design, extraction, enrichment, and embeddings: +For a first catalog, run: ```bash tldrgraph init ``` -By default TLDRGraph detects `claude`, `cursor-agent`, or `gemini`, asks once before enrichment token spend, processes every -eligible node in batches of 200, and downloads/builds the local embedding model. -Use `--yes` for non-interactive approval, `--batch N` to override the batch size, -`--embeddings off|auto|on` to override embeddings, or `--no-agent-cli` for the -manual file handoff. - -After the user approves the full run, use exactly `tldrgraph init --yes`. The -approval is saved for the current candidate set, so later `tldrgraph init` calls -must continue without asking again. `--batch 200` means all nodes in 200-node -batches; `--limit 200` means stop after only 200 nodes. Never add `--limit` or -`--embeddings off` unless the user explicitly requests a partial or no-embedding run. - -If no supported agent is available or dense embeddings cannot be built, `init` -preserves the graph and prints a resumable status. It never guesses source intent -or architectural layers. - -## `status: needs_layers` - -TLDRGraph ships **no layer templates** and will not invent an architecture. -Design one from this repository. - -1. Read `.tldrgraph/propose_layers_request.json`. It carries the symbols and - files extraction already found -- a starting point, not a substitute for - opening the code. -2. **Open real source files**: entry points first, then a representative file - from each cluster in the evidence. Work out what this codebase actually does - and where responsibility changes hands. -3. Write `.tldrgraph/propose_layers_response.json`: - -```json -{ - "utility_id": "", - "layers": [ - { - "id": "short_machine_id", - "name": "Layer 1: Human Friendly Name", - "order": 1, - "description": "One sentence on what lives here", - "rules": [ - {"file_contains": ["substring"], "exclude_file": ["optional"]}, - {"label_contains": ["SymbolNamePart"]} - ] - } - ] -} -``` - -4. Run `tldrgraph init` again; it continues with enrichment and embeddings. - -### Rules that hold for any answer - -- 3 to 6 layers, plus exactly one catch-all whose `id` equals `utility_id` and - whose `rules` are `[]`. -- Unique `id` and `name` per layer; sequential integer `order` from 1. -- Rule keys: `file_contains`, `exclude_file`, `path_regex`, `label_contains`, - `exclude_label`, `label_ends_with`, `type_in`, `id_prefix`. Values are lists of - strings. Rules are evaluated in `order` and the first match wins. -- Derive rules from paths and symbol names you actually saw. A rule matching - nothing is worse than no rule; a rule matching everything collapses the map. - -## `status: needs_confirmation` - -The output shows how many nodes need enrichment and how many agent round-trips -that implies. **Ask the user whether to proceed, and show them that estimate.** -Do not decide for them. - -- They agree: `tldrgraph init --yes` saves approval for the full campaign -- Smaller first pass: `tldrgraph init --yes --limit 100` -- They decline: stop. The graph is already built and queryable. - -## `status: needs_enrichment` - -1. Read `.tldrgraph/enrichment_request.yaml`. -2. **Open the source file of every node in it.** This is the entire point: an - intent paraphrased from a symbol name poisons semantic search with - confident-sounding noise. -3. Write `.tldrgraph/enrichment_response.yaml` -- a *different* file from the - request, which is regenerated on every run: - -```yaml -- id: "" - intent: | - What this symbol does, why it exists, and its execution logic. - input_fields: [caseId, remarks] - output_fields: [status, disposition] - calls: [ApplicationsService, pension_cases] -``` - -Every `intent` must contain **2-3 complete sentences** covering what the symbol does, -why it exists, and its source-backed behavior. Markdown headings and list markers do not -count as sentences. - -4. Run `tldrgraph init` again. Approval is already saved. If another - `needs_enrichment` batch appears, process it immediately and repeat this loop - without asking the user again. Continue until `status: done`. +When the status is `needs_feature_workflows`: -Inside an existing Codex/Claude/Cursor session, nested-agent protection may stop -the CLI from launching a second agent. In that case **you are the enrichment -agent**: process every 200-node batch yourself. A `needs_enrichment` status is a -continuation instruction, not a reason to stop or request confirmation. - -**Copy every `id` verbatim.** A constructed id matches nothing, is dropped, and -gets reported back to you -- but the work is wasted. - -**Never invent `fields` or `calls`.** Omit what you cannot verify in the code: an -empty list is a correct answer, a wrong `calls` entry becomes a real wrong edge. - -## Once it says DONE - -```bash -tldrgraph query "" -tldrgraph trace "" "" -tldrgraph layers -tldrgraph ui --serve -``` +1. Read the `source_hash` and validation status returned by the command. +2. Identify feature outcomes, define areas, and write the v4 + `.tldrgraph/features.yaml` index before delegating workflow research. +3. Spawn one fresh source-reading subagent for each indexed feature. Never assign the entire catalog to one subagent. +4. Have each worker write only its evidence-backed v4 workflow to its assigned + `.tldrgraph/workflows/.yaml` file; do not collect or rewrite it. +5. Run the same command again. Use `tldrgraph refresh` when updating an existing catalog after source changes. -Read-only, and they never trigger enrichment. Full schema: -`.tldrgraph/AGENT_CONTRACT.md`. +Continue until the status is `done`. diff --git a/.gitignore b/.gitignore index cff9126..53961f3 100644 --- a/.gitignore +++ b/.gitignore @@ -85,9 +85,7 @@ Thumbs.db !.vscode/extensions.json # BEGIN TLDRGRAPH -# TLDRGraph analysis state. Generated artifacts are ignored; the agent -# contract and layer map are committed so the whole team shares them. +# TLDRGraph generated workflow state. .tldrgraph/* !.tldrgraph/AGENT_CONTRACT.md -!.tldrgraph/layers.config.yaml # END TLDRGRAPH diff --git a/.tldrgraph/AGENT_CONTRACT.md b/.tldrgraph/AGENT_CONTRACT.md index b2655be..2589a4f 100644 --- a/.tldrgraph/AGENT_CONTRACT.md +++ b/.tldrgraph/AGENT_CONTRACT.md @@ -1,285 +1,66 @@ # TLDRGraph Agent Contract -**Audience: the coding agent with this repository open** (Codex, Claude Code, Cursor, Antigravity). - -TLDRGraph builds an architectural graph from the graphify AST export. The layer set -itself is designed by you, reading this repository. TLDRGraph ships no layer templates. -Structure *within* a layer, and the high-volume deterministic seams between layers, are -extracted automatically. What cannot be extracted automatically is: - -- indirect dispatch, queue / event hops, dynamically-built routes; -- the natural-language **intent** that makes semantic search work at all. - -That is your job. You are not a fallback for a hosted model — you are the primary -enrichment path, because **you can open the files**. The API path only ever sees a label -and a path (`snippet` is never populated), so it guesses. You do not have to guess. - ---- - -## Start here: `tldrgraph init` - -One command handles layers, extraction, enrichment, and embeddings: - -```bash -tldrgraph init +TLDRGraph builds a source-backed feature catalog and Workflow Explorer. It does +not construct an architecture graph and does not launch an AI process itself. + +## Direct artifact workflow + +Run `tldrgraph init` to create a catalog. Run `tldrgraph refresh` to update an +existing catalog after source changes. A missing, invalid, or stale catalog returns +`needs_feature_workflows` and prints the current `source_hash`. The active coding +agent identifies feature outcomes and writes the catalog index first. It then +delegates each indexed outcome to a separate source-reading subagent. Each +worker writes only its assigned workflow file; the active agent does not collect +worker objects or write workflows: + +```text +.tldrgraph/features.yaml +.tldrgraph/workflows/.yaml ``` -It asks once before enrichment token spend. Full approval is persisted for the current -candidate set until enrichment finishes, so continuation runs must not ask again. By -default it uses 200-node batches and builds dense embeddings. - -| status | what it wants | -| --- | --- | -| `needs_layers` | Read the code and design this repository's architecture. **TLDRGraph ships no layer templates**; nothing will be applied for you. The request carries sketches of how other kinds of codebase divide — for shape only, never to copy. | -| `needs_confirmation` | Show the estimate and ask once. Approval via `tldrgraph init --yes` persists until the current campaign is done. | -| `needs_enrichment` | Open, read, and describe this batch, then continue immediately without asking again. | -| `needs_embeddings` | Enrichment finished but the required dense model/index could not be built. Fix model access and rerun init. | -| `done` | Nothing left. Use `query` / `trace` / `layers`. | - -`--json` gives you the same thing machine-readably. The sections below document the file -formats `init` reads and writes; the underlying `queue-enrichment` / `apply-enrichment` -commands remain available for scripting. - -**Copy every `id` verbatim from the request.** A constructed id matches nothing, is -dropped, and will be reported back to you — but the work is wasted. - ---- - -## The loop - -```bash -tldrgraph init --yes # 1. approve every current candidate; writes a 200-node request -# 2. read every requested source and write enrichment_response.yaml -tldrgraph init # 3. applies it and emits the next batch; approval is remembered -# 4. repeat steps 2-3 without asking until status: done -``` - -Inside an existing coding-agent session, nested-agent protection can prevent the CLI -from launching another agent. In that case **you are the enrichment agent** and must -process every batch yourself. Do not stop at `needs_enrichment`. - -`--batch 200` means process all candidates in chunks of 200. `--limit 200` means stop -after only 200 candidates and is only for an explicitly requested partial run. Never add -`--limit` or `--embeddings off` unless the user explicitly requests that behavior. - -Request and response are **separate files**. Never write your answer back into -`enrichment_request.yaml`; it is regenerated on every run and your work would be lost. +Each final artifact must use the reported source hash. Run the command that +reported the status again; TLDRGraph validates the direct artifacts against the current source +inventory and generates the visualizer. -| File | Written by | Read by | -| --- | --- | --- | -| `.tldrgraph/enrichment_request.yaml` (or `enrichment_request.json`) | `queue-enrichment` | you | -| `.tldrgraph/enrichment_response.yaml` (or `enrichment_response.json`) | **you** | `apply-enrichment` | -| `.tldrgraph/enrichment_cursor.json` | both commands | both commands | -| `.tldrgraph/enrichment_approval.json` | `init --yes` | later `init` runs | -| `.tldrgraph/pending_enrichment.json` | *(legacy)* | `apply-enrichment`, only if no response file exists | +## Evidence ---- - -## Request schema (`enrichment_request.yaml`) - -```yaml -schema: codechakra/enrichment-request@1 -generated_at: "2026-08-19T00:00:00+00:00" -response_file: .tldrgraph/enrichment_response.yaml -contract: .tldrgraph/AGENT_CONTRACT.md -progress: - total_candidates: 1873 # un-enriched, non-utility nodes - already_enriched: 12 # nodes that already carry an intent - queued_now: 200 # entries in "nodes" below - remaining_after: 1673 # still waiting after this batch is applied -nodes: - - id: backend_src_applications_applications_controller_applicationscontroller - label: ApplicationsController - layer_id: api - layer: "Layer 2: API Gateway" - file: backend/src/applications/applications.controller.ts - source_location: L31 - degree: 41 # in + out edges in the AST graph - cross_layer_degree: 17 # of those, how many cross a layer boundary - rank: 1 # 1 = highest priority in this batch - existing_intent_source: heuristic # "" when the node has no intent at all -``` - -`file` is repo-relative. `source_location` is graphify's line hint and may be `null`. -`layer_id` is the stable machine key (e.g. `cli`, `engine`, `storage`, `api`, `ui`). - -`existing_intent_source` is `"heuristic"` when the node already carries an intent written -by the offline template enricher. That text was generated from the label and layer alone -— it has not read a line of source — so the node is still a candidate and your answer -should overwrite it. Applied answers are stamped `"agent"` and are never re-queued. - ---- - -## Response schema (`enrichment_response.yaml` or `enrichment_response.json`) - -A **YAML list** (preferred) or **JSON array** of objects: +Every feature and every displayed workflow step must have at least one evidence +record: ```yaml -- id: backend_src_applications_applications_controller_applicationscontroller - intent: | - ### Pension Application Lifecycle Gateway - REST gateway for the pension application lifecycle. Authorizes DEO/AAO/AO/DAG roles, - dispatches cases to ApplicationsService and records status transitions. - input_fields: - - caseId - - transitionPayload - - remarks - - sanctionOrderNo - output_fields: - - applicationStatus - - disposition - calls: - - ApplicationsService - - JwtAuthGuard - - RolesGuard - - pension_cases -``` - -Equivalent JSON format (also accepted from `.tldrgraph/enrichment_response.json` or `.tldrgraph/pending_enrichment.json`): -```json -[ - { - "id": "backend_src_applications_applications_controller_applicationscontroller", - "intent": "### Pension Application Lifecycle Gateway\nREST gateway for the pension application lifecycle. It authorizes roles and dispatches source-backed status transitions.", - "input_fields": ["caseId", "transitionPayload", "remarks", "sanctionOrderNo"], - "output_fields": ["applicationStatus", "disposition"], - "calls": ["ApplicationsService", "JwtAuthGuard", "RolesGuard", "pension_cases"] - } -] +file: relative/path/to/source.py +symbol: verified_symbol +line: 12 +code_start: 12 +code_end: 28 ``` -| Key | Type | Meaning | -| --- | --- | --- | -| `id` | string, **required** | The node id, copied **verbatim** from the request. An id that is not in the graph is skipped silently. | -| `intent` | string (Markdown) | Markdown formatted 2-3 sentence explanation: what this symbol does, why it exists, and its source-backed behavior. Headings and list markers do not count as sentences. This is the text semantic search matches against. | -| `input_fields` | array of strings | Input parameters, arguments, request body payload attributes, query filters. | -| `output_fields` | array of strings | Return types, response models, emitted event names, or mutated state attributes. | -| `fields` | array of strings (legacy) | Supported for backwards compatibility (maps to input fields). | -| `calls` | array of strings or objects | Downstream symbols, files (`file:symbol`), or node IDs this symbol calls. Cross-layer bridges are created with 100% confidence. | -| `layer_id` | string (optional) | Explicitly reassign the architectural layer ID if the AST classification miscategorized it. | - -`input_fields`, `output_fields`, and `calls` may be omitted or empty. An object with only `id` and `intent` is -valid and useful. - ---- - -## Hard rules - -1. **Open and read the actual source file before writing an intent.** You have the repo - checked out; that is the entire reason this path exists. Read `file` (use - `source_location` to find the symbol), and read enough of its imports and callees to - describe what it really does. An intent paraphrased from the label is worse than no - intent, because it poisons search with confident-sounding noise. - -2. **Write every intent in 2-3 complete sentences.** Cover what the symbol does, why it - exists, and its source-backed behavior. Markdown headings and list markers do not count - as sentences. - -3. **Do not invent fields or calls. Omit what you cannot verify in the code.** If you - read the file and it handles three params, list three. Do not pad the list with what a - symbol of that name "usually" has. `"fields": []` is a correct, honest answer. - A wrong `calls` entry creates a real, wrong edge in the graph that later queries will - follow. - -4. **`calls` entries are resolved with 2-tier high precision.** - - **Tier 1 (Exact Match, 100% confidence):** Exact symbol names (`ApplicationsService`), - function names, node IDs, file paths (`calc.ts`), or database table names (`pension_cases`). - - **Tier 2 (Vector Fallback):** Semantic search with a calibrated 0.35 score floor. - - | Good | Bad | - | --- | --- | - | `ApplicationsService` | `the application service` | - | `calc.ts` | `some calculation helper` | - | `pension_cases` | `the database` | - | `JwtAuthGuard` | `auth stuff` | - - Prefer the exact symbol name, file name, or table/model name as it appears in the source. - -5. **Copy `id` verbatim.** Do not normalize, shorten or re-case it. - -6. **Answer only the nodes in the request.** Extra ids are ignored; missing ids just come - back in a later batch. - -7. **After full approval, never ask again for the same campaign.** Continue processing - `needs_enrichment` batches until `status: done`. Do not silently add `--limit` or - `--embeddings off`. - ---- - -## Priority order in the queue - -The queue is not arbitrary — a node that many things depend on is worth more of your -attention than a leaf. A node is a **candidate** when it sits outside `General / Utility` -and either has no intent at all, or has one that came from the offline template heuristic -(`enrichment_source: "heuristic"`, i.e. nobody read the source). Candidates are sorted by: - -1. **`cross_layer_degree` descending** — neighbours that sit in a *different* layer. - These are the seams TLDRGraph exists to describe, and they are exactly where the AST - alone is weakest. -2. **`degree` descending** — total in + out edges. Hub nodes first. -3. **node id ascending** — only to make the ordering deterministic. - -Both degrees are computed from the live graph. (The `degree` key that graphify emits is -absent, so anything reading `node["degree"]` from the raw export sees `0`; TLDRGraph -recomputes it and stamps it back into `.tldrgraph/graph.json`.) - ---- - -## Paging and progress - -`queue-enrichment` remembers what it has handed out in `.tldrgraph/enrichment_cursor.json`: - -- `applied` — ids successfully merged by `apply-enrichment`. Never re-queued. -- `queued` — ids handed out but not yet applied ("in flight"). Skipped by default. - -So running `queue-enrichment` twice in a row **advances** to the next batch instead of -repeating. Two escape hatches: - -- `--requeue` — also hand out in-flight ids again (use when a batch was abandoned). -- `--reset` — clear all progress and start again from the highest-priority node. -- `--limit 0` — no cap; queue every remaining candidate at once. - ---- - -## What `apply-enrichment` does with your answer - -For each object it can match to a node: - -1. sets `intent`, rewrites `summary` to `":