From 674aafec1bef99f35f54c37af461705a557b0721 Mon Sep 17 00:00:00 2001 From: aarroyo Date: Thu, 30 Jul 2026 20:17:51 -0500 Subject: [PATCH 1/3] =?UTF-8?q?docs(gaps):=20close=20GT-602=20=E2=80=94=20?= =?UTF-8?q?the=20wasm=20parity=20block=20was=20verified=20by=20its=20negat?= =?UTF-8?q?ive=20half?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The code half landed in 086111db; the board never caught up, so the row still read IN-PROGRESS with criteria 1 and 3 unticked while the enforcement was already in `Test mcp-server`, a required check on `main`. Credited only after re-running the negative half rather than the colour: both `policy.wasm` artifacts (`src/rulesets/opa/` and `src/sdk/cli/rulesets/opa/`) were moved off disk and `abac-rego-parity.spec.ts` re-run — 11/11 green and both artifacts recompiled with fresh timestamps. That is what separates "the block no longer self-skips" from "it passed because the file happened to be there", which is the exact distinction GT-632 was registered for. Full mcp-server suite 452/452. Guards 08 (640 gaps, 583 closure records), 09 (reconciliation regenerated to match), 04 green. Co-Authored-By: Claude Opus 5 --- .../evidence/gap-closure-evidence.json | 17 +++++++++++++++++ .../gaps/gap-reference-catalog.es.md | 4 ++-- .../gaps/gap-reference-catalog.md | 4 ++-- .../core/control-center/gaps/gap-tracking.es.md | 4 ++-- .../core/control-center/gaps/gap-tracking.md | 4 ++-- .../maturity-reconciliation.json | 6 +++--- 6 files changed, 28 insertions(+), 11 deletions(-) diff --git a/reference/core/control-center/evidence/gap-closure-evidence.json b/reference/core/control-center/evidence/gap-closure-evidence.json index 13a4abd7..f5765e52 100644 --- a/reference/core/control-center/evidence/gap-closure-evidence.json +++ b/reference/core/control-center/evidence/gap-closure-evidence.json @@ -8428,6 +8428,23 @@ ], "dependencyDisposition": "none" }, + { + "id": "GT-602", + "closedAt": "2026-07-30", + "closureCommit": "086111db", + "evidence": [ + "src/packages/mcp-server/src/mcp/abac-rego-parity.spec.ts", + ".harness/scripts/compile-opa-wasm.mjs", + ".harness/scripts/generate-abac-tool-sets.mjs", + "src/rulesets/opa/abac-mcp-tool-access.rego", + ".github/workflows/ci-cd.yml" + ], + "validationCommands": [ + "npm --workspace @beyondnet/evolith-mcp test", + "node .harness/scripts/ci/08-validate-tracking.mjs" + ], + "dependencyDisposition": "none" + }, { "id": "GT-607", "evidence": [ diff --git a/reference/core/control-center/gaps/gap-reference-catalog.es.md b/reference/core/control-center/gaps/gap-reference-catalog.es.md index bf8c79ce..42d98d38 100644 --- a/reference/core/control-center/gaps/gap-reference-catalog.es.md +++ b/reference/core/control-center/gaps/gap-reference-catalog.es.md @@ -7252,9 +7252,9 @@ Serie histórica de gaps registrada en el antiguo `gap-analysis-core.es.md`, pre - **Componente:** `Evolith MCP` · **Criticidad:** P0 · **Complejidad:** M - **Procedencia:** Evaluación del código componente a componente realizada el 2026-07-26 en el repositorio compañero `why-architecture` (`docs/evolith-diagnostico-es.md`), verificada contra el código de este repositorio antes de registrarse. - **Criterios de aceptación:** - - [ ] Un test evalúa el `policy.wasm` compilado sobre todos los nombres registrados y exige ALLOW para un `architect` en `production`. + - [x] Un test evalúa el `policy.wasm` compilado sobre todos los nombres registrados y exige ALLOW para un `architect` en `production`. - [x] Los conjuntos de herramientas del rego se generan desde el registro en vez de mantenerse a mano. - - [ ] CI falla cuando una herramienta existe en el registro TypeScript y no en la política compilada. + - [x] CI falla cuando una herramienta existe en el registro TypeScript y no en la política compilada. #### GT-603 diff --git a/reference/core/control-center/gaps/gap-reference-catalog.md b/reference/core/control-center/gaps/gap-reference-catalog.md index 24e26320..c6a88377 100644 --- a/reference/core/control-center/gaps/gap-reference-catalog.md +++ b/reference/core/control-center/gaps/gap-reference-catalog.md @@ -7347,9 +7347,9 @@ Historical gap series tracked in the former `gap-analysis-core.md`, preserved fo - **Component:** `Evolith MCP` · **Criticality:** P0 · **Complexity:** M - **Provenance:** Component-by-component source assessment conducted 2026-07-26 in the companion `why-architecture` repository (`docs/evolith-assessment-en.md`), verified against this repository's code before registration. - **Acceptance criteria:** - - [ ] A test evaluates the compiled `policy.wasm` over all registered tool names and asserts ALLOW for an `architect` in `production`. + - [x] A test evaluates the compiled `policy.wasm` over all registered tool names and asserts ALLOW for an `architect` in `production`. - [x] The rego tool sets are generated from the tool registry rather than hand-maintained. - - [ ] CI fails when a tool exists in the TypeScript registry and not in the compiled policy. + - [x] CI fails when a tool exists in the TypeScript registry and not in the compiled policy. #### GT-603 diff --git a/reference/core/control-center/gaps/gap-tracking.es.md b/reference/core/control-center/gaps/gap-tracking.es.md index e2e16882..e52dcc3a 100644 --- a/reference/core/control-center/gaps/gap-tracking.es.md +++ b/reference/core/control-center/gaps/gap-tracking.es.md @@ -46,7 +46,7 @@ Este tablero es la única fuente de verdad para deuda técnica, gaps, oportunida | [`GT-620`](./gap-reference-catalog.es.md#gt-620) | `reference/core/interfaces/using-the-cli.md` — el slot en inglés — empieza con `# Cómo usar la CLI de Evolith` y mide 817 palabras funcionales españolas frente a 151 inglesas. **Verificado aquí contra el código.** El proyecto está en el umbral del open source sin punto de entrada en inglés para su superficie principal. Es además una instancia viva de un punto ciego que nombró la auditoría de madurez: `04-check-bilingual-parity` compara el CONTEO de cabeceras `##`/`###`, así que un fichero sin traducir en el slot equivocado pasa en verde — éste pasa. Fix: escribir la guía inglesa real, y extender el gate de paridad con una heurística barata de idioma para que la clase de defecto no pueda repetirse en silencio. Origen: hallazgo 3.6 del diagnóstico de producto (https://github.com/beyondnetcode/why-architecture/blob/main/docs/evolith-diagnostico-es.md). **COMPLETADO (ola 2026-07-28).** `using-the-cli.md` —el slot INGLÉS— es ahora inglés: medidas 0 palabras funcionales españolas frente a 1.222 inglesas, donde antes eran 817 frente a 151 y empezaba con `# Cómo usar la CLI de Evolith`. Los comandos documentados se comprobaron contra `--help` del CLI construido, no se transcribieron. Ambos ficheros mantienen 41 cabeceras, así que el gate de paridad sigue pasando — que es justo por lo que esto era invisible: ese gate compara CONTEOS de cabeceras, y cazar un fichero en el idioma equivocado exige la heurística de idioma entregada en follow_ups. **Reabierto a EN-PROGRESO en la reconciliación:** el defecto está arreglado y medido, pero el único criterio de aceptación de la fila pide un test que falle sin el arreglo, y no existe — `04-check-bilingual-parity.mjs` sigue con cero heurística de idioma, así que el mismo defecto puede repetirse en silencio. Marcarlo sería la sobre-afirmación que este tablero lleva pillando. **CERRADO el 2026-07-28 — y la heurística reveló que la clase era mucho mayor que esta fila.** `04-check-bilingual-parity` compara ahora el IDIOMA de cada par, no solo el número de sus títulos, vía `lib/language-heuristic.mjs`: palabras funcionales inequívocas en un idioma, contadas solo sobre PROSA (se retiran bloques de código, código en línea y destinos de enlace, para que banderas e identificadores no sesguen el recuento). Se abstiene ante una muestra de menos de 40 marcadores o una mayoría por debajo del 65% — una acusación falsa sería peor que el punto ciego, porque acabaría con la comprobación apagada. **La regresión se demuestra, no se afirma:** volver a poner el texto español en el slot inglés pone la puerta en rojo con `710 palabras funcionales en español frente a 71 en inglés`, y restaurar la guía inglesa la pone verde; 7 autotests cubren los dos veredictos, las dos abstenciones y la trampa del código. **Al activarla aparecieron 19 documentos más en el slot equivocado** — 8 slots ingleses escritos en español (entre ellos `using-the-mcp.md` con 956 palabras españolas frente a 107 inglesas, y `using-the-rest-api.md` con 1231 frente a 35) y 11 slots españoles escritos en inglés. Quedan en LÍNEA BASE POR NOMBRE, no por conteo: un ratchet numérico permitiría cambiar un fichero mal etiquetado por otro sin que el número se moviera. Registrado como `GT-628`. Este fichero está deliberadamente fuera de esa línea base, así que falla si alguna vez regresa. | | | `Documentation` | Cross | P1 | S | `COMPLETADO` | | [`GT-621`](./gap-reference-catalog.es.md#gt-621) | Existen 17 puertos y 49 adaptadores para una sola pasada de ejecución; el camino caliente depende de 9 puertos requeridos y está bien dimensionado, mientras que los bordes fríos están sobre-construidos — dos adaptadores de interacción sin llamadores, un provider completo y sin conectar. **No verificado aquí**; se registra tal como lo reporta el diagnóstico y debe confirmarse antes de dimensionar el trabajo. El error no es construirlos: es contarlos como capacidad entregada en los documentos de visión. Fix: declarar en esos documentos qué puertos están en el camino caliente y cuáles son especulativos, para que un número de inventario deje de leerse como una afirmación de capacidad. Origen: hallazgo 5.7 del diagnóstico de producto (https://github.com/beyondnetcode/why-architecture/blob/main/docs/evolith-diagnostico-es.md). **COMPLETADO (ola 2026-07-28).** Los documentos de visión ahora declaran qué puertos están en el camino caliente (9 requeridos) y cuáles son especulativos, de modo que un número de inventario deja de leerse como afirmación de capacidad entregada. Los puertos y adaptadores no se tocaron: el defecto nunca fue construirlos, fue contarlos. **Reabierto a EN-PROGRESO en la reconciliación:** el defecto está arreglado y medido, pero el único criterio de aceptación de la fila pide un test que falle sin el arreglo, y no existe — `04-check-bilingual-parity.mjs` sigue con cero heurística de idioma, así que el mismo defecto puede repetirse en silencio. Marcarlo sería la sobre-afirmación que este tablero lleva pillando. **CERRADO el 2026-07-28, con las cifras de la propia fila corregidas.** La evidencia decía 17 puertos / 49 adaptadores y un camino caliente de 9; medido hoy son **19 interfaces de puerto, 53 ficheros de adaptador y 11 en el camino caliente** (7 obligatorios, 4 opcionales). El "dos adaptadores de interacción sin llamadores" del diagnóstico también se queda corto: `InteractionAdapterPort` tiene SEIS implementaciones y el runtime no llama a ninguna, junto a `ISchedulerPort`, `ICommunicationGatewayPort`, `IAssistantTransport`, `IQualitySignalProvider` e `IStructuralReviewer` — ocho costuras declaradas que la ejecución nunca alcanza. **La peor instancia no era prosa sino un diagrama:** `master-view.svg` publicaba "AgentRuntimeService — 12 hexagonal ports · 30 adapters", cifras que no coincidían con nada y que se leen como capacidad entregada. Corregido a "11 of 19 ports on the hot path · 53 adapters". El guard `45-validate-port-inventory-honesty.mjs` DERIVA el camino caliente de `AgentRuntimeDeps` — esa interfaz ES lo que la ejecución puede alcanzar — así que no puede pudrirse como una cifra escrita a mano, y falla ante cualquier documento rastreado que publique un conteo sin la distinción o cuya anotación discrepe del código. 9 autotests, cuatro de ellos anti-vacuos: cero puertos, una forma de deps no parseable (el camino caliente debe derivarse, nunca asumirse), cero adaptadores, y un corpus de escaneo que se movió. Cableado en ci-cd.yml y registrado en `guard-classification.mjs`; 37/37 guards clasificados, 34/34 vistos fallar. | | | `agent-runtime` | Cross | P3 | S | `COMPLETADO` | | [`GT-601`](./gap-reference-catalog.es.md#gt-601) | **`canonical-result.mapper.ts` escribe tres campos de trazabilidad como arrays vacíos incondicionales y fija el motor a mano — y dos consumidores los leen.** `:130` `rulesExecuted: []`, `:134` `missingEvidence: []`, `:104`/`:133` `risks: []` son literales en cada evaluación real, y `:72` fija `engine: 'opa'` en toda referencia de `policiesApplied` corriera `NativeEvaluator` u `OpaEvaluator`. `sarif-exporter.ts:256` y `drift-gate.ts:203` derivan `evaluatedRules` de `rulesExecuted`. Consecuencia: **todo SARIF y todo manifiesto de evidencia del drift-gate que se emite hoy dice "0 reglas evaluadas"**, y en un producto cuya paridad dual es argumento de venta el único artefacto que la probaría atribuye mal la mitad de sus ejecuciones. El contrato de evidencia EVD-01..04 queda satisfecho en la forma y vacío en el fondo: el grafo de auditoría acumulado se escribe en blanco desde el origen. **+ola 2026-07-28 — 3 de 4 criterios.** Los cuatro literales desaparecieron: `rulesExecuted` es el conjunto deduplicado (regla, motor) que el pipeline ejecutó realmente, `missingEvidence` nombra artefactos declarados y no presentados, `risks` se construyen con el vocabulario de GT-569 manteniendo `skipped` (medio) distinto de `errored` (alto), y el motor se resuelve por regla en vez de hardcodearse a `'opa'`. SARIF emite ahora el CATÁLOGO de reglas de la corrida, así que una evaluación limpia deja de leerse como "0 rules evaluated". 26 tests nuevos. **NO cerrado:** `risks` sigue vacío en una corrida REAL hasta que `SatelliteEvaluationPipeline` adjunte `coverage` a su veredicto — ese productor está en `application/services/`, fuera del área del agente — y el motor se DERIVA de la ruta de la regla en vez de DECLARARSE, porque `RuleEvaluation` no tiene campo `engine`. El mapper ya prefiere un sello explícito en cuanto exista, así que la atribución es inferencia, no testimonio. **CERRADO el 2026-07-28 por verificación, no por código nuevo.** El arreglo había aterrizado en una ola anterior y nadie comprobó si satisfacía la fila. Los cuatro criterios se cumplen: `mapPipelineVerdict` puebla `rulesExecuted` desde `executedByKey` y un test comprueba el conteo contra una fixture con número de reglas conocido; `resolveRuleEngine` atribuye cada regla al evaluador que la ejecutó, con tests sobre la ruta nativa Y la de OPA más un override por regla; el exportador SARIF deriva `evaluatedRules` de `rulesExecuted` (deduplicado y ordenado) y comprueba que la lista no es vacía; y `risks` y `missingEvidence` se pueblan con el vocabulario de GT-569 — una regla que el motor NO pudo evaluar se convierte en riesgo, distinta de una que lanzó excepción, y lo que no se pudo comprobar nunca se cuenta como comprobado. **Verificado por mutación, no por lectura:** devolver `rulesExecuted` al array vacío incondicional pone 10 tests en rojo entre el mapper y el exportador SARIF. 40/40 en verde. | | | `Evolith Core` | Cross | P0 | S | `COMPLETADO` | -| [`GT-602`](./gap-reference-catalog.es.md#gt-602) | **15 de las 50 herramientas MCP están denegadas en producción por la política compilada que el dispatcher carga de verdad.** Establecido cargando `src/sdk/cli/rulesets/opa/policy.wasm` —el artefacto que se usa en dispatch, no el fuente rego— y evaluándolo para un `architect` en `production`: `evolith-adr-list`, `evolith-adr-get/create/update/matrix`, `evolith-pattern-list/get/list-by-topology`, `evolith-scaffold`, `evolith-docs-scaffold`, `evolith-init-batch`, `evolith-sdlc-generate`, `evolith-fixtures` y `evolith-upgrade-plan/apply` devuelven `ABAC-03` + `ABAC-01`. Como `mcp-tool-dispatch.ts:146` exige que permitan **ambos** motores, las quince quedan FORBIDDEN. El propio comentario del rego predijo este fallo; `abac-classification-coverage.spec.ts` guarda TS↔registro pero **nada guarda rego↔TS**. Consecuencia: un agente que pide a Evolith sus propios ADRs es rechazado, en silencio y por completo. Ola 2: el P0 está corregido — 15 pares herramienta/rol permitidos, cada uno argumentado por separado, y el ensanchamiento es estructuralmente nulo porque el dispatch exige nativo Y opa y los 15 ya estaban permitidos por el ABAC nativo. viewer/production sigue denegado en las 20 herramientas de escritura. Un spec de paridad bidireccional rego↔TypeScript lo protege. **AC1 NO acreditado:** el test de wasm SE AUTO-SALTA en CI porque el job `Test mcp-server` no ejecuta `build:policy`, así que CI verifica paridad contra el FUENTE rego, no contra el bundle compilado — un test que se salta justo donde se exige es el patrón de pase vacuo que este tablero persigue. AC2 intacto. **Criterio 2 intentado el 2026-07-28 y NO enviado, porque la versión ingenua es una regresión en producción.** Generar los tres conjuntos rego desde `TOOL_CLASSIFICATION` parece obvio y está mal: el mapa tiene 50 herramientas, el rego enumera 62, y la diferencia de 9 (`evolith-ping`, `evolith-echo`, `evolith-read-file`, `evolith-list-dir`, `evolith-read-gap-tracking`, `evolith-gate-status`, `evolith-write-file`, `evolith-replace-file`, `evolith-run-command`) la clasifica en runtime el FALLBACK de `classifyTool` — los conjuntos heredados `READ_TOOLS`/`WRITE_TOOLS`/`DEPLOY_TOOLS` más heurísticas de nombre (`includes()` over read/write/run/deploy and friends), no el mapa. Un generador alimentado solo por el mapa habría BORRADO esas nueve de la política, dejándolas sin clasificar y por tanto denegadas por ABAC-03 en producción — el mismo modo de fallo por el que se registró este gap, reintroducido por el arreglo. Detectado comparando la pertenencia antes y después en vez de dar por cosmético el diff; no se commiteó nada. **La fuente correcta es la lista de herramientas registradas clasificada a través del propio `classifyTool`, no el mapa**, lo que exige que el paquete mcp-server exponga primero esa proyección (`dist` construido, o un artefacto volcado) para que el generador consuma una sola respuesta en lugar de reimplementar el fallback y crear una tercera copia de la lógica. `abac-rego-parity.spec.ts` pasa hoy justamente porque compara contra `classifyTool`, fallback incluido — que es la razón de que nunca señalara el hueco de 12 herramientas entre el mapa y el rego. **Criterio 2 CERRADO el 2026-07-29, en el segundo diseño.** Los conjuntos rego se derivan ahora de `AbacEvaluator.toolProjection()` — la unión enumerable de todos los conjuntos que declara el código, con cada nombre clasificado LLAMANDO a `classifyTool` en vez de releyendo sus mapas. El generador ejecuta esa proyección con `ts-node` sobre el fuente TypeScript, así que no necesita build ni crea una tercera copia de la lógica del fallback. Resultado: **64 herramientas, CERO pérdidas** frente a la política mantenida a mano, más dos nombres (`read-tool`, `write-tool`) que el runtime ya clasificaba y el rego omitía. Es un ensanchamiento de dos nombres de fixture que el lado TypeScript ya permitía, así que la intersección nativo∧opa no cambia. El `--check` está cableado en el job `Test mcp-server` y el artefacto queda registrado como PRIMER eslabón de la cadena de artefactos derivados de GT-630, que ahora corre 3 eslabones y sigue en punto fijo. mcp-server 434/434 con la política generada; tests de política OPA 224/224. El primer diseño — alimentarse solo de `TOOL_CLASSIFICATION` — se descartó porque habría borrado doce herramientas y las habría denegado en producción; ese razonamiento se conserva arriba en vez de borrarlo con el arreglo que vino después. | | | `Evolith MCP` | Construction | P0 | M | `EN-PROGRESO` | +| [`GT-602`](./gap-reference-catalog.es.md#gt-602) | **15 de las 50 herramientas MCP están denegadas en producción por la política compilada que el dispatcher carga de verdad.** Establecido cargando `src/sdk/cli/rulesets/opa/policy.wasm` —el artefacto que se usa en dispatch, no el fuente rego— y evaluándolo para un `architect` en `production`: `evolith-adr-list`, `evolith-adr-get/create/update/matrix`, `evolith-pattern-list/get/list-by-topology`, `evolith-scaffold`, `evolith-docs-scaffold`, `evolith-init-batch`, `evolith-sdlc-generate`, `evolith-fixtures` y `evolith-upgrade-plan/apply` devuelven `ABAC-03` + `ABAC-01`. Como `mcp-tool-dispatch.ts:146` exige que permitan **ambos** motores, las quince quedan FORBIDDEN. El propio comentario del rego predijo este fallo; `abac-classification-coverage.spec.ts` guarda TS↔registro pero **nada guarda rego↔TS**. Consecuencia: un agente que pide a Evolith sus propios ADRs es rechazado, en silencio y por completo. Ola 2: el P0 está corregido — 15 pares herramienta/rol permitidos, cada uno argumentado por separado, y el ensanchamiento es estructuralmente nulo porque el dispatch exige nativo Y opa y los 15 ya estaban permitidos por el ABAC nativo. viewer/production sigue denegado en las 20 herramientas de escritura. Un spec de paridad bidireccional rego↔TypeScript lo protege. **AC1 NO acreditado:** el test de wasm SE AUTO-SALTA en CI porque el job `Test mcp-server` no ejecuta `build:policy`, así que CI verifica paridad contra el FUENTE rego, no contra el bundle compilado — un test que se salta justo donde se exige es el patrón de pase vacuo que este tablero persigue. AC2 intacto. **Criterio 2 intentado el 2026-07-28 y NO enviado, porque la versión ingenua es una regresión en producción.** Generar los tres conjuntos rego desde `TOOL_CLASSIFICATION` parece obvio y está mal: el mapa tiene 50 herramientas, el rego enumera 62, y la diferencia de 9 (`evolith-ping`, `evolith-echo`, `evolith-read-file`, `evolith-list-dir`, `evolith-read-gap-tracking`, `evolith-gate-status`, `evolith-write-file`, `evolith-replace-file`, `evolith-run-command`) la clasifica en runtime el FALLBACK de `classifyTool` — los conjuntos heredados `READ_TOOLS`/`WRITE_TOOLS`/`DEPLOY_TOOLS` más heurísticas de nombre (`includes()` over read/write/run/deploy and friends), no el mapa. Un generador alimentado solo por el mapa habría BORRADO esas nueve de la política, dejándolas sin clasificar y por tanto denegadas por ABAC-03 en producción — el mismo modo de fallo por el que se registró este gap, reintroducido por el arreglo. Detectado comparando la pertenencia antes y después en vez de dar por cosmético el diff; no se commiteó nada. **La fuente correcta es la lista de herramientas registradas clasificada a través del propio `classifyTool`, no el mapa**, lo que exige que el paquete mcp-server exponga primero esa proyección (`dist` construido, o un artefacto volcado) para que el generador consuma una sola respuesta en lugar de reimplementar el fallback y crear una tercera copia de la lógica. `abac-rego-parity.spec.ts` pasa hoy justamente porque compara contra `classifyTool`, fallback incluido — que es la razón de que nunca señalara el hueco de 12 herramientas entre el mapa y el rego. **Criterio 2 CERRADO el 2026-07-29, en el segundo diseño.** Los conjuntos rego se derivan ahora de `AbacEvaluator.toolProjection()` — la unión enumerable de todos los conjuntos que declara el código, con cada nombre clasificado LLAMANDO a `classifyTool` en vez de releyendo sus mapas. El generador ejecuta esa proyección con `ts-node` sobre el fuente TypeScript, así que no necesita build ni crea una tercera copia de la lógica del fallback. Resultado: **64 herramientas, CERO pérdidas** frente a la política mantenida a mano, más dos nombres (`read-tool`, `write-tool`) que el runtime ya clasificaba y el rego omitía. Es un ensanchamiento de dos nombres de fixture que el lado TypeScript ya permitía, así que la intersección nativo∧opa no cambia. El `--check` está cableado en el job `Test mcp-server` y el artefacto queda registrado como PRIMER eslabón de la cadena de artefactos derivados de GT-630, que ahora corre 3 eslabones y sigue en punto fijo. mcp-server 434/434 con la política generada; tests de política OPA 224/224. El primer diseño — alimentarse solo de `TOOL_CLASSIFICATION` — se descartó porque habría borrado doce herramientas y las habría denegado en producción; ese razonamiento se conserva arriba en vez de borrarlo con el arreglo que vino después. **Criterios 1 y 3 CERRADOS el 2026-07-30 — el bloque de wasm ya no se auto-salta.** El bloque de bundle compilado de `abac-rego-parity.spec.ts` estaba gobernado por `EVOLITH_REQUIRE_OPA_WASM`: corría en un job de CI y hacía `it.skip` en todos los demás porque `policy.wasm` es salida de build gitignorada — exactamente la forma que produjo [`GT-632`](./gap-reference-catalog.es.md#gt-632). Ahora el bloque COMPILA la política él mismo cuando falta o es más vieja que cualquier `.rego` del que se construye (vía `.harness/scripts/compile-opa-wasm.mjs`, el mismo script de `npm run build:policy`, escribiendo las mismas rutas que carga el runtime), y si esa compilación no puede ocurrir la suite falla con el diagnóstico del propio compilador: **no existe entorno en el que pase sin haber evaluado un wasm**, y la obsolescencia se trata como ausencia porque un bundle rancio ES el defecto. Los denominadores se asertan en vez de asumirse — al menos las 50 herramientas registradas medidas hoy, `evaluated.length === registered.length` y `allowed.length === registered.length` — y el control negativo ABAC-03 corre PRIMERO, porque un `evaluate()` que devolviera una forma inesperada haría vacuas todas las aserciones de ALLOW. `compile-opa-wasm.mjs` instala por rename en vez de `copyFileSync`, porque que el test compile bajo demanda hace alcanzable la lectura-durante-escritura entre workers de jest hermanos y una lectura parcial revienta en `loadPolicy`, que el evaluador reporta como OPA_ERROR — o sea, una denegación. **Verificado el 2026-07-30 por la mitad negativa, no por el color:** apartando del disco AMBOS `policy.wasm` (`src/rulesets/opa/` y `src/sdk/cli/rulesets/opa/`) y volviendo a correr la spec — 11/11 en verde y los dos artefactos recompilados con marca de tiempo nueva, que es lo que distingue «no se salta» de «pasó porque el fichero estaba ahí». | | | `Evolith MCP` | Construction | P0 | M | `COMPLETADO` | | [`GT-603`](./gap-reference-catalog.es.md#gt-603) | **El ledger de turnos de agente está completo, con tests unitarios y ausente de la inyección de dependencias, y la columna de actor no se puede tipar retroactivamente.** `AgentExecutionService.cs` valida el alcance y luego **audita antes de ejecutar y aborta el turno si falla la escritura de auditoría**; `AgentTurnAuditor.cs` registra alcances concedidos frente a usados y guarda la longitud del prompt, no su texto. `IAgentExecutionPort` aparece en **cero registros de DI y cero endpoints**, mientras `AssistantEndpoints.cs` pasa de largo por `AgentRuntimeGateway` sin persistir nada. Aparte, `AuditEntryProps.cs:11` declara `public Guid ActorId` sin `actor_type`, `agent_id`, `model_id` ni `session_id`. Como `audit_entries` es append-only por trigger de base de datos (migración `20260719202323`), **las filas escritas antes de que exista el discriminador no se pueden corregir jamás**. Complementa a [`GT-586`](./gap-reference-catalog.es.md#gt-586), que cubre el lado Core; éste es el lado de persistencia del Tracker. | | | `Evolith Tracker` | Cross | P0 | M | `PENDIENTE` | | [`GT-604`](./gap-reference-catalog.es.md#gt-604) | **Ninguna superficie escribe evidencia en el Tracker: el camino de escritura es unidireccional y apunta hacia dentro.** Un grep sobre `src/sdk/cli`, `src/packages/mcp-server` y `src/packages/core-domain` no devuelve ninguna URL base del Tracker ni cliente de ingesta; la única URL del Tracker en este repositorio es `AGENT_RUNTIME_APPROVAL_TRACKER_URL`, y los únicos escritores de `core_evaluation_transactions` son endpoints iniciados por el propio Tracker. Consecuencia: cada `evolith validate`, cada veto de `enforce edit`, cada `tools/call` de MCP y cada ejecución del drift-gate en CI se evapora al terminar el proceso. La estrategia se apoya en evidencia acumulada mientras las superficies que la producen no tienen dónde depositarla. Es un defecto de composición: ninguna revisión por componente puede verlo, porque cada componente es internamente coherente. **EN-PROGRESO 2026-07-29 — solo el criterio 1.** El contrato de ingesta aterrizó en `src/packages/contracts/src/ingest/evaluation-ingest.ts` con `correlationId` OBLIGATORIO y AMBOS responsables representados por separado: `requestedBy.actorId` es quien pidió la evaluación y `violations[].accountableOwner` es quien debe corregir. El criterio 2 no se puede completar desde aquí — no existe endpoint de ingesta en el Tracker al que llamar. El criterio 3 es 0% construible en este repositorio: RoboSoft vive en `evolith_tracker`. Se escribió un traspaso bilingüe en `reference/core/control-center/opportunities/tracker-handover-gt604.md` y su equivalente `.es.md`, con 22 encabezados cada uno. **Sin confirmar, y así queda registrado: la dependencia declarada de esta fila respecto a GT-603 parece equivocada.** GT-603 migra `audit_entries` mientras esta fila nombra `core_evaluation_transactions`, y la atribución del lado Core ya se entregó bajo GT-586 — pero el esquema del Tracker no está en este repositorio, así que es una hipótesis a confirmar por el responsable del Tracker, no un hallazgo. | | | `Evolith Suite` | Cross | P0 | L | `EN-PROGRESO` | | [`GT-605`](./gap-reference-catalog.es.md#gt-605) | **Existen dos grafos de evidencia y a cada uno le falta la mitad del otro.** `evidence-graph.ts` define aristas tipadas (`requires` / `validates` / `blocks`) y tiene **cero consumidores fuera de su propio fichero de test**. El Tracker persiste `EvidenceRecordProps.References` como `List` en una columna jsonb cuyo único consumidor no-test es un `Contains()` lineal para deduplicar identificadores externos: sin tabla de aristas, sin tipo, sin búsqueda inversa, sin consulta por profundidad. El modelo tipado vive donde nada persiste; el persistido vive donde nada está tipado. Consecuencia: "qué ADR se movió por qué decisión de puerta por qué turno de agente" es inrespondible, y esa traversal es la mitad fuerte del foso declarado. Ola 2: el modelo tipado único se publica en `@beyondnet/evolith-contracts` con las cuatro cosas que un tipo compartido por sí solo deja ambiguas — vocabulario (los tres del Core literales más `caused_by`, porque ninguno de requires/validates/blocks expresa "el ADR se movió A CAUSA DE la decisión"), dirección fijada en prosa, una codificación canónica de nodo que hace migrable mecánicamente la lista jsonb del Tracker, y un recorrido acotado en profundidad contra el que probar su CTE recursivo. **Pendiente:** tres criterios son del lado Tracker y fuera de este repositorio; aquí, core-domain sigue declarando su propio `EvidenceEdge`, así que hay dos declaraciones de tipo para un solo modelo. | | | `Evolith Suite` | Cross | P1 | M | `EN-PROGRESO` | @@ -655,7 +655,7 @@ Este tablero es la única fuente de verdad para deuda técnica, gaps, oportunida | [`GT-246`](./gap-reference-catalog.es.md#gt-246) | Implementar experimentos Chaos Mesh/Litmus | | | `QA` | Cross | P3 | L | `COMPLETADO` | -**Progreso:** 600 / 640 completados · 15 en progreso · 21 pendientes · 4 diferidos +**Progreso:** 601 / 640 completados · 14 en progreso · 21 pendientes · 4 diferidos **Oleada 2026-06-23 (auditoría profunda de Winston III):** Añadidos 14 gaps nuevos `GT-212`…`GT-225` del Winston Audit Playbook que cubren: higiene de estado ADR (GT-212), metadata + presupuestos operativos + corpus de guías por topología (GT-213, GT-217, GT-219), observabilidad + OpenAPI en controladores REST (GT-214, GT-215), paridad de input-schemas OPA + densidad de tests por topología (GT-216, GT-222), plantillas de rollback + on-call de Fase 05 (GT-218), cobertura de ramas CLI + paridad de envelope --format + limpieza de skip-list (GT-220, GT-224, GT-225), audit logging HTTP de MCP (GT-221), y tests e2e de paridad cross-surface (GT-223). diff --git a/reference/core/control-center/gaps/gap-tracking.md b/reference/core/control-center/gaps/gap-tracking.md index eef5cf68..f77453e2 100644 --- a/reference/core/control-center/gaps/gap-tracking.md +++ b/reference/core/control-center/gaps/gap-tracking.md @@ -46,7 +46,7 @@ This board is the single source of truth for technical debt, gaps, opportunities | [`GT-620`](./gap-reference-catalog.md#gt-620) | `reference/core/interfaces/using-the-cli.md` — the English slot — opens with `# Cómo usar la CLI de Evolith` and measures 817 Spanish function words against 151 English ones. **Verificado aquí contra el código.** The project is at the cusp of open source with no English entry point for its main surface. It is also a live instance of a blind spot the maturity audit named: `04-check-bilingual-parity` compares the COUNT of `##`/`###` headings, so an untranslated file in the wrong slot passes green — this file does. Fix: write the real English guide, and extend the parity gate with a cheap language heuristic so the class of defect cannot recur silently. Origin: finding 3.6 of the product diagnostic (https://github.com/beyondnetcode/why-architecture/blob/main/docs/evolith-diagnostico-es.md). **DONE (wave 2026-07-28).** `using-the-cli.md` — the ENGLISH slot — is now English: measured 0 Spanish function words against 1,222 English, where before it was 817 against 151 and opened with `# Cómo usar la CLI de Evolith`. The commands documented were checked against `--help` on the built CLI rather than transcribed. Both files keep 41 headings, so the parity gate still passes — which is exactly why this was invisible: that gate compares heading COUNTS, and catching a wrong-language file needs the language heuristic handed over in follow_ups. **Re-opened to IN-PROGRESS on reconciliation:** the defect is fixed and measured, but the row's only acceptance criterion asks for a test that fails without the fix, and there is none — `04-check-bilingual-parity.mjs` still has zero language heuristic, so the same defect can recur silently. Ticking it would be the over-claim this board keeps catching. **CLOSED 2026-07-28 — and the heuristic found the class was far bigger than this row.** `04-check-bilingual-parity` now compares the LANGUAGE of each pair, not just the count of its headings, via `lib/language-heuristic.mjs`: function words unambiguous in one language, counted over PROSE only (fences, inline code and link targets stripped, so flags and identifiers cannot skew the tally). It declines to judge a sample under 40 markers or a majority under 65% — a false accusation would be worse than the blind spot, because it would get the check switched off. **The regression is demonstrated, not asserted:** restoring the Spanish text into the English slot turns the gate red with `710 Spanish function words vs 71 English`, and restoring the English guide turns it green; 7 self-tests cover both verdicts, both refusals to judge, and the code-stripping trap. **Switching it on exposed 19 more documents in the wrong slot** — 8 English slots written in Spanish (including `using-the-mcp.md` at 956 Spanish words vs 107 English, and `using-the-rest-api.md` at 1231 vs 35) and 11 Spanish slots written in English. They are BASELINED BY NAME, not by count: a numeric ratchet would let one mislabelled file be swapped for another without the number moving. Registered as `GT-628`. This file is deliberately absent from that baseline, so it fails if it ever regresses. | | | `Documentation` | Cross | P1 | S | `DONE` | | [`GT-621`](./gap-reference-catalog.md#gt-621) | 17 ports and 49 adapters exist for a single execution pass; the hot path depends on 9 required ports and is well sized, while the cold edges are over-built — two interaction adapters with no callers, a provider complete and unconnected. **No verificado aquí**; se registra tal como lo reporta el diagnóstico y debe confirmarse antes de dimensionar el trabajo. The error is not building them: it is counting them as delivered capability in the vision documents. Fix: state in those documents which ports are on the hot path and which are speculative, so an inventory number stops reading as a capability claim. Origin: finding 5.7 of the product diagnostic (https://github.com/beyondnetcode/why-architecture/blob/main/docs/evolith-diagnostico-es.md). **DONE (wave 2026-07-28).** The vision documents now state which ports are on the hot path (9 required) and which are speculative, so an inventory number stops reading as a claim of delivered capability. The ports and adapters themselves were not touched: the defect was never building them, it was counting them. **Re-opened to IN-PROGRESS on reconciliation:** the defect is fixed and measured, but the row's only acceptance criterion asks for a test that fails without the fix, and there is none — `04-check-bilingual-parity.mjs` still has zero language heuristic, so the same defect can recur silently. Ticking it would be the over-claim this board keeps catching. **CLOSED 2026-07-28, with the row's own numbers corrected.** The evidence said 17 ports / 49 adapters and a hot path of 9; measured today it is **19 port interfaces, 53 adapter files and 11 on the hot path** (7 required, 4 optional). The diagnostic's "two interaction adapters with no callers" is also understated: `InteractionAdapterPort` has SIX implementations and the runtime calls none of them, alongside `ISchedulerPort`, `ICommunicationGatewayPort`, `IAssistantTransport`, `IQualitySignalProvider` and `IStructuralReviewer` — eight declared seams the execution pass never reaches. **The worst instance was not prose but a diagram:** `master-view.svg` published "AgentRuntimeService — 12 hexagonal ports · 30 adapters", numbers that matched nothing and read as delivered capability. Fixed to "11 of 19 ports on the hot path · 53 adapters". The guard `45-validate-port-inventory-honesty.mjs` DERIVES the hot path from `AgentRuntimeDeps` — that interface IS what the execution pass can reach — so it cannot rot the way a hand-typed count did, and it fails any tracked document that publishes a count without the split or whose annotation disagrees with the code. 9 self-tests, four of them anti-vacuous: zero ports, an unparseable deps shape (the hot path must be derived, never assumed), zero adapters, and a scan corpus that moved. Wired into ci-cd.yml and registered in `guard-classification.mjs`; 37/37 guards classified, 34/34 observed failing. | | | `agent-runtime` | Cross | P3 | S | `DONE` | | [`GT-601`](./gap-reference-catalog.md#gt-601) | **`canonical-result.mapper.ts` writes three traceability fields as unconditional empty arrays and hardcodes the engine — and two consumers read them.** `:130` `rulesExecuted: []`, `:134` `missingEvidence: []`, `:104`/`:133` `risks: []` are literals on every real evaluation, and `:72` sets `engine: 'opa'` on every `policiesApplied` ref regardless of whether `NativeEvaluator` or `OpaEvaluator` ran. `sarif-exporter.ts:256` and `drift-gate.ts:203` both derive `evaluatedRules` from `rulesExecuted`. Consequence: every SARIF log and every PR drift-gate evidence manifest emitted today states **"0 rules evaluated"**, and in a product whose dual-engine parity is a selling point the one artifact that would prove it misattributes half its runs. The EVD-01..04 evidence contract is satisfied structurally and empty in substance, so the accumulated audit graph is being written blank at the source. **+wave 2026-07-28 — 3 of 4 criteria.** All four literals are gone: `rulesExecuted` is the deduped (rule, engine) set the pipeline actually executed, `missingEvidence` names declared-but-unpresented artifacts, `risks` are built from GT-569's vocabulary keeping `skipped` (medium) distinct from `errored` (high), and the engine is resolved per rule instead of hardcoded to `'opa'`. SARIF now emits the run's rule CATALOG, so a clean evaluation stops reading as "0 rules evaluated". 26 new tests. **NOT closed:** `risks` stays empty on a REAL run until `SatelliteEvaluationPipeline` attaches `coverage` to its verdict — that producer is in `application/services/`, outside the agent's area — and the engine is DERIVED from the rule path rather than DECLARED, because `RuleEvaluation` has no `engine` field. The mapper already prefers an explicit stamp the moment one exists, so the attribution is inference, not testimony. **CLOSED 2026-07-28 by verification, not by new code.** The fix had landed in an earlier wave and nobody checked whether it satisfied the row. All four criteria are met: `mapPipelineVerdict` populates `rulesExecuted` from `executedByKey` and a test asserts the count against a fixture with a known rule count; `resolveRuleEngine` attributes each rule to the evaluator that ran it, with tests over the native AND the OPA path plus a rule-level override; the SARIF exporter derives `evaluatedRules` from `rulesExecuted` (deduped and sorted) and asserts a non-zero list; and `risks` and `missingEvidence` are populated in the GT-569 vocabulary — a rule the engine could NOT evaluate becomes a risk, distinct from one that threw, and what could not be checked is never counted as checked. **Verified by mutation rather than by reading:** reverting `rulesExecuted` to the unconditional empty array turns 10 tests red across the mapper and the SARIF exporter. 40/40 green. | | | `Evolith Core` | Cross | P0 | S | `DONE` | -| [`GT-602`](./gap-reference-catalog.md#gt-602) | **15 of the 50 MCP tools are denied in production by the compiled policy the dispatcher actually loads.** Established by loading `src/sdk/cli/rulesets/opa/policy.wasm` — the artifact used at dispatch, not the source — and evaluating it for an `architect` in `production`: `evolith-adr-list`, `evolith-adr-get/create/update/matrix`, `evolith-pattern-list/get/list-by-topology`, `evolith-scaffold`, `evolith-docs-scaffold`, `evolith-init-batch`, `evolith-sdlc-generate`, `evolith-fixtures` and `evolith-upgrade-plan/apply` all return `ABAC-03` + `ABAC-01`. Since `mcp-tool-dispatch.ts:146` requires native **and** OPA to allow, every one is FORBIDDEN. The rego's own comment predicted this; `abac-classification-coverage.spec.ts` guards TS↔registry but **nothing guards rego↔TS**. Consequence: an agent asking Evolith for its own ADRs is refused, silently and totally. Wave 2: the P0 is fixed — 15 tool/role pairs permitted, each argued individually, and the widening is structurally nil because dispatch requires native AND opa and all 15 were already allowed by native ABAC. Viewer/production still denied on all 20 write tools. A bidirectional rego↔TypeScript parity spec guards it. **AC1 NOT credited:** the wasm test SELF-SKIPS in CI because `Test mcp-server` does not run `build:policy`, so CI enforces parity against the rego SOURCE, not the compiled bundle — a test that skips where it is enforced is the vacuous-pass pattern this board tracks. AC2 untouched. **Criterion 2 attempted 2026-07-28 and NOT shipped, because the naive version is a production regression.** Generating the three rego sets from `TOOL_CLASSIFICATION` looks obvious and is wrong: the map holds 50 tools, the rego enumerates 62, and the 9-tool difference (`evolith-ping`, `evolith-echo`, `evolith-read-file`, `evolith-list-dir`, `evolith-read-gap-tracking`, `evolith-gate-status`, `evolith-write-file`, `evolith-replace-file`, `evolith-run-command`) is classified at runtime by `classifyTool`'s FALLBACK — the legacy `READ_TOOLS`/`WRITE_TOOLS`/`DEPLOY_TOOLS` sets plus name heuristics (`includes()` over read/write/run/deploy and friends), not by the map. A generator sourced from the map alone would have DELETED those nine from the policy, making them unclassified and therefore denied by ABAC-03 in production — the same failure mode this gap was registered for, reintroduced by the fix. Caught by diffing membership before and after rather than trusting the diff to be cosmetic; nothing was committed. **The correct source is the registered tool list classified through `classifyTool` itself, not the map**, which means the mcp-server package must first expose that projection (built `dist`, or a dumped artifact) so the generator consumes one answer instead of re-implementing the fallback and creating a third copy of the logic. `abac-rego-parity.spec.ts` passes today precisely because it compares against `classifyTool`, fallback included — which is why it never flagged the 12-tool gap between the map and the rego. **Criterion 2 CLOSED 2026-07-29, on the second design.** The rego tool sets are now derived from `AbacEvaluator.toolProjection()` — the enumerable union of every set the code declares, each name classified by CALLING `classifyTool` rather than by re-reading its maps. The generator runs that projection through `ts-node` against the TypeScript source, so it needs no build and creates no third copy of the fallback logic. Result: **64 tools, ZERO losses** against the hand-maintained policy, plus two names (`read-tool`, `write-tool`) the runtime already classified and the rego omitted. That is a widening of two fixture names the TypeScript side already allowed, so the native∧opa intersection is unchanged. `--check` is wired into the `Test mcp-server` job and the artifact is registered as the FIRST link of GT-630's derived-artifact chain, which now runs 3 links and stays at a fixed point. mcp-server 434/434 with the generated policy; OPA policy tests 224/224. The first design — sourcing from `TOOL_CLASSIFICATION` alone — was thrown away because it would have deleted twelve tools and denied them in production; that reasoning is preserved above rather than erased by the fix that followed. | | | `Evolith MCP` | Construction | P0 | M | `IN-PROGRESS` | +| [`GT-602`](./gap-reference-catalog.md#gt-602) | **15 of the 50 MCP tools are denied in production by the compiled policy the dispatcher actually loads.** Established by loading `src/sdk/cli/rulesets/opa/policy.wasm` — the artifact used at dispatch, not the source — and evaluating it for an `architect` in `production`: `evolith-adr-list`, `evolith-adr-get/create/update/matrix`, `evolith-pattern-list/get/list-by-topology`, `evolith-scaffold`, `evolith-docs-scaffold`, `evolith-init-batch`, `evolith-sdlc-generate`, `evolith-fixtures` and `evolith-upgrade-plan/apply` all return `ABAC-03` + `ABAC-01`. Since `mcp-tool-dispatch.ts:146` requires native **and** OPA to allow, every one is FORBIDDEN. The rego's own comment predicted this; `abac-classification-coverage.spec.ts` guards TS↔registry but **nothing guards rego↔TS**. Consequence: an agent asking Evolith for its own ADRs is refused, silently and totally. Wave 2: the P0 is fixed — 15 tool/role pairs permitted, each argued individually, and the widening is structurally nil because dispatch requires native AND opa and all 15 were already allowed by native ABAC. Viewer/production still denied on all 20 write tools. A bidirectional rego↔TypeScript parity spec guards it. **AC1 NOT credited:** the wasm test SELF-SKIPS in CI because `Test mcp-server` does not run `build:policy`, so CI enforces parity against the rego SOURCE, not the compiled bundle — a test that skips where it is enforced is the vacuous-pass pattern this board tracks. AC2 untouched. **Criterion 2 attempted 2026-07-28 and NOT shipped, because the naive version is a production regression.** Generating the three rego sets from `TOOL_CLASSIFICATION` looks obvious and is wrong: the map holds 50 tools, the rego enumerates 62, and the 9-tool difference (`evolith-ping`, `evolith-echo`, `evolith-read-file`, `evolith-list-dir`, `evolith-read-gap-tracking`, `evolith-gate-status`, `evolith-write-file`, `evolith-replace-file`, `evolith-run-command`) is classified at runtime by `classifyTool`'s FALLBACK — the legacy `READ_TOOLS`/`WRITE_TOOLS`/`DEPLOY_TOOLS` sets plus name heuristics (`includes()` over read/write/run/deploy and friends), not by the map. A generator sourced from the map alone would have DELETED those nine from the policy, making them unclassified and therefore denied by ABAC-03 in production — the same failure mode this gap was registered for, reintroduced by the fix. Caught by diffing membership before and after rather than trusting the diff to be cosmetic; nothing was committed. **The correct source is the registered tool list classified through `classifyTool` itself, not the map**, which means the mcp-server package must first expose that projection (built `dist`, or a dumped artifact) so the generator consumes one answer instead of re-implementing the fallback and creating a third copy of the logic. `abac-rego-parity.spec.ts` passes today precisely because it compares against `classifyTool`, fallback included — which is why it never flagged the 12-tool gap between the map and the rego. **Criterion 2 CLOSED 2026-07-29, on the second design.** The rego tool sets are now derived from `AbacEvaluator.toolProjection()` — the enumerable union of every set the code declares, each name classified by CALLING `classifyTool` rather than by re-reading its maps. The generator runs that projection through `ts-node` against the TypeScript source, so it needs no build and creates no third copy of the fallback logic. Result: **64 tools, ZERO losses** against the hand-maintained policy, plus two names (`read-tool`, `write-tool`) the runtime already classified and the rego omitted. That is a widening of two fixture names the TypeScript side already allowed, so the native∧opa intersection is unchanged. `--check` is wired into the `Test mcp-server` job and the artifact is registered as the FIRST link of GT-630's derived-artifact chain, which now runs 3 links and stays at a fixed point. mcp-server 434/434 with the generated policy; OPA policy tests 224/224. The first design — sourcing from `TOOL_CLASSIFICATION` alone — was thrown away because it would have deleted twelve tools and denied them in production; that reasoning is preserved above rather than erased by the fix that followed. **Criteria 1 and 3 CLOSED 2026-07-30 — the wasm block no longer self-skips.** The compiled-bundle block of `abac-rego-parity.spec.ts` was governed by `EVOLITH_REQUIRE_OPA_WASM`: it ran in one CI job and `it.skip`ped everywhere else, because `policy.wasm` is a gitignored build output — precisely the shape that produced [`GT-632`](./gap-reference-catalog.md#gt-632). The block now COMPILES the policy itself when it is missing or older than any `.rego` it is built from (via `.harness/scripts/compile-opa-wasm.mjs`, the same script `npm run build:policy` runs, writing the same paths the runtime loads), and if that compilation cannot happen the suite fails with the compiler's own diagnostic: **there is no environment in which it passes without having evaluated a wasm**, and staleness is treated as absence because a stale bundle IS the defect. Denominators are asserted rather than assumed — at least the 50 registered tools measured today, `evaluated.length === registered.length`, and `allowed.length === registered.length` — and the ABAC-03 negative control runs FIRST, since an `evaluate()` returning an unexpected shape would make every ALLOW assertion pass vacuously. `compile-opa-wasm.mjs` installs by rename instead of `copyFileSync`, because a test that compiles on demand makes read-during-write between sibling jest workers reachable, and a partial read throws in `loadPolicy`, which the evaluator reports as OPA_ERROR — a denial. **Verified 2026-07-30 by the negative half, not by the colour:** both `policy.wasm` artifacts (`src/rulesets/opa/` and `src/sdk/cli/rulesets/opa/`) were moved off disk and the spec re-run — 11/11 green and both artifacts recompiled with fresh timestamps, which is what separates "does not skip" from "passed because the file happened to be there". | | | `Evolith MCP` | Construction | P0 | M | `DONE` | | [`GT-603`](./gap-reference-catalog.md#gt-603) | **The agent-turn ledger is complete, unit-tested and absent from dependency injection, and the actor column cannot be typed retroactively.** `Tracker.Application/Integration/AgentExecution/AgentExecutionService.cs` validates scope, then **audits before executing and aborts the turn if the audit write fails**; `AgentTurnAuditor.cs` records granted-vs-used scopes and stores prompt length, not text. `IAgentExecutionPort` appears in **zero DI registrations and zero endpoints**, while `AssistantEndpoints.cs` proxies straight through `AgentRuntimeGateway` persisting nothing. Separately `AuditEntryProps.cs:11` declares `public Guid ActorId` with no `actor_type`, `agent_id`, `model_id` or `session_id`. Because `audit_entries` is append-only by database trigger (migration `20260719202323`), **rows written before the discriminator exists can never be corrected**. Complements [`GT-586`](./gap-reference-catalog.md#gt-586), which covers the Core-side `EvaluationContext`; this is the Tracker persistence side. | | | `Evolith Tracker` | Cross | P0 | M | `PENDING` | | [`GT-604`](./gap-reference-catalog.md#gt-604) | **No surface writes evidence to the Tracker: the write path is one-directional and points inward.** Grep across `src/sdk/cli`, `src/packages/mcp-server` and `src/packages/core-domain` returns no Tracker base URL and no ingest client; the only Tracker URL in the Core repo is `AGENT_RUNTIME_APPROVAL_TRACKER_URL`, and the only writers of `core_evaluation_transactions` are Tracker-initiated endpoints. Consequence: every `evolith validate`, every `enforce edit` veto, every MCP `tools/call` and every CI drift-gate run evaporates on process exit. The strategy is premised on accumulated evidence and the components that generate it have no way to deposit it. This is a composition defect: no single component review can see it, because each one is internally consistent. **IN-PROGRESS 2026-07-29 — criterion 1 only.** The ingest contract landed at `src/packages/contracts/src/ingest/evaluation-ingest.ts` with `correlationId` REQUIRED and BOTH owners carried distinctly: `requestedBy.actorId` is who asked, `violations[].accountableOwner` is who must fix. Criterion 2 cannot be completed from here — there is no Tracker ingest endpoint to call. Criterion 3 is 0% buildable in this repository: RoboSoft lives in `evolith_tracker`. A bilingual handover was written at `reference/core/control-center/opportunities/tracker-handover-gt604.md` and its `.es.md` counterpart, 22 headings each. **Unconfirmed, and recorded as such: this row's declared dependency on GT-603 appears wrong.** GT-603 migrates `audit_entries` while this row names `core_evaluation_transactions`, and the Core-side attribution already shipped under GT-586 — but the Tracker schema is not in this repository, so that is a hypothesis for the Tracker owner to confirm, not a finding. | | | `Evolith Suite` | Cross | P0 | L | `IN-PROGRESS` | | [`GT-605`](./gap-reference-catalog.md#gt-605) | **Two evidence graphs exist, each missing the other's half.** `src/packages/core-domain/src/evidence/evidence-graph.ts` defines typed edges (`requires` / `validates` / `blocks`) and has **zero consumers outside its own spec file**. The Tracker persists `EvidenceRecordProps.References` as `List` into a jsonb column, whose only non-test consumer is a linear `Contains()` used for external-id dedup — no edge table, no edge type, no reverse lookup, no depth query. The typed model lives where nothing persists; the persisted model lives where nothing is typed. Consequence: "which ADR moved because of which gate decision because of which agent turn" is unanswerable, and that traversal is the stronger half of the stated moat. Wave 2: the single typed model is published in `@beyondnet/evolith-contracts` with the four things a shared type alone leaves ambiguous — vocabulary (the Core’s three verbatim plus `caused_by`, since none of requires/validates/blocks expresses "the ADR moved BECAUSE OF the decision"), direction fixed in prose, a canonical node encoding that makes the Tracker’s jsonb list mechanically migratable, and a depth-bounded traversal its recursive CTE can be tested against. **Remaining:** three criteria are Tracker-side and out of repository; in THIS repo core-domain still declares its own `EvidenceEdge`, so there are two type declarations for one model. | | | `Evolith Suite` | Cross | P1 | M | `IN-PROGRESS` | @@ -655,7 +655,7 @@ This board is the single source of truth for technical debt, gaps, opportunities | [`GT-246`](./gap-reference-catalog.md#gt-246) | Implement Chaos Mesh/Litmus experiments | | | `QA` | Cross | P3 | L | `DONE` | -**Progress:** 600 / 640 done · 15 in progress · 21 pending · 4 deferred +**Progress:** 601 / 640 done · 14 in progress · 21 pending · 4 deferred **Wave 2026-06-23 (Winston deep audit III):** Added 14 new gaps `GT-212`…`GT-225` from the Winston Audit Playbook covering: ADR status hygiene (GT-212), topology manifest metadata + operational budgets + guidance corpus (GT-213, GT-217, GT-219), REST controller observability + OpenAPI (GT-214, GT-215), OPA input-schema parity + per-topology test density (GT-216, GT-222), SDLC Phase 05 rollback + on-call templates (GT-218), CLI branch coverage + envelope format coverage + skip-list cleanup (GT-220, GT-224, GT-225), MCP HTTP audit logging (GT-221), and cross-surface parity e2e tests (GT-223). diff --git a/reference/core/control-center/maturity-reports/maturity-reconciliation.json b/reference/core/control-center/maturity-reports/maturity-reconciliation.json index 48f12b6f..3881b592 100644 --- a/reference/core/control-center/maturity-reports/maturity-reconciliation.json +++ b/reference/core/control-center/maturity-reports/maturity-reconciliation.json @@ -4,13 +4,13 @@ "asOf": "2026-07-26", "gaps": { "total": 640, - "done": 600, + "done": 601, "pending": 21, - "inProgress": 15, + "inProgress": 14, "deferred": 4 }, "evidence": { - "closureRecords": 582, + "closureRecords": 583, "cliPackage": "@beyondnet/evolith-cli@1.2.2", "adrCount": 139, "rulesetCount": 175, From ea37bec162f88641cfd2b5cab7d9b2ab24e22fc1 Mon Sep 17 00:00:00 2001 From: aarroyo Date: Thu, 30 Jul 2026 20:55:45 -0500 Subject: [PATCH 2/3] =?UTF-8?q?docs(gaps):=20one=20status=20per=20row=20?= =?UTF-8?q?=E2=80=94=20drop=20two=20dead=20columns=20and=20stop=20restatin?= =?UTF-8?q?g=20status=20in=20prose?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A row carried its status in up to three places at once: the `Estado`/`Status` column, a bolded restatement inside the narrative, and — for 45 of 616 entries — a `**Status:**` field in the catalog. Reading a row meant reconciling them. Two mechanical changes, no content lost: `Qué significa` / `What it means` and `Ejemplo` / `Example` were empty in 599 of 614 rows (97.6%). They widened every row and pushed the status column out of view for the 2.4% that used them. Dropped; the 15 rows that did use them keep their text folded into the Gap cell behind bold labels. The narrative's status announcements are re-labelled by what they actually are — a dated record, not a second status field: `**COMPLETADO (x):**` -> `**Cierre (x):**`, `**EN-PROGRESO (x):**` -> `**Avance (x):**`, and the English twins to `**Closure**` / `**Progress**`. 71 replacements across 61 ES rows, 74 across 64 EN rows. Prose that discusses the vocabulary is untouched on purpose — "a row can sit IN-PROGRESS forever with no acceptance criteria" is an argument about the board, not a second status, and rewriting it would have destroyed the point. The one genuine restatement outside a bold label (GT-518 "Sigue EN-PROGRESO (...)") became "Sigue abierto (...)". Guard 08 finds its column by header name rather than by position, so removing columns does not touch it: 640 gaps, 616/616 catalog sections, 583 closure records, green. Bilingual parity green (1643 files), doc validation green (1491). Checked and NOT changed: the 45 catalog `**Status:**` fields contradict the board in zero cases, so they are redundant rather than wrong — left for a decision rather than deleted in a formatting commit. Co-Authored-By: Claude Opus 5 --- .../control-center/gaps/gap-tracking.es.md | 1284 ++++++++--------- .../core/control-center/gaps/gap-tracking.md | 1284 ++++++++--------- 2 files changed, 1284 insertions(+), 1284 deletions(-) diff --git a/reference/core/control-center/gaps/gap-tracking.es.md b/reference/core/control-center/gaps/gap-tracking.es.md index e52dcc3a..0abad13a 100644 --- a/reference/core/control-center/gaps/gap-tracking.es.md +++ b/reference/core/control-center/gaps/gap-tracking.es.md @@ -11,648 +11,648 @@ Este tablero es la única fuente de verdad para deuda técnica, gaps, oportunida > Una sola tabla con todos los gaps y actividades rastreadas. Los IDs `GT-*` enlazan a su detalle completo en el catálogo; los IDs `MT-A*` enlazan al plan de implementación Multi-Topology de apoyo, pero esta tabla sigue siendo la fuente canónica de estado. Orden: pendientes primero (por criticidad P0→P3, luego complejidad XS→XL), después completados (mismo criterio). GitHub renderiza Markdown de forma estática (sin orden ni búsqueda interactivos): la columna **Componente** categoriza y la búsqueda de archivo de GitHub (`/`) encuentra un ID o término. -| ID | Gap | Qué significa | Ejemplo | Componente | Fase | Criticidad | Complejidad | Estado | -|---|---|---|---|:---:|:---:|:---:|:---:|:---:| -| [`GT-640`](./gap-reference-catalog.es.md#gt-640) | **El guard comparaba el snapshot contra seis números copiados del propio snapshot, así que solo podía pasar — y el archivo que no estaba vigilando se estampa en 388 filas de un artefacto mayor.** `native-evaluability-snapshot.json` declaraba en su propia cabecera que era "un SNAPSHOT CAPTURADO, no la fuente de verdad". **No existía ningún script de captura.** Se mantenía a mano y derivó: fijaba `documentation-only: 129` mientras Core fijaba 136, y seguía llamando `unimplemented-native` a reglas que ya tenían handler. Su guard en `iso-5055-mapping.test.mjs` asertaba los seis conteos de clase del snapshot contra seis números escritos a mano en el test — los mismos seis literales que el snapshot ya contenía. **El valor esperado y el valor real eran copias el uno del otro**, así que la aserción se sostenía mientras el archivo no cambiara, fijara Core lo que fijara; informaba de una concordancia que nunca comprobó, lo cual es peor que no tener guard, porque la fila que protege se lee como verificada. **La deriva no se queda donde empieza.** `build-iso-5055-mapping.mjs` estampa `nativeEvaluability` en TODAS las filas del mapeo ISO/IEC 5055 desde ese archivo — 388 tras la recaptura, 381 antes — así que una sola clase rancia se blanquea en un artefacto derivado cinco veces mayor que su entrada, y sobredimensiona el backlog de handlers que es el entregable entero de [`GT-598`](./gap-reference-catalog.es.md#gt-598). El orden recaptura→reconstrucción no estaba escrito en ningún sitio, y el guard del propio mapeo no corría en **ningún workflow**. **CORREGIDO el 2026-07-29:** un script de captura que ejecuta el triage REAL de Core vía `ts-node` en vez de reimplementar la clasificación, la medición extraída del spec de jest a `test/rule-corpus-triage.ts` para que un script pueda alcanzarla — esa inalcanzabilidad es la razón de que el archivo se mantuviera a mano —, aserciones snapshot-contra-triage-fresco en ambas direcciones, un guard del job de documentación que lee los conteos que Core fija LEYENDO el spec y **lanza excepción** cuando el literal falta o cambia de forma, y ambos pasos declarados como eslabones de la cadena de artefactos derivados de [`GT-630`](./gap-reference-catalog.es.md#gt-630). **Recapturado el 2026-07-29 tras mergear el fix en la línea actual, y la deriva había crecido mientras estaba en vuelo:** `native-handler` 139 → 151, `documentation-only` 129 → 136, el backlog de handlers 60 → **48**, corpus 379 → 386, mapeo 381 → 388 filas. Doce de esos handlers los cerraron [`GT-595`](./gap-reference-catalog.es.md#gt-595) y [`GT-632`](./gap-reference-catalog.es.md#gt-632) mientras el snapshot mantenido a mano seguía reportándolos como backlog — el defecto demostrándose una vez más, sobre números que nadie escribió dos veces. **RENUMERADA GT-633 -> GT-640 el 2026-07-30**, después de que `49-validate-gap-id-allocation` detectara que `GT-633` ya nombraba otro gap en `develop` (el scaffolding de `evolith init`, `c8270e48`, registrado un día antes). Se mueve la fila más nueva, según el propio consejo del guard. | | | `Governance` | Cross | P1 | S | `COMPLETADO` | -| [`GT-639`](./gap-reference-catalog.es.md#gt-639) | **Nada hace visible el trabajo en vuelo entre ramas, así que el mismo gap se arregla dos veces.** El 2026-07-30 el mismo trabajo se hizo dos veces, en tres ocasiones distintas, por sesiones que no podían verse: [`GT-640`](./gap-reference-catalog.es.md#gt-640) (registrado como GT-633 entonces) arreglado por un script de captura en `main` (#276) Y por un renderer independiente dentro del spec en `develop` (#279) — los dos correctos, los dos recapturando los mismos números, y dos generadores para un artefacto es el defecto del propio GT-640 un nivel más arriba, así que una tercera sesión gastó una reconciliación entera en dejar uno; el parser del guard de evidencias arreglado dos veces (`d1ea72a3` en main, `5acb29ed` en develop); y un id de GT asignado dos veces, forzando la renumeración que hay detrás de [`GT-638`](./gap-reference-catalog.es.md#gt-638). **El coste no es el conflicto de merge** — es el trabajo duplicado antes de que nadie llegara a él, más la reconciliación, que introdujo un defecto propio (una comparación sensible al orden que reescribía la fecha de captura en cada corrida, invisible un día porque ambas estampaban la misma fecha). Una fila lleva estado pero no QUIÉN la trabaja ni DÓNDE, así que dos sesiones pueden leer ambas `PENDIENTE` y empezar ambas, correctamente. **La convención de un solo driver ya existe; lo que falta es un mecanismo que la haga observable**, así que en un día cargado se degrada en silencio. **GUARD CONSTRUIDO el 2026-07-30 — tres criterios de cuatro.** `50-validate-gap-claim` deriva las reclamaciones de los PULL REQUESTS ABIERTOS (cada `GT-*` en título, cuerpo o nombre de rama) y falla cuando un id lo reclaman dos, nombrando ambos con su rama. Derivado y no a mano a propósito: una reclamación escrita a mano se queda rancia en la dirección que importa. Cableado en `Governance guards (GT-578)` con `GH_TOKEN`; observado en rojo por `43-validate-guard-negative-fixtures` (39/39). **El criterio 1 sigue abierto a propósito:** una lista commiteada en el tablero derivaría de estado vivo de GitHub y estaría rancia en cuanto alguien abriera un PR — y la cadena de [`GT-630`](./gap-reference-catalog.es.md#gt-630) existe para exigir que los derivados alcancen un punto fijo, así que uno perpetuamente rancio sería peor que ninguno. Tampoco ve una rama sin PR abierto, y lo dice en su propio texto de fallo. | | | `Governance` | Cross | P2 | M | `EN-PROGRESO` | -| [`GT-638`](./gap-reference-catalog.es.md#gt-638) | **El tablero no tiene asignador de ids, así que dos ramas paralelas reparten el mismo número GT.** Un id se elige leyendo el `GT-*` más alto de la rama en la que uno está, y nada lo contrasta con otra rama — así que dos sesiones en paralelo asignan el mismo número y una se entera en el merge. Ya pasó: `8449af3d` en `develop` dice *"renumber the ratchet fix GT-634 -> GT-637, ID collision with develop"*; esa fila y [`GT-634`](./gap-reference-catalog.es.md#gt-634) en `main` tomaron el mismo id con horas de diferencia, desde sesiones que no podían verse. **La renumeración es manual y con pérdidas**, que cuesta más que el choque: un id vive en la fila del tablero, el ancla del catálogo, el registro de evidencia de cierre, referencias cruzadas en dos idiomas, mensajes de commit y cuerpos de PR — y sólo los tres primeros son comprobables mecánicamente. `08-validate-tracking` no puede ayudar por construcción: valida dentro de un único árbol de trabajo, y las ramas quedan fuera de su mundo. Tampoco es un desliz aislado — en el merge `develop` → `main` del 2026-07-30 el mismo patrón de duplicación convergente aparece tres veces: GT-640 (entonces GT-633) arreglado dos veces, el parser del guard de evidencias arreglado dos veces, este id asignado dos veces. **CERRADO el 2026-07-30 — y en su primera corrida real encontró una SEGUNDA colisión que nadie conocía.** `49-validate-gap-id-allocation` compara el **Title:** de cada id en HEAD contra el título que ese id lleva en la rama base; un título distinto para el mismo id es un número nombrando dos gaps. Contra `origin/develop` reportó **`GT-633`**: allí es "`evolith init` scaffolds a repository that cannot pass its own governance on the first run" (`c8270e48`, 2026-07-29) y en main era la fila del guard tautológico (`c3221276`, 2026-07-30), renumerada después a `GT-640`. El merge no lo había sacado a la luz. Se niega a adivinar qué lado está mal — un retítulo y una colisión se ven igual desde fuera —, así que nombra los dos. Cableado en `Governance guards (GT-578)` con un `git fetch` explícito de la base, porque un checkout superficial de CI no la tiene. Observado en rojo por `43-validate-guard-negative-fixtures` (38/38). Sigue sin ver un id asignado en una rama sin PR abierto, y su salida `--verbose` lo advierte en cada id nuevo. | | | `Governance` | Cross | P2 | S | `COMPLETADO` | -| [`GT-637`](./gap-reference-catalog.es.md#gt-637) | **El ratchet de referencias muertas de GT-578 estaba atascado en 305 por dos bugs reales en el propio GUARD, no por el tamaño del backlog de registros que decía medir.** Al cerrar el PR de `GT-633` apareció un fallo en `Governance guards (GT-578)` y, al investigarlo, un refactor sin commitear y sin terminar que ya estaba en el árbol de trabajo: `resolveCommand` llamaba a `npmScriptOperand(...)`, una función que nunca se definió, así que el guard crasheaba con `ReferenceError` en cada invocación — incluso en modo reporte plano. Reparar esa función (restaurar su comportamiento y luego corregirlo) bajó el conteo de 309 a 78, lo que destapó el segundo bug: el corpus usa `npm run --workspace