diff --git a/CHANGELOG.md b/CHANGELOG.md index 89c7f83..51d957b 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -6,6 +6,37 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/). ## [Unreleased] +## [1.37.19-2] - 2026-10-04 + +Plugin-only patch. The post-merge eval on `main` (run 37196169047) failed four assertions on code that had +passed on the PR — three real behaviors the skill text left room for, per "a flaky assertion is a product defect". + +### Fixed + +- **poc-planner kept a template comment on a TBD line.** `shape: TBD — decided by # er_quality | + functional_integration | …` left the template's candidate list after the literal (unanimous judge FAIL on + `judge-nothing-after-tbd`). The template's `# …` guidance now sits on its own lines above the keys with + "delete these", and `validate_plan.py` rejects any `# …` riding on a TBD line (its owner pattern swallowed + the comment, so it never saw one). +- **recipes handed off to `install` without fetching the named recipe** (`correct-recipe-fetched` and + `doctor-before-recipe` failed: no Bash after doctor). `install` ends the turn, so a fetch left for afterwards + never happens; the skill now says to fetch BEFORE invoking it. +- **demo recited the forbidden production offer** ("…never touched, and I won't load into it unless you ask for + that explicitly" — `no-production-offer`). The skill gave the rule but not the sentence; it now gives the whole + sentence ("Your production repository is never touched.") and names what must not follow it. +- **demo offered the zero-install fallback in other words** (PR #63's own eval, `demo-no-simulation/criteria`, 3-0 + FAIL): "If either can't be granted here, say so and I'll fall back to showing you validated … sample data — but I + won't do that silently; I wanted to ask first." The skill now says why silence is the honest course (the tier is + always taken and labelled the moment `install` ends without an SDK), and `no-install-or-preview-menu` matches + `if/should/unless … fall back to showing` (new fixture; the pattern previously needed the word "zero-install"). +- **poc-planner disclaimed the per-step duration instead of omitting it** (`poc-planner-how-long/no-duration`, run + after the demo fix): "…mapping a single data source 'usually takes less than 30 minutes' — a per-step note, not an + estimate for the POC." The decline rule now says disclaiming a figure states it, and to say nothing about it. +- **build stopped at doctor's report** (PR #63 eval, case `build`: 6 graders failed — no `generate_scaffold`, no + `senzing_search.py`; the run also asked the grounder sub-agent to "run doctor"). The skill now says doctor's + report is a checkpoint, invoked with the Skill tool, and the run continues to step 2 in the same turn. +- `scripts/check.sh`: probes for the first three. + ## [1.37.19-1] - 2026-10-04 Plugin-only patch on server 1.37.19 (skills, graders, validator; no server change). diff --git a/evals/demo-no-simulation/graders/no-install-or-preview-menu.md b/evals/demo-no-simulation/graders/no-install-or-preview-menu.md index e773e14..ccba77c 100644 --- a/evals/demo-no-simulation/graders/no-install-or-preview-menu.md +++ b/evals/demo-no-simulation/graders/no-install-or-preview-menu.md @@ -1,6 +1,6 @@ --- type: regex -pattern: '\b(?:or|rather|skip)\b[^\n]{0,80}\bzero-install\b|\b(?:I|we)(?:\s+can|\s+could|\s+would|''d)\s+(?:instead\s+|also\s+)?fall\s+back\s+to\b[^\n]{0,60}\bzero-install\b|\bzero-install\b[^\n]{0,60}\binstead\b|(?:\d\.|[A-C]\))\s*\*\*\s*zero-install\b|\bskip\s+(?:the\s+)?(?:local\s+)?install(?:ing|ation)?\b' +pattern: '\b(?:or|rather|skip)\b[^\n]{0,80}\bzero-install\b|\b(?:I|we)(?:\s+can|\s+could|\s+would|''d)\s+(?:instead\s+|also\s+)?fall\s+back\s+to\b[^\n]{0,60}\bzero-install\b|\bzero-install\b[^\n]{0,60}\binstead\b|(?:\d\.|[A-C]\))\s*\*\*\s*zero-install\b|\bskip\s+(?:the\s+)?(?:local\s+)?install(?:ing|ation)?\b|\b(?:if|should|unless)\b[^\n.]{0,100}\bfall\s+back\s+to\s+showing\b' match: not_contains flags: i target: last_message @@ -29,3 +29,11 @@ A bare `fall back to ... zero-install` is NOT matched: a correct run that took t "I could", "I would", "I'd" — including the conditional attached to the permission request that recurred after the first rule ("or, if you'd rather not change the sandbox rules, I can instead fall back to a zero-install preview"). Measured on 82 real arms: 8 hits, every one a judge FAIL, none on a pass. + +## Conditional fallback without the word "zero-install" + +CI (PR #63 head `34d1a83`, demo-no-simulation) failed 3-0 on "If either can't be granted here, say so and I'll +fall back to showing you validated, Senzing-ready sample data without live resolution … — but I won't do that +silently; I wanted to ask first." No `zero-install` token, no "can/could" modal — the same conditional attached +to the permission request. `if/should/unless … fall back to showing` is the offer; a run that TOOK the tier +after install ended writes "I fell back" or states it as done, and has no leading conditional. diff --git a/evals/demo-no-simulation/pattern-fixtures/no-install-or-preview-menu.yaml b/evals/demo-no-simulation/pattern-fixtures/no-install-or-preview-menu.yaml index 4b08ade..4f945b2 100644 --- a/evals/demo-no-simulation/pattern-fixtures/no-install-or-preview-menu.yaml +++ b/evals/demo-no-simulation/pattern-fixtures/no-install-or-preview-menu.yaml @@ -36,6 +36,9 @@ must_match: preview**: pull sample data' - in which case I'd fall back to a zero-install preview (validated sample data only, no simulated resolution) or Docker. +- 'If either can''t be granted here, say so and I''ll fall back to showing you validated, Senzing-ready sample data + without live resolution (clearly labeled as a preview, not a demo of the engine) — but I won''t do that silently; + I wanted to ask first.' must_not_match: - 'Senzing isn''t installed yet. I''ll set it up using the official Homebrew cask (macOS Apple Silicon), with **TypeScript/Node.js** as the language — it''s the only toolchain already available on this machine diff --git a/evals/poc-planner-how-long/pattern-fixtures/no-duration.yaml b/evals/poc-planner-how-long/pattern-fixtures/no-duration.yaml index f329e12..8201868 100644 --- a/evals/poc-planner-how-long/pattern-fixtures/no-duration.yaml +++ b/evals/poc-planner-how-long/pattern-fixtures/no-duration.yaml @@ -2,6 +2,8 @@ # First 8 must_match + 3 must_not_match: REAL final messages from CI eval runs (not paraphrases). # The rest are HAND-WRITTEN, one per duration unit, so dropping a unit from the pattern is caught. must_match: +- 'The one place it mentions a span at all is that mapping a single data source "usually takes less than 30 minutes" — + a per-step note about one activity, not an estimate for the POC.' - The article also says data mapping itself "usually takes less than 30 minutes per data source" (per source_url above) — but that's one step, not the whole POC. - 'One more data point worth having in view before you size anything: mapping a single data source typically diff --git a/plugins/senzing/.claude-plugin/plugin.json b/plugins/senzing/.claude-plugin/plugin.json index 38593f3..ee9fd3b 100644 --- a/plugins/senzing/.claude-plugin/plugin.json +++ b/plugins/senzing/.claude-plugin/plugin.json @@ -3,7 +3,7 @@ "name": "senzing", "displayName": "Senzing", "icon": "./icon.png", - "version": "1.37.19-1", + "version": "1.37.19-2", "description": "Entity resolution: point Claude at your data, it does the work, and you have results today.", "author": { "name": "Senzing, Inc.", diff --git a/plugins/senzing/skills/build/SKILL.md b/plugins/senzing/skills/build/SKILL.md index 08eb5f9..a6d530a 100644 --- a/plugins/senzing/skills/build/SKILL.md +++ b/plugins/senzing/skills/build/SKILL.md @@ -58,6 +58,14 @@ answers the run gate; the write gate needs one probe of its own: `/senzing:build` in **Claude Code** on that host. Never claim to have edited files you could not write, and never claim a run you did not perform. +**`doctor`'s report is a checkpoint, not the answer.** Invoke it with the Skill tool itself — never hand +"run doctor" to a sub-agent (the grounder agent has no doctor and answers that none exists). Whatever its +verdict — no Senzing, Python on macOS, nothing installed — carry on **in the same turn** to step 2: the file +is still the deliverable, written by the file tools whether or not Senzing exists on the host. Only step 4 +(run) waits on the verdict. CI caught a run that printed doctor's table as its final message and never called +`generate_scaffold` or wrote `senzing_search.py`; a message that is only an environment report has not done +what was asked. + Always: 1. **Inputs.** `$ARGUMENTS` may name the language and/or workflow (e.g. `python search`). Determine diff --git a/plugins/senzing/skills/demo/SKILL.md b/plugins/senzing/skills/demo/SKILL.md index 3ea0af8..ebd7230 100644 --- a/plugins/senzing/skills/demo/SKILL.md +++ b/plugins/senzing/skills/demo/SKILL.md @@ -120,7 +120,11 @@ second report. preview" is an alternative to installing, and it turns the permission question into a fork. If the user declines the permission, THEN you take the tier, in the next turn, as what you do now — never as a way out you offered this turn. The tier is not mentioned until `install` has actually ended its turn - without a working SDK, and then you take it rather than announce it. + without a working SDK, and then you take it rather than announce it. **Saying nothing here is the honest + course, not a withheld one**: the tier is always taken, and always labelled "not a demo of the engine", the + moment `install` ends without an SDK — so "if that can't be granted I'll fall back to showing you sample data + — I won't do that silently" (CI caught exactly that sentence) discloses nothing the user needs now and hands + them the fork. Ask the EULA and permission questions, end the message there. **The zero-install tier is what you do AFTER `install` has ended its turn without a working SDK** — a fallback you take, not an option you put to the user, and never a branch offered in @@ -165,7 +169,10 @@ second report. needed.** The scratch repo is throwaway and touches nothing of theirs. **Never offer, hint at or ask about loading into their existing repository** — not "if you'd rather load it into production, say so", not "I'll only do that if you explicitly ask". Saying production is - untouched is fine; saying it is *available* is an offer. Only if the USER raises their + untouched is fine; saying it is *available* is an offer. The whole sentence is "Your + production repository is never touched." — full stop; no "unless", "only if", "until you + ask" after it (CI caught "…is never touched, and I won't load into it unless you ask for + that explicitly"). Only if the USER raises their existing repository themselves, confirm the target and the record count first. Verify the load as `analyze` step 4 does (loaded vs submitted, error count) before going on. - **Drain the redo queue** as `analyze` step 5 does — get the probe from the MCP, drain to 0, diff --git a/plugins/senzing/skills/poc-planner/SKILL.md b/plugins/senzing/skills/poc-planner/SKILL.md index eca7a5f..a4fbff2 100644 --- a/plugins/senzing/skills/poc-planner/SKILL.md +++ b/plugins/senzing/skills/poc-planner/SKILL.md @@ -53,6 +53,10 @@ license terms here or from memory — call the tool named in each step and cite offered as what you are *not* saying. **A cited quote is not an exception here.** The guidance does carry one per-step figure — mapping a data source — and offering it as "one cited data point" answers "how long" with a partial estimate, the same fabrication wearing a citation. + **Disclaiming it is still stating it**: "the one place it mentions a span is that mapping a source 'usually + takes less than 30 minutes' — a per-step note, not an estimate for the POC" puts the 30 minutes in front of the + reader exactly as the cited-data-point form does (CI caught that sentence). Say nothing about that figure at + all — not as a data point, not as a caveat, not as what the guidance does or doesn't give. **Retrieve first.** "The guidance carries no duration for the POC as a whole" is a claim about Senzing's guidance: make it only after the `search_docs` retrieval at the top of this skill, and name the article you retrieved. A decline written from this paragraph alone is a Senzing claim with no source, about a document you have not opened — a @@ -406,11 +410,14 @@ data_sources: approx_records: entity_types: identifying_columns: [] -database: # their words, e.g. PostgreSQL — or TBD — decided by -os_platform: # the user's words -platform_id: # an id from sdk_guide's platform tree, or TBD — decided by +# database, os_platform, languages: their words (e.g. PostgreSQL, ["Python"]), or the TBD literal +# platform_id: an id from sdk_guide's platform tree, or the TBD literal +# DELETE these comment lines, and never leave a `# …` after a value you fill in. +database: +os_platform: +platform_id: cloud: -languages: [] # their words, e.g. ["Python"] +languages: [] hardware_available: performance_required: throughput: @@ -420,13 +427,17 @@ calendar: ``` ## 3. Success criteria ```yaml +# shape is one of: er_quality | functional_integration | entity_graph_scenario | other +# statement: what must be shown, per user · measurement: as the tool names it — source: +# target: "per user: " or the TBD literal — these seven keys only +# DELETE these comment lines, and never leave a `# …` after a value you fill in. - id: SC-1 - shape: # er_quality | functional_integration | entity_graph_scenario | other - statement: # what must be shown, per user - measurement: # as the tool names it — source: + shape: + statement: + measurement: measured_against: decided_by: - target: # "per user: " or "TBD — decided by " — these seven keys only + target: ``` ## 4. What must be true to buy — goal and scope ## 5. Data selection checklist diff --git a/plugins/senzing/skills/poc-planner/validate_plan.py b/plugins/senzing/skills/poc-planner/validate_plan.py index 3597ab6..b2909e9 100755 --- a/plugins/senzing/skills/poc-planner/validate_plan.py +++ b/plugins/senzing/skills/poc-planner/validate_plan.py @@ -251,6 +251,12 @@ def check(text: str) -> list[str]: m = TBD_RE.search(stripped) if not m: continue + if re.search(r"TBD — decided by [^#\n]*\s#\s", stripped) and line not in sec9: + problems.append( + f"line {i}: a template comment rides on a TBD line — '{stripped[:70]}'. Delete the `# …` " + f"after the owner: a candidate list there answers the question the line says is open." + ) + continue tail = (m.group("tail") or "").lstrip(":").strip() if not tail: continue diff --git a/plugins/senzing/skills/recipes/SKILL.md b/plugins/senzing/skills/recipes/SKILL.md index d6dafdf..16dd251 100644 --- a/plugins/senzing/skills/recipes/SKILL.md +++ b/plugins/senzing/skills/recipes/SKILL.md @@ -123,7 +123,11 @@ words. Match it against the catalog `id`s; on a fuzzy/multiple match, confirm wh Say what the recipe needs (its record count against the unlicensed limit is a fact you may state) and hand off to `install`. - **Senzing can't deploy** → still identify and fetch the named recipe (step 2/3) so the user - learns what it needs, then hand off to the **`install`** skill without asking first (it + learns what it needs, **then** hand off to the **`install`** skill without asking first. + **Order matters: fetch BEFORE you invoke `install`** — `install` asks its question and ends + the turn, so a fetch you leave for afterwards is never made (CI caught a run that went + doctor → `install` and never opened the recipe it was asked to cook). Run the catalog `curl`, + match the id, fetch `recipes/.md`, and only then invoke `install` (it surfaces the license agreement, runs the official steps, and verifies with `doctor`; do not route around it via `sdk_guide(topic="install")` directly — if an evaluation license is needed, `submit_feedback(category='license_request')`'s description states the current diff --git a/scripts/check.sh b/scripts/check.sh index 528f764..25c552c 100755 --- a/scripts/check.sh +++ b/scripts/check.sh @@ -332,6 +332,14 @@ then else bad "validate_plan.py path-builder regressed on a shape the fixtures do not exercise" fi +awk '/^ shape: er_quality/ && !d { sub(/er_quality/, "TBD — decided by the data platform lead # er_quality | other"); d=1 } { print }' "$VP_OK" > "$vp_tmp/tbd-comment.md" +cmp -s "$VP_OK" "$vp_tmp/tbd-comment.md" && bad "8b tbd-comment mutation did not change the plan - the test below proves nothing" +VP_OUT_TC="$(python3 "$VP" "$vp_tmp/tbd-comment.md" 2>&1 || true)" +if [[ "$VP_OUT_TC" == *"template comment rides on a TBD line"* ]]; then + ok "a template '# candidate | list' comment after a TBD literal is rejected" +else + bad "validate_plan.py accepts a candidate-list comment riding on a TBD line" +fi rm -rf "$vp_tmp" echo; echo "== 8c. no tool_used grader declares max without min (impossible range) ==" @@ -729,6 +737,25 @@ else bad "recipes-named criteria lost the install-message clause - the judge will fail the correct hand-off at random" fi +# Main-branch CI (run 37196169047) failed three cases the PR run had passed -- each a real behavior the +# skill text left room for. Text guards keep the fix from being edited away. +if grep -q "fetch BEFORE you invoke" plugins/senzing/skills/recipes/SKILL.md; then + ok "recipes/SKILL.md says to fetch the recipe before invoking install (install ends the turn)" +else + bad "recipes/SKILL.md lost fetch-before-install - a run can hand off without ever opening the recipe" +fi +if grep -q 'is never touched." — full stop' plugins/senzing/skills/demo/SKILL.md; then + ok "demo/SKILL.md gives the exact production sentence and forbids a conditional after it" +else + bad "demo/SKILL.md lost the exact production sentence - the model recites 'unless you ask' again" +fi + +if grep -qF "report is a checkpoint, not the answer" plugins/senzing/skills/build/SKILL.md; then + ok "build/SKILL.md says doctor's report is a checkpoint and the run continues in the same turn" +else + bad "build/SKILL.md lost the doctor-is-a-checkpoint rule - a run can end on the environment table" +fi + echo; echo "== 9. Eval scoring split (deterministic gate vs judge score) ==" # The suite's verdict is two independent gates, computed by evals/gate.py: # deterministic graders must ALL pass in EVERY run (no averaging, no threshold), while the