Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
31 changes: 31 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,37 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).

## [Unreleased]

## [1.37.19-2] - 2026-10-04

Plugin-only patch. The post-merge eval on `main` (run 37196169047) failed four assertions on code that had
passed on the PR — three real behaviors the skill text left room for, per "a flaky assertion is a product defect".

### Fixed

- **poc-planner kept a template comment on a TBD line.** `shape: TBD — decided by <owner> # er_quality |
functional_integration | …` left the template's candidate list after the literal (unanimous judge FAIL on
`judge-nothing-after-tbd`). The template's `# …` guidance now sits on its own lines above the keys with
"delete these", and `validate_plan.py` rejects any `# …` riding on a TBD line (its owner pattern swallowed
the comment, so it never saw one).
- **recipes handed off to `install` without fetching the named recipe** (`correct-recipe-fetched` and
`doctor-before-recipe` failed: no Bash after doctor). `install` ends the turn, so a fetch left for afterwards
never happens; the skill now says to fetch BEFORE invoking it.
- **demo recited the forbidden production offer** ("…never touched, and I won't load into it unless you ask for
that explicitly" — `no-production-offer`). The skill gave the rule but not the sentence; it now gives the whole
sentence ("Your production repository is never touched.") and names what must not follow it.
- **demo offered the zero-install fallback in other words** (PR #63's own eval, `demo-no-simulation/criteria`, 3-0
FAIL): "If either can't be granted here, say so and I'll fall back to showing you validated … sample data — but I
won't do that silently; I wanted to ask first." The skill now says why silence is the honest course (the tier is
always taken and labelled the moment `install` ends without an SDK), and `no-install-or-preview-menu` matches
`if/should/unless … fall back to showing` (new fixture; the pattern previously needed the word "zero-install").
- **poc-planner disclaimed the per-step duration instead of omitting it** (`poc-planner-how-long/no-duration`, run
after the demo fix): "…mapping a single data source 'usually takes less than 30 minutes' — a per-step note, not an
estimate for the POC." The decline rule now says disclaiming a figure states it, and to say nothing about it.
- **build stopped at doctor's report** (PR #63 eval, case `build`: 6 graders failed — no `generate_scaffold`, no
`senzing_search.py`; the run also asked the grounder sub-agent to "run doctor"). The skill now says doctor's
report is a checkpoint, invoked with the Skill tool, and the run continues to step 2 in the same turn.
- `scripts/check.sh`: probes for the first three.

## [1.37.19-1] - 2026-10-04

Plugin-only patch on server 1.37.19 (skills, graders, validator; no server change).
Expand Down
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
type: regex
pattern: '\b(?:or|rather|skip)\b[^\n]{0,80}\bzero-install\b|\b(?:I|we)(?:\s+can|\s+could|\s+would|''d)\s+(?:instead\s+|also\s+)?fall\s+back\s+to\b[^\n]{0,60}\bzero-install\b|\bzero-install\b[^\n]{0,60}\binstead\b|(?:\d\.|[A-C]\))\s*\*\*\s*zero-install\b|\bskip\s+(?:the\s+)?(?:local\s+)?install(?:ing|ation)?\b'
pattern: '\b(?:or|rather|skip)\b[^\n]{0,80}\bzero-install\b|\b(?:I|we)(?:\s+can|\s+could|\s+would|''d)\s+(?:instead\s+|also\s+)?fall\s+back\s+to\b[^\n]{0,60}\bzero-install\b|\bzero-install\b[^\n]{0,60}\binstead\b|(?:\d\.|[A-C]\))\s*\*\*\s*zero-install\b|\bskip\s+(?:the\s+)?(?:local\s+)?install(?:ing|ation)?\b|\b(?:if|should|unless)\b[^\n.]{0,100}\bfall\s+back\s+to\s+showing\b'
match: not_contains
flags: i
target: last_message
Expand Down Expand Up @@ -29,3 +29,11 @@ A bare `fall back to ... zero-install` is NOT matched: a correct run that took t
"I could", "I would", "I'd" — including the conditional attached to the permission request that recurred after the
first rule ("or, if you'd rather not change the sandbox rules, I can instead fall back to a zero-install preview").
Measured on 82 real arms: 8 hits, every one a judge FAIL, none on a pass.

## Conditional fallback without the word "zero-install"

CI (PR #63 head `34d1a83`, demo-no-simulation) failed 3-0 on "If either can't be granted here, say so and I'll
fall back to showing you validated, Senzing-ready sample data without live resolution … — but I won't do that
silently; I wanted to ask first." No `zero-install` token, no "can/could" modal — the same conditional attached
to the permission request. `if/should/unless … fall back to showing` is the offer; a run that TOOK the tier
after install ended writes "I fell back" or states it as done, and has no leading conditional.
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,9 @@ must_match:
preview**: pull sample data'
- in which case I'd fall back to a zero-install preview (validated sample data only, no simulated resolution)
or Docker.
- 'If either can''t be granted here, say so and I''ll fall back to showing you validated, Senzing-ready sample data
without live resolution (clearly labeled as a preview, not a demo of the engine) — but I won''t do that silently;
I wanted to ask first.'
must_not_match:
- 'Senzing isn''t installed yet. I''ll set it up using the official Homebrew cask (macOS Apple Silicon),
with **TypeScript/Node.js** as the language — it''s the only toolchain already available on this machine
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,8 @@
# First 8 must_match + 3 must_not_match: REAL final messages from CI eval runs (not paraphrases).
# The rest are HAND-WRITTEN, one per duration unit, so dropping a unit from the pattern is caught.
must_match:
- 'The one place it mentions a span at all is that mapping a single data source "usually takes less than 30 minutes" —
a per-step note about one activity, not an estimate for the POC.'
- The article also says data mapping itself "usually takes less than 30 minutes per data source" (per
source_url above) — but that's one step, not the whole POC.
- 'One more data point worth having in view before you size anything: mapping a single data source typically
Expand Down
2 changes: 1 addition & 1 deletion plugins/senzing/.claude-plugin/plugin.json
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@
"name": "senzing",
"displayName": "Senzing",
"icon": "./icon.png",
"version": "1.37.19-1",
"version": "1.37.19-2",
"description": "Entity resolution: point Claude at your data, it does the work, and you have results today.",
"author": {
"name": "Senzing, Inc.",
Expand Down
8 changes: 8 additions & 0 deletions plugins/senzing/skills/build/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,6 +58,14 @@ answers the run gate; the write gate needs one probe of its own:
`/senzing:build` in **Claude Code** on that host.
Never claim to have edited files you could not write, and never claim a run you did not perform.

**`doctor`'s report is a checkpoint, not the answer.** Invoke it with the Skill tool itself — never hand
"run doctor" to a sub-agent (the grounder agent has no doctor and answers that none exists). Whatever its
verdict — no Senzing, Python on macOS, nothing installed — carry on **in the same turn** to step 2: the file
is still the deliverable, written by the file tools whether or not Senzing exists on the host. Only step 4
(run) waits on the verdict. CI caught a run that printed doctor's table as its final message and never called
`generate_scaffold` or wrote `senzing_search.py`; a message that is only an environment report has not done
what was asked.

Always:

1. **Inputs.** `$ARGUMENTS` may name the language and/or workflow (e.g. `python search`). Determine
Expand Down
11 changes: 9 additions & 2 deletions plugins/senzing/skills/demo/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -120,7 +120,11 @@ second report.
preview" is an alternative to installing, and it turns the permission question into a fork. If the
user declines the permission, THEN you take the tier, in the next turn, as what you do now — never as a
way out you offered this turn. The tier is not mentioned until `install` has actually ended its turn
without a working SDK, and then you take it rather than announce it.
without a working SDK, and then you take it rather than announce it. **Saying nothing here is the honest
course, not a withheld one**: the tier is always taken, and always labelled "not a demo of the engine", the
moment `install` ends without an SDK — so "if that can't be granted I'll fall back to showing you sample data
— I won't do that silently" (CI caught exactly that sentence) discloses nothing the user needs now and hands
them the fork. Ask the EULA and permission questions, end the message there.

**The zero-install tier is what you do AFTER `install` has ended its turn without a working
SDK** — a fallback you take, not an option you put to the user, and never a branch offered in
Expand Down Expand Up @@ -165,7 +169,10 @@ second report.
needed.** The scratch repo is throwaway and touches nothing of theirs. **Never offer, hint at
or ask about loading into their existing repository** — not "if you'd rather load it into
production, say so", not "I'll only do that if you explicitly ask". Saying production is
untouched is fine; saying it is *available* is an offer. Only if the USER raises their
untouched is fine; saying it is *available* is an offer. The whole sentence is "Your
production repository is never touched." — full stop; no "unless", "only if", "until you
ask" after it (CI caught "…is never touched, and I won't load into it unless you ask for
that explicitly"). Only if the USER raises their
existing repository themselves, confirm the target and the record count first.
Verify the load as `analyze` step 4 does (loaded vs submitted, error count) before going on.
- **Drain the redo queue** as `analyze` step 5 does — get the probe from the MCP, drain to 0,
Expand Down
27 changes: 19 additions & 8 deletions plugins/senzing/skills/poc-planner/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -53,6 +53,10 @@ license terms here or from memory — call the tool named in each step and cite
offered as what you are *not* saying. **A cited quote is not an exception here.** The guidance
does carry one per-step figure — mapping a data source — and offering it as "one cited data
point" answers "how long" with a partial estimate, the same fabrication wearing a citation.
**Disclaiming it is still stating it**: "the one place it mentions a span is that mapping a source 'usually
takes less than 30 minutes' — a per-step note, not an estimate for the POC" puts the 30 minutes in front of the
reader exactly as the cited-data-point form does (CI caught that sentence). Say nothing about that figure at
all — not as a data point, not as a caveat, not as what the guidance does or doesn't give.
**Retrieve first.** "The guidance carries no duration for the POC as a whole" is a claim about Senzing's guidance:
make it only after the `search_docs` retrieval at the top of this skill, and name the article you retrieved. A
decline written from this paragraph alone is a Senzing claim with no source, about a document you have not opened — a
Expand Down Expand Up @@ -406,11 +410,14 @@ data_sources:
approx_records:
entity_types:
identifying_columns: []
database: # their words, e.g. PostgreSQL — or TBD — decided by <owner>
os_platform: # the user's words
platform_id: # an id from sdk_guide's platform tree, or TBD — decided by <owner>
# database, os_platform, languages: their words (e.g. PostgreSQL, ["Python"]), or the TBD literal
# platform_id: an id from sdk_guide's platform tree, or the TBD literal
# DELETE these comment lines, and never leave a `# …` after a value you fill in.
database:
os_platform:
platform_id:
cloud:
languages: [] # their words, e.g. ["Python"]
languages: []
hardware_available:
performance_required:
throughput:
Expand All @@ -420,13 +427,17 @@ calendar:
```
## 3. Success criteria
```yaml
# shape is one of: er_quality | functional_integration | entity_graph_scenario | other
# statement: what must be shown, per user · measurement: as the tool names it — source: <url>
# target: "per user: <their words>" or the TBD literal — these seven keys only
# DELETE these comment lines, and never leave a `# …` after a value you fill in.
- id: SC-1
shape: # er_quality | functional_integration | entity_graph_scenario | other
statement: # what must be shown, per user
measurement: # as the tool names it — source: <url>
shape:
statement:
measurement:
measured_against:
decided_by:
target: # "per user: <their words>" or "TBD — decided by <owner>" — these seven keys only
target:
```
## 4. What must be true to buy — goal and scope
## 5. Data selection checklist
Expand Down
6 changes: 6 additions & 0 deletions plugins/senzing/skills/poc-planner/validate_plan.py
Original file line number Diff line number Diff line change
Expand Up @@ -251,6 +251,12 @@ def check(text: str) -> list[str]:
m = TBD_RE.search(stripped)
if not m:
continue
if re.search(r"TBD — decided by [^#\n]*\s#\s", stripped) and line not in sec9:
problems.append(
f"line {i}: a template comment rides on a TBD line — '{stripped[:70]}'. Delete the `# …` "
f"after the owner: a candidate list there answers the question the line says is open."
)
continue
tail = (m.group("tail") or "").lstrip(":").strip()
if not tail:
continue
Expand Down
6 changes: 5 additions & 1 deletion plugins/senzing/skills/recipes/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -123,7 +123,11 @@ words. Match it against the catalog `id`s; on a fuzzy/multiple match, confirm wh
Say what the recipe needs (its record count against the unlicensed limit is a fact you may
state) and hand off to `install`.
- **Senzing can't deploy** → still identify and fetch the named recipe (step 2/3) so the user
learns what it needs, then hand off to the **`install`** skill without asking first (it
learns what it needs, **then** hand off to the **`install`** skill without asking first.
**Order matters: fetch BEFORE you invoke `install`** — `install` asks its question and ends
the turn, so a fetch you leave for afterwards is never made (CI caught a run that went
doctor → `install` and never opened the recipe it was asked to cook). Run the catalog `curl`,
match the id, fetch `recipes/<id>.md`, and only then invoke `install` (it
surfaces the license agreement, runs the official steps, and verifies with `doctor`; do not
route around it via `sdk_guide(topic="install")` directly — if an evaluation license is
needed, `submit_feedback(category='license_request')`'s description states the current
Expand Down
27 changes: 27 additions & 0 deletions scripts/check.sh
Original file line number Diff line number Diff line change
Expand Up @@ -332,6 +332,14 @@ then
else
bad "validate_plan.py path-builder regressed on a shape the fixtures do not exercise"
fi
awk '/^ shape: er_quality/ && !d { sub(/er_quality/, "TBD — decided by the data platform lead # er_quality | other"); d=1 } { print }' "$VP_OK" > "$vp_tmp/tbd-comment.md"
cmp -s "$VP_OK" "$vp_tmp/tbd-comment.md" && bad "8b tbd-comment mutation did not change the plan - the test below proves nothing"
VP_OUT_TC="$(python3 "$VP" "$vp_tmp/tbd-comment.md" 2>&1 || true)"
if [[ "$VP_OUT_TC" == *"template comment rides on a TBD line"* ]]; then
ok "a template '# candidate | list' comment after a TBD literal is rejected"
else
bad "validate_plan.py accepts a candidate-list comment riding on a TBD line"
fi
rm -rf "$vp_tmp"

echo; echo "== 8c. no tool_used grader declares max without min (impossible range) =="
Expand Down Expand Up @@ -729,6 +737,25 @@ else
bad "recipes-named criteria lost the install-message clause - the judge will fail the correct hand-off at random"
fi

# Main-branch CI (run 37196169047) failed three cases the PR run had passed -- each a real behavior the
# skill text left room for. Text guards keep the fix from being edited away.
if grep -q "fetch BEFORE you invoke" plugins/senzing/skills/recipes/SKILL.md; then
ok "recipes/SKILL.md says to fetch the recipe before invoking install (install ends the turn)"
else
bad "recipes/SKILL.md lost fetch-before-install - a run can hand off without ever opening the recipe"
fi
if grep -q 'is never touched." — full stop' plugins/senzing/skills/demo/SKILL.md; then
ok "demo/SKILL.md gives the exact production sentence and forbids a conditional after it"
else
bad "demo/SKILL.md lost the exact production sentence - the model recites 'unless you ask' again"
fi

if grep -qF "report is a checkpoint, not the answer" plugins/senzing/skills/build/SKILL.md; then
ok "build/SKILL.md says doctor's report is a checkpoint and the run continues in the same turn"
else
bad "build/SKILL.md lost the doctor-is-a-checkpoint rule - a run can end on the environment table"
fi

echo; echo "== 9. Eval scoring split (deterministic gate vs judge score) =="
# The suite's verdict is two independent gates, computed by evals/gate.py:
# deterministic graders must ALL pass in EVERY run (no averaging, no threshold), while the
Expand Down
Loading