Skip to content

Plan slice 005 — provenance-bound tool identity, against SEP-2640 read in full - #15

Open
HackTuah wants to merge 1 commit into
mainfrom
slice/005-tool-identity
Open

HackTuah wants to merge 1 commit into
mainfrom
slice/005-tool-identity

Conversation

@HackTuah

@HackTuah HackTuah commented Sep 8, 2026

Copy link
Copy Markdown
Member

PLAN only. No shipping code, no behaviour change.

SEP-2640 was read before the design, not after

Archived at slices/005-provenance-bound-tool-identity/logs/spec-sep-2640-skills.md by the command that fetched it — 644 lines — and re-measured at time of writing:

state=open  merged=False  mergeable_state=clean
Status: Accepted  Sponsor: @pja-ant

Its mergeable_state moved from blocked to clean while the survey that found it was still running. This slice treats it as landing.

2640 already states compound identity normatively — "the pair of the host's identity for the originating server and the skill's uri", "MUST NOT key any of these on the uri alone". The PLAN says outright that this slice does not get to claim the idea.

The three differences it does claim

  1. Tools, not skills. 2640 never touches Tool, tools/list or tools/call.
  2. On the wire, not a host-assigned label. 2640's origin half is the host's identity for the server — explicitly not serverInfo.name — and it never appears in any MCP message.
  3. The project's own security charter still lists "Tool identity across servers" as Open, no champion.

What happens to the claim if 2640 is extended to tools

Decided in the PLAN, before the code:

  • origin half stays a host-assigned label → the claim narrows to difference 2 and survives;
  • origin goes on the wire for tools → the claim is dead, the right response is to adopt the standard's shape, and it is recorded as retired rather than quietly restated;
  • either way the defect is unaffected — a newcomer named after an incumbent is indistinguishable today whatever any SEP says.

With a trigger rather than an intention: re-fetch 2640 and re-run the open-SEP sweep before this slice's PR.

Two dependencies, stated before the work

Verification

Files were git added before the gate ran, since an unadded file is invisible to git ls-files and the step would have printed pass over it:

reuse pass (235 tracked; 59 in scope, 57 headered + 2 sidecar; excluded 174 archive + 2 licence text)
Gate OK.  GATE_EXIT=0

SCR-254, provenance-bound tool identity. The number is 005 rather than the
"003" the issue used to carry: slice directories own one number sequence
assigned when a slice starts, 003 has shipped, and 005 is next free.

SEP-2640 IS READ AND ARCHIVED BEFORE THE DESIGN, not after it. Fetched by the
command that wrote logs/spec-sep-2640-skills.md, 644 lines, and re-measured:
state=open merged=False mergeable_state=clean, Status: Accepted,
Sponsor: @pja-ant. Its mergeable_state moved from blocked to clean while the
survey that found it was still running, so this slice treats it as landing.

2640 already states compound identity normatively -- "the pair of the host's
identity for the originating server and the skill's uri", "MUST NOT key any of
these on the uri alone". This slice does not get to claim the idea, and the
PLAN says so in those words rather than working around it.

The three differences it does claim, each checkable:

  1. tools, not skills -- 2640 never touches Tool, tools/list or tools/call
  2. on the wire, not a host-assigned label -- 2640's origin half is the
     HOST's identity for the server, explicitly not serverInfo.name, and it
     never appears in any MCP message
  3. the project's own security charter still lists "Tool identity across
     servers" as Open with no champion

And what happens to the claim if 2640 is later extended to tools, decided now
rather than under pressure: if the origin half stays a host-assigned label the
claim narrows to difference 2 and survives; if origin goes on the wire for
tools the claim is DEAD, the right response is to adopt the standard's shape,
and it is recorded as retired rather than quietly restated. A trigger, not a
vibe: re-fetch 2640 and re-run the open-SEP sweep before this slice's PR.

Two dependencies stated before the work rather than at the round that needs
them. This brief SUPERSEDES the earlier "do not spawn further subagents", per
PR #14, so the lanes may spawn. And SCR-289 blocks this slice if its record
rests on a survivor's status -- the harness can add a failure under load, which
is the direction that turns a survivor into a false kill.

No shipping code, no behaviour change.

Gate green with the new files IN the population -- added before the run, since
an unadded file is invisible to git ls-files and the step would have printed
pass over it:

    reuse pass (235 tracked; 59 in scope, 57 headered + 2 sidecar)
    Gate OK.  GATE_EXIT=0

Signed-off-by: Ayla Croft <aylacroft@proton.me>
@HackTuah

HackTuah commented Sep 8, 2026

Copy link
Copy Markdown
Member Author

Do not re-run the red check on this PR

The failing push run is deliberate standing evidence. A green re-run would destroy it. Owner ruling, 2026-09-08.

This branch changes a PLAN file and an archived specification page. No code, no test, no dependency. The pull_request run passed and the push run failed on the same commit:

run 34174978505  event=pull_request  success
run 34174975336  event=push          failure

test FAIL (exit 2) — Running ExUnit with seed: 70815, max_cases: 8
  1) HTTPBanditTest "a refusal over the adapter's drain cap now announces the close it always did"
     code: assert String.contains?(bytes, "HTTP/1.1 403")
  160 tests, 1 failure

Two runs of one commit, two answers. That is the finding, and it is filed as SCR-289.

Why re-running would be the wrong kind of tidy

A re-run would almost certainly go green, and the PR would look fine. That is how a flake becomes invisible — and this particular flake has a direction: it adds a failure, which is what turns a survivor into a false kill in a mutation table. It has already done that once in this repository, to Mc2 in slice 003 round 3, from this same component.

So the red is worth more than the green. It is the only artefact in the repository that demonstrates, on a commit containing no code, that "gate green" is currently a sample rather than a measurement.

What unblocks this PR

Not a re-run. SCR-289 lands first, as its own slice, ahead of this one:

  • remove the wall-clock window from the Bandit-backed test — no longer one candidate among several, it is the requirement;
  • demonstrate the rate before fixing, not a single transcript;
  • prove a known survivor still reports as a survivor under the conditions that produced the false kill;
  • demonstrate the fix at CI's max_cases: 8, where it was caught, not only at the development machine's 64;
  • commit the mutation harness, which currently lives in a session scratchpad — the 003 record cites $S/mut.sh <name>, a path no future reader can resolve, and a command that cannot be re-run is not a record.

This PR waits on that. Leave the check red.

HackTuah added a commit that referenced this pull request Sep 8, 2026
SCR-289. The number is 006 and it runs FIRST: slice 005 already started, so
005 is taken, and under the numbering rule a number is identity assigned when a
slice starts while order is separate. Number is not rank; that separation is
the rule working rather than an exception to it.

Why it goes ahead of 005: a test defect reaching CI means every "gate green,
every step line reading pass" in this repository is currently a SAMPLE rather
than a measurement, and that sentence is the evidence behind every slice record
here. 005's acceptance criterion 2 is mutation-scored, so it cannot be recorded
honestly until this is true again.

READ FROM THE CODE RATHER THAN FROM THE ISSUE. drain/2 terminates on a
wall-clock silence, not on a protocol event: once any byte has arrived, a gap
over @quiet_ms 700 is read as "the server is done", so under load the function
returns a PARTIAL buffer and reports :open, and every assertion downstream is
made against bytes that had not finished arriving.

This is already a repair of a previous instance of itself. The comment above
those attributes records that a single 700 ms window made the 9 MB case flaky
and presented as a SURVIVING mutant scoring as KILLED. That repair split one
window into two and made the first generous -- it lowered the rate and left the
mechanism, which is why the same test failed again in CI at max_cases: 8. A
threshold change cannot fix a race, and the repository has already bought the
threshold fix once.

Three observations, one mechanism, and the error always has the same direction
-- it ADDS a failure, which is the direction that makes a table read all-killed
and nobody investigate:

    003 round 3     dev, max_cases 64     Mc2, a survivor, scored KILLED
    003 re-score    40 back-to-back       M13rev read 2 (rec 1), Mr2 4 (rec 3)
    PR #15 push     CI, max_cases 8       test FAIL on a commit with NO code

The red is a RATE, not a transcript, measured at both max_cases profiles and
with N derived from the observed rate rather than chosen. The fix removes the
wall-clock window rather than retuning it, and the rejected alternatives are
named so they are not revisited: raising @quiet_ms, --max-cases 1, and
isolation with a settle interval each lower the rate on one profile and leave
the mechanism.

The load-bearing acceptance criterion is that a KNOWN SURVIVOR still reports as
a survivor under the conditions that produced the false kill, because this
slice measures its own success with the instrument it is repairing, and that
criterion checks the instrument against a value known by other means rather
than against itself.

Also in scope: commit the mutation harness under tools/. Slice 003's record
cites $S/mut.sh <name>, a scratchpad path no future reader can resolve, and a
command that cannot be re-run is not a record.

Not in scope: re-scoring slice 003, re-running PR #15's red check, and lib/.

No shipping code, no behaviour change.

    reuse pass (234 tracked; 59 in scope, 57 headered + 2 sidecar)
    Gate OK.  GATE_EXIT=0

Signed-off-by: Ayla Croft <aylacroft@proton.me>
HackTuah added a commit that referenced this pull request Sep 8, 2026
…ced on demand

SCR-289. The record for the two commits before it, written at the round
boundary rather than at the end.

Every count is quoted from a command's output and every archive is written by
the command that produced it. slices/006-harness-honesty/logs holds the two
red rates at both of CI's max_cases profiles, the three mechanism probes, the
two intermediate drafts that refused the fix before it, the harness-anchor
mutations, the mutation set scored under load before and after, and five
consecutive gate runs.

THE FINDING THAT MATTERS IS NOT THE FIX. Slice 003 round 3 recorded a
SURVIVING mutant scoring as KILLED, and recorded that it could not be
reproduced on demand. It can:

    tools/mutate.sh score, 4 passes per mutant, 32 busy loops on 32 cores,
    with the harness as it was --

    M2never  recorded 0 failures, a SURVIVOR   read  1 | 1 | 1 | 0
    Mc2      recorded 0 failures, a SURVIVOR   read  0 | 0 | 0 | 0

Sixteen of forty scored suites read one failure more than the mutant deserves,
and the row corrupted is a recorded survivor. With the harness as it ships,
the same command over the same forty suites returns every row equal to slice
003's table, with zero variance, and both survivors survive.

Slice 003's table is NOT re-scored and its records are NOT rewritten. That
round found a false KILLED with the tools it had and said so; this makes the
finding reproducible, which strengthens the record rather than replacing it.

Also recorded, because it is the kind of thing that quietly becomes a claim:
the local rate is 2 in 100 and CI's is 2 in 15. The disagreement is reported
rather than smoothed. Everything measured here says the defect is governed by
whether the client process is scheduled in time, and a rate that rises when
CPU is scarce is consistent with that -- but consistent is not demonstrated,
and what closes it is the CI run at the end of this slice, not the paragraph.

PR #15's and PR #17's red push runs were not re-run. lib/ is untouched.

Signed-off-by: Ayla Croft <aylacroft@proton.me>
HackTuah added a commit that referenced this pull request Sep 8, 2026
SCR-289. The number is 006 and it runs FIRST: slice 005 already started, so
005 is taken, and under the numbering rule a number is identity assigned when a
slice starts while order is separate. Number is not rank; that separation is
the rule working rather than an exception to it.

Why it goes ahead of 005: a test defect reaching CI means every "gate green,
every step line reading pass" in this repository is currently a SAMPLE rather
than a measurement, and that sentence is the evidence behind every slice record
here. 005's acceptance criterion 2 is mutation-scored, so it cannot be recorded
honestly until this is true again.

READ FROM THE CODE RATHER THAN FROM THE ISSUE. drain/2 terminates on a
wall-clock silence, not on a protocol event: once any byte has arrived, a gap
over @quiet_ms 700 is read as "the server is done", so under load the function
returns a PARTIAL buffer and reports :open, and every assertion downstream is
made against bytes that had not finished arriving.

This is already a repair of a previous instance of itself. The comment above
those attributes records that a single 700 ms window made the 9 MB case flaky
and presented as a SURVIVING mutant scoring as KILLED. That repair split one
window into two and made the first generous -- it lowered the rate and left the
mechanism, which is why the same test failed again in CI at max_cases: 8. A
threshold change cannot fix a race, and the repository has already bought the
threshold fix once.

Three observations, one mechanism, and the error always has the same direction
-- it ADDS a failure, which is the direction that makes a table read all-killed
and nobody investigate:

    003 round 3     dev, max_cases 64     Mc2, a survivor, scored KILLED
    003 re-score    40 back-to-back       M13rev read 2 (rec 1), Mr2 4 (rec 3)
    PR #15 push     CI, max_cases 8       test FAIL on a commit with NO code

The red is a RATE, not a transcript, measured at both max_cases profiles and
with N derived from the observed rate rather than chosen. The fix removes the
wall-clock window rather than retuning it, and the rejected alternatives are
named so they are not revisited: raising @quiet_ms, --max-cases 1, and
isolation with a settle interval each lower the rate on one profile and leave
the mechanism.

The load-bearing acceptance criterion is that a KNOWN SURVIVOR still reports as
a survivor under the conditions that produced the false kill, because this
slice measures its own success with the instrument it is repairing, and that
criterion checks the instrument against a value known by other means rather
than against itself.

Also in scope: commit the mutation harness under tools/. Slice 003's record
cites $S/mut.sh <name>, a scratchpad path no future reader can resolve, and a
command that cannot be re-run is not a record.

Not in scope: re-scoring slice 003, re-running PR #15's red check, and lib/.

No shipping code, no behaviour change.

    reuse pass (234 tracked; 59 in scope, 57 headered + 2 sidecar)
    Gate OK.  GATE_EXIT=0

Signed-off-by: Ayla Croft <aylacroft@proton.me>
HackTuah added a commit that referenced this pull request Sep 8, 2026
…ced on demand

SCR-289. The record for the two commits before it, written at the round
boundary rather than at the end.

Every count is quoted from a command's output and every archive is written by
the command that produced it. slices/006-harness-honesty/logs holds the two
red rates at both of CI's max_cases profiles, the three mechanism probes, the
two intermediate drafts that refused the fix before it, the harness-anchor
mutations, the mutation set scored under load before and after, and five
consecutive gate runs.

THE FINDING THAT MATTERS IS NOT THE FIX. Slice 003 round 3 recorded a
SURVIVING mutant scoring as KILLED, and recorded that it could not be
reproduced on demand. It can:

    tools/mutate.sh score, 4 passes per mutant, 32 busy loops on 32 cores,
    with the harness as it was --

    M2never  recorded 0 failures, a SURVIVOR   read  1 | 1 | 1 | 0
    Mc2      recorded 0 failures, a SURVIVOR   read  0 | 0 | 0 | 0

Sixteen of forty scored suites read one failure more than the mutant deserves,
and the row corrupted is a recorded survivor. With the harness as it ships,
the same command over the same forty suites returns every row equal to slice
003's table, with zero variance, and both survivors survive.

Slice 003's table is NOT re-scored and its records are NOT rewritten. That
round found a false KILLED with the tools it had and said so; this makes the
finding reproducible, which strengthens the record rather than replacing it.

Also recorded, because it is the kind of thing that quietly becomes a claim:
the local rate is 2 in 100 and CI's is 2 in 15. The disagreement is reported
rather than smoothed. Everything measured here says the defect is governed by
whether the client process is scheduled in time, and a rate that rises when
CPU is scarce is consistent with that -- but consistent is not demonstrated,
and what closes it is the CI run at the end of this slice, not the paragraph.

PR #15's and PR #17's red push runs were not re-run. lib/ is untouched.

Signed-off-by: Ayla Croft <aylacroft@proton.me>
HackTuah added a commit that referenced this pull request Sep 8, 2026
SCR-289. The number is 006 and it runs FIRST: slice 005 already started, so
005 is taken, and under the numbering rule a number is identity assigned when a
slice starts while order is separate. Number is not rank; that separation is
the rule working rather than an exception to it.

Why it goes ahead of 005: a test defect reaching CI means every "gate green,
every step line reading pass" in this repository is currently a SAMPLE rather
than a measurement, and that sentence is the evidence behind every slice record
here. 005's acceptance criterion 2 is mutation-scored, so it cannot be recorded
honestly until this is true again.

READ FROM THE CODE RATHER THAN FROM THE ISSUE. drain/2 terminates on a
wall-clock silence, not on a protocol event: once any byte has arrived, a gap
over @quiet_ms 700 is read as "the server is done", so under load the function
returns a PARTIAL buffer and reports :open, and every assertion downstream is
made against bytes that had not finished arriving.

This is already a repair of a previous instance of itself. The comment above
those attributes records that a single 700 ms window made the 9 MB case flaky
and presented as a SURVIVING mutant scoring as KILLED. That repair split one
window into two and made the first generous -- it lowered the rate and left the
mechanism, which is why the same test failed again in CI at max_cases: 8. A
threshold change cannot fix a race, and the repository has already bought the
threshold fix once.

Three observations, one mechanism, and the error always has the same direction
-- it ADDS a failure, which is the direction that makes a table read all-killed
and nobody investigate:

    003 round 3     dev, max_cases 64     Mc2, a survivor, scored KILLED
    003 re-score    40 back-to-back       M13rev read 2 (rec 1), Mr2 4 (rec 3)
    PR #15 push     CI, max_cases 8       test FAIL on a commit with NO code

The red is a RATE, not a transcript, measured at both max_cases profiles and
with N derived from the observed rate rather than chosen. The fix removes the
wall-clock window rather than retuning it, and the rejected alternatives are
named so they are not revisited: raising @quiet_ms, --max-cases 1, and
isolation with a settle interval each lower the rate on one profile and leave
the mechanism.

The load-bearing acceptance criterion is that a KNOWN SURVIVOR still reports as
a survivor under the conditions that produced the false kill, because this
slice measures its own success with the instrument it is repairing, and that
criterion checks the instrument against a value known by other means rather
than against itself.

Also in scope: commit the mutation harness under tools/. Slice 003's record
cites $S/mut.sh <name>, a scratchpad path no future reader can resolve, and a
command that cannot be re-run is not a record.

Not in scope: re-scoring slice 003, re-running PR #15's red check, and lib/.

No shipping code, no behaviour change.

    reuse pass (234 tracked; 59 in scope, 57 headered + 2 sidecar)
    Gate OK.  GATE_EXIT=0

Signed-off-by: Ayla Croft <aylacroft@proton.me>
HackTuah added a commit that referenced this pull request Sep 8, 2026
…ced on demand

SCR-289. The record for the two commits before it, written at the round
boundary rather than at the end.

Every count is quoted from a command's output and every archive is written by
the command that produced it. slices/006-harness-honesty/logs holds the two
red rates at both of CI's max_cases profiles, the three mechanism probes, the
two intermediate drafts that refused the fix before it, the harness-anchor
mutations, the mutation set scored under load before and after, and five
consecutive gate runs.

THE FINDING THAT MATTERS IS NOT THE FIX. Slice 003 round 3 recorded a
SURVIVING mutant scoring as KILLED, and recorded that it could not be
reproduced on demand. It can:

    tools/mutate.sh score, 4 passes per mutant, 32 busy loops on 32 cores,
    with the harness as it was --

    M2never  recorded 0 failures, a SURVIVOR   read  1 | 1 | 1 | 0
    Mc2      recorded 0 failures, a SURVIVOR   read  0 | 0 | 0 | 0

Sixteen of forty scored suites read one failure more than the mutant deserves,
and the row corrupted is a recorded survivor. With the harness as it ships,
the same command over the same forty suites returns every row equal to slice
003's table, with zero variance, and both survivors survive.

Slice 003's table is NOT re-scored and its records are NOT rewritten. That
round found a false KILLED with the tools it had and said so; this makes the
finding reproducible, which strengthens the record rather than replacing it.

Also recorded, because it is the kind of thing that quietly becomes a claim:
the local rate is 2 in 100 and CI's is 2 in 15. The disagreement is reported
rather than smoothed. Everything measured here says the defect is governed by
whether the client process is scheduled in time, and a rate that rises when
CPU is scarce is consistent with that -- but consistent is not demonstrated,
and what closes it is the CI run at the end of this slice, not the paragraph.

PR #15's and PR #17's red push runs were not re-run. lib/ is untouched.

Signed-off-by: Ayla Croft <aylacroft@proton.me>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant