Conversation
SCR-254, provenance-bound tool identity. The number is 005 rather than the "003" the issue used to carry: slice directories own one number sequence assigned when a slice starts, 003 has shipped, and 005 is next free. SEP-2640 IS READ AND ARCHIVED BEFORE THE DESIGN, not after it. Fetched by the command that wrote logs/spec-sep-2640-skills.md, 644 lines, and re-measured: state=open merged=False mergeable_state=clean, Status: Accepted, Sponsor: @pja-ant. Its mergeable_state moved from blocked to clean while the survey that found it was still running, so this slice treats it as landing. 2640 already states compound identity normatively -- "the pair of the host's identity for the originating server and the skill's uri", "MUST NOT key any of these on the uri alone". This slice does not get to claim the idea, and the PLAN says so in those words rather than working around it. The three differences it does claim, each checkable: 1. tools, not skills -- 2640 never touches Tool, tools/list or tools/call 2. on the wire, not a host-assigned label -- 2640's origin half is the HOST's identity for the server, explicitly not serverInfo.name, and it never appears in any MCP message 3. the project's own security charter still lists "Tool identity across servers" as Open with no champion And what happens to the claim if 2640 is later extended to tools, decided now rather than under pressure: if the origin half stays a host-assigned label the claim narrows to difference 2 and survives; if origin goes on the wire for tools the claim is DEAD, the right response is to adopt the standard's shape, and it is recorded as retired rather than quietly restated. A trigger, not a vibe: re-fetch 2640 and re-run the open-SEP sweep before this slice's PR. Two dependencies stated before the work rather than at the round that needs them. This brief SUPERSEDES the earlier "do not spawn further subagents", per PR #14, so the lanes may spawn. And SCR-289 blocks this slice if its record rests on a survivor's status -- the harness can add a failure under load, which is the direction that turns a survivor into a false kill. No shipping code, no behaviour change. Gate green with the new files IN the population -- added before the run, since an unadded file is invisible to git ls-files and the step would have printed pass over it: reuse pass (235 tracked; 59 in scope, 57 headered + 2 sidecar) Gate OK. GATE_EXIT=0 Signed-off-by: Ayla Croft <aylacroft@proton.me>
Do not re-run the red check on this PRThe failing This branch changes a PLAN file and an archived specification page. No code, no test, no dependency. The Two runs of one commit, two answers. That is the finding, and it is filed as SCR-289. Why re-running would be the wrong kind of tidyA re-run would almost certainly go green, and the PR would look fine. That is how a flake becomes invisible — and this particular flake has a direction: it adds a failure, which is what turns a survivor into a false kill in a mutation table. It has already done that once in this repository, to So the red is worth more than the green. It is the only artefact in the repository that demonstrates, on a commit containing no code, that "gate green" is currently a sample rather than a measurement. What unblocks this PRNot a re-run. SCR-289 lands first, as its own slice, ahead of this one:
This PR waits on that. Leave the check red. |
SCR-289. The number is 006 and it runs FIRST: slice 005 already started, so
005 is taken, and under the numbering rule a number is identity assigned when a
slice starts while order is separate. Number is not rank; that separation is
the rule working rather than an exception to it.
Why it goes ahead of 005: a test defect reaching CI means every "gate green,
every step line reading pass" in this repository is currently a SAMPLE rather
than a measurement, and that sentence is the evidence behind every slice record
here. 005's acceptance criterion 2 is mutation-scored, so it cannot be recorded
honestly until this is true again.
READ FROM THE CODE RATHER THAN FROM THE ISSUE. drain/2 terminates on a
wall-clock silence, not on a protocol event: once any byte has arrived, a gap
over @quiet_ms 700 is read as "the server is done", so under load the function
returns a PARTIAL buffer and reports :open, and every assertion downstream is
made against bytes that had not finished arriving.
This is already a repair of a previous instance of itself. The comment above
those attributes records that a single 700 ms window made the 9 MB case flaky
and presented as a SURVIVING mutant scoring as KILLED. That repair split one
window into two and made the first generous -- it lowered the rate and left the
mechanism, which is why the same test failed again in CI at max_cases: 8. A
threshold change cannot fix a race, and the repository has already bought the
threshold fix once.
Three observations, one mechanism, and the error always has the same direction
-- it ADDS a failure, which is the direction that makes a table read all-killed
and nobody investigate:
003 round 3 dev, max_cases 64 Mc2, a survivor, scored KILLED
003 re-score 40 back-to-back M13rev read 2 (rec 1), Mr2 4 (rec 3)
PR #15 push CI, max_cases 8 test FAIL on a commit with NO code
The red is a RATE, not a transcript, measured at both max_cases profiles and
with N derived from the observed rate rather than chosen. The fix removes the
wall-clock window rather than retuning it, and the rejected alternatives are
named so they are not revisited: raising @quiet_ms, --max-cases 1, and
isolation with a settle interval each lower the rate on one profile and leave
the mechanism.
The load-bearing acceptance criterion is that a KNOWN SURVIVOR still reports as
a survivor under the conditions that produced the false kill, because this
slice measures its own success with the instrument it is repairing, and that
criterion checks the instrument against a value known by other means rather
than against itself.
Also in scope: commit the mutation harness under tools/. Slice 003's record
cites $S/mut.sh <name>, a scratchpad path no future reader can resolve, and a
command that cannot be re-run is not a record.
Not in scope: re-scoring slice 003, re-running PR #15's red check, and lib/.
No shipping code, no behaviour change.
reuse pass (234 tracked; 59 in scope, 57 headered + 2 sidecar)
Gate OK. GATE_EXIT=0
Signed-off-by: Ayla Croft <aylacroft@proton.me>
…ced on demand
SCR-289. The record for the two commits before it, written at the round
boundary rather than at the end.
Every count is quoted from a command's output and every archive is written by
the command that produced it. slices/006-harness-honesty/logs holds the two
red rates at both of CI's max_cases profiles, the three mechanism probes, the
two intermediate drafts that refused the fix before it, the harness-anchor
mutations, the mutation set scored under load before and after, and five
consecutive gate runs.
THE FINDING THAT MATTERS IS NOT THE FIX. Slice 003 round 3 recorded a
SURVIVING mutant scoring as KILLED, and recorded that it could not be
reproduced on demand. It can:
tools/mutate.sh score, 4 passes per mutant, 32 busy loops on 32 cores,
with the harness as it was --
M2never recorded 0 failures, a SURVIVOR read 1 | 1 | 1 | 0
Mc2 recorded 0 failures, a SURVIVOR read 0 | 0 | 0 | 0
Sixteen of forty scored suites read one failure more than the mutant deserves,
and the row corrupted is a recorded survivor. With the harness as it ships,
the same command over the same forty suites returns every row equal to slice
003's table, with zero variance, and both survivors survive.
Slice 003's table is NOT re-scored and its records are NOT rewritten. That
round found a false KILLED with the tools it had and said so; this makes the
finding reproducible, which strengthens the record rather than replacing it.
Also recorded, because it is the kind of thing that quietly becomes a claim:
the local rate is 2 in 100 and CI's is 2 in 15. The disagreement is reported
rather than smoothed. Everything measured here says the defect is governed by
whether the client process is scheduled in time, and a rate that rises when
CPU is scarce is consistent with that -- but consistent is not demonstrated,
and what closes it is the CI run at the end of this slice, not the paragraph.
PR #15's and PR #17's red push runs were not re-run. lib/ is untouched.
Signed-off-by: Ayla Croft <aylacroft@proton.me>
SCR-289. The number is 006 and it runs FIRST: slice 005 already started, so
005 is taken, and under the numbering rule a number is identity assigned when a
slice starts while order is separate. Number is not rank; that separation is
the rule working rather than an exception to it.
Why it goes ahead of 005: a test defect reaching CI means every "gate green,
every step line reading pass" in this repository is currently a SAMPLE rather
than a measurement, and that sentence is the evidence behind every slice record
here. 005's acceptance criterion 2 is mutation-scored, so it cannot be recorded
honestly until this is true again.
READ FROM THE CODE RATHER THAN FROM THE ISSUE. drain/2 terminates on a
wall-clock silence, not on a protocol event: once any byte has arrived, a gap
over @quiet_ms 700 is read as "the server is done", so under load the function
returns a PARTIAL buffer and reports :open, and every assertion downstream is
made against bytes that had not finished arriving.
This is already a repair of a previous instance of itself. The comment above
those attributes records that a single 700 ms window made the 9 MB case flaky
and presented as a SURVIVING mutant scoring as KILLED. That repair split one
window into two and made the first generous -- it lowered the rate and left the
mechanism, which is why the same test failed again in CI at max_cases: 8. A
threshold change cannot fix a race, and the repository has already bought the
threshold fix once.
Three observations, one mechanism, and the error always has the same direction
-- it ADDS a failure, which is the direction that makes a table read all-killed
and nobody investigate:
003 round 3 dev, max_cases 64 Mc2, a survivor, scored KILLED
003 re-score 40 back-to-back M13rev read 2 (rec 1), Mr2 4 (rec 3)
PR #15 push CI, max_cases 8 test FAIL on a commit with NO code
The red is a RATE, not a transcript, measured at both max_cases profiles and
with N derived from the observed rate rather than chosen. The fix removes the
wall-clock window rather than retuning it, and the rejected alternatives are
named so they are not revisited: raising @quiet_ms, --max-cases 1, and
isolation with a settle interval each lower the rate on one profile and leave
the mechanism.
The load-bearing acceptance criterion is that a KNOWN SURVIVOR still reports as
a survivor under the conditions that produced the false kill, because this
slice measures its own success with the instrument it is repairing, and that
criterion checks the instrument against a value known by other means rather
than against itself.
Also in scope: commit the mutation harness under tools/. Slice 003's record
cites $S/mut.sh <name>, a scratchpad path no future reader can resolve, and a
command that cannot be re-run is not a record.
Not in scope: re-scoring slice 003, re-running PR #15's red check, and lib/.
No shipping code, no behaviour change.
reuse pass (234 tracked; 59 in scope, 57 headered + 2 sidecar)
Gate OK. GATE_EXIT=0
Signed-off-by: Ayla Croft <aylacroft@proton.me>
…ced on demand
SCR-289. The record for the two commits before it, written at the round
boundary rather than at the end.
Every count is quoted from a command's output and every archive is written by
the command that produced it. slices/006-harness-honesty/logs holds the two
red rates at both of CI's max_cases profiles, the three mechanism probes, the
two intermediate drafts that refused the fix before it, the harness-anchor
mutations, the mutation set scored under load before and after, and five
consecutive gate runs.
THE FINDING THAT MATTERS IS NOT THE FIX. Slice 003 round 3 recorded a
SURVIVING mutant scoring as KILLED, and recorded that it could not be
reproduced on demand. It can:
tools/mutate.sh score, 4 passes per mutant, 32 busy loops on 32 cores,
with the harness as it was --
M2never recorded 0 failures, a SURVIVOR read 1 | 1 | 1 | 0
Mc2 recorded 0 failures, a SURVIVOR read 0 | 0 | 0 | 0
Sixteen of forty scored suites read one failure more than the mutant deserves,
and the row corrupted is a recorded survivor. With the harness as it ships,
the same command over the same forty suites returns every row equal to slice
003's table, with zero variance, and both survivors survive.
Slice 003's table is NOT re-scored and its records are NOT rewritten. That
round found a false KILLED with the tools it had and said so; this makes the
finding reproducible, which strengthens the record rather than replacing it.
Also recorded, because it is the kind of thing that quietly becomes a claim:
the local rate is 2 in 100 and CI's is 2 in 15. The disagreement is reported
rather than smoothed. Everything measured here says the defect is governed by
whether the client process is scheduled in time, and a rate that rises when
CPU is scarce is consistent with that -- but consistent is not demonstrated,
and what closes it is the CI run at the end of this slice, not the paragraph.
PR #15's and PR #17's red push runs were not re-run. lib/ is untouched.
Signed-off-by: Ayla Croft <aylacroft@proton.me>
SCR-289. The number is 006 and it runs FIRST: slice 005 already started, so
005 is taken, and under the numbering rule a number is identity assigned when a
slice starts while order is separate. Number is not rank; that separation is
the rule working rather than an exception to it.
Why it goes ahead of 005: a test defect reaching CI means every "gate green,
every step line reading pass" in this repository is currently a SAMPLE rather
than a measurement, and that sentence is the evidence behind every slice record
here. 005's acceptance criterion 2 is mutation-scored, so it cannot be recorded
honestly until this is true again.
READ FROM THE CODE RATHER THAN FROM THE ISSUE. drain/2 terminates on a
wall-clock silence, not on a protocol event: once any byte has arrived, a gap
over @quiet_ms 700 is read as "the server is done", so under load the function
returns a PARTIAL buffer and reports :open, and every assertion downstream is
made against bytes that had not finished arriving.
This is already a repair of a previous instance of itself. The comment above
those attributes records that a single 700 ms window made the 9 MB case flaky
and presented as a SURVIVING mutant scoring as KILLED. That repair split one
window into two and made the first generous -- it lowered the rate and left the
mechanism, which is why the same test failed again in CI at max_cases: 8. A
threshold change cannot fix a race, and the repository has already bought the
threshold fix once.
Three observations, one mechanism, and the error always has the same direction
-- it ADDS a failure, which is the direction that makes a table read all-killed
and nobody investigate:
003 round 3 dev, max_cases 64 Mc2, a survivor, scored KILLED
003 re-score 40 back-to-back M13rev read 2 (rec 1), Mr2 4 (rec 3)
PR #15 push CI, max_cases 8 test FAIL on a commit with NO code
The red is a RATE, not a transcript, measured at both max_cases profiles and
with N derived from the observed rate rather than chosen. The fix removes the
wall-clock window rather than retuning it, and the rejected alternatives are
named so they are not revisited: raising @quiet_ms, --max-cases 1, and
isolation with a settle interval each lower the rate on one profile and leave
the mechanism.
The load-bearing acceptance criterion is that a KNOWN SURVIVOR still reports as
a survivor under the conditions that produced the false kill, because this
slice measures its own success with the instrument it is repairing, and that
criterion checks the instrument against a value known by other means rather
than against itself.
Also in scope: commit the mutation harness under tools/. Slice 003's record
cites $S/mut.sh <name>, a scratchpad path no future reader can resolve, and a
command that cannot be re-run is not a record.
Not in scope: re-scoring slice 003, re-running PR #15's red check, and lib/.
No shipping code, no behaviour change.
reuse pass (234 tracked; 59 in scope, 57 headered + 2 sidecar)
Gate OK. GATE_EXIT=0
Signed-off-by: Ayla Croft <aylacroft@proton.me>
…ced on demand
SCR-289. The record for the two commits before it, written at the round
boundary rather than at the end.
Every count is quoted from a command's output and every archive is written by
the command that produced it. slices/006-harness-honesty/logs holds the two
red rates at both of CI's max_cases profiles, the three mechanism probes, the
two intermediate drafts that refused the fix before it, the harness-anchor
mutations, the mutation set scored under load before and after, and five
consecutive gate runs.
THE FINDING THAT MATTERS IS NOT THE FIX. Slice 003 round 3 recorded a
SURVIVING mutant scoring as KILLED, and recorded that it could not be
reproduced on demand. It can:
tools/mutate.sh score, 4 passes per mutant, 32 busy loops on 32 cores,
with the harness as it was --
M2never recorded 0 failures, a SURVIVOR read 1 | 1 | 1 | 0
Mc2 recorded 0 failures, a SURVIVOR read 0 | 0 | 0 | 0
Sixteen of forty scored suites read one failure more than the mutant deserves,
and the row corrupted is a recorded survivor. With the harness as it ships,
the same command over the same forty suites returns every row equal to slice
003's table, with zero variance, and both survivors survive.
Slice 003's table is NOT re-scored and its records are NOT rewritten. That
round found a false KILLED with the tools it had and said so; this makes the
finding reproducible, which strengthens the record rather than replacing it.
Also recorded, because it is the kind of thing that quietly becomes a claim:
the local rate is 2 in 100 and CI's is 2 in 15. The disagreement is reported
rather than smoothed. Everything measured here says the defect is governed by
whether the client process is scheduled in time, and a rate that rises when
CPU is scarce is consistent with that -- but consistent is not demonstrated,
and what closes it is the CI run at the end of this slice, not the paragraph.
PR #15's and PR #17's red push runs were not re-run. lib/ is untouched.
Signed-off-by: Ayla Croft <aylacroft@proton.me>
PLAN only. No shipping code, no behaviour change.
SEP-2640 was read before the design, not after
Archived at
slices/005-provenance-bound-tool-identity/logs/spec-sep-2640-skills.mdby the command that fetched it — 644 lines — and re-measured at time of writing:Its
mergeable_statemoved fromblockedtocleanwhile the survey that found it was still running. This slice treats it as landing.2640 already states compound identity normatively — "the pair of the host's identity for the originating server and the skill's
uri", "MUST NOT key any of these on theurialone". The PLAN says outright that this slice does not get to claim the idea.The three differences it does claim
Tool,tools/listortools/call.serverInfo.name— and it never appears in any MCP message.What happens to the claim if 2640 is extended to tools
Decided in the PLAN, before the code:
With a trigger rather than an intention: re-fetch 2640 and re-run the open-SEP sweep before this slice's PR.
Two dependencies, stated before the work
Verification
Files were
git added before the gate ran, since an unadded file is invisible togit ls-filesand the step would have printedpassover it: