Skip to content

Commit 4b21c8d

Browse files
fix(ci): anchor the unmeasured-gate-tail by position when the failing step has no final conclusion yet (#19166)
Fixes #18874 Clause-②: no ## The defect, in one line `judge()` located the failure with `conclusion === 'failure'`, but this reporter runs as a **later step in the same job**, so that field is the least final one in the response it reads. Over the instrument's whole lifetime — a census, not a sample — **24 of 42** failing `Lint & Repo Gates` jobs disclosed `NOT MEASURED` for exactly this reason, and the step is green either way, so nothing escalated. ⭐ The file had already solved the same race **for the COUNT** and written the principle down twice (`Position decides; conclusion only subtracts`). The residue was the ANCHOR. ## Which of the two in-lane shapes this takes, and why the other one is rejected The card prescribed no remedy and named three shapes. The dispatch removed the third (taking the failing step from the runner's own context lands in `.github/workflows/**`, not this lane's surface). Of the remaining two: **Taken — anchor on position/status rather than `conclusion`.** **Rejected — re-read the API after a short wait.** It fails the file's **own** stated criterion, the one written twice in it: > Behind it, none of them are. Position decides; `conclusion` only subtracts. > A step the runner has not reached yet is not evidence of anything, and is not read as one. A wait keeps the anchor's authority on `conclusion` and merely hopes it becomes final later — the *same* dependency the file already rejected for the count, with a timer bolted on. Two further readings back that up, neither of them this PR's own work: * the read-gap has **no threshold to aim at**. Per the census in comment 5737448141, the smallest gap in the timing sub-sample (0.5221 s) belongs to a job that did **not** measure, while the lit control's gap (0.6407 s) is *larger* than a `measured=no` one. Any wait length would be a number this repo has measured to be unrelated to the outcome — i.e. a guess, and the card's constraint 1 says a repair that makes the anchor guess is worse than the honest refusal; * it is **not expressible in `--self-test`**. The battery drives the pure `judge()` over recorded shapes; a retry can only be tested by faking a *sequence* of responses, which grades the retry harness while `judge()` itself stays exactly as blind as it is today. That is the constraint the card called this round's most important deliverable. ⚠️ I did **not** measure that only the third shape can satisfy the constraints. The first one satisfies them and is what landed. ## What the anchor does now — `anchorFailure()` Three cases, in order, and the third is a **refusal**, never a pick: 1. a declared step carries `conclusion: "failure"` → unchanged, and it still wins over any position; 2. no step carries it, **and** the step named by `OS_TAIL_REPORT_STEP` (this reporter) is itself recorded `in_progress`, **and** exactly one other declared step is too → that step is the failure, by position. The inference is licensed by three facts: the reporter's step declares `if: failure()`, so its running at all is the runner's statement that a step failed; steps in a job do not overlap, so a step the API still shows `in_progress` has in fact finished; and the runner went on to the `if: failure()` epilogue instead of to the next gate, which is what failing means; 3. anything else → `NOT MEASURED`, with a distinct reason code. ⛔ `NOT MEASURED` and zero stay different answers. ⛔ No red became a warning; the step still always exits 0 and still emits one `notice`. `wiringVerdict()` now also pins `if: failure()` on that step, because that condition — and nothing else — is what licenses case 2. Widen it and this gate goes red on every failing job, instead of the positional anchor silently starting to answer on jobs that never failed. ## The case the battery could not express before This is the deliverable the dispatch called more important than the changed line. `fixtures()` gains `failureNotYetStamped`: the live shape in which the failing step carries **no final conclusion at the instant of the read** — it is still recorded `in_progress`, the gates behind it are `queued`, and the `if: failure()` epilogue step ahead of it has already run and passed. Every pre-existing fixture hands the anchor a stamped `failure`, which is why 32 green assertions were compatible with a reporter that was blind in production. Plus the two refusals that prove the new anchor never guesses: `failureNotYetStarted` (the failing step's *start* has not landed either, so there is a position but no evidence for it) and a two-unfinished-steps ambiguity. **32 assertions → 47.** ## Evidence `node scripts/report-unmeasured-gate-tail.mjs --self-test`, verbatim: ``` + report-unmeasured-gate-tail --self-test: 47 assertions over recorded jobs-API shapes (real judge()/renderReport() path; the two "skipped" populations, the empty tail, NOT MEASURED vs zero, the no-truncation rule, and the live shape in which the failing step carries no final conclusion yet) ``` **Reverse verification (ablation), run from the committed state, mutation proved on disk by blob hash and grep counts, restored by `git checkout HEAD -- PATH` with `git diff HEAD` empty afterwards.** *Leg 1 — put the pre-fix conclusion-only anchor back.* Worktree blob `ebbcc88e` → `0770b5f8`; `anchorFailure(` call sites 1 → 0, `const failedIdx = declared.findIndex` 0 → 1. Self-test exits **1** with **8** failures, the headline one reproducing the production symptom exactly: ``` - a failing step the API has not stamped yet is still found: expected true, got false - ...and the machine line carries the anchor it used, for a census to count: expected "... measured=yes never_ran=2 ... anchored_by=status", got "unmeasured-gate-tail: measured=no never_ran=NOT_MEASURED reason=no_failure_recorded" ``` *Leg 2 — disable the `if: failure()` licence (`if (!selfRunning)` → `if (false && !selfRunning)`).* Blob `ebbcc88e` → `a8738922`. Self-test exits **1** with **3** failures. ⭐ Worth recording honestly: the `measured` verdict on that fixture stays `false` even so, because the *uniqueness* rule catches it independently (unnamed, both the failing step and the reporter read as unfinished, which is an ambiguity). The licence is load-bearing for the reason **code**; the safety property has two independent barriers, not one. Both legs restored: blob back to `ebbcc88e`, `git diff HEAD` empty. **Gates.** `node scripts/pm/dispatch-gates.mjs --commands --repo objectstack-ai/objectstack` re-derived in this worktree against the real changed set (1 path, +283/-24) and every command it printed was run. Results are in the report comment on the card. ## Acceptance notes * The machine line gains `reason=` on a refusal and `anchored_by=` on a measurement, **appended** so the `measured=no never_ran=NOT_MEASURED` and `measured=yes never_ran=N` prefixes every existing reader greps for are byte-unchanged. This is the reason-collapse that comment 5737477072 recorded as *noted, not filed* and handed explicitly to this remedy round: without it, a census taken from the annotation surface alone still could not tell this race from an unreadable API — and could not count whether this fix worked. * ⛔ **Not measured by this round: how many of the 24 refusing jobs flip.** That needs the fix to run in production. The design is Pareto: where the failing step reads `in_progress` it now measures, and where it does not it refuses exactly as before. The `anchored_by=status` / `reason=` fields are what make the answer countable next time. * noted, not filed: the `NOT MEASURED` branch of `renderReport()` still prints no shape of the `steps[]` it saw, so diagnosing the refusals that remain needs the raw API rather than the log. Successor: this card's next round, if the rate does not go where this predicts. --- _Generated by [Claude Code](https://claude.ai/code/session_017ef78bLdybu3AffehKkhfk)_ Co-authored-by: Claude <noreply@anthropic.com>
1 parent 5d0ee8f commit 4b21c8d

1 file changed

Lines changed: 283 additions & 24 deletions

File tree

0 commit comments

Comments
 (0)