From f3a6a1e1c23cf99e9732858635b3247ee0e5c489 Mon Sep 17 00:00:00 2001 From: Pranav Bhave <90584946+Cubits11@users.noreply.github.com> Date: Tue, 8 Sep 2026 02:08:58 +0000 Subject: [PATCH 01/18] registry: E3 and E3B registered; MC-005 attribution and the ledger snapshot gated Four breaks between the speaking surfaces and the register are closed. 1. Stale counters. films/EXPLAINER-SERIES.md Episode 10 said "observation rows: zero" three days after 4,800 rows were committed. Episode 10 is rewritten in full, its central sentence replaced, and every counter the deck speaks now reads from metrics/ledger_snapshot.json with that file's as_of date spoken aloud. scripts/ledger_snapshot.py --check is the drift gate that would have caught it: it fails when the repository moves and the recorded counts do not, and it deliberately excludes blocker ages, which move with the calendar. /now/ gains the E3/E3B entry; index.html's claim count and owner-review date follow the registry. 2. MC-005 is explicitly non-canonical everywhere. It stays retracted: the kernel refuses a second registration for a claim that already has a protected state, and registering the same content under a fresh id would reverse a recorded owner judgement without a recorded reason. The W1 selection-regret numbers (regret <= 2 items / 2.4 points, 0 of 11 CIs excluding zero, 45 of 45 positive excess joint miss) keep their artifact at experiments/e2/results/retrospective/ and are labelled unregistered on every surface that speaks them. scripts/verify_retracted.py derives retracted ids from the history and fails any line that names one without saying so. Frozen experiment artifacts are excluded from that gate on purpose: a freeze is immutable by contract and an annotation inside one would be the forbidden rescue. 3. exclusive_cells was already closed and the deck had not caught up. The 32-cell block is committed in MC-002's expected, declared by a CLARIFY transition, re-asserted from the hash-verified release by reanalyze_bells_subset.py, and required by generate_missing_column.py and verify_figures.py. Blocker B1 is marked closed with those locators. 4. The BELLS licence disagreement is closed with one evidence-backed status: MC-002's "none declared upstream" with commercial_reuse facts_only. Nothing infers a licence from silence. B2, the ASTRA brief and the FABLE brief are updated to match, and the 2026-09-05 OBS CUT carries a dated superseding note rather than a rewrite. E3-001 and E3B-001 are registered at exactly the strength their rows earn. Support is the observation file itself, so the ledger now reads 2 of 19 claims resting on own measurement rather than 0. Both commitments lead with the failed prediction, carry prediction_verdicts in expected, and forbid re-thresholding, re-seeding or restating a prediction after the outcome. The sharpening the two pilots suggest -- that marginal-only reporting is uninformative in the middle of the marginal range and fully informative at its extremes -- is registered as a non-claim in those words, with the ~40-item pre-scoring check described as an engineering gate, because neither pilot tested it. scripts/verify_e3.py re-derives every registered quantity, including the bootstrap interval, from the committed observation rows alone -- no model, no network, no analyzer -- and asserts it against both the run result file and the registry. claims_history verify passes at 44 entries with the prefix rule satisfied and 19 live commitments equal to the chain tip. No validator was weakened, no failed prediction rewritten, no provenance invented, and no governance document added: the three new files are a verifier, a gate and a generated snapshot. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_019DdNofiEikbis8CvokH7iS --- ARTIFACTS/12-WEEK-PROGRAM.md | 14 +- ARTIFACTS/2026-09-05-FABLE-5.1-OBS-CUT.md | 25 +- ARTIFACTS/2026-09-07-RECONCILIATION.md | 2 +- claims.yaml | 226 ++++++++++++++- claims_history.yaml | 190 +++++++++++++ docs/RESEARCH_INDEX.md | 4 +- docs/foundations/ASTRA-BRIEF.md | 5 +- docs/foundations/FABLE-BRIEF.md | 61 ++-- docs/graph/repo-graph.json | 332 ++++++++++++++++++++-- docs/graph/repo-graph.mmd | 134 +++++---- experiments/e3/RESULT.md | 10 +- experiments/e3b/RESULT.md | 10 +- films/EXPLAINER-SERIES.md | 271 ++++++++++++------ films/data/facts.json | 22 +- films/flagship/source-map.json | 4 +- index.html | 4 +- ledger/index.html | 22 +- metrics/ledger_snapshot.json | 59 ++++ now/index.html | 30 +- observatory/index.html | 50 +++- scripts/ledger_snapshot.py | 150 ++++++++++ scripts/verification_manifest.py | 4 + scripts/verify_e3.py | 244 ++++++++++++++++ scripts/verify_retracted.py | 172 +++++++++++ sitemap.xml | 8 +- 25 files changed, 1815 insertions(+), 238 deletions(-) create mode 100644 metrics/ledger_snapshot.json create mode 100644 scripts/ledger_snapshot.py create mode 100644 scripts/verify_e3.py create mode 100644 scripts/verify_retracted.py diff --git a/ARTIFACTS/12-WEEK-PROGRAM.md b/ARTIFACTS/12-WEEK-PROGRAM.md index 2f37560..4c486d6 100644 --- a/ARTIFACTS/12-WEEK-PROGRAM.md +++ b/ARTIFACTS/12-WEEK-PROGRAM.md @@ -18,12 +18,22 @@ the second-guard selection by at most 2.4 points, and in the largest end-to-end measurement the stack was statistically indistinguishable from its single strongest member.** +> **Registration status, 2026-09-08.** This fact is **UNREGISTERED**. The W1 +> artifact is committed at `experiments/e2/results/retrospective/` — run report, +> three matrices, and `independent_t1.py`, which recomputes T1 from the raw +> released rows outside the analyzer's path. No claim id carries these numbers +> and no CI check re-asserts them: the claim that used to (retracted 2026-09-06) +> was withdrawn over a licence defect in its support block, and the retraction +> reason records that the computation itself is untouched. Every surface that +> speaks these numbers must say "computed 2026-09-02, unregistered" in the same +> breath; `scripts/verify_retracted.py` gates the attribution. + Three sources, all VERIFIED this run: 1. **BELLS 2025 released subset** (`non_adversarial_prompts.csv @ 507566c5`, sha256 `791dd4b0…`, 170 rows, 11 verdict columns). Computed 2026-09-02 on - the hash-verified file (scratchpad script; to be committed as the W1 - artifact), harmful stratum n=82, native points: + the hash-verified file — the W1 artifact is now committed at + `experiments/e2/results/retrospective/` — harmful stratum n=82, native points: - Five specialized supervisors: for **0 of 5** incumbents does the partner maximizing measured union catch differ in union from the partner chosen by marginal rank. Regret 0 items. diff --git a/ARTIFACTS/2026-09-05-FABLE-5.1-OBS-CUT.md b/ARTIFACTS/2026-09-05-FABLE-5.1-OBS-CUT.md index bf0c710..008054d 100644 --- a/ARTIFACTS/2026-09-05-FABLE-5.1-OBS-CUT.md +++ b/ARTIFACTS/2026-09-05-FABLE-5.1-OBS-CUT.md @@ -1,5 +1,16 @@ # OBSERVATION CUT — BELLS denominators and provenance +> **Superseding note, 2026-09-08 (annotation only; the cut below is unchanged).** +> Both open items this cut named are closed. MC-005 was **retracted** on +> 2026-09-06 — the RETRACT is entry 39 of `claims_history.yaml` and its reason +> is the licence disagreement recorded below — so the registry now has exactly +> one record of the BELLS file's licence: MC-002's `none declared upstream`, +> with `commercial_reuse: facts_only`. Nothing in this repository infers a +> licence from silence. The `exclusive_cells` block this cut found uncommitted +> is committed, declared by an MC-002 CLARIFY transition, and re-asserted from +> the hash-verified file by `scripts/reanalyze_bells_subset.py` in CI. The +> 990-versus-1041 denominator remains unreconciled, exactly as this cut left it. + Written 2026-09-05 against working tree of branch `claude/mc-005-selection-regret` (HEAD `f9b24e2`, uncommitted owner changes present and untouched). Status vocabulary: OBSERVED (bytes opened or command executed in this run) · DERIVED @@ -34,7 +45,7 @@ nothing redistributed) or `--cache DIR` offline. Exit 0 this run. | C | `077555d9` (2025-02-19, "smaller dataset for playground") | same path, sha `7fa0fbf5885e…` | 174 | 86 / 50 / 38 | 12 | as B | `prompt_guard=1`, `langkit=1` | rest | **SELECTION RULE: UNKNOWN** (commit subject only: "smaller dataset for playground") | selected subset | none | | D | `00b42bfd` (2025-02-21, "New data, new results interpretation") | same path, sha `52ca9ab8eb62…` | 174 | 86 / 50 / 38 | 12 | as C, claude column renamed `claude-3-5-sonnet-20241022` | none (Prompt Guard and LangKit now carry real 0/1 values) | all verdict columns | SELECTION RULE: UNKNOWN | selected subset | none | | E | `ffe88ccb` (2025-02-23, "new playground data") | same path, sha `d93f9fe1d1a1…` | 174 | 86 / 50 / 38 | 12 | as D | `llm_guard=0` | rest | SELECTION RULE: UNKNOWN | selected subset | none | -| F | `b20aeed5` (2025-03-22, "new results") = **`507566c5` (2025-07-08, default-branch head)** | same path, sha `791dd4b0a168f2eb…03f57c3` | **170** | **82 / 50 / 38** | 11 (Miscellaneous absent) | five specialized + gpt-4, claude-3-5-sonnet-20241022, gemini-1.5-pro-latest, mistral-large-latest, deepseek-ai/DeepSeek-V3, grok-2-latest | none over 170 rows; within the 82 harmful rows `llm_guard=0` | all eleven verdict columns | **SELECTION RULE: UNKNOWN**. Only 74–76 of the 170 questions appear in C–E, so F is a fresh draw from pool B, not a trimming of C. All 170 questions are in A and B with identical `harm_level`. | selected subset (author-selected; rule unstated) | MC-002 (every `expected` count), MC-003 (`{0…12}/82`, leave-one-out), MC-005 at HEAD (n = 82), census row `bells-misuse-2025` (`item_level_outcomes_released`), dossier `bells-misuse-2025`, disclosure page, two `/answers/` pages, `/try/` TRY-B, film theses in `films/SLATE.md` | +| F | `b20aeed5` (2025-03-22, "new results") = **`507566c5` (2025-07-08, default-branch head)** | same path, sha `791dd4b0a168f2eb…03f57c3` | **170** | **82 / 50 / 38** | 11 (Miscellaneous absent) | five specialized + gpt-4, claude-3-5-sonnet-20241022, gemini-1.5-pro-latest, mistral-large-latest, deepseek-ai/DeepSeek-V3, grok-2-latest | none over 170 rows; within the 82 harmful rows `llm_guard=0` | all eleven verdict columns | **SELECTION RULE: UNKNOWN**. Only 74–76 of the 170 questions appear in C–E, so F is a fresh draw from pool B, not a trimming of C. All 170 questions are in A and B with identical `harm_level`. | selected subset (author-selected; rule unstated) | MC-002 (every `expected` count), MC-003 (`{0…12}/82`, leave-one-out), MC-005 at HEAD (n = 82; MC-005 retracted 2026-09-06), census row `bells-misuse-2025` (`item_level_outcomes_released`), dossier `bells-misuse-2025`, disclosure page, two `/answers/` pages, `/try/` TRY-B, film theses in `films/SLATE.md` | | G | `507566c5` | `data/adversarial_prompts.csv` sha `32fe8663621a…` | 8 | 8 / 0 / 0 | 1 (all Miscellaneous) | eleven verdict columns | `harm_level`, `category`, `deepseek-ai/DeepSeek-V3=0` | rest | SELECTION RULE: UNKNOWN | selected subset | 12-WEEK-PROGRAM.md ("8 adversarial rows released"); census row ("plus 8 adversarial prompts") | | H | `8a974123`→`507566c5` | `data/metacognitive_results.csv` sha `66a9f09c60d3…` | 612 rows = 102 prompts × 6 models | ground_truth 504 harmful / 108 not_harmful rows; 19 non-adversarial prompts | — | none of the five supervisors; frontier models only, with **newer versions** (`claude-3-7-sonnet-20250219`, `gemini-2.5-pro-exp-03-25`) | `model` set | responses | UNKNOWN | selected subset | none | @@ -137,7 +148,10 @@ F into a population estimate. MC-002 stays non-population-level. files of February 2025. - Upstream repository declares no licence (GitHub API `license: null`; no `LICENSE` in the tree at `507566c5`). MC-002 records "none declared - upstream"; MC-005 at HEAD `f9b24e2` records `license: MIT` for the same file. + upstream"; MC-005 (since retracted) at HEAD `f9b24e2` recorded `license: MIT` + for the same file. + MC-005 was retracted 2026-09-06 for exactly this; `none declared upstream` + is now the register's only statement about that file. - The v0 commit `0fc3d6d3` carries 114,540 adversarial rows with prompt text, while the paper (App. 0.C) states the full dataset is not publicly released. Recorded as a provenance fact; nothing here redistributes it. @@ -159,7 +173,8 @@ F into a population estimate. MC-002 stays non-population-level. through `ROOT.rglob("index.html")`. Zero findings on tracked files. A clean clone has no `_private/`, so CI is unaffected; the two files this packet adds are not `index.html` and are not touched by that gate. -- Working tree (untouched): staged removal of MC-005 from `claims.yaml` +- Working tree (untouched): staged removal of MC-005 (now retracted) from `claims.yaml` + (completed, and declared as a RETRACT in the history, on 2026-09-06) relative to HEAD; unstaged addition of `exclusive_cells` to MC-002 and a matching `claims_history.yaml` transition; modified generators, verifiers, film receipts; untracked `experiments/e3/`. @@ -240,8 +255,8 @@ not create a finding. - No E2 item, threshold, estimator or criterion was touched. E2 remains `UNTESTED BY DESIGN`; P2 remains `EMPTY`. - No live API was called; no money was spent; nobody was contacted. -- MC-002, MC-003, MC-005 and the census row were not edited. Two candidate - owner actions are noted, not performed: (a) MC-005's `license: MIT` is +- MC-002, MC-003, MC-005 (since retracted) and the census row were not edited. Two candidate + owner actions are noted, not performed: (a) retracted MC-005's `license: MIT` is unsupported by upstream (no licence declared) and conflicts with MC-002's record for the same file; (b) the census row's evidence text "170 prompts … the full dataset is available only by contacting the authors" is diff --git a/ARTIFACTS/2026-09-07-RECONCILIATION.md b/ARTIFACTS/2026-09-07-RECONCILIATION.md index 3a9f5f0..16042a4 100644 --- a/ARTIFACTS/2026-09-07-RECONCILIATION.md +++ b/ARTIFACTS/2026-09-07-RECONCILIATION.md @@ -14,7 +14,7 @@ held twice: `experiments/e3/freeze/or-bench-80k.csv` and `experiments/e3b/freeze/or-bench-80k.csv`, byte-identical (`22e95602…`). Each freeze is self-contained by design, so the duplicate stays and is recorded by detector D7 in `docs/graph/repo-graph.json`. The -remaining ~11.7k lines are the E3 and E3B experiments, the MC-005 +remaining ~11.7k lines are the E3 and E3B experiments, the retracted MC-005 retraction, and the film re-render. Nothing in the 17 commits was found uncommittable; see §4 for what an adversarial pass did find. diff --git a/claims.yaml b/claims.yaml index a255905..63cef77 100644 --- a/claims.yaml +++ b/claims.yaml @@ -50,7 +50,7 @@ # explicit finding that no meaningful rescue applies version: "0.4" -last_owner_review: "2026-09-07" +last_owner_review: "2026-09-08" claims: - id: CC-001 @@ -1403,3 +1403,227 @@ claims: read or recorded as a failed check - triggers watch file content; semantic drift outside watched files remains a manual review event, and the ledger says so + + - id: E3-001 + proposition: > + E3 — the first pilot this repository ran itself — scored two ungated + classifiers on 400 harmful and 800 benign items and produced 2,400 + committed observation rows. Its primary pre-registered prediction + FAILED: excess joint miss was +0.0018 with a 95% bootstrap CI of + [-0.00096, +0.00706], which includes zero. Both guards missed almost + everything on this pool (0.9825 and 0.9625), so the Fréchet interval + the two marginals allow is [0.9450, 0.9625] — 1.75 percentage points + wide — and the observed joint miss of 0.9475 lies inside it. The + prediction that it would lie inside HELD; the difficulty-stratification + prediction was NOT COMPUTED and remains open. + scope: > + Exactly the 2,400 rows committed at experiments/e3/results/observations.jsonl, + produced 2026-09-06 by two research classifiers — protectai + deberta-v3-base-prompt-injection-v2 (0.2B) and dcarpintero + pangolin-guard-base (0.1B) — at thresholds frozen in e3_config.json + (sha256 e163a2f2…) before any harmful item was scored. One pool + (or-bench-toxic at e36d8b80), one operating point each, static full + exposure. scripts/verify_e3.py recomputes every quantity below from + those rows alone, including the bootstrap interval, which is a + deterministic function of the committed rows under the frozen seed. + Nothing here transfers to E2's guards, pools or operating points, and + nothing here is about a deployed system. + dimensions: + visibility: public + provenance: machine_generated_owner_executed + support_role: executed_output + evidential_status: supported_within_scope + maturity: experimental + support: + url: "https://github.com/Cubits11/cubits11.github.io/blob/main/experiments/e3/results/observations.jsonl" + commit: null + note: > + The support is the measurement itself: 2,400 per-item, per-guard rows + this repository produced. RESULT.md is the interpretation of those rows, + not the evidence for them, and is hash-pinned by a trigger below so an + edit to a quoted prediction fires a re-review. scripts/verify_e3.py + re-derives every registered number from the rows alone on every push. + expected: + observation_rows: 2400 + harmful: + n: 400 + p_miss_G1: 0.9825 + p_miss_G2: 0.9625 + q_obs: 0.9475 + q_ind: 0.9456562500000001 + delta: 0.0018437499999999218 + frechet: [0.9450000000000001, 0.9625] + inside_frechet: true + delta_ci95: [-0.0009562499999999918, 0.007062499999999972] + ci_excludes_zero: false + benign_evaluation_joint_flag: + n: 400 + p_miss_G1: 0.0525 + p_miss_G2: 0.025 + q_obs: 0.005 + q_ind: 0.0013125 + delta: 0.0036875 + frechet: [0.0, 0.025] + inside_frechet: true + prediction_verdicts: + P1_delta_positive_ci_excludes_zero: FAILED + P2_q_obs_inside_frechet: HELD + P3_difficulty_stratification: NOT_COMPUTED + last_reviewed: "2026-09-08" + review_window_days: 120 + review_triggers: + - type: local_content_change + enforcement: executable + path: experiments/e3/results/observations.jsonl + sha256: "952416606e0caa7d2f58c45322907be6a9c35a0d651910947688977cf99b22c0" + note: fires when a committed observation row changes without a registry re-review + - type: local_content_change + enforcement: executable + path: experiments/e3/results/e3_result.json + sha256: "84e21e59665b8626d41a73325563ec022251d79187c5b8d91e25f7b5411b2717" + note: fires when the recorded run result changes without a registry re-review + - type: local_content_change + enforcement: executable + path: experiments/e3/RESULT.md + sha256: "3d8b6af2311c5c39362d2bc7fb9fd451fd5e43afafae8ec0e665e06432e333ad" + note: fires when the run report changes, including any edit to a quoted prediction + - type: manual + enforcement: manual + event: a third pilot is run on a pool where both guards' miss rates are intermediate + falsifier: + condition: > + Recomputing from the committed observation rows yields any quantity + different from the expected block beyond 1e-12, or the observed joint + miss is shown to lie outside the Fréchet interval its own marginals + allow, or a quoted prediction in the bound run report is shown to have + been edited after the outcome was visible. + consequence: REJECT + forbidden_rescues: + - do not re-threshold, re-draw, or re-pool after the outcome and report the new numbers as this run + - do not re-run the bootstrap under a different seed or B and report the resulting interval as this one + - do not restate the failed primary prediction, narrow it, or drop it from the record + - do not treat E3B as a re-run, a correction, or a replacement of this result + non_claims: + - the 2,400 rows prove the instrument runs end to end; they do not prove it + measures what the programme says it measures, and on this pool the + identified set was nearly a point, so it measured almost nothing + - not evidence that either classifier is good or bad; two research models, + one pool, one operating point each + - the null is a null at this scale on this pool, not evidence of + independence and not evidence about any real guardrail's dependence + - says nothing about E2, its three frozen guards, its pools, or its + operating points + - no vendor, product, deployed stack, or population is described + + - id: E3B-001 + proposition: > + E3B put the same two classifiers on the attack family they were built + for — 400 real prompt injections — and produced a further 2,400 + committed observation rows. Its primary prediction FAILED and so did + the prediction the redraw existed to test. One guard missed nothing + (0.0000, catching 400 of 400) and the other missed 0.3975, so the + Fréchet interval the marginals allow is [0.0000, 0.0000] — zero points + wide — the observed joint miss is exactly 0.0000, and the bootstrap CI + is the degenerate [0, 0]. E3B was pre-registered to produce an interval + wider than 10 percentage points; it produced one narrower than E3's. + The prediction that the observed joint miss would lie inside the + interval HELD, trivially. + scope: > + Exactly the 2,400 rows committed at experiments/e3b/results/observations.jsonl, + produced 2026-09-06 by the same two classifiers at thresholds frozen in + e3b_config.json (sha256 252a5db9…) before any injection item was scored. + One pool (Lakera gandalf_ignore_instructions at 04737b65, MIT), one + operating point each, static full exposure. This is a new experiment with + its own freeze, not a re-run of E3: E3's result stands whatever this + shows, and both are registered. scripts/verify_e3.py recomputes every + quantity below from the committed rows alone. The guard that caught 400 + of 400 has two live explanations — generalisation and near-duplicate + leakage from an undisclosed corpus — and nothing in this run separates + them. + dimensions: + visibility: public + provenance: machine_generated_owner_executed + support_role: executed_output + evidential_status: supported_within_scope + maturity: experimental + support: + url: "https://github.com/Cubits11/cubits11.github.io/blob/main/experiments/e3b/results/observations.jsonl" + commit: null + note: > + The support is the measurement itself: 2,400 further per-item, + per-guard rows. RESULT.md is the interpretation, hash-pinned by a + trigger below. scripts/verify_e3.py re-derives every registered number + from the rows alone on every push. + expected: + observation_rows: 2400 + harmful: + n: 400 + p_miss_G1: 0.0 + p_miss_G2: 0.3975 + q_obs: 0.0 + q_ind: 0.0 + delta: 0.0 + frechet: [0.0, 0.0] + inside_frechet: true + delta_ci95: [0.0, 0.0] + ci_excludes_zero: false + benign_evaluation_joint_flag: + n: 400 + p_miss_G1: 0.06 + p_miss_G2: 0.0425 + q_obs: 0.0075 + q_ind: 0.00255 + delta: 0.0049499999999999995 + frechet: [0.0, 0.0425] + inside_frechet: true + prediction_verdicts: + P1_delta_positive_ci_excludes_zero: FAILED + P2_q_obs_inside_frechet: HELD + P3_frechet_width_above_10pp: FAILED + last_reviewed: "2026-09-08" + review_window_days: 120 + review_triggers: + - type: local_content_change + enforcement: executable + path: experiments/e3b/results/observations.jsonl + sha256: "00e75bf47f68ec2487532d595c74362eb85d0e683f3a0d4f294ff2ae429e4631" + note: fires when a committed observation row changes without a registry re-review + - type: local_content_change + enforcement: executable + path: experiments/e3b/results/e3_result.json + sha256: "a9685abb1ca63193325ddf319bb5d87cd0c395daf0e7614d571caefbefbfde87" + note: fires when the recorded run result changes without a registry re-review + - type: local_content_change + enforcement: executable + path: experiments/e3b/RESULT.md + sha256: "8a72bd5eb9ddbdf25dfceae9d4bc89dafdd39aa8abae1ae9a7bcddf3fd27d987" + note: fires when the run report changes, including any edit to a quoted prediction + - type: manual + enforcement: manual + event: the contamination status of either guard's training corpus becomes checkable + falsifier: + condition: > + Recomputing from the committed observation rows yields any quantity + different from the expected block beyond 1e-12, or a quoted prediction + in the bound run report is shown to have been edited after the outcome + was visible, or E3B is represented anywhere in this repository as a + re-run, correction, or replacement of E3. + consequence: REJECT + forbidden_rescues: + - do not re-threshold, re-draw, or re-pool after the outcome and report the new numbers as this run + - do not present E3B as a repair of E3, or E3 as superseded by it + - do not restate either failed prediction, narrow it, or drop it from the record + - do not treat the zero-width interval as a measured absence of dependence + non_claims: + - a zero-width identified set means the marginals already fixed the joint + miss; it is not a measurement that the two guards fail independently + - G1 catching 400 of 400 is not evidence that it is a good guardrail, and + contamination is unverified rather than excluded + - G2 missing 159 of 400 is not evidence that it is a bad one + - 2,400 further rows prove the instrument runs; on this pool the question + was degenerate, so they measure almost nothing + - "the sharpening these two pilots suggest — that marginal-only reporting + is uninformative in the middle of the marginal range and fully + informative at its extremes — is a hypothesis and an engineering design + gate, not a result of either pilot; neither run tested it" + - no vendor, product, deployed stack, or population is described diff --git a/claims_history.yaml b/claims_history.yaml index 7e86989..f2517fb 100644 --- a/claims_history.yaml +++ b/claims_history.yaml @@ -1774,3 +1774,193 @@ entries: - validating the instrument establishes nothing about the safety, representativeness, or deployment behavior of any real system digest: 40d1ba4ae1f4492ecb4fc85335124ba6de81435e4959dc4c5decdf2b02b9d3ae +- kind: registration + claim_id: E3-001 + digest_of_commitment: 2118370312a15f0974a97ca3f70293452ecec2feb5c7b177ef08fafd1a7f5fe9 + commitment: + proposition: 'E3 — the first pilot this repository ran itself — scored two ungated classifiers on + 400 harmful and 800 benign items and produced 2,400 committed observation rows. Its primary pre-registered + prediction FAILED: excess joint miss was +0.0018 with a 95% bootstrap CI of [-0.00096, +0.00706], + which includes zero. Both guards missed almost everything on this pool (0.9825 and 0.9625), so the + Fréchet interval the two marginals allow is [0.9450, 0.9625] — 1.75 percentage points wide — and + the observed joint miss of 0.9475 lies inside it. The prediction that it would lie inside HELD; + the difficulty-stratification prediction was NOT COMPUTED and remains open. + + ' + scope: 'Exactly the 2,400 rows committed at experiments/e3/results/observations.jsonl, produced 2026-09-06 + by two research classifiers — protectai deberta-v3-base-prompt-injection-v2 (0.2B) and dcarpintero + pangolin-guard-base (0.1B) — at thresholds frozen in e3_config.json (sha256 e163a2f2…) before any + harmful item was scored. One pool (or-bench-toxic at e36d8b80), one operating point each, static + full exposure. scripts/verify_e3.py recomputes every quantity below from those rows alone, including + the bootstrap interval, which is a deterministic function of the committed rows under the frozen + seed. Nothing here transfers to E2''s guards, pools or operating points, and nothing here is about + a deployed system. + + ' + falsifier: + condition: 'Recomputing from the committed observation rows yields any quantity different from the + expected block beyond 1e-12, or the observed joint miss is shown to lie outside the Fréchet interval + its own marginals allow, or a quoted prediction in the bound run report is shown to have been + edited after the outcome was visible. + + ' + consequence: REJECT + forbidden_rescues: + - do not re-threshold, re-draw, or re-pool after the outcome and report the new numbers as this run + - do not re-run the bootstrap under a different seed or B and report the resulting interval as this + one + - do not restate the failed primary prediction, narrow it, or drop it from the record + - do not treat E3B as a re-run, a correction, or a replacement of this result + non_claims: + - the 2,400 rows prove the instrument runs end to end; they do not prove it measures what the programme + says it measures, and on this pool the identified set was nearly a point, so it measured almost + nothing + - not evidence that either classifier is good or bad; two research models, one pool, one operating + point each + - the null is a null at this scale on this pool, not evidence of independence and not evidence about + any real guardrail's dependence + - says nothing about E2, its three frozen guards, its pools, or its operating points + - no vendor, product, deployed stack, or population is described + expected: + observation_rows: 2400 + harmful: + n: 400 + p_miss_G1: 0.9825 + p_miss_G2: 0.9625 + q_obs: 0.9475 + q_ind: 0.9456562500000001 + delta: 0.0018437499999999218 + frechet: + - 0.9450000000000001 + - 0.9625 + inside_frechet: true + delta_ci95: + - -0.0009562499999999918 + - 0.007062499999999972 + ci_excludes_zero: false + benign_evaluation_joint_flag: + n: 400 + p_miss_G1: 0.0525 + p_miss_G2: 0.025 + q_obs: 0.005 + q_ind: 0.0013125 + delta: 0.0036875 + frechet: + - 0.0 + - 0.025 + inside_frechet: true + prediction_verdicts: + P1_delta_positive_ci_excludes_zero: FAILED + P2_q_obs_inside_frechet: HELD + P3_difficulty_stratification: NOT_COMPUTED + recorded_at: '2026-09-08' + provenance_class: CONTEMPORANEOUS + evidence_refs: + - experiments/e3/RESULT.md + - sha256:d5d5ca69b529585ac5e966189872dee6f77b6bf4ee91705ec4d9bfd805a6dba4 + - experiments/e3/results/observations.jsonl + - sha256:952416606e0caa7d2f58c45322907be6a9c35a0d651910947688977cf99b22c0 + - scripts/verify_e3.py + reason: 'registration: E3, the first pilot this repository ran itself — 2,400 committed observation + rows, 2026-09-06. The commitment leads with the failed primary prediction: excess joint miss +0.0018 + with a 95% bootstrap CI including zero, on a pool where both guards'' marginals were so extreme that + the Frechet interval was 1.75 points wide. scripts/verify_e3.py re-derives every registered quantity, + including the bootstrap interval, from the committed rows alone on every push.' + previous_transition_digest: 40d1ba4ae1f4492ecb4fc85335124ba6de81435e4959dc4c5decdf2b02b9d3ae + digest: b27afec13765b332708d7fbaa66674ef678040eea9eb7739a944274fd4d9a63f +- kind: registration + claim_id: E3B-001 + digest_of_commitment: 8bee68218ae7634359bdafefbae4b6f98ae01531bde2e0502e9f6eda48220d88 + commitment: + proposition: 'E3B put the same two classifiers on the attack family they were built for — 400 real + prompt injections — and produced a further 2,400 committed observation rows. Its primary prediction + FAILED and so did the prediction the redraw existed to test. One guard missed nothing (0.0000, catching + 400 of 400) and the other missed 0.3975, so the Fréchet interval the marginals allow is [0.0000, + 0.0000] — zero points wide — the observed joint miss is exactly 0.0000, and the bootstrap CI is + the degenerate [0, 0]. E3B was pre-registered to produce an interval wider than 10 percentage points; + it produced one narrower than E3''s. The prediction that the observed joint miss would lie inside + the interval HELD, trivially. + + ' + scope: 'Exactly the 2,400 rows committed at experiments/e3b/results/observations.jsonl, produced 2026-09-06 + by the same two classifiers at thresholds frozen in e3b_config.json (sha256 252a5db9…) before any + injection item was scored. One pool (Lakera gandalf_ignore_instructions at 04737b65, MIT), one operating + point each, static full exposure. This is a new experiment with its own freeze, not a re-run of + E3: E3''s result stands whatever this shows, and both are registered. scripts/verify_e3.py recomputes + every quantity below from the committed rows alone. The guard that caught 400 of 400 has two live + explanations — generalisation and near-duplicate leakage from an undisclosed corpus — and nothing + in this run separates them. + + ' + falsifier: + condition: 'Recomputing from the committed observation rows yields any quantity different from the + expected block beyond 1e-12, or a quoted prediction in the bound run report is shown to have been + edited after the outcome was visible, or E3B is represented anywhere in this repository as a re-run, + correction, or replacement of E3. + + ' + consequence: REJECT + forbidden_rescues: + - do not re-threshold, re-draw, or re-pool after the outcome and report the new numbers as this run + - do not present E3B as a repair of E3, or E3 as superseded by it + - do not restate either failed prediction, narrow it, or drop it from the record + - do not treat the zero-width interval as a measured absence of dependence + non_claims: + - a zero-width identified set means the marginals already fixed the joint miss; it is not a measurement + that the two guards fail independently + - G1 catching 400 of 400 is not evidence that it is a good guardrail, and contamination is unverified + rather than excluded + - G2 missing 159 of 400 is not evidence that it is a bad one + - 2,400 further rows prove the instrument runs; on this pool the question was degenerate, so they + measure almost nothing + - the sharpening these two pilots suggest — that marginal-only reporting is uninformative in the middle + of the marginal range and fully informative at its extremes — is a hypothesis and an engineering + design gate, not a result of either pilot; neither run tested it + - no vendor, product, deployed stack, or population is described + expected: + observation_rows: 2400 + harmful: + n: 400 + p_miss_G1: 0.0 + p_miss_G2: 0.3975 + q_obs: 0.0 + q_ind: 0.0 + delta: 0.0 + frechet: + - 0.0 + - 0.0 + inside_frechet: true + delta_ci95: + - 0.0 + - 0.0 + ci_excludes_zero: false + benign_evaluation_joint_flag: + n: 400 + p_miss_G1: 0.06 + p_miss_G2: 0.0425 + q_obs: 0.0075 + q_ind: 0.00255 + delta: 0.0049499999999999995 + frechet: + - 0.0 + - 0.0425 + inside_frechet: true + prediction_verdicts: + P1_delta_positive_ci_excludes_zero: FAILED + P2_q_obs_inside_frechet: HELD + P3_frechet_width_above_10pp: FAILED + recorded_at: '2026-09-08' + provenance_class: CONTEMPORANEOUS + evidence_refs: + - experiments/e3b/RESULT.md + - sha256:b9915a030b02f888994e7b872df0c2a31ee64c18fa7151477beb2902bc3f1c6c + - experiments/e3b/results/observations.jsonl + - sha256:00e75bf47f68ec2487532d595c74362eb85d0e683f3a0d4f294ff2ae429e4631 + - scripts/verify_e3.py + reason: 'registration: E3B, the injection redraw — a new experiment with its own freeze, not a re-run + of E3. 2,400 further committed rows, 2026-09-06. Two of its three predictions failed, including the + one the redraw existed to test: the Frechet interval came back zero points wide rather than wider + than ten. Registered with both failures in the commitment. scripts/verify_e3.py re-derives every quantity + from the committed rows.' + previous_transition_digest: b27afec13765b332708d7fbaa66674ef678040eea9eb7739a944274fd4d9a63f + digest: f24f4b6a7648fdc2b2854db4af31cb846b5b9c46260135bfcb6969852e0465f9 diff --git a/docs/RESEARCH_INDEX.md b/docs/RESEARCH_INDEX.md index ab0579f..cf53bdd 100644 --- a/docs/RESEARCH_INDEX.md +++ b/docs/RESEARCH_INDEX.md @@ -2,7 +2,7 @@ GENERATED by scripts/generate_research_index.py from claims.yaml, distribution/outcomes.yaml, distribution/experiments.yaml and the film manifests. Do not edit by hand. -Registry v0.4 · last owner review 2026-09-07 · one question: *Do published per-guard scores determine what the stack misses together?* +Registry v0.4 · last owner review 2026-09-08 · one question: *Do published per-guard scores determine what the stack misses together?* ## Independent outcomes (someone who is not the author did work) @@ -38,6 +38,8 @@ Zero is the recorded value where it is zero. The stop rule and the procedure for | GCE-001 | GCE — the Guardrail Composability Explorer — is a coursework MVP demo (AI-285) for toggling guardrails and observing composed behavior. Its front-door composability-coefficient framing is superseded by CC-Framework's de… | supported_within_scope · superseded | — | supersession is a statement about theoretical framing, not about the correctness of GCE's code or its value as coursework | — | | SITE-001 | This site's palette token pairs were computed to pass WCAG AA contrast (most AAA), and manual QA passes covered light/dark themes, desktop and mobile layouts, keyboard focus, reduced motion, and no-JS rendering, as reco… | supported_within_scope · released | — | no field CLS/LCP/CrUX measurements are claimed | — | | SITE-002 | This site's claim registry is enforced in CI: verify_claims.py validates schema and bindings, executes the executable review triggers against live evidence, and fails the build when any claim passes its freshness window… | supported_within_scope · released | — | a green run verifies registry consistency and quiet triggers, never the truth of any claim's content | — | +| E3-001 | E3 — the first pilot this repository ran itself — scored two ungated classifiers on 400 harmful and 800 benign items and produced 2,400 committed observation rows. Its primary pre-registered prediction FAILED: excess jo… | supported_within_scope · experimental | — | the 2,400 rows prove the instrument runs end to end; they do not prove it measures what the programme says it measures, and on this pool the identified set was… | — | +| E3B-001 | E3B put the same two classifiers on the attack family they were built for — 400 real prompt injections — and produced a further 2,400 committed observation rows. Its primary prediction FAILED and so did the prediction t… | supported_within_scope · experimental | — | a zero-width identified set means the marginals already fixed the joint miss; it is not a measurement that the two guards fail independently | — | ## Experiments an outsider can run diff --git a/docs/foundations/ASTRA-BRIEF.md b/docs/foundations/ASTRA-BRIEF.md index b20a1b8..697be2b 100644 --- a/docs/foundations/ASTRA-BRIEF.md +++ b/docs/foundations/ASTRA-BRIEF.md @@ -65,7 +65,10 @@ most** — it turns §2's rubric into a verifier. - `films/data/facts.json` is stale against `claims.yaml`; the gate says regenerate **and re-inspect every film**. That re-inspection is human. Do not regenerate and declare it done. - `claims_history.yaml` entry 38 fails the append-only check in the working tree. Do not commit around it. -- MC-002 and MC-005 disagree on the BELLS licence (`none declared upstream` vs `MIT`). +- ~~MC-002 and MC-005 disagree on the BELLS licence~~ — retracted claim, closed below. + **Closed 2026-09-06:** MC-005 was retracted for this defect. MC-002's + `none declared upstream` / `facts_only` is the register's only statement about + that file, and it records the absence of a licence rather than inferring one. - **K7 binds:** no second device may be audited until `distribution/QUEUE.md` item 5 produces one real viewer response. ### Non-claims for anything built from this brief diff --git a/docs/foundations/FABLE-BRIEF.md b/docs/foundations/FABLE-BRIEF.md index c2ea8f3..fd27c38 100644 --- a/docs/foundations/FABLE-BRIEF.md +++ b/docs/foundations/FABLE-BRIEF.md @@ -42,31 +42,42 @@ Registers nothing. Edits no registry. Writes exactly one file: ARTIFACTS/- An empty cell ships. A forced cell is the failure this protocol exists to prevent. ``` -### First cut to run — it is already blocking the deploy - -`MC-005 removal`. Diagnosed 2026-09-05, not acted on, because the fix requires a -judgment only the owner can declare. - -`scripts/claims_history.py verify` fails: `entries[38] differs from the prior -accepted revision`. Cause, from `prior_reference()` — the kernel compares the -working tree against `HEAD^1`, and at that revision `entries[38]` is MC-005's -registration (keys `commitment`, `digest_of_commitment`). In the working tree -that slot now holds an MC-002 CLARIFY transition. **MC-005's history entry was -replaced, not appended after.** - -The kernel already names its own remedy elsewhere in the same file: -`"protected state exists but the claim is gone — declare a RETRACT"`. So the -append-only fix is to keep entry 38, append a RETRACT for MC-005, and append the -MC-002 CLARIFY after it — 40 entries, prefix rule satisfied. - -**Why an agent must not write it.** The entry's own field is -`direction_basis: DECLARED_HUMAN_JUDGMENT`. A retraction reason is a judgment -about a claim, and the schema says who declares it. There is also a live -candidate reason on record: the OBS CUT found MC-005 recording `license: MIT` -for a file whose upstream declares no licence, which MC-002 records correctly as -`none declared upstream`. Whether that is *the* reason is the owner's to say. - -Run the cut on it. Do not write the entry. +### First cut to run — CLOSED 2026-09-06 + +`MC-005 removal` (the claim is now retracted). Diagnosed 2026-09-05, decided +2026-09-06, and the kernel is +green: `scripts/claims_history.py verify` passes with the prefix rule satisfied. + +What was wrong: `entries[38]` — the registration of the now-retracted MC-005 — +had been replaced rather +than appended after, so the working tree's history was not a prefix of the +accepted one. What was done: entry 38 was kept, a **RETRACT** for MC-005 was +appended as entry 39 with `direction_basis: DECLARED_HUMAN_JUDGMENT`, and the +MC-002 CLARIFY followed it. The retraction reason is on the record and it is +narrow: the registered support block declared `license: MIT` for a BELLS file +whose upstream declares no licence, while MC-002 binds the same file and records +`none declared upstream`. Two records disagreeing about one file's licence is a +defect in the register, not a judgement about the result — and the reason says +so explicitly, so the W1 selection-regret computation is untouched by the +retraction. + +**Where that computation now lives, and what it may be called.** The numbers are +real, executed 2026-09-02, and reproducible: `experiments/e2/results/retrospective/` +holds the run report, the three matrices, and `independent_t1.py`, which +recomputes T1 from the raw released rows outside the analyzer's path. They are +**unregistered**. No claim id carries them, no CI check re-asserts them, and no +surface in this repository may attribute them to a registered claim. +`scripts/verify_retracted.py` enforces exactly that: any line naming a retracted +id must say on the same line that it is retracted, and the gate runs in the +verification manifest. + +Re-registering the same content is not available as a shortcut. The kernel +refuses a second registration for a claim that already has a protected state +(`claims_history.py:333`), and registering it under a fresh id would reverse a +recorded owner judgement without a recorded reason. If the owner wants these +numbers registered, that is a new decision with its own entry — and it needs the +licence block to match MC-002 and a CI re-assertion script before it is worth +making. --- diff --git a/docs/graph/repo-graph.json b/docs/graph/repo-graph.json index 78ca938..f4f2c3b 100644 --- a/docs/graph/repo-graph.json +++ b/docs/graph/repo-graph.json @@ -1022,6 +1022,34 @@ "rel": "reads", "to": "claims.yaml" }, + { + "basis": "STATIC_REF", + "from": "scripts/ledger_snapshot.py", + "inferred": true, + "rel": "reads", + "to": ".claude/skills/evidence-ledger/ledger.py" + }, + { + "basis": "STATIC_REF", + "from": "scripts/ledger_snapshot.py", + "inferred": true, + "rel": "reads", + "to": "ledger.py" + }, + { + "basis": "STATIC_REF", + "from": "scripts/ledger_snapshot.py", + "inferred": true, + "rel": "writes", + "to": "ledger_snapshot.json" + }, + { + "basis": "STATIC_REF", + "from": "scripts/ledger_snapshot.py", + "inferred": true, + "rel": "writes", + "to": "metrics/ledger_snapshot.json" + }, { "basis": "STATIC_REF", "from": "scripts/mixture_bounds.py", @@ -1799,6 +1827,13 @@ "rel": "reads", "to": "scripts/identification.py" }, + { + "basis": "STATIC_REF", + "from": "scripts/verification_manifest.py", + "inferred": true, + "rel": "reads", + "to": "scripts/ledger_snapshot.py" + }, { "basis": "STATIC_REF", "from": "scripts/verification_manifest.py", @@ -1904,6 +1939,13 @@ "rel": "reads", "to": "scripts/verify_consequence.py" }, + { + "basis": "STATIC_REF", + "from": "scripts/verification_manifest.py", + "inferred": true, + "rel": "reads", + "to": "scripts/verify_e3.py" + }, { "basis": "STATIC_REF", "from": "scripts/verification_manifest.py", @@ -1960,6 +2002,13 @@ "rel": "reads", "to": "scripts/verify_resume_receipt.py" }, + { + "basis": "STATIC_REF", + "from": "scripts/verification_manifest.py", + "inferred": true, + "rel": "reads", + "to": "scripts/verify_retracted.py" + }, { "basis": "STATIC_REF", "from": "scripts/verification_manifest.py", @@ -2324,6 +2373,41 @@ "rel": "reads", "to": "try/index.html" }, + { + "basis": "STATIC_REF", + "from": "scripts/verify_e3.py", + "inferred": true, + "rel": "reads", + "to": "RESULT.md" + }, + { + "basis": "STATIC_REF", + "from": "scripts/verify_e3.py", + "inferred": true, + "rel": "reads", + "to": "claims.yaml" + }, + { + "basis": "STATIC_REF", + "from": "scripts/verify_e3.py", + "inferred": true, + "rel": "reads", + "to": "e3_result.json" + }, + { + "basis": "STATIC_REF", + "from": "scripts/verify_e3.py", + "inferred": true, + "rel": "reads", + "to": "experiments" + }, + { + "basis": "STATIC_REF", + "from": "scripts/verify_e3.py", + "inferred": true, + "rel": "reads", + "to": "observations.jsonl" + }, { "basis": "STATIC_REF", "from": "scripts/verify_facts.py", @@ -2653,6 +2737,69 @@ "rel": "reads", "to": "resume/pranav-bhave-resume.receipt.json" }, + { + "basis": "STATIC_REF", + "from": "scripts/verify_retracted.py", + "inferred": true, + "rel": "reads", + "to": "*.html" + }, + { + "basis": "STATIC_REF", + "from": "scripts/verify_retracted.py", + "inferred": true, + "rel": "reads", + "to": "*.json" + }, + { + "basis": "STATIC_REF", + "from": "scripts/verify_retracted.py", + "inferred": true, + "rel": "reads", + "to": "*.md" + }, + { + "basis": "STATIC_REF", + "from": "scripts/verify_retracted.py", + "inferred": true, + "rel": "reads", + "to": "*.py" + }, + { + "basis": "STATIC_REF", + "from": "scripts/verify_retracted.py", + "inferred": true, + "rel": "reads", + "to": "*.txt" + }, + { + "basis": "STATIC_REF", + "from": "scripts/verify_retracted.py", + "inferred": true, + "rel": "reads", + "to": "*.yaml" + }, + { + "basis": "STATIC_REF", + "from": "scripts/verify_retracted.py", + "inferred": true, + "rel": "reads", + "to": "*.yml" + }, + { + "basis": "STATIC_REF", + "from": "scripts/verify_retracted.py", + "inferred": true, + "rel": "reads", + "to": "claims.yaml" + }, + { + "basis": "STATIC_REF", + "from": "scripts/verify_retracted.py", + "inferred": true, + "rel": "reads", + "to": "claims_history.yaml" + }, { "basis": "STATIC_REF", "from": "scripts/verify_spine.py", @@ -3056,6 +3203,26 @@ "rel": "runs", "to": "scripts/claims_history.py" }, + { + "args": [], + "basis": "MANIFEST", + "check": "retracted-id attribution", + "from": "scripts/verification_manifest.py", + "inferred": false, + "rel": "runs", + "to": "scripts/verify_retracted.py" + }, + { + "args": [ + "--test" + ], + "basis": "MANIFEST", + "check": "retracted-id attribution fixtures", + "from": "scripts/verification_manifest.py", + "inferred": false, + "rel": "runs", + "to": "scripts/verify_retracted.py" + }, { "args": [ "--test" @@ -3326,6 +3493,26 @@ "rel": "runs", "to": "scripts/degeneracy.py" }, + { + "args": [], + "basis": "MANIFEST", + "check": "E3 and E3B re-asserted from committed rows", + "from": "scripts/verification_manifest.py", + "inferred": false, + "rel": "runs", + "to": "scripts/verify_e3.py" + }, + { + "args": [ + "--check" + ], + "basis": "MANIFEST", + "check": "ledger snapshot drift", + "from": "scripts/verification_manifest.py", + "inferred": false, + "rel": "runs", + "to": "scripts/ledger_snapshot.py" + }, { "args": [], "basis": "MANIFEST", @@ -3636,6 +3823,60 @@ "sha256": "9736b71389bf0ff1f3a26d03a1ed437c12a3d62c24da94c71306ae30e4a8171e", "to": "scripts/verify_claims.py" }, + { + "basis": "PIN", + "from": "E3-001", + "holds": true, + "inferred": false, + "rel": "pins", + "sha256": "952416606e0caa7d2f58c45322907be6a9c35a0d651910947688977cf99b22c0", + "to": "experiments/e3/results/observations.jsonl" + }, + { + "basis": "PIN", + "from": "E3-001", + "holds": true, + "inferred": false, + "rel": "pins", + "sha256": "84e21e59665b8626d41a73325563ec022251d79187c5b8d91e25f7b5411b2717", + "to": "experiments/e3/results/e3_result.json" + }, + { + "basis": "PIN", + "from": "E3-001", + "holds": true, + "inferred": false, + "rel": "pins", + "sha256": "3d8b6af2311c5c39362d2bc7fb9fd451fd5e43afafae8ec0e665e06432e333ad", + "to": "experiments/e3/RESULT.md" + }, + { + "basis": "PIN", + "from": "E3B-001", + "holds": true, + "inferred": false, + "rel": "pins", + "sha256": "00e75bf47f68ec2487532d595c74362eb85d0e683f3a0d4f294ff2ae429e4631", + "to": "experiments/e3b/results/observations.jsonl" + }, + { + "basis": "PIN", + "from": "E3B-001", + "holds": true, + "inferred": false, + "rel": "pins", + "sha256": "a9685abb1ca63193325ddf319bb5d87cd0c395daf0e7614d571caefbefbfde87", + "to": "experiments/e3b/results/e3_result.json" + }, + { + "basis": "PIN", + "from": "E3B-001", + "holds": true, + "inferred": false, + "rel": "pins", + "sha256": "8a72bd5eb9ddbdf25dfceae9d4bc89dafdd39aa8abae1ae9a7bcddf3fd27d987", + "to": "experiments/e3b/RESULT.md" + }, { "basis": "GIT", "from": "experiments/e2", @@ -3892,8 +4133,8 @@ } ], "generated_from": { - "branch": "main", - "head": "9777b0a" + "branch": "claude/canon-reconciliation", + "head": "4b659fb" }, "nodes": [ { @@ -4106,6 +4347,13 @@ "kind": "instrument", "type": "script" }, + { + "has_check_mode": true, + "has_test_mode": false, + "id": "scripts/ledger_snapshot.py", + "kind": "instrument", + "type": "script" + }, { "has_check_mode": false, "has_test_mode": true, @@ -4267,6 +4515,13 @@ "kind": "verifier", "type": "script" }, + { + "has_check_mode": false, + "has_test_mode": false, + "id": "scripts/verify_e3.py", + "kind": "verifier", + "type": "script" + }, { "has_check_mode": true, "has_test_mode": true, @@ -4323,6 +4578,13 @@ "kind": "verifier", "type": "script" }, + { + "has_check_mode": false, + "has_test_mode": true, + "id": "scripts/verify_retracted.py", + "kind": "verifier", + "type": "script" + }, { "has_check_mode": false, "has_test_mode": false, @@ -4373,13 +4635,13 @@ "type": "script" }, { - "checks": 60, + "checks": 64, "id": "scripts/verification_manifest.py", "type": "trunk" }, { "consequence": "REJECT", - "days_left": 106, + "days_left": 105, "id": "CC-001", "last_reviewed": "2026-08-24", "review_due": "2026-12-22", @@ -4388,7 +4650,7 @@ }, { "consequence": "REJECT", - "days_left": 106, + "days_left": 105, "id": "CC-002", "last_reviewed": "2026-08-24", "review_due": "2026-12-22", @@ -4397,7 +4659,7 @@ }, { "consequence": "REJECT", - "days_left": 106, + "days_left": 105, "id": "CC-003", "last_reviewed": "2026-08-24", "review_due": "2026-12-22", @@ -4406,7 +4668,7 @@ }, { "consequence": "REJECT", - "days_left": 106, + "days_left": 105, "id": "CC-004", "last_reviewed": "2026-08-24", "review_due": "2026-12-22", @@ -4415,7 +4677,7 @@ }, { "consequence": "NARROW", - "days_left": 106, + "days_left": 105, "id": "CC-005", "last_reviewed": "2026-08-24", "review_due": "2026-12-22", @@ -4424,7 +4686,7 @@ }, { "consequence": "REJECT", - "days_left": 120, + "days_left": 119, "id": "CC-006", "last_reviewed": "2026-09-07", "review_due": "2027-01-05", @@ -4433,7 +4695,7 @@ }, { "consequence": "REJECT", - "days_left": 16, + "days_left": 15, "id": "REL-001", "last_reviewed": "2026-08-24", "review_due": "2026-09-23", @@ -4442,7 +4704,7 @@ }, { "consequence": "REJECT", - "days_left": 60, + "days_left": 59, "id": "MC-001", "last_reviewed": "2026-09-07", "review_due": "2026-11-06", @@ -4451,7 +4713,7 @@ }, { "consequence": "REJECT", - "days_left": 50, + "days_left": 49, "id": "MC-002", "last_reviewed": "2026-08-28", "review_due": "2026-10-27", @@ -4460,7 +4722,7 @@ }, { "consequence": "REJECT", - "days_left": 52, + "days_left": 51, "id": "MC-003", "last_reviewed": "2026-08-30", "review_due": "2026-10-29", @@ -4469,7 +4731,7 @@ }, { "consequence": "REJECT", - "days_left": 53, + "days_left": 52, "id": "MC-004", "last_reviewed": "2026-08-31", "review_due": "2026-10-30", @@ -4478,7 +4740,7 @@ }, { "consequence": "NARROW", - "days_left": 80, + "days_left": 79, "id": "AF-001", "last_reviewed": "2026-08-28", "review_due": "2026-11-26", @@ -4487,7 +4749,7 @@ }, { "consequence": "NARROW", - "days_left": 106, + "days_left": 105, "id": "GA-001", "last_reviewed": "2026-08-24", "review_due": "2026-12-22", @@ -4496,7 +4758,7 @@ }, { "consequence": "NARROW", - "days_left": 106, + "days_left": 105, "id": "GV-001", "last_reviewed": "2026-08-24", "review_due": "2026-12-22", @@ -4505,7 +4767,7 @@ }, { "consequence": "NARROW", - "days_left": 351, + "days_left": 350, "id": "GCE-001", "last_reviewed": "2026-08-24", "review_due": "2027-08-24", @@ -4514,7 +4776,7 @@ }, { "consequence": "NARROW", - "days_left": 113, + "days_left": 112, "id": "SITE-001", "last_reviewed": "2026-08-31", "review_due": "2026-12-29", @@ -4523,13 +4785,31 @@ }, { "consequence": "REJECT", - "days_left": 112, + "days_left": 111, "id": "SITE-002", "last_reviewed": "2026-08-30", "review_due": "2026-12-28", "support_url": "https://github.com/Cubits11/cubits11.github.io/blob/main/scripts/verify_claims.py", "type": "claim" }, + { + "consequence": "REJECT", + "days_left": 120, + "id": "E3-001", + "last_reviewed": "2026-09-08", + "review_due": "2027-01-06", + "support_url": "https://github.com/Cubits11/cubits11.github.io/blob/main/experiments/e3/results/observations.jsonl", + "type": "claim" + }, + { + "consequence": "REJECT", + "days_left": 120, + "id": "E3B-001", + "last_reviewed": "2026-09-08", + "review_due": "2027-01-06", + "support_url": "https://github.com/Cubits11/cubits11.github.io/blob/main/experiments/e3b/results/observations.jsonl", + "type": "claim" + }, { "docs": { "PREREG.md": false, @@ -4571,8 +4851,8 @@ { "ahead_of_origin_main": 0, "behind_origin_main": 0, - "head": "9777b0a", - "id": "main", + "head": "4b659fb", + "id": "claude/canon-reconciliation", "last_commit": "2026-09-07", "reachable_from_origin_main": true, "type": "branch" @@ -4580,8 +4860,8 @@ { "ahead_of_origin_main": 0, "behind_origin_main": 0, - "head": "9777b0a", - "id": "origin", + "head": "4b659fb", + "id": "main", "last_commit": "2026-09-07", "reachable_from_origin_main": true, "type": "branch" @@ -4589,7 +4869,7 @@ { "ahead_of_origin_main": 0, "behind_origin_main": 0, - "head": "9777b0a", + "head": "4b659fb", "id": "origin/main", "last_commit": "2026-09-07", "reachable_from_origin_main": true, @@ -4605,5 +4885,5 @@ } ], "schema": "repo-graph v0.1", - "stable_digest": "309fc44a19ab387d4de101f36644beb717657c3264132333a2df1cc2bcfb51ff" + "stable_digest": "bbe9895078294fdaacc1c1e086643c8ad88c7d3f0db66a1d343769503a166571" } diff --git a/docs/graph/repo-graph.mmd b/docs/graph/repo-graph.mmd index 8c2b337..fccc84d 100644 --- a/docs/graph/repo-graph.mmd +++ b/docs/graph/repo-graph.mmd @@ -34,37 +34,38 @@ graph LR n32["GCE-001"] n33["SITE-001"] n34["SITE-002"] - n35["experiments/e2"] - n36["experiments/e3"] - n37["experiments/e3b"] - n38["distribution/traction"] - n17 -->|runs| n39 - n17 -->|runs| n40 + n35["E3-001"] + n36["E3B-001"] + n37["experiments/e2"] + n38["experiments/e3"] + n39["experiments/e3b"] + n40["distribution/traction"] n17 -->|runs| n41 - n17 -->|runs| n5 n17 -->|runs| n42 n17 -->|runs| n43 - n17 -->|runs| n43 - n17 -->|runs| n43 + n17 -->|runs| n5 n17 -->|runs| n44 n17 -->|runs| n45 n17 -->|runs| n46 + n17 -->|runs| n46 + n17 -->|runs| n45 + n17 -->|runs| n45 n17 -->|runs| n47 - n17 -->|runs| n7 n17 -->|runs| n48 - n17 -->|runs| n9 - n17 -->|runs| n10 - n17 -->|runs| n8 - n17 -->|runs| n49 n17 -->|runs| n49 n17 -->|runs| n50 - n17 -->|runs| n50 - n17 -->|runs| n6 + n17 -->|runs| n7 n17 -->|runs| n51 - n17 -->|runs| n12 + n17 -->|runs| n9 + n17 -->|runs| n10 + n17 -->|runs| n8 n17 -->|runs| n52 + n17 -->|runs| n52 + n17 -->|runs| n53 n17 -->|runs| n53 + n17 -->|runs| n6 n17 -->|runs| n54 + n17 -->|runs| n12 n17 -->|runs| n55 n17 -->|runs| n56 n17 -->|runs| n57 @@ -79,59 +80,70 @@ graph LR n17 -->|runs| n66 n17 -->|runs| n67 n17 -->|runs| n68 - n17 -->|runs| n1 n17 -->|runs| n69 - n17 -->|runs| n2 - n17 -->|runs| n3 - n17 -->|runs| n15 n17 -->|runs| n70 - n17 -->|runs| n16 n17 -->|runs| n71 - n17 -->|runs| n11 n17 -->|runs| n72 n17 -->|runs| n73 + n17 -->|runs| n1 n17 -->|runs| n74 + n17 -->|runs| n2 + n17 -->|runs| n3 + n17 -->|runs| n15 n17 -->|runs| n75 - n17 -->|runs| n75 + n17 -->|runs| n16 n17 -->|runs| n76 + n17 -->|runs| n11 n17 -->|runs| n77 - n17 -->|runs| n14 n17 -->|runs| n78 n17 -->|runs| n79 - n25 -->|pins| n80 - n27 -->|pins| n54 - n29 -->|pins| n56 - n33 -->|pins| n81 - n34 -->|pins| n39 - n35 -->|governed_by| n82 - n35 -->|governed_by| n83 - n35 -->|governed_by| n84 - n35 -->|governed_by| n85 - n35 -->|governed_by| n86 - n35 -->|governed_by| n87 - n35 -->|governed_by| n88 - n35 -->|governed_by| n89 - n35 -->|governed_by| n90 - n36 -->|governed_by| n91 - n36 -->|governed_by| n92 - n36 -->|governed_by| n93 - n36 -->|governed_by| n94 - n36 -->|governed_by| n95 - n36 -->|governed_by| n96 - n36 -->|governed_by| n97 - n36 -->|governed_by| n98 - n36 -->|governed_by| n99 - n36 -->|governed_by| n100 + n17 -->|runs| n80 + n17 -->|runs| n80 + n17 -->|runs| n81 + n17 -->|runs| n82 + n17 -->|runs| n14 + n17 -->|runs| n83 + n17 -->|runs| n84 + n25 -->|pins| n85 + n27 -->|pins| n57 + n29 -->|pins| n59 + n33 -->|pins| n86 + n34 -->|pins| n41 + n35 -->|pins| n87 + n35 -->|pins| n88 + n35 -->|pins| n89 + n36 -->|pins| n90 + n36 -->|pins| n91 + n36 -->|pins| n92 + n37 -->|governed_by| n93 + n37 -->|governed_by| n94 + n37 -->|governed_by| n95 + n37 -->|governed_by| n96 + n37 -->|governed_by| n97 + n37 -->|governed_by| n98 + n37 -->|governed_by| n99 + n37 -->|governed_by| n100 n37 -->|governed_by| n101 - n37 -->|governed_by| n102 - n37 -->|governed_by| n103 - n37 -->|governed_by| n104 - n37 -->|governed_by| n105 - n37 -->|governed_by| n106 - n37 -->|governed_by| n107 - n37 -->|governed_by| n108 - n38 -->|distributes| n18 - n38 -->|distributes| n21 - n38 -->|distributes| n25 - n38 -->|distributes| n26 - n38 -->|distributes| n27 + n38 -->|governed_by| n102 + n38 -->|governed_by| n103 + n38 -->|governed_by| n104 + n38 -->|governed_by| n105 + n38 -->|governed_by| n106 + n38 -->|governed_by| n107 + n38 -->|governed_by| n108 + n38 -->|governed_by| n109 + n38 -->|governed_by| n110 + n38 -->|governed_by| n111 + n39 -->|governed_by| n112 + n39 -->|governed_by| n113 + n39 -->|governed_by| n114 + n39 -->|governed_by| n115 + n39 -->|governed_by| n116 + n39 -->|governed_by| n117 + n39 -->|governed_by| n118 + n39 -->|governed_by| n119 + n40 -->|distributes| n18 + n40 -->|distributes| n21 + n40 -->|distributes| n25 + n40 -->|distributes| n26 + n40 -->|distributes| n27 diff --git a/experiments/e3/RESULT.md b/experiments/e3/RESULT.md index e9ee213..7cf460a 100644 --- a/experiments/e3/RESULT.md +++ b/experiments/e3/RESULT.md @@ -84,4 +84,12 @@ this result stands whatever that one shows. - Nothing here transfers to E2's three guards, its pools, or its operating points. - The programme's least favourable fact is unchanged and untested by this: joint measurement changed second-guard selection by at most 2.4 points on every per-item matrix examined, with no regret interval excluding zero. - Discrepancy D1 (the FPR\* threshold direction) is declared in `e3_config.json`; it was decided before any harmful item was scored. -- Not registered in `claims.yaml`. Registration is an owner action. +- **Registered 2026-09-08 as `E3-001`.** (This line replaced "Not registered in + `claims.yaml`; registration is an owner action" — the owner action was taken.) + The registration carries the failed prediction(s) above in the commitment + itself, not only in this file, and `scripts/verify_e3.py` re-derives every + registered number — including the bootstrap interval — from + `results/observations.jsonl` alone on every push. Nothing in the numbers, + the predictions, the thresholds or the verdicts above was changed by the + registration; editing any of them now fires the claim's own local-content + trigger. diff --git a/experiments/e3b/RESULT.md b/experiments/e3b/RESULT.md index d34dd0c..5aedf87 100644 --- a/experiments/e3b/RESULT.md +++ b/experiments/e3b/RESULT.md @@ -113,4 +113,12 @@ belongs in any future prereg as a gate, not a hope. - The programme's least favourable fact is untouched: joint measurement changed second-guard selection by at most 2.4 points on every per-item matrix examined, with no regret interval excluding zero. -- Not registered in `claims.yaml`. Registration is an owner action. +- **Registered 2026-09-08 as `E3B-001`.** (This line replaced "Not registered in + `claims.yaml`; registration is an owner action" — the owner action was taken.) + The registration carries the failed prediction(s) above in the commitment + itself, not only in this file, and `scripts/verify_e3.py` re-derives every + registered number — including the bootstrap interval — from + `results/observations.jsonl` alone on every push. Nothing in the numbers, + the predictions, the thresholds or the verdicts above was changed by the + registration; editing any of them now fires the claim's own local-content + trigger. diff --git a/films/EXPLAINER-SERIES.md b/films/EXPLAINER-SERIES.md index a844105..7084aab 100644 --- a/films/EXPLAINER-SERIES.md +++ b/films/EXPLAINER-SERIES.md @@ -1,7 +1,16 @@ # The Explainer Series — ten episodes **STATUS: DRAFT SCRIPTS. Nothing recorded, nothing uploaded, nothing scheduled.** -Written 2026-09-05 on branch `claude/mc-005-selection-regret`. This file +Written 2026-09-05; **reconciled against the record 2026-09-08** after E3 and +E3B registered, MC-005 was retracted, and `exclusive_cells` was committed. +Episode 10 was rewritten in full because its central sentence had gone false. + +**Standing rule for this deck.** Any counter this deck speaks — claims, rows, +qualified outcomes, open blockers — is read from `metrics/ledger_snapshot.json` +and quoted with that file's `as_of` date, never typed from memory. +`scripts/ledger_snapshot.py --check` fails when the repository moves and the +snapshot does not, which is the gate that would have caught Episode 10 three +days earlier. Re-read it before any take. This file registers no claim, edits no registry, changes no generated page, and starts no campaign. Every numeral below carries a locator; two are flagged as **UNVERIFIED** and are barred from recording until checked. @@ -114,13 +123,14 @@ Both must exit 0. Additionally, three per-episode blocks: | # | Blocker | Affects | State on 2026-09-05 | |---|---|---|---| -| B1 | `exclusive_cells` is an **uncommitted working-tree addition** to MC-002. The 32-pattern table has no committed registry binding yet. | **Ep 2, Ep 4** | Open. I recomputed all 8 occupied patterns independently from the hash-verified released file and they agree exactly; the *fact* is solid, the *binding* is not. Commit and regenerate before filming. | -| B2 | MC-002 records the BELLS file's licence as `none declared upstream`; MC-005 at HEAD records `license: MIT` for the same file. `ARTIFACTS/2026-09-05-FABLE-5.1-OBS-CUT.md` confirms the upstream repository declares no licence. | **Ep 7** | Open, and it is a registry inconsistency, not a filming detail. Do not record Ep 7 until one of the two records changes. | -| B3 | The all-miss category split (Ep 2) and the LLM Guard file-B count (Ep 8) are computed in-session, not registered. | **Ep 2, Ep 8** | Ep 2's split is reproducible from the released file today. Ep 8's requires a count on upstream file `d6ebd0e5` that **has not been taken** — marked UNVERIFIED in place. | -| B4 | `scripts/verification_manifest.py` **does not exit 0 on this working tree.** The claim-history kernel reports `entries[38] differs from the prior accepted revision — accepted history is append-only`, from the uncommitted `claims_history.yaml` change that accompanies B1. | **all ten** | Open as of 2026-09-05. This is the owner's in-progress edit, not something this deck touched, but the pre-flight rule is unconditional: nothing is recorded while the manifest fails. An accepted transition being rewritten rather than appended is also the exact shape the registry treats as serious — resolve it as a registry question first, not as a filming blocker. | +| B1 | `exclusive_cells` had no committed registry binding. | **Ep 2, Ep 4** | **CLOSED 2026-09-06.** The 32-cell block is committed in MC-002's `expected`, declared by a CLARIFY transition in `claims_history.yaml`, re-asserted against the hash-verified file by `scripts/reanalyze_bells_subset.py` in CI, and required by `scripts/generate_missing_column.py` and `scripts/verify_figures.py`. | +| B2 | Two registry records disagreed about the BELLS file's licence. | **Ep 7** | **CLOSED 2026-09-06.** MC-005 was retracted for exactly this defect; MC-002's `none declared upstream` with `commercial_reuse: facts_only` is the register's only statement about that file, and it records the absence of a licence rather than inferring one from silence. Ep 7's own anchors changed as a result — see Ep 7. | +| B3 | The all-miss category split (Ep 2) and the LLM Guard file-B count (Ep 8) are computed in-session, not registered. | **Ep 2, Ep 8** | **Still open, and narrower than it was.** Ep 2's split is reproducible from the released file today. Ep 8's requires a count on upstream file `d6ebd0e5` that **has not been taken** — marked UNVERIFIED in place, and barred from recording until it is. | +| B4 | The claim-history kernel reported `entries[38] differs from the prior accepted revision`. | **all ten** | **CLOSED 2026-09-06.** Entry 38 was kept, a RETRACT for MC-005 appended after it, and the MC-002 CLARIFY appended after that. `scripts/claims_history.py verify` passes: 44 entries, prefix rule satisfied, 19 live commitments equal to the chain tip. | Every numeral spoken on camera must resolve from `films/data/facts.json`, from -`claims.yaml`, or from a command shown running on screen. If a spoken number +`claims.yaml`, from `metrics/ledger_snapshot.json` with its date spoken aloud, or +from a command shown running on screen. If a spoken number disagrees with the render, redo the take. No exceptions, including for a number you are certain about. @@ -372,7 +382,7 @@ LangKit, NeMo, LLM Guard, where **1 means missed**: | union / all-miss | 73 / 9 (11.0%) | MC-002 `expected.union_detection`, `all_miss` | | product of miss rates | 3.5% | MC-002 proposition | | ratio, recomputed to product | ≈3.1× | MC-002 proposition | -| the eight patterns | as tabled | MC-002 `expected.exclusive_cells` — **UNCOMMITTED (B1)**. Independently recomputed from the hash-verified file this session; the eight counts sum to 82 and each filter's row-sum equals `82 − catches`. | +| the eight patterns | as tabled | MC-002 `expected.exclusive_cells` — **committed and bound.** Re-asserted from the hash-verified released file by `scripts/reanalyze_bells_subset.py` on every push; the eight counts sum to 82 and each filter's row-sum equals `82 − catches`. | --- @@ -869,8 +879,17 @@ the finding hiding inside a number that looks like bookkeeping. **Public title:** *I Spent a Month on This. The Answer Was "It Doesn't Matter."* **Runtime:** 6:00 · **Form:** owner camera + terminal + one long unbroken take -**Teaches F7 · Retrieves F4 and F3** · **Bound to:** MC-005 -**BLOCKED BY B2 — do not record until the licence contradiction is resolved.** +**Teaches F7 · Retrieves F4 and F3** · **Bound to:** no registered claim. +The numbers come from `experiments/e2/results/retrospective/REPORT.md` and are +**unregistered**: MC-005, which used to carry them, was retracted 2026-09-06 +over a licence defect in its support block, and the retraction reason states +that the computation itself is untouched. +**RECORDING CONDITION.** Every number in this episode must be spoken as +*computed 2026-09-02, unregistered, and not re-asserted by CI* — in the +episode, out loud, not in a description. Saying them with the authority of a +registration is the failure this deck exists to prevent. B2 is closed; this +condition replaces it and does not expire until the owner registers the +computation under a new id or decides not to. This is the episode that decides whether this channel is doing research or marketing. It is the one where my own result argues against my own pitch. It @@ -1001,15 +1020,21 @@ Say it once. Do not soften it. Do not follow it with "but." | On screen | Value | Locator | |---|---|---| -| maximum selection regret | ≤ 2 items (2.4 pp) | MC-005 proposition (HEAD) | -| intervals excluding zero | 0 of 11 | MC-005 proposition | -| specialized supervisors, same pick | 5 of 5, regret 0, CI [0,0] | MC-005 proposition | -| positive excess joint miss | 45 of 45 non-degenerate pairs | MC-005 proposition | -| the ten pairs not counted | 55 total pairs (11 choose 2) − 45 | DERIVED here, not stated by MC-005: LLM Guard misses all 82, so its miss indicator is constant and its excess joint miss is identically zero against every partner — exactly 10 pairs. Verify against the run report before speaking it. | -| stratified odds ratio ≥ 1.5 | 23 of 36 defined | MC-005 proposition | -| +inf ratios counted / undefined dropped | 17 and 9 of 45 | MC-005 scope, discrepancies D3 | -| bootstrap | B = 2000, picks fixed, D4 | MC-005 scope | -| executed | owner, 2026-09-02; **not re-asserted by CI on every push, and the record says so** | MC-005 scope | +**Every row below is UNREGISTERED.** MC-005 is retracted; no claim id carries +these numbers and no CI check re-asserts them. The locator is the run report +and the matrices beside it, not the registry. + +| On screen | Value | Locator | +|---|---|---| +| maximum selection regret | ≤ 2 items (2.4 pp) | `experiments/e2/results/retrospective/REPORT.md` T1 — unregistered | +| intervals excluding zero | 0 of 11 | same run report — unregistered | +| specialized supervisors, same pick | 5 of 5, regret 0, CI [0,0] | same run report — unregistered | +| positive excess joint miss | 45 of 45 non-degenerate pairs | same run report — unregistered | +| the ten pairs not counted | 55 total pairs (11 choose 2) − 45 | DERIVED in this deck, not stated by the run report: LLM Guard misses all 82, so its miss indicator is constant and its excess joint miss is identically zero against every partner — exactly 10 pairs. Verify against the run report before speaking it. | +| stratified odds ratio ≥ 1.5 | 23 of 36 defined | same run report — unregistered | +| +inf ratios counted / undefined dropped | 17 and 9 of 45 | same run report — unregistered | +| bootstrap | B = 2000, picks fixed, seed frozen 2026-09-01 | same run report — unregistered | +| executed | owner, 2026-09-02; **unregistered, and not re-asserted by CI on any push** | the retraction of MC-005, `claims_history.yaml` entry 39 | | independent recomputation | agrees on every pick, union, regret, benign union | `experiments/e2/results/retrospective/independent_t1.py` | --- @@ -1208,7 +1233,8 @@ sentence it is *not*. | "Most guardrails are useless" | Three of five added zero **to this group, on this stratum** | | "Pairwise testing is useless" | Two **constructed** worlds show it can't determine the triple | | "The census proves joint evidence doesn't exist" | One documented search, one reviewer, found five of twenty | -| "Joint measurement is essential" | Changed the arithmetic in 45 of 45 pairs. Changed no decision. | +| "Joint measurement is essential" | Changed the arithmetic in 45 of 45 pairs. Changed no decision — and that count is unregistered. | +| "I measure guardrail dependence" | Two pilots, 4,800 rows, both primary predictions failed, both identified sets degenerate. | | "The benchmark is unreliable" | Its population is exactly recoverable. Its **selection rule** is not. | > Every one of those left-hand sentences would get more views than the @@ -1261,27 +1287,43 @@ Beat. --- -# EPISODE 10 · ZERO +# EPISODE 10 · TWO OF NINETEEN -**Public title:** *I Built a Research Program and It Has Measured Nothing* +**Public title:** *I Built a Research Program. Here Is Everything It Has Actually Measured.* **Runtime:** 6:30 · **Form:** owner camera + one live terminal command, unedited -**Mass retrieval: F1–F9** · **Bound to:** `.claude/skills/evidence-ledger/ledger.py` - -The capstone, and the only episode that is genuinely uncomfortable to publish. -It runs one command on camera and reads its output without cutting away. +**Mass retrieval: F1–F9** · **Bound to:** `E3-001`, `E3B-001`, and +`metrics/ledger_snapshot.json` + +**Rewritten 2026-09-08.** The previous cut of this episode said "observation +rows: zero" and "this repository has produced zero measurements of its own." It +was written 2026-09-05 and both sentences were false by 2026-09-06, when E3 and +E3B committed 4,800 rows. The episode that accused the whole field of quoting +stale numbers had gone stale in three days, in the one place it could least +afford to. That is why the deck now reads its counters from +`metrics/ledger_snapshot.json` and speaks the date, and why +`scripts/ledger_snapshot.py --check` fails the build when the repository moves +and the snapshot does not. + +The capstone, and the only episode that is uncomfortable to publish. It runs one +command on camera and reads its output without cutting away. ### Rejected conceits, and why - *Ending on hope.* Killed. Any "but here's what's next!" turn converts an - honest accounting into a pitch, and the whole series has been arguing that the - turn is where claims get inflated. + honest accounting into a pitch, and the whole series has argued that the turn + is where claims get inflated. - *Not making this episode.* Named because it was the real temptation. A channel that publishes nine competent explainers and hides the ledger is doing - marketing with a research aesthetic. The ledger is the differentiator; the - explainers are the doorway. -- *Making it episode one.* Killed on ordering grounds. "I've measured nothing" - as a cold open with no prior context is self-deprecation. After nine episodes - of actual findings, it's an accounting. + marketing with a research aesthetic. +- *Keeping the old "zero measurements" version because it is a better story.* + Killed, and this is the one worth naming. The old cut was more dramatic and it + is now false. A more quotable sentence that has stopped being true is exactly + the object this series exists to take apart, and it does not get an exemption + for being mine. +- *Presenting the two pilots as a comeback.* Killed. Both failed their primary + prediction. The honest shape is not "and then I measured something" — it is + "and then I measured something, and it could not answer the question I built + it to answer." ### Cold recall · 0:00–0:30 *(mass retrieval — no answers given)* @@ -1308,107 +1350,148 @@ Nine questions, fast, no pauses for answers: **PREDICTION HOLD — 3 seconds.** -### The command · 1:30–3:00 +### The command · 1:30–3:10 -Run it live. Do not cut. Read the output as it appears. +Run it live. Do not cut. Read the output as it appears, and say the date. ```bash python3 .claude/skills/evidence-ledger/ledger.py ``` -> Seventeen registered claims. Own measurements: **zero.** +> Nineteen registered claims. The number of them resting on a measurement I made +> myself: **two.** Both registered two days ago. > -> Observation rows this repository has produced: **zero.** +> Observation rows this repository has produced: **four thousand eight hundred.** +> Two experiments, twenty-four hundred rows each. > -> Qualified outcomes from the outside world: **zero.** +> Outcomes produced by anybody who is not me — a reproduction, a correction, a +> cold run: **zero.** Across five categories. Zero is recorded as zero. > -> The main experiment has **eight governing documents and zero rows.** Documents -> per row: infinity. The tool prints the infinity symbol, because I wrote it to. +> Thirteen things are blocked on a human doing something. > -> Eleven things are blocked on a human, and the oldest has been open four days. +> And the main experiment, the one this whole programme is built around, still +> has **eight governing documents and no data at all.** The tool prints +> documents-per-row as an infinity symbol, because I wrote it to. -### The honest accounting · 3:00–4:15 +### The honest accounting · 3:10–4:30 -> So what have I actually been doing? Two things, and it's worth being exact. +> Two of nineteen. That number moved off zero on the sixth of September and it is +> worth being exact about what moved it, because it is not a success story. +> +> **Two pilots. Both failed their main prediction.** +> +> The first one: twelve hundred items, two classifiers, twenty-four hundred rows. +> I'd predicted the joint miss rate would come out above what independence +> predicts, with a confidence interval clear of zero, and I had a reason — +> every per-item matrix I'd looked at went that way. It came out at plus +> nought-point-one-eight of a percentage point, interval straddling zero. > -> **Recounting other people's public files.** Every empirical number in this -> series is somebody else's measurement that I recomputed. That is real work — -> nine of nine independent checks landing on a hidden population is real, and -> three of five filters contributing zero is real. It is **arithmetic on -> released bits**, and I loaded no model to get it. +> And the reason was my fault in a specific way. Both classifiers missed +> ninety-six to ninety-eight percent of that material. They were injection +> detectors and I'd pointed them at ordinary harmful requests. When both rates +> sit at the top like that, the range the two scores leave open is one and +> three-quarter points wide. There was nothing in there for a joint measurement +> to find. > -> **Building machinery that makes my own claims prosecutable.** Forty-six -> automatic checks. A rule that already rejected one of my published numbers. A -> registry where a claim expires if I don't re-check it. +> So I ran it again on the material those detectors were built for. **It failed +> harder.** One of them caught four hundred out of four hundred. Miss rate zero. +> And if either filter misses nothing, then both-miss is exactly zero — the two +> published rates have already told you the answer completely. The range was +> **zero points wide.** > -> What I have not done is **measure anything myself.** Not once. The experiment -> that would is frozen, correctly, waiting on hardware I own and a human action I -> haven't taken. +> Two experiments. Two useless ranges. At opposite ends. -### The wrong answer on purpose · 4:15–4:45 +### The wrong answer on purpose · 4:30–5:00 -> Which means the machinery was premature. Should have measured first, built the -> apparatus after. +> So the programme's thesis is wrong. The missing column doesn't matter. Beat. -> I don't think so, and I want to be precise about why, because "I was right to -> build it" is exactly the self-serving conclusion to watch me for. +> No — and I want to be careful here, because "my thesis survived" is exactly the +> conclusion to watch me for. What the two runs suggest is narrower than my +> thesis and narrower than that objection. **The range is wide in the middle and +> degenerate at the ends.** When every filter's miss rate is somewhere ordinary, +> the two scores leave a lot open and the joint measurement earns its cost. At +> either extreme they have already answered you. > -> The rule that rejected my census number was written **three days before** it -> fired. If I'd written it after finding the missed evaluation, it would have -> been worthless — that's a threshold chosen after seeing the outcome, and it's -> the one thing this program treats as voiding the result retroactively. +> That is a hypothesis my own two runs suggest. It is not a result of either run; +> neither of them tested it. It is registered as a non-claim, in those words, on +> both experiments. > -> The apparatus has to be early or it isn't apparatus. **That's a reason for -> some of it. It is not a reason for eight documents and zero rows.** Those are -> different sentences and only the first one is defensible. +> What does fall out of it is an engineering check, and it costs about forty +> items: score a small slice, read the two miss rates, work out how wide the +> range is, and abandon the pool if it is narrow. Both of my experiments would +> have been stopped by that before I downloaded a single model. -### Weakest sentence that survives · 4:45–5:10 +### The thing that argues against me hardest · 5:00–5:30 -> **This repository has produced zero measurements of its own. Everything -> empirical in these ten episodes is a recount of somebody else's public file. -> The apparatus that would catch me being wrong exists and has fired once. The -> measurement it was built for has not started.** +> One more, and it is the least favourable number I have. +> +> Somebody took every per-item matrix available and asked whether measuring the +> joint would have changed *which* second filter you picked. At most about two +> and a half percentage points, and not one case where the effect was +> distinguishable from zero. Measuring the joint changed the arithmetic and +> changed no decision. +> +> **That number is unregistered.** It has no claim id, no CI check re-asserts it, +> and the claim that used to carry it was retracted over a defect in its own +> support block. It is computed, it is reproducible, it is in the repository — +> and it is not part of the register, and I am not allowed to say it as though +> it were. + +### Weakest sentence that survives · 5:30–5:55 -### The one action · 5:10–6:00 +> **Two of nineteen registered claims rest on measurements I made. Both are +> pilots whose primary prediction failed, on pools where the two scores had +> already fixed the answer. Everything else empirical in these ten episodes is a +> recount of somebody else's public file, and no outcome has yet been produced by +> anyone who is not me.** -> There's one thing standing between here and the first row of my own data, and -> it is not code and it is not a model. +### The one action · 5:55–6:15 + +> There is one thing standing between here and a result I would defend, and it is +> not compute and it is not funding. > > **Five people who have never seen this project need to watch one short film, > with the sound off, and answer three questions.** If two of them miss the same -> question, the film doesn't ship and I rewrite it. That gate is frozen. I -> haven't run it. -> -> That is the whole bottleneck. Not compute. Not funding. Five strangers and a -> stopwatch. Everything else in this repository is me building things I'm -> allowed to build without asking anyone. +> question, the film does not ship and I rewrite it. That gate is frozen. I have +> not run it. -### Fluency guard · 6:00–6:15 +### Fluency guard · 6:15–6:25 -> Ten episodes is a lot of fluency about a program with zero rows. If these got -> good enough that the work started sounding established — that's the failure -> mode this series was designed against, and I built it anyway, and you should -> hold me to the counters rather than the delivery. +> Ten episodes is a lot of fluency about a programme with two own-measurement +> claims and two failed predictions. If these got good enough that the work +> started sounding established — that is the failure mode this series was +> designed against, and I built it anyway, and you should hold me to the counters +> rather than the delivery. -### Exit ticket · 6:15–6:30 +### Exit ticket · 6:25–6:30 -> Run the command yourself. It's in the repository. If the zeros have changed by -> the time you're watching, that's the only evidence that any of this went +> Run the command yourself. It is in the repository. If those counters have moved +> by the time you are watching, that is the only evidence any of this went > anywhere. **cubits11.github.io** ### Bound numerals +Every counter below is read from `metrics/ledger_snapshot.json`, **as of +2026-09-08**, and the date is spoken on camera. If +`python3 scripts/ledger_snapshot.py --check` fails, this episode is stale and +does not get recorded until it is regenerated and this table is re-read. + | On screen | Value | Locator | |---|---|---| -| claims / own-measurement | 17 / **0** | `ledger.py`, 2026-09-05 run | -| observation rows | 0 | same | -| qualified outcomes | 0, across 5 categories | same; `distribution/outcomes.yaml` | -| open blockers / oldest | 11 / 4 days | same | -| e2 scaffolding | 8 documents, 0 rows | same | -| automatic checks | 46 | `scripts/verification_manifest.py` | -| external interactions before the stop rule | 3 of 12 | `distribution/QUEUE.md` item 9 | +| registered claims | 19 | `metrics/ledger_snapshot.json` → `claims.total` | +| resting on own measurement | **2** (E3-001, E3B-001) | same → `claims.own_measurement` | +| observation rows | **4,800** (2,400 + 2,400) | same → `observations.rows` | +| qualified outcomes | 0, across 5 categories | same → `outcomes.qualified_total`; `distribution/outcomes.yaml` | +| open blockers | 13 | same → `blocking.open` (ages are read live from the ledger, never typed) | +| e2 scaffolding | 8 documents, 0 rows | same → `scaffolding` | +| external interactions before the stop rule | 3 of 12 | same → `outcomes.technical_interactions` | +| E3 delta / CI | +0.0018, 95% CI [−0.00096, +0.00706], includes zero | `claims.yaml` E3-001 `expected` | +| E3 Fréchet width | 1.75 pp, from marginals 0.9825 / 0.9625 | E3-001 `expected.harmful.frechet` | +| E3B miss rates / width | 0.0000 and 0.3975 → width 0 | `claims.yaml` E3B-001 `expected` | +| both pilots' primary prediction | FAILED | E3-001 and E3B-001 `expected.prediction_verdicts` | +| selection regret ≤ 2 items (2.4 pp), 0 of 11 CIs exclude zero | **UNREGISTERED** | `experiments/e2/results/retrospective/REPORT.md`; the claim that carried it is retracted | | the cold gate | 5 minimum, 8 maximum; two failures on one question holds release | `distribution/launch-units.yaml` → `cold_test_gate`; `QUEUE.md` item 5 | --- diff --git a/films/data/facts.json b/films/data/facts.json index 457574b..a1b01de 100644 --- a/films/data/facts.json +++ b/films/data/facts.json @@ -2,7 +2,7 @@ "_generated_by": "scripts/films/bind_facts.py — DO NOT EDIT; regenerate and re-inspect the films", "_inputs": { "census.yaml": "0e130c67013fe834df32c7f8de5afac5cb8d2dc4f66b66423bd6a98016e7c598", - "claims.yaml": "3b597a200639e0c9781c59ac5b6bb6c1257cd1ec389ecfe3638bf21f52ab6357" + "claims.yaml": "0c7791565d877cc45627dc2374a842a8784c5ec39e74771a6a4c106e71a9a38b" }, "_kinds": [ "OBSERVED", @@ -1240,13 +1240,31 @@ "provenance": "owner_verified", "review_due": "2026-12-28", "review_window_days": 120 + }, + { + "evidential_status": "supported_within_scope", + "id": "E3-001", + "last_reviewed": "2026-09-08", + "maturity": "experimental", + "provenance": "machine_generated_owner_executed", + "review_due": "2027-01-06", + "review_window_days": 120 + }, + { + "evidential_status": "supported_within_scope", + "id": "E3B-001", + "last_reviewed": "2026-09-08", + "maturity": "experimental", + "provenance": "machine_generated_owner_executed", + "review_due": "2027-01-06", + "review_window_days": 120 } ] }, "REGISTRY.last_owner_review": { "kind": "REGISTRY", "source": "claims.yaml:last_owner_review", - "value": "2026-09-07" + "value": "2026-09-08" }, "REGISTRY.version": { "kind": "REGISTRY", diff --git a/films/flagship/source-map.json b/films/flagship/source-map.json index 3193e46..d986a74 100644 --- a/films/flagship/source-map.json +++ b/films/flagship/source-map.json @@ -4,8 +4,8 @@ "films/flagship/script.json": "db9bb0e5f475bf0dfc05085e71899eab9b0e47d39446e5850f22dd4404d0cca2", "scripts/films/build_flagship.py": "09248adf7cf047ffad802eadacb2122b35c5f9131fe6b6c46358aa6abfcffcfe", "films/flagship/rehearsal-template.html": "1d6e8f858360f6248108a249d62f0f206d1b9bb36e4cbaa21238ea475dc0e614", - "films/data/facts.json": "0c2591dd9cef13583fa48f0b736818cad0e9342cd00210c350bdc6bf9822a962", - "claims.yaml": "3b597a200639e0c9781c59ac5b6bb6c1257cd1ec389ecfe3638bf21f52ab6357", + "films/data/facts.json": "15136fc279b788aad616167b9cd3f6ef9b009d643c3466b04cef985719656651", + "claims.yaml": "0c7791565d877cc45627dc2374a842a8784c5ec39e74771a6a4c106e71a9a38b", "census.yaml": "0e130c67013fe834df32c7f8de5afac5cb8d2dc4f66b66423bd6a98016e7c598", "films/the-stack/manifest.yaml": "1b8acaac37e0b2ec95e59572a0b583d59fb0d37cf506806320236568f60d915f", "films/the-multiplication/manifest.yaml": "44182359d2bce6942fbcba46eaca44a495722b9521ac7d8a8c6e31cb1d1be902", diff --git a/index.html b/index.html index d3831a8..439e767 100644 --- a/index.html +++ b/index.html @@ -201,7 +201,7 @@
Public record — every marked claim bound in the ledger - Registry v0.4 · 17 claims + Registry v0.4 · 19 claims CI prosecution: weekly + every push
@@ -691,7 +691,7 @@

Contact

Ledger v0.4 — generated from claims.yaml, drift-checked Triggers: executable where CI can watch, manual declared where it can't - Last owner review 2026-09-07 + Last owner review 2026-09-08
diff --git a/ledger/index.html b/ledger/index.html index 886375a..5cf7c2a 100644 --- a/ledger/index.html +++ b/ledger/index.html @@ -35,7 +35,7 @@ "sameAs": "https://github.com/Cubits11/cubits11.github.io/blob/main/claims.yaml", "isAccessibleForFree": true, "creator": { "@type": "Person", "name": "Pranav Bhave", "url": "https://cubits11.github.io/" }, - "dateModified": "2026-09-07", + "dateModified": "2026-09-08", "distribution": [{ "@type": "DataDownload", "encodingFormat": "application/yaml", "contentUrl": "https://cubits11.github.io/claims.yaml" }] } @@ -103,8 +103,8 @@

Evidence ledger

CI's reach.

Schema v0.4 - Last owner review: 2026-09-07 - 17 claims + Last owner review: 2026-09-08 + 19 claims CI runs ↗
@@ -248,6 +248,22 @@

SITE-002

This site's claim registry is enforced in CI: verify_claims.py validates schema and bindings, executes the executable review triggers against live evidence, and fails the build when any claim passes its freshness window — on every push and weekly.

Scope
The verifier at the recorded content hash, and the public workflow that runs it. The generated-page and figure checks are separate steps in the same workflow (ledger, modules, observatory, figure geometry, internal links).
Binding
verify_claims.py
Mutable link by design; the local-content trigger below detects edits.
Dimensions
visibilityPublicprovenanceOwner-verifiedsupport roleSite documentmaturityReleased
Reviewed
2026-08-30 · window 120 days
Triggers
  • executable fires when the verifier changes without a registry re-review
  • manual workflow triggers or cadence change
Falsifier
A controlled violation of a required registry field, bound-evidence trigger, or review window reaches a successful named CI workflow, or the workflow no longer runs on pushes to main and weekly.
CONSEQUENCE REJECT
Forbidden rescues
  • do not cite a normal green run as evidence that violations are rejected
  • do not count a manual review or a generated-page drift check as automatic schema, trigger, or freshness enforcement
  • do not treat a warning-only or report-only job as a build failure
Non-claims
  • a green run verifies registry consistency and quiet triggers, never the truth of any claim's content
  • an UNDETERMINED run (exit 2) blocks the build because a source was never reached; it is not a finding about any binding, and must not be read or recorded as a failed check
  • triggers watch file content; semantic drift outside watched files remains a manual review event, and the ledger says so
+
+
+

E3-001

+ Supported within scope · Machine-generated, owner-executed +
+

E3 — the first pilot this repository ran itself — scored two ungated classifiers on 400 harmful and 800 benign items and produced 2,400 committed observation rows. Its primary pre-registered prediction FAILED: excess joint miss was +0.0018 with a 95% bootstrap CI of [-0.00096, +0.00706], which includes zero. Both guards missed almost everything on this pool (0.9825 and 0.9625), so the Fréchet interval the two marginals allow is [0.9450, 0.9625] — 1.75 percentage points wide — and the observed joint miss of 0.9475 lies inside it. The prediction that it would lie inside HELD; the difficulty-stratification prediction was NOT COMPUTED and remains open.

+
Scope
Exactly the 2,400 rows committed at experiments/e3/results/observations.jsonl, produced 2026-09-06 by two research classifiers — protectai deberta-v3-base-prompt-injection-v2 (0.2B) and dcarpintero pangolin-guard-base (0.1B) — at thresholds frozen in e3_config.json (sha256 e163a2f2…) before any harmful item was scored. One pool (or-bench-toxic at e36d8b80), one operating point each, static full exposure. scripts/verify_e3.py recomputes every quantity below from those rows alone, including the bootstrap interval, which is a deterministic function of the committed rows under the frozen seed. Nothing here transfers to E2's guards, pools or operating points, and nothing here is about a deployed system.
Binding
observations.jsonl
The support is the measurement itself: 2,400 per-item, per-guard rows this repository produced. RESULT.md is the interpretation of those rows, not the evidence for them, and is hash-pinned by a trigger below so an edit to a quoted prediction fires a re-review. scripts/verify_e3.py re-derives every registered number from the rows alone on every push.
Dimensions
visibilityPublicprovenanceMachine-generated, owner-executedsupport roleExecuted outputmaturityExperimental
Reviewed
2026-09-08 · window 120 days
Triggers
  • executable fires when a committed observation row changes without a registry re-review
  • executable fires when the recorded run result changes without a registry re-review
  • executable fires when the run report changes, including any edit to a quoted prediction
  • manual a third pilot is run on a pool where both guards' miss rates are intermediate
Falsifier
Recomputing from the committed observation rows yields any quantity different from the expected block beyond 1e-12, or the observed joint miss is shown to lie outside the Fréchet interval its own marginals allow, or a quoted prediction in the bound run report is shown to have been edited after the outcome was visible.
CONSEQUENCE REJECT
Forbidden rescues
  • do not re-threshold, re-draw, or re-pool after the outcome and report the new numbers as this run
  • do not re-run the bootstrap under a different seed or B and report the resulting interval as this one
  • do not restate the failed primary prediction, narrow it, or drop it from the record
  • do not treat E3B as a re-run, a correction, or a replacement of this result
Non-claims
  • the 2,400 rows prove the instrument runs end to end; they do not prove it measures what the programme says it measures, and on this pool the identified set was nearly a point, so it measured almost nothing
  • not evidence that either classifier is good or bad; two research models, one pool, one operating point each
  • the null is a null at this scale on this pool, not evidence of independence and not evidence about any real guardrail's dependence
  • says nothing about E2, its three frozen guards, its pools, or its operating points
  • no vendor, product, deployed stack, or population is described
+
+
+
+

E3B-001

+ Supported within scope · Machine-generated, owner-executed +
+

E3B put the same two classifiers on the attack family they were built for — 400 real prompt injections — and produced a further 2,400 committed observation rows. Its primary prediction FAILED and so did the prediction the redraw existed to test. One guard missed nothing (0.0000, catching 400 of 400) and the other missed 0.3975, so the Fréchet interval the marginals allow is [0.0000, 0.0000] — zero points wide — the observed joint miss is exactly 0.0000, and the bootstrap CI is the degenerate [0, 0]. E3B was pre-registered to produce an interval wider than 10 percentage points; it produced one narrower than E3's. The prediction that the observed joint miss would lie inside the interval HELD, trivially.

+
Scope
Exactly the 2,400 rows committed at experiments/e3b/results/observations.jsonl, produced 2026-09-06 by the same two classifiers at thresholds frozen in e3b_config.json (sha256 252a5db9…) before any injection item was scored. One pool (Lakera gandalf_ignore_instructions at 04737b65, MIT), one operating point each, static full exposure. This is a new experiment with its own freeze, not a re-run of E3: E3's result stands whatever this shows, and both are registered. scripts/verify_e3.py recomputes every quantity below from the committed rows alone. The guard that caught 400 of 400 has two live explanations — generalisation and near-duplicate leakage from an undisclosed corpus — and nothing in this run separates them.
Binding
observations.jsonl
The support is the measurement itself: 2,400 further per-item, per-guard rows. RESULT.md is the interpretation, hash-pinned by a trigger below. scripts/verify_e3.py re-derives every registered number from the rows alone on every push.
Dimensions
visibilityPublicprovenanceMachine-generated, owner-executedsupport roleExecuted outputmaturityExperimental
Reviewed
2026-09-08 · window 120 days
Triggers
  • executable fires when a committed observation row changes without a registry re-review
  • executable fires when the recorded run result changes without a registry re-review
  • executable fires when the run report changes, including any edit to a quoted prediction
  • manual the contamination status of either guard's training corpus becomes checkable
Falsifier
Recomputing from the committed observation rows yields any quantity different from the expected block beyond 1e-12, or a quoted prediction in the bound run report is shown to have been edited after the outcome was visible, or E3B is represented anywhere in this repository as a re-run, correction, or replacement of E3.
CONSEQUENCE REJECT
Forbidden rescues
  • do not re-threshold, re-draw, or re-pool after the outcome and report the new numbers as this run
  • do not present E3B as a repair of E3, or E3 as superseded by it
  • do not restate either failed prediction, narrow it, or drop it from the record
  • do not treat the zero-width interval as a measured absence of dependence
Non-claims
  • a zero-width identified set means the marginals already fixed the joint miss; it is not a measurement that the two guards fail independently
  • G1 catching 400 of 400 is not evidence that it is a good guardrail, and contamination is unverified rather than excluded
  • G2 missing 159 of 400 is not evidence that it is a bad one
  • 2,400 further rows prove the instrument runs; on this pool the question was degenerate, so they measure almost nothing
  • the sharpening these two pilots suggest — that marginal-only reporting is uninformative in the middle of the marginal range and fully informative at its extremes — is a hypothesis and an engineering design gate, not a result of either pilot; neither run tested it
  • no vendor, product, deployed stack, or population is described
+