E3–E7B registrations, S5 contract, joint-cell render, E8 runner, prepared issues - #20
Merged
Merged
Conversation
…apshot gated Four breaks between the speaking surfaces and the register are closed. 1. Stale counters. films/EXPLAINER-SERIES.md Episode 10 said "observation rows: zero" three days after 4,800 rows were committed. Episode 10 is rewritten in full, its central sentence replaced, and every counter the deck speaks now reads from metrics/ledger_snapshot.json with that file's as_of date spoken aloud. scripts/ledger_snapshot.py --check is the drift gate that would have caught it: it fails when the repository moves and the recorded counts do not, and it deliberately excludes blocker ages, which move with the calendar. /now/ gains the E3/E3B entry; index.html's claim count and owner-review date follow the registry. 2. MC-005 is explicitly non-canonical everywhere. It stays retracted: the kernel refuses a second registration for a claim that already has a protected state, and registering the same content under a fresh id would reverse a recorded owner judgement without a recorded reason. The W1 selection-regret numbers (regret <= 2 items / 2.4 points, 0 of 11 CIs excluding zero, 45 of 45 positive excess joint miss) keep their artifact at experiments/e2/results/retrospective/ and are labelled unregistered on every surface that speaks them. scripts/verify_retracted.py derives retracted ids from the history and fails any line that names one without saying so. Frozen experiment artifacts are excluded from that gate on purpose: a freeze is immutable by contract and an annotation inside one would be the forbidden rescue. 3. exclusive_cells was already closed and the deck had not caught up. The 32-cell block is committed in MC-002's expected, declared by a CLARIFY transition, re-asserted from the hash-verified release by reanalyze_bells_subset.py, and required by generate_missing_column.py and verify_figures.py. Blocker B1 is marked closed with those locators. 4. The BELLS licence disagreement is closed with one evidence-backed status: MC-002's "none declared upstream" with commercial_reuse facts_only. Nothing infers a licence from silence. B2, the ASTRA brief and the FABLE brief are updated to match, and the 2026-09-05 OBS CUT carries a dated superseding note rather than a rewrite. E3-001 and E3B-001 are registered at exactly the strength their rows earn. Support is the observation file itself, so the ledger now reads 2 of 19 claims resting on own measurement rather than 0. Both commitments lead with the failed prediction, carry prediction_verdicts in expected, and forbid re-thresholding, re-seeding or restating a prediction after the outcome. The sharpening the two pilots suggest -- that marginal-only reporting is uninformative in the middle of the marginal range and fully informative at its extremes -- is registered as a non-claim in those words, with the ~40-item pre-scoring check described as an engineering gate, because neither pilot tested it. scripts/verify_e3.py re-derives every registered quantity, including the bootstrap interval, from the committed observation rows alone -- no model, no network, no analyzer -- and asserts it against both the run result file and the registry. claims_history verify passes at 44 entries with the prefix rule satisfied and 19 live commitments equal to the chain tip. No validator was weakened, no failed prediction rewritten, no provenance invented, and no governance document added: the three new files are a verifier, a gate and a generated snapshot. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019DdNofiEikbis8CvokH7iS
scripts/distribute.py run, offline. Only generated traction artifacts and the repo graph move; no approval added, no draft approved, nothing dispatched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019DdNofiEikbis8CvokH7iS
… model Extracts the client-side model kernel that anthropic.com/institute/econ-scenarios ships, gates it against the twelve printed numbers of Table 3 of the Anthropic Institute's Working Paper 2026-02, and evaluates it only at parameter vectors built from the five marginal quantiles that paper prints in Table 2. The kernel reproduces all twelve Table 3 numbers to the printed decimal and GDP is strictly increasing in each of the five parameters over their published interquartile ranges. Holding the four parameters the note to Table 4 names at that note's values, the same five marginals admit a median 2030 GDP anywhere in [2.12, 18.07] percent above the no-AI path across the 1,296 rank-permutation couplings of their published quartiles: 9.88 at the comonotone corner, 8.32 under independence, against the 8.6 the paper obtained by running each of 3,259 respondents' own five-vector. The marginals fix an interval fifteen points wide, not a point, and the measured joint lies strictly inside it. The paper is the authority on the distinction. Its Appendix B states that the Table 2 medians are taken "item by item, so no single respondent need give all five median answers," while Table 4 requires "all five answers from the same person." The words joint, correlation and copula do not appear in it. Not preregistered, and recorded as not preregistered: the artifacts were public before this directory existed, so no prediction here held because none was made. No error in the paper is claimed — Table 4 does the joint-preserving computation and does it correctly. The site's "GDP is 10% higher" figure is recorded as an unidentified estimand, not as a composition of marginals: the vector of medians gives 9.88 and an independent-coupling mean gives 10.01, and the public record does not say which the sentence names. The pinned kernel bytes are a third-party artifact with no declared licence and are not redistributed; freeze/sources.json records that and freeze/cache/ enforces it. scripts/verify_e6.py re-derives all 34 registered quantities from the 243 committed rows alone, with no network and no re-run of the model, and is now in the verification manifest. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
films/data/facts.json rebound (96 facts; the only changes are the new E6-001 registry row, the claims.yaml input hash and last_owner_review), observatory regenerated from claims.yaml, ledger snapshot refreshed to 20 claims / 3 resting on own measurement / 5,043 observation rows. The 12 stale film render receipts this exposes are PRE-EXISTING: at ab41c8f the receipts already pinned facts.json 0c2591dd while the committed file was 15136fc2. The same 12 films fail before and after this change, and bind_facts --check goes from failing to passing. Re-rendering needs the Chrome/CDP toolchain and would touch same-scores__social-square, which is held at the cold-viewer gate for the live x-film-same-scores campaign, so it is left for a deliberate pass rather than done as a side effect here. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Seven frontier_*_bench.json judge score files from shawnray-research/certified-agent-guardrails @ 79097583 have NOT been retrieved at this commit. This file fixes the estimator (per-judge threshold at the highest value whose benign false-flag rate stays under 5%, the rule E3 used) and five predictions about pairwise joint miss rates among those seven judges, before any of their scores are visible. The nine judges sharing the 132-item pool HAVE been read and are recorded as exploratory, not as a test. The commit order is the evidence: verify_e7.py refuses to record a confirmation unless this commit is an ancestor of the results commit. No model is loaded and no API is called — every score is already published — so this is outside the scope of the host refusal in e2/run/adapters.py. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The prereg fixed the operating point as "the highest threshold whose false-flag rate ... does not exceed 5% -- the same rule E3 used". Those clauses contradict each other and the code implemented the first. Raising a flag threshold lowers the false-flag rate monotonically, so the highest threshold inside any FPR budget is the top of the score range, where the gate flags nothing: all three surviving frontier judges came back with miss rate 1.000, every pair degenerate, P1-P5 NOT_EVALUABLE. experiments/e3/run/calibrate.py had already caught this and written the fix as DISCREPANCY D1, before E3 scored a single harmful item. The instrument was right and the new preregistration was wrong. The fix is not applied to this hold-out. E7's own forbidden_rescues list bars re-thresholding after a hold-out number is seen, so the seven frontier judges are spent and no preregistered statement about them is available from here any more. EXPLORATORY.md records the nine-judge, single-threshold numbers that prompted the experiment (mean excess joint miss +0.107, median 81% of the way to the Frechet upper bound, 36 of 36 pairs positive) explicitly as exploratory, with its weaknesses stated: n=35, one global threshold across incommensurable score scales, 36 dependent pairs, and the phenomenon already found and named by the source artifact's own ensemble_robustness.py. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Nine judge score files in shawnray-research/certified-agent-guardrails @ 79097583 -- groq70b, mistral, and the seven nim_* files -- have NOT been retrieved at this commit. Only their filenames and byte sizes, from the git tree listing, are known here. States the operating-point rule once and in one direction, with the sign of the monotonicity named, so it cannot be read the two contradictory ways that voided E7: the LOWEST threshold among a judge's benign scores at which its benign false-flag rate is at or below 5%. That is what e3/run/calibrate.py implements. Adds a power floor E7 lacked: fewer than 6 surviving non-degenerate pairs is recorded UNDERPOWERED and no prediction is scored, so a thin result cannot be read as a weak confirmation. Lowering that floor after the fact is a declared forbidden rescue. The 16 read judges (E7's exploratory nine and the void E7's seven frontier judges) are excluded and may not be pooled in. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Six of nine hold-out judges reached the 5% benign false-flag budget; three were excluded by the pre-stated rule. Shared pool 96 items (21 injection goals, 75 benign), 15 non-degenerate pairs against a floor of 6. P1 median delta > 0 HELD +0.2018 P2 mean delta >= +0.05 HELD +0.1852 P3 median reach >= 0.60 HELD 0.900 P4 all q_obs inside bounds HELD 15/15 P5 q_obs > q_ind on >=2/3 HELD 15/15 The modal pair: two judges each missing about half the injection goals, independence predicting a 0.249 both-miss rate, the observed both-miss rate 0.476 -- which is exactly the Frechet upper bound. One judge's misses are a subset of the other's. Seven of fifteen pairs sit at the bound; the median pair sits 90% of the way up its interval. Adding the second judge bought 4.8 points where independence promised 27.5. This is the first non-degenerate preregistered measurement here. E3 and E3B produced 4,800 rows against marginals of 0.98 and 0.96, where the interval is 1.75 points wide and the answer is nearly forced; these marginals land between 0.43 and 0.52. No model was loaded and no API called -- every score was already published -- so the host refusal in e2/run/adapters.py is not engaged and no owner action was needed. Priority is recorded, not claimed: arXiv:2607.22868v1 states the bound as a proposition and its own ensemble_robustness.py reports the correlation. E7B measures against a prereg; it claims neither the bound nor the phenomenon. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…laim E7B-001 enters the registry with its five HELD verdicts, the prereg commit f646e13 recorded as the evidence that the hold-out was unread, and priority for the bound and the phenomenon explicitly disclaimed to arXiv:2607.22868v1. E6-001 gains the objection the adversarial sweep raised hardest against it: the five elicited quantities are conditional, so their product within one respondent is the chain rule and is exact. E6 varies the coupling across the 10,980 respondents, not the composition of one respondent's conditionals. RESULT.md gained a section stating this before a reader can raise it, and the transition is declared NARROW in the claim-history chain. Derived surfaces regenerated: 21 claims, 4 resting on own measurement. 819 manifest checks pass; the 12 stale film render receipts remain the pre-existing failure recorded at 0f75088. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The chain-rule distinction and its matching non-claim were written and the claim's sha256 trigger was re-pinned to the new content, but the file itself was left out of ea720a2's add list. verify_claims passed locally because the working tree matched the pin; a fresh clone would have failed it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ng sent anthropic-econ-scenarios-2026: the one-row ask. Opens with what the report does right (Table 4 is the joint-preserving computation; Appendix B states the distinction; keeping the joint cost 70% of the sample) and asks only for the model's output at the Table 2 median vector beside Table 4's 8.6, plus which summary the site's 10% names. Explicitly forbids sending any claim that they composed marginals or that the report contains an error. certified-agent-guardrails-2026: credit, no ask. arXiv:2607.22868v1 states the Frechet bracket for an any-flag gate as a proposition and its own ensemble_robustness.py had already named the correlated failure. Records that nothing here may claim priority for either, and that the release is a PRESENT-by-computation census row whose item sets are nested (96 in 132 in 167), so any 23-judge joint must be reported on the 96-item core. ari-defense-in-depth-2025: one clause. The 90%^5 = 0.001% sentence is the independence point where the Frechet upper bound is 0.10, 10,000x larger. Carries a fairness section requiring all three verified mitigations into any message -- the "(assuming independence)" parenthetical is theirs, the piece hedges elsewhere, and the sentence is a hypothetical illustration over six heterogeneous layers, not a measured joint statistic. Channel note forbids a public quote-post. CAMPAIGN-2026-09-10.md: sequencing (credit before claim, ask before post, E7B before Anthropic), two unapproved X thread drafts, a one-liner, and a LessWrong packet rather than a draft -- the spine, the numbers, the four strongest objections with what actually answers them, and the citations that must appear (Embrechts et al. 2014; Dung & Mai arXiv:2510.11235; the UK AISI safety-case post whose stated independence limitation is the best single hook). approvals.json is untouched: no draft here carries an owner approval, so scripts/distribute.py publish cannot dispatch any of it. The x-film-same-scores cold-viewer gate is not touched or substituted for. claims_history.yaml reconstruction count 41 -> 42 with a note; the 42nd is E6-001 at ea720a2, covered by a contemporaneous NARROW declaration. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… count at 44 At 48e98e4 the reconstruction check fails: the record states 42/9 and git yields 44/9. The two additional events are the E6-001 and E7B-001 corrections committed there. Both are covered by existing contemporaneous declarations; the strict-eligible set is unchanged. A merge commit would also have broken the count. The first-parent walk sees a merge as one step, so a merged branch's own transitions collapse into it. Rehearsed on a scratch clone: main plus a no-ff merge of 48e98e4 recomputes 41/9 against a stated 42/9. Pre-genesis events keep the declared first-parent reconstruction. After the genesis anchor, each commit is compared with its actual parents, and a state inherited from any parent is not counted again. On a linear history the two methods agree. With this change applied, the same rehearsal recomputes 44/9. tests/test_history_merges.py fixes the property: a merged branch's changes, and a merged removal, each survive exactly once. It runs in the manifest.
distribute.py run rebinds the 9 event sources that were pending_commit at the previous head; 0 remain unbound. claims_history.yaml now binds to 5a39b06. draft-history.json appends the rebound revisions (20 -> 24 entries) and leaves every earlier entry unchanged. The repo graph is regenerated at the same head. approvals.json, publications.json, metrics.json, interactions.json and deviations.json are byte-identical. Nothing was approved, dispatched or published.
…elevance DIRECTION.yaml step S5. Correction C2 asks that three quantities stop being one power number, and that they be declared before an expensive run rather than after its outcomes are visible. schema.json gains an `inference` block in three required groups — what the identified set permits before any n, what sampling is asked to resolve inside it, and what the decision requires of both. It is required at `status: frozen` only. validate_contract.py checks each group on its own terms and evaluates the declared minimum-information condition in exact decimal arithmetic; the boundary is admitted and only a strict overrun is rejected. An unimplemented condition name fails closed instead of passing silently. Two fixtures: a passing frozen control, and a design whose identified set is wider than its target admits — the case no sample size fixes, which the old "CI crossed zero therefore underpowered" reading would have mislabelled. Ten tests, including one showing the joint condition is not implied by the three group checks, and one boundary case a float implementation would reject. Each of the six guards was disabled in turn; every one is covered by a test that fails without it. The five retrospective fixtures stay `constructed_fixture` and are not retrofitted; a test holds that none of them is rejected for a missing block. Nothing here establishes that a declared identification width is the width a design actually has — the block compares declared numbers with each other. S5's status in research/DIRECTION.yaml and the grade of plane edge Q are owner records and are left unchanged.
…s prepared Render: films/lib/blender/build_frechet_slots.py draws, through the bead-cube harness, the integer band the E3 and E3B marginals fix for the both-miss count and lights the slot the rows recorded: E3 379 of 400 inside [378, 385], E3B 0 inside [0, 0]. Every count is recounted from committed rows, cross-checked against each run's result file, and pinned to the commit that last touched its source. Two renders, identical IDAT. CAPTION.md states what the frame does not show. S3: experiments/e8/run/runner.py takes comparator, direction, budget, pools and exclusion policy from a contract validate_contract.py has accepted, refuses an unfrozen one, stores per-item score vectors and derives cells. tests/test_e8_runner.py shows a changed contract changes the selection. No E8 contract is frozen; the runner has produced no row, and S3 stays `next` until the owner freezes one. Issues, prepared and not sent, in distribution/issues/: a counterexample invitation against E3-001's both-miss count, the E3-001 reproduction gap, and the GuardBench upstream ask from its dossier, each one click from the form. Found on the way: both YAML issue forms were unlisted by GitHub since 2026-09-01 because their descriptions exceeded 200 characters, so every prefilled link opened a blank issue. Descriptions shortened; verify_consequence.py now fails on the limit. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…er 200 characters GitHub lists an issue form only when its top-level description is 3–200 characters. counterexample.yml (305) and reproduction.yml (296) exceeded it from 2026-09-01, so the chooser showed only the census row-correction template and every prefilled link on /try/, the reproduce page, worldspace, README, CONTRIBUTING and DISPATCH opened a blank issue. The file page on github.com states the cause: "Description must be between 3 and 200 characters." Both descriptions are shortened; fields, ids, options and validations are unchanged. scripts/verify_consequence.py now fails when a description leaves 3–200, and checks every prefilled link on the five further surfaces, not only /try/, against the listed forms and their fields. The forms are listed only once this commit reaches the default branch. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Takes the ingress branch's verify_consequence.py, which covers every prefill surface, and regenerates the repo graph over the merged tree. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The release stack, 18 commits over main, merged with the issue-form ingress fix (#19). Contents, in order:
Gates on cc7c7ec: verification_manifest 83 checks exit 0 (venv, Python 3.14 with requirements.txt). Review as a merge commit. Nothing here sends anything, freezes E8, or touches E2 collection.
🤖 Generated with Claude Code