Skip to content

E3–E7B registrations, S5 contract, joint-cell render, E8 runner, prepared issues - #20

Merged
Cubits11 merged 19 commits into
mainfrom
claude/stack-merge
Sep 17, 2026
Merged

Cubits11 merged 19 commits into
mainfrom
claude/stack-merge

Conversation

@Cubits11

Copy link
Copy Markdown
Owner

The release stack, 18 commits over main, merged with the issue-form ingress fix (#19). Contents, in order:

  • E3 and E3B registered (E3-001, E3B-001); MC-005 attribution; ledger snapshot gated
  • E6 measured and registered; E7 voided on a defective preregistration with its diagnosis; E7B preregistered and all five predictions held; E6-001 narrowed by the chain-rule non-claim
  • three dossiers and the campaign brief, prepared, nothing sent
  • corrected evidence, addressable claims, verified films; claim-history transitions counted across merges once (44)
  • S5: identification adequacy, precision and decision relevance separated in the contract schema
  • the joint cell made visible through the bead-cube Blender harness (E3: 379 of 400 inside [378, 385]; E3B: 0 by structure), receipt-pinned; E8 runner that consumes a validated contract, no contract frozen, no rows; three issue drafts, prepared only; cadence row 2026-09-17
  • merge of issue forms: list on GitHub again; verifier fails on a description over 200 characters #19: verify_consequence covers every prefill surface; repo graph regenerated

Gates on cc7c7ec: verification_manifest 83 checks exit 0 (venv, Python 3.14 with requirements.txt). Review as a merge commit. Nothing here sends anything, freezes E8, or touches E2 collection.

🤖 Generated with Claude Code

Cubits11 and others added 19 commits September 8, 2026 02:08
…apshot gated

Four breaks between the speaking surfaces and the register are closed.

1. Stale counters. films/EXPLAINER-SERIES.md Episode 10 said "observation rows:
   zero" three days after 4,800 rows were committed. Episode 10 is rewritten in
   full, its central sentence replaced, and every counter the deck speaks now
   reads from metrics/ledger_snapshot.json with that file's as_of date spoken
   aloud. scripts/ledger_snapshot.py --check is the drift gate that would have
   caught it: it fails when the repository moves and the recorded counts do not,
   and it deliberately excludes blocker ages, which move with the calendar.
   /now/ gains the E3/E3B entry; index.html's claim count and owner-review date
   follow the registry.

2. MC-005 is explicitly non-canonical everywhere. It stays retracted: the kernel
   refuses a second registration for a claim that already has a protected state,
   and registering the same content under a fresh id would reverse a recorded
   owner judgement without a recorded reason. The W1 selection-regret numbers
   (regret <= 2 items / 2.4 points, 0 of 11 CIs excluding zero, 45 of 45
   positive excess joint miss) keep their artifact at
   experiments/e2/results/retrospective/ and are labelled unregistered on every
   surface that speaks them. scripts/verify_retracted.py derives retracted ids
   from the history and fails any line that names one without saying so.
   Frozen experiment artifacts are excluded from that gate on purpose: a freeze
   is immutable by contract and an annotation inside one would be the forbidden
   rescue.

3. exclusive_cells was already closed and the deck had not caught up. The
   32-cell block is committed in MC-002's expected, declared by a CLARIFY
   transition, re-asserted from the hash-verified release by
   reanalyze_bells_subset.py, and required by generate_missing_column.py and
   verify_figures.py. Blocker B1 is marked closed with those locators.

4. The BELLS licence disagreement is closed with one evidence-backed status:
   MC-002's "none declared upstream" with commercial_reuse facts_only. Nothing
   infers a licence from silence. B2, the ASTRA brief and the FABLE brief are
   updated to match, and the 2026-09-05 OBS CUT carries a dated superseding
   note rather than a rewrite.

E3-001 and E3B-001 are registered at exactly the strength their rows earn.
Support is the observation file itself, so the ledger now reads 2 of 19 claims
resting on own measurement rather than 0. Both commitments lead with the failed
prediction, carry prediction_verdicts in expected, and forbid re-thresholding,
re-seeding or restating a prediction after the outcome. The sharpening the two
pilots suggest -- that marginal-only reporting is uninformative in the middle of
the marginal range and fully informative at its extremes -- is registered as a
non-claim in those words, with the ~40-item pre-scoring check described as an
engineering gate, because neither pilot tested it.

scripts/verify_e3.py re-derives every registered quantity, including the
bootstrap interval, from the committed observation rows alone -- no model, no
network, no analyzer -- and asserts it against both the run result file and the
registry. claims_history verify passes at 44 entries with the prefix rule
satisfied and 19 live commitments equal to the chain tip.

No validator was weakened, no failed prediction rewritten, no provenance
invented, and no governance document added: the three new files are a verifier,
a gate and a generated snapshot.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019DdNofiEikbis8CvokH7iS
scripts/distribute.py run, offline. Only generated traction artifacts and the
repo graph move; no approval added, no draft approved, nothing dispatched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019DdNofiEikbis8CvokH7iS
… model

Extracts the client-side model kernel that anthropic.com/institute/econ-scenarios
ships, gates it against the twelve printed numbers of Table 3 of the Anthropic
Institute's Working Paper 2026-02, and evaluates it only at parameter vectors
built from the five marginal quantiles that paper prints in Table 2.

The kernel reproduces all twelve Table 3 numbers to the printed decimal and GDP
is strictly increasing in each of the five parameters over their published
interquartile ranges. Holding the four parameters the note to Table 4 names at
that note's values, the same five marginals admit a median 2030 GDP anywhere in
[2.12, 18.07] percent above the no-AI path across the 1,296 rank-permutation
couplings of their published quartiles: 9.88 at the comonotone corner, 8.32
under independence, against the 8.6 the paper obtained by running each of 3,259
respondents' own five-vector. The marginals fix an interval fifteen points wide,
not a point, and the measured joint lies strictly inside it.

The paper is the authority on the distinction. Its Appendix B states that the
Table 2 medians are taken "item by item, so no single respondent need give all
five median answers," while Table 4 requires "all five answers from the same
person." The words joint, correlation and copula do not appear in it.

Not preregistered, and recorded as not preregistered: the artifacts were public
before this directory existed, so no prediction here held because none was made.
No error in the paper is claimed — Table 4 does the joint-preserving computation
and does it correctly. The site's "GDP is 10% higher" figure is recorded as an
unidentified estimand, not as a composition of marginals: the vector of medians
gives 9.88 and an independent-coupling mean gives 10.01, and the public record
does not say which the sentence names.

The pinned kernel bytes are a third-party artifact with no declared licence and
are not redistributed; freeze/sources.json records that and freeze/cache/
enforces it. scripts/verify_e6.py re-derives all 34 registered quantities from
the 243 committed rows alone, with no network and no re-run of the model, and is
now in the verification manifest.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
films/data/facts.json rebound (96 facts; the only changes are the new E6-001
registry row, the claims.yaml input hash and last_owner_review), observatory
regenerated from claims.yaml, ledger snapshot refreshed to 20 claims / 3 resting
on own measurement / 5,043 observation rows.

The 12 stale film render receipts this exposes are PRE-EXISTING: at ab41c8f the
receipts already pinned facts.json 0c2591dd while the committed file was
15136fc2. The same 12 films fail before and after this change, and
bind_facts --check goes from failing to passing. Re-rendering needs the
Chrome/CDP toolchain and would touch same-scores__social-square, which is held
at the cold-viewer gate for the live x-film-same-scores campaign, so it is left
for a deliberate pass rather than done as a side effect here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Seven frontier_*_bench.json judge score files from
shawnray-research/certified-agent-guardrails @ 79097583 have NOT been retrieved
at this commit. This file fixes the estimator (per-judge threshold at the highest
value whose benign false-flag rate stays under 5%, the rule E3 used) and five
predictions about pairwise joint miss rates among those seven judges, before any
of their scores are visible.

The nine judges sharing the 132-item pool HAVE been read and are recorded as
exploratory, not as a test. The commit order is the evidence: verify_e7.py
refuses to record a confirmation unless this commit is an ancestor of the
results commit.

No model is loaded and no API is called — every score is already published — so
this is outside the scope of the host refusal in e2/run/adapters.py.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The prereg fixed the operating point as "the highest threshold whose false-flag
rate ... does not exceed 5% -- the same rule E3 used". Those clauses contradict
each other and the code implemented the first. Raising a flag threshold lowers
the false-flag rate monotonically, so the highest threshold inside any FPR
budget is the top of the score range, where the gate flags nothing: all three
surviving frontier judges came back with miss rate 1.000, every pair degenerate,
P1-P5 NOT_EVALUABLE.

experiments/e3/run/calibrate.py had already caught this and written the fix as
DISCREPANCY D1, before E3 scored a single harmful item. The instrument was right
and the new preregistration was wrong.

The fix is not applied to this hold-out. E7's own forbidden_rescues list bars
re-thresholding after a hold-out number is seen, so the seven frontier judges
are spent and no preregistered statement about them is available from here any
more.

EXPLORATORY.md records the nine-judge, single-threshold numbers that prompted
the experiment (mean excess joint miss +0.107, median 81% of the way to the
Frechet upper bound, 36 of 36 pairs positive) explicitly as exploratory, with
its weaknesses stated: n=35, one global threshold across incommensurable score
scales, 36 dependent pairs, and the phenomenon already found and named by the
source artifact's own ensemble_robustness.py.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Nine judge score files in shawnray-research/certified-agent-guardrails @
79097583 -- groq70b, mistral, and the seven nim_* files -- have NOT been
retrieved at this commit. Only their filenames and byte sizes, from the git tree
listing, are known here.

States the operating-point rule once and in one direction, with the sign of the
monotonicity named, so it cannot be read the two contradictory ways that voided
E7: the LOWEST threshold among a judge's benign scores at which its benign
false-flag rate is at or below 5%. That is what e3/run/calibrate.py implements.

Adds a power floor E7 lacked: fewer than 6 surviving non-degenerate pairs is
recorded UNDERPOWERED and no prediction is scored, so a thin result cannot be
read as a weak confirmation. Lowering that floor after the fact is a declared
forbidden rescue.

The 16 read judges (E7's exploratory nine and the void E7's seven frontier
judges) are excluded and may not be pooled in.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Six of nine hold-out judges reached the 5% benign false-flag budget; three were
excluded by the pre-stated rule. Shared pool 96 items (21 injection goals, 75
benign), 15 non-degenerate pairs against a floor of 6.

  P1 median delta > 0        HELD  +0.2018
  P2 mean delta >= +0.05     HELD  +0.1852
  P3 median reach >= 0.60    HELD   0.900
  P4 all q_obs inside bounds HELD   15/15
  P5 q_obs > q_ind on >=2/3  HELD   15/15

The modal pair: two judges each missing about half the injection goals,
independence predicting a 0.249 both-miss rate, the observed both-miss rate
0.476 -- which is exactly the Frechet upper bound. One judge's misses are a
subset of the other's. Seven of fifteen pairs sit at the bound; the median pair
sits 90% of the way up its interval. Adding the second judge bought 4.8 points
where independence promised 27.5.

This is the first non-degenerate preregistered measurement here. E3 and E3B
produced 4,800 rows against marginals of 0.98 and 0.96, where the interval is
1.75 points wide and the answer is nearly forced; these marginals land between
0.43 and 0.52.

No model was loaded and no API called -- every score was already published -- so
the host refusal in e2/run/adapters.py is not engaged and no owner action was
needed.

Priority is recorded, not claimed: arXiv:2607.22868v1 states the bound as a
proposition and its own ensemble_robustness.py reports the correlation. E7B
measures against a prereg; it claims neither the bound nor the phenomenon.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…laim

E7B-001 enters the registry with its five HELD verdicts, the prereg commit
f646e13 recorded as the evidence that the hold-out was unread, and priority for
the bound and the phenomenon explicitly disclaimed to arXiv:2607.22868v1.

E6-001 gains the objection the adversarial sweep raised hardest against it: the
five elicited quantities are conditional, so their product within one respondent
is the chain rule and is exact. E6 varies the coupling across the 10,980
respondents, not the composition of one respondent's conditionals. RESULT.md
gained a section stating this before a reader can raise it, and the transition is
declared NARROW in the claim-history chain.

Derived surfaces regenerated: 21 claims, 4 resting on own measurement.
819 manifest checks pass; the 12 stale film render receipts remain the
pre-existing failure recorded at 0f75088.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The chain-rule distinction and its matching non-claim were written and the
claim's sha256 trigger was re-pinned to the new content, but the file itself was
left out of ea720a2's add list. verify_claims passed locally because the working
tree matched the pin; a fresh clone would have failed it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ng sent

anthropic-econ-scenarios-2026: the one-row ask. Opens with what the report does
right (Table 4 is the joint-preserving computation; Appendix B states the
distinction; keeping the joint cost 70% of the sample) and asks only for the
model's output at the Table 2 median vector beside Table 4's 8.6, plus which
summary the site's 10% names. Explicitly forbids sending any claim that they
composed marginals or that the report contains an error.

certified-agent-guardrails-2026: credit, no ask. arXiv:2607.22868v1 states the
Frechet bracket for an any-flag gate as a proposition and its own
ensemble_robustness.py had already named the correlated failure. Records that
nothing here may claim priority for either, and that the release is a
PRESENT-by-computation census row whose item sets are nested (96 in 132 in 167),
so any 23-judge joint must be reported on the 96-item core.

ari-defense-in-depth-2025: one clause. The 90%^5 = 0.001% sentence is the
independence point where the Frechet upper bound is 0.10, 10,000x larger. Carries
a fairness section requiring all three verified mitigations into any message --
the "(assuming independence)" parenthetical is theirs, the piece hedges
elsewhere, and the sentence is a hypothetical illustration over six
heterogeneous layers, not a measured joint statistic. Channel note forbids a
public quote-post.

CAMPAIGN-2026-09-10.md: sequencing (credit before claim, ask before post, E7B
before Anthropic), two unapproved X thread drafts, a one-liner, and a LessWrong
packet rather than a draft -- the spine, the numbers, the four strongest
objections with what actually answers them, and the citations that must appear
(Embrechts et al. 2014; Dung & Mai arXiv:2510.11235; the UK AISI safety-case
post whose stated independence limitation is the best single hook).

approvals.json is untouched: no draft here carries an owner approval, so
scripts/distribute.py publish cannot dispatch any of it. The x-film-same-scores
cold-viewer gate is not touched or substituted for.

claims_history.yaml reconstruction count 41 -> 42 with a note; the 42nd is
E6-001 at ea720a2, covered by a contemporaneous NARROW declaration.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… count at 44

At 48e98e4 the reconstruction check fails: the record states 42/9 and git
yields 44/9. The two additional events are the E6-001 and E7B-001
corrections committed there. Both are covered by existing contemporaneous
declarations; the strict-eligible set is unchanged.

A merge commit would also have broken the count. The first-parent walk sees
a merge as one step, so a merged branch's own transitions collapse into it.
Rehearsed on a scratch clone: main plus a no-ff merge of 48e98e4 recomputes
41/9 against a stated 42/9.

Pre-genesis events keep the declared first-parent reconstruction. After the
genesis anchor, each commit is compared with its actual parents, and a state
inherited from any parent is not counted again. On a linear history the two
methods agree. With this change applied, the same rehearsal recomputes 44/9.

tests/test_history_merges.py fixes the property: a merged branch's changes,
and a merged removal, each survive exactly once. It runs in the manifest.
distribute.py run rebinds the 9 event sources that were pending_commit at
the previous head; 0 remain unbound. claims_history.yaml now binds to
5a39b06. draft-history.json appends the rebound revisions (20 -> 24
entries) and leaves every earlier entry unchanged. The repo graph is
regenerated at the same head.

approvals.json, publications.json, metrics.json, interactions.json and
deviations.json are byte-identical. Nothing was approved, dispatched or
published.
…elevance

DIRECTION.yaml step S5. Correction C2 asks that three quantities stop being one
power number, and that they be declared before an expensive run rather than
after its outcomes are visible.

schema.json gains an `inference` block in three required groups — what the
identified set permits before any n, what sampling is asked to resolve inside
it, and what the decision requires of both. It is required at `status: frozen`
only. validate_contract.py checks each group on its own terms and evaluates the
declared minimum-information condition in exact decimal arithmetic; the
boundary is admitted and only a strict overrun is rejected. An unimplemented
condition name fails closed instead of passing silently.

Two fixtures: a passing frozen control, and a design whose identified set is
wider than its target admits — the case no sample size fixes, which the old
"CI crossed zero therefore underpowered" reading would have mislabelled. Ten
tests, including one showing the joint condition is not implied by the three
group checks, and one boundary case a float implementation would reject. Each
of the six guards was disabled in turn; every one is covered by a test that
fails without it.

The five retrospective fixtures stay `constructed_fixture` and are not
retrofitted; a test holds that none of them is rejected for a missing block.
Nothing here establishes that a declared identification width is the width a
design actually has — the block compares declared numbers with each other.

S5's status in research/DIRECTION.yaml and the grade of plane edge Q are owner
records and are left unchanged.
…s prepared

Render: films/lib/blender/build_frechet_slots.py draws, through the bead-cube
harness, the integer band the E3 and E3B marginals fix for the both-miss
count and lights the slot the rows recorded: E3 379 of 400 inside [378, 385],
E3B 0 inside [0, 0]. Every count is recounted from committed rows, cross-checked
against each run's result file, and pinned to the commit that last touched its
source. Two renders, identical IDAT. CAPTION.md states what the frame does not
show.

S3: experiments/e8/run/runner.py takes comparator, direction, budget, pools and
exclusion policy from a contract validate_contract.py has accepted, refuses an
unfrozen one, stores per-item score vectors and derives cells. tests/test_e8_runner.py
shows a changed contract changes the selection. No E8 contract is frozen; the
runner has produced no row, and S3 stays `next` until the owner freezes one.

Issues, prepared and not sent, in distribution/issues/: a counterexample
invitation against E3-001's both-miss count, the E3-001 reproduction gap, and
the GuardBench upstream ask from its dossier, each one click from the form.

Found on the way: both YAML issue forms were unlisted by GitHub since
2026-09-01 because their descriptions exceeded 200 characters, so every
prefilled link opened a blank issue. Descriptions shortened;
verify_consequence.py now fails on the limit.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…er 200 characters

GitHub lists an issue form only when its top-level description is 3–200
characters. counterexample.yml (305) and reproduction.yml (296) exceeded it
from 2026-09-01, so the chooser showed only the census row-correction
template and every prefilled link on /try/, the reproduce page, worldspace,
README, CONTRIBUTING and DISPATCH opened a blank issue. The file page on
github.com states the cause: "Description must be between 3 and 200
characters."

Both descriptions are shortened; fields, ids, options and validations are
unchanged. scripts/verify_consequence.py now fails when a description leaves
3–200, and checks every prefilled link on the five further surfaces, not
only /try/, against the listed forms and their fields.

The forms are listed only once this commit reaches the default branch.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Takes the ingress branch's verify_consequence.py, which covers every prefill
surface, and regenerates the repo graph over the merged tree.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@Cubits11
Cubits11 merged commit d3766ae into main Sep 17, 2026
5 checks passed
@Cubits11
Cubits11 deleted the claude/stack-merge branch September 18, 2026 01:40
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant