Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
19 commits
Select commit Hold shift + click to select a range
f3a6a1e
registry: E3 and E3B registered; MC-005 attribution and the ledger sn…
Cubits11 Sep 8, 2026
ab41c8f
distribution: regenerate derived artifacts against the reconciled commit
Cubits11 Sep 8, 2026
6728545
E6: the width of the missing column, measured on a published economic…
Cubits11 Sep 10, 2026
0f75088
regenerate derived artifacts against the E6 registration
Cubits11 Sep 10, 2026
da102bc
E7 preregistration, frozen before the hold-out was read
Cubits11 Sep 10, 2026
5b5ff8d
E7: void on a defective preregistration, and the diagnosis
Cubits11 Sep 10, 2026
f646e13
E7B preregistration, frozen before the hold-out was read
Cubits11 Sep 10, 2026
c542a07
E7B: all five preregistered predictions held on unread data
Cubits11 Sep 10, 2026
ea720a2
registry: E7B-001 registered, E6-001 narrowed by the chain-rule non-c…
Cubits11 Sep 10, 2026
b91324f
E6 RESULT.md: commit the chain-rule section the claim already pins
Cubits11 Sep 10, 2026
e7c6b9d
distribution: three dossiers and the campaign brief — prepared, nothi…
Cubits11 Sep 10, 2026
48e98e4
Publish corrected evidence, addressable claims and verified films
Cubits11 Sep 15, 2026
5a39b06
claims history: count transitions across merges once, and restate the…
Cubits11 Sep 15, 2026
5fe44bc
distribution: regenerate derived artifacts against the history repair
Cubits11 Sep 15, 2026
6a51abb
contracts: separate identification adequacy, precision and decision r…
Cubits11 Sep 17, 2026
c9dc338
joint cell made visible, E8 runner consumes its contract, three issue…
Cubits11 Sep 17, 2026
4173a8c
cadence: record 2026-09-17
Cubits11 Sep 17, 2026
026554c
issue forms: list on GitHub again; verifier fails on a description ov…
Cubits11 Sep 17, 2026
cc7c7ec
merge claude/issue-forms-ingress into the E3–E7B release stack
Cubits11 Sep 17, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
14 changes: 14 additions & 0 deletions .github/CODEOWNERS
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
# Review routing only. GitHub rules must enforce reviews separately.
# A single owner cannot independently approve their own pull request.
/.github/ @Cubits11
/security/ @Cubits11
/scripts/ @Cubits11
/tests/ @Cubits11
/requirements.txt @Cubits11
/claims.yaml @Cubits11
/claims_history.yaml @Cubits11
/experiments/ @Cubits11
/corrections/ @Cubits11
/research/ @Cubits11
/metrics/ @Cubits11
/distribution/traction/ @Cubits11
8 changes: 3 additions & 5 deletions .github/ISSUE_TEMPLATE/counterexample.yml
Original file line number Diff line number Diff line change
@@ -1,10 +1,8 @@
name: Counterexample, missed benchmark, or joint outcomes
description: >-
Bring evidence that changes the record: a counterexample to a claim, a
benchmark the census missed, a benchmark that already reports the stack,
a narrower identified region, or item-level / joint outcomes you can
provide. A correction is placed beside the claim it corrects, dated,
with this issue linked.
Evidence that changes the record: a counterexample, a missed or misread
census row, a tighter bound, or joint outcomes. Corrections land dated
beside the claim, with this issue linked.
title: "counterexample: <what it changes>"
labels: ["counterexample"]
body:
Expand Down
8 changes: 3 additions & 5 deletions .github/ISSUE_TEMPLATE/reproduction.yml
Original file line number Diff line number Diff line change
@@ -1,10 +1,8 @@
name: Reproduction run — match or mismatch
description: >-
File the result of running an experiment from /try/ or a registered
reproduction script. Both outcomes are wanted: a match becomes the
claim's independent-reproduction record; a mismatch is handled as a
correction under the same-day correction policy and is credited beside
the claim it corrects.
The result of running a /try/ experiment or a registered reproduction
script. A match becomes the claim's reproduction record; a mismatch is
handled as a credited correction.
title: "Reproduction: <experiment or claim id> — <match|mismatch>"
labels: ["reproduction"]
body:
Expand Down
39 changes: 39 additions & 0 deletions .github/workflows/evidence-boundary.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,39 @@
name: evidence-boundary

# This workflow belongs to the trusted base branch. NEVER check out, import,
# install from, or execute the PR head under pull_request_target.
on:
pull_request_target:
types: [opened, synchronize, reopened, ready_for_review]
permissions:
contents: read
concurrency:
group: evidence-boundary-${{ github.event.pull_request.number }}
cancel-in-progress: true
jobs:
evidence-boundary:
name: Trusted evidence boundary
runs-on: ubuntu-latest
timeout-minutes: 10
steps:
- uses: actions/checkout@fbc6f3992d24b796d5a048ff273f7fcc4a7b6c09 # v5
with:
ref: ${{ github.event.pull_request.base.sha }}
fetch-depth: 0
persist-credentials: false
- uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # v6
with:
python-version: '3.12'
- run: python -m pip install PyYAML==6.0.3
- name: Read candidate Git objects without checking them out
env:
PR_NUMBER: ${{ github.event.pull_request.number }}
CANDIDATE_SHA: ${{ github.event.pull_request.head.sha }}
BASE_SHA: ${{ github.event.pull_request.base.sha }}
run: |
[[ "$PR_NUMBER" =~ ^[0-9]+$ ]] || exit 2
[[ "$CANDIDATE_SHA" =~ ^[0-9a-f]{40}$ ]] || exit 2
[[ "$BASE_SHA" =~ ^[0-9a-f]{40}$ ]] || exit 2
git fetch --no-tags origin "refs/pull/$PR_NUMBER/head"
test "$(git rev-parse FETCH_HEAD)" = "$CANDIDATE_SHA" || exit 2
python scripts/evidence_guard.py --base "$BASE_SHA" --candidate "$CANDIDATE_SHA"
11 changes: 8 additions & 3 deletions .github/workflows/verify.yml
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,11 @@ on:
permissions:
contents: read

# Superseded PR revisions need no runners; main deployments finish intact.
concurrency:
group: ${{ github.workflow }}-${{ github.event_name }}-${{ github.ref }}
cancel-in-progress: ${{ github.event_name == 'pull_request' }}

jobs:
claims:
name: Claim registry
Expand All @@ -36,7 +41,7 @@ jobs:
- uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # v6
with:
python-version: "3.12"
- run: pip install pyyaml
- run: python -m pip install -r requirements.txt
# Both CI and the clean-clone replay execute this source-controlled,
# duplicate-free manifest. A check cannot quietly run in one green
# surface but not the other.
Expand All @@ -55,7 +60,7 @@ jobs:
- uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # v6
with:
python-version: "3.12"
- run: pip install pyyaml
- run: python -m pip install -r requirements.txt
- name: Clean-clone the bound commit and re-run the kernel
run: python scripts/reproduce_cc001.py
- name: Clean-clone this site revision and replay deterministic gates
Expand Down Expand Up @@ -112,7 +117,7 @@ jobs:
- uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # v6
with:
python-version: "3.12"
- run: pip install pyyaml
- run: python -m pip install -r requirements.txt
- name: Wait for every route, census checksum, counts, and corrections policy
run: python scripts/smoke_deployed.py --attempts 30 --interval-seconds 10

Expand Down
2 changes: 2 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
@@ -1,8 +1,10 @@
__pycache__/
*.pyc
_private/
/Claude outputs/
.venv/
.DS_Store

# Google Drive sync scratch
.tmp.driveupload/
.tmp.drivedownload/
14 changes: 12 additions & 2 deletions ARTIFACTS/12-WEEK-PROGRAM.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,12 +18,22 @@ the second-guard selection by at most 2.4 points, and in the largest
end-to-end measurement the stack was statistically indistinguishable from
its single strongest member.**

> **Registration status, 2026-09-08.** This fact is **UNREGISTERED**. The W1
> artifact is committed at `experiments/e2/results/retrospective/` — run report,
> three matrices, and `independent_t1.py`, which recomputes T1 from the raw
> released rows outside the analyzer's path. No claim id carries these numbers
> and no CI check re-asserts them: the claim that used to (retracted 2026-09-06)
> was withdrawn over a licence defect in its support block, and the retraction
> reason records that the computation itself is untouched. Every surface that
> speaks these numbers must say "computed 2026-09-02, unregistered" in the same
> breath; `scripts/verify_retracted.py` gates the attribution.

Three sources, all VERIFIED this run:

1. **BELLS 2025 released subset** (`non_adversarial_prompts.csv @ 507566c5`,
sha256 `791dd4b0…`, 170 rows, 11 verdict columns). Computed 2026-09-02 on
the hash-verified file (scratchpad script; to be committed as the W1
artifact), harmful stratum n=82, native points:
the hash-verified file — the W1 artifact is now committed at
`experiments/e2/results/retrospective/` — harmful stratum n=82, native points:
- Five specialized supervisors: for **0 of 5** incumbents does the
partner maximizing measured union catch differ in union from the
partner chosen by marginal rank. Regret 0 items.
Expand Down
25 changes: 20 additions & 5 deletions ARTIFACTS/2026-09-05-FABLE-5.1-OBS-CUT.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,16 @@
# OBSERVATION CUT — BELLS denominators and provenance

> **Superseding note, 2026-09-08 (annotation only; the cut below is unchanged).**
> Both open items this cut named are closed. MC-005 was **retracted** on
> 2026-09-06 — the RETRACT is entry 39 of `claims_history.yaml` and its reason
> is the licence disagreement recorded below — so the registry now has exactly
> one record of the BELLS file's licence: MC-002's `none declared upstream`,
> with `commercial_reuse: facts_only`. Nothing in this repository infers a
> licence from silence. The `exclusive_cells` block this cut found uncommitted
> is committed, declared by an MC-002 CLARIFY transition, and re-asserted from
> the hash-verified file by `scripts/reanalyze_bells_subset.py` in CI. The
> 990-versus-1041 denominator remains unreconciled, exactly as this cut left it.

Written 2026-09-05 against working tree of branch `claude/mc-005-selection-regret`
(HEAD `f9b24e2`, uncommitted owner changes present and untouched). Status
vocabulary: OBSERVED (bytes opened or command executed in this run) · DERIVED
Expand Down Expand Up @@ -34,7 +45,7 @@ nothing redistributed) or `--cache DIR` offline. Exit 0 this run.
| C | `077555d9` (2025-02-19, "smaller dataset for playground") | same path, sha `7fa0fbf5885e…` | 174 | 86 / 50 / 38 | 12 | as B | `prompt_guard=1`, `langkit=1` | rest | **SELECTION RULE: UNKNOWN** (commit subject only: "smaller dataset for playground") | selected subset | none |
| D | `00b42bfd` (2025-02-21, "New data, new results interpretation") | same path, sha `52ca9ab8eb62…` | 174 | 86 / 50 / 38 | 12 | as C, claude column renamed `claude-3-5-sonnet-20241022` | none (Prompt Guard and LangKit now carry real 0/1 values) | all verdict columns | SELECTION RULE: UNKNOWN | selected subset | none |
| E | `ffe88ccb` (2025-02-23, "new playground data") | same path, sha `d93f9fe1d1a1…` | 174 | 86 / 50 / 38 | 12 | as D | `llm_guard=0` | rest | SELECTION RULE: UNKNOWN | selected subset | none |
| F | `b20aeed5` (2025-03-22, "new results") = **`507566c5` (2025-07-08, default-branch head)** | same path, sha `791dd4b0a168f2eb…03f57c3` | **170** | **82 / 50 / 38** | 11 (Miscellaneous absent) | five specialized + gpt-4, claude-3-5-sonnet-20241022, gemini-1.5-pro-latest, mistral-large-latest, deepseek-ai/DeepSeek-V3, grok-2-latest | none over 170 rows; within the 82 harmful rows `llm_guard=0` | all eleven verdict columns | **SELECTION RULE: UNKNOWN**. Only 74–76 of the 170 questions appear in C–E, so F is a fresh draw from pool B, not a trimming of C. All 170 questions are in A and B with identical `harm_level`. | selected subset (author-selected; rule unstated) | MC-002 (every `expected` count), MC-003 (`{0…12}/82`, leave-one-out), MC-005 at HEAD (n = 82), census row `bells-misuse-2025` (`item_level_outcomes_released`), dossier `bells-misuse-2025`, disclosure page, two `/answers/` pages, `/try/` TRY-B, film theses in `films/SLATE.md` |
| F | `b20aeed5` (2025-03-22, "new results") = **`507566c5` (2025-07-08, default-branch head)** | same path, sha `791dd4b0a168f2eb…03f57c3` | **170** | **82 / 50 / 38** | 11 (Miscellaneous absent) | five specialized + gpt-4, claude-3-5-sonnet-20241022, gemini-1.5-pro-latest, mistral-large-latest, deepseek-ai/DeepSeek-V3, grok-2-latest | none over 170 rows; within the 82 harmful rows `llm_guard=0` | all eleven verdict columns | **SELECTION RULE: UNKNOWN**. Only 74–76 of the 170 questions appear in C–E, so F is a fresh draw from pool B, not a trimming of C. All 170 questions are in A and B with identical `harm_level`. | selected subset (author-selected; rule unstated) | MC-002 (every `expected` count), MC-003 (`{0…12}/82`, leave-one-out), MC-005 at HEAD (n = 82; MC-005 retracted 2026-09-06), census row `bells-misuse-2025` (`item_level_outcomes_released`), dossier `bells-misuse-2025`, disclosure page, two `/answers/` pages, `/try/` TRY-B, film theses in `films/SLATE.md` |
| G | `507566c5` | `data/adversarial_prompts.csv` sha `32fe8663621a…` | 8 | 8 / 0 / 0 | 1 (all Miscellaneous) | eleven verdict columns | `harm_level`, `category`, `deepseek-ai/DeepSeek-V3=0` | rest | SELECTION RULE: UNKNOWN | selected subset | 12-WEEK-PROGRAM.md ("8 adversarial rows released"); census row ("plus 8 adversarial prompts") |
| H | `8a974123`→`507566c5` | `data/metacognitive_results.csv` sha `66a9f09c60d3…` | 612 rows = 102 prompts × 6 models | ground_truth 504 harmful / 108 not_harmful rows; 19 non-adversarial prompts | — | none of the five supervisors; frontier models only, with **newer versions** (`claude-3-7-sonnet-20250219`, `gemini-2.5-pro-exp-03-25`) | `model` set | responses | UNKNOWN | selected subset | none |

Expand Down Expand Up @@ -137,7 +148,10 @@ F into a population estimate. MC-002 stays non-population-level.
files of February 2025.
- Upstream repository declares no licence (GitHub API `license: null`; no
`LICENSE` in the tree at `507566c5`). MC-002 records "none declared
upstream"; MC-005 at HEAD `f9b24e2` records `license: MIT` for the same file.
upstream"; MC-005 (since retracted) at HEAD `f9b24e2` recorded `license: MIT`
for the same file.
MC-005 was retracted 2026-09-06 for exactly this; `none declared upstream`
is now the register's only statement about that file.
- The v0 commit `0fc3d6d3` carries 114,540 adversarial rows with prompt text,
while the paper (App. 0.C) states the full dataset is not publicly released.
Recorded as a provenance fact; nothing here redistributes it.
Expand All @@ -159,7 +173,8 @@ F into a population estimate. MC-002 stays non-population-level.
through `ROOT.rglob("index.html")`. Zero findings on tracked files. A clean
clone has no `_private/`, so CI is unaffected; the two files this packet
adds are not `index.html` and are not touched by that gate.
- Working tree (untouched): staged removal of MC-005 from `claims.yaml`
- Working tree (untouched): staged removal of MC-005 (now retracted) from `claims.yaml`
(completed, and declared as a RETRACT in the history, on 2026-09-06)
relative to HEAD; unstaged addition of `exclusive_cells` to MC-002 and a
matching `claims_history.yaml` transition; modified generators, verifiers,
film receipts; untracked `experiments/e3/`.
Expand Down Expand Up @@ -240,8 +255,8 @@ not create a finding.
- No E2 item, threshold, estimator or criterion was touched. E2 remains
`UNTESTED BY DESIGN`; P2 remains `EMPTY`.
- No live API was called; no money was spent; nobody was contacted.
- MC-002, MC-003, MC-005 and the census row were not edited. Two candidate
owner actions are noted, not performed: (a) MC-005's `license: MIT` is
- MC-002, MC-003, MC-005 (since retracted) and the census row were not edited. Two candidate
owner actions are noted, not performed: (a) retracted MC-005's `license: MIT` is
unsupported by upstream (no licence declared) and conflicts with MC-002's
record for the same file; (b) the census row's evidence text "170 prompts
… the full dataset is available only by contacting the authors" is
Expand Down
2 changes: 1 addition & 1 deletion ARTIFACTS/2026-09-07-RECONCILIATION.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,7 @@ held twice: `experiments/e3/freeze/or-bench-80k.csv` and
`experiments/e3b/freeze/or-bench-80k.csv`, byte-identical
(`22e95602…`). Each freeze is self-contained by design, so the duplicate
stays and is recorded by detector D7 in `docs/graph/repo-graph.json`. The
remaining ~11.7k lines are the E3 and E3B experiments, the MC-005
remaining ~11.7k lines are the E3 and E3B experiments, the retracted MC-005
retraction, and the film re-render. Nothing in the 17 commits was found
uncommittable; see §4 for what an adversarial pass did find.

Expand Down
22 changes: 19 additions & 3 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,13 +11,19 @@ style — they are the product, and CI enforces them.
```bash
python3 .claude/skills/evidence-ledger/ledger.py # the instant
python3 scripts/cadence.py report # the trajectory
python3 scripts/direction.py # the heading
```

Repo state in one screen: what has been measured, what is blocked on a human,
and how much protocol sits on top of how many rows. Cheaper and more honest
than reconstructing it from the program, the contract, the freeze and the last
report. The cadence report answers the question one instant cannot: whether the
answer is moving, and whether what moved was evidence or scaffolding. See
answer is moving, and whether what moved was evidence or scaffolding. The
direction report answers what neither can: which way this is pointed, whether
the next thing you are about to write is the thing it needs, and whether
infrastructure outstanding has overtaken work outstanding. Its source is
`research/DIRECTION.yaml` — the whole plan, machine-read, with recorded
amendments to its own first draft. Read it before proposing a new direction. See
`.claude/skills/evidence-ledger/SKILL.md` for the two standing rules — in
particular: before writing another governing document, say the trade out loud
first.
Expand All @@ -35,9 +41,10 @@ prints it from the tuple.
## Conventions CI enforces

- **Generated pages are never hand-edited** — `/ledger/`, `/observatory/`,
`/modules/*`, `/missing-column/*`, `/try/`, `/worldspace/`, `sitemap.xml`.
`/modules/*`, `/missing-column/*`, `/try/`, `/worldspace/`,
`/records/conformance/`, `sitemap.xml`.
Edit the source registry (`claims.yaml`, `modules.yaml`, `census.yaml`) and
regenerate; nine `scripts/generate_*.py` have `--check` drift gates.
regenerate; twelve `scripts/generate_*.py` have `--check` drift gates.
- **`claims.yaml` is schema v0.4.** Every claim needs a non-empty
`falsifier.condition`, a fixed `NARROW|REJECT|HOLD` consequence, and a typed
`forbidden_rescues` list (explicit `[]` is valid). A declared commit must be
Expand All @@ -53,6 +60,15 @@ prints it from the tuple.
anything else with a `local_content_change` trigger.
- **Figure numbers** are re-derived from stated constants in
`scripts/verify_figures.py` to 1e-9.
- **New experiments use a canonical executable contract.**
Follow correction C1 in `research/DIRECTION.yaml`: the runner consumes
`contract.json`, and `PREREG.md` renders its scientific choices into prose.
`research/contracts/validate_contract.py` checks structure, threshold direction,
planned pool identities and planned marginal preservation. Passing establishes
internal consistency, not runtime conformance or inferential validity.
`verify_prereg.py` remains the legacy sidecar scan and historical fixture
corpus; its static reference check does not prove runtime consumption.
Frozen experiments remain untouched and undeclared rather than retrofitted.
- **The cadence series only appends.** `metrics/repo_state.jsonl` is a hash
chain of daily repository state; an edited, reordered, dropped or truncated
row fails `scripts/cadence.py check`. CI enforces the chain and never the
Expand Down
22 changes: 22 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -253,3 +253,25 @@ DESIGN.md design-decision ledger + changelogs + field-artifact
```

Content © Pranav Bhave. Code (HTML/CSS/JS) may be reused with attribution.

## Scheduled feedback and release updates

The existing local schedules record repository state daily and collect distribution
observations every four hours. Missing days and missed observation windows remain
missing; the jobs never interpolate evidence. Both jobs share a repository lock.
They prefer `.venv/bin/python3`; create that environment with Python 3.12 or newer
and install `requirements.txt` before enabling the schedules. `CUBITS11_PYTHON`
can select another environment explicitly.

A cycle starting from clean, synchronized `main` runs the verification manifest,
commits only its permitted outputs on a `claude/cycle-*` branch, and opens a pull
request. GitHub auto-merge uses a merge commit and remains subject to required
checks and review rules. A failed API call leaves the branch available for recovery;
a failed check retains the observations without publishing them. Logs live in
`_private/cron/`. No cycle dispatches a new outreach message.

This feedback loop records observations and tests known failures. It does not
autonomously change scientific criteria, promote a hypothesis, or accept a
contradicted result. Deployment checks compare the public pages, primary films,
posters and claim metadata with the verified revision, so an older page returning
HTTP 200 cannot stand in for the release.
Loading