diff --git a/distribution/QUEUE.md b/distribution/QUEUE.md index d2ba5f6..ebc14f7 100644 --- a/distribution/QUEUE.md +++ b/distribution/QUEUE.md @@ -11,10 +11,10 @@ message was sent during reconciliation. | 3 | GuardBench joint-reporter ask | **WAIT until September 9**, one week after IBM. Recheck the target and duplicate issues before sending the existing [dossier](dossiers/guardbench-2024.md) with [reporter](../contrib/guardbench_joint.py). No matching Cubits11 request found among the public issues inspected September 5. | A reply, accepted patch, or usable joint-results release. | | 4 | Dispatch `x-film-same-scores` | **HOLD for the cold-viewer gate below and account baseline.** Current `campaigns.yaml` requires **5–8** cold viewers, not three. Then re-derive numerals at the dispatch commit; post the existing 11-second square cut and registered copy once. Record its actual permalink/time and platform denominators; never invent a dispatch date. | Verified independent reproduction. Link clicks / impressions are a verification-intent proxy, not observed `/try/` landings. | | 5 | Run the frozen Same Scores comprehension trial | **NEXT HUMAN ACTION.** Five minimum, eight maximum, people unfamiliar with the project. One viewing, sound off, no explanation; collect all three answers verbatim before scoring. Rules remain in `launch-units.yaml → cold_test_gate`. Two failures on the same question hold release; revise once and use a fresh cold set, never rescore. | Real answers and pass/fail diagnostics. No trial has been fabricated or recorded as completed. | -| 6 | BELLS source-selection question | **PREPARED, separate weekly card.** Inspect the actual author announcement and existing replies first. Respect the current PRESENT-first selection rule; do not strengthen the frozen row or publish private denominator findings through an outreach shortcut. | A checkable source confirmation or correction. | +| 6 | BELLS source-selection question | **SENT September 16:** [BELLS issue #1](https://github.com/CentreSecuriteIA/bells_leaderboard/issues/1). Exact posted text recorded in the dossier and dispatch log; awaiting an answer about selection. Do not duplicate. | A checkable source confirmation or correction. | | 7 | Subsequent film dispatches | **WAIT for Same Scores' actual T+14d decision.** Then change one representation at a time, at least 14 days apart, under the existing rule. No fixed dates until the first dispatch exists. | Reproduction or joint release; attention stays diagnostic. | | 8 | Process an external outcome | Follow [EXTERNAL_EVENTS.md](EXTERNAL_EVENTS.md) when a real artifact arrives. The three existing tracked issues had no comments at the September 5 check. An issue opened by the author is not a qualified outcome. | Verified entry, relevant regeneration, credit, and any required correction. | -| 9 | Apply the stop rule | **NOT TRIGGERED: 3/12 technical interactions, 0 qualified outcomes.** The third is IBM's already-sent request, recorded retrospectively. Nine interactions remain. Do not infer unobserved funnel stages. | A supported diagnosis, with exactly one changed variable when the rule fires. | +| 9 | Apply the stop rule | **NOT TRIGGERED: 4/12 technical interactions, 0 qualified outcomes.** The third is IBM's already-sent request; the fourth is the BELLS selection question. Eight interactions remain. Do not infer unobserved funnel stages. | A supported diagnosis, with exactly one changed variable when the rule fires. | | 10 | Cohort B films | **HOLD.** Only the existing A–D evidence gates in the E2 brief justify further films. | The evidence that opens a gate, not another film. | ## Dates that can actually be scheduled diff --git a/distribution/dispatch-log.yaml b/distribution/dispatch-log.yaml index 5e1982b..5363ede 100644 --- a/distribution/dispatch-log.yaml +++ b/distribution/dispatch-log.yaml @@ -43,3 +43,20 @@ dispatches: - {at: "T+7d", rule: "same snapshot; no second film until the T+14d decision"} - {at: "T+14d", rule: "apply the campaign's decision rule; silence is NO_OBSERVED_RESPONSE"} outcome: null + + - id: bells-selection-rule + sent_at_commit: null # owner browser submission; checkout at send time not established + reproduction_commit: 4b659fb97bf50aff9793ff4a26c3410b7cf0a53c + adjudication: + - {on: substantive_response, rule: "Verify selection rule or correction against the cited artifact before awarding outcome credit."} + - {on: silence, rule: "Retain NO_OBSERVED_RESPONSE; do not duplicate the issue."} + dossier: distribution/dossiers/bells-misuse-2025.md + target: "CentreSecuriteIA/bells_leaderboard — issue tracker" + status: SENT_PENDING_RESPONSE + sent_by: owner + reported: "2026-09-16" + sent_at_utc: "2026-09-17T03:19:51Z" + permalink: https://github.com/CentreSecuriteIA/bells_leaderboard/issues/1 + observed_comments: 0 + message_as_sent: "Thank you for releasing per-item supervisor verdicts in `data/non_adversarial_prompts.csv`.\n\nAt commit `507566c5a4606c8e3dec0bd59a5c5fde62594951` the file has 170 rows, 82 of them with `harm_level` = `harmful`. On 9 of those 82 rows, all five specialized-supervisor columns (`lakera_guard`, `prompt_guard`, `langkit`, `nemo`, `llm_guard`) are 0. This script downloads the file at that commit, checks its SHA-256, and recomputes the count: https://github.com/Cubits11/cubits11.github.io/blob/4b659fb97bf50aff9793ff4a26c3410b7cf0a53c/scripts/reanalyze_bells_subset.py\n\nYour FAQ and the paper's Dataset Access appendix call the playground rows \"representative examples\". How were these 170 prompts selected from the full non-adversarial set? A sampling rule, a filter, manual choice, or a link to an existing description would answer it.\n\nI report 9/82 only as a count on these 170 released rows, not as an estimate for the full benchmark or for any deployed system. If I have misread the release, a correction is just as useful." + outcome: null diff --git a/distribution/dossiers/bells-misuse-2025.md b/distribution/dossiers/bells-misuse-2025.md index 087540b..ebf088b 100644 --- a/distribution/dossiers/bells-misuse-2025.md +++ b/distribution/dossiers/bells-misuse-2025.md @@ -1,6 +1,8 @@ # Dossier — bells-misuse-2025 -Row: `bells-misuse-2025` · classification PRESENT (computable_via_item_release) · reconstruction PARTIALLY_IDENTIFIED · status: PREPARED, nothing sent. +Row: `bells-misuse-2025` · classification PRESENT (computable_via_item_release) · reconstruction PARTIALLY_IDENTIFIED · status: SENT September 16, 2026 (America/New_York), RESPONSE PENDING. + +Issue: https://github.com/CentreSecuriteIA/bells_leaderboard/issues/1 · sent 2026-09-17T03:19:51Z · zero comments when verified. **1. What the source publishes.** Per-supervisor results for twelve systems, and a released per-item subset: 170 non-adversarial prompts (82 harmful, 50 benign, 38 borderline) with binary verdict columns for five specialized supervisors, at a pinned commit (`507566c5…`, hash-verified in CI). Nothing joint is printed; the headline population's per-item outcomes and the ~4,165 adversarial prompts have no per-item release. The subset's selection rule is unstated. @@ -18,8 +20,14 @@ Row: `bells-misuse-2025` · classification PRESENT (computable_via_item_release) **8. Correction condition.** Evidence that a released column is not the labelled system's verdict, that the subset was released with a stated rule the row misreports, or that a fuller per-item release existed on or before 2026-08-27. Any of these corrects MC-002 or the row under their registered falsifiers. -**The ask, verbatim:** +**The ask, verbatim (as posted):** -> Your released 170-prompt subset lets anyone recompute what the five specialized supervisors miss *together*: 9 of 82 harmful prompts, against an independence plug-in of 2.87 — and the leave-one-out unions show three of the five add no exclusive coverage on that stratum. One question so the row about your work is right: how were the 170 prompts selected? If the row is wrong, say so — corrections land dated in the public file. +> Thank you for releasing per-item supervisor verdicts in `data/non_adversarial_prompts.csv`. +> +> At commit `507566c5a4606c8e3dec0bd59a5c5fde62594951` the file has 170 rows, 82 of them with `harm_level` = `harmful`. On 9 of those 82 rows, all five specialized-supervisor columns (`lakera_guard`, `prompt_guard`, `langkit`, `nemo`, `llm_guard`) are 0. This script downloads the file at that commit, checks its SHA-256, and recomputes the count: https://github.com/Cubits11/cubits11.github.io/blob/4b659fb97bf50aff9793ff4a26c3410b7cf0a53c/scripts/reanalyze_bells_subset.py +> +> Your FAQ and the paper's Dataset Access appendix call the playground rows "representative examples". How were these 170 prompts selected from the full non-adversarial set? A sampling rule, a filter, manual choice, or a link to an existing description would answer it. +> +> I report 9/82 only as a count on these 170 released rows, not as an estimate for the full benchmark or for any deployed system. If I have misread the release, a correction is just as useful. **Channel.** GitHub issues on `CentreSecuriteIA/bells_leaderboard` (the row's recorded route). One message. diff --git a/distribution/outcomes.yaml b/distribution/outcomes.yaml index 2b1a7e7..209825f 100644 --- a/distribution/outcomes.yaml +++ b/distribution/outcomes.yaml @@ -46,8 +46,16 @@ diagnostics: cold_comprehension_trials: [] # An invitation or open issue is diagnostic only; it is never credited as # an external reproduction, correction, release, PR, or cold run. - technical_interactions: 3 # replies that carried a number, source, or correction + technical_interactions: 4 # replies that carried a number, source, or correction technical_interaction_log: + - date: "2026-09-16" + kind: upstream_issue + url: https://github.com/CentreSecuriteIA/bells_leaderboard/issues/1 + summary: > + Owner asked how the 170 released BELLS rows were selected, acknowledging + the published representative-examples description. Posted September 16 + in America/New_York (September 17 UTC). Open with no comments when verified; + diagnostic only, not an independent reproduction or source correction. - date: "2026-09-02" kind: upstream_issue url: https://github.com/IBM/Adversarial-Prompt-Evaluation/issues/7 diff --git a/distribution/traction/events.json b/distribution/traction/events.json index 6b2590c..ee7e55e 100644 --- a/distribution/traction/events.json +++ b/distribution/traction/events.json @@ -52,13 +52,13 @@ "novelty": "snapshot; not asserted newly published" }, { - "id": "70050c08430f7c32", + "id": "c93d03ba49aff94d", "kind": "metrics", "source": { "path": "distribution/outcomes.yaml", - "sha256": "2cefa442652376b870e95f86840e5716cc91da5ea91f9be46585d2a19ff23b66", - "commit": "26a9d578ffe77deecc191f02ae578e6f8ae6e2c0", - "url": "https://github.com/Cubits11/cubits11.github.io/blob/26a9d578ffe77deecc191f02ae578e6f8ae6e2c0/distribution/outcomes.yaml", + "sha256": "91f5decd619427fdb700ce5a458f73453402fc06c9b25a5c1ed039a74c64ef38", + "commit": "baf018bfd3997716cc8800f51861bb81edd98652", + "url": "https://github.com/Cubits11/cubits11.github.io/blob/baf018bfd3997716cc8800f51861bb81edd98652/distribution/outcomes.yaml", "binding": "committed" }, "draft_ids": [], diff --git a/distribution/traction/manifest.json b/distribution/traction/manifest.json index cabbe64..fe7d442 100644 --- a/distribution/traction/manifest.json +++ b/distribution/traction/manifest.json @@ -2,7 +2,7 @@ "version": 1, "generator": "scripts/distribute.py", "outputs": { - "events.json": "5f7baafe4f8a98f8823333777efbf0899ef1375a6e18c23e66abd78a062fe0a1", + "events.json": "343833167a7794e4269e40c21868b52fcd3055d9ce358c593c1646cf3538e8e6", "drafts.json": "6916cc6433575eace3023a124bcbf93ecffe3ac5ce6637e6b4661c6ff1128419", "queue.json": "ca0558ccf563ec99695cfc68a39a2a5c7e19a9c33dbaf232f09689c589bf7789", "dashboard.json": "988051f7c9d5b82684071ab29fdb52a891ce0f93470405f402cc6beff62d04bb", diff --git a/docs/RESEARCH_INDEX.md b/docs/RESEARCH_INDEX.md index 2c74ff4..89da169 100644 --- a/docs/RESEARCH_INDEX.md +++ b/docs/RESEARCH_INDEX.md @@ -13,7 +13,7 @@ Registry v0.4 · last owner review 2026-09-14 · one question: *Do published per | paired outcome releases | 0 | | upstream prs | 0 | | human cold runs | 0 | -| technical interactions (diagnostic, not an outcome) | 3 | +| technical interactions (diagnostic, not an outcome) | 4 | Zero is the recorded value where it is zero. The stop rule and the procedure for recording an outcome are in `distribution/EXTERNAL_EVENTS.md`. diff --git a/docs/graph/repo-graph.json b/docs/graph/repo-graph.json index bb9447d..9cb0b0a 100644 --- a/docs/graph/repo-graph.json +++ b/docs/graph/repo-graph.json @@ -4902,66 +4902,17 @@ "severity": "RECORD", "sha256": "22e956020cdfdd43c328f02441f003c933ee4e7083d02d9d24e414e4b42ae2e6" }, - { - "ahead": 15, - "behind": 0, - "branch": "claude/inferential-adequacy", - "detector": "D4", - "severity": "TOPOLOGY" - }, { "ahead": 1, "behind": 0, - "branch": "claude/issue-forms-ingress", - "detector": "D4", - "severity": "TOPOLOGY" - }, - { - "ahead": 17, - "behind": 0, - "branch": "claude/joint-cell-visible", - "detector": "D4", - "severity": "TOPOLOGY" - }, - { - "ahead": 14, - "behind": 0, - "branch": "claude/seo-falsifiable-web", - "detector": "D4", - "severity": "TOPOLOGY" - }, - { - "ahead": 17, - "behind": 0, - "branch": "claude/stack-merge", - "detector": "D4", - "severity": "TOPOLOGY" - }, - { - "ahead": 15, - "behind": 0, - "branch": "origin/claude/inferential-adequacy", + "branch": "claude/bells-sent-recorded", "detector": "D4", "severity": "TOPOLOGY" }, { "ahead": 1, "behind": 0, - "branch": "origin/claude/issue-forms-ingress", - "detector": "D4", - "severity": "TOPOLOGY" - }, - { - "ahead": 17, - "behind": 0, - "branch": "origin/claude/joint-cell-visible", - "detector": "D4", - "severity": "TOPOLOGY" - }, - { - "ahead": 14, - "behind": 0, - "branch": "origin/claude/seo-falsifiable-web", + "branch": "origin/claude/bells-sent-recorded", "detector": "D4", "severity": "TOPOLOGY" }, @@ -4973,8 +4924,8 @@ } ], "generated_from": { - "branch": "claude/stack-merge", - "head": "4173a8c" + "branch": "claude/bells-sent-recorded", + "head": "baf018b" }, "nodes": [ { @@ -5825,110 +5776,137 @@ "type": "experiment" }, { - "ahead_of_origin_main": 15, + "ahead_of_origin_main": 1, "behind_origin_main": 0, + "head": "baf018b", + "id": "claude/bells-sent-recorded", + "last_commit": "2026-09-17", + "reachable_from_origin_main": false, + "type": "branch" + }, + { + "ahead_of_origin_main": 0, + "behind_origin_main": 5, "head": "6a51abb", "id": "claude/inferential-adequacy", "last_commit": "2026-09-17", - "reachable_from_origin_main": false, + "reachable_from_origin_main": true, "type": "branch" }, { - "ahead_of_origin_main": 1, - "behind_origin_main": 0, + "ahead_of_origin_main": 0, + "behind_origin_main": 19, "head": "026554c", "id": "claude/issue-forms-ingress", "last_commit": "2026-09-17", - "reachable_from_origin_main": false, + "reachable_from_origin_main": true, "type": "branch" }, { - "ahead_of_origin_main": 17, - "behind_origin_main": 0, + "ahead_of_origin_main": 0, + "behind_origin_main": 3, "head": "4173a8c", "id": "claude/joint-cell-visible", "last_commit": "2026-09-17", - "reachable_from_origin_main": false, + "reachable_from_origin_main": true, "type": "branch" }, { - "ahead_of_origin_main": 14, - "behind_origin_main": 0, + "ahead_of_origin_main": 0, + "behind_origin_main": 6, "head": "5fe44bc", "id": "claude/seo-falsifiable-web", "last_commit": "2026-09-15", - "reachable_from_origin_main": false, + "reachable_from_origin_main": true, "type": "branch" }, { - "ahead_of_origin_main": 17, - "behind_origin_main": 0, - "head": "4173a8c", + "ahead_of_origin_main": 0, + "behind_origin_main": 1, + "head": "cc7c7ec", "id": "claude/stack-merge", "last_commit": "2026-09-17", - "reachable_from_origin_main": false, + "reachable_from_origin_main": true, "type": "branch" }, { "ahead_of_origin_main": 0, "behind_origin_main": 0, - "head": "4b659fb", + "head": "d3766ae", "id": "main", - "last_commit": "2026-09-07", + "last_commit": "2026-09-17", "reachable_from_origin_main": true, "type": "branch" }, { "ahead_of_origin_main": 0, "behind_origin_main": 0, - "head": "4b659fb", + "head": "d3766ae", "id": "origin", - "last_commit": "2026-09-07", + "last_commit": "2026-09-17", "reachable_from_origin_main": true, "type": "branch" }, { - "ahead_of_origin_main": 15, + "ahead_of_origin_main": 1, "behind_origin_main": 0, + "head": "baf018b", + "id": "origin/claude/bells-sent-recorded", + "last_commit": "2026-09-17", + "reachable_from_origin_main": false, + "type": "branch" + }, + { + "ahead_of_origin_main": 0, + "behind_origin_main": 5, "head": "6a51abb", "id": "origin/claude/inferential-adequacy", "last_commit": "2026-09-17", - "reachable_from_origin_main": false, + "reachable_from_origin_main": true, "type": "branch" }, { - "ahead_of_origin_main": 1, - "behind_origin_main": 0, + "ahead_of_origin_main": 0, + "behind_origin_main": 19, "head": "026554c", "id": "origin/claude/issue-forms-ingress", "last_commit": "2026-09-17", - "reachable_from_origin_main": false, + "reachable_from_origin_main": true, "type": "branch" }, { - "ahead_of_origin_main": 17, - "behind_origin_main": 0, + "ahead_of_origin_main": 0, + "behind_origin_main": 3, "head": "4173a8c", "id": "origin/claude/joint-cell-visible", "last_commit": "2026-09-17", - "reachable_from_origin_main": false, + "reachable_from_origin_main": true, "type": "branch" }, { - "ahead_of_origin_main": 14, - "behind_origin_main": 0, + "ahead_of_origin_main": 0, + "behind_origin_main": 6, "head": "5fe44bc", "id": "origin/claude/seo-falsifiable-web", "last_commit": "2026-09-15", - "reachable_from_origin_main": false, + "reachable_from_origin_main": true, + "type": "branch" + }, + { + "ahead_of_origin_main": 0, + "behind_origin_main": 1, + "head": "cc7c7ec", + "id": "origin/claude/stack-merge", + "last_commit": "2026-09-17", + "reachable_from_origin_main": true, "type": "branch" }, { "ahead_of_origin_main": 0, "behind_origin_main": 0, - "head": "4b659fb", + "head": "d3766ae", "id": "origin/main", - "last_commit": "2026-09-07", + "last_commit": "2026-09-17", "reachable_from_origin_main": true, "type": "branch" }, diff --git a/films/flagship/source-map.json b/films/flagship/source-map.json index bb663e1..48944fe 100644 --- a/films/flagship/source-map.json +++ b/films/flagship/source-map.json @@ -17,7 +17,7 @@ "films/the-falsifier-that-fired/manifest.yaml": "89814cb92b88aa87ab546798f737a9d8ff3a4a80548d8375cd67d2b3d3148acd", "scripts/reanalyze_bells_subset.py": "c39643ca63146931aaaa8614477453adcd5d411501d251d33d12bdb811759a88", "films/thirteen-worlds/manifest.yaml": "351adf9ccbd42f54c53645251378aaeb0255e5507874cfe78d59009476d79828", - "missing-column/reproduce/index.html": "80280667a8e99287856837e17586b12d63106c2790107395761ffb17863ff29f", + "missing-column/reproduce/index.html": "1dcdcb4f7c87b1e9968a2435036a3c49a253baef911eecfe050832279e49ca92", "films/the-correction-invitation/manifest.yaml": "9b87c56e01095ea1db9152d67c12f193afc670a0ed44a54bc664652b5680602d" }, "sections": [ diff --git a/metrics/ledger_snapshot.json b/metrics/ledger_snapshot.json index 0c767d8..783c371 100644 --- a/metrics/ledger_snapshot.json +++ b/metrics/ledger_snapshot.json @@ -85,7 +85,7 @@ "outcomes": { "qualified_total": 0, "categories": 5, - "technical_interactions": 3, + "technical_interactions": 4, "stop_threshold": 12 } } diff --git a/missing-column/reproduce/index.html b/missing-column/reproduce/index.html index 45e6e52..154ae2e 100644 --- a/missing-column/reproduce/index.html +++ b/missing-column/reproduce/index.html @@ -217,7 +217,7 @@
| file | sha256 |
|---|---|
research/DIRECTION.yaml | 92ed64289f8dae5bbc577053e7f11d9e4551a0cc415b0c28bb62150291d691d1 |
metrics/ledger_snapshot.json | e18dfb22fd0b4016d0e1e90e20481e2b932991efbf201811d913d5c79af13c15 |
corrections/records/2026-09-10.json | a2cca203fd92ffe17ec09d841e90154cfd9cf966f4aeb1b1a62ad0c223cb2cc0 |
corrections/records/2026-09-10-preserved-inputs.json | b20ad976edcfcccf2a8fa7fa8db609a3b735bb5a2ab1a75c459395fe437a100d |
scripts/verify_prereg.py | e58aba8bb699a1691bb9accd266e356584d5d2d488053f074dc64b944d7870b1 |
claims.yaml | 9964c0137292e4338246d423c7124a0109b28973f4ca00f94317e855cd4dcc94 |
research/DIRECTION.yaml | 92ed64289f8dae5bbc577053e7f11d9e4551a0cc415b0c28bb62150291d691d1 |
metrics/ledger_snapshot.json | 852b271b4404dd4309d004b969c968427675e06e4acc7fd41fdbf463eb436b52 |
corrections/records/2026-09-10.json | a2cca203fd92ffe17ec09d841e90154cfd9cf966f4aeb1b1a62ad0c223cb2cc0 |
corrections/records/2026-09-10-preserved-inputs.json | b20ad976edcfcccf2a8fa7fa8db609a3b735bb5a2ab1a75c459395fe437a100d |
scripts/verify_prereg.py | e58aba8bb699a1691bb9accd266e356584d5d2d488053f074dc64b944d7870b1 |
claims.yaml | 9964c0137292e4338246d423c7124a0109b28973f4ca00f94317e855cd4dcc94 |
QUALIFIED OUTCOMES — work done by someone who is not the author · bound to distribution/outcomes.yaml
| independent reproductions | 0 |
| source corrections | 0 |
| paired outcome releases | 0 |
| upstream prs | 0 |
| human cold runs | 0 |
Diagnostics, kept apart and never counted as outcomes: technical interactions 3; blinded comprehension trials 0. Zero is the recorded value. Be the first independent rerun → experiment A.
+Diagnostics, kept apart and never counted as outcomes: technical interactions 4; blinded comprehension trials 0. Zero is the recorded value. Be the first independent rerun → experiment A.