the program map: 13 distinct questions across 47 arcs, and the audit that says so - #51
Merged
Conversation
Two new files. Nothing else is touched: no receipt, no certificate, no published paper, nothing under .github/. papers/INDEX_program_map_2026_09_01.md — one row per arc for all 47, each headline recorded as it terminally stands rather than as it was first hoped; an IDEA INDEX in date order with CITED/RE-DERIVED marked per row; a RECEIPTS INDEX from headline number to the file holding it. papers/AUDIT_the_whole_program_2026_09_01.md — an AUDIT, not a RESULT. No preregistration covers it and it carries no headline finding. What it found, in three lines: - 47 arcs resolve to 13 distinct research questions and 6 recurring mechanisms, with the sensitivity of that number printed next to it. The arc-to-arc citation graph has density 0.0407; 14 arcs name no other arc, 10 are named by none. - mention-versus-use was operationally present on 2026-05-25 under the name "FORM impersonating MEANING" and the 2026-08-26 synthesis that named the class cites none of the five arcs that had it. The handed-target mechanism has been found eleven times and named zero. - No deposited DOI is known to assert something since refuted. One open item: CITATION.cff still names the record whose repo copy carries a scope erratum, and there is no in-repo evidence of a re-deposit. The prose-reading hypothesis is tested rather than assumed and comes back SUPPORTED-WITH-QUALIFICATION, with counterexamples in both directions and the objection about verdict-free instruments argued rather than waved away. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Both 2026-09-01 research briefs open with 'read papers/INDEX.md first'. The file did not exist under that name and nothing enforced it, so the map would have rotted exactly the way the audit says the arcs did. Renamed INDEX_program_map -> papers/INDEX.md and added papers/build_index.py plus tests/test_index.py. WHAT IS ENFORCED, and the line matters: every arc directory has exactly one row; every row names a directory that exists; every tag comes from the closed vocabulary; every status comes from the closed vocabulary; every module named in ships-in exists. You cannot add an arc to this repository without adding a row. WHAT IS NOT: the one-sentence terminal claim, the idea index and the receipts index are AUTHORED. A script that regenerated them would be inventing the thing it claims to check -- the handed-target defect this lab spent the week measuring. The receipt says so in a not_checked field. The checker found three real defects in the index on its first run: papers/arxiv/ has no row (it is submission artifacts, not an arc -- now excluded), showcase-viz was wrongly excluded (33 papers, it is an arc), and resonance_profiler.py resolves inside its own arc directory rather than at the repo root, which is itself the audit's finding that it was never promoted into styxx/. Mutation-tested: an arc directory with no row turns the suite red, and the phantom arc was removed and its absence asserted. Full suite 2893. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Two new files under
papers/. Nothing else is touched — no receipt, no certificate, no seal, no published paper, nothing under.github/.git show --statis two additions and zero modifications.Neither file is a RESULT. The AUDIT says so in its first paragraph: no preregistration covers it, no bar was frozen, and it therefore carries no headline finding. Every number in it is either a shell command printed beside it or a quotation from a document named in place.
Lead: how many distinct ideas this program actually contains
47 arcs resolve to 13 distinct research questions and 6 recurring mechanisms.
The clustering rule is printed before the clusters so it can be argued with: two arcs ask the same question if a decisive answer to one would change what the other is allowed to claim. The sensitivity is printed next to the number — splitting by sub-question gives about 22, collapsing to the README's own three layers gives 3 — because quoting 13 without that paragraph would be quoting a correlate.
The mechanical part is not a judgement call. Building the arc→arc citation graph by searching every arc's markdown for every other arc's directory name gives 88 directed links out of 2,162 possible, density 0.0407, mean out-degree 1.87. Fourteen arcs name no other arc. Ten arcs are named by no other arc. That is the operator's feeling in a form anyone can re-run in twenty lines of Python.
Two structural facts underneath it: 29 of the 45 dated arcs span two days or fewer, and 22 of the 47 begin on one of three days (2026-05-25, 2026-06-02, 2026-06-03). Arc count is a directory count.
Lead: published risk
No deposited DOI is known to assert something this lab has since refuted. Stating that plainly, because it is a good result and it was not guaranteed.
The record does the hard thing correctly in three places: the frame-locality circularity was deposited as its own erratum DOI (
10.5281/zenodo.21679805) and then as a corrected edition (21693636); the read≠write line was re-versioned v26 → v27 → v28 withERRATUM_v26_adaptive_claim.mdtelling readers to cite v26 with the erratum; and two claims that were killed before deposit (calib-poison's coupling constant, anchored-validity's reserved21520429) are staged and gated rather than published.One open item.
CITATION.cffnames10.5281/zenodo.19777921as the preferred citation for this repository. That essay's repo copy carries a scope erratum dated 2026-06-21 bounding or falsifying two of its central claims. Whether the deposited record carries that erratum is UNCHECKABLE from this repository and is recorded as UNCHECKABLE. What is checkable:scripts/zenodo_version_emlv.pywas last modified 2026-04-26, eight weeks before the erratum, and no deposit artifact dated after it names that DOI. The asymmetry is the finding — this lab has in-repo evidence of re-versioning a corrected paper twice, and none of doing it here.Also standing, and correctly described as disclosed rather than as risk: the tool-call-drift AUC inconsistency across permanent records (0.916 vs 0.943), named as uneditable in the 2026-05-17 review. That review is the only whole-corpus permanent-record review in the repository and it predates the construct ceiling, the erratum, the mention-and-use synthesis, EXTERNAL-1 and V14. Re-running it is the cheapest item in the recommendation.
The three specific checks
1. Re-invention. The operator's "at least five instruments" is an undercount, and the sharper finding is when the class was already known.
On 2026-05-25,
knowledge-boundary-calibration/SYNTHESIS_behavioral_knowledge_boundary_2026_05_25.mdcarries a section headed "The recurring adversary: FORM impersonating MEANING", listing three instances inside the instruments. The same defect appears the same day in tier3, council-reference-free-truth, deception-correction-gate and decoupled-diagonal-capstone.SYNTHESIS_mention_and_use_2026_08_26.mdcatalogues ten instances, concludes "claimhood needs its own predicate", and cites none of those five arcs. Checks:grep -rn "impersonating meaning" papers/returns one file;grep -c "sycophancy\|deception-correction\|promptopinion\|truthground\|decoupled" papers/SYNTHESIS_mention_and_use_2026_08_26.mdreturns 0; the stringmention-vs-useappears nowhere in the repo before 2026-08-26.The handed-target pattern is worse: found eleven times across nine arcs, named zero times. The oldest specimen is also the cleanest —
sycophancy-target-gate/FINDING_promptopinion_2026_05_24.mdseparates 1.00 on a fixed-template holdout from 0.47 on fresh varied phrasing, caught by the arc's own preregistered bar before ship.anchored-validity/PAPER_gold_anchors_license_nothing_2026_07_21.mdstates the general law, andpapers/closed-model-frontier/*.mdnames neither anchored-validity nor auditor-ceiling (whose 0.126 false-accusation rate on TriviaQA is the same construct as the 0.23 and the 0.16, one altitude down).Two more independent re-derivations, both cross-arc and both uncited: showcase-viz's "the wall is in the whitening metric, not the channel" (06-12) and disjoint-worlds' "the cliff was the linear map class, not the minds" (08-01); and mind-instrument declaring the unsupervised matcher the bottleneck (06-10) while disjoint-worlds b34v3/b41 break exactly that bottleneck without citing it. Even the doctrine re-derives: read-neq-write's bite gate and anchored-validity's void panel are the same primitive with no cross-reference either way. And the plumbing: 16 deposit scripts, ~3,166 lines, including four near-identical read≠write scripts and one script whose only job is repairing an orphan record the neighbouring script produced.
The count is attacked from both sides in §2. Two places where it is too generous are named, and the harsher reading — 10 questions, 8 mechanisms — is stated as available on the same receipts.
2. The
SYNTHESIS_connection_of_mindscheck — the numbers reconcile, the certificate does not cover the sentence.§2's gemma cliff and the 0.7857 / 0.5714 reads are all true of three different map classes;
b31v2_result.jsoncarriesM0_linear_top1 = 0.0143andM1_mlp_top1 = 0.7857side by side. But §2's sentence says "label-free map", not "linear map" — the word linear is absent from that paragraph — and §3, which never mentions gemma, does not retract it. The reconciliation is one paragraph away, in §2's own appended text, with no inline marker on the superseded sentence.The load-bearing part is checkable in ten lines: the certificate's 94 ledger entries cover lines
[3, 23, 25, 26, 32, 33, …]and there is no entry of any status for line 27, the line carrying "reads at exactly chance 0.014". The document isOATH-HELD,0 UNGROUNDED, andSEALED. Two binder observations from the same file: line 32's392is bound tob41_result.json:n_anchor_rowsrather than theb31v2_result.jsonthe sentence names, and line 33's70is bound toverifier_7b_result.json:not_gated.coverage_curve[1].n. So on this documentOATH-HELDmeans every audited numeral appears somewhere in the receipt bundle, not every claim is supported by the receipt it cites — the same mention-vs-use shape, one level up. The unmarked sentence ships verbatim inpapers/arxiv/connection-of-minds/main.texline 53 and in the submission copy.Eighteen more unmarked kills are tabulated in §5, ranked by how outward-facing they are. Tier 1 is the urgent one: six files under
release/, all dated April 2026, quote AUC 0.998 / 0.976 / 0.943 as empirical validation with no construct-ceiling caveat — six weeks before that ceiling was measured, and none carries an erratum marker. Two of them are addressed to external bodies.The shape is narrower than "the lab is careless", and §5.2 states it: a claim killed by a document that names it, in a file that does not name the killer. The killer almost always exists, is almost always correct, and is usually written the same day. Arcs that carry their own erratum are safe, and §4 exists to say how many do.
3. The prose-reading hypothesis: SUPPORTED-WITH-QUALIFICATION.
Direction 1 holds on a narrower class than stated: every instrument that had to decide claimhood or register from open-ended prose and was then measured against a blind panel or a held-out corpus failed — eleven of them, no exceptions found.
Direction 2 is false. At least six artifact-only instruments failed outright and none for a prose reason: representational-integrity (0.63/0.67, near chance), depth-truth (0.5468, anti-signal OOD), first-afference's coupling detector (0.0083 against a 0.10 bar, withdrawn), mind-instrument cross-species (0.0292 vs a null p95 of 0.0312), conscience-mount B37/B38/B39 (three VOIDs), ancient-question's synthetic RSA (self-retracted as circular). And a counterexample in the other direction: LLM-judge equivalence clustering reads prose and beat the artifact alternative it replaced, 0.948–1.000 against cosine's 0.573.
Both objections are argued rather than waved away. The artifact-only instruments genuinely do postdate the 2026-08-06 standing rule, so the comparison is confounded with date and discipline and cannot be unconfounded on this corpus. And the verdict-free objection is sharper than it looks, because this lab already named the defect: counting a no-verdict instrument as a working instrument is the vacuous-pass error the mention-and-use addendum catalogues and that
dogfood-self-auditmeasured at a 40.4% dead-term rate. On present evidence the artifact-only successors are not known to work; they are known not to have been measured.What survives is better than the hypothesis: instruments that must locate their own target have failed every time they were measured against readers who did not write the prose, and that is the handed-target mechanism with prose/artifact as its proxy. It also explains why direction 2 fails — an artifact instrument that must find its own target fails the same way, and artifact instruments handed their target post this program's largest numbers.
What would settle it is one preregistered head-to-head:
styxx.undeclaredand the retired path-claim accuser, on the same external AIDev PRs, same blind panel, same frozen floor, with extraction measured separately from adjudication. No measurement in this corpus currently separates prose-vs-artifact from handed-vs-found.The recommendation (§8), in one line each
three-axis-sendtime-gate(pre-data since May, three of six modules imported by nothing),cooperative-agent-regime(sign the topic-control prereg or mark the drift-axis positive unlicensed — its deposited receipt'spreregistration_lock_hashis the literal string"TBD-after-operator-signs"), andwhite-box-vs-text-map(its abstract and its own status line cannot both be current).LEDGER.mdalready names as owed and that 0 of 163 cycles carry.CORPUS_STATE_2026_08_31.md, and it is the general repair for the defect that broke three documents in two days.Coverage, disclosed
§9 states what this audit did not do. Roughly 200–250 of 1,135 markdown files were read by nine parallel readers, weighted to late-dated SYNTHESIS/RESULT/FINDING/erratum files; coverage ranged from complete on the two-file arcs to about 8% on
grounded-honesty-axis(16 of 208). Confidence is given per section — §1, §4 and §5-existence are HIGH; §2's clustering, §6 and §7 are MEDIUM with the uncertain parts named. Four arcs are recorded UNCLEAR because their terminal state could not be established, which is the verdict rather than a placeholder.benchmarks/,bench/,demo/,telescope/,web/,integrations/,packages/,hooks/, 172 test files and 197 of the 213 certificates were not examined, so §5 is a lower bound and not a census.Do not merge this on the strength of the summary above. §2 and §7 are judgement calls and the audit says so.
🤖 Generated with Claude Code