Skip to content

docs: bind published counts to the command that re-derives them - #66

Merged
ManSio merged 1 commit into
mainfrom
docs/benchmark-decay-claims
Oct 3, 2026
Merged

ManSio merged 1 commit into
mainfrom
docs/benchmark-decay-claims

Conversation

@ManSio

@ManSio ManSio commented Oct 3, 2026

Copy link
Copy Markdown
Owner

What

Four published numbers had drifted from the code they describe:

number stated live (measured on this tree)
README test badge 1965 passed 2007 collected / 2001 passed
README architecture row ~1889 tests 2007 collected
WISDOM intel census intel_*=14 20
WISDOM test census tests=1180 2007 collected

This is a named class, not sloppiness: benchmark decay — a public static
number whose meaning erodes while the number stays put.

Why a guard and not just an edit

A manual edit fixes it once. Each claim now carries the command that re-derives
it, and disagreement exits rc=3, so the next drift is caught by
un_all.py
instead of by a reader noticing.

The comparison is symmetric. The first cut used live <= stated, which only
catches understatement; its own selftest caught an inflated badge passing through.
Two of the four synthetic cases exist to be rejected, one per direction.

History is superseded, not rewritten

The dated 2026-08-12 snapshot (intel_*=14, tests=1180) was correct when
written
. It is now labelled as a snapshot and left intact; the live census is a
separate section. Published history gets a new record that supersedes it, not a
silent numeric edit.

Also

run_all.py reported the protocol guards' false-positive share as UNMEASURED
after it had been measured. Re-derived from raw: 8 reported, 5 false positives
= 62.5%, 3 actionable. Quoting the raw count overstated defects by 2.7x.

Verified on this tree

  • pytest tests/ → 2001 passed, 6 skipped, exit 0
  • tools/verification/run_all.py → 14/14 provability steps OK, exit 0
  • verify_public_claims.py --selftest → 4/4, including 2 that must be rejected
  • verify_public_claims.py → 4/4 published numbers reproduce
  • pre-commit: 9/9 hooks OK

The guard was confirmed failing on the known-bad state (4/4 WRONG) before the fix
and passing after — held-out on the original defect, not generalised.

Sources

  • GTM-Bench, "Keeping a Benchmark Honest" (2026-09-06) — benchmark decay;
    "every reported score names the version it was run against; a number without a
    version is uninterpretable".
  • READU (arXiv 2607.15780) — an alert judge to remove false positives is part
    of the construction, not an optional extra.
  • driftmd — "badge versions" as its own named check.

Not in this PR

  • G5 denominator n_reviewed is still 0. The tri-state claim classification
    (CLAIM / NOT_A_CLAIM / DERIVED with machine-checked reason codes) is not done
    and is not claimed here.
  • 7 tracked files outside the personal-path guard's declared scope still contain a
    local username (.local/*.py, docs/archive/*.md,
    scripts/reconstruct_judge_cot.py). The guard's boundary is declared in its
    docstring, so this is unreported territory rather than a hole in the guard —
    left for a separate decision, and archives must not be rewritten.

🤖 Generated with opencode

Benchmark decay: the badge read 1965 while the suite collects 2007, the
architecture table read ~1889, and WISDOM's census read intel_*=14 / tests=1180
against a live 20 / 2007. Every one of those is a public number a reader could cite.

A manual edit fixes the symptom once. The guard is the fix: each claim now
carries the command that re-derives it, and disagreement exits rc=3.

The comparison is symmetric on purpose. The first cut used `live <= stated`,
which only catches understatement — its own selftest caught an inflated badge
sliding through. Two synthetic cases must be rejected, and both directions are
now covered.

The dated 2026-08-12 snapshot (intel_*=14, tests=1180) was correct on the day it
was written; it is marked as a snapshot and left intact. The live census is a
separate section, so published history is superseded by a new record rather than
silently rewritten.

Also records the measured false-positive share of the protocol guards (62.5%,
5 of 8) in the suite output that previously still said UNMEASURED — quoting the
raw count overstated defects by 2.7x.

Verified on this tree:
  pytest tests/                     2001 passed, 6 skipped, exit 0
  tools/verification/run_all.py     14/14 provability steps OK, exit 0
  verify_public_claims --selftest   4/4, incl. 2 that must be rejected
  verify_public_claims              4/4 published numbers reproduce

Sources for naming the class: GTM-Bench "Keeping a Benchmark Honest" (2026-09-06)
— benchmark decay, "a number without a version is uninterpretable"; READU
(arXiv 2607.15780) — an alert judge to remove false positives is part of the
construction, not an option; driftmd — "badge versions" as its own check.
@coderabbitai

coderabbitai Bot commented Oct 3, 2026

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 480abcc2-0320-4e70-b81b-0c1fd1906754
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@ManSio
ManSio merged commit 0130850 into main Oct 3, 2026
12 of 13 checks passed
@ManSio
ManSio deleted the docs/benchmark-decay-claims branch October 3, 2026 06:16
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant