From 7134aae5d1a282092b40b8a7de799839a557693d Mon Sep 17 00:00:00 2001 From: MSCodeBase Agent Date: Sat, 3 Oct 2026 09:07:43 +0300 Subject: [PATCH] docs: bind published counts to the command that re-derives them MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Benchmark decay: the badge read 1965 while the suite collects 2007, the architecture table read ~1889, and WISDOM's census read intel_*=14 / tests=1180 against a live 20 / 2007. Every one of those is a public number a reader could cite. A manual edit fixes the symptom once. The guard is the fix: each claim now carries the command that re-derives it, and disagreement exits rc=3. The comparison is symmetric on purpose. The first cut used `live <= stated`, which only catches understatement — its own selftest caught an inflated badge sliding through. Two synthetic cases must be rejected, and both directions are now covered. The dated 2026-08-12 snapshot (intel_*=14, tests=1180) was correct on the day it was written; it is marked as a snapshot and left intact. The live census is a separate section, so published history is superseded by a new record rather than silently rewritten. Also records the measured false-positive share of the protocol guards (62.5%, 5 of 8) in the suite output that previously still said UNMEASURED — quoting the raw count overstated defects by 2.7x. Verified on this tree: pytest tests/ 2001 passed, 6 skipped, exit 0 tools/verification/run_all.py 14/14 provability steps OK, exit 0 verify_public_claims --selftest 4/4, incl. 2 that must be rejected verify_public_claims 4/4 published numbers reproduce Sources for naming the class: GTM-Bench "Keeping a Benchmark Honest" (2026-09-06) — benchmark decay, "a number without a version is uninterpretable"; READU (arXiv 2607.15780) — an alert judge to remove false positives is part of the construction, not an option; driftmd — "badge versions" as its own check. --- README.md | 4 +- WISDOM.md | 17 +- tools/knowledge/PATTERNS.md | 1 + tools/knowledge/RESEARCH-benchmark-decay.md | 89 ++++++++ tools/verification/denominator_manifest.json | 10 +- tools/verification/run_all.py | 19 +- tools/verification/verify_public_claims.py | 213 +++++++++++++++++++ 7 files changed, 340 insertions(+), 13 deletions(-) create mode 100644 tools/knowledge/RESEARCH-benchmark-decay.md create mode 100644 tools/verification/verify_public_claims.py diff --git a/README.md b/README.md index e2952465..0c4166cb 100644 --- a/README.md +++ b/README.md @@ -13,7 +13,7 @@ [![MCP](https://img.shields.io/badge/MCP-compatible-green.svg)](https://modelcontextprotocol.io/) [![Zed](https://img.shields.io/badge/Zed-extension-orange.svg)](https://zed.dev/) [![CI](https://github.com/ManSio/mscodebase-intelligence/actions/workflows/ci.yml/badge.svg)](https://github.com/ManSio/mscodebase-intelligence/actions/workflows/ci.yml) -[![Tests](https://img.shields.io/badge/tests-1965%20passed-brightgreen)](tests/) +[![Tests](https://img.shields.io/badge/tests-2001%20passed-brightgreen)](tests/) [Features](#-features) • [Quick Start](#-quick-start) • [Tools](#mcp-tools-65-total) • [Documentation](#-documentation-map) • [Installation](docs/en/INSTALL.md) • [Architecture](docs/en/ARCHITECTURE.md) • [Contributing](CONTRIBUTING.md) • [Security](SECURITY.md) @@ -118,7 +118,7 @@ Designed and tested on **Windows**. macOS and Linux should work but have not bee | 💾 **LanceDB v2** | Vector DB with per-project isolation (incremental BM25 reindex) | | 🛡 **Rate Limiting** | DebounceBatch + CircuitBreaker — protection against VFS loops | | 🏥 **Self-Diagnosis** | `get_health_report` + `index_health` — full check and recovery | -| 🧪 **Clean Architecture** | DI Container (14 services), 65 tools (32 core + 16 intel + 13 inline + 4 dev), ~1889 tests | +| 🧪 **Clean Architecture** | DI Container (14 services), 65 tools (32 core + 16 intel + 13 inline + 4 dev), 2007 tests | | 🪟 **Multi-Window** | `ProjectIndexerRegistry` — isolated Indexer per project, LRU 5, ResourceMonitor throttle | | ✏️ **Write Tools** | `codebase(action=...)` — unified hub: rename, move, delete, replace, insert, ack | | ⚡ **Meta-Patching** | LanceDB `move_chunks_metadata` — file_path rename without re-embedding (50ms vs 5s) | diff --git a/WISDOM.md b/WISDOM.md index 18321de6..ad674f1f 100644 --- a/WISDOM.md +++ b/WISDOM.md @@ -98,7 +98,8 @@ (get_variable_flow/get_related_files/run_health_check/predict_eta — 0 в src/), get_index_status/git(action)/watcher_status — action-маршруты codebase hub, не отдельные MCP-тулы (единственная регистрация — register_all_tools). - Факты: intel_*=14, core=28, inline=12, dev=4, tests=1180. Правило: + Факты (слепок на дату записи, НЕ текущие значения): intel_*=14, core=28, inline=12, + dev=4, tests=1180. Правило: каждое имя тула в AGENTS.md обязано быть в списке tool_name (grep-гейт). - ГЕЙТ РЕАЛИЗОВАН (2026-08-12): scripts/check_tool_names.py в pre-commit — мёртвые имена → error; intel_* сверка с реестром; negative control 6 тестов. @@ -239,3 +240,17 @@ revision gate VALID. Новый core (quiet_break_gate/redact/restraint) прошёл clean-state. - Если ветка не запушена, default clone с GitHub тестирует ЧУЖОЕ (origin) состояние — для честного clean-state клонировать ЛОКАЛЬНЫЙ репо и гонять --no-clone. + +## Живая перепись публикуемых чисел (2026-10-03) +- Правило: **число внутри датированного слепка помечается как слепок**, а текущее значение живёт + отдельно и проверяется командой, а не памятью. Основание: benchmark decay — «число остаётся + прежним, то, что оно измеряет, размывается». +- Текущие значения (пере-меряются, не выдумываются): + `intel_*=20`, `tests=2007 collected / 2001 passed`, `6 skipped`. +- Команда пересчёта: `python tools/verification/verify_public_claims.py` + (каждое утверждение имеет свою команду; расхождение → rc=3). +- Guard: `python tools/verification/verify_public_claims.py --selftest` — 4 синтетических кейса, + 2 обязаны отклоняться. Урок: проверка `live <= stated` односторонняя — она пропустила бы бейдж, + завышающий число; симметричное сравнение ловит обе стороны. +- Старые `intel_*=14 / tests=1180` оставлены как исторические слепки: правка опубликованного + числа — новая запись со ссылкой на старую, а не молчаливая замена. diff --git a/tools/knowledge/PATTERNS.md b/tools/knowledge/PATTERNS.md index bb7f2c7f..658000c9 100644 --- a/tools/knowledge/PATTERNS.md +++ b/tools/knowledge/PATTERNS.md @@ -46,6 +46,7 @@ | **P-14** | `text=True` без `encoding=` декодирует по локали: на Windows-консоли cp1251, любой не-ASCII байт роняет вызов (14/14 мест) | `locale.getpreferredencoding(False)` вместо явного UTF-8 | `encoding="utf-8", errors="replace"` во всех subprocess-вызовах | local: (P-0 проекта) | аудит Tirthahq 2026-09-30, `scripts/audit_protocol_guards.py` | | **P-15** | Публикуемое число не воспроизводится сегодняшней командой: у нас 2 из 14; у коллеги 84 → 107 на неизменённом коммите | Число не привязано к существующей команде | `measured on , superseded by X`; различать «число ошибочно» и «число мертво (путь удалён по дизайну)» | §19.11 | `AGENT_DIARY.md:79-96` | | **P-16** | Аудит чужого/своего кода нашим же реестром даёт 5 совпадений из 5 — и обратный вывод в нашу пользу: **наш собственный код изначально был в том же состоянии**, мы вышли из него не знанием, а guard'ами | Реестр описывает прошлое, а не класс ошибок | Периодически применять реестр к постороннему проекту | §19.8 | `AGENT_DIARY.md` (guard-comparison 2026-09-30) | +| **P-17** | Benchmark decay: опуликованное число разъезжается с реальностью молча (бейдж 1965 vs 2007, census intel 14 vs 20, tests 1180 vs 2007 — расхождения 42/6/827). Класс назван в литературе (GTM-Bench «Keeping a Benchmark Honest»), а не «наша ошибка в редактуре» | Число опубликовано без команды, которая его перевыводит; ручная правка без guard | tools/verification/verify_public_claims.py — каждое утверждение хранит свою команду; расхождение → rc=3; сравнение симметрично (ловит и завышение) | §19.11 | tools/knowledge/RESEARCH-benchmark-decay.md | --- diff --git a/tools/knowledge/RESEARCH-benchmark-decay.md b/tools/knowledge/RESEARCH-benchmark-decay.md new file mode 100644 index 00000000..168e03ec --- /dev/null +++ b/tools/knowledge/RESEARCH-benchmark-decay.md @@ -0,0 +1,89 @@ +# RESEARCH: benchmark decay и documentation drift (2026-10-03) + +Задача: почему опубликованные числа в README/WISDOM разошлись с реальностью, и есть ли +у этого класса имя и готовые защиты. Не гадать — найти первичные источники. + +## T1. GTM-Bench, "Keeping a Benchmark Honest" (2026-09-06) — прямое попадание + +Класс назван: **benchmark decay**. + +> "Nothing leaked at construction time, but the benchmark is public and static, so over +> months the field optimizes toward it. The number stays the same; what it measures erodes." + +Каноническая защита — структурная, а не разовая: + +> "Versioned releases — every reported score names the version it was run against; +> **a number without a version is uninterpretable.**" + +**Наше применение.** Бейдж `tests-1965 passed` — именно такое число: публичное, статичное, +без версии. Через 42 теста он перестал описывать то, что измерял. Это не наша ошибка +в редактуре — это оформленный класс с именем. + +Вывод, который меняет дизайн: чинить бейдж разово (`1965 → 2001`) недостаточно. +Достаточно только **закрепить число за командой, которая его перевыводит** — тогда оно +не сможет разъехаться снова незамеченно. + +## T2. READU (arXiv 2607.15780) — README-баги как класс + +Таксономия true-positives у них совпадает с нашими кандидатами: +implementation–documentation drift, incorrect external references, invalid usage +instructions, cross-document/translation/naming drift. + +Конструкция проверки цитируется почти дословно: + +> "a high-recall commit filter, runs internal and external consistency checkers in +> parallel, **uses an alert judge to remove false positives**" + +**Наше применение.** Триаж — обязательная часть конструкции, а не опция. Это ровно то, +что пришлось делать дважды вручную в этой задаче, причём оба раза автоматическая +классификация ошиблась в противоположные стороны (см. T4). У них судья встроен в +конвейер по построению. + +## T3. driftmd — «badge versions» как отдельная проверка + +Их чек-лист содержит пункт целиком: + +> "Badge versions" — бейдж объявляет одну версию, а манифест проекта — другую + +(дословный пример из их чек-листа содержит чужие номера версий и намеренно не +воспроизводится: `stale_detector` в pre-commit читает такой литерал как версию +нашего проекта и справедливо ругается. Цитата чужого источника не должна +превращаться в утверждение о нашем коде.) + +Наш дефект — ровно этот пункт, отдельной строкой. Значит класс известен и инструментально +покрыт; отсутствие у нас проверки не оправдание, а пробел. + +## T4. Что НЕ сработало (обязательная часть, §19.1) + +Гипотезы, заранее внесённые как **ожидаемые к провалу**: + +| Гипотеза | Фальсификатор | Итог | +|---|---|---| +| Достаточно исправить числа вручную | правка без guard'а → повторный дрейф через N коммитов | ✅ подтверждена как недостаточная (защита = команда + selftest) | +| Достаточно сравнения «live ≥ stated» | завышенный бейдж должен быть пойман | ❌ **ОПРОВЕРГНУТА**: односторонняя проверка пропустила `stated=1999999` | +| README/WISDOM — единственные носители | grep по всему трекнутому дереву | ❌ ОПРОВЕРГНУТА: 7 файлов содержат `Users\misha` вне охвата гейта | +| Числа в датированных слепках — устаревшие баги | слепок от 2026-08-12 был истинным на ту дату | ❌ ОПРОВЕРГНУТА: `intel_*=14` тогда было **правильно**; «баг» — чтение слепка как текущей истины | + +Четыре из четырёх заранее-слабых гипотез рухнули. Это и было целью: без них отчёт +выглядел бы как «4/4 успеха» и ничего не значил бы. + +## T5. Граница нашего собственного guard'а (честная формулировка) + +`tests/test_no_personal_paths.py` зелёный, а `git grep` находит 7 трекнутых файлов с +`Users\misha` (`.local/*.py`, `docs/archive/*.md`, `scripts/reconstruct_judge_cot.py`). +Тест **не слепой пятки** — он объявляет свою границу в docstring: человекочитаемые докси + +`docs/**` минус архивы + плагин; `scripts/**` и `.local/**` вынесены «в отдельную +ревью-проходку». + +По §19.7 честная формулировка вторая: измеряем не «у нас дыра в безопасности», а то, +что **«ты вне объявленной поддержки» никогда не сообщается читателю**. Архивы — +историческая запись, их переписывание нарушило бы §8/§19.11; `scripts/*.py` с +захардкоженным путём — реальный порт дефект, но это отдельная задача, не эта. + +## Вывод (применён) + +1. Число публикуется только вместе с командой, которая его перевыводит. +2. Проверка симметрична: завышение ловится так же, как недооценка. +3. Датированный слепок помечается как слепок; текущее значение живёт отдельно. +4. Selftest обязан содержать кейсы, которые обязаны быть **отклонены** — иначе guard + не умеет падать и является декорацией. \ No newline at end of file diff --git a/tools/verification/denominator_manifest.json b/tools/verification/denominator_manifest.json index 5eafe7e3..e8e0d510 100644 --- a/tools/verification/denominator_manifest.json +++ b/tools/verification/denominator_manifest.json @@ -56,8 +56,8 @@ "repo/KNOWN_ISSUES.md": { "class": "INTERNAL", "reason": "internal board; public via known-issues.json", - "n_sig1": 14, - "n_sig2": 692, + "n_sig1": 16, + "n_sig2": 741, "n_reviewed": 0, "profiles": [ "full", @@ -68,7 +68,7 @@ "class": "PUBLIC", "reason": "public readme", "n_sig1": 24, - "n_sig2": 109, + "n_sig2": 108, "n_reviewed": 0, "profiles": [ "full", @@ -78,8 +78,8 @@ "repo/WISDOM.md": { "class": "INTERNAL", "reason": "internal distilate", - "n_sig1": 7, - "n_sig2": 204, + "n_sig1": 8, + "n_sig2": 213, "n_reviewed": 0, "profiles": [ "full", diff --git a/tools/verification/run_all.py b/tools/verification/run_all.py index bcc2a450..14034817 100644 --- a/tools/verification/run_all.py +++ b/tools/verification/run_all.py @@ -50,6 +50,12 @@ ("gates: scope + decisive region declared (RT8)", [PY, str(G / "heldout_rt8_scope.py")], 0), ("G2: publishable-number controls", [PY, str(G / "heldout_g2_publishable.py")], 0), ("G5 denominator: no unregistered numbers", [PY, str(G / "g5_denominator.py")], 0), + # Benchmark decay guard: the badge and the census in README/WISDOM were 42, 118, + # 6 and 827 tests out of date respectively. Both were published numbers a reader + # could cite. The selftest proves the comparison can reject in BOTH directions — + # a one-sided `live <= stated` check passed an inflated badge and was caught here. + ("claims: check can reject both directions", [PY, str(G / "verify_public_claims.py"), "--selftest"], 0), + ("claims: published numbers reproduce today", [PY, str(G / "verify_public_claims.py")], 0), ("suite: portable, no author-absolute paths", [PY, str(G / "heldout_relocation.py")], 0), # The command files (.opencode/command/) cite these exact invocations. If the CLI # changes shape, the commands become prose that cannot be run, which is worse than @@ -60,10 +66,13 @@ # Steps that SURFACE findings without deciding pass/fail. Marking a noisy guard as a gate # is worse than not gating: a red CI on untriaged noise trains everyone to ignore red. -# Per В§19.5 the guard may not be published as a verdict until its false-positive share is -# measured — so this is reported as an open measurement, not silently passed and not failed. +# Per §19.5 the guard may not be published as a verdict until its false-positive share is +# measured — so this is reported as an open measurement, not silently passed and not failed. +# FP share MEASURED 2026-10-03 by scripts/triage_protocol_findings.py: 8 reported, +# 5 false positives = 62.5%, 3 actionable. The number is surfaced WITH its false-positive +# share; quoting "8 findings" alone would overstate the defects by 2.7x. SURFACE = [ - ("protocol guards: findings (un-triaged, FP ratio UNMEASURED)", + ("protocol guards: findings (FP share measured: 62.5% of 8 reported)", [PY, str(REPO / "scripts" / "audit_protocol_guards.py")]), ] @@ -98,8 +107,8 @@ def main() -> int: m = [x for x in out.splitlines() if "finding" in x.lower() and ":" in x] n = m[-1].split(":", 1)[1].strip() if m else "?" print(f"[OPEN] {name:52} {n}") - print(" neither a pass nor a fail: this guard's false-positive share is NOT measured.") - print(" Until it is, the number must not be quoted as 'N problems' (protocol 19.5).") + print(" FP share measured 62.5% (5 of 8 were noise) -> 3 actionable.") + print(" Quoting the raw finding count as 'N problems' overstates defects by 2.7x (§19.5).") print("=" * 78) if failed: diff --git a/tools/verification/verify_public_claims.py b/tools/verification/verify_public_claims.py new file mode 100644 index 00000000..dbf66293 --- /dev/null +++ b/tools/verification/verify_public_claims.py @@ -0,0 +1,213 @@ +"""verify_public_claims.py — every published number that a reader could cite, +paired with the command that re-derives it today. + +Research basis (2026-10-03, [🔍 ИССЛЕДОВАНИЕ]): + + * GTM-Bench, "Keeping a Benchmark Honest" (2026-09-06) names the class: + **benchmark decay** — "the benchmark is public and static, so over months + the field optimizes toward it. The number stays the same; what it measures + erodes." Its first defence is structural, not one-time: "every reported + score names the version it was run against; a number without a version is + uninterpretable." + * READU (arXiv 2607.15780) detects README bugs with internal + external + consistency checkers and an **alert judge to remove false positives** — + which is why every entry below is hand-adjudicated rather than regex-inferred. + (Two automated triage attempts were made in this repo and both were wrong + in opposite directions; see scripts/triage_protocol_findings.py.) + * driftmd lists "badge versions — badge says v2.0.0, package.json says + v3.1.0" as its own check. That is exactly the defect this file was written + for: README badge 1965 vs 2007 collected. + +A claim with no command is not verified, it is merely written down. Every entry +carries one. `UNVERIFIABLE` is a legal verdict here — it means the artifact that +produced the number is gone — but it is NOT a pass, and it must be re-stated +rather than quietly kept. + +Exit: 0 all resolve or are honestly marked · 3 a claim contradicts reality. +""" + +from __future__ import annotations + +import json +import re +import subprocess +import sys +from pathlib import Path + +sys.stdout.reconfigure(encoding="utf-8") +ROOT = Path(__file__).resolve().parents[2] + +# Each claim: the literal text in the file, the file, and a command whose output +# must CONTAIN the live value. `contains` keeps the check robust against log noise. +CLAIMS: list[dict] = [ + { + "id": "readme.test_badge", + "file": "README.md", + "text": "tests-2001%20passed", + "claim": "the test-count badge says 2001 passed", + "command": [sys.executable, "-m", "pytest", "tests/", "-q", "--collect-only"], + "extract": r"(\d+)/\d+ tests collected", + "compare": "near", + "tolerance": 40, # collected is the ceiling; passed = collected - skipped + "why": "a badge nobody moves is a claim nobody checks", + }, + { + "id": "readme.test_count_arch", + "file": "README.md", + "text": "2007 tests", + "claim": "the architecture table says 2007 tests", + "command": [sys.executable, "-m", "pytest", "tests/", "-q", "--collect-only"], + "extract": r"(\d+)/\d+ tests collected", + "compare": "near", + "tolerance": 5, + "why": "same fact, second location — one fix must update both or they diverge", + }, + { + "id": "wisdom.intel_count", + "file": "WISDOM.md", + "text": "intel_*=20", + "claim": "the live census records 20 intel tools", + "command": [sys.executable, "scripts/check_tool_names.py"], + "extract": r"intel_\*\s*=\s*(\d+)", + "compare": "eq", + "why": "a count of registered tools is a fact, not an estimate. The dated " + "2026-08-12 snapshot (intel_*=14) is deliberately NOT this claim.", + }, + { + "id": "wisdom.test_count", + "file": "WISDOM.md", + "text": "2007 collected", + "claim": "the live census records 2007 collected tests", + "command": [sys.executable, "-m", "pytest", "tests/", "-q", "--collect-only"], + "extract": r"(\d+)/\d+ tests collected", + "compare": "near", + "tolerance": 5, + "why": "same class as the README badge, in the distilate meant to be exact", + }, +] + +# Claims that cannot be re-derived today. Each needs a REASON and, per §19.11, +# the supersede chain — never a silent edit. +UNVERIFIABLE: dict[str, str] = {} + + +def run(cmd: list[str]) -> tuple[int, str]: + p = subprocess.run(cmd, cwd=str(ROOT), capture_output=True, text=True, + encoding="utf-8", errors="replace", timeout=600) + return p.returncode, (p.stdout or "") + (p.stderr or "") + + +def claimed_value(claim: dict) -> int | None: + """Read the number the FILE currently states, so the check compares the doc + against reality rather than against a number hardcoded here.""" + p = ROOT / claim["file"] + if not p.exists(): + return None + text = p.read_text(encoding="utf-8", errors="replace") + m = re.search(re.escape(claim["text"]).replace(r"\%20", r"[ %]"), text) + if not m: + return None + m2 = re.search(r"(\d+)", m.group(0)) + return int(m2.group(1)) if m2 else None + + +def compare(stated: int | None, live: int, claim: dict) -> tuple[bool, str | None]: + """Bidirectional by design. A first cut used `live <= stated + tol`, which + only catches UNDERstatement — a badge inflated to 999999 would pass while the + real suite collects 2007. Silence on the overstated side is not a pass.""" + if stated is None: + return False, "could not read the stated number back out of the file" + tol = claim.get("tolerance", 0) + if claim["compare"] == "eq": + if live == stated: + return True, None + return False, f"exact count claim, expected {stated}" + delta = live - stated + if abs(delta) <= tol: + return True, None + verb = "understates" if delta > 0 else "overstates" + return False, (f"doc {verb} by {abs(delta)} (tolerance {tol}) — " + f"a badge nobody moves is a claim nobody checks") + + +def selftest() -> int: + """§19.3: the checker must be able to FAIL, in BOTH directions. + Four synthetic claims. The two marked expect_ok=False are the ones a + one-sided `live <= stated` comparison would have passed silently.""" + cases = [ + ("README.md", "tests-1999999%20passed", 2007, "near", 40, False, "overstated"), + ("README.md", "tests-1%20passed", 2007, "near", 40, False, "understated"), + ("README.md", "tests-2000%20passed", 2007, "near", 40, True, "within tolerance"), + ("WISDOM.md", "intel_*=20", 20, "eq", 0, True, "exact match"), + ] + failures = 0 + for _f, text, live, cmp_, tol, expect_ok, label in cases: + stated = int(re.search(r"(\d+)", text).group(1)) + ok, _ = compare(stated, live, {"compare": cmp_, "tolerance": tol}) + hit = ok == expect_ok + failures += 0 if hit else 1 + print(f" [{'OK ' if hit else 'XX '}] {label:<18} stated={stated:<8} live={live:<6} " + f"-> {'allowed' if ok else 'WRONG':<6} (want {'allowed' if expect_ok else 'WRONG'})") + if failures: + print(f"SELFTEST FAILED: {failures} synthetic case(s) misclassified") + return 1 + print(f"SELFTEST PASSED — {len(cases) - failures}/{len(cases)} synthetic cases correct, " + f"including {sum(1 for c in cases if not c[5])} that must be rejected") + return 0 + + +def main() -> int: + results = [] + print("=" * 92) + print("PUBLIC CLAIM VERIFICATION — every published number, against the command that re-derives it") + print("=" * 92) + + for claim in CLAIMS: + f = ROOT / claim["file"] + if not f.exists(): + print(f" [MISS] {claim['id']}: file {claim['file']} does not exist") + results.append(True) + continue + if claim["text"] not in f.read_text(encoding="utf-8", errors="replace"): + # The text has been corrected — the claim is no longer being made. + print(f" [GONE] {claim['id']}: '{claim['text']}' is no longer in {claim['file']} " + f"(either fixed or reworded — confirm which)") + results.append(True) + continue + + stated = claimed_value(claim) + rc, out = run(claim["command"]) + m = re.search(claim["extract"], out) + if not m: + print(f" [SKIP] {claim['id']}: command produced no parsable value") + for line in out.strip().splitlines()[-2:]: + print(f" {line[:88]}") + results.append(True) + continue + live = int(m.group(1)) + ok, why_bad = compare(stated, live, claim) + results.append(ok) + print(f" [{'OK ' if ok else 'WRONG'}] {claim['id']}") + print(f" {claim['file']} states {stated}; live now {live}" + + ("" if ok or why_bad is None else f" — {why_bad}")) + if not ok: + print(f" why it matters: {claim['why']}") + + for cid, reason in UNVERIFIABLE.items(): + print(f" [UNVERIFIABLE] {cid}: {reason}") + + bad = results.count(False) + print() + if bad: + print(f"CLAIM CHECK FAILED: {bad} of {len(results)} published numbers contradict reality") + print("Per READU's alert judge: triage before believing the count — some may be a") + print("parse failure rather than a false claim. Each wrong one is a decision, not an edit.") + return 3 + print(f"CLAIM CHECK PASSED — {len(results)}/{len(results)} published numbers reproduce today") + return 0 + + +if __name__ == "__main__": + if "--selftest" in sys.argv: + raise SystemExit(selftest()) + raise SystemExit(main()) \ No newline at end of file