diff --git a/docs/audit/2026-08-26-audit.html b/docs/audit/2026-08-26-audit.html new file mode 100644 index 00000000..fb5bafa8 --- /dev/null +++ b/docs/audit/2026-08-26-audit.html @@ -0,0 +1,272 @@ + + + + + +Audit — checkmydata-ai + + + +
+

project-audit · 2026-08-26T16:33:49Z · read-only

+

checkmydata-ai

+

Every number below was produced by a command this run executed. What could not be measured is listed as such, never omitted and never counted as clean.

+
+
+
+
29findings
+
7probes run
+
2blind
+
0closed since last
+
0new
+
0open 3+ runs
+
+

First run. There is no earlier sidecar in this directory, so nothing is reported as closed or new — a diff against a run that never happened would be a claim about nothing. The next audit will have both columns.

+

What this project is

+
+ + + + + + + + + +
version— (no manifest version)
languages
package managers
monorepoyes
submodules0
CIgithub-actions
deploy targetscompose, procfile
error telemetrynone found in manifests
tracked files1500
+

Findings

+
+

critical A private key block is committed in the tree

+

backend/tests/integration/test_analytics_collect_e2e.py:126 · P=3.0 · first seen 2026-08-26

+

Remedy. rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not

+
+
+

critical A private key block is committed in the tree

+

backend/tests/integration/test_analytics_connection_security.py:49 · P=3.0 · first seen 2026-08-26

+

Remedy. rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not

+
+
+

critical A private key block is committed in the tree

+

backend/tests/integration/test_analytics_connections_api.py:48 · P=3.0 · first seen 2026-08-26

+

Remedy. rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not

+
+
+

critical A private key block is committed in the tree

+

backend/tests/integration/test_security_rbac.py:302 · P=3.0 · first seen 2026-08-26

+

Remedy. rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not

+
+
+

critical A private key block is committed in the tree

+

backend/tests/integration/test_vendor_credentials_api.py:31 · P=3.0 · first seen 2026-08-26

+

Remedy. rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not

+
+
+

critical A private key block is committed in the tree

+

backend/tests/integration/test_vendor_credentials_api.py:185 · P=3.0 · first seen 2026-08-26

+

Remedy. rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not

+
+
+

critical A private key block is committed in the tree

+

backend/tests/unit/analytics/test_ga4_adapter.py:933 · P=3.0 · first seen 2026-08-26

+

Remedy. rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not

+
+
+

critical A private key block is committed in the tree

+

backend/tests/unit/services/test_vendor_credential_no_secret_logging.py:36 · P=3.0 · first seen 2026-08-26

+

Remedy. rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not

+
+
+

critical A private key block is committed in the tree

+

backend/tests/unit/services/test_vendor_credential_service.py:39 · P=3.0 · first seen 2026-08-26

+

Remedy. rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not

+
+
+

critical A private key block is committed in the tree

+

backend/tests/unit/services/test_vendor_credential_service.py:139 · P=3.0 · first seen 2026-08-26

+

Remedy. rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not

+
+
+

critical A private key block is committed in the tree

+

backend/tests/unit/services/test_vendor_credential_service.py:161 · P=3.0 · first seen 2026-08-26

+

Remedy. rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not

+
+
+

critical A private key block is committed in the tree

+

backend/tests/unit/test_connection_config_secrecy.py:19 · P=3.0 · first seen 2026-08-26

+

Remedy. rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not

+
+
+

critical A private key block is committed in the tree

+

backend/tests/unit/test_connection_lifecycle.py:90 · P=3.0 · first seen 2026-08-26

+

Remedy. rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not

+
+
+

critical A private key block is committed in the tree

+

backend/tests/unit/test_ssh_key_routes.py:60 · P=3.0 · first seen 2026-08-26

+

Remedy. rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not

+
+
+

critical A private key block is committed in the tree

+

backend/tests/unit/test_ssh_tunnel.py:326 · P=3.0 · first seen 2026-08-26

+

Remedy. rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not

+
+
+

critical A private key block is committed in the tree

+

frontend/src/__tests__/components/GA4ConnectionForm.test.tsx:288 · P=3.0 · first seen 2026-08-26

+

Remedy. rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not

+
+
+

critical A private key block is committed in the tree

+

frontend/src/__tests__/components/VendorCredentialsPanel.test.tsx:122 · P=3.0 · first seen 2026-08-26

+

Remedy. rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not

+
+
+

critical A private key block is committed in the tree

+

frontend/src/components/ssh/SshKeyManager.tsx:109 · P=3.0 · first seen 2026-08-26

+

Remedy. rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not

+
+
+

critical A private key block is committed in the tree

+

frontend/src/components/ssh/SshKeyManager.tsx:258 · P=3.0 · first seen 2026-08-26

+

Remedy. rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not

+
+
+

critical A private key block is committed in the tree

+

frontend/src/lib/api/vendor-credentials.ts:52 · P=3.0 · first seen 2026-08-26

+

Remedy. rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not

+
+
+

critical A private key block appears in git history

+

commit c273fa761381 · P=1.5 · first seen 2026-08-26

+

Remedy. rotate the credential at its issuer

+
+
+

critical A private key block appears in git history

+

commit 5acb82d1fd17 · P=1.5 · first seen 2026-08-26

+

Remedy. rotate the credential at its issuer

+
+
+

critical A private key block appears in git history

+

commit d6388423fb55 · P=1.5 · first seen 2026-08-26

+

Remedy. rotate the credential at its issuer

+
+
+

critical A private key block appears in git history

+

commit 4fac5a7d44ad · P=1.5 · first seen 2026-08-26

+

Remedy. rotate the credential at its issuer

+
+
+

critical A private key block appears in git history

+

commit 3e34591e9045 · P=1.5 · first seen 2026-08-26

+

Remedy. rotate the credential at its issuer

+
+
+

critical A private key block appears in git history

+

commit d3f7dc9cd8c0 · P=1.5 · first seen 2026-08-26

+

Remedy. rotate the credential at its issuer

+
+
+

critical A private key block appears in git history

+

commit c9c40974bea5 · P=1.5 · first seen 2026-08-26

+

Remedy. rotate the credential at its issuer

+
+
+

critical A private key block appears in git history

+

commit fbf811253b42 · P=1.5 · first seen 2026-08-26

+

Remedy. rotate the credential at its issuer

+
+
+

medium A deployed surface reports no errors anywhere its maintainer can see

+

manifests · P=1.0 · first seen 2026-08-26

+

Deploy targets declared (compose, procfile) with no telemetry dependency in any manifest. A failure on a user's machine is invisible.

+
languages= deploy=compose,procfile telemetry=[]
+

Remedy. add an error reporter, or record the decision not to — the gap worth closing is that nobody wrote down which it is

+
+

What was not looked at

+

A probe that could not run returns blind, never clean. An empty section here would mean every probe answered — not that nothing is wrong.

+
+ + +
ProbePhaseWhy not
channel-divergenceprodno version in a manifest to make a claim about
published-versionprodno name+version in a manifest
+

Probes

+
+ + + + + + + + + +
ProbePhaseVerdictNote
secrets-treeprobefinding20 credential pattern(s)
secrets-historyprobefinding8 in history
worktreeprobecleanworking tree clean
telemetryprodfindingno error reporting found
ci-presentprodcleanCI configured: github-actions
docs-presentseamsclean7 documentation marker(s): README.md, docs, CLAUDE.md, CONTRIBUTING.md, CHANGELOG.md, docs/adr, ARCHITECTURE.md
gitignore-secretsprobecleanno credential-shaped file is tracked
channel-divergenceprodblindno version in a manifest to make a claim about
published-versionprodblindno name+version in a manifest
+ +
+ + \ No newline at end of file diff --git a/docs/audit/2026-08-26-audit.json b/docs/audit/2026-08-26-audit.json new file mode 100644 index 00000000..89cef0bf --- /dev/null +++ b/docs/audit/2026-08-26-audit.json @@ -0,0 +1,602 @@ +{ + "capabilities": [ + "cargo", + "docker", + "gh", + "git", + "network", + "node", + "npm", + "python3" + ], + "counts": { + "findings": 29, + "probes_blind": 2, + "probes_run": 7 + }, + "findings": [ + { + "blast": 3, + "detail": "", + "effort": 1, + "evidence": "", + "first_seen": "2026-08-26", + "id": "f-cd7a032e015a", + "p": 3.0, + "probe": "secrets-tree", + "remedy": "rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not", + "runs_open": 0, + "severity": "critical", + "title": "A private key block is committed in the tree", + "where": "backend/tests/integration/test_analytics_collect_e2e.py:126" + }, + { + "blast": 3, + "detail": "", + "effort": 1, + "evidence": "", + "first_seen": "2026-08-26", + "id": "f-466ab3c7e8d3", + "p": 3.0, + "probe": "secrets-tree", + "remedy": "rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not", + "runs_open": 0, + "severity": "critical", + "title": "A private key block is committed in the tree", + "where": "backend/tests/integration/test_analytics_connection_security.py:49" + }, + { + "blast": 3, + "detail": "", + "effort": 1, + "evidence": "", + "first_seen": "2026-08-26", + "id": "f-15bc59dcad13", + "p": 3.0, + "probe": "secrets-tree", + "remedy": "rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not", + "runs_open": 0, + "severity": "critical", + "title": "A private key block is committed in the tree", + "where": "backend/tests/integration/test_analytics_connections_api.py:48" + }, + { + "blast": 3, + "detail": "", + "effort": 1, + "evidence": "", + "first_seen": "2026-08-26", + "id": "f-ffa99d263ebb", + "p": 3.0, + "probe": "secrets-tree", + "remedy": "rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not", + "runs_open": 0, + "severity": "critical", + "title": "A private key block is committed in the tree", + "where": "backend/tests/integration/test_security_rbac.py:302" + }, + { + "blast": 3, + "detail": "", + "effort": 1, + "evidence": "", + "first_seen": "2026-08-26", + "id": "f-91fc925fc4af", + "p": 3.0, + "probe": "secrets-tree", + "remedy": "rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not", + "runs_open": 0, + "severity": "critical", + "title": "A private key block is committed in the tree", + "where": "backend/tests/integration/test_vendor_credentials_api.py:31" + }, + { + "blast": 3, + "detail": "", + "effort": 1, + "evidence": "", + "first_seen": "2026-08-26", + "id": "f-9e004d63832f", + "p": 3.0, + "probe": "secrets-tree", + "remedy": "rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not", + "runs_open": 0, + "severity": "critical", + "title": "A private key block is committed in the tree", + "where": "backend/tests/integration/test_vendor_credentials_api.py:185" + }, + { + "blast": 3, + "detail": "", + "effort": 1, + "evidence": "", + "first_seen": "2026-08-26", + "id": "f-c3b1f338966d", + "p": 3.0, + "probe": "secrets-tree", + "remedy": "rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not", + "runs_open": 0, + "severity": "critical", + "title": "A private key block is committed in the tree", + "where": "backend/tests/unit/analytics/test_ga4_adapter.py:933" + }, + { + "blast": 3, + "detail": "", + "effort": 1, + "evidence": "", + "first_seen": "2026-08-26", + "id": "f-4bd961a817ba", + "p": 3.0, + "probe": "secrets-tree", + "remedy": "rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not", + "runs_open": 0, + "severity": "critical", + "title": "A private key block is committed in the tree", + "where": "backend/tests/unit/services/test_vendor_credential_no_secret_logging.py:36" + }, + { + "blast": 3, + "detail": "", + "effort": 1, + "evidence": "", + "first_seen": "2026-08-26", + "id": "f-036f91903887", + "p": 3.0, + "probe": "secrets-tree", + "remedy": "rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not", + "runs_open": 0, + "severity": "critical", + "title": "A private key block is committed in the tree", + "where": "backend/tests/unit/services/test_vendor_credential_service.py:39" + }, + { + "blast": 3, + "detail": "", + "effort": 1, + "evidence": "", + "first_seen": "2026-08-26", + "id": "f-cce03106df85", + "p": 3.0, + "probe": "secrets-tree", + "remedy": "rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not", + "runs_open": 0, + "severity": "critical", + "title": "A private key block is committed in the tree", + "where": "backend/tests/unit/services/test_vendor_credential_service.py:139" + }, + { + "blast": 3, + "detail": "", + "effort": 1, + "evidence": "", + "first_seen": "2026-08-26", + "id": "f-5d2bca4dbe9c", + "p": 3.0, + "probe": "secrets-tree", + "remedy": "rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not", + "runs_open": 0, + "severity": "critical", + "title": "A private key block is committed in the tree", + "where": "backend/tests/unit/services/test_vendor_credential_service.py:161" + }, + { + "blast": 3, + "detail": "", + "effort": 1, + "evidence": "", + "first_seen": "2026-08-26", + "id": "f-bb6c1a1fc96f", + "p": 3.0, + "probe": "secrets-tree", + "remedy": "rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not", + "runs_open": 0, + "severity": "critical", + "title": "A private key block is committed in the tree", + "where": "backend/tests/unit/test_connection_config_secrecy.py:19" + }, + { + "blast": 3, + "detail": "", + "effort": 1, + "evidence": "", + "first_seen": "2026-08-26", + "id": "f-8bfec0c17d8e", + "p": 3.0, + "probe": "secrets-tree", + "remedy": "rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not", + "runs_open": 0, + "severity": "critical", + "title": "A private key block is committed in the tree", + "where": "backend/tests/unit/test_connection_lifecycle.py:90" + }, + { + "blast": 3, + "detail": "", + "effort": 1, + "evidence": "", + "first_seen": "2026-08-26", + "id": "f-260ea1b3052f", + "p": 3.0, + "probe": "secrets-tree", + "remedy": "rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not", + "runs_open": 0, + "severity": "critical", + "title": "A private key block is committed in the tree", + "where": "backend/tests/unit/test_ssh_key_routes.py:60" + }, + { + "blast": 3, + "detail": "", + "effort": 1, + "evidence": "", + "first_seen": "2026-08-26", + "id": "f-a232815713d1", + "p": 3.0, + "probe": "secrets-tree", + "remedy": "rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not", + "runs_open": 0, + "severity": "critical", + "title": "A private key block is committed in the tree", + "where": "backend/tests/unit/test_ssh_tunnel.py:326" + }, + { + "blast": 3, + "detail": "", + "effort": 1, + "evidence": "", + "first_seen": "2026-08-26", + "id": "f-5b63f326702b", + "p": 3.0, + "probe": "secrets-tree", + "remedy": "rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not", + "runs_open": 0, + "severity": "critical", + "title": "A private key block is committed in the tree", + "where": "frontend/src/__tests__/components/GA4ConnectionForm.test.tsx:288" + }, + { + "blast": 3, + "detail": "", + "effort": 1, + "evidence": "", + "first_seen": "2026-08-26", + "id": "f-5611e0b62c8e", + "p": 3.0, + "probe": "secrets-tree", + "remedy": "rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not", + "runs_open": 0, + "severity": "critical", + "title": "A private key block is committed in the tree", + "where": "frontend/src/__tests__/components/VendorCredentialsPanel.test.tsx:122" + }, + { + "blast": 3, + "detail": "", + "effort": 1, + "evidence": "", + "first_seen": "2026-08-26", + "id": "f-e14137c9263d", + "p": 3.0, + "probe": "secrets-tree", + "remedy": "rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not", + "runs_open": 0, + "severity": "critical", + "title": "A private key block is committed in the tree", + "where": "frontend/src/components/ssh/SshKeyManager.tsx:109" + }, + { + "blast": 3, + "detail": "", + "effort": 1, + "evidence": "", + "first_seen": "2026-08-26", + "id": "f-60caab69e785", + "p": 3.0, + "probe": "secrets-tree", + "remedy": "rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not", + "runs_open": 0, + "severity": "critical", + "title": "A private key block is committed in the tree", + "where": "frontend/src/components/ssh/SshKeyManager.tsx:258" + }, + { + "blast": 3, + "detail": "", + "effort": 1, + "evidence": "", + "first_seen": "2026-08-26", + "id": "f-332dd7570e06", + "p": 3.0, + "probe": "secrets-tree", + "remedy": "rotate the credential at its issuer, then remove it from the tree; if it is in history, rotation is the fix and rewriting history is not", + "runs_open": 0, + "severity": "critical", + "title": "A private key block is committed in the tree", + "where": "frontend/src/lib/api/vendor-credentials.ts:52" + }, + { + "blast": 3, + "detail": "", + "effort": 2, + "evidence": "", + "first_seen": "2026-08-26", + "id": "f-5491ab525a11", + "p": 1.5, + "probe": "secrets-history", + "remedy": "rotate the credential at its issuer", + "runs_open": 0, + "severity": "critical", + "title": "A private key block appears in git history", + "where": "commit c273fa761381" + }, + { + "blast": 3, + "detail": "", + "effort": 2, + "evidence": "", + "first_seen": "2026-08-26", + "id": "f-c7e9cce3d6cd", + "p": 1.5, + "probe": "secrets-history", + "remedy": "rotate the credential at its issuer", + "runs_open": 0, + "severity": "critical", + "title": "A private key block appears in git history", + "where": "commit 5acb82d1fd17" + }, + { + "blast": 3, + "detail": "", + "effort": 2, + "evidence": "", + "first_seen": "2026-08-26", + "id": "f-6188a3f57810", + "p": 1.5, + "probe": "secrets-history", + "remedy": "rotate the credential at its issuer", + "runs_open": 0, + "severity": "critical", + "title": "A private key block appears in git history", + "where": "commit d6388423fb55" + }, + { + "blast": 3, + "detail": "", + "effort": 2, + "evidence": "", + "first_seen": "2026-08-26", + "id": "f-c4be3b834dd6", + "p": 1.5, + "probe": "secrets-history", + "remedy": "rotate the credential at its issuer", + "runs_open": 0, + "severity": "critical", + "title": "A private key block appears in git history", + "where": "commit 4fac5a7d44ad" + }, + { + "blast": 3, + "detail": "", + "effort": 2, + "evidence": "", + "first_seen": "2026-08-26", + "id": "f-a7627092d471", + "p": 1.5, + "probe": "secrets-history", + "remedy": "rotate the credential at its issuer", + "runs_open": 0, + "severity": "critical", + "title": "A private key block appears in git history", + "where": "commit 3e34591e9045" + }, + { + "blast": 3, + "detail": "", + "effort": 2, + "evidence": "", + "first_seen": "2026-08-26", + "id": "f-eb0c4ec18199", + "p": 1.5, + "probe": "secrets-history", + "remedy": "rotate the credential at its issuer", + "runs_open": 0, + "severity": "critical", + "title": "A private key block appears in git history", + "where": "commit d3f7dc9cd8c0" + }, + { + "blast": 3, + "detail": "", + "effort": 2, + "evidence": "", + "first_seen": "2026-08-26", + "id": "f-b6fe551371a5", + "p": 1.5, + "probe": "secrets-history", + "remedy": "rotate the credential at its issuer", + "runs_open": 0, + "severity": "critical", + "title": "A private key block appears in git history", + "where": "commit c9c40974bea5" + }, + { + "blast": 3, + "detail": "", + "effort": 2, + "evidence": "", + "first_seen": "2026-08-26", + "id": "f-11c39415a44e", + "p": 1.5, + "probe": "secrets-history", + "remedy": "rotate the credential at its issuer", + "runs_open": 0, + "severity": "critical", + "title": "A private key block appears in git history", + "where": "commit fbf811253b42" + }, + { + "blast": 2, + "detail": "Deploy targets declared (compose, procfile) with no telemetry dependency in any manifest. A failure on a user's machine is invisible.", + "effort": 2, + "evidence": "languages= deploy=compose,procfile telemetry=[]", + "first_seen": "2026-08-26", + "id": "f-a825549fc486", + "p": 1.0, + "probe": "telemetry", + "remedy": "add an error reporter, or record the decision not to — the gap worth closing is that nobody wrote down which it is", + "runs_open": 0, + "severity": "medium", + "title": "A deployed surface reports no errors anywhere its maintainer can see", + "where": "manifests" + } + ], + "generated_at": "2026-08-26T16:33:49Z", + "probes": [ + { + "id": "secrets-tree", + "needs": [ + "git" + ], + "phase": "probe", + "reason": "20 credential pattern(s)", + "verdict": "finding" + }, + { + "id": "secrets-history", + "needs": [ + "git" + ], + "phase": "probe", + "reason": "8 in history", + "verdict": "finding" + }, + { + "id": "worktree", + "needs": [ + "git" + ], + "phase": "probe", + "reason": "working tree clean", + "verdict": "clean" + }, + { + "id": "telemetry", + "needs": [], + "phase": "prod", + "reason": "no error reporting found", + "verdict": "finding" + }, + { + "id": "ci-present", + "needs": [], + "phase": "prod", + "reason": "CI configured: github-actions", + "verdict": "clean" + }, + { + "id": "docs-present", + "needs": [], + "phase": "seams", + "reason": "7 documentation marker(s): README.md, docs, CLAUDE.md, CONTRIBUTING.md, CHANGELOG.md, docs/adr, ARCHITECTURE.md", + "verdict": "clean" + }, + { + "id": "gitignore-secrets", + "needs": [ + "git" + ], + "phase": "probe", + "reason": "no credential-shaped file is tracked", + "verdict": "clean" + }, + { + "id": "channel-divergence", + "needs": [ + "git" + ], + "phase": "prod", + "reason": "no version in a manifest to make a claim about", + "verdict": "blind" + }, + { + "id": "published-version", + "needs": [ + "npm", + "network" + ], + "phase": "prod", + "reason": "no name+version in a manifest", + "verdict": "blind" + } + ], + "profile": { + "ci": [ + "github-actions" + ], + "deploy": [ + "compose", + "procfile" + ], + "docs": [ + "README.md", + "docs", + "CLAUDE.md", + "CONTRIBUTING.md", + "CHANGELOG.md", + "docs/adr", + "ARCHITECTURE.md" + ], + "languages": [], + "managers": [], + "monorepo": true, + "name": "checkmydata-ai", + "root": "/Users/sshlg/DATA/checkmydata-ai", + "submodules": [], + "telemetry": [], + "tracked_files": 1500, + "vcs": "git", + "version": null, + "workspaces": false + }, + "ratchet": { + "carried": [], + "closed": [], + "first_run": true, + "new": [], + "unranked": [ + "f-036f91903887", + "f-11c39415a44e", + "f-15bc59dcad13", + "f-260ea1b3052f", + "f-332dd7570e06", + "f-466ab3c7e8d3", + "f-4bd961a817ba", + "f-5491ab525a11", + "f-5611e0b62c8e", + "f-5b63f326702b", + "f-5d2bca4dbe9c", + "f-60caab69e785", + "f-6188a3f57810", + "f-8bfec0c17d8e", + "f-91fc925fc4af", + "f-9e004d63832f", + "f-a232815713d1", + "f-a7627092d471", + "f-a825549fc486", + "f-b6fe551371a5", + "f-bb6c1a1fc96f", + "f-c3b1f338966d", + "f-c4be3b834dd6", + "f-c7e9cce3d6cd", + "f-cce03106df85", + "f-cd7a032e015a", + "f-e14137c9263d", + "f-eb0c4ec18199", + "f-ffa99d263ebb" + ] + }, + "root": "/Users/sshlg/DATA/checkmydata-ai", + "schema": "project-audit/1", + "sidecar": "/Users/sshlg/DATA/checkmydata-ai/docs/audit/2026-08-26-audit.json" +} diff --git a/docs/audit/2026-08-26-cold-audit.html b/docs/audit/2026-08-26-cold-audit.html new file mode 100644 index 00000000..951e94ca --- /dev/null +++ b/docs/audit/2026-08-26-cold-audit.html @@ -0,0 +1,314 @@ +Cold Audit — checkmydata-ai + + +
+ +

Cold audit · read-only · nothing written

+

What is actually true of checkmydata-ai

+

2026-08-26 · after the 2026-08-25/26 remediation programme · production +checkmydata-api v274, checkmydata-web v222, both healthy

+ +

01The headline

+

One finding is worth acting on today, and it is the same shape as a defect +this programme fixed yesterday — which is why it is worth naming rather than filing.

+ +
+ HighA-1 +

A public, indexable page sells plans the deployment cannot charge for

+
BILLING_ENABLED     = set          heroku config -a checkmydata-api
+STRIPE_SECRET_KEY   = <UNSET>      heroku config:get STRIPE_SECRET_KEY
+/pricing            → 200          <meta name="robots" content="index, follow">
+/api/billing/plans  → 200          returns the real plan catalogue
+subscriptions       = 0 rows       stripe_events = 0 rows
+

A signed-in visitor who clicks Upgrade reaches POST /api/billing/checkout, + which calls a Stripe API with no key. It degrades honestly — + billing_service.py:45-46 raises BillingError and the route returns + 400, not a crash — but the message the customer sees is + "Stripe is not configured (STRIPE_SECRET_KEY missing)": an operator's + diagnostic, on the money path, on a page search engines are invited to index.

+
Why this is not just a to-do. It is precisely the shape of + RERANKER_ENABLED — a flag advertising a capability the runtime does not carry — + which yesterday's app/ops/capability_report.py was built to catch at every boot. + The gate makes exactly three claims (reranker_enabled, + chroma_embedding_model, chroma_server_url) and this is not one of + them. A mechanism built for a class of defect that does not cover the class's clearest + member is the interesting part.
+

Remedy, in order of cost. Add the claim to CLAIMS — + billing_enabled asserted, stripe_secret_key provided — so every boot + says it. Then either unset BILLING_ENABLED until the keys exist, or give the + checkout route a customer-facing message. Whichever is chosen, the boot line is what stops + it being rediscovered.

+
+ +

02What is written, enabled, and has never run

+

The 2026-08-23 report named six such subsystems. Counted from production's +own empty tables, it is eleven. For all of them a green CI run is the only +evidence there is, and the first real user is the first integration test.

+ +
+ + + + + + + + + + + + + +
SubsystemEvidence (0 rows in production)Reachable by a user?
Billingsubscriptions, stripe_eventsYes — see A-1
GA4 analyticsanalytics_imports, all five ga4_*, vendor_credentialsYes — the [1.16.0] headline feature has never ingested a row
Scheduled queriesscheduled_queries, schedule_runsYes
NotificationsnotificationsYes
Investigationsdata_investigationsAuto-triggered — orchestrator_auto_investigate_enabled defaults on
Learning voteslearning_votesYes
Batch queriesbatch_queriesYes
Semantic layermetric_definitions, metric_relationshipsYes
RAG feedbackrag_feedbackYes
Validation feedbackdata_validation_feedbackYes
Project repositoriesproject_repositoriesThe table is unused entirely — its dead encrypted column was dropped yesterday (F-REPO-04)
+ +
+ MediumA-2 +

A fix shipped yesterday changed a mechanism that has never been used

+

learning_votes is empty, and a5f6888 (#216) corrected how a + vote is weighed against a re-derivation. The reasoning holds and the tests are + real, but no human has ever voted on a learning in this deployment.

+
Not a defect — a calibration point. Effort spent on paths with zero + traffic is effort not spent on the chat path, which carries all 142 messages. Worth + knowing before the next priority call, not worth undoing.
+

Remedy. None in code. Record the exercised/unexercised split beside + the board so priority arguments start from it.

+
+ +

03The mechanical pass found nothing, loudly

+ +
+ InfoA-3 +

28 of 29 findings are false, and the scanner's own profile says why

+
29 findings — 28 critical "private key committed", 1 medium
+profile: "languages": []  "managers": []  "telemetry": []  "version": null
+         "monorepo": true
+

It detected a monorepo and then read manifests only at the root, which has none. + backend/pyproject.toml and frontend/package.json were never opened. + That single gap produced the empty language and manager lists, the false telemetry + finding — Sentry is in both manifests and live in production — and both declared-blind + probes.

+
Three probes were blind silently. Two said blind + and gave a reason. telemetry returned a finding instead, and the two + empty profile fields returned nothing at all. A probe that cannot see and says so is a + question; a probe that cannot see and answers anyway is a wrong answer with a clean + verdict attached.
+
+ +
+ InfoA-4 +

Zero real key material, in the tree or in history — verified, not assumed

+
16 files matched -----BEGIN … PRIVATE KEY-----   (17 of 20 hits in tests/)
+3 hits in source: a <textarea> placeholder ×2 and a `secretHint` string
+strict PEM test (matching END, no code in the span, >16 distinct chars):
+  tree     → 1 candidate    backend/tests/integration/test_security_rbac.py:302
+  history  → 0 candidates   across all 8 flagged commits
+ssh-keygen -y on that candidate → "invalid format"
+

The one survivor is high-entropy filler shaped like a key, used by + test_ssh_key_not_in_response to prove the API never echoes a submitted key + back. It is not a key.

+
My first measurement of this was also wrong, and the way it + was wrong is worth keeping. A lazy (.*?)-----END spanned several unrelated tests + and accumulated 836 characters of code as "base64", which read as a second real key. The + strict form — matching END of the same kind, rejecting spans containing code — + removed it. A scanner and a verifier can fail identically.
+
+ +

04What yesterday's changes actually did in production

+

Each of these was a claim when it shipped. These are the measurements.

+ +
+ VerifiedN1 · heartbeat +

The daily sync completed end to end, and a 42-minute index survived

+
08-26 00:00 UTC  index_repo    completed  (record_index)   00:00 → 00:02
+08-26 00:02      db_index      completed                   00:02 → 00:27
+08-26 00:27      code_db_sync  completed                   00:27 → 00:28
+08-26 00:00      daily_sync    completed                   00:00 → 00:28
+08-25 22:00      index_repo    completed                   22:00 → 22:42  (42 min)
+
+stale run reaped: 0    R14: 0    R15: 0
+

Baseline from the 08-23 report: failing every day since 08-07, 64 of 70 + failures with error='stale run reaped'. Before the fix any single step over + 300 s was killed while working; a 42-minute run across many steps is the direct + counter-evidence.

+
One reap did happen after the fix, and its cause is named. + A db_index started 00:42 CEST and was reaped at 00:51; Heroku released + v269 at 00:46 CEST — my own deploy restarted the worker mid-run. The process + was genuinely gone, so the reaper was correct. I can tell only because N3 recorded the + step.
+
+ +
+ VerifiedN3 · error catalog +

The reaper's failures now reach the product's own log, with the step named

+
run | db_index   | fatal | stale run reaped (step: fetch_samples) | 08-25 22:51
+run | daily_sync | fatal | stale run reaped (step: db_index)      | 08-25 22:51
+

Before: 3 rows, newest 08-17, against 143 failed runs. The catalog now records what it + exists to record, and the step is what makes a concentration diagnosable.

+
+ +
+ VerifiedN4 · N8 · capability gate +

Live, and its first act was to state a truth the audit had to dig for

+

Three boots, three warnings: the configuration names a 768-d embedding model while + Chroma embeds at 384-d, and the vectors are not comparable. reranker_enabled + dropped out of the warnings because it is now unset. make config-drift → exit 0. + npm audit --audit-level=high → exit 0. Gate at 80% against a real 82%.

+
+ +
+ InfoA-5 +

149 orphaned cluster rows — checked, and benign

+

Disabling CLUSTERING_ENABLED left 149 code_clusters rows written + 08-25 06:53. They are read by the SQL agent — but the read is behind the same flag + (sql_agent.py:189-193), so nothing consumes them, and a re-enable deletes before + inserting (code_graph_service.py:404). Dead bytes, not stale data that is + trusted.

+
Raised because yesterday's change created the doubt and did not resolve it. + A side effect named as "clustering off" is not the same as one measured.
+
+ +

04bFound while trying to merge this audit

+

Both of these surfaced because the audit's own pull request would not +build. Neither was on any checklist.

+ +
+ HighA-6 +

main is not protected, so yesterday's CI fix has no teeth

+
GET /repos/CheckMyData-AI/checkmydata-ai/branches/main/protection
+  → 404  "Branch not protected"
+

#223 made CI run on every pull request, including the stacked ones it used to skip. + That closed the blindness. It did not make the result binding: with no protection + rule, a pull request can be merged with failing checks, or with none at all — which is + how every merge in this programme happened, mine included.

+
The finding is the pair. A gate that runs and cannot block is a gate an + operator believes in. #223's own lesson was that an unmeasured branch and a passing one + look identical on the page; an unenforced check and an enforced one look identical too.
+

Remedy. Require backend-lint-test and + frontend-build on main. It costs one setting and makes the eleven + ratchets this programme added actually load-bearing.

+
+ +
+ MediumA-7 +

GitHub has dispatched no workflow for 16 hours, and this PR gets none

+
PR #233  →  0 check-runs on head 25ea94c6db3d, combined state "pending", 0 statuses
+triggers attempted: opened, reopened, synchronize   →  0 runs each
+actions/permissions  → enabled, allowed_actions: all
+workflow "CI"        → state: active
+last CI run anywhere → 2026-08-26T00:30:49Z   (now 16:46Z)
+

The configuration is not the cause: on.pull_request parses to + None, which is the canonical "all pull requests", and the same file dispatched + correctly for feat/frontend-sentry-dsn hours earlier.

+
Not diagnosed, and stopped deliberately. The one + hypothesis that fits "worked, then stopped" is an Actions minutes or spending limit on the + organisation, and reading that needs the admin:org scope this session does not + have. Five rounds were spent; a sixth would be guessing.
+

Remedy. Check the organisation's Actions billing. If it is not that, + the evidence above is what a support request needs.

+
+ +

05Proposed board rows

+

Nothing was written. These are the rows this audit would file, with the +evidence each already carries.

+ +
+ + + + + + + + +
IdSevRow
A-1🟠 HighA public indexable page sells plans the deployment cannot charge for. BILLING_ENABLED set, STRIPE_SECRET_KEY unset, /pricing 200 with robots: index, follow. Checkout degrades honestly to 400 but shows the customer "Stripe is not configured (STRIPE_SECRET_KEY missing)". Same shape as RERANKER_ENABLED; capability_report.py makes three claims and this is not one. Fix: add the claim, then unset the flag or give the route a customer-facing message.
A-6🟠 Highmain is not protected. #223 made CI run on every PR; nothing makes the result binding, so a PR can merge with failing or absent checks — as every merge in this programme did. A gate that runs and cannot block is a gate an operator believes in. Fix: require backend-lint-test and frontend-build on main.
A-7🟡 MediumNo workflow dispatched for 16 hours. PR #233 gets 0 check-runs across opened, reopened and synchronize; Actions is enabled and the workflow active; the same file dispatched correctly hours earlier. Undiagnosed — the fitting hypothesis is an org Actions limit, which needs admin:org to read. Fix: check org Actions billing.
A-2⚪ InfoEleven subsystems have never run on live data, not six. Counted from production's empty tables. For each, a green CI run is the only evidence, and the first real user is the first integration test. Fix: record the exercised/unexercised split so priority arguments start from it; nothing in code.
A-3⚪ InfoThe mechanical audit is blind on this monorepo. monorepo: true yet manifests read only at the root → empty languages/managers/telemetry, a false telemetry finding, two blind probes. Three probes were blind silently. Fix: belongs upstream in the audit script, not here — recorded so the next run's numbers are not read as measurements.
A-5⚪ Info149 orphaned code_clusters rows. Verified benign: the read is behind the same flag and a re-enable deletes before inserting. Fix: none required; delete on flag-off if tidiness is wanted.
+ +

06What this audit did not look at

+

Naming these is the point of the section; a report that lists only what it +found reads as coverage it does not have.

+ + + + +
diff --git a/docs/qa-audit/issues.md b/docs/qa-audit/issues.md index 70731e52..098247ba 100644 --- a/docs/qa-audit/issues.md +++ b/docs/qa-audit/issues.md @@ -81,15 +81,23 @@ frontend **A** (563 smells, 7 SOLID). ## 1. Open severity tally +> The 2026-08-25/26 remediation programme worked from +> [`docs/reports/full-analysis-2026-08-23.html`](../reports/full-analysis-2026-08-23.html) — +> committed so the claims that cite it can be checked rather than taken. Thirteen of its +> findings are closed; four defects it did **not** contain were found while closing them +> (`CB-CI1`, `CB-COV1`, `CB-SEN1`, `CB-SEN2`), and two of its numbers were themselves +> measurement artefacts — see `CB-COV1`. + + | Severity | Open | |---|---| | 🔴 Critical | 0 | | 🟠 High | **0** | | 🟡 Medium | 0 | | 🟢 Low | 25 | -| ⚪ Info | 11 | +| ⚪ Info | 13 | -*Counted 2026-08-25, not estimated: **33 open `F-` rows and 76 struck** by `grep -cE '^\| F-'` / `grep -cE '^\| ~~F-'` over this file, plus **3 open `CB-` rows** those two commands do not see. The severity table above counts all 36 open rows of both kinds, which is why it does not match the `F-` figure — the two measure different sets and each says which. Both are derived from the rows themselves.* +*Counted 2026-08-26, not estimated: **33 open `F-` rows and 76 struck** by `grep -cE '^\| F-'` / `grep -cE '^\| ~~F-'` over this file, plus **5 open `CB-` rows** those two commands do not see. The severity table above counts all 38 open rows of both kinds, which is why it does not match the `F-` figure — the two measure different sets and each says which. Both are derived from the rows themselves.* *(R1+R2 closed 4 High + 8 Medium + 3 Low. R3 (`fbf8112`) closed 2 High (F-SSH-08, F-RULE-01) + 5 Medium (F-RULE-05, F-DG-07/09, F-GRAPH-01, F-LEARN-07) + 1 Low (F-SSH-06). The 2026-07-19 UX @@ -388,6 +396,12 @@ maintainability / reliability risks. | CB-M4 | 🟢 | **God-files / long functions** (backend grade B, 1,934 smells): `agents/orchestrator.py` (2,525 LOC), `agents/sql_agent.py` (2,017), `knowledge/pipeline_runner.py` (1,769), `api/routes/chat.py` (1,724), `services/agent_learning_service.py` (1,259), `main.py` (1,248), `api/routes/connections.py` (1,199). High blast-radius, hard to test in isolation. | Decompose by responsibility (extract per-stage / per-concern modules); fold into the ongoing dashboard-rebuild altitude work. | | CB-M5 | ⚪ | **Silent error swallowing** — counted, not quoted: the live figures are in `tests/unit/docs/test_suppression_debt_ratchet.py`, which fails when they grow (2026-08-25: 53 `except …: pass`, 611 `except Exception`; the 516 written here before had drifted 18%). Some benign (cache writes), but blanket swallowing in request paths (`chat.py`, `health_monitor.py`) hides real failures (vision invariant #5, honest degradation). Overlaps **F-CHAT-05**. | Narrow exception types + log at appropriate level; reserve bare `pass` for provably-benign cleanup. | | CB-L1 | ⚪ | **Suppression debt** — counted, not quoted: ceilings live in `tests/unit/docs/test_suppression_debt_ratchet.py` and going up fails the suite (2026-08-25: 49 `# type: ignore`, 128 `# noqa`; the 47/104 written here had drifted). (0 TODO/FIXME/HACK markers — clean.) | Revisit and remove suppressions where feasible; track the rest. | +| ~~CB-CI1~~ | ✅ | ~~**CI did not run on a stacked pull request**~~ — `ci.yml` triggered on `pull_request: branches: [main]`, so a PR based on another feature branch ran **zero** checks over its whole life, and its page showed no red because nothing had run. That is how three raw git conflict markers reached `proj/leave-and-caps` and survived review. **Fixed 2026-08-26** (#223): the base-branch filter is gone; `push` stays scoped to `main`. `deploy.yml` tightened in the same change — its `workflow_run` filter matches the HEAD branch, so a PR opened *from* `main` could otherwise have reached the deploy job; it now also requires `workflow_run.event == 'push'`. Six tests over both workflow files, two verified red first (`tests/unit/docs/test_ci_covers_every_pull_request.py`). | +| ~~CB-COV1~~ | ✅ | ~~**Coverage stopped tracing at the first database `await`**~~ — `[tool.coverage.run]` carried no `concurrency`, so coverage used its default `thread`. SQLAlchemy's async layer switches through a greenlet, which coverage does not follow unless told, so every statement after the first `await` against the database **inside the same frame** executed and was reported as uncovered — most of the body of most routes. Measured on `tests/integration` (673 tests, identical run twice): `chat.py` **35% → 46%**, 88 statements on one file. Project-wide, same 41 023 statements: **79% → 82%**, 1 065 statements. Two decisions had already been taken on the wrong number — an audit finding of "chat.py 35%", and a `fail_under` gate ten points below reality. **Fixed 2026-08-26** (#228, #229): `concurrency = ["greenlet", "thread"]`, gate raised 72 → 80, and the threshold — which lived in four places — is now bound by `tests/unit/docs/test_coverage_gate_is_stated_once.py`. | +| ~~CB-SEN1~~ | ✅ | ~~**Sentry was reachable by two secrets neither scrubbing layer could see**~~ — the built-in `EventScrubber` matches key names and its 33-key default carries neither `dsn` nor `database_url`; `before_send` matched values but walked only `exception.values`, `logentry` and `breadcrumbs`. A key of either name in `extra` or `contexts` was caught by **neither**. Urgent rather than theoretical from the moment `SENTRY_DSN` was set in production. **Fixed 2026-08-26** (#230): layer 1 wired with the denylist extended 33 → 39, layer 2 walks `extra` and `contexts` recursively and depth-bounded, host preserved. Fifteen tests, one of them an assertion about *Sentry* — that layer 1 alone still leaks values — so the redundancy question re-opens from a red test rather than from memory. | +| ~~CB-SEN2~~ | ✅ | ~~**The Sentry release would have been blank on the container stack**~~ — `HEROKU_SLUG_COMMIT`, the value every guide names, is populated only for slug (buildpack) deploys. This app is on the **container** stack, where the variable exists and is always **empty** (measured on v271 *after* `runtime-dyno-metadata` was enabled). Issues would attach to a release with no commits and suspect-commit attribution would silently do nothing. Enabling the labs feature was necessary and not sufficient, and nothing would have said so. **Fixed 2026-08-26** (#231): the commit is baked into the image via `--build-arg GIT_SHA` → `ENV RELEASE`; verified in production, `RELEASE == main` HEAD. The empty string is the trap — `os.getenv` returns `""` there, not `None`, so an `is None` check would have accepted it; a test catches that form. | +| CB-UX1 | ⚪ | **102 UX scenarios carry a verification older than 30 days.** 110 of 127 were dated 2026-07-19 while 152 commits had landed since; five were re-audited 2026-08-26 and the ceiling now stands at 105, of which 102 still have a changed Coverage file under them. Ordered and computable: `python3 scripts/ux_verification_status.py --backlog 2026-07-19`. The ceiling in `tests/unit/docs/test_ux_scenarios.py` may fall but not rise. | Re-audit in batches, worst first; date each verdict and add an `SCN-NNN` anchor so a machine can check it (21 of 127 have one). | +| CB-OPS1 | ⚪ | **Worker peak memory is unverified at full load.** The Standard-1X → Standard-2X resize removed R14/R15 (170/2 in a 6.5 h window before; **0/0** since, including the 2026-08-26 02:00 CEST daily sync which completed all four runs). But the 1 143 MB peak came from `graph_build` / `generate_docs` over 25 421 symbols, and no run since has rebuilt the graph — the index has been incremental with no qualifying changes, and `clustering_enabled` is now off. The quota is 1 024 MB, so a full rebuild is still the open question. | Force one full re-index (`force_full`) on a quiet window and watch for `mem=` lines; the absence of any is the evidence, since Heroku emits them only over quota. | **Verified-good in the codebase audit (no issue):** SQL identifier quoting (`connectors/base.py:262` doubles quotes correctly), credential exposure (`ConnectionResponse` returns no secrets; Fernet at diff --git a/docs/reports/full-analysis-2026-08-23.html b/docs/reports/full-analysis-2026-08-23.html new file mode 100644 index 00000000..3814d14a --- /dev/null +++ b/docs/reports/full-analysis-2026-08-23.html @@ -0,0 +1,878 @@ +CheckMyData Full Analysis + + +
+ +
+
Инженерный аудит · измерено, не оценено
+

CheckMyData.ai — полный разбор состояния

+

Что готово, что не готово, что недоделано и что ломается в проде. Каждое утверждение + ниже несёт команду, файл со строкой или запрос, которым оно получено. Там, где измерить не удалось, + так и написано — вместо оценки.

+
+ 23 августа 2026 + main @ a5f6888 + prod v259 · 21.08 09:03 + версия в файлах 1.16.0 + последний тег v1.15.1 · 09.07 +
+
+ +
+

Короткий вывод в одном абзаце

+

Код в очень хорошем состоянии; развёрнутая система — нет. 6 741 бэкенд-тест зелёный, + покрытие 79% против гейта 72%, на доске нет ни одной находки выше Low, дизайн-система соблюдается + почти идеально, секретов в репозитории нет. При этом в проде: индексация репозитория падает 85% + прогонов подряд шестнадцать дней из-за отсутствующего heartbeat, восемь флагов включены + вопреки коду и документации (один из них назван инвариантом vision), трекинга ошибок нет + вообще — Sentry не сконфигурирован, а собственный журнал ошибок содержит три записи и не поймал + ни одного из 143 падений. Продукт задеплоен, но не эксплуатируется: 8 пользователей, 2 проекта, + 0 подписок, 0 GA4-подключений. Разрыв не в качестве кода, а в наблюдаемости и конфигурации прода.

+
+ +
+
    +
  1. Как это измерялось
  2. +
  3. Цифры
  4. +
  5. Матрица готовности по подсистемам
  6. +
  7. Прод: реальное состояние
  8. +
  9. Новые находки этого разбора
  10. +
  11. Что не доделано
  12. +
  13. Существующая доска находок
  14. +
  15. Документация и UX: честность утверждений
  16. +
  17. Безопасность и зависимости
  18. +
  19. Ловушки измерения, в которые я попал
  20. +
  21. Что делать, по порядку
  22. +
+
+ +

01Как это измерялось

+

Одиннадцать срезов, все — на живых артефактах: рабочее дерево, прод-приложение Heroku, +прод-Postgres (только SELECT), OSV API, OpenAPI-схема живого сервиса. Ни одно число ниже не +перенесено из документации — все пересчитаны.

+ +
+ + + + + + + + + + + + +
СрезИсточник истиныКак получено
Тесты и покрытиеполный прогон, не CI-кэшpytest tests/ --cov=app · npx vitest run
Гейты фронтендарабочее деревоtsc --noEmit · eslint --max-warnings=0
Прод-логи1 500 строк, окно 07:42–11:27 UTCheroku logs -n 1500
Прод-состояние66 таблиц Postgresheroku pg:psql (только SELECT)
Прод-конфигурация68 переменных окруженияheroku config:get (значения скрыты)
Прод-образone-off dynoheroku run python -c "find_spec(...)"
API-поверхностьOpenAPI живого сервисаcurl /openapi.json → 209 операций
УязвимостиOSV.dev, 180 пакетовPOST api.osv.dev/v1/querybatch
Дефолты флаговconfig.py против CLAUDE.mdregex-сверка 20 ключей
Доска и сценариистроки файлов, не сводкиPython-разбор таблиц Markdown
Секретыиндекс gitgit grep по 5 шаблонам
+ +

02Цифры

+

Слева — то, что измерено сегодня. Где документация утверждает другое, это отмечено.

+ +
+
6 741бэкенд-тестов прошло
0 упало · 4 skip · 1 xfail · exit 0
11 мин 40 с
+
79%покрытие бэкенда
40 963 инструкции, 8 445 непокрыто
гейт CI — 72%
+
707 / 709фронтенд-тестов
2 падают в полном прогоне,
проходят по одиночке — флак
+
0TODO / FIXME / HACK
во всём backend/app и frontend/src
+
+ +
+
99 464строк Python в 381 модуле
+ 118 678 строк тестов в 515 файлах
+
51 331строк TS/TSX в 312 файлах
94 тест-файла
+
209HTTP-операций на 172 путях
по OpenAPI живого прода
+
799коммитов · 146 за 30 дней
71 смерженный PR
+
+ +
+
143упавших фоновых прогона в проде
из них 70 — индексация репозитория
+
3записи в журнале ошибок продукта
Sentry не сконфигурирован
+
8флагов включены в проде
вопреки дефолтам кода
+
0открытых находок
Critical / High / Medium на доске
+
+ +
+

Расхождение с документацией, которое стоит поправить

+

CLAUDE.md утверждает «7 009 тестов — 6 326 бэкенд + 683 фронтенд по 93 файлам, + покрытие 78%». Измерено сегодня: 6 746 бэкенд (6 741 + 4 skip + 1 xfail) и 709 фронтенд + по 94 файлам, покрытие 79%. Цифра помечена «measured 2026-08-20» и просто устарела на три + дня — но она подаётся как текущая, а не как снимок.

+
+ +

03Матрица готовности по подсистемам

+

«Готово» здесь означает: код есть, тесты зелёные, и в проде есть данные, +доказывающие, что подсистема работала. «Не проверено в проде» — код и тесты есть, а прод-данных ноль; +это не обвинение, а точное описание того, чего мы про подсистему не знаем.

+ +
+ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
ПодсистемаСтатусДоказательство из прода
Аутентификация (email/пароль, Google, верификация, сброс)готовоusers=8 · audit_logs=52
Мультиарендность — проекты, участники, роли, инвайтыготовоprojects=2 · project_members=6
Чат-агент (REST / SSE / WS, оркестратор, гейты)готовоsessions=22 · messages=116
traces=212 · spans=9 526
Учёт токенов и бюджетыготовоtoken_usage=5 953
Индексация схемы БДготово37 completed / 5 failed
последнее падение 18.07
Синхронизация код↔БДготово29 completed / 0 failed · rows=354
MCP-сервер (смонтированный, персональные токены)готоворезолв токена в логах 11:24 сегодня
Заголовки безопасности CSP / HSTS / XFOготовопроверено curl -I на живом хосте
Восстановление BM25 на диске (F-KNOW-12)готовоrebuilt=1 schema_rebuilt=1 failed=0
Обучение агента (память по подключению)готово80 записей, все с connection_id
Инсайты и trust-скорыготовоinsight_records=125 · trust_scores=125
Индексация репозитория (M1–M6)ломается12 completed / 70 failed
64 из них на шаге graph_build
Код-графчастично25 421 символ, но всего 2 121 ребро
≈0,08 ребра на символ
Кластеризация графачастичноcode_clusters=147 (значит, доходит иногда)
Реранкер (кросс-энкодер)no-opфлаг=true, но sentence_transformers ABSENT
Эмбеддер 768-d (BAAI/bge-base)no-opмолча падает на MiniLM 384-d
Биллинг Stripeне проверено в продеbilling_enabled=true, но
subscriptions=0 · stripe_events=0
Аналитика GA4не проверено в продеcollect_enabled=true, но
vendor_credentials=0 · imports=0 · факты=0
Расписания запросовне проверено в продеscheduled_queries=0
Расследования данных (InvestigationAgent)не проверено в продеdata_investigations=0
Голосование за обученияне проверено в продеlearning_votes=0
Уведомленияне проверено в продеnotifications=0
Наблюдаемость / трекинг ошибокотсутствуетSENTRY_DSN не задан · error_log=3 строки
+ +
+

Что эта таблица на самом деле говорит

+

Шесть подсистем — биллинг, GA4, расписания, расследования, голосование, уведомления — написаны, + покрыты тестами, включены в проде и ни разу не выполнялись на живых данных. Это не баг: продукт + ещё не в коммерческой эксплуатации. Но это значит, что для них «зелёный CI» — единственное + свидетельство, а первый реальный пользователь окажется первым интеграционным тестом. Про эти шесть + честная формулировка — «не проверено», а не «готово».

+
+ +

04Прод: реальное состояние

+ +

Масштаб эксплуатации

+

Приложение checkmydata-api, стек container, регион us, +web + worker по одному Standard-1X дино, Postgres essential-1, Redis mini. +Релиз v259 от 21 августа 09:03 — и это ровно HEAD ветки main +(a5f68880, деплой-workflow success), то есть прод не отстаёт от кода.

+ +
+
8пользователей
+
2проекта · 2 подключения
+
22чат-сессии · 116 сообщений
+
0подписок · 0 событий Stripe
+
+ +

Хроническое падение индексации — главная проблема прода

+

Индексация репозитория падает каждый день с 7 августа, всегда на одном и том же шаге, +всегда с одной и той же причиной:

+ +
+ + + + + + + + + +
ДатаПаденийШаг, на котором умерло
22.081graph_build
21.085generate_docs, graph_build
20.084generate_docs
19.085graph_build
18.085graph_build
17.085graph_build
16.081graph_build
13.08 → 07.0830graph_build (каждый день)
+ +

Итог по всей истории: 64 из 70 падений — на graph_build, +failure_kind='fatal', error='stale run reaped'. Причина установлена и +описана ниже как находка N1 — это не исключение в коде графа, а отсутствующий heartbeat.

+ +

Память воркера: ровно 899 МБ, 150% квоты, без движения

+

За окно 3 ч 45 мин: 611 × Error R14 (Memory quota exceeded), 0 × R15 (SIGKILL), +0 × H12/H13/H10. Значение памяти — плоское mem=899M(150.4%) во всех 611 замерах, +минимум равен максимуму.

+ +
+

Это меняет диагноз, который был у задачи #12

+

Плоские 899 МБ без роста и без единого R15 означают, что 899 МБ — это базовый след воркера в + покое, а не утечка индексации. Воркер превышает квоту, ничего не делая. Фикс размера батча + (EMBEDDING_UPSERT_BATCH_SIZE=8) держится — SIGKILL исчез полностью. Но задача + #12 формулировала блокер как «частота деплоев»: деплои прекратились 21 августа в 09:03, а падение + graph_build случилось и 22-го. Частота деплоев была не при чём. Настоящая + причина — N1.

+
+ +

Чего в логах нет

+

За всё окно: 0 × pipeline_end, 0 × pipeline_start, +0 × code_symbol_embed, 0 × Traceback, 0 × DuplicateIDError. +Судить по этому нельзя: окно покрывает 09:42–13:27 по местному, а суточная синхронизация +запускается в 00:00 и 02:00. Строки skipped=2 в кроне — не баг: оба проекта +просто ждут своего часа. Heroku хранит 1 500 строк, за пределы окна не заглянуть.

+ +
+

Что в проде работает подтверждённо

+

Восстановление BM25 после рестарта (F-KNOW-12) — в логах старта сегодня: + bm25_local_reconcile: rebuilt=1 schema_rebuilt=1 present=0 no_docs=1 failed=0. + Оба индекса пересобрались на web-дино за 13 секунд. Это фикс из этой серии работ, и он держится.

+

Заголовки безопасности — CSP с frame-ancestors 'none' и + object-src 'none', HSTS max-age=31536000; includeSubDomains, + X-Frame-Options: DENY, nosniff, + Referrer-Policy: strict-origin-when-cross-origin. Полный образцовый набор.

+
+ +

05Новые находки этого разбора

+

Тринадцать находок, которых нет на доске. Каждая — с доказательством и с фиксом. +Severity расставлена по достижимости в текущей конфигурации, а не по громкости названия.

+ +
+
N1У run_repo_index нет heartbeat — и 300 секунд превратились в лимит на один шагHIGH
+

Из трёх долгих фоновых функций в worker.py две обёрнуты в + heartbeat(...), а третья — нет. heartbeat_at у строки прогона обновляется + только на входе в шаг и на выходе из него; внутри шага ничто не тикает. Значит любой шаг, + идущий дольше stale_running_heartbeat_timeout_seconds, объявляется мёртвым, пока он + работает. graph_build на 25 421 символе идёт дольше — и его добивают. Каждый день.

+

Это объясняет и кажущееся противоречие: успешные прогоны длились 417 с и 2 568 с и не были + добиты, потому что у них много шагов, и каждый — короче 300 с.

+
Код: backend/app/worker.py:221-233 — обёртки нет; + сравнить с run_db_index :114-128 и run_code_db_sync + :190-201, где она есть · app/services/run_coordinator.py:250-283 — + heartbeat_at = _now() на входе и выходе шага · app/config.py:537 — + stale_running_heartbeat_timeout_seconds: int = 300
+ Прод: 64/70 падений с current_step='graph_build', + error='stale run reaped', ежедневно с 07.08
+
Фикс: обернуть run_repo_index в + app.core.heartbeat.heartbeat с тикером на строку indexing_runs, ровно как + сделано для двух других. Тест, который должен упасть до фикса: шаг-заглушка длиннее таймаута не + должен быть добит reaper'ом.
+
+ +
+
N2Прод включает восемь флагов, которые код и документация объявляют выключеннымиHIGH
+

Развёрнутая конфигурация существенно отличается от задокументированной. Особенно важен последний + ряд: CLAUDE.md:274 прямо называет его инвариантом vision, а vision.md:73 + формулирует так — «знание об одной базе никогда не протекает в запросы к другой».

+
+ + + + + + + + + + +
ФлагПродДефолт кодаЧем это грозит
CLUSTERING_ENABLEDtrueFalse+1 CPU-тяжёлый шаг в манифест
GIT_POLL_ENABLEDtrueFalseиз семьи «ingestion automation», по правилу дома — off
GIT_WEBHOOK_ENABLEDtrueFalseто же
AUTO_SYNC_AFTER_INDEXtrueFalseто же
FRESHNESS_RECONCILER_ENABLEDtrueFalseто же
SCHEMA_CHANGE_ALERTS_ENABLEDtrueFalseто же
ANALYTICS_COLLECT_ENABLEDtrueFalseпо расписанию зовёт сторонние API
DATA_GATE_LLM_SEMANTICStrueFalseлишний LLM-вызов на каждый гейт
CROSS_CONNECTION_LEARNINGS_ENABLEDtrueFalseснимает защиту, названную инвариантом vision
+

В данных инвариант пока не нарушен: все 80 строк agent_learnings имеют + непустой connection_id, глобальных нет. Выключена именно защита, а не соблюдение.

+
Проверка: heroku config:get <KEY> против + grep -m1 '^\s*<key>:' backend/app/config.py по 20 ключам · + SELECT (connection_id IS NULL), count(*) FROM agent_learnings GROUP BY 1f | 80
+
Фикс: решить по каждому флагу отдельно — либо привести прод к дефолту, либо + изменить дефолт в коде и объяснить в CLAUDE.md, почему прежнее правило больше не + действует. Третий вариант — оставить как есть — допустим только для тех, где расхождение осознанно + и записано. CROSS_CONNECTION_LEARNINGS_ENABLED под это исключение не попадает: пока + vision.md называет его инвариантом, прод не должен его снимать молча.
+
+ +
+
N3В проде нет трекинга ошибок — ни внешнего, ни внутреннегоMEDIUM
+

SENTRY_DSN отсутствует в списке 68 переменных окружения, то есть Sentry не + инициализируется вовсе — при том что sentry-sdk[fastapi] в зависимостях, + @sentry/nextjs в package.json, а код инициализации со скрабингом PII + написан и лежит в app/core/sentry.py. Готовая наблюдаемость просто не подключена.

+

Собственный журнал продукта error_log содержит три записи, самая свежая от + 17 августа. Ни одно из 143 падений фоновых прогонов в него не попало. То есть самая частая + производственная ошибка системы не видна ни в одном из двух предназначенных для этого каналов.

+
Проверка: heroku config — ключа SENTRY_DSN нет · + SELECT count(*) FROM error_log3 · + SELECT status, count(*) FROM indexing_runs GROUP BY status → failed 143
+ Ирония в данных: одна из трёх записей — «Stale: pipeline_end never received», + 12 повторов, статус open. Журнал поймал проблему чата, но не индексации.
+
Фикс: (1) задать SENTRY_DSN — код уже готов; (2) писать в + error_log из пути reaper'а: строка, добитая как fatal, — это ровно то + событие, для которого журнал существует.
+
+ +
+
N4Конфигурация утверждает две возможности, которых в образе нетMEDIUM
+

RERANKER_ENABLED=true в проде, а в образе нет ни + sentence_transformers, ни torch — проверено запуском внутри прода. + Реранкер честно деградирует в NoopReranker и один раз пишет в лог, так что аварии нет. + Но оператор, читающий heroku config, видит включённую фичу, которой не существует.

+

Та же форма у эмбеддера: дефолт chroma_embedding_model — + BAAI/bge-base-en-v1.5 (768-d), в проде переменная не задана, библиотеки нет, и Chroma + молча берёт свой ONNX MiniLM (384-d). Предупреждение при импорте воспроизводится дословно.

+
Прод: heroku run python -c "find_spec(...)" → + sentence_transformers → ABSENT, torch → ABSENT, + onnxruntime → PRESENT
+ Почему: Dockerfile.backend:27pip install ".[redis]", а + sentence-transformers живёт в extra ml (pyproject.toml:59-62)
+ Дословно из лога импорта: Embedding model BAAI/bge-base-en-v1.5 requires the optional + 'sentence-transformers' package, which is not installed; using ChromaDB's built-in default + (all-MiniLM-L6-v2, 384-dim)
+
Фикс: либо снять RERANKER_ENABLED с прода (честнее — сейчас он + ничего не даёт), либо поставить extra ml вместе с полным переиндексом, потому что + векторы 384-d и 768-d несопоставимы. Отдельно стоит сделать так, чтобы флаг, включённый без + зависимости, отказывался стартовать или писал предупреждение уровня WARNING при + каждом запуске, а не один раз.
+
+ +
+
N5В файле доски закоммичены неразрешённые конфликт-маркеры gitMEDIUM
+

docs/qa-audit/issues.md на ветке proj/leave-and-caps — той, что несёт + открытый PR #218 — содержит три сырых маркера конфликта на строках 91, 97 и 100. Абзац с итогом + доски продублирован четыре раза с противоречащими числами. На main маркеров + нет, но дубль абзаца уже смержен — там их два.

+

Отдельно неприятно то, почему это прошло: doc-ratchet проверяет, что фраза «N open rows + and M struck» в файле присутствует, а не что она присутствует однажды. Проверка, + которая ловит отсутствие, но не ловит дублирование, — и есть тот случай, когда зелёный гейт хуже + отсутствующего.

+
Проверка: git grep -nE '^(<<<<<<< |=======$|>>>>>>> )' + → docs/qa-audit/issues.md:91,97,100 · вхождений 'not estimated:': ветка + 4, main 2, должно быть 1
+
Фикс: вычистить маркеры и оставить один абзац с числами, выведенными из + строк; добавить в ratchet два утверждения — «маркеров конфликта нет ни в одном отслеживаемом + файле» и «абзац итога встречается ровно один раз». Первое стоит сделать общим для всего + репозитория, а не только для доски.
+
+ +
+
N6Завершённая индексация репозитория навсегда показывает «шаг 9 из 15»LOW
+

Манифест ставит record_index на позицию 9 и дописывает шесть флаговых шагов на + позиции 10–15. Раннер же выполняет record_index последним — после + bm25_build. В результате все 12 успешных прогонов в проде стоят на + step_index=9, total_steps=15. Процент при этом честный — 100, его выставляет + finish(), — так что пользователь видит «100%, шаг 9 из 15» одновременно.

+
Код: app/knowledge/run_manifests.py:21-30 — порядок манифеста · + app/knowledge/pipeline_runner.py:1779record_index против + :1254bm25_build
+ Прод: SELECT status,current_step,step_index,total_steps,progress_pct,count(*) + → completed | record_index | 9 | 15 | 100 | 12
+
Фикс: привести порядок манифеста к порядку исполнения — флаговые шаги перед + record_index, а не после. Тест: step_position каждого шага должен + монотонно возрастать в том порядке, в котором раннер их вызывает.
+
+ +
+
N7/docs и /openapi.json открыты в проде без аутентификацииLOW
+

Оба отдают HTTP 200 анонимно; схема — 227 139 байт, полный перечень 209 операций и всех моделей. + Для FastAPI это поведение по умолчанию, и многие команды сознательно его оставляют. Находка не в + том, что это опасно, а в том, что это не выглядит решением: ни в + SECURITY.md, ни в docs/DEPLOYMENT.md нет строки, объясняющей выбор.

+
Проверка: + curl -o /dev/null -w '%{http_code}' https://<prod>/openapi.json200, + 227 139 байт · /docs200
+
Фикс: либо закрыть в production через + FastAPI(docs_url=None, openapi_url=None) под флагом, либо записать решение оставить + открытым — с причиной. Молчание здесь и есть дефект.
+
+ +
+
N8chromadb 1.5.9 в проде несёт CRITICAL-уязвимость без доступного фикса — но вектор недостижимLOW
+

OSV даёт для установленной версии GHSA-f4j7-r4q5-qw2c / + PYSEC-2026-311 — «pre-authentication code injection», severity CRITICAL, + и исправленной версии не существует.

+

Почему severity здесь всё-таки Low: CHROMA_SERVER_URL в проде не задан, + CHROMA_PERSIST_DIR=/app/data/chroma, и vector_store.py:154-157 берёт + HttpClient только при заданном URL — иначе embedded PersistentClient. + HTTP-слушателя нет, значит и «pre-auth» вектора нет. Но severity станет Critical в тот момент, + когда кто-то настроит удалённую Chroma, и произойдёт это одной переменной окружения.

+
OSV: 180 пакетов через querybatch; у chromadb==1.5.9 + 2 advisory, fixed-in: NONE · + Конфигурация: heroku config:get CHROMA_SERVER_URL → пусто
+
Фикс: добавить boot-проверку — при заданном + CHROMA_SERVER_URL и уязвимой версии chromadb отказываться стартовать или + писать CRITICAL в лог. Это тот случай, когда защищать надо не текущую конфигурацию, а + переход в следующую.
+
+ +
+
N925 эндпоинтов вообще не упомянуты в API.mdLOW
+

В проде 172 пути и 209 операций. API.md называет 130 путей; 49 не совпадают + дословно, из них 25 не упомянуты никак — ни путём, ни последним сегментом. Среди них целые + группы: пять эндпоинтов обучений по подключению, три управления прогонами + (cancel / retry / GET), три журнала ошибок и отказов + запросов, /api/chat/search, /api/chat/explain-sql, + /api/projects/access-requests.

+
Проверка: OpenAPI живого прода против путей в API.md, + нормализация {param}→{id}; вторым проходом отсеяны те, что упомянуты прозой или + wildcard'ом вроде /api/auth/*
+
Фикс: дописать 25 строк — или, надёжнее, проверять покрытие тестом, + который сравнивает app.openapi() с API.md и падает на новом + недокументированном пути. Документация, синхронность которой доказывается кодом возврата, + не расходится.
+
+ +
+
N10Все 127 UX-сценариев заявлены как «implemented / PASS», но 110 верификаций старше пяти недельLOW
+

В индексе docs/ux/scenarios.md — 127 строк, у всех 127 статус + implemented, у 125 вердикт PASS. Даты проверки: 110 из них — + 19 июля, то есть 35 дней назад. За это время на main легло 152 коммита, включая + изменения пользовательского поведения. Плюс: только 22 из 127 сценариев упоминаются + где-либо в коде или тестах — у остальных 105 нет якоря, по которому реализацию можно проверить + автоматически.

+

Формулировка «100% implemented, 98% PASS» здесь — утверждение, а не измерение.

+
Проверка: разбор таблицы-индекса Python'ом → {'implemented': 127}, + {'PASS': 125, 'other': 2}, даты {'2026-07-19': 110, '2026-08-16': 5, + '2026-08-19': 9, '2026-08-20': 1, '2026-08-21': 2} · перекрёстная сверка + SCN-\d+ по всем .py/.ts/.tsx → 22 из 127
+
Фикс: отделить «реализовано» от «проверено» — второе с датой и с тем, чем + проверено. Дальше — прогонять /ux-audit партиями и ставить якоря + SCN-NNN в тесты, начиная со сценариев, которых коснулись последние 152 коммита.
+
+ +
+
N11Тегирование релизов остановилось; 175 записей висят в [Unreleased]INFO
+

pyproject.toml и package.json согласованно говорят 1.16.0, в + CHANGELOG.md есть датированная секция [1.16.0]но тега + v1.16.0 не существует. Последний тег — v1.15.1 от 9 июля, + с тех пор на main 152 коммита, а в [Unreleased] накопилось + 175 пунктов на 1 176 строк.

+

Причина понятна и сама по себе не порочна: авто-деплой на Heroku по мержу расцепил «деплой» и + «релиз». Но версия в файлах теперь не соответствует ни одному тегу, и «какой код в v259» отвечается + только по SHA.

+
Проверка: git tag --sort=-creatordatev1.15.1 + первый · git rev-parse -q --verify refs/tags/v1.16.0 → пусто · + git rev-list --count v1.15.1..origin/main152 · пунктов + '^- ' в [Unreleased]175
+
Фикс: либо тегировать 1.16.0 на том коммите, который ей соответствует, и + вести теги дальше, либо перейти на схему «версия = SHA деплоя» и убрать номер из файлов. Сейчас + действуют обе схемы одновременно, и ни одна не отвечает на вопрос «что в проде».
+
+ +
+
N12Фронтенд-сьют флакует под нагрузкойINFO
+

В полном прогоне (94 файла) падают 2 теста из 709 — GA4ConnectionForm.test.tsx, + оба по Test timed out in 5000ms. Тот же файл в одиночку: 10 passed. + То есть тест не сломан, он не успевает при параллельной нагрузке. CI зелёный — значит там либо + быстрее, либо повезло, и это ровно тот класс дефекта, который проявляется у кого-то другого.

+
Полный прогон: Test Files 2 failed | 92 passed (94), + Tests 2 failed | 707 passed (709) · + Одиночный: npx vitest run src/__tests__/components/GA4ConnectionForm.test.tsx + → 1 passed (1) · 10 passed (10)
+
Фикс: поднять testTimeout для этого файла или убрать из него + ожидание, зависящее от планировщика. Флак в CI дороже, чем кажется: он учит игнорировать красное.
+
+ +
+
N13Долг подавлений вырос за числа, записанные на доскеINFO
+

Строки CB-M5 и CB-L1 фиксируют конкретные числа. Пересчёт сегодня + даёт больше по всем четырём метрикам — сильнее всего по except Exception, +18%.

+
+ + + + + +
МетрикаНа доскеСегодняΔ
except Exception516611+95
except …: pass5154+3
# type: ignore4749+2
# noqa104128+24
+
Фикс: превратить эти четыре числа в ratchet-тест — «не больше, чем + сегодня». Число в документе, которое никто не пересчитывает, дрейфует; число в тесте — нет.
+
+ +

06Что не доделано

+ +

Мёртвый код, который решает не разработчик

+

F-REPO-04ProjectRepository.auth_token_encrypted: колонка объявлена в +модели, принимается параметром в RepositoryService.create, не передаётся ни одним +вызывающим, отсутствует в ALLOWED_UPDATE_FIELDS и не читается нигде. Это либо +незаконченная HTTPS-аутентификация репозитория, либо балласт. Колонка, предназначенная хранить +секреты, не должна сидеть в схеме без объяснения — но выбор между «дописать» и «снять миграцией» +продуктовый, не мой.

+ +

Находки, которым нужен не патч, а решение

+

F-LEARN-04 и F-LEARN-05 — семантические противоречия между обучениями и +дедупликация перефразировок. Обе требуют эмбеддингов или LLM-вызова там, где сейчас нет ни того, ни +другого. Приклеить их к обычной правке значило бы спрятать стоимость: это выбор архитектуры и бюджета +токенов, а не строчка кода.

+ +

Открытые PR

+
+ + + + + + +
PRЧто несётБазаCIБлокер
#217домен роли: F-PROJ-07/08/11mainSUCCESS ×2готов к мержу
#218выход из проекта + ограниченные списки: F-PROJ-12/13proj/role-domainнет вердиктаN5 — конфликт-маркеры в файле; мержить после #217
+

Стек на #217 сделан осознанно: #218 нужны logger и VALID_ROLES из него. +Порядок — сначала #217, затем #218 перецелится на main автоматически. Но #218 нельзя мержить как +есть: находка N5 — в его собственном диффе.

+ +

Самые непокрытые модули — там же, где болит

+

Покрытие 79% в целом, но распределено неровно, и наименее покрытые модули — это точка входа +продукта и тот самый пайплайн, который падает в проде.

+
+ + + + + + + + + +
МодульПокрытиеНе покрыто строкПочему это важно
app/api/routes/chat.py35%493главная точка входа продукта
app/main.py40%492старт, кроны, lifespan
app/knowledge/pipeline_runner.py53%326падает в проде каждый день
app/api/routes/repos.py50%229запуск индексации
app/knowledge/code_db_sync_pipeline.py56%208код↔БД
app/api/routes/chat_utility.py31%192explain-sql, estimate, summarize
app/connectors/ssh_exec.py47%158только что переработан (F-SSH-07)
app/api/routes/feed.py13%135самое низкое покрытие в проекте
+

Файлов с нулевым покрытием — ноль из 380. Это хороший знак: непокрытое здесь — это ветки, +а не забытые модули.

+ +

07Существующая доска находок

+

Числа пересчитаны из строк файла, не взяты из его сводки — и они не совпали с +тем, что я сам сообщал в предыдущем отчёте.

+ +
+
+ origin/main +
+ + + + +
открыто42 (39 F- + 3 CB-)
закрыто70
Critical / High / Medium0 / 0 / 0
Low / Info30 / 12
+
+
+ после мержа #217 + #218 +
+ + + + +
открыто37 (34 F- + 3 CB-)
закрыто75
Critical / High / Medium0 / 0 / 0
Low / Info25 / 12
+
+
+ +
+

Поправка к моему предыдущему отчёту

+

В отчёте от 21 августа я написал «осталось 39 (27 Low, 12 Info)». Правильно: 42 открытых + строки — 30 Low и 12 Info. Ошибка возникла так: grep -oE по emoji в этой локали + отдаёт неверный результат — на 42 строках он вернул «42 🟢» при том, что в файле явно есть строки с + ⚪. Пересчёт Python'ом по позиции ячейки даёт 30/12. Числа выше и в остальном отчёте получены + вторым способом.

+
+ +

Тринадцать находок из этой серии работ оказались неверны как написаны — и каждый раз это +вскрывалось одинаково: тест, написанный чтобы упасть, проходил. Из тринадцати разобранных подробно +восемь пришлось переформулировать инверсиями — «200 вместо 500», «сигнал неразличим» вместо «сигнала +нет», «всегда» вместо «после рестарта», гонка, которой не существует. Это единственная процедура из +той серии, которую стоит закрепить правилом: находка не считается подтверждённой, пока тест, +написанный чтобы упасть, действительно не упал.

+ +

08Документация и UX: честность утверждений

+ +
+
108md-файлов в docs/
43 408 строк
+
18корневых документов
7 795 строк
+
127UX-сценариев
2 156 строк
+
20/20дефолтов флагов в CLAUDE.md
совпали с config.py
+
+ +
+

Что в документации действительно хорошо

+

Сверка двадцати дефолтов флагов из CLAUDE.md с config.py дала + полное совпадение — включая тонкие места вроде reranker_enabled=False (после + исправления 10 августа), clustering_enabled=False, + embedding_upsert_batch_size=8. Документация здесь не дрейфует, и это редкость на + 100 тысячах строк кода.

+

CHANGELOG.md на 3 468 строк, ADR на месте, у каждой находки на доске — id, severity + и путь. Инженерная гигиена документов заметно выше среднего.

+
+ +

Расхождения, которые нашлись, собраны в находках N5 (маркеры конфликта и дубли абзаца), +N9 (25 эндпоинтов), N10 (устаревшие верификации сценариев) и в поправке к числу тестов +выше. Ни одно из них не про неверное описание архитектуры — все про числа и списки, которые никто +не пересчитывает. Это ровно тот класс, который лечится ratchet-тестом, а не вычиткой.

+ +

09Безопасность и зависимости

+ +
+

Прод-образ чище локальной среды разработки

+

Это важно и легко перепутать. Локальный venv несёт 32 advisory в 6 пакетах, из них 17 + исправимы обновлением. Но Dockerfile ставит зависимости по нижним границам, поэтому + прод, собранный 21 августа, получил уже исправленные версии: aiohttp 3.14.3 + (fixed), cryptography 50.0.0 (fixed), GitPython 3.1.59 (fixed). + То есть 17 advisory, которые видит локальный аудит, прода не касаются — их закрыли до сборки. + Устарел локальный venv, а не проект.

+
+ +
+ + + + + + +
ПакетЛокальноВ продеЧто реально висит на проде
aiohttp3.14.13.14.3закрыто 6 advisory исправлены
cryptography49.0.050.0.0закрыто 2 advisory исправлены
GitPython3.1.553.1.59закрыто 9 advisory исправлены
chromadb1.5.91.5.9висит CRITICAL, фикса нет — см. N8
ecdsa0.19.20.19.2висит HIGH, фикса нет — но не на пути авторизации
+ +

Про ecdsa честно: GHSA-wj6h-64fc-37mp — тайминг-атака Minerva на P-256, +исправленной версии нет. Пакет приходит транзитивно через python-jose, но +jwt_algorithm = "HS256" (config.py:148) — алгоритм симметричный, ECDSA +на P-256 в аутентификации не участвует. Severity в этой конфигурации — низкая, и это вывод из +конфигурации, а не из желания.

+ +

Фронтенд: 7 high-уязвимостей

+

npm audit — 0 critical, 7 high, 0 moderate/low. Среди них сам +next, а также sharp (наследует четыре CVE libvips), +postcss (XSS через неэкранированный </style>), +js-yaml, brace-expansion, fast-uri, nanoid. +Часть — только сборочные, но next и sharp — рантайм.

+ +

Секреты и соблюдение дизайн-системы

+
+
Секретов в репозитории нет.
+ Поиск по 5 шаблонам (sk-…, AKIA…, PEM-заголовки, + ghp_…, lin_api_…) дал только placeholder-строки в UI + (SshKeyManager.tsx, vendor-credentials.ts). Отслеживаемых + .env — только два .example. Присваиваний + SECRET=/TOKEN= с высокой энтропией — ноль.
+
Дизайн-система соблюдается почти идеально.
+ Запрещённых сырых палитровых классов Tailwind + (bg-slate-500 и подобных) во всём frontend/src2. + console.log в продакшн-коде — 0. Использований типа any1. + @ts-ignore0; шесть eslint-disable-next-line, все точечные и + обоснованные. tsc --noEmit и eslint --max-warnings=0 — оба чистые.
+
+ +

10Ловушки измерения, в которые я попал

+

Шесть раз за этот разбор моя собственная команда соврала — и каждый раз в сторону +более уверенного вывода. Записываю их, потому что вывод, полученный сломанным измерением, выглядит +точно так же, как правильный.

+ +
+ + + + + + + + + + + + + + + + + + + +
Что я сделалЧто получилКак поймал
npx vitest run … | tail -40VITEST_EXIT=0 при двух упавших тестах — код возврата принадлежал tailувидел 2 failed в тексте, который сам же напечатал
grep -oE '🔴|🟠|🟡|🟢|⚪'«42 🟢» при том, что в файле явно есть ⚪прочитал строки глазами; пересчитал Python'ом по позиции ячейки
uvx pip-audit без аргументов«No known vulnerabilities found» — проверена среда самого uvx, не проект0 пакетов вместо 180; перешёл на OSV API по списку из importlib.metadata
regex по @router.… для путей API«159 недокументированных» — префиксы роутеров не захватывалисьпути вышли как /ask вместо /api/chat/ask; взял OpenAPI живого прода
regex (GET|POST) /api/ по API.md«19 упоминаний из 208» — документ использует табличный форматцифра была слишком плохой, чтобы быть правдой; пересчёт дал 130 из 172
вывод «шаги 10–15 пропущены» из step 9/15серьёзное обвинение пайплайну в молчаливом частичном успехепроверил порядок вызовов в раннере: record_index — последний. Осталась находка + N6, но она Low, а не High
+ +
+

Общая форма всех шести

+

Пять из шести — это сломанное измерение, выдающее правдоподобный результат. Ни одно не + выглядело как ошибка: exit=0, «уязвимостей нет», «42 строки» — всё это нормальные + ответы. Единственное, что их вскрыло, — сверка второго измерения с первым и чтение исходных строк + глазами. Отсюда практическое правило: число, на котором держится вывод, должно быть получено + дважды разными путями — особенно когда первое измерение подтверждает то, что ты и так думал.

+
+ +

11Что делать, по порядку

+

Порядок — по отношению «стоимость исправления к тому, что оно перестанет скрывать», +а не по severity.

+ +
+ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
#ДействиеПочему первымОбъём
1Задать SENTRY_DSN в продекод инициализации со скрабингом PII уже написан. Пока трекинга нет, каждая следующая находка + будет добываться так же вручную, как этаодна переменная
2Обернуть run_repo_index в heartbeat (N1)снимает 85% отказов индексации, которые идут 16 дней. Паттерн уже есть в том же файле для + двух других функций~10 строк + тест
3Вычистить конфликт-маркеры и мержить #217 → #218 (N5)#218 нельзя мержить как есть; #217 зелёный и ждётминуты
4Писать в error_log из пути reaper'а (N3)журнал ошибок продукта не поймал ни одного из 143 падений — он существует ровно для этогонебольшой
5Решить по девяти флагам прода (N2)сначала CROSS_CONNECTION_LEARNINGS_ENABLED: пока vision называет его инвариантом, + прод не должен его снимать молчарешение + запись
6Снять RERANKER_ENABLED или поставить extra ml (N4)конфигурация не должна утверждать возможность, которой нет в образепеременная, либо образ + переиндекс
7Ratchet-тесты на числа: конфликт-маркеры, дубль абзаца, покрытие API.md, + долг подавлений (N5, N9, N13)все четыре расхождения этого разбора — числа, которые никто не пересчитывает. Лечится + тестом, а не вычиткойсредний
8Поднять покрытие chat.py (35%) и + pipeline_runner.py (53%)точка входа продукта и модуль, падающий в проде, — наименее покрытые в проектекрупный
9Разгрузить воркер: 899 МБ при квоте 512 МБ в покоеR15 ушли, но 611 × R14 за 4 часа — это постоянный свап. Resize или вынос индексацииинфраструктурный
10Прогнать /ux-audit по сценариям, которых коснулись + 152 коммита (N10)«100% implemented / PASS» — утверждение пятинедельной давности, а не измерениекрупный
11Решить судьбу auth_token_encrypted (F-REPO-04)колонка под секреты без писателя и читателя. Выбор «дописать или снять» — продуктовыйрешение
12Разобраться с 7 high в npm, начиная с next и sharpединственные два из семи, что попадают в рантаймсредний
+ +
+

+Отчёт собран 23 августа 2026 из рабочего дерева proj/leave-and-caps, +origin/main @ a5f6888 и живого прода checkmydata-api v259.
+Прод-Postgres читался только запросами SELECT; персональные данные не извлекались — +только агрегаты. Значения переменных окружения нигде не печатались; для SENTRY_DSN +проверялось лишь наличие.
+Всё, что не удалось измерить, названо в тексте неизмеренным. Шесть сломанных измерений, +встретившихся по ходу, перечислены в разделе 10 вместе с тем, как каждое было поймано. +

+ +
diff --git a/docs/reports/session-2026-08-21.html b/docs/reports/session-2026-08-21.html new file mode 100644 index 00000000..3e0fa115 --- /dev/null +++ b/docs/reports/session-2026-08-21.html @@ -0,0 +1,195 @@ +Аудит-сессия 2026-08-21 — отчёт + + +
+
+

Аудит-сессия: отчёт по остановке

+

checkmydata-ai · 21 августа 2026 · цикл остановлен по запросу · ветка feat/ledger-redesign → 26 PR в main

+
+ +

За сессию закрыто 75 находок из аудита, смержено 26 PR, два PR открыты и ждут мержа. Критических, высоких и средних находок на доске не осталось. Ниже — что сделано, что нет, и что выяснилось про сам аудит.

+ +
+
75находок закрыто
+
39открыто в main
+
0Critical / High / Medium
+
26PR смержено
+
2PR ждут мержа
+
6740тестов бэкенда
+
709тестов фронта
+
+ +

Главный результат — не количество, а то, что нашлось в самих находках

+ +

Из тринадцати разобранных подробно находок восемь оказались неверны как написаны — и не в мелочах, а инверсиями. Это не упрёк автору доски: так выглядит любой аудит, написанный по чтению кода, а не по его исполнению.

+ +
+ + + + + + + + + + + + +
СтрокаЧто было написаноЧто оказалось
F-PROJ-06«partial-success 500s»500 не бывает вовсе: 200, скрывающая провал — письмо не ушло, ответ успешный
F-VIZ-01«нет признака свежести»Признак есть. Он не различает «никто не открывал» и «обещанное обновление сломано»
F-KNOW-07«добавить метрику промаха»Метрика была — и срабатывала на здоровом пути, поэтому не значила ничего
F-KNOW-12«dense-only после рестарта»Всегда: писатель в воркере, читатель в веб-дино, разные диски
F-SSH-04«db_port не экранирован»Верно про код, но недостижимо — тип закрывает все пути
F-SSH-05«гонка check-then-pin»Гонки нет: между проверкой и записью ноль await
F-GIT-02«наследует риск RCE»Не наследует: проверка стоит в вызываемом, её наследуют все
F-PROJ-07«не валидирует строку роли»Хуже: fail-open — опечатка в требуемой роли открывала эндпоинт всем
+
+ +
Каждый раз это выяснялось одним и тем же способом: тест, написанный чтобы упасть, проходил. Это дешевле любого чтения кода, и это единственная процедура из сессии, которую стоит закрепить как правило.
+ +

Сквозная форма дефекта

+ +

Больше половины закрытых находок — один и тот же дефект в разных местах: сигнал, который не может сказать, какая из двух вещей произошла.

+ +
+ + + + + + + + + + +
НаходкаСигнал существовалДва состояния, которые он сливал
F-PROJ-06_send логировал сбой«отправлено» против «залогировано и всё равно 200»
F-VIZ-01timeAgo(last_executed_at)«никто не открывал» против «обновление сломано»
F-KNOW-07retrieval_degraded_total«лексически не совпало» против «индекса нет на этой машине»
F-DG-08кросс-стадийная проверка«пропорции в норме» против «оба числа обрезаны»
F-SCHED-03предикат жнеца«свежая» против «возраст неизвестен»
F-PROJ-13список участников«вся команда» против «первые 500 из 5000»
+
+ +

Опознаётся по двум приметам: метка причины с единственным значением, и булево там, где у явления три состояния.

+ +

Что осталось: 39 находок в main

+ +

Ни одной Critical, High или Medium. Из 39: 27 Low и 12 Info. Две из них я завёл сам в этой сессии, намеренно не закрывая.

+ +

Заведено мной и оставлено открытым сознательно

+ + +

Требуют проектного решения, а не правки

+ + +

Остальное — обычная работа в порядке приоритета

+

F-LLM-05 (мутация под правами viewer — её независимо подтвердила развёртка по 143 точкам вызова), F-BILL-08 (чарджбэк не отзывает доступ), F-CONN-07 (исчерпание удалённых соединений), F-CHAT-04 (устаревший конфиг в WS-сессии), F-CHAT-08 (fatal-стадия не отменяет соседние), F-SCHED-06 (нет минимального интервала расписаний), F-FE-03 (a11y и токены не проаудированы) и далее.

+ +

Два PR открыты

+
+ + + + + + +
PRЧтоСостояние
#217F-PROJ-07/08/11 — домен роли: UnknownRoleError, AssignableRole, защита upsert от гонкиперебазирован, ждёт CI
#218F-PROJ-12/13 — выход из проекта, ограниченные списки с честной меткой обрезанияв стеке на #217
+
+

#218 стоит на #217 осознанно: ему нужны logger и VALID_ROLES оттуда. Порядок мержа — сначала #217, затем #218 сам перенацелится на main.

+ +

Нерешённая проблема, которая не в коде

+ +
+Задача #12 — репо-индекс в проде не доходит до pipeline_end. +

Формулировка исправлена по измерению, а не по памяти: за реальное окно 3.5 часа Error R15 (фикс размера батча держится, никого не убивают), но 204× R14 и pipeline_end по-прежнему ноль. Стадия code_symbol_embed стартует на 25 270 символах и через 23 минуты воркер перезапускается посреди неё — без убийства и без ошибки. Все пять перезапусков в окне совпали с релизами v240–v244, каждый через 12–14 с после релиза.

+

Блокирует частота моих же деплоев, а не память. Прогону нужно больше времени, чем интервал между мержами, и цикл его не давал. Теперь цикл остановлен — окно появится само. Resize воркера (heroku ps:resize worker=standard-2x) остаётся правильным долгим решением, но гейтом он не является.

+
+ +

Ловушки, которые стоит знать про эту машину и этот процесс

+ + +

Тесты, которые закрепляли дефект

+

Три раза существующий тест защищал находку, а не пропускал её:

+ +
Утверждать слово — не значит утверждать смысл. Утверждать значение — не значит утверждать, что значение верно.
+ +

Проверяемость самого отчёта

+ + + +