-
Notifications
You must be signed in to change notification settings - Fork 6
ci(backend): replace shards with exact-custody semantic lanes #198
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Merged
abiorh-claw
merged 34 commits into
main
from
codex/ws-ci-001-02b-exact-custody-semantic-lanes
Jul 25, 2026
Merged
Changes from all commits
Commits
Show all changes
34 commits
Select commit
Hold shift + click to select a range
161417f
ci(backend): replace shards with exact-custody semantic lanes
Abiorh001 b771c64
chore(agent-loop): declare semantic lanes merge outcome
Abiorh001 a3ad393
fix(ci): close semantic lane custody gaps
Abiorh001 848fa03
fix(ci): stabilize independent node inventory
Abiorh001 cda3f1f
docs(ci): clarify local diagnostic boundary
Abiorh001 18ead0b
fix(ci): confine deterministic UUIDs to collection
Abiorh001 2dfcf06
docs(ci): align hosted proof sequence
Abiorh001 d3b2a65
docs(agent-loop): record semantic lane trust evidence
Abiorh001 cc8ee4f
fix(ci): initialize semantic lane evidence root
Abiorh001 cb63719
docs(agent-loop): record semantic lane bootstrap repair
Abiorh001 239adb1
fix(ci): preserve schema lane coverage custody
Abiorh001 80dd8c8
docs(agent-loop): record coverage custody repair
Abiorh001 4afb6fb
perf(ci): add exact-custody semantic execution units
Abiorh001 ad7270d
fix(ci): bind nodes to semantic resource units
Abiorh001 4661681
test(ci): reject shared semantic unit namespaces
Abiorh001 8bc7e9f
docs(agent-loop): record semantic unit review evidence
Abiorh001 31516a3
Revert "docs(agent-loop): record semantic unit review evidence"
Abiorh001 f134920
Revert "test(ci): reject shared semantic unit namespaces"
Abiorh001 eeab0fb
Revert "fix(ci): bind nodes to semantic resource units"
Abiorh001 f9ddc55
Revert "perf(ci): add exact-custody semantic execution units"
Abiorh001 15e3589
perf(ci): rebalance semantic test lanes
Abiorh001 bf16f1a
test(ci): align semantic lane fixtures
Abiorh001 f5d2abd
docs(agent-loop): record lane rebalance evidence
Abiorh001 24f3b63
fix(ci): preserve semantic lane failure evidence
Abiorh001 cd1e8a8
docs(agent-loop): reconcile external CI review
Abiorh001 06709b0
fix(ci): separate accepted timing target from gates
Abiorh001 a44c516
docs(ci): align accepted timing semantics
Abiorh001 bf3ea27
docs(agent-loop): bind accepted timing review
Abiorh001 2292ed6
Merge remote-tracking branch 'origin/main' into codex/ws-ci-001-02b-e…
Abiorh001 398b25d
Merge remote-tracking branch 'origin/main' into codex/ws-ci-001-02b-e…
Abiorh001 1526ade
Merge remote-tracking branch 'origin/main' into codex/ws-ci-001-02b-e…
Abiorh001 f601266
docs(agent-loop): bind main integration review
Abiorh001 400be48
fix(ci): preserve orchestration failure causes
Abiorh001 fb6a7ff
docs(agent-loop): bind final external repair review
Abiorh001 File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
104 changes: 104 additions & 0 deletions
104
...I-001-backend-ci-acceleration/reviews/WS-CI-001-02B-external-review-response.md
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,104 @@ | ||
| # External Review Response | ||
|
|
||
| ## Chunk | ||
|
|
||
| `WS-CI-001-02B` — Exact-Custody Semantic Test Lanes | ||
|
|
||
| ## Source | ||
|
|
||
| CodeRabbit inline review on PR #198, posted 2026-07-24 against `f5d2abd7`. | ||
|
|
||
| ## Comments addressed | ||
|
|
||
| All fourteen comments were valid and addressed: | ||
|
|
||
| 1. Standardized `parameterized` and `full-service` wording. | ||
| 2. Clarified that the rebalance preserved the final workflow contract and gates. | ||
| 3. Replaced the absolute Ruff claim with the observed resolver-drift failure. | ||
| 4. Synchronized status history with reviewed and hosted node evidence. | ||
| 5. Initially clarified the 480-second failure behavior. After exact-head hosted | ||
| proof exceeded the target while every correctness and coverage gate passed, | ||
| the owner explicitly accepted the measured risk and required the accepted | ||
| performance target not to leave the required check permanently red. | ||
| 6. Mapped `git rev-parse` failures to the stable isolated-runner error contract. | ||
| 7. Preserved `KeyboardInterrupt` and `SystemExit` during MinIO bucket creation. | ||
| 8. Preserved failed lane evidence when isolation metadata was never created. | ||
| 9. Reported pytest collection failure before manifest validation. | ||
| 10. Preserved exactly four failed lane rows after runtime and partial-startup | ||
| orchestration failures. | ||
| 11. Removed inherited `PYTEST_ADDOPTS` and `PYTEST_PLUGINS` from independent | ||
| collection. | ||
| 12. Made UUID restoration unconditional in the validator regression test. | ||
| 13. Documented PostgreSQL as the service container and MinIO as an in-step, | ||
| loopback-published container with a health loop. | ||
| 14. Added an Agent Gate assertion for the step that reads `run.exit` and fails | ||
| the required check. | ||
|
|
||
| The refreshed review on exact head `f601266a` added three findings, all | ||
| addressed: | ||
|
|
||
| 15. Replaced placeholder command arguments with the exact focused pytest nodes | ||
| and Markdown path used for verification. | ||
| 16. Preserved the traceback and exception message for unexpected lane | ||
| orchestration failures while retaining cleanup and four-lane failure | ||
| evidence. | ||
| 17. Allowed `asyncio.CancelledError` and other non-`Exception` cancellation | ||
| signals to propagate during MinIO bucket creation instead of translating | ||
| them into namespace collisions. | ||
|
|
||
| ## Comments deferred | ||
|
|
||
| None. | ||
|
|
||
| ## Human decisions needed | ||
|
|
||
| The repository owner explicitly accepted the measured timing risk from hosted | ||
| runs `30121249272` and `30123755007` and deferred further speed optimization. | ||
| The workflow continues to record whether the 480-second target was met, but an | ||
| accepted performance miss no longer overrides successful test, custody, | ||
| coverage, isolation, and service-contract gates. | ||
|
|
||
| ## Commands rerun | ||
|
|
||
| ```bash | ||
| cd backend | ||
| .venv/bin/python -m pip install ruff==0.15.22 | ||
| .venv/bin/ruff check app tests scripts | ||
| .venv/bin/python -m pytest -q \ | ||
| tests/test_ci_test_lanes.py::test_unexpected_runner_failure_force_kills_and_records_every_lane \ | ||
| tests/test_ci_test_lanes.py::test_partial_startup_failure_records_exactly_four_failed_lanes \ | ||
| tests/test_isolated_database_runner.py::test_minio_creation_preserves_process_interrupts \ | ||
| tests/test_isolated_database_runner.py::test_minio_creation_preserves_async_cancellation \ | ||
| tests/test_isolated_database_runner.py::test_minio_probe_cleans_up_and_preserves_async_cancellation | ||
| cd .. | ||
| python3 scripts/test_agent_gates.py | ||
| python3 scripts/check_internal_review_evidence.py | ||
| python3 scripts/check_markdown_links.py \ | ||
| .agent-loop/initiatives/WS-CI-001-backend-ci-acceleration/reviews/WS-CI-001-02B-external-review-response.md | ||
| git diff --check | ||
| ``` | ||
|
|
||
| Latest focused repair result: exact Ruff passed; the five named traceback, | ||
| partial-startup, process-interrupt, and async-cancellation tests passed; all 100 | ||
| Agent Gate tests passed. The earlier broader local run passed 89 focused | ||
| non-service tests, while 11 service-backed tests remained mandatory for hosted | ||
| CI. | ||
|
|
||
| ## Internal repair review | ||
|
|
||
| Historical CodeRabbit repair review SHA: `24f3b638b175352ddce3548d8c247b65c3328087` | ||
|
|
||
| Final refreshed repair review SHA: `400be486863e6eb83a6343e872763b9076770537` | ||
|
|
||
| Senior engineering, QA/test, security/auth, product/ops, architecture, CI | ||
| integrity, docs, reuse/dedup, and test-delta tracks passed. Low residual risks | ||
| are limited to an allowlist-based traceback redactor; current workflow secret | ||
| inputs are covered and regression tested, while broader pattern-based hardening | ||
| is deferred. | ||
|
|
||
| ## Remaining risks | ||
|
|
||
| - The final repair head still requires fresh GitHub Agent Gates, Backend, and | ||
| external review. | ||
| - The eight-minute goal remains unmet and is explicitly visible in hosted | ||
| evidence through `timing_target_met: false`; optimization remains deferred. | ||
138 changes: 138 additions & 0 deletions
138
...I-001-backend-ci-acceleration/reviews/WS-CI-001-02B-internal-review-evidence.md
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,138 @@ | ||
| # Internal Review Evidence | ||
|
|
||
| ## Chunk | ||
|
|
||
| `WS-CI-001-02B` — Exact-Custody Semantic Test Lanes | ||
|
|
||
| ## Required Statements | ||
|
|
||
| open sub-agent sessions: none | ||
|
|
||
| valid findings addressed: yes | ||
|
|
||
| ## Reviewed Revision | ||
|
|
||
| Reviewed code SHA: 400be486863e6eb83a6343e872763b9076770537 | ||
|
|
||
| Reviewed at: 2026-07-25T15:42:43Z | ||
|
|
||
| Reviewer run IDs: ci02b_cr_senior, ci02b_cr_qa, ci02b_cr_security, | ||
| ci02b_cr_ops, ci02b_cr_arch, ci02b_cr_ci, ci02b_cr_docs, | ||
| ci02b_cr_reuse, ci02b_cr_test_delta | ||
|
|
||
| The reviewed SHA includes current trusted main through PR #200 and the final | ||
| CodeRabbit repair. After it, only this final evidence reconciliation changed. | ||
|
|
||
| ## Reviewer Results | ||
|
|
||
| | Reviewer | Result | Blocking findings | Notes | | ||
| |---|---:|---|---| | ||
| | senior engineering | PASS | None | Redacted tracebacks and cancellation cleanup are maintainable. | | ||
| | QA/test | PASS | None | Failure diagnostics, redaction, and cancellation are regression protected. | | ||
| | security/auth | PASS WITH LOW RISKS | None | Known workflow secrets are redacted; broader pattern hardening is deferred. | | ||
| | product/ops | PASS | None | No product, compensation, review-decision, or reputation behavior changed. | | ||
| | architecture | PASS | None | Runner ownership and failure-evidence boundaries remain intact. | | ||
| | CI integrity | PASS WITH LOW RISKS | None | Gate posture is unchanged; current secret inputs are covered. | | ||
| | docs | PASS | None | Exact commands and all seventeen findings are reconciled. | | ||
| | reuse/dedup | PASS | None | Existing cleanup and finalization paths remain reused. | | ||
| | test delta | PASS | None | No tests weakened; five focused repair cases pass. | | ||
|
|
||
| The bootstrap repair review initially blocked publication because the reviewed | ||
| SHA was stale and this file had an extra blank line at EOF. Both evidence | ||
| defects are corrected here. All six repair reviewers accepted the fixed-path, | ||
| mode-700 evidence-root initialization; CI integrity found no weakened gate. | ||
|
|
||
| The exact-head coverage repair review found no code blocker. Successful lanes | ||
| still require non-symlink ordinary coverage; an admin runner self-test coverage | ||
| file is combined only when it exists because that direct self-test process can | ||
| legitimately collect no `app` data. Failed lanes retain nonzero exit evidence | ||
| instead of allowing missing coverage to mask the original failure. Optional | ||
| admin coverage absence remains a documented low diagnostic risk, not a product | ||
| coverage or exact-node custody bypass. | ||
|
|
||
| ## Valid Findings Addressed | ||
|
|
||
| - Replaced invalid underscore-bearing MinIO bucket names with S3-valid, | ||
| collision-tested lane namespaces and bound the real S3 test bucket to its | ||
| owning lane. | ||
| - Removed the isolated-runner test exclusion. All runner self-test nodes remain | ||
| canonical and execute under the explicit `admin_runner_self_test` kind while | ||
| ordinary children never receive the admin database URL. | ||
| - Added an independent recursive pytest collection that rejects a | ||
| self-consistent but missing or foreign runner manifest. | ||
| - Preserved full parameterized node IDs while stabilizing import-time UUIDv4 | ||
| parameters by exact head, callsite, line, and ordinal. Both plugins restore | ||
| `uuid.uuid4` and repository import aliases before test bodies execute. | ||
| - Made negative process exits, skip, deselection, interruption, partial | ||
| completion, digest drift, and shared database/storage/coverage custody fail. | ||
| - Removed intermediate coverage files after authenticated per-lane combination | ||
| so the final workflow accepts exactly four public lane artifacts before one | ||
| literal `coverage combine`. | ||
| - Added hosted evidence for total Backend wall time, slowest lane, aggregate | ||
| runner seconds, exact node counts, coverage percentage, raw digests, and the | ||
| explicit 480-second target outcome. Correctness and coverage remain | ||
| fail-closed; the owner-accepted performance miss is recorded rather than | ||
| overriding those results. | ||
| - Redacted direct admin-runner logs and aligned the operations runbook with the | ||
| actual hosted sequence and local diagnostic boundary. | ||
| - Added explicit fixed-path initialization of `.ci/test-lanes` before the first | ||
| hosted collection. This closes the observed `invalid_lane_outputs` bootstrap | ||
| failure without allowing the runner to create an unowned parent directory. | ||
| - Repaired schema-lane coverage finalization after hosted run `30109561363` | ||
| proved that its ordinary unit emitted coverage while the direct admin | ||
| self-test unit legitimately emitted none. Missing ordinary coverage still | ||
| fails closed, and regression tests cover both paths. | ||
| - Rejected the six-process execution-unit experiment after hosted run | ||
| `30118538144` proved it increased CPU contention. The experiment and its | ||
| temporary tests were reverted without weakening the four-lane contract. | ||
| - Rebalanced the four isolated processes around measured hotspots: | ||
| `project_lifecycle`, `task_lifecycle`, `schema_contracts`, and | ||
| `shared_foundations`. Retired lane names were removed from the runner, | ||
| focused tests, runbook, and current status. | ||
| - Reconciled all fourteen CodeRabbit findings. Failure paths now preserve a | ||
| stable Git error, process interrupts, missing isolation metadata, collection | ||
| precedence, and exactly four failed lane rows across runtime and partial | ||
| startup failures. Independent recollection clears inherited pytest injection, | ||
| UUID teardown is unconditional, and Agent Gates bind the explicit `run.exit` | ||
| failure step. | ||
|
|
||
| ## Commands Run | ||
|
|
||
| ```bash | ||
| cd backend | ||
| ruff check app tests scripts | ||
| python -m pytest -q tests/test_ci_test_lanes.py tests/test_test_lane_evidence.py | ||
| python scripts/run_test_lanes.py --collect-only --metadata-dir "$tmp/collect" --summary-json "$tmp/collect-summary.json" | ||
| python scripts/validate_test_lane_evidence.py --metadata-dir "$tmp/collect" --summary-json "$tmp/collect-summary.json" | ||
| cd .. | ||
| python3 scripts/test_agent_gates.py | ||
| python3 scripts/update_post_merge_memory.py validate-merge-intent --repository-root . --base-ref origin/main | ||
| python3 scripts/check_markdown_links.py docs/operations_backend_testing.md .agent-loop/initiatives/WS-CI-001-backend-ci-acceleration/STATUS.md | ||
| git diff --check origin/main...HEAD | ||
| ``` | ||
|
|
||
| ## Results | ||
|
|
||
| - Ruff passed with exact local `ruff 0.15.22`. | ||
| - 89 focused lane-runner, isolated-runner, and independent-validator tests | ||
| passed without local service authority; 11 service-backed cases remain | ||
| mandatory in hosted CI and are not skipped by the workflow. | ||
| - 100 Agent Gate tests passed. | ||
| - Exact collection and independent recollection agreed on 2,056 pytest nodes | ||
| before the timing-only workflow repair; final exact-head hosted recollection | ||
| remains required. | ||
| - Merge intent, Markdown links, stale wording, and diff integrity passed. | ||
| - Local full-service execution was not used as hosted performance evidence. | ||
|
|
||
| ## Remaining Risks | ||
|
|
||
| - The exact GitHub Backend job must still prove real PostgreSQL and MinIO | ||
| concurrency, API E2E, 78/90 coverage gates, and record its 480-second target | ||
| outcome on the final PR head. Prior run `30118538144` passed functional | ||
| and coverage custody but failed timing for the now-reverted six-process | ||
| experiment; it is diagnostic evidence, not completion proof. | ||
| - A force-kill after the bounded cleanup grace can leave runner-owned resources | ||
| on persistent local services. Evidence fails closed; operators must inspect | ||
| and remove only exact recorded resources. | ||
| - The runner and independent validator intentionally duplicate the stable UUID | ||
| collection specification. Their separate tests must prevent common-mode drift. |
Oops, something went wrong.
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
Uh oh!
There was an error while loading. Please reload this page.