Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
34 commits
Select commit Hold shift + click to select a range
161417f
ci(backend): replace shards with exact-custody semantic lanes
Abiorh001 Jul 24, 2026
b771c64
chore(agent-loop): declare semantic lanes merge outcome
Abiorh001 Jul 24, 2026
a3ad393
fix(ci): close semantic lane custody gaps
Abiorh001 Jul 24, 2026
848fa03
fix(ci): stabilize independent node inventory
Abiorh001 Jul 24, 2026
cda3f1f
docs(ci): clarify local diagnostic boundary
Abiorh001 Jul 24, 2026
18ead0b
fix(ci): confine deterministic UUIDs to collection
Abiorh001 Jul 24, 2026
2dfcf06
docs(ci): align hosted proof sequence
Abiorh001 Jul 24, 2026
d3b2a65
docs(agent-loop): record semantic lane trust evidence
Abiorh001 Jul 24, 2026
cc8ee4f
fix(ci): initialize semantic lane evidence root
Abiorh001 Jul 24, 2026
cb63719
docs(agent-loop): record semantic lane bootstrap repair
Abiorh001 Jul 24, 2026
239adb1
fix(ci): preserve schema lane coverage custody
Abiorh001 Jul 24, 2026
80dd8c8
docs(agent-loop): record coverage custody repair
Abiorh001 Jul 24, 2026
4afb6fb
perf(ci): add exact-custody semantic execution units
Abiorh001 Jul 24, 2026
ad7270d
fix(ci): bind nodes to semantic resource units
Abiorh001 Jul 24, 2026
4661681
test(ci): reject shared semantic unit namespaces
Abiorh001 Jul 24, 2026
8bc7e9f
docs(agent-loop): record semantic unit review evidence
Abiorh001 Jul 24, 2026
31516a3
Revert "docs(agent-loop): record semantic unit review evidence"
Abiorh001 Jul 24, 2026
f134920
Revert "test(ci): reject shared semantic unit namespaces"
Abiorh001 Jul 24, 2026
eeab0fb
Revert "fix(ci): bind nodes to semantic resource units"
Abiorh001 Jul 24, 2026
f9ddc55
Revert "perf(ci): add exact-custody semantic execution units"
Abiorh001 Jul 24, 2026
15e3589
perf(ci): rebalance semantic test lanes
Abiorh001 Jul 24, 2026
bf16f1a
test(ci): align semantic lane fixtures
Abiorh001 Jul 24, 2026
f5d2abd
docs(agent-loop): record lane rebalance evidence
Abiorh001 Jul 24, 2026
24f3b63
fix(ci): preserve semantic lane failure evidence
Abiorh001 Jul 24, 2026
cd1e8a8
docs(agent-loop): reconcile external CI review
Abiorh001 Jul 24, 2026
06709b0
fix(ci): separate accepted timing target from gates
Abiorh001 Jul 25, 2026
a44c516
docs(ci): align accepted timing semantics
Abiorh001 Jul 25, 2026
bf3ea27
docs(agent-loop): bind accepted timing review
Abiorh001 Jul 25, 2026
2292ed6
Merge remote-tracking branch 'origin/main' into codex/ws-ci-001-02b-e…
Abiorh001 Jul 25, 2026
398b25d
Merge remote-tracking branch 'origin/main' into codex/ws-ci-001-02b-e…
Abiorh001 Jul 25, 2026
1526ade
Merge remote-tracking branch 'origin/main' into codex/ws-ci-001-02b-e…
Abiorh001 Jul 25, 2026
f601266
docs(agent-loop): bind main integration review
Abiorh001 Jul 25, 2026
400be48
fix(ci): preserve orchestration failure causes
Abiorh001 Jul 25, 2026
fb6a7ff
docs(agent-loop): bind final external repair review
Abiorh001 Jul 25, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -2,13 +2,21 @@

- Phase: signed implementation
- Active planning chunk: none
- Active implementation chunk: `WS-CI-001-02A`
- Proposed implementation successor: `WS-CI-001-02B`
- Later proposed chunk: `WS-CI-001-02B` (cannot start before 02A evidence)
- Active implementation chunk: `WS-CI-001-02B`
- Proposed implementation successor: none
- Later proposed chunk: `WS-CI-001-03` planning only; not a declared successor
- Human direction: preserve Konan's authorship and measured work from PR #180,
but adopt it only under prospective zero-trust scope and evidence
- Current gate: runtime schema recovery reviewed at `cf91bb81`; full Backend suite,
canonical fingerprint, and all coverage gates remain before user-owned merge
decision
- Current gate: reviewed repair SHA `24f3b638` independently collected and
reconciled 2,056 nodes after addressing all fourteen CodeRabbit findings.
Prior PR head `f5d2abd7` passed its 2,049 hosted nodes, API E2E, resource
custody, and coverage floors in run `30121249272`; its required check failed
because 9m39s exceeded the hard 480-second bound. A later exact-head run
again passed all functional and coverage gates in about nine minutes. The
repository owner accepted that measured performance risk and deferred
optimization, so the workflow now records the eight-minute target result
without letting that accepted miss override the required correctness gates.
The repair still requires fresh hosted and external review before the
user-owned merge decision
- Deferred option: routing/cache/timing reassessment is future `WS-CI-001-03`,
with no start or successor declaration
Original file line number Diff line number Diff line change
@@ -0,0 +1,104 @@
# External Review Response

## Chunk

`WS-CI-001-02B` — Exact-Custody Semantic Test Lanes

## Source

CodeRabbit inline review on PR #198, posted 2026-07-24 against `f5d2abd7`.

## Comments addressed

All fourteen comments were valid and addressed:

1. Standardized `parameterized` and `full-service` wording.
2. Clarified that the rebalance preserved the final workflow contract and gates.
3. Replaced the absolute Ruff claim with the observed resolver-drift failure.
4. Synchronized status history with reviewed and hosted node evidence.
5. Initially clarified the 480-second failure behavior. After exact-head hosted
proof exceeded the target while every correctness and coverage gate passed,
the owner explicitly accepted the measured risk and required the accepted
performance target not to leave the required check permanently red.
6. Mapped `git rev-parse` failures to the stable isolated-runner error contract.
7. Preserved `KeyboardInterrupt` and `SystemExit` during MinIO bucket creation.
8. Preserved failed lane evidence when isolation metadata was never created.
9. Reported pytest collection failure before manifest validation.
10. Preserved exactly four failed lane rows after runtime and partial-startup
orchestration failures.
11. Removed inherited `PYTEST_ADDOPTS` and `PYTEST_PLUGINS` from independent
collection.
12. Made UUID restoration unconditional in the validator regression test.
13. Documented PostgreSQL as the service container and MinIO as an in-step,
loopback-published container with a health loop.
14. Added an Agent Gate assertion for the step that reads `run.exit` and fails
the required check.

The refreshed review on exact head `f601266a` added three findings, all
addressed:

15. Replaced placeholder command arguments with the exact focused pytest nodes
and Markdown path used for verification.
16. Preserved the traceback and exception message for unexpected lane
orchestration failures while retaining cleanup and four-lane failure
evidence.
17. Allowed `asyncio.CancelledError` and other non-`Exception` cancellation
signals to propagate during MinIO bucket creation instead of translating
them into namespace collisions.

## Comments deferred

None.

## Human decisions needed

The repository owner explicitly accepted the measured timing risk from hosted
runs `30121249272` and `30123755007` and deferred further speed optimization.
The workflow continues to record whether the 480-second target was met, but an
accepted performance miss no longer overrides successful test, custody,
coverage, isolation, and service-contract gates.

## Commands rerun

```bash
cd backend
.venv/bin/python -m pip install ruff==0.15.22
.venv/bin/ruff check app tests scripts
.venv/bin/python -m pytest -q \
tests/test_ci_test_lanes.py::test_unexpected_runner_failure_force_kills_and_records_every_lane \
tests/test_ci_test_lanes.py::test_partial_startup_failure_records_exactly_four_failed_lanes \
tests/test_isolated_database_runner.py::test_minio_creation_preserves_process_interrupts \
tests/test_isolated_database_runner.py::test_minio_creation_preserves_async_cancellation \
tests/test_isolated_database_runner.py::test_minio_probe_cleans_up_and_preserves_async_cancellation
cd ..
python3 scripts/test_agent_gates.py
python3 scripts/check_internal_review_evidence.py
python3 scripts/check_markdown_links.py \
.agent-loop/initiatives/WS-CI-001-backend-ci-acceleration/reviews/WS-CI-001-02B-external-review-response.md
git diff --check
Comment thread
coderabbitai[bot] marked this conversation as resolved.
```

Latest focused repair result: exact Ruff passed; the five named traceback,
partial-startup, process-interrupt, and async-cancellation tests passed; all 100
Agent Gate tests passed. The earlier broader local run passed 89 focused
non-service tests, while 11 service-backed tests remained mandatory for hosted
CI.

## Internal repair review

Historical CodeRabbit repair review SHA: `24f3b638b175352ddce3548d8c247b65c3328087`

Final refreshed repair review SHA: `400be486863e6eb83a6343e872763b9076770537`

Senior engineering, QA/test, security/auth, product/ops, architecture, CI
integrity, docs, reuse/dedup, and test-delta tracks passed. Low residual risks
are limited to an allowlist-based traceback redactor; current workflow secret
inputs are covered and regression tested, while broader pattern-based hardening
is deferred.

## Remaining risks

- The final repair head still requires fresh GitHub Agent Gates, Backend, and
external review.
- The eight-minute goal remains unmet and is explicitly visible in hosted
evidence through `timing_target_met: false`; optimization remains deferred.
Original file line number Diff line number Diff line change
@@ -0,0 +1,138 @@
# Internal Review Evidence

## Chunk

`WS-CI-001-02B` — Exact-Custody Semantic Test Lanes

## Required Statements

open sub-agent sessions: none

valid findings addressed: yes

## Reviewed Revision

Reviewed code SHA: 400be486863e6eb83a6343e872763b9076770537

Reviewed at: 2026-07-25T15:42:43Z

Reviewer run IDs: ci02b_cr_senior, ci02b_cr_qa, ci02b_cr_security,
ci02b_cr_ops, ci02b_cr_arch, ci02b_cr_ci, ci02b_cr_docs,
ci02b_cr_reuse, ci02b_cr_test_delta

The reviewed SHA includes current trusted main through PR #200 and the final
CodeRabbit repair. After it, only this final evidence reconciliation changed.

## Reviewer Results

| Reviewer | Result | Blocking findings | Notes |
|---|---:|---|---|
| senior engineering | PASS | None | Redacted tracebacks and cancellation cleanup are maintainable. |
| QA/test | PASS | None | Failure diagnostics, redaction, and cancellation are regression protected. |
| security/auth | PASS WITH LOW RISKS | None | Known workflow secrets are redacted; broader pattern hardening is deferred. |
| product/ops | PASS | None | No product, compensation, review-decision, or reputation behavior changed. |
| architecture | PASS | None | Runner ownership and failure-evidence boundaries remain intact. |
| CI integrity | PASS WITH LOW RISKS | None | Gate posture is unchanged; current secret inputs are covered. |
| docs | PASS | None | Exact commands and all seventeen findings are reconciled. |
| reuse/dedup | PASS | None | Existing cleanup and finalization paths remain reused. |
| test delta | PASS | None | No tests weakened; five focused repair cases pass. |

The bootstrap repair review initially blocked publication because the reviewed
SHA was stale and this file had an extra blank line at EOF. Both evidence
defects are corrected here. All six repair reviewers accepted the fixed-path,
mode-700 evidence-root initialization; CI integrity found no weakened gate.

The exact-head coverage repair review found no code blocker. Successful lanes
still require non-symlink ordinary coverage; an admin runner self-test coverage
file is combined only when it exists because that direct self-test process can
legitimately collect no `app` data. Failed lanes retain nonzero exit evidence
instead of allowing missing coverage to mask the original failure. Optional
admin coverage absence remains a documented low diagnostic risk, not a product
coverage or exact-node custody bypass.

## Valid Findings Addressed

- Replaced invalid underscore-bearing MinIO bucket names with S3-valid,
collision-tested lane namespaces and bound the real S3 test bucket to its
owning lane.
- Removed the isolated-runner test exclusion. All runner self-test nodes remain
canonical and execute under the explicit `admin_runner_self_test` kind while
ordinary children never receive the admin database URL.
- Added an independent recursive pytest collection that rejects a
self-consistent but missing or foreign runner manifest.
- Preserved full parameterized node IDs while stabilizing import-time UUIDv4
parameters by exact head, callsite, line, and ordinal. Both plugins restore
`uuid.uuid4` and repository import aliases before test bodies execute.
- Made negative process exits, skip, deselection, interruption, partial
completion, digest drift, and shared database/storage/coverage custody fail.
- Removed intermediate coverage files after authenticated per-lane combination
so the final workflow accepts exactly four public lane artifacts before one
literal `coverage combine`.
- Added hosted evidence for total Backend wall time, slowest lane, aggregate
runner seconds, exact node counts, coverage percentage, raw digests, and the
explicit 480-second target outcome. Correctness and coverage remain
fail-closed; the owner-accepted performance miss is recorded rather than
overriding those results.
- Redacted direct admin-runner logs and aligned the operations runbook with the
actual hosted sequence and local diagnostic boundary.
- Added explicit fixed-path initialization of `.ci/test-lanes` before the first
hosted collection. This closes the observed `invalid_lane_outputs` bootstrap
failure without allowing the runner to create an unowned parent directory.
- Repaired schema-lane coverage finalization after hosted run `30109561363`
proved that its ordinary unit emitted coverage while the direct admin
self-test unit legitimately emitted none. Missing ordinary coverage still
fails closed, and regression tests cover both paths.
- Rejected the six-process execution-unit experiment after hosted run
`30118538144` proved it increased CPU contention. The experiment and its
temporary tests were reverted without weakening the four-lane contract.
- Rebalanced the four isolated processes around measured hotspots:
`project_lifecycle`, `task_lifecycle`, `schema_contracts`, and
`shared_foundations`. Retired lane names were removed from the runner,
focused tests, runbook, and current status.
- Reconciled all fourteen CodeRabbit findings. Failure paths now preserve a
stable Git error, process interrupts, missing isolation metadata, collection
precedence, and exactly four failed lane rows across runtime and partial
startup failures. Independent recollection clears inherited pytest injection,
UUID teardown is unconditional, and Agent Gates bind the explicit `run.exit`
failure step.

## Commands Run

```bash
cd backend
ruff check app tests scripts
python -m pytest -q tests/test_ci_test_lanes.py tests/test_test_lane_evidence.py
python scripts/run_test_lanes.py --collect-only --metadata-dir "$tmp/collect" --summary-json "$tmp/collect-summary.json"
python scripts/validate_test_lane_evidence.py --metadata-dir "$tmp/collect" --summary-json "$tmp/collect-summary.json"
cd ..
python3 scripts/test_agent_gates.py
python3 scripts/update_post_merge_memory.py validate-merge-intent --repository-root . --base-ref origin/main
python3 scripts/check_markdown_links.py docs/operations_backend_testing.md .agent-loop/initiatives/WS-CI-001-backend-ci-acceleration/STATUS.md
git diff --check origin/main...HEAD
```

## Results

- Ruff passed with exact local `ruff 0.15.22`.
- 89 focused lane-runner, isolated-runner, and independent-validator tests
passed without local service authority; 11 service-backed cases remain
mandatory in hosted CI and are not skipped by the workflow.
- 100 Agent Gate tests passed.
- Exact collection and independent recollection agreed on 2,056 pytest nodes
before the timing-only workflow repair; final exact-head hosted recollection
remains required.
- Merge intent, Markdown links, stale wording, and diff integrity passed.
- Local full-service execution was not used as hosted performance evidence.

## Remaining Risks

- The exact GitHub Backend job must still prove real PostgreSQL and MinIO
concurrency, API E2E, 78/90 coverage gates, and record its 480-second target
outcome on the final PR head. Prior run `30118538144` passed functional
and coverage custody but failed timing for the now-reverted six-process
experiment; it is diagnostic evidence, not completion proof.
- A force-kill after the bounded cleanup grace can leave runner-owned resources
on persistent local services. Evidence fails closed; operators must inspect
and remove only exact recorded resources.
- The runner and independent validator intentionally duplicate the stable UUID
collection specification. Their separate tests must prevent common-mode drift.
Loading
Loading