From 5165e0750ace10b5b38c7b6b797bced5284c8dd8 Mon Sep 17 00:00:00 2001 From: forhappy Date: Wed, 30 Sep 2026 16:40:24 -0700 Subject: [PATCH 01/60] Record failed three-node load and acknowledged creation recovery --- docs/performance-plan.md | 10 +- .../2026-09-30-three-node-proxy.md | 147 ++++++++++++++++-- 2 files changed, 146 insertions(+), 11 deletions(-) diff --git a/docs/performance-plan.md b/docs/performance-plan.md index 5b2f5c8..5359cf4 100644 --- a/docs/performance-plan.md +++ b/docs/performance-plan.md @@ -4,9 +4,15 @@ Use this plan to design capacity work and interpret Canopy benchmark results. It separates demonstrated behavior from proposed targets. The current Cellule dependency candidate is `e07670e2348231ed401cc7280a47e3ab97596ffe`, fetched from `origin/main` during the [three-node proxy campaign](performance/2026-09-30-three-node-proxy.md#prepare-the-next-upstream-candidate). -Its build, end-to-end and performance verification remain separate open gates. +One hosted workflow passed its build, workspace tests and isolated RustFS +qualification; a sibling failed a delayed-startup test with `AddrInUse`. +Local three-node end-to-end and matched performance verification remain open. The [workspace/RustFS verification](performance/2026-09-30-workspace-rustfs.md) -and running three-node baseline retain `70bd25f142f1976fdd63ffe60e46e15ae276ffdc`. +and the retained three-node baseline use `70bd25f142f1976fdd63ffe60e46e15ae276ffdc`. +That baseline stopped after 20 fully bound windows and one interrupted creation +window when all three owners fenced. All 136 acknowledged creations and both +critical Git fixtures passed fresh-state recovery. The full original corpus +recheck remains in progress before further load; failed results stay in the record. Earlier results do not establish the new candidate's density or latency. The [chunk-count follow-up](performance/2026-09-30-chunk-count.md) removes diff --git a/docs/performance/2026-09-30-three-node-proxy.md b/docs/performance/2026-09-30-three-node-proxy.md index e4ad1d5..b7db41c 100644 --- a/docs/performance/2026-09-30-three-node-proxy.md +++ b/docs/performance/2026-09-30-three-node-proxy.md @@ -1,6 +1,6 @@ # Qualify three Canopy nodes behind a proxy -**Status: in progress.** This campaign preserves the documented 10,000-repository +**Status: baseline failed; fresh-state recovery in progress.** This campaign preserves the documented 10,000-repository reference target. A ready fleet, small socket tests or an incomplete seed does not qualify throughput, latency, recovery or production capacity. @@ -59,15 +59,20 @@ Cellule `origin/main` advanced during the running baseline. Canopy's manifests and six Cellule lockfile entries now select `e07670e2348231ed401cc7280a47e3ab97596ffe`. Dependency resolution and locked metadata validation passed without updating unrelated packages or creating a -local Cargo target directory. Compilation and verification of this candidate -remain open; the measurements in this document still use the retained +local Cargo target directory. [Workflow 36789746786](https://github.com/crabbuild/canopy/actions/runs/36789746786) +passed formatting, clippy, workspace tests, isolated RustFS compatibility and +server build. Its [sibling workflow](https://github.com/crabbuild/canopy/actions/runs/36789752428) +failed `lifecycle::lease::startup_preflight_does_not_consume_the_node_lease` +with `AddrInUse`; provider qualification and build were skipped in that job. +Both used source `a2c59a3`; neither is a local three-node performance comparison. +Local candidate verification remains open; the measurements in this document use the retained `70bd25f` binary, not the new source pin. The intervening commits add [fenced route reuse and coalesced owner discovery](https://github.com/crabbuild/cellule/pull/32) and [single-enrollment-read follower append authorization](https://github.com/crabbuild/cellule/pull/31). Their relevance to the observed directory/latency failures is a hypothesis, -not a measured fix. The live campaign's scripts, binary and provider configuration -remain unchanged. CI verification can run independently; local compilation, +not a measured fix. The baseline's scripts, binary and provider configuration +were unchanged during load. CI verification can run independently; local compilation, fresh-state recovery and matched performance comparisons must follow the recorded baseline and produce their own artifact bindings. @@ -85,11 +90,11 @@ recorded baseline and produce their own artifact bindings. | Replacement fleet | Three live nodes behind one proxy; matching UUID read through each node | | Fresh 10,000 seed | Completed with exit 0: 10,000 distinct canonical repository UUIDs, 100 two-commit Git fixtures and 100 exact 1 MiB LFS fixtures; 8,567.224 seconds | | Critical stock-Git probe through ingress | Passed 17 functional steps against the live RustFS fixture; exact nine-ref inventories in two repositories, v0/v2 mirror clones and strict full fsck | -| Scheduled load matrix | Not yet qualified | +| Scheduled load matrix | Failed: 20 fully bound windows and one interrupted creation window; all 21 complete arrival ledgers retained. Git/LFS load phases had not started | | Initial three-owner loss | All three verified seeding-node PIDs killed; process absence recorded, followed by 32.006 seconds of unchanged lease-expiry wait; no data deleted | | Critical-fixture fresh-state recovery | Passed: both original identities, four exact nine-ref inventories across v0/v2, payload, Git notes and strict full fsck | | Full-corpus recovery / preflight | Passed: all 10,000 original identities and all 100 populated fixtures through three fresh node directories and the new ingress; v0/v2 clones, exact Git/incremental/LFS bytes and strict full fsck; four verification workers, no retries | -| Post-load acknowledged-write recovery | Not yet run | +| Post-load acknowledged-write recovery | All 136 acknowledged creations passed after observed three-owner failure and fresh-state startup. No Git/LFS load writes existed because those phases had not started; full original-corpus recheck remains in progress | The old seed is not resumed or merged into a passing manifest. The replacement uses an independent provider, bucket and deployment prefix. The original 30-second @@ -126,8 +131,9 @@ configurations into 108 windows: preparation and drain. It covers metadata with 100/500/1,000-identity uniform and skewed sets, creation, discovery, stock Git reads and incremental pull/fetch, ref-only pushes, fresh 256 KiB/1 MiB child-commit pushes and 1 MiB LFS transfers. -Git concurrency and offered rate vary independently. Load clocks have started -after complete-corpus preflight; the matrix is not complete. Higher admission +Git concurrency and offered rate vary independently. Load clocks started +after complete-corpus preflight; the fleet later fenced during creation and the +matrix is not complete. Higher admission profiles, rate envelopes, further skewed Git tests and complete acknowledged-write recovery remain required. @@ -186,6 +192,129 @@ this is a correlated failure path, not a proven Cellule bottleneck. Same-set repetition improves the uniform measurements, but the skewed repeats still fail. No runtime optimization, retry, timeout or admission increase was applied. +## Retain the failed baseline and recover acknowledged writes + +At approximately `2026-09-30T23:28:18Z`, all three baseline nodes exited with +code 1. Node logs record Cell fencing and +`Runtime(Node("node lease bounds are invalid"))`; the launcher records +`node-1 exited with 1`. The proxy closed, and the campaign terminated with +`resource observation failed` rather than restarting or retrying a window. +RustFS still had the same container identity and start time, without a restart +or OOM kill. No provider or node data was deleted. + +The terminal audit verifies all 21 report/sample digests, complete unique +arrival sequence coverage, outcome totals, in-window successes and nearest-rank +latency quantiles. The first 20 windows also retain bound resource boundaries. +The interrupted third creation window has a complete report and sample ledger, +but no valid after-boundary: it is **not resource-qualified** and is absent from +the completed-report list. Its acknowledgement ledger must still be included in +recovery checks. + +### Metadata with larger selected sets + +Each window offered 2,400 arrivals at 20 requests/s with 32 driver slots and +the same 100-entry admission limit per node. The larger selected sets do not +raise that limit. Attempt percentiles include completed errors; driver drops +have no fabricated latency. In-window rates exclude successful drain completions. + +| Selected identities | Access | Repeat | OK / drops / HTTP 503 / transport errors | Successful req/s in window | Attempt p95 / p99 (ms) | +| --- | --- | --- | --- | --- | --- | +| 500 | Uniform | 1 | 1,982 / 401 / 17 / 0 | 16.483 | 3,708.603 / 4,579.120 | +| 500 | Uniform | 2 | 2,028 / 362 / 10 / 0 | 16.875 | 3,857.033 / 6,677.808 | +| 500 | Uniform | 3 | 1,977 / 391 / 32 / 0 | 16.225 | 3,832.183 / 6,385.354 | +| 500 | Skewed | 1 | 1,896 / 493 / 11 / 0 | 15.800 | 2,903.415 / 19,529.582 | +| 500 | Skewed | 2 | 2,186 / 210 / 4 / 0 | 17.975 | 2,567.978 / 6,886.433 | +| 500 | Skewed | 3 | 1,770 / 586 / 44 / 0 | 14.550 | 5,200.413 / 8,219.480 | +| 1,000 | Uniform | 1 | 280 / 2,096 / 6 / 18 | 2.075 | 30,004.694 / 30,015.112 | +| 1,000 | Uniform | 2 | 531 / 1,855 / 9 / 5 | 4.167 | 20,763.359 / 29,897.866 | +| 1,000 | Uniform | 3 | 881 / 1,497 / 22 / 0 | 7.083 | 9,454.380 / 20,007.887 | +| 1,000 | Skewed | 1 | 2,064 / 336 / 0 / 0 | 17.192 | 3,862.827 / 6,904.098 | +| 1,000 | Skewed | 2 | 2,368 / 32 / 0 / 0 | 19.725 | 519.354 / 2,599.916 | +| 1,000 | Skewed | 3 | 2,302 / 40 / 58 / 0 | 19.183 | 1,586.558 / 4,136.254 | + +A user-requested UI preview began during the final metadata windows and used +its own RustFS instance and deployment, without touching benchmark storage. +It still consumed the shared Mac/Colima host. The audit conservatively flags +the last two metadata windows and all creation windows as possible overlaps; +those are not uncontended capacity measurements. The earlier broad-set failures +predate the preview. Preview activity is retained in `ui-preview-intervention.json`. + +### Creation before owner failure + +All three windows offered 120 unique new names at 1 repository/s and concurrency +16. Creation acknowledges a canonical repository UUID, not just an accepted +HTTP request. A lost or refused response is not an acknowledgement, even if +storage might contain an unacknowledged reservation. + +| Repeat | Acknowledged | Driver drops | HTTP 503 | Transport errors | Successful repos/s in window | Attempt p95 / p99 (ms) | +| --- | --- | --- | --- | --- | --- | --- | +| 1 | 85 | 14 | 5 | 16 | 0.708 | 30,006.512 / 30,006.951 | +| 2 | 51 | 27 | 11 | 31 | 0.425 | 30,071.556 / 30,163.645 | +| 3, interrupted | 0 | 0 | 1 | 119 | 0.000 | 9.802 / 22.701 | + +The interrupted window's short error latencies are failed connection attempts, +**not an improvement**. None establishes a stable creation envelope. The audit +retains 136 distinct acknowledged name/UUID pairs across the complete creation +ledgers; the 224 failed arrivals are not asserted to have been committed. + +### Fresh-state recovery gate + +```mermaid +flowchart LR + failure[Three owners fenced] --> absent[Confirm all old PIDs absent] + absent --> expiry[Wait 32 seconds] + expiry --> fresh[Three new IDs and local directories] + fresh --> acknowledgements[Verify all 136 creation ACKs] + acknowledgements --> critical[Verify both critical Git fixtures] + critical --> corpus[Verify all 10,000 original identities and 100 Git/LFS fixtures] +``` + +`failed-owner-exit.json` records all four old process IDs absent followed by +32.076 seconds of conservative lease-expiry wait. This boundary is an observed +failure, not a deliberate SIGKILL. `fleet-node100-after-failure` starts the same +retained release binary and durable deployment with three new node IDs, signing +keys and empty local directories, still capped at 100 active entries per node. +Its proxy is separate from the user's UI preview. + +The recovery run passed all 136 creation ACKs and both critical Git fixtures: +four exact ref inventories across v0/v2 mirror clones, exact payload/notes and +strict full fsck. It is now checking the full original corpus with four bounded +workers and no retries. The critical receipt's digest matches the earlier +recovery result because its expected fixture content did not change; the new +owner boundary and fleet binding are recorded separately, not inferred from +that content hash. Successful creation/critical recovery does not close the +full-corpus gate or qualify any +unstarted Git/LFS load phase. The new dependency candidate must subsequently +produce its own release artifact, end-to-end and matched performance evidence; +it cannot replace verification of writes made by this failed baseline. + +| Artifact under `canopy-three-proxy-q3FO2z` | SHA-256 | +| --- | --- | +| `failed-baseline-audit.json` | `398fb729e1017ddd956a6fdc6e1c5dc528f1efc86dbbeeec4035fd4e7ecd6b62` | +| `fleet-node100-matched/outcome.json` | `153174bf051cd8795039f43c6b38a6fad4416645244c58aeb71a853bba5f136f` | +| `failed-owner-exit.json` | `16b300d2724ec52a73f6dab656c682c70567a6669a6612bd0fb8baa920a5d889` | +| `fleet-node100-after-failure/ready.json` | `47bd95391d667be68403dd55d02f0349d3dc874f908b64ced3ab112611056062` | +| `creation-after-failure-verification.json` | `45127974487f1b0e272a1d7603886b84b2afe5c3255a8a04653e380c5a41db14` | +| `critical-after-failure-verification.json` | `2aac52dcca5973fbe9ed8badebd5eda2c97563e091d94816df5735f76f9c0c5d` | + +### Diagnosis boundaries + +The independent provider observer compares adjacent UTC and monotonic sample +clocks; their largest difference over the observation was 0.007746 seconds. +There is no sampled large wall-clock jump. RustFS CPU observations near fencing +include 214.24% and 141.15% under a two-core quota, followed by reduced activity. +Earlier node warnings record successful but slow lease refreshes. These narrow +the next investigation toward provider/publication pressure and the lease +refresh path, but do not isolate its cause or prove a Cellule optimization. +The production 30-second authority interval, client deadlines, admission bounds +and fencing checks remain unchanged. + +The mixed hosted CI result is a separate problem: the delayed-startup test +obtains and releases an ephemeral listener before its 31-second pause, then +binds that same address at startup. `AddrInUse` is consistent with reuse by a +parallel test; it is not evidence that the lease assertion failed. A +deterministic occupied-port reproduction and fixture repair are still required. + ## Measure the requested operations Complete and verify the corpus before interpreting scheduled load results. From 59e780b099f12ebaeb1e104bfcc0708871c305db Mon Sep 17 00:00:00 2001 From: forhappy Date: Wed, 30 Sep 2026 16:51:57 -0700 Subject: [PATCH 02/60] Preserve failed post-load corpus recovery and diagnostic reads --- docs/performance-plan.md | 3 +- .../2026-09-30-three-node-proxy.md | 36 ++++++++++++++++--- 2 files changed, 34 insertions(+), 5 deletions(-) diff --git a/docs/performance-plan.md b/docs/performance-plan.md index 5359cf4..349f05c 100644 --- a/docs/performance-plan.md +++ b/docs/performance-plan.md @@ -12,7 +12,8 @@ and the retained three-node baseline use `70bd25f142f1976fdd63ffe60e46e15ae276ff That baseline stopped after 20 fully bound windows and one interrupted creation window when all three owners fenced. All 136 acknowledged creations and both critical Git fixtures passed fresh-state recovery. The full original corpus -recheck remains in progress before further load; failed results stay in the record. +recheck failed with a 30-second identity-read timeout; separate successful read +probes do not turn that attempt into a pass. Failed results stay in the record. Earlier results do not establish the new candidate's density or latency. The [chunk-count follow-up](performance/2026-09-30-chunk-count.md) removes diff --git a/docs/performance/2026-09-30-three-node-proxy.md b/docs/performance/2026-09-30-three-node-proxy.md index b7db41c..93831e6 100644 --- a/docs/performance/2026-09-30-three-node-proxy.md +++ b/docs/performance/2026-09-30-three-node-proxy.md @@ -1,6 +1,6 @@ # Qualify three Canopy nodes behind a proxy -**Status: baseline failed; fresh-state recovery in progress.** This campaign preserves the documented 10,000-repository +**Status: baseline and post-load full-corpus recovery failed; qualification in progress.** This campaign preserves the documented 10,000-repository reference target. A ready fleet, small socket tests or an incomplete seed does not qualify throughput, latency, recovery or production capacity. @@ -94,7 +94,7 @@ recorded baseline and produce their own artifact bindings. | Initial three-owner loss | All three verified seeding-node PIDs killed; process absence recorded, followed by 32.006 seconds of unchanged lease-expiry wait; no data deleted | | Critical-fixture fresh-state recovery | Passed: both original identities, four exact nine-ref inventories across v0/v2, payload, Git notes and strict full fsck | | Full-corpus recovery / preflight | Passed: all 10,000 original identities and all 100 populated fixtures through three fresh node directories and the new ingress; v0/v2 clones, exact Git/incremental/LFS bytes and strict full fsck; four verification workers, no retries | -| Post-load acknowledged-write recovery | All 136 acknowledged creations passed after observed three-owner failure and fresh-state startup. No Git/LFS load writes existed because those phases had not started; full original-corpus recheck remains in progress | +| Post-load acknowledged-write recovery | All 136 acknowledged creations passed after observed three-owner failure and fresh-state startup. No Git/LFS load writes existed because those phases had not started; full original-corpus recheck failed with an identity-read timeout | The old seed is not resumed or merged into a passing manifest. The replacement uses an independent provider, bucket and deployment prefix. The original 30-second @@ -278,8 +278,14 @@ Its proxy is separate from the user's UI preview. The recovery run passed all 136 creation ACKs and both critical Git fixtures: four exact ref inventories across v0/v2 mirror clones, exact payload/notes and -strict full fsck. It is now checking the full original corpus with four bounded -workers and no retries. The critical receipt's digest matches the earlier +strict full fsck. Its full original-corpus check then failed with four bounded +workers and no retries. The last complete progress checkpoint was 300/10,000; +the identity request for `density-ea9283d2648c-00355` +(`ed794e34-158b-4cc1-9e71-d21d7405ab2b`, request ID +`b8333796-0a12-47a9-b37a-0f4ffa56fbd0`) timed out at the unchanged 30-second +deadline. The verifier process exited with code 1 and its result records +`completed: false`, `error: RuntimeError`; this is not a complete recovery pass. +The critical receipt's digest matches the earlier recovery result because its expected fixture content did not change; the new owner boundary and fleet binding are recorded separately, not inferred from that content hash. Successful creation/critical recovery does not close the @@ -296,6 +302,28 @@ it cannot replace verification of writes made by this failed baseline. | `fleet-node100-after-failure/ready.json` | `47bd95391d667be68403dd55d02f0349d3dc874f908b64ced3ab112611056062` | | `creation-after-failure-verification.json` | `45127974487f1b0e272a1d7603886b84b2afe5c3255a8a04653e380c5a41db14` | | `critical-after-failure-verification.json` | `2aac52dcca5973fbe9ed8badebd5eda2c97563e091d94816df5735f76f9c0c5d` | +| `recovery-after-failure.json`, terminal failed attempt | `5113e16b0fee7c59a5165c98d48810e47090a16df6a44425b53e97150d791430` | +| `verify-failed-baseline.log` | `af6ac6a69c47acc935aa7aa0b8d4ce3926111b043bb438d5394ec6cd26ec36e5` | +| `failed-recovery-read-probes.json`, separate diagnostic | `596888de9b9db6184b7fd5fe4cc72b13ea08941a3fa10bf1f8406eb19d2a1844` | + +After the failed verifier exited, all three recovery nodes remained live and +readiness succeeded. Separate, explicitly diagnostic requests each checked +Directory authentication and the failed repository identity, with new request +IDs and no hidden retries. All six returned 200; identity checks matched exactly. +Directory reads took 1,303.672, 6.145 and 399.635 ms; repository reads took +789.856, 1,572.026 and 841.882 ms. The earlier timed-out read may have left the +repository resident. These measurements occurred while the separate candidate +build ran on the shared host: they are neither a cold recovery retry nor a warm +reference-capacity result. The failed full-corpus record remains unchanged. + +The new dependency build uses a separate artifact-volume target directory and +one Cargo build job; it does not replace any running server or recreate a +workspace `target` directory. Release admission binds the Cargo lockfile and +compiled release. The new binary therefore requires its own fresh prefix; +old-release migration is unsupported and admission must not be bypassed to +reuse the baseline's store. The original 10,000 identities and creation ACKs +remain retained for baseline recovery diagnosis, not reseeded as a passing +replacement. ### Diagnosis boundaries From b3ea671a0d112e3757f121dbc2eed4219e1ce977 Mon Sep 17 00:00:00 2001 From: forhappy Date: Wed, 30 Sep 2026 17:16:40 -0700 Subject: [PATCH 03/60] Fix delayed-startup fixture port race and bind Cellule candidate A competing listener reproduced AddrInUse after the original preflight pause. Allocate the server port at bind time and retain a listener on the old address while preserving all lease assertions. Record the immutable production build and functional Git checks separately from incomplete candidate recovery and performance qualification. --- .../tests/multi_server/lifecycle/lease.rs | 15 +++- docs/performance-plan.md | 6 +- .../2026-09-30-three-node-proxy.md | 80 +++++++++++++++++-- 3 files changed, 91 insertions(+), 10 deletions(-) diff --git a/crates/canopy-server/tests/multi_server/lifecycle/lease.rs b/crates/canopy-server/tests/multi_server/lifecycle/lease.rs index 215644c..1ed44fd 100644 --- a/crates/canopy-server/tests/multi_server/lifecycle/lease.rs +++ b/crates/canopy-server/tests/multi_server/lifecycle/lease.rs @@ -113,8 +113,12 @@ async fn slow_deployment_read_does_not_block_node_lease_renewal() -> Result { async fn startup_preflight_does_not_consume_the_node_lease() -> Result { let files = tempfile::TempDir::new()?; let store = Arc::new(PausedStore::default()); - let address = available_address().await?; - let settings = config(address, files.path().join("node")); + let competing_listener = TcpListener::bind("127.0.0.1:0").await?; + let address = competing_listener.local_addr()?; + let mut settings = config(address, files.path().join("node")); + // Bind only when preflight finishes; an address observed before the + // 31-second pause is not a reservation against parallel tests. + settings.listen.set_port(0); let layout = cellule_runtime::ltx::CellStorageLayout::new( cellule_store::Store::new(store.clone()), settings.store_prefix.clone(), @@ -123,6 +127,8 @@ async fn startup_preflight_does_not_consume_the_node_lease() -> Result { *store.read.lock().unwrap() = Some((layout.release_path(), 1)); let startup = tokio::spawn(CanopyServer::start(settings, store.clone())); store.wait().await?; + // Keep the originally selected port occupied throughout startup so this + // regression cannot pass by merely winning the released-port race. // Deployment validation happens before this node owns any Cell. Its I/O // must not spend the node authority lease used by subsequent startup. tokio::time::sleep(Duration::from_secs(31)).await; @@ -133,6 +139,8 @@ async fn startup_preflight_does_not_consume_the_node_lease() -> Result { )?; store.proceed.notify_one(); let server = timeout(Duration::from_secs(10), startup).await???; + let actual_address = server.local_addr(); + assert_ne!(actual_address, address); let initial = store .initial_advertisement .lock() @@ -143,8 +151,9 @@ async fn startup_preflight_does_not_consume_the_node_lease() -> Result { .as_str() .ok_or("advertisement issue time missing")? .parse::()?; - create_repository(address, "after-slow-preflight").await?; + create_repository(actual_address, "after-slow-preflight").await?; server.shutdown().await?; + drop(competing_listener); assert!( issued >= preflight_finished, "preflight consumed the initial node lease" diff --git a/docs/performance-plan.md b/docs/performance-plan.md index 349f05c..bf0b95f 100644 --- a/docs/performance-plan.md +++ b/docs/performance-plan.md @@ -6,7 +6,11 @@ dependency candidate is `e07670e2348231ed401cc7280a47e3ab97596ffe`, fetched from `origin/main` during the [three-node proxy campaign](performance/2026-09-30-three-node-proxy.md#prepare-the-next-upstream-candidate). One hosted workflow passed its build, workspace tests and isolated RustFS qualification; a sibling failed a delayed-startup test with `AddrInUse`. -Local three-node end-to-end and matched performance verification remain open. +That fixture race was reproduced and repaired without changing production +authority checks; all four lease lifecycle tests passed locally. A separately +bound production release passed 17 critical stock-Git checks through three +nodes and a proxy against RustFS. Its new 10,000-repository seed is in progress; +complete-corpus recovery and matched performance verification remain open. The [workspace/RustFS verification](performance/2026-09-30-workspace-rustfs.md) and the retained three-node baseline use `70bd25f142f1976fdd63ffe60e46e15ae276ffdc`. That baseline stopped after 20 fully bound windows and one interrupted creation diff --git a/docs/performance/2026-09-30-three-node-proxy.md b/docs/performance/2026-09-30-three-node-proxy.md index 93831e6..44f3665 100644 --- a/docs/performance/2026-09-30-three-node-proxy.md +++ b/docs/performance/2026-09-30-three-node-proxy.md @@ -65,8 +65,10 @@ server build. Its [sibling workflow](https://github.com/crabbuild/canopy/actions failed `lifecycle::lease::startup_preflight_does_not_consume_the_node_lease` with `AddrInUse`; provider qualification and build were skipped in that job. Both used source `a2c59a3`; neither is a local three-node performance comparison. -Local candidate verification remains open; the measurements in this document use the retained -`70bd25f` binary, not the new source pin. +The candidate's local release build and 17-step critical Git check now passed. +Its 10,000-repository seed is in progress; full-corpus recovery and matched load +remain open. The baseline measurements below use the retained `70bd25f` binary, +not the new source pin. The intervening commits add [fenced route reuse and coalesced owner discovery](https://github.com/crabbuild/cellule/pull/32) and [single-enrollment-read follower append authorization](https://github.com/crabbuild/cellule/pull/31). @@ -76,6 +78,51 @@ were unchanged during load. CI verification can run independently; local compila fresh-state recovery and matched performance comparisons must follow the recorded baseline and produce their own artifact bindings. +### Bind the local candidate separately + +The locked production release built with one Cargo job in 578.311 seconds. +Its production source binding is `5165e075`, equivalent to merged `d8a6e114` +for the bound manifests, lockfile and production files. Build time is setup +evidence, not server throughput. Cargo integration tests can overwrite +`target/release/canopy` with a test-feature variant, so the fleet uses a retained +immutable copy of the original production executable, not that mutable path. + +| Candidate input or gate | Evidence / status | +| --- | --- | +| Cellule revision | `e07670e2348231ed401cc7280a47e3ab97596ffe` | +| Production executable SHA-256 | `e90728c60cbb941cc8caa1c698bc24bcf6edda1935d98853f0c0319933668f16` | +| RustFS | New container `canopy-candidate-rustfs-w0sfigmz`, same pinned ARM64 manifest as the baseline, two CPU quota / 2 GiB memory; fresh bucket and prefix | +| Fleet | Three new node identities/directories behind one proxy; unchanged 100-entry admission and 1.5 GiB configured local disk limit per node | +| Critical Git behavior | All 17 steps passed: atomic publication/refusal, mixed refusal, correct/stale force-with-lease, shallow/deepen/unshallow, filtered lazy fetch v0/v2, new-object push, fast-forward pull, deletion/pruning, mirror push, credential refusal and exact mirror inventories with strict full fsck | +| Corpus setup | New 10,000 identities with 100 two-commit Git and 1 MiB LFS fixtures requested; seed in progress, not yet complete | +| Recovery and performance | Candidate owner-loss, full-corpus recovery, scheduled load and matched comparisons remain unqualified | + +Artifacts are retained under `canopy-e07670e-candidate-W0SFiGMz` on the +qualification volume: + +| Artifact | SHA-256 | +| --- | --- | +| `build.json` | `f4a8b63fede3fc0556f9d956513bbe9fc85ad3db9ec8b9f3421aba7aae4f6e9b` | +| `candidate-binding.json` | `9f4be2686d6f74fe1c8f45ae6b2afa6b6e9d108d7a3b4ddde5839e0d0e9ec517` | +| `fleet-candidate-matched/ready.json` | `975088199e3d1b2ba1aa6ce234253915dca0dd885566843bcc88f29413232822` | +| `critical-candidate.json` | `48fc4f70c291aa3770d4cf093c3945603ddcaf3ef393a7ded3f66e10139fb9dd` | + +The candidate's two critical repositories are separate from its density corpus: +`910d4eae-ca4a-4495-bb10-06a8cd9e9d23` and +`092a2349-8cf8-44ed-8b31-43ccae05d7d5`. The receipt binds the executable and +driver digests. Compound functional-step durations are not scheduled operation +latency or throughput. A separate observer retains provider gauges and node +RSS/cumulative CPU during seeding; it excludes native Git child CPU and does not +measure S3 API cost. + +Failed setup remains visible. `fleet-candidate-seed` failed its storage probe +with `NotFound` before readiness; the cause is not isolated. After explicit +bucket readiness, `fleet-candidate-ready` started a test-feature executable, +then all three nodes shut down gracefully with exit 0 before any corpus was +seeded. Neither setup is qualification evidence. The current fleet uses the +immutable production artifact and its own binding. Successful setup does not +erase the earlier baseline's failed full-corpus recovery. + ## Retain failed setup and incomplete work | Check | Evidence / status | @@ -316,7 +363,7 @@ repository resident. These measurements occurred while the separate candidate build ran on the shared host: they are neither a cold recovery retry nor a warm reference-capacity result. The failed full-corpus record remains unchanged. -The new dependency build uses a separate artifact-volume target directory and +The new dependency build used a separate artifact-volume target directory and one Cargo build job; it does not replace any running server or recreate a workspace `target` directory. Release admission binds the Cargo lockfile and compiled release. The new binary therefore requires its own fresh prefix; @@ -325,6 +372,12 @@ reuse the baseline's store. The original 10,000 identities and creation ACKs remain retained for baseline recovery diagnosis, not reseeded as a passing replacement. +After preserving the failed recovery and diagnostic probes, all three baseline +recovery nodes drained gracefully with exit 0 and confirmed process absence. +`failed-recovery-drain.json` records the boundary. The dedicated baseline RustFS +container was then stopped, not deleted; its durable data and node artifacts +remain available for diagnosis. The independent user UI preview stays running. + ### Diagnosis boundaries The independent provider observer compares adjacent UTC and monotonic sample @@ -339,9 +392,24 @@ and fencing checks remain unchanged. The mixed hosted CI result is a separate problem: the delayed-startup test obtains and releases an ephemeral listener before its 31-second pause, then -binds that same address at startup. `AddrInUse` is consistent with reuse by a -parallel test; it is not evidence that the lease assertion failed. A -deterministic occupied-port reproduction and fixture repair are still required. +binds that same address at startup. Holding a competing listener on the released +port reproduced `AddrInUse` in the exact test after the original 31-second +pause. The fixture now holds its initially selected port and requests port zero +for the server, then uses `server.local_addr()` and asserts that the server +bound a different port. This also avoids a released-port race in the regression +itself. The 31-second pause, advertisement issue-time assertion, real repository +creation and production authority/fencing behavior are unchanged. + +The final exact test passed in 31.17 seconds; all four lease lifecycle tests +passed together in 35.24 seconds (100 other integration tests filtered out). +This is fixture verification, not a local full-workspace or performance pass. +Retained candidate-directory logs bind the failure and final results: + +| Test log | SHA-256 | +| --- | --- | +| `startup-port-race-reproduction.log`, failed | `19a99536e5c7a119fd0a675b64e7b171c25fb5918e34560c2e8f933061c24c7f` | +| `startup-port-race-atomic-fixed.log`, passed | `585d00b55b4f4ceedf05a216df20635c347d330e86305efb47f20cea0475e04d` | +| `lease-suite-after-fixture-fix.log`, four passed | `27c4a6d42108bd57ac8c006d48f5b39b28ceb473454bc80d41eb40adced8bf28` | ## Measure the requested operations From 878d1113bcbad0579ea97238815f7c4c306e9c45 Mon Sep 17 00:00:00 2001 From: forhappy Date: Wed, 30 Sep 2026 17:28:26 -0700 Subject: [PATCH 04/60] Measure validated Git pack first-byte and transfer timing Keep the read-only v0 no-haves transfer probe separate from the bound live campaign. Count actual pack-channel bytes, reject malformed and failed responses, retain incomplete receipts, and require strict stock-Git indexing and full graph verification. Record serial transfer scope without claiming clone throughput or cold recovery. All 52 harness tests pass on Python 3.12 and 3.14. --- .../2026-09-30-three-node-proxy.md | 70 ++++- scripts/measure_git_pack.py | 246 ++++++++++++++++++ scripts/test_measure_git_pack.py | 245 +++++++++++++++++ 3 files changed, 560 insertions(+), 1 deletion(-) create mode 100644 scripts/measure_git_pack.py create mode 100644 scripts/test_measure_git_pack.py diff --git a/docs/performance/2026-09-30-three-node-proxy.md b/docs/performance/2026-09-30-three-node-proxy.md index 44f3665..349b536 100644 --- a/docs/performance/2026-09-30-three-node-proxy.md +++ b/docs/performance/2026-09-30-three-node-proxy.md @@ -436,7 +436,75 @@ successful completions inside the offered-load window and successful throughput including drain. The driver observes the complete offered window; a quick final response must not inflate the rate by omitting the last inter-arrival interval. Git timings include client setup/validation where declared. Pack first-byte, -bytes/s, CPU and provider cost still need independent measurements. +bytes/s, CPU and provider cost still need independent live measurements. + +### Separate pack waiting from transfer and validation + +`scripts/measure_git_pack.py` adds a bounded, read-only v0 upload-pack probe. +It requests the critical receipt's exact main commit with no haves and +`side-band-64k`, using the [Git pack protocol](https://git-scm.com/docs/pack-protocol). +The measurement starts before a new HTTP connection/POST and records header +time, the first channel-1 pack byte, last pack byte and completion through final +flush/EOF. NAK and progress-channel bytes cannot count as pack data. First-byte +time is taken after reading one data byte, before waiting for the remainder of +that packet; it is a client-observed boundary, not kernel arrival or server CPU. + +```mermaid +sequenceDiagram + participant Probe + participant Proxy + participant Canopy + participant Git as Local validation + Note over Probe,Canopy: Identity and live fleet/artifact checks outside POST clock + Probe->>Proxy: POST want exact commit, no haves + Proxy->>Canopy: Forward unchanged + Canopy-->>Probe: Headers, NAK, progress (not pack data) + Canopy-->>Probe: First channel-1 pack byte + Note over Probe: Record first-pack-byte time before reading remainder + Canopy-->>Probe: Remaining pack, flush, EOF + Note over Probe: End POST clock; record bytes and transfer rate + Probe->>Git: index-pack --strict into empty object database + Probe->>Git: Exact wanted commit + fsck --strict --full + Note over Probe,Git: Validation timed separately; failure leaves incomplete receipt +``` + +Run only against the explicitly bound fleet, outside a scheduled load window: + +```sh +export CANOPY_GIT_TOKEN='' +python3 -B scripts/measure_git_pack.py \ + --fleet-dir /dedicated-volume/candidate-fleet \ + --receipt /dedicated-volume/critical-candidate.json \ + --output-dir /dedicated-volume/new-pack-measurement \ + --samples 3 --timeout 30 --max-pack-bytes 268435456 +``` + +The output binds the fleet, production executable, critical receipt and source +digests. Each serial sample uses a new connection and retains its pack and +validation database. HTTP errors, redirects, wrong content types, compression, +fatal sidebands, missing/truncated/oversized data, invalid pack checksums and +missing wanted commits cannot produce a completed sample. There are no retries +or automatic warmup. Failed sample receipts remain incomplete, not low-latency +successes. + +Eight regression tests include a real stock `git upload-pack` stream over +fragmented loopback HTTP, complete serial measurement receipts, strict +indexing/fsck and corrupt/missing-graph pack rejection. The complete 52-test +Python harness passed on both Python 3.12 and 3.14. Retained final logs under +`canopy-e07670e-candidate-W0SFiGMz` bind those results: + +| Final suite log | SHA-256 | +| --- | --- | +| `python312-pack-probe-final-suite.log` | `9e6e2636cef36997c027c275ae4f32baf6ec4d6b72717624b1784eb1ddc4b725` | +| `python314-pack-probe-final-suite.log` | `344cc7ad10890a9c8bca494609c2f79c72cd3339b6f0b61c1644f1fea16e2d73` | + +This probe has not yet produced candidate RustFS measurements. Its transfer +rate is pack bytes divided by the entire POST duration, not Ethernet bytes/s or +steady-state capacity. Serial repetitions can warm server state; neither the +identity check nor a new HTTP connection establishes cold ownership. Discovery, +local validation, v2 negotiation, stock clone performance, native Git child CPU +and provider request-cost accounting require their own evidence. The running +candidate seed and its bound harness are unchanged. The campaign samples each node, the proxy launcher and its own driver once per second with `ps`: RSS, cumulative process CPU seconds and raw CPU percentage. diff --git a/scripts/measure_git_pack.py b/scripts/measure_git_pack.py new file mode 100644 index 0000000..1201386 --- /dev/null +++ b/scripts/measure_git_pack.py @@ -0,0 +1,246 @@ +#!/usr/bin/env python3 +"""Measure a read-only v0, no-haves, side-band pack POST through the ingress. + +This is not a clone benchmark: discovery, identity checks and local Git validation +are outside the POST clock. No retries, remote mutations or automatic warmup. +First-pack-byte excludes HTTP headers, NAK and progress sidebands. Local files +and a failed receipt remain available for inspection if validation fails. +""" + +import argparse +import http.client +import json +import os +from pathlib import Path +import re +import subprocess +import time +from urllib.parse import urlsplit +import uuid + +import benchmark_repositories as benchmark +import benchmark_three_node_campaign as campaign + + +def require(condition, message): + if not condition: + raise ValueError(message) + + +def packet(payload): + require(0 < len(payload) <= 65516, "invalid packet payload size") + return f"{len(payload) + 4:04x}".encode() + payload + + +def request_body(oid): + require(isinstance(oid, str) and re.fullmatch(r"[0-9a-f]{40}", oid), + "require a SHA-1 commit object ID") + return packet(f"want {oid} side-band-64k ofs-delta\n".encode()) + b"0000" + packet(b"done\n") + + +def stream_pack(response, destination, started, *, max_pack_bytes, timeout, clock=time.monotonic): + """Consume bounded pkt-lines; time channel-1's first byte before its remainder. + + read(1) at that boundary avoids waiting for an entire sideband packet. Times + are client-observed read boundaries, not kernel arrival or server CPU time. + """ + wire_bytes = pack_bytes = progress_bytes = 0 + first_pack = last_pack = None + signature = bytearray() + nak_seen = False + + def read(size): + nonlocal wire_bytes + require(clock() - started < timeout, "pack POST exceeded total deadline") + value = response.read(size) + require(len(value) == size, "truncated pkt-line") + wire_bytes += len(value) + require(wire_bytes <= max_pack_bytes + 16 * 1024 * 1024, + "response framing/progress exceeds wire bound") + require(clock() - started < timeout, "pack POST exceeded total deadline") + return value + + while True: + header = read(4) + require(re.fullmatch(b"[0-9a-fA-F]{4}", header), "invalid pkt-line header") + size = int(header, 16) + if size == 0: + break + require(5 <= size <= 65520, "invalid v0 sideband packet size") + channel = read(1) + remaining = size - 5 + if not nak_seen: + require(channel + read(remaining) == b"NAK\n", "expected no-haves NAK") + nak_seen = True + continue + require(channel in (b"\x01", b"\x02", b"\x03"), "invalid sideband channel") + if channel == b"\x03": + # Do not put server-controlled text (possibly secrets) in reports. + raise ValueError("remote fatal sideband") + if channel == b"\x02": + progress_bytes += len(read(remaining)) + continue + require(remaining > 0, "empty pack-data packet") + require(pack_bytes + remaining <= max_pack_bytes, "pack exceeds declared bound") + first = read(1) + if first_pack is None: + first_pack = clock() + chunks = [first, read(remaining - 1)] if remaining > 1 else [first] + last_pack = clock() + for chunk in chunks: + signature.extend(chunk[:max(0, 4 - len(signature))]) + destination.write(chunk) + pack_bytes += len(chunk) + require(nak_seen and first_pack is not None and signature == b"PACK", "missing valid pack prefix") + require(response.read(1) == b"", "trailing bytes after final flush") + finished = clock() + require(finished - started < timeout, "pack POST exceeded total deadline") + seconds = finished - started + require(seconds > 0 and started <= first_pack <= last_pack <= finished, + "invalid measurement clock") + return {"started_monotonic": started, "first_pack_byte_monotonic": first_pack, + "last_pack_byte_monotonic": last_pack, "completed_monotonic": finished, + "response_body_bytes": wire_bytes, "pack_bytes": pack_bytes, + "progress_payload_bytes": progress_bytes, + "post_first_pack_byte_ms": (first_pack - started) * 1000, + "post_last_pack_byte_ms": (last_pack - started) * 1000, + "post_complete_ms": seconds * 1000, + "pack_bytes_per_post_second": pack_bytes / seconds, + "timing_scope": "fresh HTTP connection + POST through final flush/EOF; no discovery or local validation; client read boundaries"} + + +def download(base_url, entry, oid, token, pack_path, *, timeout, max_pack_bytes): + # URL validation is also used by the bound benchmark harness. + benchmark.Client(base_url, token, timeout).close() + url = urlsplit(base_url) + require(all(isinstance(entry.get(key), str) and re.fullmatch(r"[a-z0-9_-]{1,64}", entry[key]) + for key in ("owner", "name")), "invalid repository path") + body = request_body(oid) + connection_type = http.client.HTTPSConnection if url.scheme == "https" else http.client.HTTPConnection + connection = connection_type(url.hostname, url.port, timeout=timeout) + request_id = str(uuid.uuid4()) + started = time.monotonic() + try: + connection.request("POST", url.path.rstrip("/") + f"/{entry['owner']}/{entry['name']}.git/git-upload-pack", + body=body, headers={"Authorization": f"Bearer {token}", + "Content-Type": "application/x-git-upload-pack-request", + "Accept": "application/x-git-upload-pack-result", + "Accept-Encoding": "identity", "Connection": "close", + "X-Request-ID": request_id}) + response = connection.getresponse() + headers_ms = (time.monotonic() - started) * 1000 + require(response.status == 200, f"pack POST HTTP {response.status}") + require(response.getheader("Content-Type", "").split(";")[0].strip() + == "application/x-git-upload-pack-result", "unexpected pack content type") + require(response.getheader("Content-Encoding", "identity") == "identity", + "encoded pack response not supported") + with pack_path.open("xb") as output: + result = stream_pack(response, output, started, max_pack_bytes=max_pack_bytes, timeout=timeout) + return {**result, "http_headers_ms": headers_ms, "request_id": request_id, + "pack_sha256": benchmark.file_sha256(pack_path)} + finally: + connection.close() + + +def validate_pack(pack_path, oid, destination, token): + """Validate the pack and the wanted commit's full closure in an empty ODB.""" + destination.mkdir(parents=True, exist_ok=False) + benchmark.git("init", "--bare", str(destination), cwd=destination.parent, token=token) + environment = {key: value for key, value in os.environ.items() + if not key.startswith(("GIT_", "AWS_", "RUSTFS_", "CANOPY_"))} + environment.update(GIT_CONFIG_NOSYSTEM="1", GIT_CONFIG_GLOBAL=os.devnull, + GIT_TERMINAL_PROMPT="0") + with pack_path.open("rb") as source: + result = subprocess.run(["git", "index-pack", "--strict", "--stdin"], + cwd=destination, env=environment, stdin=source, + capture_output=True, timeout=120, check=False) + require(result.returncode == 0, "strict pack indexing failed") + require(benchmark.git("cat-file", "-t", oid, cwd=destination, token=token) == "commit", + "wanted object is not a commit") + benchmark.git("update-ref", "refs/heads/measured", oid, cwd=destination, token=token) + benchmark.git("symbolic-ref", "HEAD", "refs/heads/measured", cwd=destination, token=token) + benchmark.git("fsck", "--strict", "--full", cwd=destination, token=token) + return {"index_pack_strict": True, "fsck_strict_full": True, "wanted_commit": oid} + + +def measure(args, token): + receipt = campaign.read_json(args.receipt) + require(receipt.get("complete") is True and receipt.get("error") is None, + "require a complete critical Git receipt") + entries = receipt.get("repositories") + require(isinstance(entries, list) and len(entries) == 2, "require two critical repositories") + entry = entries[args.repository_index] + require(benchmark.canonical_repository_uuid(entry.get("repository_id")), "invalid repository UUID") + oid = receipt.get("refs", {}).get("refs/heads/main") + request_body(oid) + ready = campaign.read_json(args.fleet_dir / "ready.json") + ready = campaign.validate_fleet(args.fleet_dir, + {"node_active_limit": ready.get("max_active_repositories_per_node")}) + require(receipt.get("base_url") == ready["proxy_url"] + and receipt.get("binary_sha256") == ready["binary_sha256"], "critical receipt/fleet binding differs") + args.output_dir.mkdir(parents=True, exist_ok=False) + result = {"version": 1, "complete": False, "error": None, "repository": entry, + "wanted_commit": oid, "protocol": 0, "haves": 0, + "driver_sha256": benchmark.file_sha256(Path(__file__)), + "git_driver_sha256": benchmark.file_sha256(Path(benchmark.__file__)), + "fleet_validator_sha256": benchmark.file_sha256(Path(campaign.__file__)), + "receipt_sha256": benchmark.file_sha256(args.receipt), + "fleet_ready_sha256": benchmark.file_sha256(args.fleet_dir / "ready.json"), + "binary_sha256": ready["binary_sha256"], "samples": [], + "declared": {"samples": args.samples, "timeout_seconds": args.timeout, + "max_pack_bytes": args.max_pack_bytes}, + "qualification": "serial no-haves v0 transfer observations; not stock clone, scheduled load, cold ownership, recovery, CPU or provider cost proof"} + path = args.output_dir / "measurement.json" + benchmark.save(path, result) + client = benchmark.Client(ready["proxy_url"], token, args.timeout) + try: + result["validation_git_version"] = benchmark.git("--version", cwd=args.output_dir, token=token) + status, body = client.request(f"/api/repositories/{entry['name']}", request_id=str(uuid.uuid4())) + require(status == 200 and json.loads(body).get("repository_id") == entry["repository_id"], + "repository identity differs") + for index in range(args.samples): + # Revalidate live identity/binary/script binding before each sample. + campaign.validate_fleet(args.fleet_dir, + {"node_active_limit": ready["max_active_repositories_per_node"]}) + pack_path = args.output_dir / f"sample-{index}.pack" + sample = {"index": index, "complete": False, "error": None} + result["samples"].append(sample) + benchmark.save(path, result) + sample.update(download(ready["proxy_url"], entry, oid, token, pack_path, + timeout=args.timeout, max_pack_bytes=args.max_pack_bytes)) + checking = time.monotonic() + sample["validation"] = validate_pack(pack_path, oid, args.output_dir / f"sample-{index}.git", token) + sample["validation_ms"] = (time.monotonic() - checking) * 1000 + sample["complete"] = True + benchmark.save(path, result) + result["complete"] = True + except BaseException as error: + result["error"] = type(error).__name__ + if result["samples"] and not result["samples"][-1]["complete"]: + result["samples"][-1]["error"] = type(error).__name__ + raise + finally: + client.close() + benchmark.save(path, result) + return result + + +def main(): + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("--fleet-dir", type=Path, required=True) + parser.add_argument("--receipt", type=Path, required=True) + parser.add_argument("--output-dir", type=Path, required=True) + parser.add_argument("--repository-index", type=int, choices=(0, 1), default=0) + parser.add_argument("--samples", type=int, default=3) + parser.add_argument("--timeout", type=float, default=30) + parser.add_argument("--max-pack-bytes", type=int, default=256 * 1024 * 1024) + args = parser.parse_args() + require(1 <= args.samples <= 100 and 0 < args.timeout <= 300 + and 32 <= args.max_pack_bytes <= 1024 * 1024 * 1024, "invalid measurement bounds") + token = os.environ.get("CANOPY_GIT_TOKEN") + require(bool(token), "set CANOPY_GIT_TOKEN") + measure(args, token) + + +if __name__ == "__main__": + main() diff --git a/scripts/test_measure_git_pack.py b/scripts/test_measure_git_pack.py new file mode 100644 index 0000000..0954e8a --- /dev/null +++ b/scripts/test_measure_git_pack.py @@ -0,0 +1,245 @@ +"""Wire parsing and stock-Git validation; not a live RustFS capacity test.""" + +import io +import json +import os +from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer +from pathlib import Path +import subprocess +import tempfile +import threading +from types import SimpleNamespace +import unittest +from unittest.mock import patch +import uuid + +import benchmark_repositories as benchmark +import measure_git_pack as probe + + +class PackTests(unittest.TestCase): + def parse(self, wire, *, bound=4096): + destination = io.BytesIO() + clock = iter(index / 1000 for index in range(1000)) + result = probe.stream_pack(io.BytesIO(wire), destination, 0, + max_pack_bytes=bound, timeout=30, + clock=lambda: next(clock)) + return result, destination.getvalue() + + def wire(self, data=b"PACKpayload"): + return probe.packet(b"NAK\n") + probe.packet(b"\x02progress\n") + probe.packet(b"\x01" + data) + b"0000" + + def test_first_byte_precedes_packet_remainder_and_ignores_progress(self): + destination = io.BytesIO() + observations = [] + class Reader(io.BytesIO): + def read(self, size): + value = super().read(size) + observations.append((size, value, len(destination.getvalue()))) + return value + ticks = iter(index / 1000 for index in range(1000)) + result = probe.stream_pack(Reader(self.wire()), destination, 0, + max_pack_bytes=4096, timeout=30, + clock=lambda: next(ticks)) + first = next(index for index, (_, value, _) in enumerate(observations) if value == b"P") + self.assertEqual(observations[first][0], 1) + self.assertEqual(observations[first + 1][1], b"ACKpayload") + self.assertEqual(result["pack_bytes"], len(b"PACKpayload")) + self.assertEqual(result["progress_payload_bytes"], len(b"progress\n")) + self.assertEqual(result["response_body_bytes"], len(self.wire())) + self.assertLess(result["post_first_pack_byte_ms"], result["post_last_pack_byte_ms"]) + self.assertEqual(destination.getvalue(), b"PACKpayload") + + def test_split_pack_signature_and_multiple_data_packets(self): + wire = probe.packet(b"NAK\n") + b"".join(probe.packet(b"\x01" + value) + for value in (b"P", b"AC", b"Kpayload")) + b"0000" + result, body = self.parse(wire) + self.assertEqual(body, b"PACKpayload") + self.assertEqual(result["pack_bytes"], len(body)) + + def test_invalid_wire_never_becomes_a_successful_sample(self): + invalid = [b"", b"zzzz", b"0001", b"0004", b"ffff", self.wire()[:-1], + self.wire() + b"extra", probe.packet(b"NAK\n") + b"0000", + probe.packet(b"NAK\n") + probe.packet(b"\x02progress") + b"0000", + probe.packet(b"NAK\n") + probe.packet(b"\x03secret remote error") + b"0000", + probe.packet(b"ACK " + b"a" * 40 + b"\n") + self.wire(), + probe.packet(b"NAK\n") + probe.packet(b"\x04PACKbad") + b"0000", + self.wire(b"not a pack"), probe.packet(b"NAK\n") + probe.packet(b"\x01") + b"0000"] + for wire in invalid: + with self.subTest(wire=wire), self.assertRaises(ValueError): + self.parse(wire) + with self.assertRaisesRegex(ValueError, "declared bound"): + self.parse(self.wire(), bound=4) + + def test_deadline_is_not_reset_by_progress_packets(self): + with self.assertRaisesRegex(ValueError, "total deadline"): + probe.stream_pack(io.BytesIO(self.wire()), io.BytesIO(), 0, + max_pack_bytes=4096, timeout=1, clock=lambda: 2) + + def test_want_and_repository_path_validation(self): + oid = "a" * 40 + self.assertEqual(probe.request_body(oid), probe.packet( + f"want {oid} side-band-64k ofs-delta\n".encode()) + b"0000" + probe.packet(b"done\n")) + for invalid in (None, "a" * 39, "a" * 64, "g" * 40, "a" * 40 + "\ndone"): + with self.assertRaises(ValueError): + probe.request_body(invalid) + with tempfile.TemporaryDirectory() as directory, \ + patch.object(probe.http.client, "HTTPConnection", side_effect=AssertionError("unexpected connection")): + with self.assertRaises(ValueError): + probe.download("http://127.0.0.1:1", {"owner": "canopy", "name": "../escape"}, + oid, "token", Path(directory) / "unused", timeout=30, max_pack_bytes=4096) + + def test_http_errors_redirects_and_compression_are_not_pack_success(self): + for status, content_type, encoding in ((503, "application/x-git-upload-pack-result", "identity"), + (302, "application/x-git-upload-pack-result", "identity"), + (200, "text/html", "identity"), + (200, "application/x-git-upload-pack-result", "gzip")): + response = SimpleNamespace(status=status, getheader=lambda key, default: { + "Content-Type": content_type, "Content-Encoding": encoding}.get(key, default)) + connection = SimpleNamespace(request=lambda *args, **kwargs: None, + getresponse=lambda: response, close=lambda: None) + with self.subTest(status=status, content_type=content_type, encoding=encoding), \ + tempfile.TemporaryDirectory() as directory, \ + patch.object(probe.http.client, "HTTPConnection", return_value=connection), \ + patch.object(probe, "stream_pack", side_effect=AssertionError("unexpected pack parsing")): + path = Path(directory) / "unused.pack" + with self.assertRaises(ValueError): + probe.download("http://127.0.0.1:1", {"owner": "canopy", "name": "source"}, + "a" * 40, "token", path, timeout=30, max_pack_bytes=4096) + self.assertFalse(path.exists()) + + def test_failed_sample_is_retained_without_token_or_error_text(self): + with tempfile.TemporaryDirectory() as directory: + root = Path(directory) + receipt = root / "critical.json" + receipt.write_text(json.dumps({"complete": True, "error": None, + "repositories": [{"repository_id": str(uuid.uuid4()), "owner": "canopy", "name": name} + for name in ("source", "mirror")], + "refs": {"refs/heads/main": "a" * 40}, "base_url": "http://127.0.0.1:1", + "binary_sha256": "b" * 64})) + fleet = root / "fleet" + fleet.mkdir() + ready = {"max_active_repositories_per_node": 100, + "proxy_url": "http://127.0.0.1:1", "binary_sha256": "b" * 64} + (fleet / "ready.json").write_text(json.dumps(ready)) + args = SimpleNamespace(receipt=receipt, fleet_dir=fleet, repository_index=0, + output_dir=root / "output", samples=3, timeout=30, max_pack_bytes=4096) + identity = json.loads(receipt.read_text())["repositories"][0]["repository_id"] + with patch.object(probe.campaign, "validate_fleet", return_value=ready), \ + patch.object(benchmark.Client, "request", return_value=(200, json.dumps( + {"repository_id": identity}).encode())), \ + patch.object(probe, "download", side_effect=ValueError("private-token-sensitive")) as download: + with self.assertRaises(ValueError): + probe.measure(args, "private-token-sensitive") + self.assertEqual(download.call_count, 1) + output = (args.output_dir / "measurement.json").read_text() + record = json.loads(output) + self.assertFalse(record["complete"]) + self.assertEqual(record["error"], "ValueError") + self.assertEqual(record["samples"], [{"index": 0, "complete": False, "error": "ValueError"}]) + self.assertNotIn("private-token-sensitive", output) + + def test_real_git_pack_over_fragmented_http_and_strict_validation(self): + with tempfile.TemporaryDirectory() as directory: + root = Path(directory) + source = root / "source" + token = "local-fixture-token" + benchmark.git("init", "-b", "main", str(source), cwd=root, token=token) + (source / "fixture.txt").write_text("complete graph\n") + benchmark.git("add", ".", cwd=source, token=token) + benchmark.git("-c", "user.name=Fixture", "-c", "user.email=fixture@example.invalid", + "commit", "-m", "fixture", cwd=source, token=token) + oid = benchmark.git("rev-parse", "HEAD", cwd=source, token=token) + body = probe.request_body(oid) + environment = {"PATH": os.environ["PATH"], "GIT_CONFIG_NOSYSTEM": "1", + "GIT_CONFIG_GLOBAL": os.devnull} + wire = subprocess.run(["git", "upload-pack", "--stateless-rpc", str(source)], + input=body, capture_output=True, check=True, env=environment).stdout + failures = [] + identity = str(uuid.uuid4()) + class Handler(BaseHTTPRequestHandler): + def log_message(self, *_): + pass + def do_GET(self): + if self.path != "/api/repositories/source": + self.send_error(404) + return + content = json.dumps({"repository_id": identity}).encode() + self.send_response(200) + self.send_header("Content-Length", str(len(content))) + self.end_headers() + self.wfile.write(content) + def do_POST(self): + try: + self.assertions() + self.send_response(200) + self.send_header("Content-Type", "application/x-git-upload-pack-result") + self.send_header("Content-Length", str(len(wire))) + self.end_headers() + for index in range(0, len(wire), 3): + self.wfile.write(wire[index:index + 3]) + self.wfile.flush() + except BaseException as error: + failures.append(type(error).__name__) + def assertions(self): + if self.path != "/canopy/source.git/git-upload-pack" \ + or self.headers["Authorization"] != "Bearer " + token \ + or self.rfile.read(int(self.headers["Content-Length"])) != body: + raise AssertionError("request mismatch") + server = ThreadingHTTPServer(("127.0.0.1", 0), Handler) + worker = threading.Thread(target=server.serve_forever, daemon=True) + worker.start() + try: + path = root / "measured.pack" + result = probe.download(f"http://127.0.0.1:{server.server_port}", + {"owner": "canopy", "name": "source"}, oid, token, path, + timeout=30, max_pack_bytes=4096) + self.assertEqual(failures, []) + self.assertGreater(result["pack_bytes"], 32) + self.assertLessEqual(result["http_headers_ms"], result["post_first_pack_byte_ms"]) + verified = probe.validate_pack(path, oid, root / "verified.git", token) + self.assertTrue(verified["fsck_strict_full"]) + self.assertEqual(benchmark.git("show", "measured:fixture.txt", + cwd=root / "verified.git", token=token), "complete graph") + corrupt = root / "corrupt.pack" + data = bytearray(path.read_bytes()) + data[-1] ^= 1 + corrupt.write_bytes(data) + with self.assertRaisesRegex(ValueError, "indexing failed"): + probe.validate_pack(corrupt, oid, root / "corrupt.git", token) + with self.assertRaisesRegex(RuntimeError, "cat-file"): + probe.validate_pack(path, "0" * 40, root / "wrong-tip.git", token) + incomplete_graph = root / "incomplete-graph.pack" + incomplete_graph.write_bytes(subprocess.run( + ["git", "pack-objects", "--stdout"], cwd=source, env=environment, + input=(oid + "\n").encode(), capture_output=True, check=True).stdout) + with self.assertRaises((ValueError, RuntimeError)): + probe.validate_pack(incomplete_graph, oid, root / "incomplete-graph.git", token) + ready = {"max_active_repositories_per_node": 100, + "proxy_url": f"http://127.0.0.1:{server.server_port}", "binary_sha256": "b" * 64} + fleet = root / "fleet" + fleet.mkdir() + (fleet / "ready.json").write_text(json.dumps(ready)) + receipt = root / "critical.json" + receipt.write_text(json.dumps({"complete": True, "error": None, + "repositories": [{"repository_id": identity, "owner": "canopy", "name": "source"}, + {"repository_id": str(uuid.uuid4()), "owner": "canopy", "name": "mirror"}], + "refs": {"refs/heads/main": oid}, "base_url": ready["proxy_url"], + "binary_sha256": ready["binary_sha256"]})) + args = SimpleNamespace(receipt=receipt, fleet_dir=fleet, repository_index=0, + output_dir=root / "serial", samples=2, timeout=30, max_pack_bytes=4096) + with patch.object(probe.campaign, "validate_fleet", return_value=ready): + measured = probe.measure(args, token) + self.assertTrue(measured["complete"]) + self.assertTrue(all(sample["complete"] for sample in measured["samples"])) + self.assertEqual(len(measured["samples"]), 2) + self.assertEqual(json.loads((args.output_dir / "measurement.json").read_text()), measured) + self.assertEqual(failures, []) + finally: + server.shutdown() + worker.join(timeout=5) + server.server_close() + + +if __name__ == "__main__": + unittest.main() From 6a58063f7568c49299fa1fbc463f3e2afd4c1ece Mon Sep 17 00:00:00 2001 From: forhappy Date: Wed, 30 Sep 2026 17:31:02 -0700 Subject: [PATCH 05/60] Record live candidate pack probe under explicit seed load --- .../2026-09-30-three-node-proxy.md | 26 +++++++++++++++++-- 1 file changed, 24 insertions(+), 2 deletions(-) diff --git a/docs/performance/2026-09-30-three-node-proxy.md b/docs/performance/2026-09-30-three-node-proxy.md index 349b536..a8b294c 100644 --- a/docs/performance/2026-09-30-three-node-proxy.md +++ b/docs/performance/2026-09-30-three-node-proxy.md @@ -498,8 +498,30 @@ Python harness passed on both Python 3.12 and 3.14. Retained final logs under | `python312-pack-probe-final-suite.log` | `9e6e2636cef36997c027c275ae4f32baf6ec4d6b72717624b1784eb1ddc4b725` | | `python314-pack-probe-final-suite.log` | `344cc7ad10890a9c8bca494609c2f79c72cd3339b6f0b61c1644f1fea16e2d73` | -This probe has not yet produced candidate RustFS measurements. Its transfer -rate is pack bytes divided by the entire POST duration, not Ethernet bytes/s or +The committed probe (`878d111`) subsequently passed three serial samples +through the candidate proxy/RustFS while the 10,000-repository seed was active. +Each downloaded pack contained 1,153 bytes and passed strict indexing, exact +wanted-commit validation and full strict fsck with Apple Git 2.50.1. These are +setup-load functional/timing observations, not scheduled load or matched +baseline comparisons. The compressed critical fixture is too small to measure +bulk transfer capacity. + +| Serial sample | Headers (ms) | First pack byte (ms) | Complete POST (ms) | Local validation (ms, outside POST) | +| --- | ---: | ---: | ---: | ---: | +| 0 | 227.072 | 255.891 | 257.780 | 166.786 | +| 1 | 490.265 | 507.593 | 509.568 | 144.591 | +| 2 | 121.930 | 139.526 | 140.933 | 150.131 | + +The candidate-directory artifact `pack-probe-during-seed/measurement.json` +has SHA-256 `59801101b210e16254f5f6ae2d6460eef954ab402326db5369d733ad4bb1dcf3`. +It binds the immutable production binary, critical receipt, fleet and probe +digest `82a66363dd2242719fdc5a259d74bc051f03ad6168538064fe404e7f1609e512`. +Raw monotonic boundaries, request IDs, pack digests and declared bounds are +retained. The overlapping seed is an explicit setup-load intervention, not +background-free performance evidence; no seed result is promoted to scheduled +creation throughput. + +The reported transfer rate is pack bytes divided by the entire POST duration, not Ethernet bytes/s or steady-state capacity. Serial repetitions can warm server state; neither the identity check nor a new HTTP connection establishes cold ownership. Discovery, local validation, v2 negotiation, stock clone performance, native Git child CPU From 6199a73d2a23ec99a809f625ee5d41331588b7a2 Mon Sep 17 00:00:00 2001 From: forhappy Date: Wed, 30 Sep 2026 17:39:57 -0700 Subject: [PATCH 06/60] Observe kernel self and reaped-child CPU with calibrated units Keep the immutable running fleet and campaign unchanged. Read Darwin rusage v2 with its Mach timebase and Linux proc stat ticks; preserve raw counters and process identities and reject discontinuities. Verify native CPU-clock calibration and delayed child rollup, while explicitly excluding live-tree totals and Git-only attribution. All 59 harness tests pass on Python 3.12 and 3.14. --- .../2026-09-30-three-node-proxy.md | 45 +++++ scripts/fleet_cpu.py | 177 ++++++++++++++++++ scripts/test_fleet_cpu.py | 142 ++++++++++++++ 3 files changed, 364 insertions(+) create mode 100644 scripts/fleet_cpu.py create mode 100644 scripts/test_fleet_cpu.py diff --git a/docs/performance/2026-09-30-three-node-proxy.md b/docs/performance/2026-09-30-three-node-proxy.md index a8b294c..edd8bbe 100644 --- a/docs/performance/2026-09-30-three-node-proxy.md +++ b/docs/performance/2026-09-30-three-node-proxy.md @@ -542,6 +542,51 @@ the running campaign's bound source is not changed mid-measurement. These process samples also exclude native Git child CPU, so they cannot support a complete server CPU-per-GiB claim. +### Observe kernel self and reaped-child CPU separately + +`scripts/fleet_cpu.py` adds a read-only observer without changing the bound +campaign or production executable. macOS reads `proc_pid_rusage` v2; Linux reads +the per-process `utime`, `stime`, `cutime` and `cstime` counters. Each record keeps +raw counters, time scale, RSS and a process-start identity. A changed identity, +decreased counter, missing process or invalid interval fails the observation; +earlier samples and an incomplete receipt remain retained. + +On this ARM64 Mac, the Mach timebase is 125/3. A 0.150020-second process-CPU +calibration produced 3,602,003 raw units: scaling yielded 0.150083 seconds; +dividing by 1e9 alone would incorrectly report 0.003602 seconds. The reader +queries the platform timebase rather than hard-coding this host's ratio. +[XNU's task accounting](https://github.com/apple-oss-distributions/xnu/blob/main/osfmk/kern/task.c) +uses Mach-time counters. [Linux proc stat](https://man7.org/linux/man-pages/man5/proc_pid_stat.5.html) +uses clock ticks, scaled with `SC_CLK_TCK`. + +Child counters cover rolled-up usage, not instantaneous live-tree CPU. +[XNU's reap path](https://github.com/apple-oss-distributions/xnu/blob/main/bsd/kern/kern_exit.c) +adds child usage to its parent; Linux's child fields cover waited-for children. +A native regression burns child CPU, holds that child alive, proves the parent's +child counter is unchanged, then waits and proves the counter increased. +Seven tests also check SDK layout, Linux names with parentheses, unit +calibration, identity/clock/counter discontinuities and retained failures. +The full 59-test Python harness passed on Python 3.12 and 3.14. Candidate-volume +logs `python312-kernel-cpu-suite.log` and `python314-kernel-cpu-suite.log` have +SHA-256 `9d316452976529027fe905cdf349710ae4966638f8b87c6fe82a634df809864c` +and `8abca5145c727e3391d0b9949a6fd61b18e696219526f276b28bc5d005e74649`. + +```sh +python3 -B scripts/fleet_cpu.py \ + --fleet-dir /dedicated-volume/candidate-fleet \ + --output-dir /dedicated-volume/new-kernel-cpu-observation \ + --samples 120 --interval 1 +``` + +The output binds the fleet, executable and observer/validator digests. Intervals +separate self CPU percentage from reaped-child CPU seconds. Short-lived reaped +children need not be caught by a periodic process-list sample, but live and +unreaped children remain excluded. A child started before a load window may be +charged when reaped inside it. These counters include all charged children, not +only Git, and do not close operation-specific CPU/GiB accounting without +independently verified child boundaries. Historical `ps` results are unchanged; +no child CPU is retroactively imputed to the failed baseline. + A separate read-only observer records dedicated RustFS container CPU/memory samples about every six seconds, beginning during full-corpus preflight. Its binding includes the observer digest, campaign PID, container ID and start time; diff --git a/scripts/fleet_cpu.py b/scripts/fleet_cpu.py new file mode 100644 index 0000000..9c55d88 --- /dev/null +++ b/scripts/fleet_cpu.py @@ -0,0 +1,177 @@ +#!/usr/bin/env python3 +"""Read kernel CPU counters for the three bound Canopy processes. + +Self CPU and reaped-child CPU are separate. Child counters are charged at reap, +not continuously while children run: this does NOT establish total live process +tree CPU, Git-only attribution, or operation-window CPU without closed child +boundaries. No process attachment, signals, sampling of payloads, or retries. +""" + +import argparse +import ctypes +from datetime import datetime, timezone +import json +import math +import os +from pathlib import Path +import sys +import time + +import benchmark_repositories as benchmark +import benchmark_three_node_campaign as campaign + + +COUNTERS = ("user", "system", "reaped_child_user", "reaped_child_system") + + +class DarwinUsage(ctypes.Structure): + # rusage_info_v2 from the platform SDK's sys/resource.h. Time fields use + # Mach absolute units; calibrate through mach_timebase_info, not /1e9 alone. + _fields_ = [("uuid", ctypes.c_ubyte * 16)] + [(name, ctypes.c_uint64) for name in ( + "user_time", "system_time", "pkg_idle_wkups", "interrupt_wkups", "pageins", + "wired_size", "resident_size", "phys_footprint", "proc_start_abstime", + "proc_exit_abstime", "child_user_time", "child_system_time", + "child_pkg_idle_wkups", "child_interrupt_wkups", "child_pageins", + "child_elapsed_abstime", "diskio_bytesread", "diskio_byteswritten")] + + +class Timebase(ctypes.Structure): + _fields_ = [("numer", ctypes.c_uint32), ("denom", ctypes.c_uint32)] + + +def require(condition, message): + if not condition: + raise ValueError(message) + + +def parse_linux_stat(text, pid, ticks_per_second, page_size): + # comm may contain whitespace and ')'; fields after its last ')' are fixed. + head, marker, tail = text.rpartition(")") + require(marker and head.split("(", 1)[0].strip() == str(pid), "invalid proc stat PID") + fields = tail.split() + require(len(fields) >= 22 and ticks_per_second > 0 and page_size > 0, "invalid proc stat layout") + require(fields[0] not in ("Z", "X", "x"), "process is not live") + values = [int(fields[index]) for index in (11, 12, 13, 14, 19, 21)] + require(all(value >= 0 for value in values) and values[4] > 0, "invalid proc stat counters") + return {"pid": pid, "identity": {"start_ticks": values[4]}, + "backend": "linux-proc-stat", "seconds_per_unit": 1 / ticks_per_second, + "scale": {"clock_ticks_per_second": ticks_per_second}, + "cpu_units": dict(zip(COUNTERS, values[:4])), "rss_bytes": values[5] * page_size} + + +class Reader: + def __init__(self): + self.platform = sys.platform + if self.platform == "darwin": + self.library = ctypes.CDLL("/usr/lib/libproc.dylib", use_errno=True) + self.library.proc_pid_rusage.argtypes = [ctypes.c_int, ctypes.c_int, ctypes.c_void_p] + self.library.proc_pid_rusage.restype = ctypes.c_int + system = ctypes.CDLL("/usr/lib/libSystem.B.dylib") + system.mach_timebase_info.argtypes = [ctypes.POINTER(Timebase)] + system.mach_timebase_info.restype = ctypes.c_int + base = Timebase() + require(system.mach_timebase_info(ctypes.byref(base)) == 0 + and base.numer > 0 and base.denom > 0, "invalid Mach timebase") + self.scale = {"mach_numer": base.numer, "mach_denom": base.denom} + self.seconds_per_unit = base.numer / base.denom / 1_000_000_000 + elif self.platform == "linux": + self.ticks = os.sysconf("SC_CLK_TCK") + self.page_size = os.sysconf("SC_PAGE_SIZE") + else: + raise RuntimeError("kernel CPU reader requires macOS or Linux") + + def snapshot(self, pid): + require(isinstance(pid, int) and not isinstance(pid, bool) and pid > 0, "invalid PID") + started = time.monotonic() + if self.platform == "darwin": + usage = DarwinUsage() + if self.library.proc_pid_rusage(pid, 2, ctypes.byref(usage)) != 0: + raise OSError(ctypes.get_errno(), "kernel CPU snapshot failed") + require(usage.proc_start_abstime > 0 and usage.proc_exit_abstime == 0, + "process is not live") + result = {"pid": pid, "identity": {"start_abstime": usage.proc_start_abstime, + "executable_uuid": bytes(usage.uuid).hex()}, + "backend": "darwin-proc-pid-rusage-v2", "scale": self.scale, + "seconds_per_unit": self.seconds_per_unit, + "cpu_units": dict(zip(COUNTERS, (usage.user_time, usage.system_time, + usage.child_user_time, usage.child_system_time))), + "rss_bytes": usage.resident_size} + else: + result = parse_linux_stat(Path(f"/proc/{pid}/stat").read_text(), pid, self.ticks, self.page_size) + finished = time.monotonic() + return {**result, "read_started_monotonic": started, "read_finished_monotonic": finished, + "captured_monotonic": (started + finished) / 2} + + +def delta(before, after): + require(all(before[key] == after[key] for key in ("pid", "identity", "backend", "scale", "seconds_per_unit")), + "process identity or CPU scale changed") + seconds = after["captured_monotonic"] - before["captured_monotonic"] + require(math.isfinite(seconds) and seconds > 0, "invalid CPU observation interval") + units = {name: after["cpu_units"][name] - before["cpu_units"][name] for name in COUNTERS} + require(all(value >= 0 for value in units.values()), "CPU counter decreased") + cpu = {name: value * after["seconds_per_unit"] for name, value in units.items()} + return {"pid": after["pid"], "interval_seconds": seconds, "cpu_seconds": cpu, + "self_cpu_percent": 100 * (cpu["user"] + cpu["system"]) / seconds, + "reaped_child_cpu_seconds": cpu["reaped_child_user"] + cpu["reaped_child_system"], + "child_scope": "charged at reap; excludes live/unreaped children, not Git-only attribution"} + + +def observe(args): + ready = campaign.read_json(args.fleet_dir / "ready.json") + plan = {"node_active_limit": ready.get("max_active_repositories_per_node")} + ready = campaign.validate_fleet(args.fleet_dir, plan) + reader = Reader() + args.output_dir.mkdir(parents=True, exist_ok=False) + path = args.output_dir / "observation.json" + samples_path = args.output_dir / "samples.jsonl" + result = {"version": 1, "complete": False, "error": None, "samples": 0, + "declared": {"samples": args.samples, "interval_seconds": args.interval}, + "driver_sha256": benchmark.file_sha256(Path(__file__)), + "fleet_validator_sha256": benchmark.file_sha256(Path(campaign.__file__)), + "fleet_ready_sha256": benchmark.file_sha256(args.fleet_dir / "ready.json"), + "binary_sha256": ready["binary_sha256"], "node_pids": [node["pid"] for node in ready["nodes"]], + "platform": sys.platform, + "scope": "kernel self and reaped-child CPU only; no live-tree total, Git-only attribution, child-boundary closure, operation CPU, throughput or provider-cost proof"} + benchmark.save(path, result) + previous = None + try: + with samples_path.open("x") as output: + for index in range(args.samples): + current = [reader.snapshot(pid) for pid in result["node_pids"]] + intervals = [delta(old, new) for old, new in zip(previous, current)] if previous else [] + sample = {"index": index, "utc": datetime.now(timezone.utc).isoformat(), + "processes": current, "intervals": intervals} + output.write(json.dumps(sample) + "\n") + output.flush() + result["samples"] += 1 + previous = current + if index + 1 < args.samples: + time.sleep(args.interval) + campaign.validate_fleet(args.fleet_dir, plan) + result["samples_sha256"] = benchmark.file_sha256(samples_path) + result["complete"] = True + except BaseException as error: + result["error"] = type(error).__name__ + raise + finally: + if samples_path.exists(): + result["samples_sha256"] = benchmark.file_sha256(samples_path) + benchmark.save(path, result) + return result + + +def main(): + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("--fleet-dir", type=Path, required=True) + parser.add_argument("--output-dir", type=Path, required=True) + parser.add_argument("--samples", type=int, default=120) + parser.add_argument("--interval", type=float, default=1) + args = parser.parse_args() + require(2 <= args.samples <= 86400 and math.isfinite(args.interval) and 0.1 <= args.interval <= 30, + "invalid observation bounds") + observe(args) + + +if __name__ == "__main__": + main() diff --git a/scripts/test_fleet_cpu.py b/scripts/test_fleet_cpu.py new file mode 100644 index 0000000..758dd71 --- /dev/null +++ b/scripts/test_fleet_cpu.py @@ -0,0 +1,142 @@ +"""Kernel CPU units, child rollup, identity boundaries and retained failures.""" + +import ctypes +import json +import os +from pathlib import Path +import select +import subprocess +import sys +import tempfile +import time +from types import SimpleNamespace +import unittest +from unittest.mock import patch + +import fleet_cpu as cpu + + +class CpuTests(unittest.TestCase): + def snapshot(self, pid=123, capture=1, user=10, child=2): + return {"pid": pid, "identity": {"start_ticks": 456}, "backend": "fixture", + "scale": {"clock_ticks_per_second": 100}, "seconds_per_unit": .01, + "cpu_units": dict(zip(cpu.COUNTERS, (user, 3, child, 1))), + "rss_bytes": 4096, "captured_monotonic": capture} + + def test_darwin_sdk_structure_layout(self): + self.assertEqual(ctypes.sizeof(cpu.DarwinUsage), 160) + self.assertEqual(cpu.DarwinUsage.child_user_time.offset, 96) + self.assertEqual(cpu.DarwinUsage.proc_start_abstime.offset, 80) + + def linux_stat(self, state="S"): + fields = [state] + ["0"] * 21 + for index, value in zip((11, 12, 13, 14, 19, 21), (10, 20, 30, 40, 500, 2)): + fields[index] = str(value) + return "123 (odd ) name)) " + " ".join(fields) + + def test_linux_ticks_and_comm_parentheses(self): + sample = cpu.parse_linux_stat(self.linux_stat(), 123, 100, 4096) + self.assertEqual(sample["cpu_units"], dict(zip(cpu.COUNTERS, (10, 20, 30, 40)))) + self.assertEqual(sample["identity"], {"start_ticks": 500}) + self.assertEqual(sample["seconds_per_unit"], .01) + self.assertEqual(sample["rss_bytes"], 8192) + for text, pid, rate in ((self.linux_stat("Z"), 123, 100), (self.linux_stat("X"), 123, 100), + (self.linux_stat(), 124, 100), ("123 (truncated) S", 123, 100), + (self.linux_stat(), 123, 0)): + with self.assertRaises(ValueError): + cpu.parse_linux_stat(text, pid, rate, 4096) + + def test_delta_has_separate_self_and_reaped_child_cpu(self): + result = cpu.delta(self.snapshot(), self.snapshot(capture=3, user=30, child=52)) + self.assertEqual(result["cpu_seconds"]["user"], .2) + self.assertEqual(result["self_cpu_percent"], 10) + self.assertEqual(result["reaped_child_cpu_seconds"], .5) + self.assertIn("excludes live", result["child_scope"]) + + def test_identity_scale_clock_and_counter_discontinuities_fail(self): + before = self.snapshot() + mutations = [{"pid": 124}, {"identity": {"start_ticks": 999}}, + {"backend": "different"}, {"scale": {"clock_ticks_per_second": 1000}}, + {"seconds_per_unit": .001}, {"captured_monotonic": 1}, + {"captured_monotonic": float("nan")}, + {"cpu_units": dict(zip(cpu.COUNTERS, (9, 3, 2, 1)))}] + for mutation in mutations: + with self.subTest(mutation=mutation), self.assertRaises(ValueError): + cpu.delta(before, {**self.snapshot(capture=2), **mutation}) + + @unittest.skipUnless(sys.platform in ("darwin", "linux"), "requires native kernel reader") + def test_native_units_calibrate_against_process_cpu_clock(self): + reader = cpu.Reader() + before = reader.snapshot(os.getpid()) + started = time.process_time() + while time.process_time() - started < .15: + sum(range(10000)) + expected = time.process_time() - started + after = reader.snapshot(os.getpid()) + observed = cpu.delta(before, after)["cpu_seconds"] + self.assertAlmostEqual(observed["user"] + observed["system"], expected, delta=.04) + + @unittest.skipUnless(sys.platform in ("darwin", "linux"), "requires native kernel reader") + def test_child_cpu_is_absent_while_live_and_rolls_up_after_wait(self): + reader = cpu.Reader() + before = reader.snapshot(os.getpid()) + command = "import time,sys\nstart=time.process_time()\nwhile time.process_time()-start < .15: sum(range(10000))\nprint('ready',flush=True)\nsys.stdin.readline()\n" + child = subprocess.Popen([sys.executable, "-B", "-c", command], stdin=subprocess.PIPE, + stdout=subprocess.PIPE, stderr=subprocess.PIPE, text=True) + try: + self.assertTrue(select.select([child.stdout], [], [], 10)[0], "child readiness timed out") + self.assertEqual(child.stdout.readline(), "ready\n") + live = reader.snapshot(os.getpid()) + child_usage = reader.snapshot(child.pid) + child_self = sum(child_usage["cpu_units"][name] for name in ("user", "system")) * child_usage["seconds_per_unit"] + self.assertGreater(child_self, .12) + self.assertEqual(cpu.delta(before, live)["reaped_child_cpu_seconds"], 0) + child.communicate(input="finish\n", timeout=10) + self.assertEqual(child.returncode, 0) + after = reader.snapshot(os.getpid()) + self.assertGreater(cpu.delta(live, after)["reaped_child_cpu_seconds"], .12) + with self.assertRaises((OSError, ValueError)): + reader.snapshot(child.pid) + finally: + if child.poll() is None: + child.kill() + child.wait(timeout=10) + for stream in (child.stdin, child.stdout, child.stderr): + stream.close() + + def test_observer_complete_and_failure_receipts_preserve_samples(self): + for failure in (False, True): + with self.subTest(failure=failure), tempfile.TemporaryDirectory() as directory: + root = Path(directory) + fleet = root / "fleet" + fleet.mkdir() + ready = {"max_active_repositories_per_node": 100, "binary_sha256": "a" * 64, + "nodes": [{"pid": pid} for pid in (123, 124, 125)]} + (fleet / "ready.json").write_text(json.dumps(ready)) + args = SimpleNamespace(fleet_dir=fleet, output_dir=root / "observed", samples=2, interval=.1) + calls = [] + def snapshot(pid): + calls.append(pid) + if failure and len(calls) == 4: + raise OSError("sensitive diagnostic text") + return self.snapshot(pid, capture=1 if len(calls) <= 3 else 2) + with patch.object(cpu.campaign, "validate_fleet", return_value=ready), \ + patch.object(cpu, "Reader", return_value=SimpleNamespace(snapshot=snapshot)), \ + patch.object(cpu.time, "sleep"): + if failure: + with self.assertRaises(OSError): + cpu.observe(args) + else: + self.assertTrue(cpu.observe(args)["complete"]) + record = json.loads((args.output_dir / "observation.json").read_text()) + samples = (args.output_dir / "samples.jsonl").read_text().splitlines() + self.assertEqual(record["samples"], len(samples)) + self.assertEqual(len(samples), 1 if failure else 2) + self.assertEqual(record["complete"], not failure) + self.assertEqual(record["error"], "OSError" if failure else None) + self.assertNotIn("sensitive diagnostic text", json.dumps(record)) + self.assertEqual(record["samples_sha256"], cpu.benchmark.file_sha256(args.output_dir / "samples.jsonl")) + + +if __name__ == "__main__": + unittest.main() From 59a73078cc5841b407f791885c8eb9fa5e90f087 Mon Sep 17 00:00:00 2001 From: forhappy Date: Wed, 30 Sep 2026 17:47:06 -0700 Subject: [PATCH 07/60] Record live kernel CPU observation and pack rollup calibration --- .../2026-09-30-three-node-proxy.md | 34 +++++++++++++++++++ 1 file changed, 34 insertions(+) diff --git a/docs/performance/2026-09-30-three-node-proxy.md b/docs/performance/2026-09-30-three-node-proxy.md index edd8bbe..d86a001 100644 --- a/docs/performance/2026-09-30-three-node-proxy.md +++ b/docs/performance/2026-09-30-three-node-proxy.md @@ -587,6 +587,40 @@ only Git, and do not close operation-specific CPU/GiB accounting without independently verified child boundaries. Historical `ps` results are unchanged; no child CPU is retroactively imputed to the failed baseline. +The committed reader (`6199a73`) completed 120 live samples during seeding, from +00:40:22 to 00:42:22 UTC on October 1. Independent checks confirmed exact +indices 0..119, all three bound PIDs, unchanged process identities and raw +counter/timebase-derived totals over about 120.551 seconds: + +| Candidate node PID | Self CPU seconds | Mean self CPU (%) | Reaped-child CPU seconds | +| --- | ---: | ---: | ---: | +| 87409 | 5.483475 | 4.548687 | 0 | +| 88175 | 0.125026 | 0.103713 | 0 | +| 88533 | 0.108677 | 0.090150 | 0 | + +Zero charged child CPU in this interval is not proof that no children were +live. A separate retained `check_cpu_during_pack.py` calibration then bracketed +three verified, read-only pack requests during the still-active seed. All three +node identities stayed unchanged and their reaped-child counters increased: +48.400, 45.548 and 39.485 ms, respectively, over about 4.525 seconds. Every pack +was 1,153 bytes and passed strict indexing, exact wanted-commit validation and +full strict fsck. This verifies that the reader observes Canopy child rollup; +the overlapping seed and unclosed child boundaries prevent Git-only, +per-request or CPU/GiB attribution. No writes or fleet restarts were introduced. + +| Candidate-volume artifact | SHA-256 | +| --- | --- | +| `kernel-cpu-during-seed/observation.json` | `b48364c7c908da7789481d181d8c82d0026284e3a4c1d621e60625e0ae0feea5` | +| `kernel-cpu-during-seed/samples.jsonl` | `482ce4ebe4a5d2cb18fe3f2f88253dd6418b8e17f2bd799a4320a70d7eb2bca8` | +| `kernel-cpu-pack-calibration.json` | `0be6c1653fbd7dd34b2d12638b063d2a86c122d47379a2f43d39d7bebb5ff007` | +| `pack-probe-under-cpu-check/measurement.json` | `85bef58670c8e797375f4719d37766846788615be748b106362d51be1f50aab3` | + +Both Linux-hosted harness jobs for `6199a73` also passed all 59 tests: +[workflow 36797306682](https://github.com/crabbuild/canopy/actions/runs/36797306682) +and [workflow 36797311677](https://github.com/crabbuild/canopy/actions/runs/36797311677). +That includes native Linux tick calibration and waited-child rollup; it does +not qualify the hosted Rust/provider jobs while those are still running. + A separate read-only observer records dedicated RustFS container CPU/memory samples about every six seconds, beginning during full-corpus preflight. Its binding includes the observer digest, campaign PID, container ID and start time; From 9c29627efdf0235208938f25898bf9ba9591bf51 Mon Sep 17 00:00:00 2001 From: forhappy Date: Wed, 30 Sep 2026 18:09:52 -0700 Subject: [PATCH 08/60] Measure restart-bound RustFS request counters without reconfiguration --- .../2026-09-30-three-node-proxy.md | 51 +++++ scripts/provider_requests.py | 179 ++++++++++++++++++ scripts/test_provider_requests.py | 142 ++++++++++++++ 3 files changed, 372 insertions(+) create mode 100644 scripts/provider_requests.py create mode 100644 scripts/test_provider_requests.py diff --git a/docs/performance/2026-09-30-three-node-proxy.md b/docs/performance/2026-09-30-three-node-proxy.md index d86a001..ef8b04a 100644 --- a/docs/performance/2026-09-30-three-node-proxy.md +++ b/docs/performance/2026-09-30-three-node-proxy.md @@ -628,6 +628,57 @@ errors and sampling gaps remain in `provider-samples.jsonl`. Docker's raw displa units and rounding are retained. These container samples are not whole-VM CPU, S3 request cost or exact per-window wire-byte counts; `NetIO` is cumulative. +### Provider request counters without a restart + +The dedicated RustFS instance exposes signed, read-only console metrics at +`GET /rustfs/admin/v3/metrics?n=1&types=512`. This is NDJSON, not a Prometheus +endpoint. The image's labelled revision documents the +[authenticated handler](https://github.com/rustfs/rustfs/blob/d47f54bfb2f39f48bd1adda334bd27e151fe85b8/rustfs/src/admin/handlers/metrics.rs) +and [HTTP-only metric selection](https://github.com/rustfs/rustfs/blob/d47f54bfb2f39f48bd1adda334bd27e151fe85b8/crates/ecstore/src/services/metrics_realtime.rs). +An actual signed read returned HTTP 200 and S3 operation/outcome counters; +no telemetry configuration change or provider restart was needed. + +`scripts/provider_requests.py` records bounded snapshots and deltas by HTTP +method, S3 operation and outcome. It verifies the bound container, image, start +time and restart count around each read. Missing/decreasing series, restarted +providers and failed reads leave an incomplete receipt with earlier samples +retained. Credentials go to curl on stdin, never command arguments; only explicit +loopback HTTP endpoints are accepted, without redirects or proxy configuration. +This helper is scoped to the dedicated `three-node-candidate-e07670e` fixture. + +```sh +# Disposable fixture credentials must already be in the environment. +python3 -B scripts/provider_requests.py \ + --binding /dedicated-volume/candidate-binding.json \ + --docker-config /path/to/disposable-docker-config \ + --docker-host unix:///path/to/docker.sock \ + --output-dir /dedicated-volume/new-provider-request-observation \ + --samples 12 --interval 5 +``` + +Counts include background work and retries. Unknown labels and non-2xx outcomes +are retained: a provider 4xx can be a conditional-write rejection or missing +object, not necessarily a failed Canopy operation. Observer overhead is not +subtracted. The pinned [request-counter implementation](https://github.com/rustfs/rustfs/blob/d47f54bfb2f39f48bd1adda334bd27e151fe85b8/crates/io-metrics/src/s3_http_metrics.rs) +counts once at response headers, service error or cancellation; later body-stream +failures are outside this counter. Snapshots are not atomic across series. +These counts are request-volume evidence, not billed dollars, +request latency, wire bytes or operation-attributed cost. Seed-load observations +cannot close matched load-window accounting. + +Seven provider-reader tests cover signed-read isolation, schema/outcome labels, +new series, counter reset/disappearance, restart/identity boundaries and retained +partial failures. All 66 harness tests passed on Python 3.12 and 3.14. Retained +candidate-volume logs `python312-provider-final-suite.log` and +`python314-provider-final-suite.log` have SHA-256 +`027b64abca6de9f9cf0aa52b0bacf83c4862e9e5b466b514ff939f102f5d27e0` +and `6b4e22260518592ec0de75175f356aed856738c1c76fb5c157bb3fa4ee70ead6`. +Both full hosted workflows for the preceding `59a7307` commit also completed +successfully, including Rust tests, RustFS compatibility and server build: +[36797898373](https://github.com/crabbuild/canopy/actions/runs/36797898373) +and [36797902238](https://github.com/crabbuild/canopy/actions/runs/36797902238). +Those earlier runs did not include this new provider reader. + `push_branch` reuses one commit and therefore measures ref publication, not fresh pack ingestion. `push_commit` clones prepared base objects locally, creates a new child of the original corpus tip, and sends a distinct payload on a unique ref. diff --git a/scripts/provider_requests.py b/scripts/provider_requests.py new file mode 100644 index 0000000..d49dc5d --- /dev/null +++ b/scripts/provider_requests.py @@ -0,0 +1,179 @@ +#!/usr/bin/env python3 +"""Read RustFS console S3 counters without changing provider configuration. + +Uses signed GET /rustfs/admin/v3/metrics?n=1&types=512 (NDJSON, not +Prometheus). Requests, outcomes and restart boundaries are retained. Counts +include background ownership work and retries; they are not billed USD, bytes, +latency, or per-Canopy-operation cost without matched workload boundaries. +Credentials are read from environment and passed to curl on stdin, never argv. +""" + +import argparse +from datetime import datetime, timezone +import hashlib +import json +import math +import os +from pathlib import Path +import re +import subprocess +import time +from urllib.parse import urlsplit + +import benchmark_repositories as benchmark + + +def require(condition, message): + if not condition: + raise ValueError(message) + + +def parse_metrics(body): + require(len(body) <= 1024 * 1024, "metrics response exceeds bound") + lines = body.splitlines() + require(len(lines) == 1, "require exactly one NDJSON sample") + record = json.loads(lines[0]) + require(record.get("errors") == [] and record.get("final") is True, + "provider metrics incomplete or errored") + http = record.get("aggregated", {}).get("http") + require(isinstance(http, dict) and isinstance(http.get("collected"), str), + "missing provider HTTP counters") + counters = http.get("requests") + require(isinstance(counters, list) and len(counters) <= 4096, "invalid provider request counters") + seen = set() + for counter in counters: + require(isinstance(counter, dict), "invalid request counter") + require(counter.get("method") in ("GET", "HEAD", "PUT", "POST", "DELETE", "OPTIONS", "PATCH", + "CONNECT", "TRACE", "OTHER"), + "invalid request method") + require(isinstance(counter.get("operation"), str) + and re.fullmatch(r"(?:s3:[A-Za-z0-9_]{1,64}|unknown)", counter["operation"]), + "invalid operation label") + require(counter.get("outcome") in ("1xx", "2xx", "3xx", "4xx", "5xx", "cancelled", + "service_error", "unknown"), + "invalid outcome label") + require(isinstance(counter.get("total"), int) and not isinstance(counter["total"], bool) + and 0 <= counter["total"] < 2**64, "invalid request count") + key = tuple(counter[name] for name in ("method", "operation", "outcome")) + require(key not in seen, "duplicate counter series") + seen.add(key) + return record + + +def fetch(endpoint, access_key, secret_key, region): + url = urlsplit(endpoint) + require(url.scheme == "http" and url.hostname in ("127.0.0.1", "::1") + and url.port and not url.username and not url.password and not url.query + and not url.fragment and url.path in ("", "/"), "require explicit loopback provider endpoint") + # Bound configuration syntax and credentials so neither can inject curl + # directives. Do not include curl stderr or config text in error reports. + require(all(isinstance(value, str) and re.fullmatch(r"[A-Za-z0-9_+/=.-]{1,256}", value) + for value in (access_key, secret_key)), "invalid fixture credentials") + require(re.fullmatch(r"[a-z0-9-]{1,64}", region), "invalid region") + config = f'user = "{access_key}:{secret_key}"\naws-sigv4 = "aws:amz:{region}:s3"\n' + started = time.monotonic() + # Redirects are not followed. max-filesize bounds the body, including + # chunked responses; max-time bounds the whole request. Status trailer is + # split from the body, which is never logged alongside signing headers. + result = subprocess.run(["curl", "--disable", "--config", "-", "--noproxy", "*", "--silent", "--show-error", + "--max-time", "10", "--max-filesize", "1048576", "--write-out", "\n%{http_code}\n%{content_type}", + endpoint.rstrip("/") + "/rustfs/admin/v3/metrics?n=1&types=512"], + input=config.encode(), capture_output=True, timeout=15, check=False) + finished = time.monotonic() + require(result.returncode == 0, "provider metrics transfer failed") + body, status, content_type = result.stdout.rsplit(b"\n", 2) + require(status == b"200" and content_type.split(b";")[0].strip() == b"application/x-ndjson", + "provider metrics HTTP status/content type differs") + metrics = parse_metrics(body) + return {"utc": datetime.now(timezone.utc).isoformat(), + "read_started_monotonic": started, "read_finished_monotonic": finished, + "body_sha256": hashlib.sha256(body).hexdigest(), "metrics": metrics} + + +def provider_state(binding, docker_config, docker_host): + result = subprocess.run(["docker", "--config", str(docker_config), "--host", docker_host, + "inspect", binding["provider_container_id"]], capture_output=True, timeout=10, check=False) + require(result.returncode == 0, "provider inspect failed") + values = json.loads(result.stdout) + require(isinstance(values, list) and len(values) == 1, "ambiguous provider identity") + value = values[0] + require(value["Id"] == binding["provider_container_id"] and value["State"]["Running"] is True + and value["State"]["StartedAt"] == binding["provider_started_at"] + and value["Image"] == binding["provider_image_id"], "provider identity/start changed") + require(value["Config"]["Labels"].get("canopy.purpose") == "three-node-candidate-e07670e", + "provider purpose differs") + # Never return Docker's full environment: it contains provider credentials. + return {"id": value["Id"], "started_at": value["State"]["StartedAt"], + "image": value["Image"], "restart_count": value["RestartCount"]} + + +def changes(before, after): + def series(sample): + return {tuple(row[key] for key in ("method", "operation", "outcome")): row["total"] + for row in sample["metrics"]["aggregated"]["http"]["requests"]} + old, new = series(before), series(after) + require(old.keys() <= new.keys(), "counter series disappeared") + result = [] + for key, value in sorted(new.items()): + increment = value - old.get(key, 0) + require(increment >= 0, "provider counter decreased") + result.append(dict(zip(("method", "operation", "outcome", "requests"), (*key, increment)))) + return result + + +def observe(args): + require(2 <= args.samples <= 86400 and math.isfinite(args.interval) and .1 <= args.interval <= 30, + "invalid observation bounds") + binding = json.loads(args.binding.read_text()) + access_key, secret_key = os.environ.get("AWS_ACCESS_KEY_ID"), os.environ.get("AWS_SECRET_ACCESS_KEY") + region = os.environ.get("AWS_REGION", "us-east-1") + initial_state = provider_state(binding, args.docker_config, args.docker_host) + args.output_dir.mkdir(parents=True, exist_ok=False) + path = args.output_dir / "observation.json" + samples_path = args.output_dir / "samples.jsonl" + result = {"version": 1, "complete": False, "error": None, "samples": 0, + "declared": {"samples": args.samples, "interval_seconds": args.interval}, + "binding_sha256": benchmark.file_sha256(args.binding), "provider": initial_state, + "driver_sha256": benchmark.file_sha256(Path(__file__)), + "scope": "provider-wide S3 request/outcome counters including background work/retries; not priced cost, bytes, request latency or per-Canopy-operation attribution"} + benchmark.save(path, result) + previous = None + try: + with samples_path.open("x") as output: + for index in range(args.samples): + state = provider_state(binding, args.docker_config, args.docker_host) + sample = fetch(binding["provider_endpoint"], access_key, secret_key, region) + require(provider_state(binding, args.docker_config, args.docker_host) == state == initial_state, + "provider restarted during observation") + sample["index"] = index + sample["changes"] = changes(previous, sample) if previous else [] + output.write(json.dumps(sample) + "\n") + output.flush() + previous = sample + result["samples"] += 1 + if index + 1 < args.samples: + time.sleep(args.interval) + result["complete"] = True + except BaseException as error: + result["error"] = type(error).__name__ + raise + finally: + if samples_path.exists(): + result["samples_sha256"] = benchmark.file_sha256(samples_path) + benchmark.save(path, result) + return result + + +def main(): + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("--binding", type=Path, required=True) + parser.add_argument("--output-dir", type=Path, required=True) + parser.add_argument("--docker-config", type=Path, required=True) + parser.add_argument("--docker-host", required=True) + parser.add_argument("--samples", type=int, default=12) + parser.add_argument("--interval", type=float, default=5) + observe(parser.parse_args()) + + +if __name__ == "__main__": + main() diff --git a/scripts/test_provider_requests.py b/scripts/test_provider_requests.py new file mode 100644 index 0000000..6cecd1b --- /dev/null +++ b/scripts/test_provider_requests.py @@ -0,0 +1,142 @@ +"""Validate provider counters, safe signed reads and restart/failure boundaries.""" + +import copy +import json +from pathlib import Path +import tempfile +from types import SimpleNamespace +import unittest +from unittest.mock import patch + +import provider_requests as provider + + +class ProviderTests(unittest.TestCase): + def metrics(self, total=10): + return {"errors": [], "final": True, "hosts": [":9000"], "aggregated": {"http": { + "collected": "2026-10-01T00:00:00Z", "requests": [{"method": "GET", + "operation": "s3:GetObject", "outcome": "2xx", "total": total}]}}} + + def sample(self, total=10): + return {"metrics": self.metrics(total)} + + def test_valid_counter_snapshot_preserves_unknown_operations_and_outcomes(self): + record = self.metrics() + record["aggregated"]["http"]["requests"].append( + {"method": "PUT", "operation": "unknown", "outcome": "4xx", "total": 1}) + for method in ("CONNECT", "TRACE", "OTHER"): + for outcome in ("cancelled", "service_error", "unknown"): + record["aggregated"]["http"]["requests"].append( + {"method": method, "operation": "unknown", "outcome": outcome, "total": 1}) + self.assertEqual(provider.parse_metrics(json.dumps(record).encode() + b"\n"), record) + + def test_incomplete_duplicate_malformed_or_oversized_metrics_fail(self): + mutations = [lambda value: value.update(final=False), + lambda value: value.update(errors=["collection failed"]), + lambda value: value["aggregated"].update(http=None), + lambda value: value["aggregated"]["http"]["requests"].append( + copy.deepcopy(value["aggregated"]["http"]["requests"][0])), + lambda value: value["aggregated"]["http"]["requests"][0].update(total=True), + lambda value: value["aggregated"]["http"]["requests"][0].update(total=-1), + lambda value: value["aggregated"]["http"]["requests"][0].update(total=2**64), + lambda value: value["aggregated"]["http"]["requests"][0].update(operation="private\nlabel"), + lambda value: value["aggregated"]["http"]["requests"][0].update(method="BAD"), + lambda value: value["aggregated"]["http"]["requests"][0].update(outcome="invalid")] + for mutate in mutations: + record = self.metrics() + mutate(record) + with self.assertRaises(ValueError): + provider.parse_metrics(json.dumps(record).encode()) + for body in (b"x" * (1024 * 1024 + 1), b"", b"{}\n{}\n"): + with self.assertRaises(ValueError): + provider.parse_metrics(body) + + def test_counter_deltas_preserve_new_series_and_reject_reset_or_disappearance(self): + before, after = self.sample(10), self.sample(20) + after["metrics"]["aggregated"]["http"]["requests"].append( + {"method": "PUT", "operation": "s3:PutObject", "outcome": "4xx", "total": 2}) + self.assertEqual([row["requests"] for row in provider.changes(before, after)], [10, 2]) + with self.assertRaisesRegex(ValueError, "decreased"): + provider.changes(before, self.sample(9)) + after["metrics"]["aggregated"]["http"]["requests"] = [] + with self.assertRaisesRegex(ValueError, "disappeared"): + provider.changes(before, after) + + def test_fetch_is_credential_safe_bounded_and_configuration_isolated(self): + body = json.dumps(self.metrics()).encode() + b"\n" + result = SimpleNamespace(returncode=0, stdout=body + b"\n200\napplication/x-ndjson", stderr=b"") + with patch.object(provider.subprocess, "run", return_value=result) as run: + sample = provider.fetch("http://127.0.0.1:32776", "private-access", "private-secret", "us-east-1") + args, kwargs = run.call_args + argv = args[0] + self.assertEqual(argv[0:2], ["curl", "--disable"]) + self.assertIn("--noproxy", argv) + self.assertNotIn("--location", argv) + self.assertNotIn("private-secret", " ".join(argv)) + self.assertIn(b"private-access:private-secret", kwargs["input"]) + self.assertNotIn("private-secret", json.dumps(sample)) + self.assertEqual(sample["metrics"], self.metrics()) + + def test_unsafe_endpoint_credentials_and_failed_http_are_rejected(self): + for endpoint in ("http://example.com:9000", "https://127.0.0.1:9000", "http://127.0.0.1", + "http://name:secret@127.0.0.1:9000", "http://127.0.0.1:9000/?bad=1"): + with patch.object(provider.subprocess, "run", side_effect=AssertionError("unexpected request")), \ + self.assertRaises(ValueError): + provider.fetch(endpoint, "access", "secret", "us-east-1") + with self.assertRaises(ValueError): + provider.fetch("http://127.0.0.1:9000", "access", 'inject"\nheader=bad', "us-east-1") + for returncode, ending in ((1, b""), (0, b"\n302\napplication/x-ndjson"), + (0, b"\n401\napplication/xml"), (0, b"\n200\ntext/html")): + result = SimpleNamespace(returncode=returncode, stdout=b"{}" + ending, + stderr=b"private signed header") + with patch.object(provider.subprocess, "run", return_value=result), self.assertRaises(ValueError) as error: + provider.fetch("http://127.0.0.1:9000", "access", "secret", "us-east-1") + self.assertNotIn("private signed header", str(error.exception)) + + def test_provider_restart_identity_and_purpose_are_enforced(self): + binding = {"provider_container_id": "id", "provider_started_at": "start", "provider_image_id": "image"} + value = {"Id": "id", "State": {"Running": True, "StartedAt": "start"}, "Image": "image", + "RestartCount": 0, "Config": {"Labels": {"canopy.purpose": "three-node-candidate-e07670e"}}} + for mutate in (None, lambda item: item["State"].update(StartedAt="new"), + lambda item: item["State"].update(Running=False), + lambda item: item.update(Id="different"), + lambda item: item.update(Image="different"), + lambda item: item["Config"]["Labels"].update({"canopy.purpose": "unrelated"})): + changed = copy.deepcopy(value) + if mutate: + mutate(changed) + result = SimpleNamespace(returncode=0, stdout=json.dumps([changed]).encode()) + with patch.object(provider.subprocess, "run", return_value=result): + if mutate: + with self.assertRaises(ValueError): + provider.provider_state(binding, Path("unused"), "unused") + else: + self.assertEqual(provider.provider_state(binding, Path("unused"), "unused")["id"], "id") + + def test_observation_restarts_and_failed_reads_leave_incomplete_receipts(self): + for failure in (None, "read", "restart"): + with tempfile.TemporaryDirectory() as directory: + root = Path(directory) + binding = root / "binding.json" + binding.write_text(json.dumps({"provider_endpoint": "http://127.0.0.1:9000"})) + args = SimpleNamespace(binding=binding, docker_config=root / "docker", docker_host="unused", + output_dir=root / "observation", samples=2, interval=.1) + samples = [self.sample(10), ValueError("private diagnostic") if failure == "read" else self.sample(20)] + states = [{"id": "id", "restart_count": 0}] * 4 + states += [{"id": "id", "restart_count": 1 if failure == "restart" else 0}] + with patch.object(provider, "provider_state", side_effect=states), \ + patch.object(provider, "fetch", side_effect=samples), patch.object(provider.time, "sleep"): + if failure: + with self.assertRaises(ValueError): + provider.observe(args) + else: + self.assertTrue(provider.observe(args)["complete"]) + record = json.loads((args.output_dir / "observation.json").read_text()) + self.assertEqual(record["complete"], failure is None) + self.assertEqual(record["samples"], 1 if failure else 2) + self.assertNotIn("private diagnostic", json.dumps(record)) + self.assertEqual(record["samples_sha256"], provider.benchmark.file_sha256(args.output_dir / "samples.jsonl")) + + +if __name__ == "__main__": + unittest.main() From f7fc26f1de5b8f52debd8845777b6aca6e665ad8 Mon Sep 17 00:00:00 2001 From: forhappy Date: Wed, 30 Sep 2026 18:12:17 -0700 Subject: [PATCH 09/60] Record live provider counter observation and independent audit --- .../2026-09-30-three-node-proxy.md | 30 +++++++++++++++++++ 1 file changed, 30 insertions(+) diff --git a/docs/performance/2026-09-30-three-node-proxy.md b/docs/performance/2026-09-30-three-node-proxy.md index ef8b04a..eeb5146 100644 --- a/docs/performance/2026-09-30-three-node-proxy.md +++ b/docs/performance/2026-09-30-three-node-proxy.md @@ -679,6 +679,36 @@ successfully, including Rust tests, RustFS compatibility and server build: and [36797902238](https://github.com/crabbuild/canopy/actions/runs/36797902238). Those earlier runs did not include this new provider reader. +The committed reader (`9c29627`) completed 12 live samples at five-second +intervals, from 01:10:09 to 01:11:06 UTC on October 1. An independent audit +checked exact indices, retained sample/binding/driver digests and every interval +delta. Container identity, start time and restart count stayed unchanged. +Across the roughly 56.743-second client read span, first-to-last counter changes +were: + +| Provider operation | Outcome | Additional requests | +| --- | --- | ---: | +| `GetObject` | 2xx | 444 | +| `GetObject` | 4xx | 52 | +| `HeadObject` | 2xx | 174 | +| `PutObject` | 2xx | 2,679 | + +All other existing series had zero delta: total 3,349. No scrape overhead was +subtracted. Seeding/background ownership continued throughout; this is not a +per-repository cost, a throughput qualification or a matched improvement. + +| Candidate-volume artifact | SHA-256 | +| --- | --- | +| `provider-request-during-seed/observation.json` | `2d6a80d22480a5d56321fe3cf6e6352317810b2fd1d2e87ed0819239c232ce28` | +| `provider-request-during-seed/samples.jsonl` | `a034331099cd4e278dddd967de08d3203d9aa8d3f378c3be327aa3efc604e4a5` | +| `provider-request-audit.json` | `9c00996008d05bb0682154d907ee15383954b5af897003f75683c15da8531de3` | +| `audit_provider_observation.py` | `414bae05eae1cca67982e68c5e879d164dd892b8ff44a2a7a569ba632a3853ce` | + +Both Linux harness jobs for `9c29627` passed all 66 tests: +[36799769344](https://github.com/crabbuild/canopy/actions/runs/36799769344) +and [36799773408](https://github.com/crabbuild/canopy/actions/runs/36799773408). +Their Rust/provider jobs were still running at this observation. + `push_branch` reuses one commit and therefore measures ref publication, not fresh pack ingestion. `push_commit` clones prepared base objects locally, creates a new child of the original corpus tip, and sends a distinct payload on a unique ref. From 39829f72850b1b48bfbabd49dfcf75585d4fa2b2 Mon Sep 17 00:00:00 2001 From: forhappy Date: Wed, 30 Sep 2026 18:38:43 -0700 Subject: [PATCH 10/60] Verify expired renewal replies stay fenced and arm full seed recovery --- .../tests/multi_server/lifecycle/lease.rs | 61 +++++++++++++++++++ .../2026-09-30-three-node-proxy.md | 57 +++++++++++++++++ 2 files changed, 118 insertions(+) diff --git a/crates/canopy-server/tests/multi_server/lifecycle/lease.rs b/crates/canopy-server/tests/multi_server/lifecycle/lease.rs index 1ed44fd..a2bf498 100644 --- a/crates/canopy-server/tests/multi_server/lifecycle/lease.rs +++ b/crates/canopy-server/tests/multi_server/lifecycle/lease.rs @@ -206,3 +206,64 @@ async fn delayed_renewal_reply_does_not_extend_confirmed_authority() -> Result { ); Ok(()) } + +#[tokio::test(flavor = "multi_thread")] +async fn expired_successful_renewal_reply_cannot_reopen_ingress() -> Result { + let files = tempfile::TempDir::new()?; + let store = Arc::new(PausedStore::default()); + let mut settings = config(available_address().await?, files.path().join("node")); + settings.listen.set_port(0); + let server = CanopyServer::start(settings, store.clone()).await?; + let address = server.local_addr(); + store.delay_renewals.store(true, Ordering::SeqCst); + store.wait().await?; + let expires = store + .renewal_expiry + .lock() + .unwrap() + .ok_or("renewal expiry missing")?; + let now = || -> Result { + Ok(i64::try_from( + std::time::SystemTime::now() + .duration_since(std::time::UNIX_EPOCH)? + .as_millis(), + )?) + }; + // The conditional write has succeeded. Hold only its reply past the + // signed replacement expiry, not merely past the previous lease. + let remaining = u64::try_from(expires - now()? + 200)?; + tokio::time::sleep(Duration::from_millis(remaining)).await; + let expired_at_reply = now()?; + store.delay_renewals.store(false, Ordering::SeqCst); + store.proceed.notify_one(); + // No caller stop signal: only the server's own terminal fence may end + // supervision. An explicit shutdown could hide an unsafe revival. + let shutdown = timeout( + Duration::from_secs(10), + server.serve_until(std::future::pending::>()), + ) + .await?; + let readiness = reqwest::Client::new() + .get(format!("http://{address}/readyz")) + .timeout(Duration::from_secs(2)) + .send() + .await; + assert!(expired_at_reply > expires); + assert!( + matches!( + shutdown, + Err(canopy_server::server::ServerError::Runtime( + cellule_runtime::Error::Fenced + )) + ), + "expired successful refresh did not fence: {shutdown:?}" + ); + assert!( + match &readiness { + Ok(response) => response.status() == reqwest::StatusCode::SERVICE_UNAVAILABLE, + Err(error) => error.is_connect(), + }, + "expired refresh reopened ingress: {readiness:?}" + ); + Ok(()) +} diff --git a/docs/performance/2026-09-30-three-node-proxy.md b/docs/performance/2026-09-30-three-node-proxy.md index eeb5146..1a68118 100644 --- a/docs/performance/2026-09-30-three-node-proxy.md +++ b/docs/performance/2026-09-30-three-node-proxy.md @@ -411,6 +411,63 @@ Retained candidate-directory logs bind the failure and final results: | `startup-port-race-atomic-fixed.log`, passed | `585d00b55b4f4ceedf05a216df20635c347d330e86305efb47f20cea0475e04d` | | `lease-suite-after-fixture-fix.log`, four passed | `27c4a6d42108bd57ac8c006d48f5b39b28ceb473454bc80d41eb40adced8bf28` | +### Reject expired successful replies + +A new real-server/in-memory-store probe holds a successful conditional renewal +reply past the replacement advertisement's signed expiry. It then releases the +reply and waits for the server's own supervisor with **no caller stop signal**. +The supervisor must terminate with `Runtime(Fenced)` and ingress must remain +unavailable. Explicitly shutting down first could hide an unsafe revival. +Production lease durations, deadlines and fencing are unchanged. + +The final supervised probe passed in 33.25 seconds. All five lease lifecycle +tests then passed together in 39.95 seconds, with 100 other integration tests +filtered out. This checks the expired-reply safety branch; it does not reproduce +the baseline's provider latency or prove a performance fix. The Cellule +`node/lease.rs` source at `70bd25f` and `e07670e` is byte-identical, SHA-256 +`a903184abec46b149ab90c889699c9820a9023d9635f4d52b67ad6fdacdf9967`. + +| Candidate-volume test log | SHA-256 | +| --- | --- | +| `expired-renewal-supervised.log` | `006f0cb3208e331372d7432d682b8ac2983dcda8262763e1d2f82b9fe4663d83` | +| `lease-suite-with-expired-reply.log` | `86250ccc3c174bebaf7b8fa3f8d20ffd04b9ca497f0b6e51d23caec3732e883e` | + +### Advance only from a complete candidate seed + +The retained one-off `finish_seed_recovery.py` transition is now running, bound +to seed PID 48549, the immutable candidate binary, current provider and original +three owner process identities. Its live read-only binding check passed before +execution. It waits for that exact seed to end and requires 10,000 distinct +canonical UUID/name pairs and the same 100 two-commit Git/1 MiB LFS fixtures +before any signals. Incomplete seed, changed source/processes, missing owners +or a provider restart stop it without a retry, reseed or replacement fixture. + +After successful seeding, it records three verified-owner SIGKILLs and process +absence, waits at least 32 seconds, starts fresh directories on the same durable +deployment, and verifies both critical fixtures and the complete original +corpus. HTTP timeout stays 30 seconds, verification concurrency stays four, +and admission stays 100/node. It does not start scheduled load. At this +observation its phase is `wait_bound_seed`, not a recovery pass. + +Four simulated control-flow tests passed: complete transition, incomplete seed, +provider restart and missing node. They send no real signals or provider/Git +requests and are not recovery evidence. The live transition retains +`seed-to-recovery.json`, fault boundaries and individual verification receipts. + +| Candidate-volume control artifact | SHA-256 | +| --- | --- | +| `finish_seed_recovery.py` | `a53eef39b0357972204248460896ef6edc6332c9533156263e4bbfdc578ec3bf` | +| `test_finish_seed_recovery.py` | `60633e5411929d04670a7ea730d766241cb65d93317cc13e4aef79b3b216185f` | +| `seed-to-recovery-bound-control-tests.log` | `ac30ceeb7a4a6343e722a531212ec3211adc5351aa60b9138d212d97dfc1760a` | + +Both full hosted workflows for `9c29627` and `f7fc26f` have now passed, including +Rust tests, isolated RustFS compatibility, server build and the 66-test harness: +[36799769344](https://github.com/crabbuild/canopy/actions/runs/36799769344), +[36799773408](https://github.com/crabbuild/canopy/actions/runs/36799773408), +[36799971965](https://github.com/crabbuild/canopy/actions/runs/36799971965) and +[36799975428](https://github.com/crabbuild/canopy/actions/runs/36799975428). +Those earlier commits did not contain the new expired-reply test. + ## Measure the requested operations Complete and verify the corpus before interpreting scheduled load results. From 90d3c39a108c391200ace8c76e41e66714ac7b37 Mon Sep 17 00:00:00 2001 From: forhappy Date: Wed, 30 Sep 2026 19:33:20 -0700 Subject: [PATCH 11/60] Record candidate seed failure and bound ACK recovery diagnostics --- .../2026-09-30-three-node-proxy.md | 128 +++++++++++++++--- 1 file changed, 112 insertions(+), 16 deletions(-) diff --git a/docs/performance/2026-09-30-three-node-proxy.md b/docs/performance/2026-09-30-three-node-proxy.md index 1a68118..6710553 100644 --- a/docs/performance/2026-09-30-three-node-proxy.md +++ b/docs/performance/2026-09-30-three-node-proxy.md @@ -1,6 +1,6 @@ # Qualify three Canopy nodes behind a proxy -**Status: baseline and post-load full-corpus recovery failed; qualification in progress.** This campaign preserves the documented 10,000-repository +**Status: baseline, post-load full-corpus recovery and candidate seed failed; qualification in progress.** This campaign preserves the documented 10,000-repository reference target. A ready fleet, small socket tests or an incomplete seed does not qualify throughput, latency, recovery or production capacity. @@ -66,8 +66,9 @@ failed `lifecycle::lease::startup_preflight_does_not_consume_the_node_lease` with `AddrInUse`; provider qualification and build were skipped in that job. Both used source `a2c59a3`; neither is a local three-node performance comparison. The candidate's local release build and 17-step critical Git check now passed. -Its 10,000-repository seed is in progress; full-corpus recovery and matched load -remain open. The baseline measurements below use the retained `70bd25f` binary, +Its 10,000-repository seed subsequently failed with HTTP 503 after 3,561 recorded +acknowledgements; full-corpus recovery and matched load remain open. The baseline +measurements below use the retained `70bd25f` binary, not the new source pin. The intervening commits add [fenced route reuse and coalesced owner discovery](https://github.com/crabbuild/cellule/pull/32) @@ -94,7 +95,7 @@ immutable copy of the original production executable, not that mutable path. | RustFS | New container `canopy-candidate-rustfs-w0sfigmz`, same pinned ARM64 manifest as the baseline, two CPU quota / 2 GiB memory; fresh bucket and prefix | | Fleet | Three new node identities/directories behind one proxy; unchanged 100-entry admission and 1.5 GiB configured local disk limit per node | | Critical Git behavior | All 17 steps passed: atomic publication/refusal, mixed refusal, correct/stale force-with-lease, shallow/deepen/unshallow, filtered lazy fetch v0/v2, new-object push, fast-forward pull, deletion/pruning, mirror push, credential refusal and exact mirror inventories with strict full fsck | -| Corpus setup | New 10,000 identities with 100 two-commit Git and 1 MiB LFS fixtures requested; seed in progress, not yet complete | +| Corpus setup | Failed: 3,561 of 10,000 identities recorded, including 39 completed two-commit Git/1 MiB LFS fixtures; manifest remains incomplete | | Recovery and performance | Candidate owner-loss, full-corpus recovery, scheduled load and matched comparisons remain unqualified | Artifacts are retained under `canopy-e07670e-candidate-W0SFiGMz` on the @@ -120,9 +121,97 @@ with `NotFound` before readiness; the cause is not isolated. After explicit bucket readiness, `fleet-candidate-ready` started a test-feature executable, then all three nodes shut down gracefully with exit 0 before any corpus was seeded. Neither setup is qualification evidence. The current fleet uses the -immutable production artifact and its own binding. Successful setup does not +immutable production artifact and its own binding; its later failure is recorded +below. Successful setup does not erase the earlier baseline's failed full-corpus recovery. +### Preserve the candidate seed failure + +The original seed terminated with HTTP 503 after 5,885.969 seconds. Its manifest +remains `complete: false`: 3,561 distinct UUID/name pairs and 39 completed Git/LFS +fixtures, not a smaller passing corpus. The creation log records a pending +Directory mutation; an unacknowledged request is not proof of rollback. + +| Boundary, UTC on 2026-10-01 | Observation | +| --- | --- | +| 01:43:24.215207 | Node 2 reported a 14,651 ms successful renewal with 15,347 ms remaining | +| 01:43:33.898105 | Node 0 logged the pending creation result | +| 01:44:01.014171 | Last retained Mac-side seed/provider sample | +| 01:44:11.394524 | Independent UI node stopped lease maintenance with invalid bounds | +| 01:44:11.581900–11.602699 | All three candidate nodes stopped lease maintenance with invalid bounds; 20.799 ms spread | +| After process exit | Launcher retained three exit-1 outcomes, none forcibly killed; seed-to-recovery controller rejected the incomplete manifest before any planned owner-loss action | + +The UI used a separate RustFS container, bucket and identities on the same +Mac/Colima VM. Its lease failure preceded the first candidate failure by +187.376 ms. Both providers were responding after the fault with unchanged +container start times and zero restarts. These are correlations, not proof that +providers were responsive during the stall or that no pause/update occurred. +The UI was restored from preserved data; its later repository refresh is a +separate shared-host intervention, not a benchmark window. + +An independent offline audit checked all 934 provider samples. The maximum +adjacent start gap was 13.037 seconds. In the 17 near-failure samples, UTC-minus- +monotonic variation was 4.057 ms (119.533 ms across the entire observation). +These Mac-side samples cannot exclude VM/storage stalls, and end about +10.568 seconds before the first candidate lease termination. No gap is +interpolated. Rounded Docker gauges do not establish exact request completion, +CPU usage during every interval or historical quota configuration. + +The scoped historical Docker-event query returned no pause/unpause events; +the provider's stdout was empty in the requested failure window, and the +filtered host power query returned no sleep/wake events. None excludes an +unobserved infrastructure interruption. In particular, +[Docker returns only its last 256 historical events](https://docs.docker.com/reference/cli/docker/system/events/), +so a later empty filtered query is not a complete history. + +The candidate provider uses a Mac-backed bind mount at `/data`. A Linux-owned +Docker volume is a falsifiable next storage comparison: retain the exact binary, +image, two-CPU/2-GiB provider limits, admission, leases, 30-second HTTP deadline, +10,000 identities and 100 same-sized Git/LFS fixtures, changing the filesystem +backing in a new disposable deployment. That trial has **not** run. Shared-host +and renewal-contention hypotheses remain open; no Cellule bottleneck or fix is +established. + +### Recover every recorded candidate acknowledgement + +After independently confirming the old owners absent, a new boundary record +waited 32.301 seconds without sending any signals. A fresh three-node fleet +then started against the same durable prefix and immutable binary, with new +identities and local directories. The two critical repositories passed their +four exact v0/v2 ref inventories, payload/notes checks and strict full fsck. + +At the 02:25:27 UTC ledger checkpoint, verification of all 3,561 recorded seed ACKs was +live, with 600 verified. It uses four workers, the unchanged 30-second HTTP +deadline and no retries. Every UUID must match; each of the 39 populated +fixtures must also recover exact Git/LFS bytes, both protocols and strict full +fsck. The original incomplete manifest is never rewritten. Even a complete +ACK recovery cannot qualify a 10,000-repository gate or scheduled performance. + +A separate `observe_ack_recovery.py` companion is now collecting bounded +provider health timings, signed metrics, current pause/quota/restart state, +VM uptime and VM-wide CPU/I/O pressure. A continuous event stream is scoped +to the two fixture container IDs. It stops after 2,400 seconds, verifier exit +or a caller stop; probe errors and unexpected stream exit are retained. The +first 11 samples had no probe errors, but the observation remains incomplete. +Health responses do not establish durable-write latency, VM pressure is not +container attribution, and probes add shared-host overhead. This is diagnostic +evidence, not a performance comparison. The companion's source SHA-256 is +`3e7233d357c41f9f76194c268def6d259cc5cc56202cc73be1bb5b8181876e3a`; +its append-only samples and event stream are under +`ack-recovery-boundary-observer`, separate from the closed failure audit. + +| Closed candidate artifact | SHA-256 | +| --- | --- | +| `corpus-candidate-10000.json` | `5c7d527dc0199d55bbbef30e3a6b707dbdfc661e1d91602d3398f98fe624d7d5` | +| `seed-candidate.log` | `c25f649c85df7c9bc53a605eada98ac581c460ccee84c656c028f31c28c9fe85` | +| `fleet-candidate-matched/outcome.json` | `08318903e813bb24a865ea605157533bef77c0db143486f4215689cedb42ef41` | +| `seed-to-recovery.json` | `7e92791a2217668ef045a431ab210f13c56710ee2631b37d031aed4ef47e071c` | +| `failed-candidate-owner-exit.json` | `7e3e10b51c3c90c950566b0710121e311dadc9de4391e6e8e14d204d1ff44f28` | +| `failed-candidate-seed-audit.json` | `c73c9b2a9ae6242b1275d961b11c946f779ab580f76ce5c4f81913c6dd5bb9a6` | +| `audit_seed_failure.py` | `42daca70ada182e272a198aceb4568834a9a9fa206ac729193f9d64ebd1558c1` | +| `fleet-after-seed-failure/ready.json` | `aaa089a05e1b140ea9b0973248290aa07d6fcb11309eb29e05f570bb5bec6920` | +| `critical-after-failed-seed-verification.json` | `1031bf601b1e63970198a385fe4f7bacae0124b6d9861a8254729f371967ef31` | + ## Retain failed setup and incomplete work | Check | Evidence / status | @@ -434,24 +523,24 @@ the baseline's provider latency or prove a performance fix. The Cellule ### Advance only from a complete candidate seed -The retained one-off `finish_seed_recovery.py` transition is now running, bound +The retained one-off `finish_seed_recovery.py` transition was armed, bound to seed PID 48549, the immutable candidate binary, current provider and original three owner process identities. Its live read-only binding check passed before -execution. It waits for that exact seed to end and requires 10,000 distinct +execution. It required that exact seed to end with 10,000 distinct canonical UUID/name pairs and the same 100 two-commit Git/1 MiB LFS fixtures -before any signals. Incomplete seed, changed source/processes, missing owners -or a provider restart stop it without a retry, reseed or replacement fixture. +before any signals. The incomplete seed stopped it with `ValueError` in +`wait_bound_seed`, without any signals, retry, reseed or replacement fixture. -After successful seeding, it records three verified-owner SIGKILLs and process -absence, waits at least 32 seconds, starts fresh directories on the same durable -deployment, and verifies both critical fixtures and the complete original +Its unexecuted complete-seed path would record three verified-owner SIGKILLs and +process absence, wait at least 32 seconds, start fresh directories on the same +durable deployment, and verify both critical fixtures and the complete original corpus. HTTP timeout stays 30 seconds, verification concurrency stays four, and admission stays 100/node. It does not start scheduled load. At this -observation its phase is `wait_bound_seed`, not a recovery pass. +terminal observation its phase remained `wait_bound_seed`; it is not a recovery pass. Four simulated control-flow tests passed: complete transition, incomplete seed, provider restart and missing node. They send no real signals or provider/Git -requests and are not recovery evidence. The live transition retains +requests and are not recovery evidence. The terminated transition retains `seed-to-recovery.json`, fault boundaries and individual verification receipts. | Candidate-volume control artifact | SHA-256 | @@ -468,6 +557,12 @@ Rust tests, isolated RustFS compatibility, server build and the 66-test harness: [36799975428](https://github.com/crabbuild/canopy/actions/runs/36799975428). Those earlier commits did not contain the new expired-reply test. +Both full workflows for `39829f7` also passed, including the new expired-reply +test, workspace checks, isolated RustFS compatibility/build and 66-test harness: +[36802086544](https://github.com/crabbuild/canopy/actions/runs/36802086544) and +[36802089950](https://github.com/crabbuild/canopy/actions/runs/36802089950). +CI success does not convert the local candidate seed failure to a performance pass. + ## Measure the requested operations Complete and verify the corpus before interpreting scheduled load results. @@ -582,8 +677,9 @@ The reported transfer rate is pack bytes divided by the entire POST duration, no steady-state capacity. Serial repetitions can warm server state; neither the identity check nor a new HTTP connection establishes cold ownership. Discovery, local validation, v2 negotiation, stock clone performance, native Git child CPU -and provider request-cost accounting require their own evidence. The running -candidate seed and its bound harness are unchanged. +and provider request-cost accounting require their own evidence. The candidate +seed's bound harness was unchanged during these observations; its later terminal +failure is recorded above. The campaign samples each node, the proxy launcher and its own driver once per second with `ps`: RSS, cumulative process CPU seconds and raw CPU percentage. From dafabf8fdc9acf0eb65c75f285054094425f8b96 Mon Sep 17 00:00:00 2001 From: forhappy Date: Wed, 30 Sep 2026 19:37:14 -0700 Subject: [PATCH 12/60] Track latest Cellule docs-only revision without replacing bound runtime --- Cargo.lock | 12 +++++----- crates/canopy-server/Cargo.toml | 10 ++++---- .../2026-09-30-three-node-proxy.md | 23 ++++++++++++++++++- 3 files changed, 33 insertions(+), 12 deletions(-) diff --git a/Cargo.lock b/Cargo.lock index 9852050..30771b7 100644 --- a/Cargo.lock +++ b/Cargo.lock @@ -448,7 +448,7 @@ dependencies = [ [[package]] name = "cellule-app" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=e07670e2348231ed401cc7280a47e3ab97596ffe#e07670e2348231ed401cc7280a47e3ab97596ffe" +source = "git+https://github.com/crabbuild/cellule.git?rev=0573f48998c4e5343cd8b463d79b7bc1820c923c#0573f48998c4e5343cd8b463d79b7bc1820c923c" dependencies = [ "blake3", "cellule-runtime", @@ -457,7 +457,7 @@ dependencies = [ [[package]] name = "cellule-host" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=e07670e2348231ed401cc7280a47e3ab97596ffe#e07670e2348231ed401cc7280a47e3ab97596ffe" +source = "git+https://github.com/crabbuild/cellule.git?rev=0573f48998c4e5343cd8b463d79b7bc1820c923c#0573f48998c4e5343cd8b463d79b7bc1820c923c" dependencies = [ "cellule-app", "cellule-runtime", @@ -471,7 +471,7 @@ dependencies = [ [[package]] name = "cellule-ltx" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=e07670e2348231ed401cc7280a47e3ab97596ffe#e07670e2348231ed401cc7280a47e3ab97596ffe" +source = "git+https://github.com/crabbuild/cellule.git?rev=0573f48998c4e5343cd8b463d79b7bc1820c923c#0573f48998c4e5343cd8b463d79b7bc1820c923c" dependencies = [ "async-trait", "blake3", @@ -493,7 +493,7 @@ dependencies = [ [[package]] name = "cellule-runtime" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=e07670e2348231ed401cc7280a47e3ab97596ffe#e07670e2348231ed401cc7280a47e3ab97596ffe" +source = "git+https://github.com/crabbuild/cellule.git?rev=0573f48998c4e5343cd8b463d79b7bc1820c923c#0573f48998c4e5343cd8b463d79b7bc1820c923c" dependencies = [ "blake3", "bytes", @@ -519,7 +519,7 @@ dependencies = [ [[package]] name = "cellule-store" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=e07670e2348231ed401cc7280a47e3ab97596ffe#e07670e2348231ed401cc7280a47e3ab97596ffe" +source = "git+https://github.com/crabbuild/cellule.git?rev=0573f48998c4e5343cd8b463d79b7bc1820c923c#0573f48998c4e5343cd8b463d79b7bc1820c923c" dependencies = [ "async-trait", "blake3", @@ -542,7 +542,7 @@ dependencies = [ [[package]] name = "cellule-types" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=e07670e2348231ed401cc7280a47e3ab97596ffe#e07670e2348231ed401cc7280a47e3ab97596ffe" +source = "git+https://github.com/crabbuild/cellule.git?rev=0573f48998c4e5343cd8b463d79b7bc1820c923c#0573f48998c4e5343cd8b463d79b7bc1820c923c" dependencies = [ "schemars", "serde", diff --git a/crates/canopy-server/Cargo.toml b/crates/canopy-server/Cargo.toml index 0b97c61..1d91493 100644 --- a/crates/canopy-server/Cargo.toml +++ b/crates/canopy-server/Cargo.toml @@ -15,11 +15,11 @@ axum = "0.8.9" base64 = "0.22" blake3 = "1.8" bytes = "1.11" -cellule-app = { git = "https://github.com/crabbuild/cellule.git", rev = "e07670e2348231ed401cc7280a47e3ab97596ffe" } -cellule-host = { git = "https://github.com/crabbuild/cellule.git", rev = "e07670e2348231ed401cc7280a47e3ab97596ffe" } -cellule-ltx = { git = "https://github.com/crabbuild/cellule.git", rev = "e07670e2348231ed401cc7280a47e3ab97596ffe", features = ["replica"] } -cellule-runtime = { git = "https://github.com/crabbuild/cellule.git", rev = "e07670e2348231ed401cc7280a47e3ab97596ffe" } -cellule-store = { git = "https://github.com/crabbuild/cellule.git", rev = "e07670e2348231ed401cc7280a47e3ab97596ffe" } +cellule-app = { git = "https://github.com/crabbuild/cellule.git", rev = "0573f48998c4e5343cd8b463d79b7bc1820c923c" } +cellule-host = { git = "https://github.com/crabbuild/cellule.git", rev = "0573f48998c4e5343cd8b463d79b7bc1820c923c" } +cellule-ltx = { git = "https://github.com/crabbuild/cellule.git", rev = "0573f48998c4e5343cd8b463d79b7bc1820c923c", features = ["replica"] } +cellule-runtime = { git = "https://github.com/crabbuild/cellule.git", rev = "0573f48998c4e5343cd8b463d79b7bc1820c923c" } +cellule-store = { git = "https://github.com/crabbuild/cellule.git", rev = "0573f48998c4e5343cd8b463d79b7bc1820c923c" } ed25519-dalek = "2" flate2 = "1.1" futures-core = "0.3" diff --git a/docs/performance/2026-09-30-three-node-proxy.md b/docs/performance/2026-09-30-three-node-proxy.md index 6710553..a1b31ed 100644 --- a/docs/performance/2026-09-30-three-node-proxy.md +++ b/docs/performance/2026-09-30-three-node-proxy.md @@ -56,7 +56,7 @@ credentials were not changed. No authority check or TLS verification was disable ## Prepare the next upstream candidate Cellule `origin/main` advanced during the running baseline. Canopy's manifests -and six Cellule lockfile entries now select +and six Cellule lockfile entries first selected `e07670e2348231ed401cc7280a47e3ab97596ffe`. Dependency resolution and locked metadata validation passed without updating unrelated packages or creating a local Cargo target directory. [Workflow 36789746786](https://github.com/crabbuild/canopy/actions/runs/36789746786) @@ -79,6 +79,27 @@ were unchanged during load. CI verification can run independently; local compila fresh-state recovery and matched performance comparisons must follow the recorded baseline and produce their own artifact bindings. +### Track the latest source without replacing the bound runtime + +Upstream subsequently advanced to `0573f48998c4e5343cd8b463d79b7bc1820c923c` +with [the Cellule website/documentation change](https://github.com/crabbuild/cellule/pull/34). +Canopy's five direct dependency declarations and six lockfile source entries +now select that revision, with no unrelated dependency changes. Locked Cargo +metadata resolved all six packages to the exact new commit; no workspace +`target` directory or new executable was created. + +`git diff --exit-code e07670e 0573f489 -- Cargo.toml Cargo.lock crates` passed +in the Cellule checkout. Both commits have the same `crates` tree object +`587e3215ce1bc0c67db3390a137623b6fead01ac`, root Cargo manifest blob +`6b7e13d8439558455ea04458357f3a70992fa367` and lockfile blob +`b512761a55bba5165101ce9905dfe7f603f5b079`. This upstream update is not a +runtime performance fix. A new production build, hosted verification and +end-to-end qualification for the new source pin remain open; existing binary +bindings and the live ACK recovery still identify `e07670e`. + +The candidate-volume `metadata-0573f489.json` SHA-256 is +`8cbbca85086cd846f265c97d9c0ca50f85c33b795c87ec2e40405a7db8c879e9`. + ### Bind the local candidate separately The locked production release built with one Cargo job in 578.311 seconds. From ef55688e200a5b07bc53ebc44be93515eb7a281e Mon Sep 17 00:00:00 2001 From: forhappy Date: Wed, 30 Sep 2026 20:28:39 -0700 Subject: [PATCH 13/60] docs: record audited ACK recovery and native storage trial --- docs/performance-plan.md | 27 +++-- .../2026-09-30-three-node-proxy.md | 44 +++++--- .../2026-10-01-native-filesystem.md | 104 ++++++++++++++++++ 3 files changed, 152 insertions(+), 23 deletions(-) create mode 100644 docs/performance/2026-10-01-native-filesystem.md diff --git a/docs/performance-plan.md b/docs/performance-plan.md index bf0b95f..3ad4503 100644 --- a/docs/performance-plan.md +++ b/docs/performance-plan.md @@ -2,15 +2,24 @@ Use this plan to design capacity work and interpret Canopy benchmark results. It separates demonstrated behavior from proposed targets. The current Cellule -dependency candidate is `e07670e2348231ed401cc7280a47e3ab97596ffe`, fetched from -`origin/main` during the [three-node proxy campaign](performance/2026-09-30-three-node-proxy.md#prepare-the-next-upstream-candidate). -One hosted workflow passed its build, workspace tests and isolated RustFS -qualification; a sibling failed a delayed-startup test with `AddrInUse`. -That fixture race was reproduced and repaired without changing production -authority checks; all four lease lifecycle tests passed locally. A separately -bound production release passed 17 critical stock-Git checks through three -nodes and a proxy against RustFS. Its new 10,000-repository seed is in progress; -complete-corpus recovery and matched performance verification remain open. +dependency pin is `0573f48998c4e5343cd8b463d79b7bc1820c923c`, a docs-only +upstream advance from `e07670e` with identical Rust workspace/crates. Both full +hosted workflows passed; that does not establish three-node capacity. +The [three-node proxy campaign](performance/2026-09-30-three-node-proxy.md) +retains the earlier fixture race and its regression, plus the candidate's failed +seed and successful recovery of every recorded ACK. + +| Current gate | Evidence / status | +| --- | --- | +| Candidate Git behavior | Bound `e07670e` production artifact passed all 17 critical stock-Git steps through three nodes/proxy against RustFS | +| Mac-backed candidate seed | Failed: 3,561/10,000 identities and 39 Git/LFS fixtures; incomplete manifest retained | +| Every candidate seed ACK | All 3,561 identities and 39 populated fixtures passed fresh-owner recovery; independent ledger/binding and 78-clone audit passed | +| Native-volume comparison | Same runtime, provider image/limits and full corpus requested; critical Git setup passed, full seed running | +| Full recovery and scheduled load | Still unqualified; no incomplete seed or narrow functional test closes these gates | + +The [native-filesystem trial](performance/2026-10-01-native-filesystem.md) +records the storage hypothesis, graceful diagnostic drain and new bindings. +No Cellule bottleneck or matched performance improvement is proven. The [workspace/RustFS verification](performance/2026-09-30-workspace-rustfs.md) and the retained three-node baseline use `70bd25f142f1976fdd63ffe60e46e15ae276ffdc`. That baseline stopped after 20 fully bound windows and one interrupted creation diff --git a/docs/performance/2026-09-30-three-node-proxy.md b/docs/performance/2026-09-30-three-node-proxy.md index a1b31ed..1e665e1 100644 --- a/docs/performance/2026-09-30-three-node-proxy.md +++ b/docs/performance/2026-09-30-three-node-proxy.md @@ -93,8 +93,11 @@ in the Cellule checkout. Both commits have the same `crates` tree object `587e3215ce1bc0c67db3390a137623b6fead01ac`, root Cargo manifest blob `6b7e13d8439558455ea04458357f3a70992fa367` and lockfile blob `b512761a55bba5165101ce9905dfe7f603f5b079`. This upstream update is not a -runtime performance fix. A new production build, hosted verification and -end-to-end qualification for the new source pin remain open; existing binary +runtime performance fix. Both full hosted workflows for `0573f489` passed: +[36806695489](https://github.com/crabbuild/canopy/actions/runs/36806695489) and +[36806698917](https://github.com/crabbuild/canopy/actions/runs/36806698917). +A new local production build and three-node end-to-end qualification for the +new source pin remain open; existing binary bindings and the live ACK recovery still identify `e07670e`. The candidate-volume `metadata-0573f489.json` SHA-256 is @@ -189,7 +192,8 @@ The candidate provider uses a Mac-backed bind mount at `/data`. A Linux-owned Docker volume is a falsifiable next storage comparison: retain the exact binary, image, two-CPU/2-GiB provider limits, admission, leases, 30-second HTTP deadline, 10,000 identities and 100 same-sized Git/LFS fixtures, changing the filesystem -backing in a new disposable deployment. That trial has **not** run. Shared-host +backing in a new disposable deployment. The [native-volume trial](2026-10-01-native-filesystem.md) +is now running after audited ACK recovery and graceful diagnostic drain. Shared-host and renewal-contention hypotheses remain open; no Cellule bottleneck or fix is established. @@ -201,26 +205,33 @@ then started against the same durable prefix and immutable binary, with new identities and local directories. The two critical repositories passed their four exact v0/v2 ref inventories, payload/notes checks and strict full fsck. -At the 02:25:27 UTC ledger checkpoint, verification of all 3,561 recorded seed ACKs was -live, with 600 verified. It uses four workers, the unchanged 30-second HTTP -deadline and no retries. Every UUID must match; each of the 39 populated -fixtures must also recover exact Git/LFS bytes, both protocols and strict full -fsck. The original incomplete manifest is never rewritten. Even a complete -ACK recovery cannot qualify a 10,000-repository gate or scheduled performance. +Verification completed with exit 0: all 3,561 recorded ACKs and all 39 populated +fixtures passed with four workers, the unchanged 30-second HTTP deadline and +no retries. Each UUID and exact Git/LFS payload matched, with both protocols and +strict full fsck. The independent audit checked exact unique ledger sets and +bindings, then rechecked all 78 local Git clones. The original incomplete +manifest was never rewritten. This ACK recovery cannot qualify a +10,000-repository gate or scheduled performance. -A separate `observe_ack_recovery.py` companion is now collecting bounded +A separate `observe_ack_recovery.py` companion collected bounded provider health timings, signed metrics, current pause/quota/restart state, VM uptime and VM-wide CPU/I/O pressure. A continuous event stream is scoped -to the two fixture container IDs. It stops after 2,400 seconds, verifier exit -or a caller stop; probe errors and unexpected stream exit are retained. The -first 11 samples had no probe errors, but the observation remains incomplete. +to the two fixture container IDs. It completed after verifier exit with 173 +samples and zero probe errors; the event stream was then stopped intentionally. +An offline audit checked the digests, probe bounds, state snapshots and events. Health responses do not establish durable-write latency, VM pressure is not container attribution, and probes add shared-host overhead. This is diagnostic evidence, not a performance comparison. The companion's source SHA-256 is `3e7233d357c41f9f76194c268def6d259cc5cc56202cc73be1bb5b8181876e3a`; -its append-only samples and event stream are under +its closed samples and event stream are under `ack-recovery-boundary-observer`, separate from the closed failure audit. +The diagnostic fleet subsequently drained with three unforced exit-0 outcomes. +Only its verified launcher received SIGTERM. The old dedicated provider was +stopped without deleting its data; the separate UI stayed live. See the +[native-volume trial record](2026-10-01-native-filesystem.md) for the completed +ACK audit, drain boundary and new fixture bindings. + | Closed candidate artifact | SHA-256 | | --- | --- | | `corpus-candidate-10000.json` | `5c7d527dc0199d55bbbef30e3a6b707dbdfc661e1d91602d3398f98fe624d7d5` | @@ -232,6 +243,11 @@ its append-only samples and event stream are under | `audit_seed_failure.py` | `42daca70ada182e272a198aceb4568834a9a9fa206ac729193f9d64ebd1558c1` | | `fleet-after-seed-failure/ready.json` | `aaa089a05e1b140ea9b0973248290aa07d6fcb11309eb29e05f570bb5bec6920` | | `critical-after-failed-seed-verification.json` | `1031bf601b1e63970198a385fe4f7bacae0124b6d9861a8254729f371967ef31` | +| `failed-seed-ack-verification.json` | `e9d6e6c804367a69f93c59fae57f98f90062b7fa48393e107ae8759f8f473167` | +| `failed-seed-ack-verification.jsonl` | `9290073f8b460d31a2cb21fcd62805759cf3e31a413c681cb9a3c195cf4d2008` | +| `failed-seed-ack-recovery-audit.json` | `01f2daa462f7f5ee321fc30dba813013ee2fe00415ff60d210224cace0355350` | +| `ack-recovery-observer-audit.json` | `866419ea5503df3ca685789f333f8d8fed03f5ed734b2f8bf8ac822099ef9dec` | +| `ack-recovery-graceful-drain.json` | `23a29d2a479e814b117e610a580a504cc4a7e442871e1908a40694eacaf44c11` | ## Retain failed setup and incomplete work diff --git a/docs/performance/2026-10-01-native-filesystem.md b/docs/performance/2026-10-01-native-filesystem.md new file mode 100644 index 0000000..aa517e1 --- /dev/null +++ b/docs/performance/2026-10-01-native-filesystem.md @@ -0,0 +1,104 @@ +# Compare RustFS filesystem backing + +**Status: full native-volume seed running; recovery and performance unqualified.** +This shared-Mac/Colima diagnostic keeps the 10,000-repository target. It tests +filesystem backing without replacing the bound runtime or weakening correctness. + +## Preserve the failed run first + +| Gate | Observed result | +| --- | --- | +| Original candidate seed | Failed with HTTP 503 after 3,561 identities and 39 Git/LFS fixtures; original manifest remains incomplete | +| ACK recovery | All 3,561 exact UUID/name pairs and all 39 populated fixtures passed through fresh owners, including v0/v2 clones, exact Git/LFS bytes and strict full fsck | +| Independent ACK audit | Exact unique ledger sets and artifact bindings passed; all 78 local Git clones were rechecked with strict full fsck | +| Critical fixtures | Both identities and four exact v0/v2 ref inventories, payload and notes passed | +| Diagnostic drain | Three unforced exit-0 outcomes after verified launcher SIGTERM; old RustFS stopped with data preserved | +| UI | Separate preview remains live; no UI process or provider was stopped | + +The ACK verifier's final report is `e9d6e6c804367a69f93c59fae57f98f90062b7fa48393e107ae8759f8f473167`; +its ledger is `9290073f8b460d31a2cb21fcd62805759cf3e31a413c681cb9a3c195cf4d2008`. +The independent audit is `01f2daa462f7f5ee321fc30dba813013ee2fe00415ff60d210224cace0355350`. +These SHA-256 bindings refer to artifacts in `canopy-e07670e-candidate-W0SFiGMz`, +not a passing 10,000-repository run. Unacknowledged creations are not rollback +assertions. The [failed campaign record](2026-09-30-three-node-proxy.md#preserve-the-candidate-seed-failure) +retains the original failure and process-loss boundary. + +## Change the storage backing, not the workload + +| Input | Native-volume trial | +| --- | --- | +| Runtime | Same immutable production executable, Cellule `e07670e`; binary SHA-256 `e90728c60cbb941cc8caa1c698bc24bcf6edda1935d98853f0c0319933668f16` | +| RustFS | Same pinned ARM64 manifest `sha256:0c3c7030ffb93afde8d359fb1db957b85033ede05115518bd0dede51f4353f6a` | +| Provider limits | Two CPU quota, 2 GiB memory, same swap/restart/network settings | +| Changed backing | Docker local volume `canopy-native-data-cexfad8i` at `/data`, replacing the Mac-backed bind mount in a new fixture | +| Topology | Three new Canopy processes behind one loopback TCP proxy; verified TLS peer transport | +| Node limits | 100 active repository entries and 1.5 GiB configured local disk per node | +| Corpus | 10,000 identities; 100 two-commit Git fixtures and 100 1 MiB LFS bodies; same RNG seed `20260926` | +| Deadlines | Unchanged 30-second HTTP and 120-second stock-Git deadlines | +| Deployment | New bucket/prefix and identities; failed data is neither resumed nor deleted | + +The first cold-preparation script failed an exact environment-list ordering +check before RustFS started. An independent reconciliation checked unique +environment names and identical values, image, entrypoint, arguments, resource +settings and owned volume. Only environment ordering differed. The failed +preparation receipt remains preserved; it was not rerun or overwritten. + +The native fleet passed all 17 critical Git checks. At **03:10:16 UTC on +2026-10-01**, its incomplete seed contained **3,575 identities and 40 populated +fixtures**. Passing the old failure's repository count does not prove a fix. +Different wall-clock load and observer overhead remain shared-host variables. + +```mermaid +flowchart LR + ack[Every failed-seed ACK verified] --> drain[Graceful diagnostic drain] + drain --> seed[Native full 10000 seed
running] + seed --> loss[Record three-owner loss] + loss --> recovery[Verify all original identities
and Git/LFS fixtures] + recovery --> load[Declared 108-window load plan] + load --> final[Second owner loss
verify every load ACK and corpus] +``` + +Only the first two stages are complete. Critical Git setup also passed; +native owner-loss recovery, the [declared load matrix](three-node-baseline.json), +higher admission profiles, additional critical-operation load/fault coverage +and matched performance comparisons remain open. No stage can substitute a +smaller corpus or retries for its required evidence. + +## Separate diagnostics from performance + +The closed ACK-recovery observer retained 173 samples with zero probe errors. +During that window, maximum provider health latency was 457.624 ms, signed +metrics latency 265.171 ms and VM probe latency 354.850 ms. It captured 522 +scoped container events without observed pause/unpause/update/restart/kill/OOM +events; 29 state snapshots had unchanged starts, quotas and zero restarts. +Its independent audit is +`866419ea5503df3ca685789f333f8d8fed03f5ed734b2f8bf8ac822099ef9dec`. +This window was after the failed seed, not an observation of the original stall. +Health and metrics timings do not establish durable-write latency. + +The separate native-seed observer started after 1,250 identities, retaining +provider gauges, signed counters, health timings, VM uptime/CPU/I/O pressure, +kernel self/reaped-child CPU and continuous scoped events. Errors and gaps stay +visible. It adds host overhead and cannot account for the whole seed, live child +CPU, per-operation CPU/GiB or priced object-store cost. + +## Bind the new trial + +Artifacts are retained under `canopy-native-filesystem-ceXFad8I`: + +| Closed input or receipt | SHA-256 | +| --- | --- | +| `provider-binding.json` | `aa6a531365f550611618c91ce8b6cd88974811f543765915f1436bf71a8ec570` | +| `start_full_trial.py` | `973e1c3cdefaf6ef221d9351970d481e7f8ba1c5773b55ae9ea86d74ed1b45d2` | +| `fleet-native-seed/ready.json` | `aaa9769fa8dfa6fa85491a32881448e9b71fe9941ff798b5b8058f8a29720ef9` | +| `critical-native.json` | `2b019dea0a6e89fe26f22d20860775860469ae2a2fd743f65d836e95d03ca04d` | +| `observe_native_seed.py` | `621ee99b88bb839541165417bb3a4c2c333edadfdbff0003f70d025498f119b3` | + +The corpus, controller report and observer outputs are still live and do not +have final completion digests. Setup timings are not scheduled throughput. +This is not isolated Linux reference capacity, a proven Cellule bottleneck or +a passing latest-pin executable comparison. Both hosted workflows for source +pin `0573f489` passed +([36806695489](https://github.com/crabbuild/canopy/actions/runs/36806695489), +[36806698917](https://github.com/crabbuild/canopy/actions/runs/36806698917)); +the running comparison binary remains the original `e07670e` artifact. From b44fac81c6afc00c865a009cf4f475ca41c364ff Mon Sep 17 00:00:00 2001 From: forhappy Date: Wed, 30 Sep 2026 20:42:56 -0700 Subject: [PATCH 14/60] docs: bind full native recovery and load transitions --- .../2026-10-01-native-filesystem.md | 40 ++++++++++++++++++- 1 file changed, 38 insertions(+), 2 deletions(-) diff --git a/docs/performance/2026-10-01-native-filesystem.md b/docs/performance/2026-10-01-native-filesystem.md index aa517e1..f5205ae 100644 --- a/docs/performance/2026-10-01-native-filesystem.md +++ b/docs/performance/2026-10-01-native-filesystem.md @@ -43,8 +43,8 @@ environment names and identical values, image, entrypoint, arguments, resource settings and owned volume. Only environment ordering differed. The failed preparation receipt remains preserved; it was not rerun or overwritten. -The native fleet passed all 17 critical Git checks. At **03:10:16 UTC on -2026-10-01**, its incomplete seed contained **3,575 identities and 40 populated +The native fleet passed all 17 critical Git checks. At **03:41:52 UTC on +2026-10-01**, its incomplete seed contained **9,275 identities and 95 populated fixtures**. Passing the old failure's repository count does not prove a fix. Different wall-clock load and observer overhead remain shared-host variables. @@ -64,6 +64,34 @@ higher admission profiles, additional critical-operation load/fault coverage and matched performance comparisons remain open. No stage can substitute a smaller corpus or retries for its required evidence. +## Gate the transition and complete load plan + +Two separately bound controllers are armed, not completed. The recovery +controller waits for the exact seed and its setup controller to exit, then +rereads their final receipt. A failed/incomplete corpus, changed process +identity, provider restart/pause/quota/mount change or source drift stops it +before any owner-loss action. + +Only a full 10,000-identity manifest with the original RNG-selected 100 exact +two-commit/LFS fixtures can reach the fault stage. The controller verifies the +whole PID/command/kernel-identity batch before sending three SIGKILLs, records +process absence and the unchanged 32-second expiry wait, and starts fresh +owners from the preserved deployment. Every original identity and Git/LFS +fixture, plus both critical repositories, must pass without retries. + +The load controller requires that closed recovery evidence, all source and +receipt digests, three recorded SIGKILLs and fresh owners. Its plan is the +unchanged 108 windows: 114,960 offered arrivals over 8,640 offered seconds. +Failed arrivals, interrupted windows and diagnostic gaps remain visible; +post-load verification of every ACK is a separate, still-open gate. + +Offline guard tests passed: four recovery tests cover the simulated complete +path, eight failure-control cases, provider-boundary checks and actual full +manifest/fixture validation; two load tests cover closed-recovery evidence and +five no-launch failure cases. All signals, provider calls, Git and load traffic +were mocked in control tests. Live read-only binding checks also passed before +arming. These are controller guards, not actual owner-loss or performance proof. + ## Separate diagnostics from performance The closed ACK-recovery observer retained 173 samples with zero probe errors. @@ -93,6 +121,14 @@ Artifacts are retained under `canopy-native-filesystem-ceXFad8I`: | `fleet-native-seed/ready.json` | `aaa9769fa8dfa6fa85491a32881448e9b71fe9941ff798b5b8058f8a29720ef9` | | `critical-native.json` | `2b019dea0a6e89fe26f22d20860775860469ae2a2fd743f65d836e95d03ca04d` | | `observe_native_seed.py` | `621ee99b88bb839541165417bb3a4c2c333edadfdbff0003f70d025498f119b3` | +| `finish_native_recovery.py` | `2dac2dd18bdbd416c95ab9ff77cbfb8566cad7bb4ae89dc5c6fa55b51c00500e` | +| `test_finish_native_recovery.py` | `bd919dfdff557d6e5fa59a6647d7731661efa37ec7854efd093cf0a08ab4b837` | +| `transition-guard-tests-final.log` | `9155d1771a85043e16da955d5f7f61465a4ca3da7df0518f116cc2093bce6568` | +| `transition-check-final.json` | `d93ac46cef8392eb61dbc666d49a69be89a46841b5e22b9d0bf464aafb1ac865` | +| `start_native_load.py` | `61d7828f2441c2d12061fce9fdb2c366ceb60c746e75322ee8b615e410c4645e` | +| `test_start_native_load.py` | `9a1009d2f57987931beeeead14cb8e1482e3ffd3bbf6f9bb7520a5cd77ae0e0b` | +| `load-guard-tests-final.log` | `6b307282780ae288f168cfceef2ff771a215534923036eb6875258d24e00622d` | +| `load-check-final.json` | `3076fafdfc8881746f728f8db498b08731676e33cba1e3181df7a7092e2f4a54` | The corpus, controller report and observer outputs are still live and do not have final completion digests. Setup timings are not scheduled throughput. From db569d47a4ef5780c24b68f251ea584f5c749012 Mon Sep 17 00:00:00 2001 From: forhappy Date: Wed, 30 Sep 2026 20:50:19 -0700 Subject: [PATCH 15/60] docs: record complete native seed and owner-loss boundary --- docs/performance-plan.md | 2 +- .../2026-10-01-native-filesystem.md | 59 ++++++++++++++----- 2 files changed, 44 insertions(+), 17 deletions(-) diff --git a/docs/performance-plan.md b/docs/performance-plan.md index 3ad4503..58cec27 100644 --- a/docs/performance-plan.md +++ b/docs/performance-plan.md @@ -14,7 +14,7 @@ seed and successful recovery of every recorded ACK. | Candidate Git behavior | Bound `e07670e` production artifact passed all 17 critical stock-Git steps through three nodes/proxy against RustFS | | Mac-backed candidate seed | Failed: 3,561/10,000 identities and 39 Git/LFS fixtures; incomplete manifest retained | | Every candidate seed ACK | All 3,561 identities and 39 populated fixtures passed fresh-owner recovery; independent ledger/binding and 78-clone audit passed | -| Native-volume comparison | Same runtime, provider image/limits and full corpus requested; critical Git setup passed, full seed running | +| Native-volume comparison | Same runtime/provider limits; all 10,000 identities and 100 Git/LFS fixtures seeded in 3,179.240 s; full owner-loss recovery running | | Full recovery and scheduled load | Still unqualified; no incomplete seed or narrow functional test closes these gates | The [native-filesystem trial](performance/2026-10-01-native-filesystem.md) diff --git a/docs/performance/2026-10-01-native-filesystem.md b/docs/performance/2026-10-01-native-filesystem.md index f5205ae..b719124 100644 --- a/docs/performance/2026-10-01-native-filesystem.md +++ b/docs/performance/2026-10-01-native-filesystem.md @@ -1,6 +1,6 @@ # Compare RustFS filesystem backing -**Status: full native-volume seed running; recovery and performance unqualified.** +**Status: full native-volume seed complete; owner-loss recovery running; performance unqualified.** This shared-Mac/Colima diagnostic keeps the 10,000-repository target. It tests filesystem backing without replacing the bound runtime or weakening correctness. @@ -43,32 +43,36 @@ environment names and identical values, image, entrypoint, arguments, resource settings and owned volume. Only environment ordering differed. The failed preparation receipt remains preserved; it was not rerun or overwritten. -The native fleet passed all 17 critical Git checks. At **03:41:52 UTC on -2026-10-01**, its incomplete seed contained **9,275 identities and 95 populated -fixtures**. Passing the old failure's repository count does not prove a fix. +The native fleet passed all 17 critical Git checks. The full seed completed +with exit 0 on 2026-10-01: **10,000 identities and 100 populated Git/LFS +fixtures in 3,179.240 seconds**. An independent offline audit checked exact +corpus identities, fixture selection/digests and controller/source bindings. +Seed wall time is not scheduled creation, push or fetch throughput. A complete +seed alone does not prove recovery, a root-cause fix or matched improvement. Different wall-clock load and observer overhead remain shared-host variables. ```mermaid flowchart LR ack[Every failed-seed ACK verified] --> drain[Graceful diagnostic drain] - drain --> seed[Native full 10000 seed
running] - seed --> loss[Record three-owner loss] - loss --> recovery[Verify all original identities
and Git/LFS fixtures] + drain --> seed[Native full 10000 seed
complete] + seed --> loss[Three-owner loss and wait
recorded] + loss --> recovery[All identities and Git/LFS
verification running] recovery --> load[Declared 108-window load plan] load --> final[Second owner loss
verify every load ACK and corpus] ``` -Only the first two stages are complete. Critical Git setup also passed; -native owner-loss recovery, the [declared load matrix](three-node-baseline.json), +The seed and recorded fault boundary are also complete. Critical Git setup and +critical-fixture recovery passed; full native owner-loss recovery, the +[declared load matrix](three-node-baseline.json), higher admission profiles, additional critical-operation load/fault coverage and matched performance comparisons remain open. No stage can substitute a smaller corpus or retries for its required evidence. ## Gate the transition and complete load plan -Two separately bound controllers are armed, not completed. The recovery -controller waits for the exact seed and its setup controller to exit, then -rereads their final receipt. A failed/incomplete corpus, changed process +Two separately bound controllers are running, not completed. The recovery +controller waited for the exact seed and its setup controller to exit, then +reread their final receipt. A failed/incomplete corpus, changed process identity, provider restart/pause/quota/mount change or source drift stops it before any owner-loss action. @@ -79,6 +83,15 @@ process absence and the unchanged 32-second expiry wait, and starts fresh owners from the preserved deployment. Every original identity and Git/LFS fixture, plus both critical repositories, must pass without retries. +The actual fault sent SIGKILL to PIDs 64706/64748/64836 in one verified batch. +All three exited -9; the launcher recorded no forced fallback. After recorded +process absence, the expiry wait lasted **32.037 seconds**. A new fleet started +at 03:45:31 UTC with three new IDs/PIDs and local directories, the same durable +deployment, admission and immutable binary. Both critical repositories passed +four exact v0/v2 ref inventories, payload/notes and strict full fsck. Full +10,000-identity/100-fixture recovery was live at the 1,500-identity checkpoint; +no load window had started at that checkpoint. + The load controller requires that closed recovery evidence, all source and receipt digests, three recorded SIGKILLs and fresh owners. Its plan is the unchanged 108 windows: 114,960 offered arrivals over 8,640 offered seconds. @@ -104,7 +117,13 @@ Its independent audit is This window was after the failed seed, not an observation of the original stall. Health and metrics timings do not establish durable-write latency. -The separate native-seed observer started after 1,250 identities, retaining +The separate native-seed observer started after 1,250 identities and closed +after seed exit: 459 samples, zero probe errors, 77 unchanged provider-state +snapshots and 1,377 scoped events with no observed pause/unpause/update/restart/ +kill/OOM events in that observation window. An offline integrity audit checked +timing bounds, kernel identities/counter continuity and signed-counter deltas. +This excludes the later intentional owner loss, not all possible shared-host +interventions or the old failed run. It retained provider gauges, signed counters, health timings, VM uptime/CPU/I/O pressure, kernel self/reaped-child CPU and continuous scoped events. Errors and gaps stay visible. It adds host overhead and cannot account for the whole seed, live child @@ -129,9 +148,17 @@ Artifacts are retained under `canopy-native-filesystem-ceXFad8I`: | `test_start_native_load.py` | `9a1009d2f57987931beeeead14cb8e1482e3ffd3bbf6f9bb7520a5cd77ae0e0b` | | `load-guard-tests-final.log` | `6b307282780ae288f168cfceef2ff771a215534923036eb6875258d24e00622d` | | `load-check-final.json` | `3076fafdfc8881746f728f8db498b08731676e33cba1e3181df7a7092e2f4a54` | - -The corpus, controller report and observer outputs are still live and do not -have final completion digests. Setup timings are not scheduled throughput. +| `corpus-native-10000.json` | `32e7db9cf06a7b4bb57c906f0f2e70c761ed3ae207c95c7057302b55d50a9a32` | +| `trial.json` | `b618be95af59980fac8f3576f208d17ea83c43ab266c388c82288f0037a5ad12` | +| `seed-native.log` | `c9585a548c6e9c5f2a7b4fedfb8e1c54e59e7b2707e6fc194e9cec1b8612747d` | +| `native-seed-audit.json` | `355912c5fb4f21fa15eec96bdfd5301a967a4676a771a4d05307c7aa1fa49db8` | +| `audit_native_seed.py` | `78e24e4993abb9cfaff1c6121e9ac2b2773d9e30601bf88719d8e0d009a965c1` | +| `native-seed-observer/observation.json` | `7569827510978ba20827e57d1629b72dfc370afa8fb452773844e40c931137b8` | +| `native-initial-owner-loss.json` | `3ed0fdb7df452bf586c1c2908f9b66546e73a8177fbbf088531b07ea73baefe7` | +| `fleet-native-recovery/ready.json` | `2823b1e4b9ab4143fb7c9813e2717cf89d099212f50a5e8f8ad305a1d4043154` | + +The recovery and load controllers are still live and do not have final +completion digests. Setup timings are not scheduled throughput. This is not isolated Linux reference capacity, a proven Cellule bottleneck or a passing latest-pin executable comparison. Both hosted workflows for source pin `0573f489` passed From 966da48c7e4ffecb748faee146171d2dd6531c37 Mon Sep 17 00:00:00 2001 From: forhappy Date: Wed, 30 Sep 2026 21:37:54 -0700 Subject: [PATCH 16/60] perf: audit campaign ledgers and record native recovery --- docs/performance-plan.md | 40 ++- .../2026-10-01-native-filesystem.md | 81 ++++- scripts/audit_three_node_campaign.py | 290 ++++++++++++++++++ scripts/test_audit_three_node_campaign.py | 259 ++++++++++++++++ 4 files changed, 658 insertions(+), 12 deletions(-) create mode 100644 scripts/audit_three_node_campaign.py create mode 100644 scripts/test_audit_three_node_campaign.py diff --git a/docs/performance-plan.md b/docs/performance-plan.md index 58cec27..0bed468 100644 --- a/docs/performance-plan.md +++ b/docs/performance-plan.md @@ -14,8 +14,9 @@ seed and successful recovery of every recorded ACK. | Candidate Git behavior | Bound `e07670e` production artifact passed all 17 critical stock-Git steps through three nodes/proxy against RustFS | | Mac-backed candidate seed | Failed: 3,561/10,000 identities and 39 Git/LFS fixtures; incomplete manifest retained | | Every candidate seed ACK | All 3,561 identities and 39 populated fixtures passed fresh-owner recovery; independent ledger/binding and 78-clone audit passed | -| Native-volume comparison | Same runtime/provider limits; all 10,000 identities and 100 Git/LFS fixtures seeded in 3,179.240 s; full owner-loss recovery running | -| Full recovery and scheduled load | Still unqualified; no incomplete seed or narrow functional test closes these gates | +| Native-volume comparison | Same runtime/provider limits; all 10,000 identities and 100 Git/LFS fixtures seeded in 3,179.240 s and passed initial three-owner-loss recovery; independent 200-clone audit passed | +| Native scheduled load | First attempt failed during mandatory preflight before any window; process-lifetime reproduction retained. Replacement fleet is independently hosted; full preflight is running | +| Performance qualification | Full scheduled load, every post-load ACK recovery, higher admission profiles and matched comparisons remain open | The [native-filesystem trial](performance/2026-10-01-native-filesystem.md) records the storage hypothesis, graceful diagnostic drain and new bindings. @@ -782,6 +783,41 @@ differences, corrupt LFS, lost acknowledgements, tampered samples and inconsiste accounting. These driver checks do not themselves qualify production recovery or the 10,000-repository target. +### Audit closed three-node campaign ledgers + +After the campaign process has exited, independently reconcile its reports: + +```sh +python3 -B scripts/audit_three_node_campaign.py \ + --directory /path/to/closed-campaign \ + --manifest /path/to/canopy-corpus.json \ + --plan docs/performance/three-node-baseline.json \ + --fleet-dir /path/to/original-fleet \ + --output /path/to/new-ledger-audit.json \ + --require-complete +``` + +The auditor sends no Git or provider traffic. It checks input/artifact digests, +full-corpus preflight counts, deterministic repository selection, declared +rates/concurrency/deadlines, exact arrival sequences and outcomes, nearest-rank +percentiles and both in-window and drained throughput. Completed errors remain +in the main latency population; success-only latency is reported separately. +Busy arrivals have no fabricated latency. Interrupted reports without resource +bindings remain unqualified; sample ledgers without closed reports stop the +audit and must be preserved for separate reconciliation. + +Resource summaries cover observed node self CPU/RSS and proxy protocol bytes, +including setup/drain, not all live Git children or payload-only throughput. +`--require-complete` writes the integrity result before rejecting an incomplete +declared schedule. Integrity can pass while arrivals failed: read completion, +errors and resource coverage separately. Neither this audit nor a complete +schedule establishes owner-loss recovery. Run the full seeded-corpus, +critical-fixture and every-ACK checks after the separately recorded owner loss. + +The retained failed baseline passed this audit for all 21 closed ledgers, +20 resource-bound windows and 136 acknowledged creations. Its 108-window +schedule is still incomplete; none of the failed recovery evidence changes. + ## Measurement history Each section below records one implementation or measurement slice. A passing local or mostly empty workload is evidence only for that revision and fixture. The current status table at the top of this page summarizes what the results establish. diff --git a/docs/performance/2026-10-01-native-filesystem.md b/docs/performance/2026-10-01-native-filesystem.md index b719124..cb86e0b 100644 --- a/docs/performance/2026-10-01-native-filesystem.md +++ b/docs/performance/2026-10-01-native-filesystem.md @@ -1,6 +1,6 @@ # Compare RustFS filesystem backing -**Status: full native-volume seed complete; owner-loss recovery running; performance unqualified.** +**Status: full native seed and initial owner-loss recovery passed; replacement load preflight running; performance unqualified.** This shared-Mac/Colima diagnostic keeps the 10,000-repository target. It tests filesystem backing without replacing the bound runtime or weakening correctness. @@ -56,13 +56,13 @@ flowchart LR ack[Every failed-seed ACK verified] --> drain[Graceful diagnostic drain] drain --> seed[Native full 10000 seed
complete] seed --> loss[Three-owner loss and wait
recorded] - loss --> recovery[All identities and Git/LFS
verification running] + loss --> recovery[All identities and Git/LFS
verification passed] recovery --> load[Declared 108-window load plan] load --> final[Second owner loss
verify every load ACK and corpus] ``` The seed and recorded fault boundary are also complete. Critical Git setup and -critical-fixture recovery passed; full native owner-loss recovery, the +critical-fixture and full native initial owner-loss recovery passed; the [declared load matrix](three-node-baseline.json), higher admission profiles, additional critical-operation load/fault coverage and matched performance comparisons remain open. No stage can substitute a @@ -70,7 +70,7 @@ smaller corpus or retries for its required evidence. ## Gate the transition and complete load plan -Two separately bound controllers are running, not completed. The recovery +The initial recovery controller completed successfully. It controller waited for the exact seed and its setup controller to exit, then reread their final receipt. A failed/incomplete corpus, changed process identity, provider restart/pause/quota/mount change or source drift stops it @@ -89,10 +89,13 @@ process absence, the expiry wait lasted **32.037 seconds**. A new fleet started at 03:45:31 UTC with three new IDs/PIDs and local directories, the same durable deployment, admission and immutable binary. Both critical repositories passed four exact v0/v2 ref inventories, payload/notes and strict full fsck. Full -10,000-identity/100-fixture recovery was live at the 1,500-identity checkpoint; -no load window had started at that checkpoint. +10,000-identity/100-fixture recovery completed without retries. Its independent +terminal audit rechecked all 200 local v0/v2 clones: exact HEAD/base commits, +README/incremental bytes and strict full fsck. LFS network bytes are bound to +the inspected verifier, not downloaded again by that local audit. -The load controller requires that closed recovery evidence, all source and +The replacement load controller requires that closed recovery evidence, its +independent audit, all source and receipt digests, three recorded SIGKILLs and fresh owners. Its plan is the unchanged 108 windows: 114,960 offered arrivals over 8,640 offered seconds. Failed arrivals, interrupted windows and diagnostic gaps remain visible; @@ -105,6 +108,48 @@ five no-launch failure cases. All signals, provider calls, Git and load traffic were mocked in control tests. Live read-only binding checks also passed before arming. These are controller guards, not actual owner-loss or performance proof. +### Preserve the failed preflight and fix fixture lifetime + +The first load attempt failed before any scheduled window. Its mandatory +identity preflight received `RemoteDisconnected` for +`density-b7c039e34b54-00528` / `1096f600-8ebe-4648-ac1b-d9395ead175f`. +The launcher and all three nodes disappeared with empty node logs and no +launcher outcome. This is not evidence of the earlier invalid-lease-bounds +fencing branch. The original failed campaign, controller and observer remain +preserved; the successful initial recovery does not reclassify them. + +Three harmless real-tool process-lifetime trials reproduced a relevant fixture +hazard: an attached child disappeared after its tool command ended, while a +`start_new_session=True` child survived to its natural 45-second exit. An +independent offline audit checked all three process/session layouts and logs. +This supports session cleanup as the fixture explanation; it does not prove +the particular signal that removed the real fleet or globally exclude host +interventions. No production or Cellule behavior was changed. + +Actual absence of the old fleet and both controllers was recorded without +sending signals, followed by a fresh **32.040-second** conservative wait. +The replacement fleet is now its own persistent foreground server session, +not a child of a finite recovery controller. It uses new node IDs, signing keys +and local directories, but the same executable, deployment, provider, corpus, +admission and deadlines. Its full mandatory preflight was live at the +3,500-identity checkpoint; no timed window had started at that checkpoint. +The unchanged 108-window attempt has a new output directory. It neither resumes +the failed attempt nor removes a failing window. + +A read-only binding check and eight mocked rejection cases passed for the +replacement controller. These reject incomplete independent recovery, a short +expiry wait, a still-present old PID, provider changes, the wrong launcher, +binary/deployment changes and reused node IDs. Re-running the full original +scenario remains necessary before declaring fixture-lifetime repair verified +end to end. Every post-load creation, push and LFS ACK still requires recovery. + +The separate committed ledger auditor passed against the retained failed +baseline: 21 closed arrival ledgers, 20 bound resource windows and 136 creation +ACKs, with the incomplete 108-window schedule explicitly retained. Nine new +offline auditor tests brought the harness to **75 passing tests** on Python +3.12 and 3.14. Local tests ran during recovery/preflight, not measured load +windows; they are shared-host work, not service performance proof. + ## Separate diagnostics from performance The closed ACK-recovery observer retained 173 samples with zero probe errors. @@ -156,9 +201,25 @@ Artifacts are retained under `canopy-native-filesystem-ceXFad8I`: | `native-seed-observer/observation.json` | `7569827510978ba20827e57d1629b72dfc370afa8fb452773844e40c931137b8` | | `native-initial-owner-loss.json` | `3ed0fdb7df452bf586c1c2908f9b66546e73a8177fbbf088531b07ea73baefe7` | | `fleet-native-recovery/ready.json` | `2823b1e4b9ab4143fb7c9813e2717cf89d099212f50a5e8f8ad305a1d4043154` | - -The recovery and load controllers are still live and do not have final -completion digests. Setup timings are not scheduled throughput. +| `native-seed-to-recovery.json` | `0e243c3df72c789321ebc6dbf8b25c5c5dd22cea0a304ebceb6da32fea5cfe10` | +| `full-native-recovered.json` | `49d8acf40cc8e6cc253f7ac4db0cc8ddff143c913000b4e42ee868a40d2915de` | +| `critical-native-recovered.json` | `1031bf601b1e63970198a385fe4f7bacae0124b6d9861a8254729f371967ef31` | +| `initial-native-recovery-audit.json` | `0115a21f2070999189c70ecdb860af9b6f807de9b1f3fb645e53e2a2419edd20` | +| `native-load-108/campaign.json` (failed preflight) | `f27ea086b94048d54a171ee6c93507a814b427f6140887dc2a27743d9b732e9b` | +| `native-load-controller.json` (failed attempt) | `d646486befbf202fae8a5fb6696a1706b5fda82e39b79332936229d7a0548e4c` | +| `lifetime-probe-RIRZo2/audit.json` | `04cec20ab1c97805ddb40355cf2399648b3ec9c2204b237e7cf69921d57c7b04` | +| `fixture-failure-owner-exit.json` | `8deb78805eac55e248ce461955cf9078e99f851ba9ec6ed00fb952d6f2c9caf2` | +| `serve_native_load_v2.py` | `a2c41160d59b7e7cae45164f0a8030766b4eaadc293d1cc71d5ae36846469a96` | +| `start_native_load_v2.py` | `4585e0e0564edfee476e45a4d287ff061181fe5edd40daaadaae67475245ae1f` | +| `fleet-native-load-v2/ready.json` | `2741766230b975811ef46b104b06a2a3904ceea4902537ba18191e73d7c7aa9b` | +| `load-v2-check.json` | `d3a9d6522b1a7fde45b551c8bafafe485711a81fd194559aa35bcaa1aaeb4549` | +| `harness-75-python314.log` | `07cece1c7015f5df010cca8c9b2ada28cd78fdfb33a7eb08c5d7dd63daa0cecb` | +| `harness-75-python312.log` | `77ea19d83080755caf641792dac1c1dd093100d7131e814b6b8b807cfd0d9389` | + +The baseline's new independent ledger audit is +`5416540c9b78b23e5c89ff24e771ab58012af847df4ec961adfc597bff010639` +in `canopy-three-proxy-q3FO2z`. The replacement load controller is live and +has no final completion digest. Setup timings are not scheduled throughput. This is not isolated Linux reference capacity, a proven Cellule bottleneck or a passing latest-pin executable comparison. Both hosted workflows for source pin `0573f489` passed diff --git a/scripts/audit_three_node_campaign.py b/scripts/audit_three_node_campaign.py new file mode 100644 index 0000000..a6e9d30 --- /dev/null +++ b/scripts/audit_three_node_campaign.py @@ -0,0 +1,290 @@ +#!/usr/bin/env python3 +"""Independently audit closed campaign ledgers; never send Git/provider traffic. + +Failed arrivals and interrupted/unbound windows remain visible. Ledger integrity +does not establish owner recovery, matched improvement or isolated capacity. +""" +import argparse +from collections import Counter +import hashlib +import json +import math +from pathlib import Path +import random +import re +import uuid + +import benchmark_repositories as benchmark +import benchmark_three_node_campaign as campaign + + +def require(condition, message): + if not condition: + raise ValueError(message) + + +def finite(value): + return isinstance(value, (int, float)) and not isinstance(value, bool) and math.isfinite(value) and value >= 0 + + +def quantiles(values): + ordered = sorted(values) + return {name: None if not ordered else round(ordered[math.ceil(q * len(ordered)) - 1], 3) + for name, q in (("p50", .5), ("p95", .95), ("p99", .99), ("max", 1))} + + +def selected_repositories(entries, window, seed): + operation = window["operation"] + if operation == "create": + return {sequence: None for sequence in range(window["rate"] * window["duration"])} + if operation == "lfs_download": + eligible = [entry for entry in entries if entry.get("lfs_oid") is not None] + elif operation in ("incremental_fetch", "incremental_pull"): + eligible = [entry for entry in entries if entry.get("base_commit") is not None] + elif operation in campaign.GIT_READS | {"push_commit"}: + eligible = [entry for entry in entries if entry["commit"] is not None] + else: + eligible = entries + generator = random.Random(seed) + active = generator.sample(eligible, window["active_repositories"]) + hot = active[:max(1, len(active) // 10)] + result = {} + for sequence in range(window["rate"] * window["duration"]): + population = hot if window["distribution"] == "skewed" and generator.random() < .9 else active + result[sequence] = generator.choice(population)["repository_id"] + return result + + +def audit_window(path, window, manifest_digest, driver_digest, repository_ids, expected_selection=None): + raw = path.read_bytes() + report = json.loads(raw) + require(report["version"] == 1 and report["manifest_sha256"] == manifest_digest + and report["driver_sha256"] == driver_digest, "window provenance differs") + for key, value in (("operation", window["operation"]), ("distribution", window["distribution"]), + ("active_repositories", window["active_repositories"]), ("offered_rps", window["rate"]), + ("schedule_seconds", window["duration"]), ("concurrency", window["concurrency"])): + require(report[key] == value, f"window declaration differs: {key}") + total = window["rate"] * window["duration"] + require(isinstance(report["scheduled"], int) and not isinstance(report["scheduled"], bool) + and report["scheduled"] == total and report["request_timeout_seconds"] == 30 + and report["seed"] == 20260926, "arrival clock/deadline/seed differs") + if window["operation"] in campaign.GIT_READS | {"push_branch", "push_commit"}: + require(report["git_timeout_seconds"] == 120, "Git deadline differs") + if window["operation"] == "push_commit": + require(report["git_payload_size_bytes"] == window["git_payload_bytes"], "Git payload differs") + if window["operation"] == "lfs_upload": + require(report["lfs_size_bytes"] == window["lfs_bytes"], "LFS payload differs") + if window["operation"] == "create": + require(isinstance(report["create_run_id"], str) and re.fullmatch(r"[0-9a-f]{32}", report["create_run_id"]), "invalid creation run ID") + samples_path = path.with_suffix(".samples.jsonl") + rows, seen, requests, digest = [], set(), set(), hashlib.sha256() + with samples_path.open("rb") as samples: + for line in iter(lambda: samples.readline(16 * 1024 + 1), b""): + require(len(line) <= 16 * 1024, "oversized arrival sample") + digest.update(line) + row = json.loads(line) + sequence = row["sequence"] + require(isinstance(sequence, int) and not isinstance(sequence, bool) + and 0 <= sequence < total and sequence not in seen, "duplicate/out-of-range sequence") + seen.add(sequence) + require(type(row["ingress_index"]) is int and row["ingress_index"] == 0 + and isinstance(row["result"], str), "unexpected ingress/outcome") + if expected_selection is not None: + require(row["repository_id"] == expected_selection[sequence], "deterministic active-set/distribution differs") + if window["operation"] != "create": + require(row["repository_id"] in repository_ids, "arrival repository outside corpus") + else: + require(row["repository_id"] is None, "creation selected an existing repository") + require(row["created_name"] == f"create-{report['create_run_id']}-{sequence:07d}", "creation name differs") + if row["result"] == "ok": + require(benchmark.canonical_repository_uuid(row.get("created_repository_id")), "creation ACK has invalid UUID") + if row["result"] == "driver_busy": + require(row["elapsed_ms"] is None, "busy arrival has fabricated latency") + else: + request_id = row["request_id"] + require(isinstance(request_id, str) and str(uuid.UUID(request_id)) == request_id + and request_id not in requests, "invalid/duplicate request ID") + requests.add(request_id) + require(all(finite(row[key]) for key in ("elapsed_ms", "service_ms", "dispatch_delay_ms", "completion_offset_seconds")), + "invalid arrival timing") + require(math.isclose(row["elapsed_ms"], row["service_ms"] + row["dispatch_delay_ms"], abs_tol=.001), + "elapsed/service/dispatch identity differs") + expected_elapsed = (row["completion_offset_seconds"] - sequence / window["rate"]) * 1000 + require(math.isclose(row["elapsed_ms"], expected_elapsed, abs_tol=.01), "arrival clock identity differs") + for key in ("git_push_command_ms", "git_client_preparation_ms"): + require(row.get(key) is None or finite(row[key]), "invalid push phase timing") + rows.append(row) + require(digest.hexdigest() == report["samples_sha256"] == benchmark.file_sha256(samples_path), "arrival digest changed") + require(len(rows) == len(seen) == total, "missing arrival sequences") + outcomes = Counter(row["result"] for row in rows) + require(all(type(value) is int and value >= 0 for value in report["outcomes"].values()) + and type(report["failed_arrivals"]) is int, "outcome counters are not integers") + require(dict(outcomes) == report["outcomes"] and total - outcomes["ok"] == report["failed_arrivals"], "outcome accounting differs") + require(report["error_fraction"] == (total - outcomes["ok"]) / total, "error fraction differs") + if window["operation"] == "create": + identifiers = [row["created_repository_id"] for row in rows if row["result"] == "ok"] + require(len(set(identifiers)) == len(identifiers) == report["acknowledged_created_repositories"], "creation ACK count/uniqueness differs") + dispatched = [row for row in rows if row["result"] != "driver_busy"] + for field, key in (("elapsed_ms", "scheduled_latency_ms"), ("service_ms", "service_ms"), + ("dispatch_delay_ms", "dispatch_delay_ms")): + require(quantiles([row[field] for row in dispatched]) == report[key], f"percentiles differ: {key}") + require(report["ingresses"] == [{"index": 0, "outcomes": dict(outcomes), + "scheduled_latency_ms": report["scheduled_latency_ms"], "service_ms": report["service_ms"]}], + "per-ingress accounting differs") + if window["operation"] == "push_commit": + for key in ("git_push_command_ms", "git_client_preparation_ms"): + require(quantiles([row[key] for row in rows if row.get(key) is not None]) == report[key], "push percentiles differ") + require(report["acknowledged_new_git_payload_bytes"] == outcomes["ok"] * window["git_payload_bytes"], "acknowledged payload bytes differ") + successes = sum(row["result"] == "ok" and row["completion_offset_seconds"] <= window["duration"] for row in rows) + rate = round(successes / window["duration"], 3) + require(successes == report["successful_completions_in_schedule_window"] + and rate == report["successful_rps_in_schedule_window"], "in-window throughput differs") + require(finite(report["elapsed_including_drain_seconds"]) and report["elapsed_including_drain_seconds"] >= window["duration"], + "drain duration invalid") + duration = report["elapsed_including_drain_seconds"] + require(round(outcomes["ok"] / (duration + .0005), 3) <= report["successful_rps_including_drain"] + <= round(outcomes["ok"] / (duration - .0005), 3), "drained throughput differs from rounded duration bounds") + require(benchmark.file_sha256(path) == hashlib.sha256(raw).hexdigest(), "report changed during audit") + return {"path": path.name, "sha256": hashlib.sha256(raw).hexdigest(), "samples_sha256": digest.hexdigest(), + "operation": window["operation"], "distribution": window["distribution"], + "active_repositories": window["active_repositories"], "offered_rps": window["rate"], + "concurrency": window["concurrency"], "scheduled": total, "outcomes": dict(outcomes), + "successful_rps_in_schedule_window": rate, "scheduled_latency_ms": report["scheduled_latency_ms"], + "successful_only_scheduled_latency_ms": quantiles([row["elapsed_ms"] for row in rows if row["result"] == "ok"]), + "service_ms": report["service_ms"], "dispatch_delay_ms": report["dispatch_delay_ms"]} + + +def audit_resources(directory, item, ready): + for key in ("resources", "resource_boundary"): + path = directory / item[key + "_path"] + require(path.parent == directory and benchmark.file_sha256(path) == item[key + "_sha256"], "resource binding differs") + boundary = json.loads((directory / item["resource_boundary_path"]).read_text()) + rows = [json.loads(line) for line in (directory / item["resources_path"]).read_text().splitlines()] + require(rows and all(b["monotonic"] > a["monotonic"] for a, b in zip(rows, rows[1:])), "resource samples missing/unordered") + required_pids = {ready["launcher_pid"], *(node["pid"] for node in ready["nodes"])} + pids = {p["pid"] for p in boundary["before"]["processes"]} + require(required_pids < pids and len(pids) == 5, "resource process inventory differs") + for sample in [boundary["before"], *rows, boundary["after"]]: + require(finite(sample["monotonic"]) and sample["proxy"]["fixture_id"] == ready["fixture_id"], "resource fixture differs") + captured = sample["proxy"]["captured_monotonic"] + require(finite(captured) and 0 <= sample["monotonic"] - captured <= 10, "proxy sample stale") + require(len(sample["processes"]) == 5 and {p["pid"] for p in sample["processes"]} == pids + and all(finite(p[key]) for p in sample["processes"] for key in ("cpu_seconds", "rss_kib", "ps_lifetime_cpu_percent")), + "resource PID/counter inventory differs") + require(boundary["after"]["monotonic"] > boundary["before"]["monotonic"], "resource boundary order differs") + old = {p["pid"]: p for p in boundary["before"]["processes"]} + require(all(p["cpu_seconds"] >= old[p["pid"]]["cpu_seconds"] for p in boundary["after"]["processes"]), "self CPU decreased") + current = {p["pid"]: p for p in boundary["after"]["processes"]} + node_cpu = {str(n["pid"]): current[n["pid"]]["cpu_seconds"] - old[n["pid"]]["cpu_seconds"] for n in ready["nodes"]} + rss_peak = {str(n["pid"]): max(p["rss_kib"] for sample in [boundary["before"], *rows, boundary["after"]] + for p in sample["processes"] if p["pid"] == n["pid"]) for n in ready["nodes"]} + protocol_bytes = {} + for key in ("client_bytes", "backend_bytes"): + first = boundary["before"]["proxy"]["front"][key] + last = boundary["after"]["proxy"]["front"][key] + require(type(first) is int and type(last) is int and 0 <= first <= last, "proxy byte counter decreased") + protocol_bytes[key] = last - first + return {"samples": len(rows), "resource_boundary_sha256": item["resource_boundary_sha256"], + "resources_sha256": item["resources_sha256"], + "boundary_seconds_including_setup_and_drain": boundary["after"]["monotonic"] - boundary["before"]["monotonic"], + "node_self_cpu_seconds": node_cpu, "node_peak_observed_rss_kib": rss_peak, + "front_proxy_protocol_bytes": protocol_bytes, + "scope": "ps self CPU/RSS and proxy protocol counters; includes client setup/drain, not live-child or Git-only CPU"} + + +def audit(directory, manifest_path, plan_path, fleet_dir): + index_path = directory / "campaign.json" + index_digest = benchmark.file_sha256(index_path) + index = json.loads(index_path.read_text()) + manifest = benchmark.corpus(manifest_path) + plan = json.loads(plan_path.read_text()) + schedule = campaign.windows(plan, manifest) + ready_path = fleet_dir / "ready.json" + ready = json.loads(ready_path.read_text()) + expected = {"manifest_sha256": benchmark.file_sha256(manifest_path), "plan_sha256": benchmark.file_sha256(plan_path), + "fleet_ready_sha256": benchmark.file_sha256(ready_path), "driver_sha256": benchmark.file_sha256(Path(benchmark.__file__)), + "campaign_sha256": benchmark.file_sha256(Path(campaign.__file__))} + require(index["version"] == 1 and index["bindings"] == expected and index["fleet"] == ready, "campaign provenance differs") + preflight = index["preflight"] + require(preflight is not None and preflight["verified_repositories"] == len(manifest["repositories"]) + and preflight["git_v0_v2_samples"] == sum(e["commit"] is not None for e in manifest["repositories"]) + and all(preflight[key] == value for key, value in expected.items()), "full corpus preflight differs") + require(ready["ready"] is True and ready["error"] is None and not ready["shutdown"] + and len(ready["nodes"]) == 3 and len({n["node_id"] for n in ready["nodes"]}) == 3 + and [n["index"] for n in ready["nodes"]] == [0, 1, 2] + and len({n["pid"] for n in ready["nodes"]}) == 3 + and ready["max_active_repositories_per_node"] == plan["node_active_limit"] + and ready["public_url"] == ready["proxy_url"], "fleet topology differs") + require(benchmark.file_sha256(Path(ready["binary_path"])) == ready["binary_sha256"], "binary changed") + require(set(ready["fixture_scripts_sha256"]) == {"serve_three_gateways.py", "local_tcp_proxy.py", "smoke_s3_process.py", "smoke_s3_peers.py"} + and all(benchmark.file_sha256(Path(__file__).parent / name) == digest + for name, digest in ready["fixture_scripts_sha256"].items()), "fixture script binding differs") + bound = {item["path"]: item for item in index["reports"]} + require(len(bound) == len(index["reports"]), "duplicate bound report") + windows, paths = [], [] + for ordinal, window in enumerate(schedule): + name = f"{ordinal:04d}-{window['id']}-r{window['repetition']}.json" + path = directory / name + if not path.exists(): + require(name not in bound, "bound report missing") + continue + checked = audit_window(path, window, expected["manifest_sha256"], expected["driver_sha256"], + {e["repository_id"] for e in manifest["repositories"]}, + selected_repositories(manifest["repositories"], window, 20260926)) + item = bound.get(name) + checked["resource_binding_complete"] = item is not None + if item is not None: + require(item["sha256"] == checked["sha256"] and item["operation"] == checked["operation"] + and item["outcomes"] == checked["outcomes"] + and item["failed_arrivals"] == checked["scheduled"] - checked["outcomes"].get("ok", 0), "index/window accounting differs") + checked["resources"] = audit_resources(directory, item, ready) + else: + require(index["completed"] is False and index["current_window"] == path.stem, "unexpected unbound report") + windows.append(checked) + paths.append(path) + require(set(bound) <= {p.name for p in paths}, "index references undeclared report") + declared_names = {p.stem + ".samples.jsonl" for p in paths} + orphan_ledgers = [p.name for p in directory.glob("*.samples.jsonl") if p.name not in declared_names] + require(not orphan_ledgers, "arrival ledgers lack closed reports; retain and audit separately") + creations = [p for p, row in zip(paths, windows) if row["operation"] == "create"] + writes = [p for p, row in zip(paths, windows) if row["operation"] in ("push_branch", "push_commit", "lfs_upload")] + creation_runs = benchmark.creation_runs(creations, manifest_path) if creations and any( + row["outcomes"].get("ok", 0) for row in windows if row["operation"] == "create") else [] + acknowledged_creations = sum(len(rows) for _, rows, _ in creation_runs) + write_counts = Counter() + if writes and any(row["outcomes"].get("ok", 0) for row in windows if row["operation"] in ("push_branch", "push_commit", "lfs_upload")): + runs = benchmark.write_runs(writes, manifest_path, {e["repository_id"]: e for e in manifest["repositories"]}) + for report, rows, _ in runs: + write_counts[report["operation"]] += len(rows) + complete = index["completed"] is True and len(windows) == len(bound) == len(schedule) and index["current_window"] is None + require(index["completed"] is not True or complete, "completed campaign omitted declared windows") + if complete: + total_outcomes = Counter() + for row in windows: + total_outcomes.update(row["outcomes"]) + require(dict(total_outcomes) == index["outcomes"] and index["all_arrivals_succeeded"] is + all(row["outcomes"].get("ok", 0) == row["scheduled"] for row in windows), "campaign outcome totals differ") + require(benchmark.file_sha256(index_path) == index_digest, "campaign changed during audit") + return {"version": 1, "integrity_verified": True, "schedule_completed": complete, + "all_observed_arrivals_succeeded": all(row["outcomes"].get("ok", 0) == row["scheduled"] for row in windows), + "declared_windows": len(schedule), "observed_windows": len(windows), "resource_bound_windows": len(bound), + "acknowledged_creations": acknowledged_creations, "acknowledged_writes": dict(write_counts), + "campaign_sha256": index_digest, "bindings": expected, "auditor_sha256": benchmark.file_sha256(Path(__file__)), + "windows": windows, "scope": "Closed ledgers/resources only; failures retained. No owner-loss, delivered Git/LFS body, matched improvement or isolated-capacity proof."} + + +def main(): + parser = argparse.ArgumentParser(description=__doc__) + for name in ("directory", "manifest", "plan", "fleet-dir", "output"): + parser.add_argument("--" + name, type=Path, required=True) + parser.add_argument("--require-complete", action="store_true") + args = parser.parse_args() + require(not args.output.exists(), "audit requires new output") + result = audit(args.directory, args.manifest, args.plan, args.fleet_dir) + benchmark.save(args.output, result) + print(json.dumps({key: value for key, value in result.items() if key != "windows"}, indent=2)) + require(not args.require_complete or result["schedule_completed"], "full declared schedule did not complete") + + +if __name__ == "__main__": + main() diff --git a/scripts/test_audit_three_node_campaign.py b/scripts/test_audit_three_node_campaign.py new file mode 100644 index 0000000..ac51b13 --- /dev/null +++ b/scripts/test_audit_three_node_campaign.py @@ -0,0 +1,259 @@ +"""Offline audit fixtures only; no server, process signals or remote traffic.""" +import copy +import json +from pathlib import Path +import tempfile +import unittest +import uuid + +import audit_three_node_campaign as auditor +import benchmark_repositories as benchmark +import benchmark_three_node_campaign as campaign + + +class AuditTests(unittest.TestCase): + def fixture(self, root): + identifier = str(uuid.UUID(int=1, version=4)) + window = {"id": "metadata", "operation": "metadata", "distribution": "uniform", "active_repositories": 1, + "rate": 4, "duration": 1, "concurrency": 1, "repetitions": 1} + rows = [] + for sequence, outcome, elapsed, service, dispatch, completion in ( + (0, "ok", 100, 90, 10, .1), (1, "driver_busy", None, None, None, None), + (2, "transport_error", 20, 15, 5, .52), (3, "ok", 270, 260, 10, 1.02)): + row = {"sequence": sequence, "repository_id": identifier, "ingress_index": 0, "result": outcome, "elapsed_ms": elapsed} + if elapsed is not None: + row.update(request_id=str(uuid.UUID(int=sequence+10, version=4)), service_ms=service, + dispatch_delay_ms=dispatch, completion_offset_seconds=completion) + rows.append(row) + report = {"version": 1, "manifest_sha256": "manifest", "driver_sha256": "driver", "seed": 20260926, + "operation": "metadata", "distribution": "uniform", "active_repositories": 1, "offered_rps": 4, + "schedule_seconds": 1, "concurrency": 1, "scheduled": 4, "request_timeout_seconds": 30, + "outcomes": {"ok": 2, "driver_busy": 1, "transport_error": 1}, "failed_arrivals": 2, "error_fraction": .5, + "scheduled_latency_ms": {"p50": 100, "p95": 270, "p99": 270, "max": 270}, + "service_ms": {"p50": 90, "p95": 260, "p99": 260, "max": 260}, + "dispatch_delay_ms": {"p50": 10, "p95": 10, "p99": 10, "max": 10}, + "successful_completions_in_schedule_window": 1, "successful_rps_in_schedule_window": 1., + "elapsed_including_drain_seconds": 1.05, "successful_rps_including_drain": 1.905} + report["ingresses"] = [{"index": 0, "outcomes": report["outcomes"], + "scheduled_latency_ms": report["scheduled_latency_ms"], "service_ms": report["service_ms"]}] + return root / "0000-metadata-r1.json", window, rows, report, identifier + + def write(self, path, rows, report): + with path.with_suffix(".samples.jsonl").open("w") as out: + for row in rows: + out.write(json.dumps(row) + "\n") + report["samples_sha256"] = benchmark.file_sha256(path.with_suffix(".samples.jsonl")) + benchmark.save(path, report) + + def test_latency_populations_busy_and_late_success(self): + with tempfile.TemporaryDirectory() as directory: + path, window, rows, report, identifier = self.fixture(Path(directory)) + self.write(path, rows, report) + result = auditor.audit_window(path, window, "manifest", "driver", {identifier}, dict.fromkeys(range(4), identifier)) + self.assertEqual(result["outcomes"]["driver_busy"], 1) + self.assertEqual(result["successful_rps_in_schedule_window"], 1) + self.assertEqual(result["scheduled_latency_ms"]["p50"], 100) + self.assertEqual(result["successful_only_scheduled_latency_ms"]["p50"], 100) + self.assertEqual(result["successful_only_scheduled_latency_ms"]["p95"], 270) + + def test_corrupt_accounting_timing_and_selection_rejected(self): + for change in ("missing", "duplicate", "unknown_repo", "wrong_selection", "busy_latency", "negative", "timing_identity", + "request_duplicate", "rate", "concurrency", "quantile", "throughput", "drained_rate", "outcomes", "seed", "ingress"): + with self.subTest(change=change), tempfile.TemporaryDirectory() as directory: + path, window, rows, report, identifier = self.fixture(Path(directory)) + selection = dict.fromkeys(range(4), identifier) + if change == "missing": + rows.pop() + elif change == "duplicate": + rows[3]["sequence"] = 0 + elif change == "unknown_repo": + rows[3]["repository_id"] = "outside" + elif change == "wrong_selection": + selection[3] = "other" + elif change == "busy_latency": + rows[1]["elapsed_ms"] = 0 + elif change == "negative": + rows[0]["service_ms"] = -1 + elif change == "timing_identity": + rows[0]["completion_offset_seconds"] = .2 + elif change == "request_duplicate": + rows[3]["request_id"] = rows[0]["request_id"] + elif change == "rate": + report["offered_rps"] = 5 + elif change == "concurrency": + report["concurrency"] = 2 + elif change == "quantile": + report["scheduled_latency_ms"]["p50"] = 20 + elif change == "throughput": + report["successful_completions_in_schedule_window"] = 2 + elif change == "drained_rate": + report["successful_rps_including_drain"] = 2. + elif change == "outcomes": + report["outcomes"]["ok"] = 3 + elif change == "seed": + report["seed"] = 1 + else: + report["ingresses"][0]["index"] = 1 + self.write(path, rows, report) + with self.assertRaises(ValueError): + auditor.audit_window(path, window, "manifest", "driver", {identifier}, selection) + + def test_changed_samples_digest_is_rejected(self): + with tempfile.TemporaryDirectory() as directory: + path, window, rows, report, identifier = self.fixture(Path(directory)) + self.write(path, rows, report) + with path.with_suffix(".samples.jsonl").open("a") as out: + out.write("{}\n") + with self.assertRaises((ValueError, KeyError)): + auditor.audit_window(path, window, "manifest", "driver", {identifier}) + + def test_deterministic_selection_preserves_uniform_skew_and_git_eligibility(self): + entries = [{"repository_id": str(i), "commit": "tip" if i < 10 else None, "base_commit": "base" if i < 10 else None} + for i in range(20)] + window = {"operation": "clone", "active_repositories": 10, "distribution": "uniform", "rate": 1000, "duration": 1} + uniform = auditor.selected_repositories(entries, window, 20260926) + skewed = auditor.selected_repositories(entries, {**window, "distribution": "skewed"}, 20260926) + self.assertEqual(uniform, auditor.selected_repositories(entries, window, 20260926)) + self.assertEqual(set(uniform.values()), {str(i) for i in range(10)}) + self.assertGreater(max(list(skewed.values()).count(i) for i in set(skewed.values())), 850) + self.assertEqual(set(auditor.selected_repositories(entries, {**window, "operation": "create"}, 20260926).values()), {None}) + + def test_creation_ack_counts_and_duplicate_uuid_rejected(self): + for corrupt in (False, True): + with self.subTest(corrupt=corrupt), tempfile.TemporaryDirectory() as directory: + path, window, rows, report, identifier = self.fixture(Path(directory)) + window.update(operation="create", active_repositories=None) + report.update(operation="create", active_repositories=None, create_run_id="1"*32, acknowledged_created_repositories=2) + for row in rows: + row.update(repository_id=None, created_name=f"create-{'1'*32}-{row['sequence']:07d}", + created_repository_id=str(uuid.UUID(int=row["sequence"]+100, version=4)) if row["result"] == "ok" else None) + if corrupt: + rows[3]["created_repository_id"] = rows[0]["created_repository_id"] + self.write(path, rows, report) + if corrupt: + with self.assertRaises(ValueError): + auditor.audit_window(path, window, "manifest", "driver", {identifier}) + else: + self.assertEqual(auditor.audit_window(path, window, "manifest", "driver", {identifier})["outcomes"]["ok"], 2) + + def test_fresh_push_phase_population_and_payload_accounting(self): + for corrupt in (None, "payload", "phase_quantile", "git_deadline"): + with self.subTest(corrupt=corrupt), tempfile.TemporaryDirectory() as directory: + path, window, rows, report, identifier = self.fixture(Path(directory)) + window.update(operation="push_commit", git_payload_bytes=256) + report.update(operation="push_commit", git_timeout_seconds=120, git_payload_size_bytes=256, + acknowledged_new_git_payload_bytes=512) + for row in rows: + if row["result"] == "ok": + row.update(git_push_command_ms=10, git_client_preparation_ms=5) + report["git_push_command_ms"] = dict.fromkeys(("p50", "p95", "p99", "max"), 10) + report["git_client_preparation_ms"] = dict.fromkeys(("p50", "p95", "p99", "max"), 5) + if corrupt == "payload": + report["acknowledged_new_git_payload_bytes"] = 768 + elif corrupt == "phase_quantile": + report["git_push_command_ms"]["p95"] = 11 + elif corrupt == "git_deadline": + report["git_timeout_seconds"] = 240 + self.write(path, rows, report) + if corrupt is not None: + with self.assertRaises(ValueError): + auditor.audit_window(path, window, "manifest", "driver", {identifier}) + else: + self.assertEqual(auditor.audit_window(path, window, "manifest", "driver", {identifier})["operation"], "push_commit") + + def campaign_fixture(self, root): + path, window, rows, report, identifier = self.fixture(root) + manifest_path, plan_path, fleet = root / "manifest.json", root / "plan.json", root / "fleet" + fleet.mkdir() + benchmark.save(manifest_path, {"version": 1, "complete": True, "requested_repositories": 1, + "repositories": [{"name": "repo", "owner": "canopy", "repository_id": identifier, "commit": None}]}) + benchmark.save(plan_path, {"version": 1, "corpus_repositories": 1, "node_active_limit": 100, "windows": [window]}) + binary = root / "fake-binary" + binary.write_bytes(b"offline fixture; never executed") + ready = {"ready": True, "error": None, "shutdown": [], "fixture_id": "fixture", "launcher_pid": 1, + "nodes": [{"index": i, "pid": i+2, "node_id": str(uuid.uuid4())} for i in range(3)], + "max_active_repositories_per_node": 100, "public_url": "http://127.0.0.1:1", "proxy_url": "http://127.0.0.1:1", + "binary_path": str(binary), "binary_sha256": benchmark.file_sha256(binary), + "fixture_scripts_sha256": {name: benchmark.file_sha256(Path(__file__).parent / name) for name in + ("serve_three_gateways.py", "local_tcp_proxy.py", "smoke_s3_process.py", "smoke_s3_peers.py")}} + benchmark.save(fleet / "ready.json", ready) + bindings = {"manifest_sha256": benchmark.file_sha256(manifest_path), "plan_sha256": benchmark.file_sha256(plan_path), + "fleet_ready_sha256": benchmark.file_sha256(fleet / "ready.json"), + "driver_sha256": benchmark.file_sha256(Path(benchmark.__file__)), + "campaign_sha256": benchmark.file_sha256(Path(campaign.__file__))} + report.update(manifest_sha256=bindings["manifest_sha256"], driver_sha256=bindings["driver_sha256"]) + self.write(path, rows, report) + def resource(when, cpu): + return {"monotonic": when, "processes": [{"pid": pid, "cpu_seconds": cpu, "rss_kib": 1024, + "ps_lifetime_cpu_percent": 1.} for pid in range(1, 6)], + "proxy": {"fixture_id": "fixture", "captured_monotonic": when-.1, + "front": {"client_bytes": int(cpu*100), "backend_bytes": int(cpu*200)}}} + boundary = path.with_name(path.stem + ".resource-boundary.json") + resources = path.with_name(path.stem + ".resources.jsonl") + benchmark.save(boundary, {"before": resource(10, 1), "after": resource(11.1, 2)}) + with resources.open("w") as out: + out.write(json.dumps(resource(10.25, 1.25)) + "\n") + out.write(json.dumps(resource(10.75, 1.75)) + "\n") + item = {"path": path.name, "sha256": benchmark.file_sha256(path), "operation": "metadata", + "outcomes": report["outcomes"], "failed_arrivals": 2, "resources_path": resources.name, + "resources_sha256": benchmark.file_sha256(resources), "resource_boundary_path": boundary.name, + "resource_boundary_sha256": benchmark.file_sha256(boundary)} + index = {"version": 1, "bindings": bindings, "fleet": ready, "completed": True, "current_window": None, + "preflight": {"verified_repositories": 1, "git_v0_v2_samples": 0, **bindings}, + "reports": [item], "outcomes": report["outcomes"], "all_arrivals_succeeded": False} + benchmark.save(root / "campaign.json", index) + return manifest_path, plan_path, fleet, index + + def test_complete_schedule_can_retain_failed_arrivals_without_claiming_recovery(self): + with tempfile.TemporaryDirectory() as directory: + root = Path(directory) + manifest, plan, fleet, _ = self.campaign_fixture(root) + result = auditor.audit(root, manifest, plan, fleet) + self.assertTrue(result["schedule_completed"]) + self.assertFalse(result["all_observed_arrivals_succeeded"]) + self.assertEqual(result["windows"][0]["resources"]["node_self_cpu_seconds"], {"2": 1, "3": 1, "4": 1}) + self.assertEqual(result["windows"][0]["resources"]["front_proxy_protocol_bytes"], {"client_bytes": 100, "backend_bytes": 200}) + + def test_partial_window_retained_without_resource_qualification(self): + with tempfile.TemporaryDirectory() as directory: + root = Path(directory) + manifest, plan, fleet, index = self.campaign_fixture(root) + index.update(completed=False, current_window="0000-metadata-r1", reports=[]) + benchmark.save(root / "campaign.json", index) + result = auditor.audit(root, manifest, plan, fleet) + self.assertFalse(result["schedule_completed"]) + self.assertFalse(result["windows"][0]["resource_binding_complete"]) + + def test_campaign_binding_preflight_and_resource_corruption_rejected(self): + for change in ("preflight", "omitted", "resource_digest", "resource_pid", "resource_clock", "resource_cpu", "binary", "orphan"): + with self.subTest(change=change), tempfile.TemporaryDirectory() as directory: + root = Path(directory) + manifest, plan, fleet, index = self.campaign_fixture(root) + if change == "preflight": + index["preflight"]["verified_repositories"] = 0 + elif change == "omitted": + index["reports"] = [] + elif change == "resource_digest": + index["reports"][0]["resources_sha256"] = "wrong" + elif change.startswith("resource_"): + path = root / index["reports"][0]["resource_boundary_path"] + boundary = json.loads(path.read_text()) + if change == "resource_pid": + boundary["after"]["processes"][0]["pid"] = 99 + elif change == "resource_clock": + boundary["after"]["proxy"]["captured_monotonic"] = 0 + else: + boundary["after"]["processes"][0]["cpu_seconds"] = .5 + benchmark.save(path, boundary) + index["reports"][0]["resource_boundary_sha256"] = benchmark.file_sha256(path) + elif change == "binary": + (root / "fake-binary").write_bytes(b"changed") + else: + (root / "9999-unbound.samples.jsonl").write_bytes(b"{}\n") + benchmark.save(root / "campaign.json", index) + with self.assertRaises(ValueError): + auditor.audit(root, manifest, plan, fleet) + + +if __name__ == "__main__": + unittest.main() From 772d1dc07b6f008c339879703cd0dfbc894da583 Mon Sep 17 00:00:00 2001 From: forhappy Date: Wed, 30 Sep 2026 21:48:19 -0700 Subject: [PATCH 17/60] docs: bind post-load every-ACK recovery controller --- .../2026-10-01-native-filesystem.md | 45 +++++++++++++++++-- 1 file changed, 42 insertions(+), 3 deletions(-) diff --git a/docs/performance/2026-10-01-native-filesystem.md b/docs/performance/2026-10-01-native-filesystem.md index cb86e0b..31b144d 100644 --- a/docs/performance/2026-10-01-native-filesystem.md +++ b/docs/performance/2026-10-01-native-filesystem.md @@ -71,7 +71,7 @@ smaller corpus or retries for its required evidence. ## Gate the transition and complete load plan The initial recovery controller completed successfully. It -controller waited for the exact seed and its setup controller to exit, then +waited for the exact seed and its setup controller to exit, then reread their final receipt. A failed/incomplete corpus, changed process identity, provider restart/pause/quota/mount change or source drift stops it before any owner-loss action. @@ -150,6 +150,40 @@ offline auditor tests brought the harness to **75 passing tests** on Python 3.12 and 3.14. Local tests ran during recovery/preflight, not measured load windows; they are shared-host work, not service performance proof. +### Arm every-ACK recovery without shortening the load + +A separately bound post-load controller is live in `wait_bound_load`. Its +read-only live binding check and four offline guard tests passed. During load +it only inspects the two bound process identities every 15 seconds; it does not +poll provider state or send additional Git traffic. This still adds some host +overhead and is not an isolated reference measurement. + +```mermaid +flowchart LR + endload[Actual campaign and observer exit] --> audit[Audit all closed ledgers
retain failed arrivals] + audit --> loss[Verified owner loss
actual absence and 32s wait] + loss --> fresh[New IDs and local directories
independent process session] + fresh --> scopes[Full 10000 corpus and critical fixtures
every creation, push and LFS ACK] +``` + +Before any signal, it requires closed controller/campaign/observer evidence, +unchanged source/provider bindings and an independent complete-ledger audit. +Missing or orphan ledgers stop it for reconciliation, not silent omission. +If all old owners remain live, it verifies the entire PID/command/kernel +identity batch before three SIGKILLs. Already-lost fleets get no signals and +must become authoritatively absent before the same conservative wait. A +partially lost fleet that does not close stops takeover. Even a rare mid-batch +exit retains the signals already sent. + +Fresh recovery keeps the same durable deployment, executable, provider and +admission. Its launcher uses a new process session so completion of the finite +verifier does not orphan an attached fixture. Critical fixtures, full seeded +corpus, acknowledged creations and acknowledged Git/LFS writes have separate +verification stages; one failed stage does not skip the other scopes. Counts +must match the audited ACK ledgers. There are no retries or failed-arrival +rollback assertions. This controller is armed, not completed, and a recovery +pass cannot turn a failed performance campaign into a pass. + ## Separate diagnostics from performance The closed ACK-recovery observer retained 173 samples with zero probe errors. @@ -215,11 +249,16 @@ Artifacts are retained under `canopy-native-filesystem-ceXFad8I`: | `load-v2-check.json` | `d3a9d6522b1a7fde45b551c8bafafe485711a81fd194559aa35bcaa1aaeb4549` | | `harness-75-python314.log` | `07cece1c7015f5df010cca8c9b2ada28cd78fdfb33a7eb08c5d7dd63daa0cecb` | | `harness-75-python312.log` | `77ea19d83080755caf641792dac1c1dd093100d7131e814b6b8b807cfd0d9389` | +| `finish_native_load_v2.py` | `596a713aebe893796b9425d18dcfead2987f03728ac2ccb0cb63216bcb76a036` | +| `test_finish_native_load_v2.py` | `2ad420292ccfd4fe66072b216cb816603dbeeb3c20dbb86568b56cec10495762` | +| `post-load-v2-guard-tests.log` | `28157ddb518f5c1d4f887b07be9710a779324b00d8f73347311839e52606bd59` | +| `post-load-v2-check.json` | `57befacbc9f31ac6564b8cab618c45be0851b8f81c3996c1334eadae485cb281` | The baseline's new independent ledger audit is `5416540c9b78b23e5c89ff24e771ab58012af847df4ec961adfc597bff010639` -in `canopy-three-proxy-q3FO2z`. The replacement load controller is live and -has no final completion digest. Setup timings are not scheduled throughput. +in `canopy-three-proxy-q3FO2z`. The replacement load and post-load recovery +controllers are live and have no final completion digests. Setup timings are +not scheduled throughput. This is not isolated Linux reference capacity, a proven Cellule bottleneck or a passing latest-pin executable comparison. Both hosted workflows for source pin `0573f489` passed From 669584a73c87aa0f9336ed94a70b0e258ebb4329 Mon Sep 17 00:00:00 2001 From: forhappy Date: Wed, 30 Sep 2026 22:44:00 -0700 Subject: [PATCH 18/60] docs(perf): retain RustFS OOM and bind full recovery diagnostic --- docs/performance-plan.md | 7 +- .../2026-10-01-native-filesystem.md | 153 +++++++++++++++++- 2 files changed, 150 insertions(+), 10 deletions(-) diff --git a/docs/performance-plan.md b/docs/performance-plan.md index 0bed468..9bba33d 100644 --- a/docs/performance-plan.md +++ b/docs/performance-plan.md @@ -15,12 +15,15 @@ seed and successful recovery of every recorded ACK. | Mac-backed candidate seed | Failed: 3,561/10,000 identities and 39 Git/LFS fixtures; incomplete manifest retained | | Every candidate seed ACK | All 3,561 identities and 39 populated fixtures passed fresh-owner recovery; independent ledger/binding and 78-clone audit passed | | Native-volume comparison | Same runtime/provider limits; all 10,000 identities and 100 Git/LFS fixtures seeded in 3,179.240 s and passed initial three-owner-loss recovery; independent 200-clone audit passed | -| Native scheduled load | First attempt failed during mandatory preflight before any window; process-lifetime reproduction retained. Replacement fleet is independently hosted; full preflight is running | +| Native scheduled load | First attempt failed before any window. Independent-session V2 passed full preflight, then RustFS was cgroup-OOM killed: six closed windows, one interrupted prefix and 101 unstarted windows retained | +| Post-OOM diagnostic | Three old owners killed as a verified batch; 32.033-s absence wait recorded. Same provider/data restarted with memory raised from 2 to 4 GiB; critical-2 recovery passed, full 10K/100 recovery running | | Performance qualification | Full scheduled load, every post-load ACK recovery, higher admission profiles and matched comparisons remain open | The [native-filesystem trial](performance/2026-10-01-native-filesystem.md) records the storage hypothesis, graceful diagnostic drain and new bindings. -No Cellule bottleneck or matched performance improvement is proven. +No Cellule bottleneck or matched performance improvement is proven. The 4-GiB +diagnostic is a new resource envelope, not a passing 2-GiB result. Offline plans +for 500/1,000-entry admission retain the full matrix but have not been executed. The [workspace/RustFS verification](performance/2026-09-30-workspace-rustfs.md) and the retained three-node baseline use `70bd25f142f1976fdd63ffe60e46e15ae276ffdc`. That baseline stopped after 20 fully bound windows and one interrupted creation diff --git a/docs/performance/2026-10-01-native-filesystem.md b/docs/performance/2026-10-01-native-filesystem.md index 31b144d..c152de7 100644 --- a/docs/performance/2026-10-01-native-filesystem.md +++ b/docs/performance/2026-10-01-native-filesystem.md @@ -1,6 +1,6 @@ # Compare RustFS filesystem backing -**Status: full native seed and initial owner-loss recovery passed; replacement load preflight running; performance unqualified.** +**Status: full native seed and initial recovery passed; V2 load failed after RustFS cgroup OOM; fresh 4-GiB full recovery running; performance unqualified.** This shared-Mac/Colima diagnostic keeps the 10,000-repository target. It tests filesystem backing without replacing the bound runtime or weakening correctness. @@ -131,8 +131,8 @@ sending signals, followed by a fresh **32.040-second** conservative wait. The replacement fleet is now its own persistent foreground server session, not a child of a finite recovery controller. It uses new node IDs, signing keys and local directories, but the same executable, deployment, provider, corpus, -admission and deadlines. Its full mandatory preflight was live at the -3,500-identity checkpoint; no timed window had started at that checkpoint. +admission and deadlines. Its mandatory preflight subsequently passed all +10,000 identities and all 100 Git v0/v2/LFS fixtures before timed load began. The unchanged 108-window attempt has a new output directory. It neither resumes the failed attempt nor removes a failing window. @@ -152,7 +152,7 @@ windows; they are shared-host work, not service performance proof. ### Arm every-ACK recovery without shortening the load -A separately bound post-load controller is live in `wait_bound_load`. Its +A separately bound post-load controller was armed in `wait_bound_load`. Its read-only live binding check and four offline guard tests passed. During load it only inspects the two bound process identities every 15 seconds; it does not poll provider state or send additional Git traffic. This still adds some host @@ -181,8 +181,121 @@ verifier does not orphan an attached fixture. Critical fixtures, full seeded corpus, acknowledged creations and acknowledged Git/LFS writes have separate verification stages; one failed stage does not skip the other scopes. Counts must match the audited ACK ledgers. There are no retries or failed-arrival -rollback assertions. This controller is armed, not completed, and a recovery -pass cannot turn a failed performance campaign into a pass. +rollback assertions. This controller was armed, not completed, and a recovery +pass cannot turn a failed performance campaign into a pass. In the actual V2 +attempt, RustFS OOM changed the provider boundary, so this controller stopped +before any owner-loss or recovery action. Its closed failure receipt is retained. + +### Retain the V2 OOM and all failed arrivals + +All three repetitions below offered 20 metadata requests/s for 120 seconds, +uniformly across 100 repositories, with concurrency 32. Latencies include every +scheduled arrival, including failures and driver drops; throughput counts only +successes completed inside the offered window. + +| Repetition | Outcomes (2,400 offered) | Successful requests/s | p50 / p95 / p99 (ms) | +| --- | --- | --- | --- | +| 1 | 2,400 OK | 19.983 | 59.220 / 463.515 / 1,511.428 | +| 2 | 2,400 OK | 20.000 | 49.283 / 131.151 / 205.672 | +| 3 | 2,328 OK; 29 driver busy; 43 HTTP 503 | 19.400 | 62.648 / 808.954 / 16,883.284 | + +The independent audits bind complete arrival sequences, deterministic selection, +latencies and resource boundaries. First-scheduled-hit/time-bin diagnostics +support cold activation contributing to the first tail, but do not explain the +whole failure. These groups do not replace the all-arrival metrics. Installed +CPython asyncio already sets TCP_NODELAY; no proxy-flag optimization is supported. +Small offline audits ran on the shared host during load and add overhead. + +The scoped Docker stream recorded OOM at **05:01:19.533453747 UTC** on +2026-10-01, followed by container death. Docker reported `OOMKilled=true`, +exit 137 and zero restarts. The VM kernel independently identified +`CONSTRAINT_MEMCG` for the exact fixture cgroup and killed `rustfs` PID 12755 +with 1,976,556 KiB anonymous RSS and 83,976 KiB file RSS. The VM had no swap; +a configured 4-GiB `MemorySwap` value did not provide usable swap. This proves +provider cgroup OOM, not a Cellule bottleneck, RustFS allocation mechanism, +or the cause of the earlier Mac-backed failure. + +After preserving that boundary, only the verified campaign PID 82894 received +SIGINT. The final record contains: + +| Scope | Retained accounting | +| --- | --- | +| Six closed windows | 14,400 arrivals: 7,128 OK, 29 driver busy, 403 HTTP 503, 6,840 transport errors | +| Interrupted seventh window | Exact 991-arrival sequence prefix, all transport errors; resource boundary incomplete | +| Original matrix | 108 declared: six closed, one interrupted, 101 unstarted | +| Load-write ACKs | Zero: no creation, push or LFS upload window started | + +The partial prefix does not fabricate outcomes for its remaining 1,409 declared +arrivals. The generic whole-campaign auditor rejects its orphan ledger; a +separate reconciliation audits the six closed ledgers and retains the exact +interrupted prefix. Neither turns this into a completed or passing matrix. +The diagnostic sidecar closed with 261 probe errors; gaps remain unqualified. + +### Change only the provider memory envelope for the next diagnostic + +After both controllers actually exited, a guarded batch sent SIGKILL to the +three old owners (80868/80928/80936). All exited -9 without forced fallback. +Recorded launcher/owner absence was followed by a **32.033208-second** wait. +The original OOM state, kernel evidence and all failed ledgers remain preserved. + +Only the owned, stopped provider's `HostConfig.Memory` changed: **2 → 4 GiB**. +CPU quota remains two; image, complete container Config, durable local volume, +other HostConfig values, binary, corpus, node admission and deadlines are +unchanged. The same container restarted at `2026-10-01T05:17:59.758811988Z`. +Its dynamic loopback endpoint changed from port 32787 to **32799**; a new +binding records that start/endpoint/resource envelope. The old trial and binding +were not overwritten. No data, unrelated container or UI service was removed. + +```mermaid +flowchart LR + failed[2 GiB V2 OOM
failed ledgers retained] --> loss[Verified owner loss
absence + 32.033s] + loss --> provider[Same data/image/2 CPUs
new 4 GiB envelope] + provider --> fresh[Independent fleet
fresh owners/local directories] + fresh --> verify[Full 10000/100 + critical 2
no retry or reseed] + verify --> next[New complete 108-window attempt
only after full recovery passes] +``` + +The new independent foreground fleet started at 05:19:35 UTC with fresh owner +IDs. Both critical repositories have passed four exact ref inventories, bytes, +notes and strict full fsck. The full 10,000-identity/100-fixture verifier is +running with four workers, 30-second HTTP and 120-second Git deadlines. +Cgroup current/max/peak/events/stat, CPU and shared-VM pressure are retained +with raw file order and host-monotonic probe bounds. These non-atomic snapshots +and self/reaped-child CPU are diagnostics, not live-child or priced cost proof. +The 4-GiB change is not yet a verified fix and cannot qualify the failed 2-GiB +attempt or demonstrate a code improvement. + +A separate read-only controller is now armed in `wait_bound_full_recovery`. +Its live binding check and two offline tests (23 rejection cases) passed. +It waits for the exact verifier's actual exit, rereads its terminal receipt, +checks full 10K/100, critical-2, original deadlines and closed sidecar digests, +then starts a **new entire 108-window attempt**, including mandatory preflight, +in a new directory. It sends no signals and changes no provider settings. +An incomplete recovery, PID reuse, changed provider/source, reused output or +shortened plan prevents launch. It is waiting, not a load or recovery pass. + +At 05:40:44 UTC, the recovery sidecar had 132 samples: cgroup memory had reached +the 4-GiB limit, anonymous memory was 1,658,716,160 bytes and file memory was +1,831,895,040 bytes, with zero observed OOM kills. The separate process probe +identified PID 1 as RustFS with 1,620,432 KiB anonymous RSS. Its signed `types=1` +console response was **scanner metrics, not allocator metrics**: a deep scan +was active, with 98,771 objects and 130,058 directories scanned. The raw response +is retained without relabelling it as allocation evidence. Reclaim counters +show pressure, but RSS does not distinguish live allocations from freed pages +retained by the allocator; scanner activity is not a proven root cause. + +The source at the image-reported revision `d47f54b` already +[defaults allocator reclaim to enabled](https://github.com/rustfs/rustfs/blob/d47f54bfb2f39f48bd1adda334bd27e151fe85b8/crates/config/src/constants/runtime.rs) +and [makes object-cache memory resolution container-aware](https://github.com/rustfs/rustfs/blob/d47f54bfb2f39f48bd1adda334bd27e151fe85b8/crates/object-data-cache/src/runtime_memory.rs). +Those defaults are not proof of effective runtime settings or a fix here; +old upstream reports do not justify blindly toggling either knob in this bound +attempt. No live limit, environment or source was changed during verification. + +Two offline admission-plan tests passed, including exact full-matrix retention +and rejection of corpus/rate/duration/trimming changes. Independent normalized +diffs confirm that the 500- and 1,000-entry variants change only admission. +Each keeps 108 windows, 8,640 offered seconds and 114,960 offered arrivals. +Neither variant has been executed; larger caps do not prove residency or capacity. ## Separate diagnostics from performance @@ -253,11 +366,31 @@ Artifacts are retained under `canopy-native-filesystem-ceXFad8I`: | `test_finish_native_load_v2.py` | `2ad420292ccfd4fe66072b216cb816603dbeeb3c20dbb86568b56cec10495762` | | `post-load-v2-guard-tests.log` | `28157ddb518f5c1d4f887b07be9710a779324b00d8f73347311839e52606bd59` | | `post-load-v2-check.json` | `57befacbc9f31ac6564b8cab618c45be0851b8f81c3996c1334eadae485cb281` | +| `native-load-v2/preflight.json` | `9564ba67057758324738dfbf5223ec79c172df1378458145ed586bd786551770` | +| `native-load-v2/campaign.json` | `3dda98f8c3820e151bc2f7eb2c0d72273972a6cbeb3250a84bacce160c2427af` | +| `native-load-v2-controller.json` | `3f82e91c1c4524937bb30f31c78ef39e55579f93e928b2fedd16978c66e909d9` | +| `native-post-load-v2.json` | `968c99e941dbd442326bfeb575a9ba46e46b562f17f0b584dbc857ac331fa0d3` | +| `native-windows-0000-0002-audit.json` | `ebf2ca6aaa189a4facdb0cc5033e76ebbcd9cfef5cd77cc26efbc889a401a7b2` | +| `native-v2-oom-ledger-audit.json` | `619c3ed1e0be121178a7b92379a1ce8b924e9c8d3ee71efc886757580f1e6314` | +| `native-v2-provider-oom-kernel.log` | `496836038f63914a34342a89d0895242d4c997f3dd7c31eb8470beee9304e847` | +| `native-v2-provider-oom-cancellation.json` | `7eca8217db7b9a06246626c1c1cfc652b5d4148f4c04397684eb204fd4ca4105` | +| `native-v2-oom-owner-loss.json` | `c63b5e7d4433e0130a899b0bf0efdb8d44d5b94ad08b6d5ffb5b319d9c6f5720` | +| `provider-recovery-4g.json` | `565cf80795d7d9da78eb4525f247ad3e3a0f687ae6b5804392d09304137d2da0` | +| `fleet-native-recovery-4g/ready.json` | `04b630ac1ad771a62d936071c58823c74243968f350bca14bfb15fa5d38d78af` | +| `verify_native_recovery_4g.py` | `a194d74fed89cc45048fc3378416877e3413cc1cc16c824c0446244f562be916` | +| `critical-recovered-4g.json` | `1031bf601b1e63970198a385fe4f7bacae0124b6d9861a8254729f371967ef31` | +| `start_native_load_4g.py` | `9e3ba38b22b1188923d06a4b01125cdc7921763402368815fc91b41ae10c2cef` | +| `test_start_native_load_4g.py` | `8113463ea6ad6c9e5076a94e05b5c45821f5e040ac9ca1405bd0d223f698b6b0` | +| `load-4g-guard-tests.log` | `ad1dabaf6fd44b97c2ba48c91cc833bc69a88c2e8baf813908fc24f5d1465527` | +| `load-4g-check.json` | `6865f1c3646548feed611843d84a5eb1daa635801e5df78e3c3862ad63b93a61` | +| `native-recovery-4g-memory-probe.json` | `b078266a61591ff18db05e0efabdca1e11433452e8898ffe65b9bf1e6ae91129` | +| `higher-admission-plans-v1/preparation.json` | `f2fec01bc047cbe800580b3c8515fcaf2ab29c3af1aee8fee0ba9327a163bc85` | The baseline's new independent ledger audit is `5416540c9b78b23e5c89ff24e771ab58012af847df4ec961adfc597bff010639` -in `canopy-three-proxy-q3FO2z`. The replacement load and post-load recovery -controllers are live and have no final completion digests. Setup timings are +in `canopy-three-proxy-q3FO2z`. V2 load and its stopped post-load controller have +closed failure digests above; the separate 4-GiB full verifier has no completion +digest yet. Setup timings are not scheduled throughput. This is not isolated Linux reference capacity, a proven Cellule bottleneck or a passing latest-pin executable comparison. Both hosted workflows for source @@ -265,3 +398,7 @@ pin `0573f489` passed ([36806695489](https://github.com/crabbuild/canopy/actions/runs/36806695489), [36806698917](https://github.com/crabbuild/canopy/actions/runs/36806698917)); the running comparison binary remains the original `e07670e` artifact. +Both harness (75 tests) and Rust jobs also passed for published head `772d1dc` +([36816818943](https://github.com/crabbuild/canopy/actions/runs/36816818943), +[36816813337](https://github.com/crabbuild/canopy/actions/runs/36816813337)). +Hosted verification is not the outstanding local full-matrix qualification. From e5ccb62ebeb791446090af6d1c9c0417bf3450e3 Mon Sep 17 00:00:00 2001 From: forhappy Date: Wed, 30 Sep 2026 22:47:52 -0700 Subject: [PATCH 19/60] docs(perf): record RustFS allocator instrumentation gap --- docs/performance/2026-10-01-native-filesystem.md | 8 ++++++++ 1 file changed, 8 insertions(+) diff --git a/docs/performance/2026-10-01-native-filesystem.md b/docs/performance/2026-10-01-native-filesystem.md index c152de7..890cc42 100644 --- a/docs/performance/2026-10-01-native-filesystem.md +++ b/docs/performance/2026-10-01-native-filesystem.md @@ -291,6 +291,13 @@ Those defaults are not proof of effective runtime settings or a fix here; old upstream reports do not justify blindly toggling either knob in this bound attempt. No live limit, environment or source was changed during verification. +The pinned [console collector](https://github.com/rustfs/rustfs/blob/d47f54bfb2f39f48bd1adda334bd27e151fe85b8/crates/ecstore/src/services/metrics_realtime.rs) +defines MEM as `1 << 6` but leaves that collection branch unimplemented. A +separate signed `types=64` request completed with empty aggregated/by-host/by-disk +samples. That is an **instrumentation gap**, not zero allocator usage or evidence +that memory is safe. Its raw receipt is retained separately from the scanner +probe; allocator attribution remains open. + Two offline admission-plan tests passed, including exact full-matrix retention and rejection of corpus/rate/duration/trimming changes. Independent normalized diffs confirm that the 500- and 1,000-entry variants change only admission. @@ -384,6 +391,7 @@ Artifacts are retained under `canopy-native-filesystem-ceXFad8I`: | `load-4g-guard-tests.log` | `ad1dabaf6fd44b97c2ba48c91cc833bc69a88c2e8baf813908fc24f5d1465527` | | `load-4g-check.json` | `6865f1c3646548feed611843d84a5eb1daa635801e5df78e3c3862ad63b93a61` | | `native-recovery-4g-memory-probe.json` | `b078266a61591ff18db05e0efabdca1e11433452e8898ffe65b9bf1e6ae91129` | +| `native-recovery-4g-allocator-probe.json` | `47da1e7f9ade991395715eaa211bdd830e48578f214622c85c4c9519e258381d` | | `higher-admission-plans-v1/preparation.json` | `f2fec01bc047cbe800580b3c8515fcaf2ab29c3af1aee8fee0ba9327a163bc85` | The baseline's new independent ledger audit is From 38a51dddf91994982ee4ccc0f44e58dc6e09138a Mon Sep 17 00:00:00 2001 From: forhappy Date: Wed, 30 Sep 2026 23:09:34 -0700 Subject: [PATCH 20/60] chore(deps): sync Cellule main and bind every-ACK recovery --- Cargo.lock | 12 ++--- crates/canopy-server/Cargo.toml | 10 ++-- docs/contracts.md | 6 ++- docs/delivery-plan.md | 5 +- docs/performance-plan.md | 9 ++-- .../2026-10-01-native-filesystem.md | 53 +++++++++++++++++++ 6 files changed, 78 insertions(+), 17 deletions(-) diff --git a/Cargo.lock b/Cargo.lock index 30771b7..5e81865 100644 --- a/Cargo.lock +++ b/Cargo.lock @@ -448,7 +448,7 @@ dependencies = [ [[package]] name = "cellule-app" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=0573f48998c4e5343cd8b463d79b7bc1820c923c#0573f48998c4e5343cd8b463d79b7bc1820c923c" +source = "git+https://github.com/crabbuild/cellule.git?rev=a4500add51764fa0415791aefbfa561db6ada203#a4500add51764fa0415791aefbfa561db6ada203" dependencies = [ "blake3", "cellule-runtime", @@ -457,7 +457,7 @@ dependencies = [ [[package]] name = "cellule-host" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=0573f48998c4e5343cd8b463d79b7bc1820c923c#0573f48998c4e5343cd8b463d79b7bc1820c923c" +source = "git+https://github.com/crabbuild/cellule.git?rev=a4500add51764fa0415791aefbfa561db6ada203#a4500add51764fa0415791aefbfa561db6ada203" dependencies = [ "cellule-app", "cellule-runtime", @@ -471,7 +471,7 @@ dependencies = [ [[package]] name = "cellule-ltx" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=0573f48998c4e5343cd8b463d79b7bc1820c923c#0573f48998c4e5343cd8b463d79b7bc1820c923c" +source = "git+https://github.com/crabbuild/cellule.git?rev=a4500add51764fa0415791aefbfa561db6ada203#a4500add51764fa0415791aefbfa561db6ada203" dependencies = [ "async-trait", "blake3", @@ -493,7 +493,7 @@ dependencies = [ [[package]] name = "cellule-runtime" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=0573f48998c4e5343cd8b463d79b7bc1820c923c#0573f48998c4e5343cd8b463d79b7bc1820c923c" +source = "git+https://github.com/crabbuild/cellule.git?rev=a4500add51764fa0415791aefbfa561db6ada203#a4500add51764fa0415791aefbfa561db6ada203" dependencies = [ "blake3", "bytes", @@ -519,7 +519,7 @@ dependencies = [ [[package]] name = "cellule-store" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=0573f48998c4e5343cd8b463d79b7bc1820c923c#0573f48998c4e5343cd8b463d79b7bc1820c923c" +source = "git+https://github.com/crabbuild/cellule.git?rev=a4500add51764fa0415791aefbfa561db6ada203#a4500add51764fa0415791aefbfa561db6ada203" dependencies = [ "async-trait", "blake3", @@ -542,7 +542,7 @@ dependencies = [ [[package]] name = "cellule-types" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=0573f48998c4e5343cd8b463d79b7bc1820c923c#0573f48998c4e5343cd8b463d79b7bc1820c923c" +source = "git+https://github.com/crabbuild/cellule.git?rev=a4500add51764fa0415791aefbfa561db6ada203#a4500add51764fa0415791aefbfa561db6ada203" dependencies = [ "schemars", "serde", diff --git a/crates/canopy-server/Cargo.toml b/crates/canopy-server/Cargo.toml index 1d91493..744ee1b 100644 --- a/crates/canopy-server/Cargo.toml +++ b/crates/canopy-server/Cargo.toml @@ -15,11 +15,11 @@ axum = "0.8.9" base64 = "0.22" blake3 = "1.8" bytes = "1.11" -cellule-app = { git = "https://github.com/crabbuild/cellule.git", rev = "0573f48998c4e5343cd8b463d79b7bc1820c923c" } -cellule-host = { git = "https://github.com/crabbuild/cellule.git", rev = "0573f48998c4e5343cd8b463d79b7bc1820c923c" } -cellule-ltx = { git = "https://github.com/crabbuild/cellule.git", rev = "0573f48998c4e5343cd8b463d79b7bc1820c923c", features = ["replica"] } -cellule-runtime = { git = "https://github.com/crabbuild/cellule.git", rev = "0573f48998c4e5343cd8b463d79b7bc1820c923c" } -cellule-store = { git = "https://github.com/crabbuild/cellule.git", rev = "0573f48998c4e5343cd8b463d79b7bc1820c923c" } +cellule-app = { git = "https://github.com/crabbuild/cellule.git", rev = "a4500add51764fa0415791aefbfa561db6ada203" } +cellule-host = { git = "https://github.com/crabbuild/cellule.git", rev = "a4500add51764fa0415791aefbfa561db6ada203" } +cellule-ltx = { git = "https://github.com/crabbuild/cellule.git", rev = "a4500add51764fa0415791aefbfa561db6ada203", features = ["replica"] } +cellule-runtime = { git = "https://github.com/crabbuild/cellule.git", rev = "a4500add51764fa0415791aefbfa561db6ada203" } +cellule-store = { git = "https://github.com/crabbuild/cellule.git", rev = "a4500add51764fa0415791aefbfa561db6ada203" } ed25519-dalek = "2" flate2 = "1.1" futures-core = "0.3" diff --git a/docs/contracts.md b/docs/contracts.md index 23fb473..35dab29 100644 --- a/docs/contracts.md +++ b/docs/contracts.md @@ -1186,8 +1186,10 @@ and the operation-5/8/9 codec changes require a fresh development storage prefix; there is no upgrade reader for older development databases. The module descriptor and object paths will become compatibility boundaries at the first -persistent preview. The current build pins Cellule revision -`a28de7bc09ce36d87e642adc4f4b6be50d6fcb69`. The earlier entity-partition +persistent preview. The current source pins Cellule revision +`a4500add51764fa0415791aefbfa561db6ada203`; running qualification artifacts +remain bound to their recorded revisions, not silently replaced by this pin. +The earlier entity-partition cutover also made pre-cutover prefixes incompatible; no migration is available. ## Accounts and authentication diff --git a/docs/delivery-plan.md b/docs/delivery-plan.md index ce7f66e..ee2c60d 100644 --- a/docs/delivery-plan.md +++ b/docs/delivery-plan.md @@ -72,9 +72,12 @@ These are the next implementation slices. Each slice must update its contract, b 1. Expand the S3-compatible process smoke into a node/lease fault matrix and test the target production object store. Run the checked-in CI workflow on - a Canopy remote. The current build pins Cellule revision `a28de7b` and + a Canopy remote. The current source pins Cellule revision `a4500add` and owns a startup storage probe with cleanup on failure. Historical dependency qualification runs below remain evidence for their recorded revisions only. + The [live three-node trial](performance/2026-10-01-native-filesystem.md) + retains its separately bound `e07670e` executable; latest source is not a + claim that the running artifact was rebuilt or performance-qualified. 2. Execute the [repository-density performance plan](performance-plan.md), then qualify the bounded Linux deployment and residency under faults and larger hot sets. Each node reserves one SQL slot for Directory takeover and admits diff --git a/docs/performance-plan.md b/docs/performance-plan.md index 9bba33d..18ef606 100644 --- a/docs/performance-plan.md +++ b/docs/performance-plan.md @@ -2,9 +2,12 @@ Use this plan to design capacity work and interpret Canopy benchmark results. It separates demonstrated behavior from proposed targets. The current Cellule -dependency pin is `0573f48998c4e5343cd8b463d79b7bc1820c923c`, a docs-only -upstream advance from `e07670e` with identical Rust workspace/crates. Both full -hosted workflows passed; that does not establish three-node capacity. +dependency pin is `a4500add51764fa0415791aefbfa561db6ada203`, an upstream +web/documentation advance with the same Rust crate tree and root Cargo objects +as `0573f489` and `e07670e`. Locked metadata resolves all six packages to the +new commit without unrelated dependency changes. Hosted verification for the +previous source pin passed; the latest-pin production artifact and three-node +qualification remain separate open gates. None establishes reference capacity. The [three-node proxy campaign](performance/2026-09-30-three-node-proxy.md) retains the earlier fixture race and its regression, plus the candidate's failed seed and successful recovery of every recorded ACK. diff --git a/docs/performance/2026-10-01-native-filesystem.md b/docs/performance/2026-10-01-native-filesystem.md index 890cc42..59574ae 100644 --- a/docs/performance/2026-10-01-native-filesystem.md +++ b/docs/performance/2026-10-01-native-filesystem.md @@ -274,6 +274,33 @@ in a new directory. It sends no signals and changes no provider settings. An incomplete recovery, PID reuse, changed provider/source, reused output or shortened plan prevents launch. It is waiting, not a load or recovery pass. +A second independently hosted controller is armed in `wait_bound_load` for +every-ACK recovery. Three offline tests passed, covering 28 rejection cases, +and its actual read-only binding check passed. It captures the real campaign +PID/command/kernel identity when load begins; a never-observed or replaced +campaign, live controller/event stream, incomplete sidecars, changed input or +provider, and orphan ledger stop before owner signals. + +```mermaid +flowchart LR + closed[Bound campaign/controller/event stream
actually absent] --> ledger[Audit all closed ledgers
retain failures and every ACK] + ledger --> loss[Validate whole owner batch
record loss and 32s absence wait] + loss --> fresh[New owner IDs and directories
independent process session] + fresh --> critical[Critical fixtures] + fresh --> corpus[Full 10000/100 corpus] + fresh --> creations[Every creation ACK] + fresh --> writes[Every Git/LFS write ACK] +``` + +The same guarded signal/fresh-boundary functions used by the preserved V2 +controller are digest-bound inputs, not reimplemented weaker guards. Failed +performance does not skip recovery of readable ACK ledgers. Each of the four +verification scopes retains failure and continues to the other scopes; counts +must match the audited ledgers. No-ACK stages explicitly say they are not traffic +passes. Provider/binary/admission/deadlines stay fixed; no automatic retries, +restarts or data deletion are permitted. This is an armed gate, not a recovered +ACK claim or completed load result. + At 05:40:44 UTC, the recovery sidecar had 132 samples: cgroup memory had reached the 4-GiB limit, anonymous memory was 1,658,716,160 bytes and file memory was 1,831,895,040 bytes, with zero observed OOM kills. The separate process probe @@ -304,6 +331,26 @@ diffs confirm that the 500- and 1,000-entry variants change only admission. Each keeps 108 windows, 8,640 offered seconds and 114,960 offered arrivals. Neither variant has been executed; larger caps do not prove residency or capacity. +### Track latest Cellule source without changing the live artifact + +At the source audit, `origin/main` was +`a4500add51764fa0415791aefbfa561db6ada203`, one upstream +[web/documentation commit](https://github.com/crabbuild/cellule/pull/35) after +`0573f489`. Five direct declarations and six lockfile source entries now select +that exact commit, with no other manifest/lockfile changes. Locked full Cargo +metadata resolved all six packages to the new SHA. No new executable or +workspace `target` directory was created; metadata briefly waited on the shared +Cargo package-cache lock. Other host work remains a shared-machine variable. + +The actual cached checkout and upstream Git trees have the same `crates` object +`587e3215ce1bc0c67db3390a137623b6fead01ac`, Cargo manifest object +`6b7e13d8439558455ea04458357f3a70992fa367` and lock object +`b512761a55bba5165101ce9905dfe7f603f5b079` as `0573f489`/`e07670e`. +The audit also rechecked every frozen input in the live full-recovery, load and +post-load controllers. All matched after source-pin editing; none now runs a +newly built binary. This advance is not a Cellule performance fix, a passing +latest-pin local artifact, or a reason to skip original end-to-end qualification. + ## Separate diagnostics from performance The closed ACK-recovery observer retained 173 samples with zero probe errors. @@ -390,9 +437,15 @@ Artifacts are retained under `canopy-native-filesystem-ceXFad8I`: | `test_start_native_load_4g.py` | `8113463ea6ad6c9e5076a94e05b5c45821f5e040ac9ca1405bd0d223f698b6b0` | | `load-4g-guard-tests.log` | `ad1dabaf6fd44b97c2ba48c91cc833bc69a88c2e8baf813908fc24f5d1465527` | | `load-4g-check.json` | `6865f1c3646548feed611843d84a5eb1daa635801e5df78e3c3862ad63b93a61` | +| `finish_native_load_4g.py` | `b2fc51fe7cf336d2c3e64f2807fd52744f5561efd8b078446350422d3f37f570` | +| `test_finish_native_load_4g.py` | `4d198e1148ce20c599b927b8ed33cf498d606f9aa5b3d8c7057d40ba2cb31b3f` | +| `post-load-4g-guard-tests.log` | `e8b87fc14d3c7fb0ea88fff398ba3c257e73c6ca0cc85a63fa696cc53345a07c` | +| `post-load-4g-check.json` | `b5a85134289a9689cbe63e8bf4c0ae9590cfc4965212a317ed66133c5b34af59` | | `native-recovery-4g-memory-probe.json` | `b078266a61591ff18db05e0efabdca1e11433452e8898ffe65b9bf1e6ae91129` | | `native-recovery-4g-allocator-probe.json` | `47da1e7f9ade991395715eaa211bdd830e48578f214622c85c4c9519e258381d` | | `higher-admission-plans-v1/preparation.json` | `f2fec01bc047cbe800580b3c8515fcaf2ab29c3af1aee8fee0ba9327a163bc85` | +| `metadata-a4500add.json` | `78cda3f194d9147dfabc8f1d139de6e2c43ed30f4c3b9736642c9601136dd4a9` | +| `cellule-a450-source-audit.json` | `87bac9bc1c54e8fba393cfec424436532086e0810f03f48b957c5d6cfb5d01a3` | The baseline's new independent ledger audit is `5416540c9b78b23e5c89ff24e771ab58012af847df4ec961adfc597bff010639` From a41131a01225ccc0e98d2ce52d449f509308f1ab Mon Sep 17 00:00:00 2001 From: forhappy Date: Wed, 30 Sep 2026 23:33:38 -0700 Subject: [PATCH 21/60] docs: record full post-OOM recovery and load transition --- docs/performance-plan.md | 10 ++- .../2026-10-01-native-filesystem.md | 81 +++++++++++++------ 2 files changed, 64 insertions(+), 27 deletions(-) diff --git a/docs/performance-plan.md b/docs/performance-plan.md index 18ef606..501cb51 100644 --- a/docs/performance-plan.md +++ b/docs/performance-plan.md @@ -5,9 +5,10 @@ It separates demonstrated behavior from proposed targets. The current Cellule dependency pin is `a4500add51764fa0415791aefbfa561db6ada203`, an upstream web/documentation advance with the same Rust crate tree and root Cargo objects as `0573f489` and `e07670e`. Locked metadata resolves all six packages to the -new commit without unrelated dependency changes. Hosted verification for the -previous source pin passed; the latest-pin production artifact and three-node -qualification remain separate open gates. None establishes reference capacity. +new commit without unrelated dependency changes. Both hosted Rust and harness +jobs passed for source head `38a51dd` with this pin; the latest-pin local +production artifact and three-node qualification remain separate open gates. +None establishes reference capacity. The [three-node proxy campaign](performance/2026-09-30-three-node-proxy.md) retains the earlier fixture race and its regression, plus the candidate's failed seed and successful recovery of every recorded ACK. @@ -19,7 +20,8 @@ seed and successful recovery of every recorded ACK. | Every candidate seed ACK | All 3,561 identities and 39 populated fixtures passed fresh-owner recovery; independent ledger/binding and 78-clone audit passed | | Native-volume comparison | Same runtime/provider limits; all 10,000 identities and 100 Git/LFS fixtures seeded in 3,179.240 s and passed initial three-owner-loss recovery; independent 200-clone audit passed | | Native scheduled load | First attempt failed before any window. Independent-session V2 passed full preflight, then RustFS was cgroup-OOM killed: six closed windows, one interrupted prefix and 101 unstarted windows retained | -| Post-OOM diagnostic | Three old owners killed as a verified batch; 32.033-s absence wait recorded. Same provider/data restarted with memory raised from 2 to 4 GiB; critical-2 recovery passed, full 10K/100 recovery running | +| Post-OOM diagnostic | Three old owners killed as a verified batch; 32.033-s absence wait recorded. Same provider/data restarted with memory raised from 2 to 4 GiB; full 10K/100 and critical-2 recovery passed, with independent 200-clone/four-mirror audits | +| New complete load attempt | Original 108-window/114,960-arrival matrix launched after terminal recovery; mandatory full preflight running, no timed window completed at this checkpoint; every-ACK controller captured the actual campaign identity | | Performance qualification | Full scheduled load, every post-load ACK recovery, higher admission profiles and matched comparisons remain open | The [native-filesystem trial](performance/2026-10-01-native-filesystem.md) diff --git a/docs/performance/2026-10-01-native-filesystem.md b/docs/performance/2026-10-01-native-filesystem.md index 59574ae..7d0e9b9 100644 --- a/docs/performance/2026-10-01-native-filesystem.md +++ b/docs/performance/2026-10-01-native-filesystem.md @@ -1,6 +1,6 @@ # Compare RustFS filesystem backing -**Status: full native seed and initial recovery passed; V2 load failed after RustFS cgroup OOM; fresh 4-GiB full recovery running; performance unqualified.** +**Status: full seed, initial recovery and post-OOM 4-GiB recovery passed; V2 load failed; new complete load attempt in mandatory preflight; performance unqualified.** This shared-Mac/Colima diagnostic keeps the 10,000-repository target. It tests filesystem backing without replacing the bound runtime or weakening correctness. @@ -255,30 +255,41 @@ flowchart LR verify --> next[New complete 108-window attempt
only after full recovery passes] ``` -The new independent foreground fleet started at 05:19:35 UTC with fresh owner -IDs. Both critical repositories have passed four exact ref inventories, bytes, -notes and strict full fsck. The full 10,000-identity/100-fixture verifier is -running with four workers, 30-second HTTP and 120-second Git deadlines. -Cgroup current/max/peak/events/stat, CPU and shared-VM pressure are retained -with raw file order and host-monotonic probe bounds. These non-atomic snapshots -and self/reaped-child CPU are diagnostics, not live-child or priced cost proof. -The 4-GiB change is not yet a verified fix and cannot qualify the failed 2-GiB -attempt or demonstrate a code improvement. - -A separate read-only controller is now armed in `wait_bound_full_recovery`. -Its live binding check and two offline tests (23 rejection cases) passed. -It waits for the exact verifier's actual exit, rereads its terminal receipt, -checks full 10K/100, critical-2, original deadlines and closed sidecar digests, -then starts a **new entire 108-window attempt**, including mandatory preflight, -in a new directory. It sends no signals and changes no provider settings. +The independent foreground fleet started at 05:19:35 UTC with fresh owner IDs. +The full verifier exited **0 at 06:27:18 UTC**, retaining four workers, +30-second HTTP and 120-second Git deadlines, with no retry or reseed. + +| Closed post-OOM recovery gate | Result | +| --- | --- | +| Original corpus identities | All 10,000 exact UUID/name pairs passed | +| Git/LFS fixtures | All 100 passed stock-Git v0/v2 clones, exact commits/content, strict full fsck and exact network LFS size/hash | +| Critical fixtures | Both identities and four exact ref inventories, payload, notes and strict full fsck passed | +| Independent local Git audit | All 200 corpus clones rechecked for HEAD/base commits, committed/working-copy bytes and strict full fsck; four critical mirrors independently rechecked | +| Closed observer integrity | All 631 samples, process identities, timing bounds and sidecar digests passed; zero probe errors and zero observed OOM/oom_kill counters | + +The local audits do not download LFS or recheck network identity again; those +checks remain bound to the inspected original verifier. Cgroup current/max/peak/ +events/stat, CPU and shared-VM pressure retain raw file order and host-monotonic +probe bounds. Non-atomic snapshots and self/reaped-child CPU are diagnostics, +not live-child or priced cost proof. This recovery pass cannot qualify the failed +2-GiB load, prove the new full matrix, or demonstrate a code improvement. + +A separately bound controller passed its live check and two offline tests +(23 rejection cases). It observed the exact verifier's actual exit, reread the +terminal full 10K/100 and critical-2 receipts, original deadlines and closed +sidecar digests, then launched a **new entire 108-window attempt** in a new +directory. The real campaign PID is 5740; its mandatory full preflight is +running, with no timed window completed at this checkpoint. It sends no signals +and changes no provider settings. An incomplete recovery, PID reuse, changed provider/source, reused output or -shortened plan prevents launch. It is waiting, not a load or recovery pass. +shortened plan prevents launch. Launch/preflight is not a load pass: all +114,960 declared arrivals over 8,640 offered seconds remain required. A second independently hosted controller is armed in `wait_bound_load` for every-ACK recovery. Three offline tests passed, covering 28 rejection cases, -and its actual read-only binding check passed. It captures the real campaign -PID/command/kernel identity when load begins; a never-observed or replaced -campaign, live controller/event stream, incomplete sidecars, changed input or +and its actual read-only binding check passed. It captured campaign PID 5740's +actual command/kernel identity and remains in `wait_bound_load`; a +never-observed or replaced campaign, live controller/event stream, incomplete sidecars, changed input or provider, and orphan ledger stop before owner signals. ```mermaid @@ -301,6 +312,14 @@ passes. Provider/binary/admission/deadlines stay fixed; no automatic retries, restarts or data deletion are permitted. This is an armed gate, not a recovered ACK claim or completed load result. +The closed observer reached a sampled maximum `memory.current` of 4,294,967,296 +bytes. At its final sample (06:27:17 UTC), current memory was 4,108,148,736 bytes, +anonymous memory 1,552,654,336 bytes and file memory 1,404,837,888 bytes. The +anonymous counter fell from the earlier 06:16 observation; neither growth nor +that fall distinguishes live allocation from allocator retention. No scoped +OOM, die, pause, update or restart event was observed in this recovery window. +The longest raw probe was 3.026 seconds; observer overhead remains explicit. + At 05:40:44 UTC, the recovery sidecar had 132 samples: cgroup memory had reached the 4-GiB limit, anonymous memory was 1,658,716,160 bytes and file memory was 1,831,895,040 bytes, with zero observed OOM kills. The separate process probe @@ -433,6 +452,12 @@ Artifacts are retained under `canopy-native-filesystem-ceXFad8I`: | `fleet-native-recovery-4g/ready.json` | `04b630ac1ad771a62d936071c58823c74243968f350bca14bfb15fa5d38d78af` | | `verify_native_recovery_4g.py` | `a194d74fed89cc45048fc3378416877e3413cc1cc16c824c0446244f562be916` | | `critical-recovered-4g.json` | `1031bf601b1e63970198a385fe4f7bacae0124b6d9861a8254729f371967ef31` | +| `critical-recovered-4g-local-audit.json` | `2970d432a30b34574c4afa9f2dd08b91015bc9e7364e53fc5a597efeee92fd96` | +| `native-recovery-4g.json` | `b3a1f3679238cd2a8d9fa5772fe6a099f0ae5e54b4d3b4a4d1ff36a65d4458d5` | +| `full-recovered-4g.json` | `49d8acf40cc8e6cc253f7ac4db0cc8ddff143c913000b4e42ee868a40d2915de` | +| `full-recovered-4g-local-audit.json` | `2a503ba240bf59be51e222e3cb34b9cdca1892ed518f66f758922d70b00150f3` | +| `native-recovery-4g-observer/observation.json` | `8cc916668c7bc5e9ff0f427c32fe3956d936f7dd5976dd9681bee2950bdb2e25` | +| `native-recovery-4g-observer-audit.json` | `c1c91bf017c83a5b2e477b1fa5da5d5fb8906b23a760723383e7f72baba34b01` | | `start_native_load_4g.py` | `9e3ba38b22b1188923d06a4b01125cdc7921763402368815fc91b41ae10c2cef` | | `test_start_native_load_4g.py` | `8113463ea6ad6c9e5076a94e05b5c45821f5e040ac9ca1405bd0d223f698b6b0` | | `load-4g-guard-tests.log` | `ad1dabaf6fd44b97c2ba48c91cc833bc69a88c2e8baf813908fc24f5d1465527` | @@ -446,12 +471,15 @@ Artifacts are retained under `canopy-native-filesystem-ceXFad8I`: | `higher-admission-plans-v1/preparation.json` | `f2fec01bc047cbe800580b3c8515fcaf2ab29c3af1aee8fee0ba9327a163bc85` | | `metadata-a4500add.json` | `78cda3f194d9147dfabc8f1d139de6e2c43ed30f4c3b9736642c9601136dd4a9` | | `cellule-a450-source-audit.json` | `87bac9bc1c54e8fba393cfec424436532086e0810f03f48b957c5d6cfb5d01a3` | +| `harness-38a51dd-linux-36823311167-job.log` | `963a1636810e0efdc768e1d6f4c8297f3d3b07c7d441e99919ee8969d8ab0bcf` | +| `rust-38a51dd-linux-36823311167-job.log` | `2384cf85dc36736bb8560902e71227ba823b3120cbb109856b2a040159f99842` | The baseline's new independent ledger audit is `5416540c9b78b23e5c89ff24e771ab58012af847df4ec961adfc597bff010639` in `canopy-three-proxy-q3FO2z`. V2 load and its stopped post-load controller have -closed failure digests above; the separate 4-GiB full verifier has no completion -digest yet. Setup timings are +closed failure digests above. The separate 4-GiB full verifier and independent +audits now have closed digests; the new load and post-load controllers are still +live and have no completion digests. Setup timings are not scheduled throughput. This is not isolated Linux reference capacity, a proven Cellule bottleneck or a passing latest-pin executable comparison. Both hosted workflows for source @@ -463,3 +491,10 @@ Both harness (75 tests) and Rust jobs also passed for published head `772d1dc` ([36816818943](https://github.com/crabbuild/canopy/actions/runs/36816818943), [36816813337](https://github.com/crabbuild/canopy/actions/runs/36816813337)). Hosted verification is not the outstanding local full-matrix qualification. +Both Rust and harness jobs also passed for source head `38a51dd`, selecting +Cellule `a4500add` +([36823311167](https://github.com/crabbuild/canopy/actions/runs/36823311167), +[36823307410](https://github.com/crabbuild/canopy/actions/runs/36823307410)). +The retained PR-run logs explicitly show 75 harness tests passing and the +locked Rust workspace tests using that revision. This is hosted correctness, +not a new local release executable or performance comparison. From 8e183294d33eaffe3e72651b1df35d4ffe81b348 Mon Sep 17 00:00:00 2001 From: forhappy Date: Wed, 30 Sep 2026 23:51:59 -0700 Subject: [PATCH 22/60] test: add scheduled critical Git workflows and recovery guards --- docs/performance-plan.md | 1 + .../2026-10-01-native-filesystem.md | 41 +++ scripts/benchmark_critical_git.py | 271 ++++++++++++++++++ scripts/test_benchmark_critical_git.py | 187 ++++++++++++ 4 files changed, 500 insertions(+) create mode 100644 scripts/benchmark_critical_git.py create mode 100644 scripts/test_benchmark_critical_git.py diff --git a/docs/performance-plan.md b/docs/performance-plan.md index 501cb51..07e704e 100644 --- a/docs/performance-plan.md +++ b/docs/performance-plan.md @@ -22,6 +22,7 @@ seed and successful recovery of every recorded ACK. | Native scheduled load | First attempt failed before any window. Independent-session V2 passed full preflight, then RustFS was cgroup-OOM killed: six closed windows, one interrupted prefix and 101 unstarted windows retained | | Post-OOM diagnostic | Three old owners killed as a verified batch; 32.033-s absence wait recorded. Same provider/data restarted with memory raised from 2 to 4 GiB; full 10K/100 and critical-2 recovery passed, with independent 200-clone/four-mirror audits | | New complete load attempt | Original 108-window/114,960-arrival matrix launched after terminal recovery; mandatory full preflight running, no timed window completed at this checkpoint; every-ACK controller captured the actual campaign identity | +| Concurrent critical Git workflows | [Separate driver](../scripts/benchmark_critical_git.py) prepared with nine offline tests; retains the full 17-step suite and every attempted receipt, but no live critical-load/fault result is claimed | | Performance qualification | Full scheduled load, every post-load ACK recovery, higher admission profiles and matched comparisons remain open | The [native-filesystem trial](performance/2026-10-01-native-filesystem.md) diff --git a/docs/performance/2026-10-01-native-filesystem.md b/docs/performance/2026-10-01-native-filesystem.md index 7d0e9b9..6d4e7ba 100644 --- a/docs/performance/2026-10-01-native-filesystem.md +++ b/docs/performance/2026-10-01-native-filesystem.md @@ -312,6 +312,46 @@ passes. Provider/binary/admission/deadlines stay fixed; no automatic retries, restarts or data deletion are permitted. This is an armed gate, not a recovered ACK claim or completed load result. +### Prepare concurrent critical workflows as a separate gate + +The new [critical-workflow driver](../../scripts/benchmark_critical_git.py) +schedules the complete existing 17-step stock-Git suite through a validated +three-node/proxy fleet. Each workflow creates two unique disposable repositories +and covers atomic publication/refusal, mixed refusal, correct/stale force leases, +shallow history, v0/v2 filtered lazy fetch, incremental push/pull, deletion, +pruning, mirroring, invalid credentials and exact mirror/fsck checks. + +```sh +# After the current complete matrix and every-ACK recovery gates close: +python3 -B scripts/benchmark_critical_git.py \ + --fleet-dir "$CRITICAL_FLEET_DIR" --node-active-limit 100 \ + run --output-dir "$CRITICAL_OUTPUT_DIR" \ + --duration 300 --interval 15 --concurrency 4 +``` + +This example offers 20 whole workflows over 300 seconds with up to four in +flight. It is **prepared, not executed**; it adds no traffic to the live matrix. +It cannot replace the original 108 windows or higher-admission comparisons. + +| Driver evidence | Boundary | +| --- | --- | +| Complete arrival ledger | Every workflow or busy drop retained; no backpressure-induced clock slowdown or automatic retry | +| Workflow throughput/latency | Separate offered-window and drained throughput; latency includes failed attempts, dispatch delay and client/validation work | +| Critical step timings | All recorded successes and failures retained; grouped command wall times, not isolated RPC throughput or server-only latency | +| ACK evidence | Each attempted receipt is digest-bound; partial workflows and orphan receipts stop recovery for reconciliation rather than silently skipping writes | +| Fresh-owner verification | All complete workflows checked; a failure does not skip other complete workflows; refused/reused owners and failed load remain explicit | + +The driver's verification command sends no signals. The caller must separately +record actual old-process loss, expiry wait, fresh local state, unchanged provider +and complete original-corpus recovery. In-flight fault/partial-ACK reconciliation, +resource/cost boundaries and an actual concurrent critical run remain open. +Nine new offline scheduler and receipt/recovery tests passed, and all **84 +harness tests** passed on local Python 3.14 during mandatory preflight, not +measured load. These tests mock Git/provider operations and do not establish +any live result. The real scheduler test retains a simulated failed worker's +partial receipt and verifies client closure; recovery tests check refusal of +reused owners and continued verification after another workflow fails. + The closed observer reached a sampled maximum `memory.current` of 4,294,967,296 bytes. At its final sample (06:27:17 UTC), current memory was 4,108,148,736 bytes, anonymous memory 1,552,654,336 bytes and file memory 1,404,837,888 bytes. The @@ -473,6 +513,7 @@ Artifacts are retained under `canopy-native-filesystem-ceXFad8I`: | `cellule-a450-source-audit.json` | `87bac9bc1c54e8fba393cfec424436532086e0810f03f48b957c5d6cfb5d01a3` | | `harness-38a51dd-linux-36823311167-job.log` | `963a1636810e0efdc768e1d6f4c8297f3d3b07c7d441e99919ee8969d8ab0bcf` | | `rust-38a51dd-linux-36823311167-job.log` | `2384cf85dc36736bb8560902e71227ba823b3120cbb109856b2a040159f99842` | +| `harness-84-critical-driver.log` | `ca0bd3cb47be225672e608b5b7a60f036ea82ef15df2d3b58b36ecf8dcafc437` | The baseline's new independent ledger audit is `5416540c9b78b23e5c89ff24e771ab58012af847df4ec961adfc597bff010639` diff --git a/scripts/benchmark_critical_git.py b/scripts/benchmark_critical_git.py new file mode 100644 index 0000000..0b765a5 --- /dev/null +++ b/scripts/benchmark_critical_git.py @@ -0,0 +1,271 @@ +#!/usr/bin/env python3 +"""Schedule complete critical Git workflows; preserve every attempted receipt. + +Workflow throughput and step wall time include client work and validation. They +are not isolated RPC rates or server-only latency. Recovery requires every +attempted workflow to be complete; partial write ACKs cannot be silently skipped. +Owner loss, expiry, provider boundaries and full-corpus recovery are caller gates. +""" + +import argparse +from collections import Counter +from concurrent.futures import ThreadPoolExecutor +from contextlib import closing +from datetime import datetime, timezone +import json +import math +import os +from pathlib import Path +import threading +import time +from types import SimpleNamespace + +import benchmark_repositories as benchmark +import benchmark_three_node_campaign as campaign +import check_proxy_git as critical + + +STEPS = ["create-source", "create-mirror", "atomic-multiple-ref-publication", + "mixed-refusal", "atomic-refusal", "force-with-correct-lease", "stale-lease-refusal", + "shallow-deepen-unshallow", "filtered-lazy-fetch-v0", "filtered-lazy-fetch-v2", + "incremental-push", "fast-forward-pull", "delete-branch", "fetch-prune", + "mirror-push", "invalid-credential-refusal", "exact-v0-v2-mirror-fsck"] + + +def validate_receipt(receipt): + critical.require(receipt.get("complete") is True and receipt.get("error") is None, + "partial critical workflow; reconcile every write ACK before recovery") + steps = receipt.get("steps", []) + critical.require([step.get("name") for step in steps] == STEPS + and all(step.get("ok") is True for step in steps), "critical operation scope differs") + entries = receipt.get("repositories", []) + critical.require(len(entries) == 2 and len({entry["repository_id"] for entry in entries}) == 2 + and all(benchmark.canonical_repository_uuid(entry["repository_id"]) for entry in entries), + "critical repository identities differ") + + +def summarize(samples, duration, elapsed): + counts = Counter(sample["result"] for sample in samples) + attempted = [sample for sample in samples if sample["result"] != "driver_busy"] + in_window = sum(sample["result"] == "ok" and sample["completion_offset_seconds"] <= duration + for sample in attempted) + steps = {} + for sample in attempted: + for step in sample.get("steps", []): + critical.require(step.get("name") in STEPS and isinstance(step.get("ok"), bool) + and isinstance(step.get("wall_seconds"), (int, float)) + and not isinstance(step["wall_seconds"], bool) + and math.isfinite(step["wall_seconds"]) and step["wall_seconds"] >= 0, + "invalid critical step measurement") + times, outcomes = steps.setdefault(step["name"], ([], Counter())) + times.append(step["wall_seconds"] * 1000) + outcomes["ok" if step["ok"] else "failed"] += 1 + return {"outcomes": dict(counts), "failed_arrivals": len(samples) - counts["ok"], + "successful_workflows_in_schedule_window": in_window, + "successful_workflows_per_second_in_schedule_window": in_window / duration, + "successful_workflows_per_second_including_drain": counts["ok"] / elapsed, + "scheduled_workflow_latency_ms": benchmark.percentiles([sample["elapsed_ms"] for sample in attempted]), + "service_workflow_ms": benchmark.percentiles([sample["service_ms"] for sample in attempted]), + "step_wall_ms": {name: {"attempts": len(times), "outcomes": dict(outcomes), + "percentiles": benchmark.percentiles(times)} + for name, (times, outcomes) in steps.items()}, + "step_timing_scope": "all recorded step attempts including failures; grouped Git commands/client validation, not isolated RPC/server latency", + "latency_population": "all attempted workflows, including failures; busy arrivals have no fabricated latency"} + + +def bindings(directory, ready, binary): + paths = [Path(__file__), Path(critical.__file__), Path(benchmark.__file__), Path(campaign.__file__), + binary, directory / "ready.json", directory / "deployment.json", + *(directory / f"node-{index}.json" for index in range(3))] + return {str(path.resolve()): benchmark.file_sha256(path) for path in paths} + + +def run(args, token): + critical.require(args.duration % args.interval == 0 and 1 <= args.concurrency <= 16 + and args.duration // args.interval <= 1000, "declare bounded complete arrival clock") + ready = campaign.validate_fleet(args.fleet_dir, {"node_active_limit": args.node_active_limit}) + critical.require(not (args.fleet_dir / "outcome.json").exists(), "fleet is terminal") + binary = Path(ready["binary_path"]) + bound = bindings(args.fleet_dir, ready, binary) + args.output_dir.mkdir(parents=True, exist_ok=False) + total = args.duration // args.interval + record = {"version": 1, "completed": False, "error": None, "bindings": bound, + "fleet": ready, "scheduled": total, "interval_seconds": args.interval, + "schedule_seconds": args.duration, "concurrency": args.concurrency, + "http_timeout_seconds": args.timeout, "git_timeout_seconds": 120, + "steps_per_workflow": STEPS, "started_at_utc": datetime.now(timezone.utc).isoformat(), + "scope": "Complete concurrent critical workflows through three nodes/proxy. Not isolated operation RPS, " + "full-corpus recovery, owner-loss proof, provider cost or latest-source artifact qualification."} + report = args.output_dir / "report.json" + ledger = args.output_dir / "samples.jsonl" + benchmark.save(report, record) + slots = threading.BoundedSemaphore(args.concurrency) + lock = threading.Lock() + samples = [] + started = time.monotonic() + with ledger.open("x") as out: + def save(sample): + with lock: + samples.append(sample) + out.write(json.dumps(sample) + "\n") + out.flush() + + def execute(sequence, scheduled): + dispatched = time.monotonic() + receipt_path = args.output_dir / f"workflow-{sequence:04d}.json" + sample = {"sequence": sequence, "result": "workflow_error", "receipt": receipt_path.name} + try: + with closing(benchmark.Client(ready["proxy_url"], token, args.timeout)) as client: + receipt = critical.seed(SimpleNamespace(receipt=receipt_path, + work_dir=args.output_dir / f"workflow-{sequence:04d}", binary=binary, + base_url=ready["proxy_url"], timeout=args.timeout), client, token) + validate_receipt(receipt) + sample["result"] = "ok" + except Exception as error: + sample["error"] = type(error).__name__ + finally: + finished = time.monotonic() + if receipt_path.exists(): + sample["receipt_sha256"] = benchmark.file_sha256(receipt_path) + sample["steps"] = campaign.read_json(receipt_path).get("steps", []) + sample.update(completion_offset_seconds=finished - started, + elapsed_ms=(finished - scheduled) * 1000, + service_ms=(finished - dispatched) * 1000) + save(sample) + slots.release() + + try: + with ThreadPoolExecutor(max_workers=args.concurrency) as pool: + for sequence in range(total): + scheduled = started + sequence * args.interval + delay = scheduled - time.monotonic() + if delay > 0: + time.sleep(delay) + if slots.acquire(blocking=False): + pool.submit(execute, sequence, scheduled) + else: + save({"sequence": sequence, "result": "driver_busy"}) + remaining = started + args.duration - time.monotonic() + if remaining > 0: + time.sleep(remaining) + critical.require(len(samples) == total and {sample["sequence"] for sample in samples} == set(range(total)), + "arrival ledger is incomplete") + critical.require(all(benchmark.file_sha256(Path(path)) == digest for path, digest in bound.items()), + "bound workflow inputs changed") + record["completed"] = True + except BaseException as error: + record["error"] = type(error).__name__ + raise + finally: + elapsed = time.monotonic() - started + out.flush() + record.update(summarize(samples, args.duration, elapsed), + elapsed_including_drain_seconds=elapsed, + samples_sha256=benchmark.file_sha256(ledger)) + benchmark.save(report, record) + return record + + +def load_receipts(report_path): + record = campaign.read_json(report_path) + critical.require(record.get("version") == 1 and record.get("completed") is True + and record.get("error") is None and record.get("steps_per_workflow") == STEPS, + "require a closed complete critical schedule") + critical.require(all(benchmark.file_sha256(Path(path)) == digest for path, digest in record["bindings"].items()), + "critical source/fleet/binary inputs changed") + directory = report_path.parent + ledger = directory / "samples.jsonl" + critical.require(benchmark.file_sha256(ledger) == record["samples_sha256"], "critical ledger changed") + samples = [json.loads(line) for line in ledger.read_text().splitlines()] + critical.require(len(samples) == record["scheduled"] + and {sample["sequence"] for sample in samples} == set(range(record["scheduled"])), + "critical arrival inventory differs") + totals = summarize(samples, record["schedule_seconds"], record["elapsed_including_drain_seconds"]) + critical.require(all(record.get(key) == value for key, value in totals.items()), + "critical report differs from arrival ledger") + receipts, identities, expected_paths = [], set(), set() + for sample in samples: + if sample["result"] == "driver_busy": + critical.require("receipt" not in sample, "busy arrival cannot discard a receipt") + continue + name = f"workflow-{sample['sequence']:04d}.json" + critical.require(sample.get("receipt") == name and sample["result"] in ("ok", "workflow_error"), + "unexpected receipt path/outcome") + path = directory / name + expected_paths.add(name) + critical.require(benchmark.file_sha256(path) == sample["receipt_sha256"], "critical receipt changed") + receipt = campaign.read_json(path) + critical.require(sample.get("steps") == receipt.get("steps"), "step ledger differs from receipt") + validate_receipt(receipt) + critical.require(receipt["binary_sha256"] == record["fleet"]["binary_sha256"] + and receipt["driver_sha256"] == record["bindings"][str(Path(critical.__file__).resolve())] + and receipt["git_driver_sha256"] == record["bindings"][str(Path(benchmark.__file__).resolve())], + "receipt source/binary differs") + for entry in receipt["repositories"]: + critical.require(entry["repository_id"] not in identities, "duplicate critical workflow identity") + identities.add(entry["repository_id"]) + receipts.append((path, receipt)) + critical.require({path.name for path in directory.glob("workflow-*.json")} == expected_paths, + "orphan critical receipt; reconcile before recovery") + return record, receipts + + +def verify(args, token): + record, receipts = load_receipts(args.report) + critical.require(not args.output.exists(), "requires new recovery output") + ready = campaign.validate_fleet(args.fleet_dir, {"node_active_limit": args.node_active_limit}) + old = record["fleet"]["nodes"] + critical.require(not {node["node_id"] for node in old}.intersection(node["node_id"] for node in ready["nodes"]) + and not {node["pid"] for node in old}.intersection(node["pid"] for node in ready["nodes"]) + and ready["binary_sha256"] == record["fleet"]["binary_sha256"], + "requires distinct owners with the original binary") + args.work_dir.mkdir(parents=True, exist_ok=False) + checks = [] + with closing(benchmark.Client(ready["proxy_url"], token, args.timeout)) as client: + for index, (path, receipt) in enumerate(receipts): + try: + value = {"verified": True, **critical.verify_receipt(receipt, ready["proxy_url"], + args.work_dir / f"workflow-{index:04d}", client, token)} + except Exception as error: + value = {"verified": False, "error": type(error).__name__} + checks.append({"receipt": path.name, "receipt_sha256": benchmark.file_sha256(path), **value}) + load_receipts(args.report) + result = {"version": 1, "completed": True, "report_sha256": benchmark.file_sha256(args.report), + "attempted_workflow_verifications": len(checks), + "all_workflows_verified": bool(checks) and all(check["verified"] for check in checks), + "verified_workflows": sum(check["verified"] for check in checks), + "verified_repositories": sum(check["verified"] for check in checks) * 2, + "checks": checks, "failed_load_arrivals": record["failed_arrivals"], + "owner_recovery": "not established here: caller must bind actual old-process absence, expiry wait, " + "fresh local state, unchanged provider and full original corpus"} + benchmark.save(args.output, result) + return result + + +def main(): + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("--fleet-dir", required=True, type=Path) + parser.add_argument("--node-active-limit", required=True, type=benchmark.positive) + parser.add_argument("--timeout", type=benchmark.positive, default=30) + actions = parser.add_subparsers(dest="action", required=True) + create = actions.add_parser("run") + create.add_argument("--output-dir", type=Path, required=True) + create.add_argument("--duration", type=benchmark.positive, default=300) + create.add_argument("--interval", type=benchmark.positive, default=15) + create.add_argument("--concurrency", type=benchmark.positive, default=4) + recover = actions.add_parser("verify") + recover.add_argument("--report", type=Path, required=True) + recover.add_argument("--work-dir", type=Path, required=True) + recover.add_argument("--output", type=Path, required=True) + args = parser.parse_args() + token = os.environ.get("CANOPY_GIT_TOKEN") + if not token: + parser.error("CANOPY_GIT_TOKEN is required") + result = run(args, token) if args.action == "run" else verify(args, token) + print(json.dumps(result, indent=2)) + if result.get("failed_arrivals", result.get("failed_load_arrivals", 0)) or result.get("all_workflows_verified") is False: + raise SystemExit(1) + + +if __name__ == "__main__": + main() diff --git a/scripts/test_benchmark_critical_git.py b/scripts/test_benchmark_critical_git.py new file mode 100644 index 0000000..3dc7c06 --- /dev/null +++ b/scripts/test_benchmark_critical_git.py @@ -0,0 +1,187 @@ +"""Offline scheduler/recovery guards; no real Git/provider/owner loss proof.""" +import json +from pathlib import Path +import tempfile +from types import SimpleNamespace +import unittest +from unittest.mock import patch +import uuid + +import benchmark_repositories as benchmark +import benchmark_critical_git as driver + + +def receipt(): + return {"version": 1, "complete": True, "error": None, + "steps": [{"name": name, "ok": True, "wall_seconds": 0.1} for name in driver.STEPS], + "repositories": [{"repository_id": str(uuid.uuid4()), "name": name, "owner": "canopy"} + for name in ("source", "mirror")], "binary_sha256": "a" * 64, + "driver_sha256": "b" * 64, "git_driver_sha256": "c" * 64} + + +class CriticalLoadTests(unittest.TestCase): + def test_partial_or_shortened_workflow_is_never_skipped(self): + for mutate in (lambda r: r.update(complete=False), lambda r: r.update(error="TimeoutError"), + lambda r: r["steps"].pop(), lambda r: r["steps"][3].update(ok=False), + lambda r: r["repositories"][1].update(repository_id=r["repositories"][0]["repository_id"])): + value = receipt() + mutate(value) + with self.assertRaises(RuntimeError): + driver.validate_receipt(value) + + def test_failed_attempt_latency_is_retained_busy_latency_is_not_fabricated(self): + samples = [{"result": "ok", "elapsed_ms": 10, "service_ms": 8, "completion_offset_seconds": 0.2}, + {"result": "workflow_error", "elapsed_ms": 100, "service_ms": 80, "completion_offset_seconds": 1.5}, + {"result": "driver_busy"}, + {"result": "ok", "elapsed_ms": 200, "service_ms": 160, "completion_offset_seconds": 2.5}] + value = driver.summarize(samples, 2, 4) + self.assertEqual(value["failed_arrivals"], 2) + self.assertEqual(value["successful_workflows_per_second_in_schedule_window"], 0.5) + self.assertEqual(value["successful_workflows_per_second_including_drain"], 0.5) + self.assertEqual(value["scheduled_workflow_latency_ms"], benchmark.percentiles([10, 100, 200])) + + def fixture(self, root): + sample_path = root / "workflow-0000.json" + value = receipt() + benchmark.save(sample_path, value) + sample = {"sequence": 0, "result": "ok", "receipt": sample_path.name, + "receipt_sha256": benchmark.file_sha256(sample_path), "elapsed_ms": 100, + "service_ms": 90, "completion_offset_seconds": 0.1, "steps": value["steps"]} + ledger = root / "samples.jsonl" + ledger.write_text(json.dumps(sample) + "\n") + report = {"version": 1, "completed": True, "error": None, "scheduled": 1, + "schedule_seconds": 1, "elapsed_including_drain_seconds": 1, + "steps_per_workflow": driver.STEPS, "samples_sha256": benchmark.file_sha256(ledger), + "bindings": {str(Path(driver.critical.__file__).resolve()): "b" * 64, + str(Path(benchmark.__file__).resolve()): "c" * 64}, + "fleet": {"binary_sha256": "a" * 64}} + report.update(driver.summarize([sample], 1, 1)) + path = root / "report.json" + benchmark.save(path, report) + return path, report, sample, value + + def test_closed_receipt_inventory_and_report_are_audited(self): + with tempfile.TemporaryDirectory() as temp: + root = Path(temp) + path, record, sample, value = self.fixture(root) + hashes = {root / sample["receipt"]: sample["receipt_sha256"], + root / "samples.jsonl": record["samples_sha256"], + **{Path(p): d for p, d in record["bindings"].items()}} + with patch.object(benchmark, "file_sha256", side_effect=lambda p: hashes[Path(p)]): + self.assertEqual(len(driver.load_receipts(path)[1]), 1) + benchmark.save(root / "workflow-9999.json", value) + with self.assertRaisesRegex(RuntimeError, "orphan"): + driver.load_receipts(path) + + def test_digest_bound_partial_receipt_and_forged_report_are_rejected(self): + with tempfile.TemporaryDirectory() as temp: + root = Path(temp) + path, record, sample, value = self.fixture(root) + hashes = {root / sample["receipt"]: sample["receipt_sha256"], + root / "samples.jsonl": record["samples_sha256"], + **{Path(p): d for p, d in record["bindings"].items()}} + with patch.object(benchmark, "file_sha256", side_effect=lambda p: hashes[Path(p)]): + record["failed_arrivals"] = 1 + benchmark.save(path, record) + with self.assertRaisesRegex(RuntimeError, "differs from arrival"): + driver.load_receipts(path) + record["failed_arrivals"] = 0 + benchmark.save(path, record) + value["complete"] = False + benchmark.save(root / sample["receipt"], value) + with self.assertRaisesRegex(RuntimeError, "partial"): + driver.load_receipts(path) + + def test_step_latency_includes_refusals_and_failed_work(self): + attempts = [{"result": result, "completion_offset_seconds": 0.2, + "elapsed_ms": elapsed, "service_ms": elapsed, + "steps": [{"name": "atomic-refusal", "ok": ok, "wall_seconds": elapsed / 1000}]} + for result, ok, elapsed in (("ok", True, 10), ("workflow_error", False, 90))] + value = driver.summarize(attempts, 1, 1)["step_wall_ms"]["atomic-refusal"] + self.assertEqual(value["outcomes"], {"ok": 1, "failed": 1}) + self.assertEqual(value["percentiles"], benchmark.percentiles([10, 90])) + + def test_ledger_cannot_escape_paths_or_silently_drop_arrivals(self): + for change in (lambda sample: sample.update(receipt="../escape.json"), + lambda sample: sample.update(sequence=3), + lambda sample: sample.update(result="ignored")): + with self.subTest(change=change), tempfile.TemporaryDirectory() as temp: + root = Path(temp) + path, record, sample, value = self.fixture(root) + change(sample) + (root / "samples.jsonl").write_text(json.dumps(sample) + "\n") + record["samples_sha256"] = benchmark.file_sha256(root / "samples.jsonl") + record.update(driver.summarize([sample], 1, 1)) + benchmark.save(path, record) + hashes = {root / sample["receipt"]: sample["receipt_sha256"], + root / "samples.jsonl": record["samples_sha256"], + **{Path(p): d for p, d in record["bindings"].items()}} + with patch.object(benchmark, "file_sha256", side_effect=lambda p: hashes[Path(p)]): + with self.assertRaises(RuntimeError): + driver.load_receipts(path) + + def test_reused_owner_is_refused_before_network_git(self): + with tempfile.TemporaryDirectory() as temp: + root = Path(temp) + nodes = [{"node_id": f"old-{i}", "pid": i + 1} for i in range(3)] + record = {"fleet": {"nodes": nodes, "binary_sha256": "a" * 64}, "failed_arrivals": 0} + args = SimpleNamespace(report=root / "report.json", output=root / "result.json", + work_dir=root / "clones", fleet_dir=root / "fleet", node_active_limit=100, timeout=30) + with patch.object(driver, "load_receipts", return_value=(record, [])), \ + patch.object(driver.campaign, "validate_fleet", return_value=record["fleet"]), \ + patch.object(driver.benchmark, "Client", side_effect=AssertionError("unexpected network")): + with self.assertRaisesRegex(RuntimeError, "distinct owners"): + driver.verify(args, "fixture-token") + self.assertFalse(args.work_dir.exists()) + + def test_real_scheduler_retains_failed_worker_and_closes_client(self): + with tempfile.TemporaryDirectory() as temp: + root = Path(temp) + fleet = root / "fleet" + fleet.mkdir() + ready = {"binary_path": str(root / "binary"), "proxy_url": "http://127.0.0.1:1"} + args = SimpleNamespace(duration=1, interval=1, concurrency=1, fleet_dir=fleet, + node_active_limit=100, output_dir=root / "output", timeout=30) + closed = [] + client = SimpleNamespace(close=lambda: closed.append(True)) + def failed(options, client, token): + benchmark.save(options.receipt, {"complete": False, "repositories": [{"name": "partial-ack"}]}) + raise RuntimeError("simulated write failure") + with patch.object(driver.campaign, "validate_fleet", return_value=ready), \ + patch.object(driver, "bindings", return_value={}), \ + patch.object(driver.benchmark, "Client", return_value=client), \ + patch.object(driver.critical, "seed", side_effect=failed): + result = driver.run(args, "fixture-token") + self.assertTrue(result["completed"]) + self.assertEqual(result["failed_arrivals"], 1) + self.assertEqual(result["outcomes"], {"workflow_error": 1}) + self.assertEqual(closed, [True]) + sample = json.loads((args.output_dir / "samples.jsonl").read_text()) + self.assertEqual(sample["receipt_sha256"], benchmark.file_sha256(args.output_dir / sample["receipt"])) + + def test_recovery_continues_other_complete_workflows_after_one_failure(self): + with tempfile.TemporaryDirectory() as temp: + root = Path(temp) + old = [{"node_id": f"old-{i}", "pid": i + 1} for i in range(3)] + fresh = [{"node_id": f"new-{i}", "pid": i + 10} for i in range(3)] + record = {"fleet": {"nodes": old, "binary_sha256": "a" * 64}, "failed_arrivals": 0} + entries = [(root / "one.json", receipt()), (root / "two.json", receipt())] + ready = {"nodes": fresh, "binary_sha256": "a" * 64, "proxy_url": "http://127.0.0.1:1"} + args = SimpleNamespace(report=root / "report.json", output=root / "result.json", + work_dir=root / "clones", fleet_dir=root / "fleet", node_active_limit=100, timeout=30) + client = SimpleNamespace(close=lambda: None) + with patch.object(driver, "load_receipts", return_value=(record, entries)), \ + patch.object(driver.campaign, "validate_fleet", return_value=ready), \ + patch.object(driver.benchmark, "Client", return_value=client), \ + patch.object(driver.benchmark, "file_sha256", return_value="f" * 64), \ + patch.object(driver.critical, "verify_receipt", side_effect=[RuntimeError("one failed"), + {"verified_repositories": 2, "protocols": [0, 2], "exact_ref_inventories": 4}]) as verify: + result = driver.verify(args, "fixture-token") + self.assertEqual(verify.call_count, 2) + self.assertFalse(result["all_workflows_verified"]) + self.assertEqual(result["verified_workflows"], 1) + self.assertEqual(result["attempted_workflow_verifications"], 2) + + +if __name__ == "__main__": + unittest.main() From 5b2f95c941b89c5ef7c27ce4adbe861903a5982c Mon Sep 17 00:00:00 2001 From: forhappy Date: Thu, 1 Oct 2026 10:20:00 -0700 Subject: [PATCH 23/60] build: sync Cellule routing changes and record failed load status --- Cargo.lock | 12 +- crates/canopy-server/Cargo.toml | 10 +- docs/performance-plan.md | 23 +-- .../2026-10-01-evidence-loss-and-c51dd121.md | 138 ++++++++++++++++++ .../2026-10-01-native-filesystem.md | 101 ++++++++++++- 5 files changed, 258 insertions(+), 26 deletions(-) create mode 100644 docs/performance/2026-10-01-evidence-loss-and-c51dd121.md diff --git a/Cargo.lock b/Cargo.lock index 5e81865..40b95af 100644 --- a/Cargo.lock +++ b/Cargo.lock @@ -448,7 +448,7 @@ dependencies = [ [[package]] name = "cellule-app" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=a4500add51764fa0415791aefbfa561db6ada203#a4500add51764fa0415791aefbfa561db6ada203" +source = "git+https://github.com/crabbuild/cellule.git?rev=c51dd121284ecc8878b75d32717a4dfbe2c406c2#c51dd121284ecc8878b75d32717a4dfbe2c406c2" dependencies = [ "blake3", "cellule-runtime", @@ -457,7 +457,7 @@ dependencies = [ [[package]] name = "cellule-host" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=a4500add51764fa0415791aefbfa561db6ada203#a4500add51764fa0415791aefbfa561db6ada203" +source = "git+https://github.com/crabbuild/cellule.git?rev=c51dd121284ecc8878b75d32717a4dfbe2c406c2#c51dd121284ecc8878b75d32717a4dfbe2c406c2" dependencies = [ "cellule-app", "cellule-runtime", @@ -471,7 +471,7 @@ dependencies = [ [[package]] name = "cellule-ltx" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=a4500add51764fa0415791aefbfa561db6ada203#a4500add51764fa0415791aefbfa561db6ada203" +source = "git+https://github.com/crabbuild/cellule.git?rev=c51dd121284ecc8878b75d32717a4dfbe2c406c2#c51dd121284ecc8878b75d32717a4dfbe2c406c2" dependencies = [ "async-trait", "blake3", @@ -493,7 +493,7 @@ dependencies = [ [[package]] name = "cellule-runtime" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=a4500add51764fa0415791aefbfa561db6ada203#a4500add51764fa0415791aefbfa561db6ada203" +source = "git+https://github.com/crabbuild/cellule.git?rev=c51dd121284ecc8878b75d32717a4dfbe2c406c2#c51dd121284ecc8878b75d32717a4dfbe2c406c2" dependencies = [ "blake3", "bytes", @@ -519,7 +519,7 @@ dependencies = [ [[package]] name = "cellule-store" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=a4500add51764fa0415791aefbfa561db6ada203#a4500add51764fa0415791aefbfa561db6ada203" +source = "git+https://github.com/crabbuild/cellule.git?rev=c51dd121284ecc8878b75d32717a4dfbe2c406c2#c51dd121284ecc8878b75d32717a4dfbe2c406c2" dependencies = [ "async-trait", "blake3", @@ -542,7 +542,7 @@ dependencies = [ [[package]] name = "cellule-types" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=a4500add51764fa0415791aefbfa561db6ada203#a4500add51764fa0415791aefbfa561db6ada203" +source = "git+https://github.com/crabbuild/cellule.git?rev=c51dd121284ecc8878b75d32717a4dfbe2c406c2#c51dd121284ecc8878b75d32717a4dfbe2c406c2" dependencies = [ "schemars", "serde", diff --git a/crates/canopy-server/Cargo.toml b/crates/canopy-server/Cargo.toml index 744ee1b..439cbbf 100644 --- a/crates/canopy-server/Cargo.toml +++ b/crates/canopy-server/Cargo.toml @@ -15,11 +15,11 @@ axum = "0.8.9" base64 = "0.22" blake3 = "1.8" bytes = "1.11" -cellule-app = { git = "https://github.com/crabbuild/cellule.git", rev = "a4500add51764fa0415791aefbfa561db6ada203" } -cellule-host = { git = "https://github.com/crabbuild/cellule.git", rev = "a4500add51764fa0415791aefbfa561db6ada203" } -cellule-ltx = { git = "https://github.com/crabbuild/cellule.git", rev = "a4500add51764fa0415791aefbfa561db6ada203", features = ["replica"] } -cellule-runtime = { git = "https://github.com/crabbuild/cellule.git", rev = "a4500add51764fa0415791aefbfa561db6ada203" } -cellule-store = { git = "https://github.com/crabbuild/cellule.git", rev = "a4500add51764fa0415791aefbfa561db6ada203" } +cellule-app = { git = "https://github.com/crabbuild/cellule.git", rev = "c51dd121284ecc8878b75d32717a4dfbe2c406c2" } +cellule-host = { git = "https://github.com/crabbuild/cellule.git", rev = "c51dd121284ecc8878b75d32717a4dfbe2c406c2" } +cellule-ltx = { git = "https://github.com/crabbuild/cellule.git", rev = "c51dd121284ecc8878b75d32717a4dfbe2c406c2", features = ["replica"] } +cellule-runtime = { git = "https://github.com/crabbuild/cellule.git", rev = "c51dd121284ecc8878b75d32717a4dfbe2c406c2" } +cellule-store = { git = "https://github.com/crabbuild/cellule.git", rev = "c51dd121284ecc8878b75d32717a4dfbe2c406c2" } ed25519-dalek = "2" flate2 = "1.1" futures-core = "0.3" diff --git a/docs/performance-plan.md b/docs/performance-plan.md index 07e704e..0236fe6 100644 --- a/docs/performance-plan.md +++ b/docs/performance-plan.md @@ -2,12 +2,12 @@ Use this plan to design capacity work and interpret Canopy benchmark results. It separates demonstrated behavior from proposed targets. The current Cellule -dependency pin is `a4500add51764fa0415791aefbfa561db6ada203`, an upstream -web/documentation advance with the same Rust crate tree and root Cargo objects -as `0573f489` and `e07670e`. Locked metadata resolves all six packages to the -new commit without unrelated dependency changes. Both hosted Rust and harness -jobs passed for source head `38a51dd` with this pin; the latest-pin local -production artifact and three-node qualification remain separate open gates. +dependency pin is `c51dd121284ecc8878b75d32717a4dfbe2c406c2`, the latest +observed upstream main on October 1. Unlike the earlier `a4500add` documentation +advance, this changes routing and resource-permit implementation. Locked +metadata resolves all six packages; the lockfile changes only their sources. +The new release build and its RustFS/three-node qualification are separate +gates. The 84 Python harness tests passed; this is not Rust runtime verification. None establishes reference capacity. The [three-node proxy campaign](performance/2026-09-30-three-node-proxy.md) retains the earlier fixture race and its regression, plus the candidate's failed @@ -21,12 +21,17 @@ seed and successful recovery of every recorded ACK. | Native-volume comparison | Same runtime/provider limits; all 10,000 identities and 100 Git/LFS fixtures seeded in 3,179.240 s and passed initial three-owner-loss recovery; independent 200-clone audit passed | | Native scheduled load | First attempt failed before any window. Independent-session V2 passed full preflight, then RustFS was cgroup-OOM killed: six closed windows, one interrupted prefix and 101 unstarted windows retained | | Post-OOM diagnostic | Three old owners killed as a verified batch; 32.033-s absence wait recorded. Same provider/data restarted with memory raised from 2 to 4 GiB; full 10K/100 and critical-2 recovery passed, with independent 200-clone/four-mirror audits | -| New complete load attempt | Original 108-window/114,960-arrival matrix launched after terminal recovery; mandatory full preflight running, no timed window completed at this checkpoint; every-ACK controller captured the actual campaign identity | -| Concurrent critical Git workflows | [Separate driver](../scripts/benchmark_critical_git.py) prepared with nine offline tests; retains the full 17-step suite and every attempted receipt, but no live critical-load/fault result is claimed | -| Performance qualification | Full scheduled load, every post-load ACK recovery, higher admission profiles and matched comparisons remain open | +| Complete 4-GiB load attempt | All 108 windows finished: 60,197 OK and 54,763 failed arrivals. The original attempt remains failed; local raw evidence is now unavailable | +| Post-load ACK recovery | Recorded three-owner loss and 32.249-s wait; critical fixtures passed. Fresh fleet then recorded ENOSPC and shut down; full corpus and every-ACK verification are incomplete | +| Concurrent critical Git workflows | [Separate driver](../scripts/benchmark_critical_git.py) retains the 17-step suite. The versioned gate separates failed performance from mandatory correctness recovery; no live critical-load/fault result is established | +| Evidence availability | Benchmark directories and retained executable disappeared during the later space check. No receipt backup found in checked locations; surviving RustFS data is not an ACK ledger | +| Performance qualification | Latest artifact/runtime, every post-load ACK, critical-operation load/fault coverage, higher admission profiles and matched comparisons remain open | The [native-filesystem trial](performance/2026-10-01-native-filesystem.md) records the storage hypothesis, graceful diagnostic drain and new bindings. +The [evidence-loss checkpoint](performance/2026-10-01-evidence-loss-and-c51dd121.md) +supersedes its earlier running/retained-artifact status. Historical observations +are not currently replayable; a new run cannot certify the old ACKs. No Cellule bottleneck or matched performance improvement is proven. The 4-GiB diagnostic is a new resource envelope, not a passing 2-GiB result. Offline plans for 500/1,000-entry admission retain the full matrix but have not been executed. diff --git a/docs/performance/2026-10-01-evidence-loss-and-c51dd121.md b/docs/performance/2026-10-01-evidence-loss-and-c51dd121.md new file mode 100644 index 0000000..6e7a93f --- /dev/null +++ b/docs/performance/2026-10-01-evidence-loss-and-c51dd121.md @@ -0,0 +1,138 @@ +# Preserve failures and rebuild the Cellule candidate + +**Status: performance failed; post-load correctness recovery incomplete; old raw +evidence unavailable. Latest-Cellule candidate verification is in progress.** + +This checkpoint supersedes earlier running/retained-artifact statements in the +[native-filesystem trial](2026-10-01-native-filesystem.md). It keeps the original +10,000-repository target, all 108 load windows and separate correctness gates. + +## What was observed before artifact loss + +The three-node/proxy campaign finished its original 114,960 offered arrivals. +The closed controller reported exit 1 and zero diagnostic-probe errors. The +independent auditor reported all 108 ledgers and resource boundaries reconciled. +Neither result was a performance pass. + +| Outcome | Arrivals | +| --- | ---: | +| OK | 60,197 | +| Driver busy | 19,137 | +| HTTP 503 | 35,387 | +| Transport error | 15 | +| Git error | 224 | +| Total | 114,960 | + +The first 18 metadata windows offered 43,200 arrivals: 30,521 OK, 12,408 busy and +271 HTTP 503. A separate read-only audit replayed those windows and the first +creation window before the files disappeared. Each metadata group below contains +three 120-second repetitions at 20 offered requests/s and 32 clients. Ranges are +the minimum/maximum individual repetition results, not pooled percentiles: + +| Active set / selection | OK / busy / HTTP 503 | Successful in-window requests/s | All-attempt p99 range (ms) | +| --- | --- | --- | --- | +| 100 / uniform | 6,466 / 678 / 56 | 14.833–19.683 | 1,681.714–9,139.462 | +| 100 / skewed | 6,678 / 476 / 46 | 17.833–18.908 | 3,821.461–4,428.670 | +| 500 / uniform | 2,769 / 4,379 / 52 | 4.808–10.833 | 8,644.185–19,291.809 | +| 500 / skewed | 6,323 / 834 / 43 | 15.425–19.375 | 3,842.231–11,061.363 | +| 1,000 / uniform | 2,458 / 4,694 / 48 | 3.008–10.458 | 9,178.527–25,248.536 | +| 1,000 / skewed | 5,827 / 1,347 / 26 | 13.575–17.533 | 8,666.377–10,186.262 | + +These are active sets inside the full 10,000-identity corpus, with unchanged +100-entry per-node admission; they are not smaller replacement corpora or proof +that every selected repository was resident. The first creation window offered +1 request/s for 120 seconds with 16 clients: + +| Creation measurement | Observed value | +| --- | --- | +| OK / busy / HTTP 503 / transport errors | 60 / 19 / 32 / 9 | +| Successful completions within the schedule | 0.500/s | +| All-attempt p50 / p95 / p99 | 16,407.401 / 30,012.684 / 30,020.859 ms | +| Dispatch-delay p99 | 19.254 ms | + +Busy drops have no invented latency. These failed-window observations are not +success-only latency, sustained capacity or a matched improvement. The raw +ledgers needed to replay these measurements are no longer locally available. + +## Correctness recovery did not finish + +The complete ledger recorded 1,547 creation ACKs, 1,184 ref-only push ACKs, +957 fresh-object push ACKs and 942 LFS-upload ACKs. Recovery must cover every +one of these, plus the full original corpus and critical fixtures. + +The post-load controller recorded three original owners exiting -9 without +forced fallback and an actual 32.249-second absence wait. Fresh owners started +at 10:17:26 UTC. Critical recovery passed both repositories and four exact +Git-v0/v2 ref inventories. The next full-corpus stage did not close successfully. +The fresh fleet's outcome recorded `No space left on device` while writing proxy +metrics, followed by three graceful exits. Its incomplete controller receipt is +not proof of any remaining recovery scope. + +```mermaid +flowchart LR + load[108 windows complete
performance failed] --> audit[Closed ledgers reconciled] + audit --> loss[Three-owner loss
32.249-second wait] + loss --> critical[Critical recovery passed] + critical --> disk[ENOSPC
fresh fleet shut down] + disk --> open[Full corpus and every ACK
remain unverified] +``` + +## Evidence availability changed during inspection + +At 16:54 UTC the old receipts and ENOSPC outcome could still be read. During +the subsequent free-space check, the external Canopy benchmark directories and +retained executable disappeared. The scoped Cargo dry run had identified +3,552 files/1.1 GiB; the later scoped clean reported **zero files removed**. +It does not explain the disappearance of the other directories. + +No receipt copy was found in the checked Trash, preview, checkout or temporary +locations. The cause and recoverability of the removal are not established. +The native RustFS container/data volume was still running with its original +start time and zero restarts. Stored state alone cannot reconstruct which +individual requests received ACKs or certify the old recovery. + +Historical digest tables record expected values, not available raw files. +Do not regenerate files under their old names, invent missing observations, +mark old recovery passing, or combine a new run with the old campaign. + +## Latest Cellule candidate + +The dependency now pins observed upstream main +`c51dd121284ecc8878b75d32717a4dfbe2c406c2`. Five direct declarations and six +lockfile sources changed; unrelated dependencies did not. Locked metadata +resolves all six Cellule packages to that revision. + +Unlike the earlier documentation-only `a4500add` advance, this includes upstream +routing and host-permit changes. It reuses resident catalog identity for unleased +requests while still observing fresh authority, and pairs resource charges with +semaphore permits. Source inspection is not a Canopy latency or correctness result. + +All 84 Python harness tests passed in 44.556 seconds. They exercise benchmark +accounting and guards, not the upgraded Rust implementation. A new locked release +build is running; Rust tests, RustFS end-to-end checks and full three-node/load/fault +verification remain open. New evidence belongs outside Cargo target directories. + +New closed files are under +`/Users/haipingfu/.codex/canopy-three-node-evidence-BwYz7P`, separately from the +rebuildable Cargo target: + +| Closed new artifact | SHA-256 | +| --- | --- | +| `harness.log` | `5696368dfdbb6537716b8e35dddc719726c848369dc47df293e4e092319f0176` | +| `metadata.json` | `a4a33f617545d8b7796719b791c9da241b94eb5e89db85fb3f0bb61765418a69` | +| `build_candidate.py` | `1e458a80735bdb8c83edf58755ee3933d6a403704b6dc0e8e8bbec89428f1b62` | + +`build.json` and `build.log` remain changing outputs until the actual compiler +process exits; they are not listed as closed evidence here. + +| Required gate | Current status | +| --- | --- | +| Exact new artifact and source provenance | Build in progress | +| Rust workspace correctness and RustFS Git/LFS compatibility | Open | +| Full 10,000 identities/100 populated fixtures | Open for the new artifact | +| Original matrix, critical Git load and every new ACK after owner loss | Open | +| Higher admission profiles and matched comparisons | Open | +| Old campaign every-ACK recovery | Unverified; original raw inputs unavailable | + +The [performance plan](../performance-plan.md) remains the scope. No proven +Cellule bottleneck, matched speedup or isolated Linux reference capacity is claimed. diff --git a/docs/performance/2026-10-01-native-filesystem.md b/docs/performance/2026-10-01-native-filesystem.md index 6d4e7ba..6fae14c 100644 --- a/docs/performance/2026-10-01-native-filesystem.md +++ b/docs/performance/2026-10-01-native-filesystem.md @@ -1,9 +1,17 @@ # Compare RustFS filesystem backing -**Status: full seed, initial recovery and post-OOM 4-GiB recovery passed; V2 load failed; new complete load attempt in mandatory preflight; performance unqualified.** +**Status: all 108 4-GiB load windows finished and failed; post-load recovery was interrupted; local raw artifacts are unavailable; performance unqualified.** This shared-Mac/Colima diagnostic keeps the 10,000-repository target. It tests filesystem backing without replacing the bound runtime or weakening correctness. +The [later evidence-loss checkpoint](2026-10-01-evidence-loss-and-c51dd121.md) +supersedes running/retained-artifact statements below. At the final inspection, +the campaign retained 60,197 OK and 54,763 failed arrivals. Fresh post-load owners +passed the critical fixtures, then their launcher recorded ENOSPC and graceful +shutdown; full-corpus/every-ACK recovery did not complete. The external benchmark +directories and executable subsequently disappeared. Historical digests below +are expected values, not proof that their local files remain available. + ## Preserve the failed run first | Gate | Observed result | @@ -278,9 +286,11 @@ A separately bound controller passed its live check and two offline tests (23 rejection cases). It observed the exact verifier's actual exit, reread the terminal full 10K/100 and critical-2 receipts, original deadlines and closed sidecar digests, then launched a **new entire 108-window attempt** in a new -directory. The real campaign PID is 5740; its mandatory full preflight is -running, with no timed window completed at this checkpoint. It sends no signals -and changes no provider settings. +directory. The real campaign PID is 5740; its mandatory preflight passed all +10,000 identities/100 Git-v0/v2/LFS fixtures, and timed load began at +**07:03:58 UTC**. The first three closed metadata windows failed; all 108 windows +subsequently completed without retrying or dropping those failures. The load +controller sent no signals and changed no provider settings. An incomplete recovery, PID reuse, changed provider/source, reused output or shortened plan prevents launch. Launch/preflight is not a load pass: all 114,960 declared arrivals over 8,640 offered seconds remain required. @@ -330,9 +340,32 @@ python3 -B scripts/benchmark_critical_git.py \ ``` This example offers 20 whole workflows over 300 seconds with up to four in -flight. It is **prepared, not executed**; it adds no traffic to the live matrix. +flight. It was **armed, not executed** at the earlier checkpoint; it added no +traffic to the matrix. Those waiter processes are no longer present. It cannot replace the original 108 windows or higher-admission comparisons. +A separate controller binds the live post-load verifier PID/command/kernel +identity and only reads that process/receipt every 15 seconds while prior load +runs. Before launching, it requires actual verifier/controller/campaign/event +stream absence, the entire passing 108-window matrix, all four recovery scopes, +closed sidecar digests, exact ACK counts, recorded owner loss/expiry and unchanged +provider/source/fresh-owner bindings. Three offline tests passed, including 28 +rejection cases, and its live read-only binding check passed. These are terminal +gate tests, not a mocked full workflow or actual critical-load result. + +The first timed failures prevented this attempt from satisfying that original +launch gate, even if later windows succeeded. A separately versioned gate later +passed seven offline tests/49 rejection cases and a live binding check. It allowed +failed performance to remain failed while still requiring the entire schedule, +every recovery scope, exact ACK counts and unchanged provider/owner boundaries. +Neither gate produced a live critical-load result: the post-load recovery did +not finish, and its evidence was subsequently lost. The critical live run remains open. +On an eligible future completion, the controller establishes receipt of a scoped +Docker event with a read-only `/proc/uptime` marker before traffic, runs the new +workflow process in an independent session, retains closed provider events and +rechecks the provider/fleet/source boundary. It sends no owner signals or provider +restarts. New critical ACKs still require a separate recorded fault and recovery. + | Driver evidence | Boundary | | --- | --- | | Complete arrival ledger | Every workflow or busy drop retained; no backpressure-induced clock slowdown or automatic retry | @@ -352,6 +385,52 @@ any live result. The real scheduler test retains a simulated failed worker's partial receipt and verifies client closure; recovery tests check refusal of reused owners and continued verification after another workflow fails. +### Preserve the first three 4-GiB timed failures + +All three repetitions offered 20 metadata requests/s for 120 seconds, with 32 +clients and the same deterministic uniform 100-repository selection in the full +10,000-identity corpus. Independent closed-ledger replay verified all 7,200 +arrival sequences, selections, outcome counters, nearest-rank percentiles, +in-window/drained throughput and resource/receipt digests. It deliberately does +not hash the changing campaign or active later ledgers as closed evidence. + +| Repetition | OK / busy / HTTP 503 | Successful in-window requests/s | All-attempt p50 / p95 / p99 (ms) | +| --- | --- | --- | --- | +| 1 | 2,269 / 128 / 3 | 18.833 | 137.376 / 1,610.926 / 4,621.883 | +| 2 | 2,392 / 8 / 0 | 19.683 | 135.785 / 872.798 / 1,681.714 | +| 3 | 1,805 / 542 / 53 | 14.833 | 639.841 / 5,893.097 / 9,139.462 | + +That is **6,466 OK, 678 busy drops and 56 HTTP 503s**: 734 failed arrivals. +Busy drops have no invented latency, and success-only latency does not replace +the all-attempt population. These metadata windows are not creation or Git/LFS +write throughput, and no later success can reclassify this campaign as passing. + +Three hypotheses remain open: provider pressure; directory/activation contention; +and driver/shared-host congestion. Dispatch p99 was only 10.511/19.826/20.323 ms, +while service tails were seconds. Failures recur late in windows and in repetition +3, so first-window startup alone is insufficient, and scheduler dispatch alone +does not explain the measured tails. This is captured-ledger replay, not a new +isolated replay of the failing server call or a verified causal fix. + +A separate audit froze only complete existing observer samples wholly within +the three closed resource boundaries (22/22/19 samples), retaining non-atomic +raw cgroup files and exact kernel identities. The sampled intervals cover +116.032/114.444/113.105 seconds, excluding edges without interpolation. Provider +CPU averaged 1.290/1.084/1.506 cores; throttled-time counter increments were +10.158/11.573/46.453 seconds. Memory approached the 4-GiB limit in each interval, +with zero observed `oom`/`oom_kill` counters. Some provider-wide 4xx/5xx counters +also increased; these include background work/retries and do not identify a +specific frontend request or establish billed/per-operation cost. + +The corresponding server logs report authentication/metadata failures with the +opaque `repository directory operation failed` message. The +[error wrapper](../../crates/canopy-server/src/server/mod.rs) retains a typed +invocation failure, but the [HTTP logs](../../crates/canopy-server/src/repository_http/mod.rs) +print only its outer Display value. That observation cannot distinguish a +durable refusal, pending outcome, invalid published result or not-started failure; +it is not evidence of bad user credentials or a proven Cellule bottleneck. No +live limit, source, binary, lease, deadline or provider setting was tuned. + The closed observer reached a sampled maximum `memory.current` of 4,294,967,296 bytes. At its final sample (06:27:17 UTC), current memory was 4,108,148,736 bytes, anonymous memory 1,552,654,336 bytes and file memory 1,404,837,888 bytes. The @@ -436,7 +515,9 @@ CPU, per-operation CPU/GiB or priced object-store cost. ## Bind the new trial -Artifacts are retained under `canopy-native-filesystem-ceXFad8I`: +These historical artifacts were inspected under `canopy-native-filesystem-ceXFad8I`. +That directory is now absent; this table records expected digests, not currently +available or newly reverified local files: | Closed input or receipt | SHA-256 | | --- | --- | @@ -514,6 +595,14 @@ Artifacts are retained under `canopy-native-filesystem-ceXFad8I`: | `harness-38a51dd-linux-36823311167-job.log` | `963a1636810e0efdc768e1d6f4c8297f3d3b07c7d441e99919ee8969d8ab0bcf` | | `rust-38a51dd-linux-36823311167-job.log` | `2384cf85dc36736bb8560902e71227ba823b3120cbb109856b2a040159f99842` | | `harness-84-critical-driver.log` | `ca0bd3cb47be225672e608b5b7a60f036ea82ef15df2d3b58b36ecf8dcafc437` | +| `start_critical_load_4g.py` | `b0ffd626d390a59a0b94e1ce48abd7388561454354dd8fca2ae8979ae5718a5e` | +| `test_start_critical_load_4g.py` | `c1e3bdc631a206cbaa43d681d30cdedc7ec4bca429adda6f111df4cc26b35fc2` | +| `critical-4g-gate-tests.log` | `efae8d6479515f878d6a459abc70a5880a93ef66c8c9712cc86c1d65ddf45fa9` | +| `critical-4g-check.json` | `31424b96b9c18a12dad50bd8688193f871351c9625c24277acb6a7f0a368d993` | +| `native-load-4g/preflight.json` | `78e3950590462eeca9ae8b1df729294ac3ce247056eccc8c342c542e1ca7847d` | +| `load-4g-first-three-audit.json` | `8396bd9f6d29eb5680f791b8288ac8b1e357ed7af9ebdb6861341930fc628f8c` | +| `load-4g-first-three-sidecar-audit.json` | `36602007172205cbec6d0ba1ed7ad1099cf605302ab8bd2dd132f6eb46740b35` | +| `load-4g-first-three-sidecar-prefix.jsonl` | `25162b403cecfdc12de98b1ca06a2da10c8a778c4ee9ab4c053fae1d3fe5982b` | The baseline's new independent ledger audit is `5416540c9b78b23e5c89ff24e771ab58012af847df4ec961adfc597bff010639` From 4a727fdf43220ada8df849821079a29c135bef7b Mon Sep 17 00:00:00 2001 From: forhappy Date: Thu, 1 Oct 2026 11:04:18 -0700 Subject: [PATCH 24/60] docs: record latest Cellule release and RustFS verification --- docs/performance-plan.md | 12 ++- .../2026-10-01-evidence-loss-and-c51dd121.md | 82 ++++++++++++++++--- 2 files changed, 78 insertions(+), 16 deletions(-) diff --git a/docs/performance-plan.md b/docs/performance-plan.md index 0236fe6..d1d36fa 100644 --- a/docs/performance-plan.md +++ b/docs/performance-plan.md @@ -6,15 +6,19 @@ dependency pin is `c51dd121284ecc8878b75d32717a4dfbe2c406c2`, the latest observed upstream main on October 1. Unlike the earlier `a4500add` documentation advance, this changes routing and resource-permit implementation. Locked metadata resolves all six packages; the lockfile changes only their sources. -The new release build and its RustFS/three-node qualification are separate -gates. The 84 Python harness tests passed; this is not Rust runtime verification. -None establishes reference capacity. +The new locked release build passed. Its release workspace suite passed 227 +tests with 9 ignored gates; all eight real-RustFS compatibility gates subsequently +passed, as did the production binary's initial 17-step check through three nodes +and a proxy. The 84 Python harness tests cover accounting and guards. Full-corpus +recovery, scheduled load and every-ACK verification remain separate gates; none +of these functional checks establishes reference capacity. The [three-node proxy campaign](performance/2026-09-30-three-node-proxy.md) retains the earlier fixture race and its regression, plus the candidate's failed seed and successful recovery of every recorded ACK. | Current gate | Evidence / status | | --- | --- | +| Latest Cellule candidate | `c51dd121` locked release build, 227 Rust tests, eight real-RustFS compatibility gates and initial 17-step three-node/proxy Git check passed; new 10K/100 seed running | | Candidate Git behavior | Bound `e07670e` production artifact passed all 17 critical stock-Git steps through three nodes/proxy against RustFS | | Mac-backed candidate seed | Failed: 3,561/10,000 identities and 39 Git/LFS fixtures; incomplete manifest retained | | Every candidate seed ACK | All 3,561 identities and 39 populated fixtures passed fresh-owner recovery; independent ledger/binding and 78-clone audit passed | @@ -25,7 +29,7 @@ seed and successful recovery of every recorded ACK. | Post-load ACK recovery | Recorded three-owner loss and 32.249-s wait; critical fixtures passed. Fresh fleet then recorded ENOSPC and shut down; full corpus and every-ACK verification are incomplete | | Concurrent critical Git workflows | [Separate driver](../scripts/benchmark_critical_git.py) retains the 17-step suite. The versioned gate separates failed performance from mandatory correctness recovery; no live critical-load/fault result is established | | Evidence availability | Benchmark directories and retained executable disappeared during the later space check. No receipt backup found in checked locations; surviving RustFS data is not an ACK ledger | -| Performance qualification | Latest artifact/runtime, every post-load ACK, critical-operation load/fault coverage, higher admission profiles and matched comparisons remain open | +| Performance qualification | Latest full-corpus recovery, scheduled load, every post-load ACK, critical-operation load/fault coverage, higher admission profiles and matched comparisons remain open | The [native-filesystem trial](performance/2026-10-01-native-filesystem.md) records the storage hypothesis, graceful diagnostic drain and new bindings. diff --git a/docs/performance/2026-10-01-evidence-loss-and-c51dd121.md b/docs/performance/2026-10-01-evidence-loss-and-c51dd121.md index 6e7a93f..942a51b 100644 --- a/docs/performance/2026-10-01-evidence-loss-and-c51dd121.md +++ b/docs/performance/2026-10-01-evidence-loss-and-c51dd121.md @@ -1,7 +1,8 @@ # Preserve failures and rebuild the Cellule candidate -**Status: performance failed; post-load correctness recovery incomplete; old raw -evidence unavailable. Latest-Cellule candidate verification is in progress.** +**Status: the old performance attempt failed and its post-load recovery is +incomplete. The new Cellule candidate passed release tests and initial Git +verification; its full corpus setup and performance qualification remain open.** This checkpoint supersedes earlier running/retained-artifact statements in the [native-filesystem trial](2026-10-01-native-filesystem.md). It keeps the original @@ -107,10 +108,54 @@ routing and host-permit changes. It reuses resident catalog identity for unlease requests while still observing fresh authority, and pairs resource charges with semaphore permits. Source inspection is not a Canopy latency or correctness result. -All 84 Python harness tests passed in 44.556 seconds. They exercise benchmark -accounting and guards, not the upgraded Rust implementation. A new locked release -build is running; Rust tests, RustFS end-to-end checks and full three-node/load/fault -verification remain open. New evidence belongs outside Cargo target directories. +The locked release build completed successfully. Its retained executable has +SHA-256 `32b114119960608c0a91d1c783bb69eafec432831bfa452d54d8950b09bc0e99`. +The build receipt binds 256 production-source and harness files; those bindings +still match the published candidate. Evidence and the executable are outside the +rebuildable Cargo target directory. + +| Verification | Result and scope | +| --- | --- | +| Python harness | 84 tests passed in 44.556 s; accounting and guards, not Rust runtime proof | +| Release Rust workspace | 227 tests passed, 9 ignored; the nested isolated native-Git child test is not counted twice | +| Isolated real RustFS compatibility | All eight exact ignored provider tests passed with `scripts/qualify_size.py --provider-only --release` | +| Three-node proxy Git behavior | The retained production executable passed all 17 critical steps against the separately bound RustFS provider | +| CI at `5b2f95c` | Both Rust and harness jobs passed in the [PR workflow](https://github.com/crabbuild/canopy/actions/runs/36898604222) and [branch workflow](https://github.com/crabbuild/canopy/actions/runs/36898599990) | + +The eight provider gates cover SHA-256 round trips and native merge candidates, +signed HTTP push options, signed SHA-256 SSH, stock SSH, bulk refs, partial clones +and SSH-issued LFS access. They ran in the qualification script's own disposable +RustFS fixture. That script selects `rustfs/rustfs:1.0.0-beta.8-glibc`; this is +not a matched provider-envelope comparison with the performance fixture. The +non-sparse 5-GiB size gate was explicitly excluded and remains open for this pin. + +The 17-step production check covers atomic multi-ref publication, mixed and +atomic refusal, correct and stale leases, shallow/deepen/unshallow, filtered +lazy fetches, incremental push/pull, branch deletion and pruning, mirror push, +invalid credentials and exact v0/v2 ref inventories with strict fsck. It passed +at 17:39:09 UTC. These are functional checks, not scheduled throughput, owner-loss +recovery or packet-level confirmation of negotiated protocol versions. + +### Full corpus setup + +The new performance provider has a fresh bucket/prefix on its own native Docker +volume, with the pinned RustFS image, 2 CPUs and 4 GiB memory. Its original +five-second bucket-startup command timed out. A separate startup-completion +receipt first confirmed the bucket absent, then created and checked it without +restarting or replacing the provider. The original failed receipt remains intact. + +Three independent foreground nodes use the retained executable, 100-entry +per-node admission and a loopback TCP proxy. Their launcher has its own session +and is not owned by the finite verification controller. The 10,000-repository +seed started after the eight provider tests exited successfully. It keeps seed +`20260926`, 100 two-commit Git fixtures and 100 one-MiB LFS objects, unchanged +30-second HTTP/120-second Git deadlines, and no retry or reseed. Partial manifests +are retained on failure. Serial seed time is not scheduled creation throughput. + +Four offline seed-guard tests passed with 18 rejection cases for scope trimming, +invalid or duplicate ACK identities and changed Git/LFS expectations. They do not +prove that the live corpus completed. Full verification, fresh-owner recovery, +the original 108-window matrix and every new ACK after load remain separate gates. New closed files are under `/Users/haipingfu/.codex/canopy-three-node-evidence-BwYz7P`, separately from the @@ -121,15 +166,28 @@ rebuildable Cargo target: | `harness.log` | `5696368dfdbb6537716b8e35dddc719726c848369dc47df293e4e092319f0176` | | `metadata.json` | `a4a33f617545d8b7796719b791c9da241b94eb5e89db85fb3f0bb61765418a69` | | `build_candidate.py` | `1e458a80735bdb8c83edf58755ee3933d6a403704b6dc0e8e8bbec89428f1b62` | - -`build.json` and `build.log` remain changing outputs until the actual compiler -process exits; they are not listed as closed evidence here. +| `build.json` | `818077ca334759e90e16fd3326406831f1840ba2f93b3b93952c3b8e81f9f70c` | +| `build.log` | `63c8b2f9eac5b7a144f4f5bfcb61bcbc02c3cc94d61c4bfcffce107125c87748` | +| `canopy-c51dd121` | `32b114119960608c0a91d1c783bb69eafec432831bfa452d54d8950b09bc0e99` | +| `rust-tests.log` | `8dbf6d432544e39510a7b314a6bf8a4291e008ea3f8664ada2c0d3e6a860a42b` | +| `provider-ready.json` | `dd243add0de22069284ddf3181af1f3ed1c94247d1758e27730de53c0e1e6de0` | +| `critical-controller.json` | `402ac5ced76f9585c57e25fc7642cc066b5c5ec3196441b4135373fdd93193a1` | +| `critical.json` | `84ea23bf95c5b9cc448f1e520c663c05d5cca0a952be1056774e02b34d5d2a03` | +| `provider-tests.json` | `cd94183c06648d26a44da368c8cb3cd49f597124e3b37c338673794eac1e6412` | +| `provider-tests.log` | `27f48704797584c5b08f1297c9e1cae2ef0c5f1ff8965a557a07a5f9eefbc2dd` | +| `run_provider_gates.py` | `c1f0e2a52c474fea0adeb4b671495f2d0980120b3067ebb2b38245a3f3492deb` | +| `seed_full_corpus.py` | `7bbd79a6dae40fec4809720970d80a8465be4c8df4d63c242a4d59ef63a6fc0e` | +| `seed-guard-tests.log` | `3c82e7c0a92ea5b6efa559b91d414936635fdf034fc5182b260ab7d2c0d07cc8` | + +The live seed manifest, controller receipt and log are changing outputs, not +closed evidence. Do not use their intermediate counts as a complete-corpus result. | Required gate | Current status | | --- | --- | -| Exact new artifact and source provenance | Build in progress | -| Rust workspace correctness and RustFS Git/LFS compatibility | Open | -| Full 10,000 identities/100 populated fixtures | Open for the new artifact | +| Exact new artifact and source provenance | Passed locked build and current binding checks | +| Rust workspace correctness and RustFS Git/LFS compatibility | Passed release suite, eight provider gates and initial 17-step proxy check | +| New-pin non-sparse 5-GiB transfer | Open; excluded by `--provider-only` | +| Full 10,000 identities/100 populated fixtures | Seed running; full verification and fresh-owner recovery open | | Original matrix, critical Git load and every new ACK after owner loss | Open | | Higher admission profiles and matched comparisons | Open | | Old campaign every-ACK recovery | Unverified; original raw inputs unavailable | From 7d0959746b0158fc359669ec83591e97a4eef037 Mon Sep 17 00:00:00 2001 From: forhappy Date: Thu, 1 Oct 2026 11:07:28 -0700 Subject: [PATCH 25/60] docs: distinguish frozen candidate from advancing upstream --- docs/performance-plan.md | 7 +++++-- .../2026-10-01-evidence-loss-and-c51dd121.md | 12 +++++++++++- 2 files changed, 16 insertions(+), 3 deletions(-) diff --git a/docs/performance-plan.md b/docs/performance-plan.md index d1d36fa..860fbad 100644 --- a/docs/performance-plan.md +++ b/docs/performance-plan.md @@ -2,8 +2,8 @@ Use this plan to design capacity work and interpret Canopy benchmark results. It separates demonstrated behavior from proposed targets. The current Cellule -dependency pin is `c51dd121284ecc8878b75d32717a4dfbe2c406c2`, the latest -observed upstream main on October 1. Unlike the earlier `a4500add` documentation +dependency pin is `c51dd121284ecc8878b75d32717a4dfbe2c406c2`, observed upstream +main at build preparation on October 1. Unlike the earlier `a4500add` documentation advance, this changes routing and resource-permit implementation. Locked metadata resolves all six packages; the lockfile changes only their sources. The new locked release build passed. Its release workspace suite passed 227 @@ -12,6 +12,9 @@ passed, as did the production binary's initial 17-step check through three nodes and a proxy. The 84 Python harness tests cover accounting and guards. Full-corpus recovery, scheduled load and every-ACK verification remain separate gates; none of these functional checks establishes reference capacity. +Upstream subsequently advanced to `0dc04a6` with atomic API and test changes. +That revision needs a separate artifact and runtime qualification; the running +seed and its frozen source remain on `c51dd121`. The [three-node proxy campaign](performance/2026-09-30-three-node-proxy.md) retains the earlier fixture race and its regression, plus the candidate's failed seed and successful recovery of every recorded ACK. diff --git a/docs/performance/2026-10-01-evidence-loss-and-c51dd121.md b/docs/performance/2026-10-01-evidence-loss-and-c51dd121.md index 942a51b..c612e76 100644 --- a/docs/performance/2026-10-01-evidence-loss-and-c51dd121.md +++ b/docs/performance/2026-10-01-evidence-loss-and-c51dd121.md @@ -98,7 +98,7 @@ mark old recovery passing, or combine a new run with the old campaign. ## Latest Cellule candidate -The dependency now pins observed upstream main +The dependency pins upstream main observed at build preparation, `c51dd121284ecc8878b75d32717a4dfbe2c406c2`. Five direct declarations and six lockfile sources changed; unrelated dependencies did not. Locked metadata resolves all six Cellule packages to that revision. @@ -108,6 +108,15 @@ routing and host-permit changes. It reuses resident catalog identity for unlease requests while still observing fresh authority, and pairs resource charges with semaphore permits. Source inspection is not a Canopy latency or correctness result. +Upstream advanced at 17:43:36 UTC to +[`0dc04a658bd99668936f7ec58032d054f6fbc141`](https://github.com/crabbuild/cellule/commit/0dc04a658bd99668936f7ec58032d054f6fbc141). +The inspected diff replaces deprecated atomic `fetch_update` calls with +`try_update` in LTX accounting, host disk budgeting and runtime admission, as +well as tests and the website-example checker. It is not documentation-only. +The running experiment remains bound to `c51dd121`; its results do not qualify +this newer SHA. A separately built and verified candidate is required before +upgrading the pin or changing the live experiment. + The locked release build completed successfully. Its retained executable has SHA-256 `32b114119960608c0a91d1c783bb69eafec432831bfa452d54d8950b09bc0e99`. The build receipt binds 256 production-source and harness files; those bindings @@ -190,6 +199,7 @@ closed evidence. Do not use their intermediate counts as a complete-corpus resul | Full 10,000 identities/100 populated fixtures | Seed running; full verification and fresh-owner recovery open | | Original matrix, critical Git load and every new ACK after owner loss | Open | | Higher admission profiles and matched comparisons | Open | +| New upstream `0dc04a6` artifact and runtime qualification | Open; separate from the bound `c51dd121` experiment | | Old campaign every-ACK recovery | Unverified; original raw inputs unavailable | The [performance plan](../performance-plan.md) remains the scope. No proven From 194503c51a0dc567c63899f2a54b23ee2b1c4e4c Mon Sep 17 00:00:00 2001 From: forhappy Date: Thu, 1 Oct 2026 12:09:49 -0700 Subject: [PATCH 26/60] docs: record full seed and verified evidence backups --- docs/performance-plan.md | 8 +-- .../2026-10-01-evidence-loss-and-c51dd121.md | 59 +++++++++++++++++-- 2 files changed, 57 insertions(+), 10 deletions(-) diff --git a/docs/performance-plan.md b/docs/performance-plan.md index 860fbad..37a7a61 100644 --- a/docs/performance-plan.md +++ b/docs/performance-plan.md @@ -13,15 +13,15 @@ and a proxy. The 84 Python harness tests cover accounting and guards. Full-corpu recovery, scheduled load and every-ACK verification remain separate gates; none of these functional checks establishes reference capacity. Upstream subsequently advanced to `0dc04a6` with atomic API and test changes. -That revision needs a separate artifact and runtime qualification; the running -seed and its frozen source remain on `c51dd121`. +That revision needs a separate artifact and runtime qualification; the current +experiment and its frozen source remain on `c51dd121`. The [three-node proxy campaign](performance/2026-09-30-three-node-proxy.md) retains the earlier fixture race and its regression, plus the candidate's failed seed and successful recovery of every recorded ACK. | Current gate | Evidence / status | | --- | --- | -| Latest Cellule candidate | `c51dd121` locked release build, 227 Rust tests, eight real-RustFS compatibility gates and initial 17-step three-node/proxy Git check passed; new 10K/100 seed running | +| Latest Cellule candidate | `c51dd121` locked release build, 227 Rust tests, eight real-RustFS compatibility gates and initial 17-step three-node/proxy Git check passed; new full 10K/100 seed closed and remote verification running | | Candidate Git behavior | Bound `e07670e` production artifact passed all 17 critical stock-Git steps through three nodes/proxy against RustFS | | Mac-backed candidate seed | Failed: 3,561/10,000 identities and 39 Git/LFS fixtures; incomplete manifest retained | | Every candidate seed ACK | All 3,561 identities and 39 populated fixtures passed fresh-owner recovery; independent ledger/binding and 78-clone audit passed | @@ -31,7 +31,7 @@ seed and successful recovery of every recorded ACK. | Complete 4-GiB load attempt | All 108 windows finished: 60,197 OK and 54,763 failed arrivals. The original attempt remains failed; local raw evidence is now unavailable | | Post-load ACK recovery | Recorded three-owner loss and 32.249-s wait; critical fixtures passed. Fresh fleet then recorded ENOSPC and shut down; full corpus and every-ACK verification are incomplete | | Concurrent critical Git workflows | [Separate driver](../scripts/benchmark_critical_git.py) retains the 17-step suite. The versioned gate separates failed performance from mandatory correctness recovery; no live critical-load/fault result is established | -| Evidence availability | Benchmark directories and retained executable disappeared during the later space check. No receipt backup found in checked locations; surviving RustFS data is not an ACK ledger | +| Evidence availability | Old benchmark directories and executable disappeared; old receipts remain unavailable. New closed artifacts, all build bindings and the completed seed have verified backups on a separate filesystem outside Cargo targets; changing outputs are excluded | | Performance qualification | Latest full-corpus recovery, scheduled load, every post-load ACK, critical-operation load/fault coverage, higher admission profiles and matched comparisons remain open | The [native-filesystem trial](performance/2026-10-01-native-filesystem.md) diff --git a/docs/performance/2026-10-01-evidence-loss-and-c51dd121.md b/docs/performance/2026-10-01-evidence-loss-and-c51dd121.md index c612e76..f75b7b6 100644 --- a/docs/performance/2026-10-01-evidence-loss-and-c51dd121.md +++ b/docs/performance/2026-10-01-evidence-loss-and-c51dd121.md @@ -2,7 +2,8 @@ **Status: the old performance attempt failed and its post-load recovery is incomplete. The new Cellule candidate passed release tests and initial Git -verification; its full corpus setup and performance qualification remain open.** +verification. Its full corpus seed completed; full verification, recovery and +performance qualification remain open.** This checkpoint supersedes earlier running/retained-artifact statements in the [native-filesystem trial](2026-10-01-native-filesystem.md). It keeps the original @@ -130,6 +131,7 @@ rebuildable Cargo target directory. | Isolated real RustFS compatibility | All eight exact ignored provider tests passed with `scripts/qualify_size.py --provider-only --release` | | Three-node proxy Git behavior | The retained production executable passed all 17 critical steps against the separately bound RustFS provider | | CI at `5b2f95c` | Both Rust and harness jobs passed in the [PR workflow](https://github.com/crabbuild/canopy/actions/runs/36898604222) and [branch workflow](https://github.com/crabbuild/canopy/actions/runs/36898599990) | +| CI at `7d09597` | All four Rust/harness checks passed in the [PR workflow](https://github.com/crabbuild/canopy/actions/runs/36904766930) and [branch workflow](https://github.com/crabbuild/canopy/actions/runs/36904761443) | The eight provider gates cover SHA-256 round trips and native merge candidates, signed HTTP push options, signed SHA-256 SSH, stock SSH, bulk refs, partial clones @@ -163,9 +165,30 @@ are retained on failure. Serial seed time is not scheduled creation throughput. Four offline seed-guard tests passed with 18 rejection cases for scope trimming, invalid or duplicate ACK identities and changed Git/LFS expectations. They do not -prove that the live corpus completed. Full verification, fresh-owner recovery, +prove remote Git/LFS delivery or recovery. Full verification, fresh-owner recovery, the original 108-window matrix and every new ACK after load remain separate gates. +The seed process exited successfully at **19:02:14 UTC**. Its closed manifest +contains all **10,000 identities, 100 populated Git fixtures and 100 LFS objects**; +the setup controller validated their expected identities and payload declarations. +The serial seed took 3,852.843 seconds. This is setup duration, not scheduled +creation throughput. At 19:09 UTC, the separate full-corpus verifier is running against the +same three owners and unchanged RustFS provider; setup alone does not establish +that all remote Git/LFS bytes survive owner loss. + +The initial fault, independent fresh fleet and full recovery controllers are +armed behind that verifier. The original 108-window matrix waits for full initial +recovery. Separate post-load fault, independent fresh fleet and every-ACK recovery +controllers are also armed, but have not signaled any post-load owner or passed +recovery. They require an independent replay of all closed ledgers and resource +boundaries. Failed performance does not discard ACKs or prevent their recovery +check; an interrupted or reduced matrix cannot satisfy this gate. + +Eleven offline post-load guard tests passed, covering complete arrival/ACK +accounting, changed kernel identities, launcher resumption after a signal or +receipt-write failure, and rejection of missing corpus/critical/ACK coverage. +These are local helper tests, not actual owner-loss or delivered-body proof. + New closed files are under `/Users/haipingfu/.codex/canopy-three-node-evidence-BwYz7P`, separately from the rebuildable Cargo target: @@ -188,16 +211,40 @@ rebuildable Cargo target: | `seed_full_corpus.py` | `7bbd79a6dae40fec4809720970d80a8465be4c8df4d63c242a4d59ef63a6fc0e` | | `seed-guard-tests.log` | `3c82e7c0a92ea5b6efa559b91d414936635fdf034fc5182b260ab7d2c0d07cc8` | -The live seed manifest, controller receipt and log are changing outputs, not -closed evidence. Do not use their intermediate counts as a complete-corpus result. +The seed manifest, controller receipt and log are now closed: + +| Closed seed artifact | SHA-256 | +| --- | --- | +| `corpus-10000.json` | `9ce7da04a7642bc7342d204c874cabf39ceea6c40ffd51211797687f4dd46397` | +| `seed-controller.json` | `b93d3e4ab71f2f8e98f63efd29b67c2477bd7089e8658a941368d81a162a2f1c` | +| `seed.log` | `be3fdb0b01b04c6205ce456ec8c33230cac9a7f27175c8645ebd63606d62a7b7` | + +### Independent evidence backup + +A new backup under +`/Volumes/Workspace/CrabData/canopy-evidence-backup-BwYz7P-bYyTQg` +is outside Cargo targets and on a different filesystem from the evidence root. +The initial copy contains 270 files: all 15 then-published closed artifacts and +all 256 build bindings, with one shared helper counted once. Every source and +copy digest matched, and a separate read replayed all copied digests. The closed +full seed was copied separately only after its process actually exited. + +`closed-backup.json` has SHA-256 +`4622fa3b31ca5eeb76ccfd48e4a352480e262f173efeb0dce676ff0afd879a41`; +`full-seed-backup.json` has SHA-256 +`8f95fc0073c6156792c766b7539ca1e3fb1e7ce4908f0559ababdfcab0b49521`. +This local preservation does not recover the old missing ledgers or constitute +an off-machine backup. Private fleet configs and changing verifier/load receipts +are excluded. Their existence or intermediate counts are not passing results. | Required gate | Current status | | --- | --- | | Exact new artifact and source provenance | Passed locked build and current binding checks | | Rust workspace correctness and RustFS Git/LFS compatibility | Passed release suite, eight provider gates and initial 17-step proxy check | | New-pin non-sparse 5-GiB transfer | Open; excluded by `--provider-only` | -| Full 10,000 identities/100 populated fixtures | Seed running; full verification and fresh-owner recovery open | -| Original matrix, critical Git load and every new ACK after owner loss | Open | +| Full 10,000 identities/100 populated fixtures | Full seed closed; remote verification running and fresh-owner recovery pending | +| Original matrix and every new ACK after owner loss | Controllers armed behind prerequisite gates; no timed-load or recovery pass | +| Critical concurrent Git load and fault coverage | Open; separate from the initial 17-step functional check | | Higher admission profiles and matched comparisons | Open | | New upstream `0dc04a6` artifact and runtime qualification | Open; separate from the bound `c51dd121` experiment | | Old campaign every-ACK recovery | Unverified; original raw inputs unavailable | From d6d63a398693a503aeafc5488e6afb2e0703ef89 Mon Sep 17 00:00:00 2001 From: forhappy Date: Thu, 1 Oct 2026 12:20:49 -0700 Subject: [PATCH 27/60] docs: retain failed full-corpus verification and diagnostic scope --- docs/performance-plan.md | 2 +- .../2026-10-01-evidence-loss-and-c51dd121.md | 53 +++++++++++++++---- 2 files changed, 43 insertions(+), 12 deletions(-) diff --git a/docs/performance-plan.md b/docs/performance-plan.md index 37a7a61..f6338ea 100644 --- a/docs/performance-plan.md +++ b/docs/performance-plan.md @@ -21,7 +21,7 @@ seed and successful recovery of every recorded ACK. | Current gate | Evidence / status | | --- | --- | -| Latest Cellule candidate | `c51dd121` locked release build, 227 Rust tests, eight real-RustFS compatibility gates and initial 17-step three-node/proxy Git check passed; new full 10K/100 seed closed and remote verification running | +| Latest Cellule candidate | `c51dd121` build, 227 Rust tests, eight RustFS gates and initial 17-step proxy check passed; full 10K/100 seed closed, but remote verification failed HTTP 503. Downstream fault/load gates stopped without owner signals; cause unresolved | | Candidate Git behavior | Bound `e07670e` production artifact passed all 17 critical stock-Git steps through three nodes/proxy against RustFS | | Mac-backed candidate seed | Failed: 3,561/10,000 identities and 39 Git/LFS fixtures; incomplete manifest retained | | Every candidate seed ACK | All 3,561 identities and 39 populated fixtures passed fresh-owner recovery; independent ledger/binding and 78-clone audit passed | diff --git a/docs/performance/2026-10-01-evidence-loss-and-c51dd121.md b/docs/performance/2026-10-01-evidence-loss-and-c51dd121.md index f75b7b6..130e867 100644 --- a/docs/performance/2026-10-01-evidence-loss-and-c51dd121.md +++ b/docs/performance/2026-10-01-evidence-loss-and-c51dd121.md @@ -2,8 +2,8 @@ **Status: the old performance attempt failed and its post-load recovery is incomplete. The new Cellule candidate passed release tests and initial Git -verification. Its full corpus seed completed; full verification, recovery and -performance qualification remain open.** +verification. Its full corpus seed completed, but remote verification failed +with HTTP 503. Recovery and timed load did not start.** This checkpoint supersedes earlier running/retained-artifact statements in the [native-filesystem trial](2026-10-01-native-filesystem.md). It keeps the original @@ -176,19 +176,44 @@ creation throughput. At 19:09 UTC, the separate full-corpus verifier is running same three owners and unchanged RustFS provider; setup alone does not establish that all remote Git/LFS bytes survive owner loss. -The initial fault, independent fresh fleet and full recovery controllers are -armed behind that verifier. The original 108-window matrix waits for full initial -recovery. Separate post-load fault, independent fresh fleet and every-ACK recovery -controllers are also armed, but have not signaled any post-load owner or passed -recovery. They require an independent replay of all closed ledgers and resource -boundaries. Failed performance does not discard ACKs or prevent their recovery -check; an interrupted or reduced matrix cannot satisfy this gate. +The initial fault, independent fresh fleet, full recovery and original 108-window +matrix controllers were armed behind that verifier. Separate post-load fault, +independent fresh fleet and every-ACK recovery controllers were also armed. +All stopped after the verification failure, without owner signals or timed load. +The post-load gate requires an independent replay of all closed ledgers and +resource boundaries. Failed performance would not discard ACKs or prevent their +recovery check; an interrupted or reduced matrix cannot satisfy this gate. Eleven offline post-load guard tests passed, covering complete arrival/ACK accounting, changed kernel identities, launcher resumption after a signal or receipt-write failure, and rejection of missing corpus/critical/ACK coverage. These are local helper tests, not actual owner-loss or delivered-body proof. +### Remote verification failed + +At **19:10:19 UTC**, the verifier exited 1 on HTTP 503 for +`density-3a92b05e1d80-07076` (`f2a0f25e-9db4-4a16-8a5b-1db1d43aeac9`), +request ID `6a4a923c-fbe1-419a-8943-930c7a02a67c`. Its last progress line was +6,800/10,000 identities; that is not the exact count of completed concurrent +requests or a passing prefix. No complete Git/LFS result was returned. +At the same second, node 0 logged `repository directory operation failed` +without an underlying error chain or request ID. That message does not establish +the root cause or identify a Cellule bottleneck. + +All eight finite gate/launcher processes subsequently exited. The original +three owners remained live, and the rejected post-load receipt contains an +empty signal list. There is no new recovery fleet or timed campaign directory. +Their terminal receipts and logs have verified independent-filesystem copies; +the provider retained its original start time, zero restarts and zero recorded +cgroup OOM events. Neither the failures nor the data were discarded or reseeded. + +A separate read-only replay of the failed identity and its original 16-entry +batch returned matching identities on all 67 requests. It did not reproduce +the 503 and does not overturn the original failure. A metadata-only diagnostic +across all 10,000 identities is running with concurrency 16 and the same +30-second HTTP deadline. It excludes Git/LFS body checks and is not a replacement +qualification run. The cause remains unresolved. + New closed files are under `/Users/haipingfu/.codex/canopy-three-node-evidence-BwYz7P`, separately from the rebuildable Cargo target: @@ -219,6 +244,12 @@ The seed manifest, controller receipt and log are now closed: | `seed-controller.json` | `b93d3e4ab71f2f8e98f63efd29b67c2477bd7089e8658a941368d81a162a2f1c` | | `seed.log` | `be3fdb0b01b04c6205ce456ec8c33230cac9a7f27175c8645ebd63606d62a7b7` | +| Closed failed verification or diagnostic | SHA-256 | +| --- | --- | +| `full-seed-verification.json` | `981dd75f4bc0fc1605cc06b579d9966b0c597ac64e2c5e4c6e3a29a9561fdf01` | +| `full-seed-verification.log` | `273e945deb34c5792e12da3d7d4a01819cc80e402a55ea64d04346c1a25be593` | +| `identity-503-replay.json` | `2410d2c63edcd1d4700b20bb795dc680b012190517a3ee063f3f87904c174caf` | + ### Independent evidence backup A new backup under @@ -242,8 +273,8 @@ are excluded. Their existence or intermediate counts are not passing results. | Exact new artifact and source provenance | Passed locked build and current binding checks | | Rust workspace correctness and RustFS Git/LFS compatibility | Passed release suite, eight provider gates and initial 17-step proxy check | | New-pin non-sparse 5-GiB transfer | Open; excluded by `--provider-only` | -| Full 10,000 identities/100 populated fixtures | Full seed closed; remote verification running and fresh-owner recovery pending | -| Original matrix and every new ACK after owner loss | Controllers armed behind prerequisite gates; no timed-load or recovery pass | +| Full 10,000 identities/100 populated fixtures | Full seed closed; remote verification failed HTTP 503; cause unresolved | +| Original matrix and every new ACK after owner loss | Downstream gates stopped without owner signals or timed load; recovery open | | Critical concurrent Git load and fault coverage | Open; separate from the initial 17-step functional check | | Higher admission profiles and matched comparisons | Open | | New upstream `0dc04a6` artifact and runtime qualification | Open; separate from the bound `c51dd121` experiment | From f0c5d093091ba11ecbf0fb0400e65dddbbf8ea6e Mon Sep 17 00:00:00 2001 From: forhappy Date: Thu, 1 Oct 2026 13:22:41 -0700 Subject: [PATCH 28/60] docs: record closed directory diagnostics and test verification --- docs/performance-plan.md | 8 +- .../2026-10-01-evidence-loss-and-c51dd121.md | 78 +++++++++++++++++-- 2 files changed, 78 insertions(+), 8 deletions(-) diff --git a/docs/performance-plan.md b/docs/performance-plan.md index f6338ea..7427dd7 100644 --- a/docs/performance-plan.md +++ b/docs/performance-plan.md @@ -21,7 +21,7 @@ seed and successful recovery of every recorded ACK. | Current gate | Evidence / status | | --- | --- | -| Latest Cellule candidate | `c51dd121` build, 227 Rust tests, eight RustFS gates and initial 17-step proxy check passed; full 10K/100 seed closed, but remote verification failed HTTP 503. Downstream fault/load gates stopped without owner signals; cause unresolved | +| Latest Cellule candidate | `c51dd121` build, 227 Rust tests, eight RustFS gates and initial 17-step proxy check passed; full 10K/100 seed closed, but remote verification failed HTTP 503. A separate 10K metadata diagnostic returned 9,999 matching identities and one 503; downstream fault/load gates stopped without owner signals. Cause unresolved | | Candidate Git behavior | Bound `e07670e` production artifact passed all 17 critical stock-Git steps through three nodes/proxy against RustFS | | Mac-backed candidate seed | Failed: 3,561/10,000 identities and 39 Git/LFS fixtures; incomplete manifest retained | | Every candidate seed ACK | All 3,561 identities and 39 populated fixtures passed fresh-owner recovery; independent ledger/binding and 78-clone audit passed | @@ -39,6 +39,12 @@ records the storage hypothesis, graceful diagnostic drain and new bindings. The [evidence-loss checkpoint](performance/2026-10-01-evidence-loss-and-c51dd121.md) supersedes its earlier running/retained-artifact status. Historical observations are not currently replayable; a new run cannot certify the old ACKs. +The separate diagnostic build passed 230 release workspace tests with 9 ignored +after isolating a concurrent-fork cleanup test and adding fence-safety coverage. +All 20 subsequent full-library repetitions passed at four test threads. +That local test finding does not explain the Directory 503, change production +cleanup or qualify the PR's frozen artifact; temporary diagnostics remain outside +the PR. No Cellule bottleneck or matched performance improvement is proven. The 4-GiB diagnostic is a new resource envelope, not a passing 2-GiB result. Offline plans for 500/1,000-entry admission retain the full matrix but have not been executed. diff --git a/docs/performance/2026-10-01-evidence-loss-and-c51dd121.md b/docs/performance/2026-10-01-evidence-loss-and-c51dd121.md index 130e867..2c8f682 100644 --- a/docs/performance/2026-10-01-evidence-loss-and-c51dd121.md +++ b/docs/performance/2026-10-01-evidence-loss-and-c51dd121.md @@ -132,6 +132,7 @@ rebuildable Cargo target directory. | Three-node proxy Git behavior | The retained production executable passed all 17 critical steps against the separately bound RustFS provider | | CI at `5b2f95c` | Both Rust and harness jobs passed in the [PR workflow](https://github.com/crabbuild/canopy/actions/runs/36898604222) and [branch workflow](https://github.com/crabbuild/canopy/actions/runs/36898599990) | | CI at `7d09597` | All four Rust/harness checks passed in the [PR workflow](https://github.com/crabbuild/canopy/actions/runs/36904766930) and [branch workflow](https://github.com/crabbuild/canopy/actions/runs/36904761443) | +| CI at `d6d63a3` | All four Rust/harness checks passed in the [PR workflow](https://github.com/crabbuild/canopy/actions/runs/36913619350) and [branch workflow](https://github.com/crabbuild/canopy/actions/runs/36913616080); later heads need their own results | The eight provider gates cover SHA-256 round trips and native merge candidates, signed HTTP push options, signed SHA-256 SSH, stock SSH, bulk refs, partial clones @@ -172,9 +173,9 @@ The seed process exited successfully at **19:02:14 UTC**. Its closed manifest contains all **10,000 identities, 100 populated Git fixtures and 100 LFS objects**; the setup controller validated their expected identities and payload declarations. The serial seed took 3,852.843 seconds. This is setup duration, not scheduled -creation throughput. At 19:09 UTC, the separate full-corpus verifier is running against the -same three owners and unchanged RustFS provider; setup alone does not establish -that all remote Git/LFS bytes survive owner loss. +creation throughput. The separate full-corpus verifier subsequently failed +against the same three owners and unchanged RustFS provider. Setup alone does +not establish that all remote Git/LFS bytes survive owner loss. The initial fault, independent fresh fleet, full recovery and original 108-window matrix controllers were armed behind that verifier. Separate post-load fault, @@ -209,10 +210,57 @@ cgroup OOM events. Neither the failures nor the data were discarded or reseeded. A separate read-only replay of the failed identity and its original 16-entry batch returned matching identities on all 67 requests. It did not reproduce -the 503 and does not overturn the original failure. A metadata-only diagnostic -across all 10,000 identities is running with concurrency 16 and the same -30-second HTTP deadline. It excludes Git/LFS body checks and is not a replacement -qualification run. The cause remains unresolved. +the 503 and does not overturn the original failure. The metadata-only diagnostic +closed at **19:27:31 UTC**, retaining all 10,000 observations at concurrency 16 +and the same 30-second HTTP deadline: + +| Diagnostic outcome | Observations | +| --- | ---: | +| HTTP 200 with matching repository identity | 9,999 | +| HTTP 503 | 1 | +| Total | 10,000 | + +The failure was for `density-3a92b05e1d80-04322` +(`f39019bc-2063-438c-8f0a-10fa01f17cea`), request ID +`b221bf05-2df9-4b44-9b48-8d95fb4ffa88`, at 19:22:01 UTC. Node 2 logged +`authentication failed` with `repository directory operation failed` at that +time. The original verifier failed at the metadata-read call site; this is a +related Directory failure, not proof of an identical underlying cause. +The diagnostic excludes Git/LFS bodies and is neither a qualification retry +nor a scheduled throughput measurement. **The root cause remains unresolved.** + +### Separate diagnostic build and cleanup test investigation + +Temporary error-only classification lives in a separate diagnostic worktree, +not this PR's production candidate. It records bounded request IDs and static +error categories without raw headers, private error payloads or provider URLs. +It does not change routing, retries, deadlines or HTTP responses, and has not +yet been exercised against the original three-node corpus. + +Its first release workspace suite failed the existing +`failed_spawn_releases_parent_fence_before_cache_cleanup` test. Twenty full +library repetitions reproduced that failure four times. A controlled fork +reproduction showed that an unrelated child can inherit the Git cache fence +before executing, so cleanup conservatively retains the files and their +132-byte budget charge. This is separate from the Directory HTTP 503; no causal +link or production performance improvement is established. + +The local test correction isolates the parent-fence assertion in a subprocess +and adds a deterministic inherited-fence safety check. Production cleanup stays +unchanged: files remain charged while the fence is busy. The corrected diagnostic +build at local commit `0ee8f69` passed the full locked release workspace suite +at **20:11:04 UTC**: **230 top-level tests passed, 9 ignored**. Nested subprocess +tests are counted once. Its executable and library-test artifact are retained +outside Cargo targets. These local results do not qualify the PR artifact, +newer Cellule upstream, RustFS recovery or performance. Temporary diagnostic +source and the test correction remain separate from this PR while the original +experiment's source bindings stay frozen. + +A subsequent check closed at **20:21:06 UTC**: all **20 full-library repetitions** +passed at four test threads, with 115 tests in each repetition and zero failed +attempts. Every attempt log and digest is retained. This verifies the local test +correction without erasing the original failure; it is not runtime recovery or +proof that an intermittent Directory error is fixed. New closed files are under `/Users/haipingfu/.codex/canopy-three-node-evidence-BwYz7P`, separately from the @@ -249,6 +297,20 @@ The seed manifest, controller receipt and log are now closed: | `full-seed-verification.json` | `981dd75f4bc0fc1605cc06b579d9966b0c597ac64e2c5e4c6e3a29a9561fdf01` | | `full-seed-verification.log` | `273e945deb34c5792e12da3d7d4a01819cc80e402a55ea64d04346c1a25be593` | | `identity-503-replay.json` | `2410d2c63edcd1d4700b20bb795dc680b012190517a3ee063f3f87904c174caf` | +| `metadata-churn-diagnostic.json` | `9dfb64b72791e252bcbe1ebb8cb1e6192dd12c4e8bd07f45b77ca8155b9a2337` | +| `metadata-churn-diagnostic.samples.jsonl` | `a6b1827ca695e7ea0eea47ce1b96e535d1a14b87d02ebc18752d6cb06906b8d0` | + +The separate local diagnostic evidence is under +`/Users/haipingfu/.codex/canopy-directory-diagnostics-v2-rNr2qi`: + +| Closed diagnostic artifact | SHA-256 | +| --- | --- | +| `build-tests.json` | `03aaefd41ba4dc594b635921a7a3cd5087a2f973dd8dae0e1ad995ca08f8e263` | +| `rust-tests.log` | `a91eeeb87b1f232d692cfa5a2121b63a1f52cd78cfa73e6c2b17c7542898d665` | +| `canopy-c51-directory-diagnostics` | `e2e5c12e504b9e73d11e7b6f650fa3f35e46b137136e7eddaaf6ed7b339638a2` | +| `canopy-c51-diagnostic-library-tests` | `795b6ff0abc2d1eee01c83d927b8e8b0a2b0827293b2f42acaee8f7e1496980f` | +| `test-artifact.json` | `fbdd719e2dd5e0d12dfab935ff470eff5b5f87206354462fd5df97e844ad8ae9` | +| `cleanup-reproduction/reproduction.json` | `7e71923d85025177a7b0a273656a4daccded11ba6c9f0e174814cfc4aaec50ed` | ### Independent evidence backup @@ -259,6 +321,8 @@ The initial copy contains 270 files: all 15 then-published closed artifacts and all 256 build bindings, with one shared helper counted once. Every source and copy digest matched, and a separate read replayed all copied digests. The closed full seed was copied separately only after its process actually exited. +The closed metadata diagnostic, all 10,000 samples and its helper also have +verified copies in the backup's `closed-metadata-churn-diagnostic` directory. `closed-backup.json` has SHA-256 `4622fa3b31ca5eeb76ccfd48e4a352480e262f173efeb0dce676ff0afd879a41`; From 4651aaa2cc53ff4b42d8d0b4d691dcfe9affde45 Mon Sep 17 00:00:00 2001 From: forhappy Date: Thu, 1 Oct 2026 14:19:25 -0700 Subject: [PATCH 29/60] docs: record closed authentication diagnosis and candidate failure --- .../2026-10-01-evidence-loss-and-c51dd121.md | 87 ++++++++++++++++++- 1 file changed, 83 insertions(+), 4 deletions(-) diff --git a/docs/performance/2026-10-01-evidence-loss-and-c51dd121.md b/docs/performance/2026-10-01-evidence-loss-and-c51dd121.md index 2c8f682..773fffa 100644 --- a/docs/performance/2026-10-01-evidence-loss-and-c51dd121.md +++ b/docs/performance/2026-10-01-evidence-loss-and-c51dd121.md @@ -133,6 +133,7 @@ rebuildable Cargo target directory. | CI at `5b2f95c` | Both Rust and harness jobs passed in the [PR workflow](https://github.com/crabbuild/canopy/actions/runs/36898604222) and [branch workflow](https://github.com/crabbuild/canopy/actions/runs/36898599990) | | CI at `7d09597` | All four Rust/harness checks passed in the [PR workflow](https://github.com/crabbuild/canopy/actions/runs/36904766930) and [branch workflow](https://github.com/crabbuild/canopy/actions/runs/36904761443) | | CI at `d6d63a3` | All four Rust/harness checks passed in the [PR workflow](https://github.com/crabbuild/canopy/actions/runs/36913619350) and [branch workflow](https://github.com/crabbuild/canopy/actions/runs/36913616080); later heads need their own results | +| CI at `f0c5d09` | All four Rust/harness checks passed in the [PR workflow](https://github.com/crabbuild/canopy/actions/runs/36921149190) and [branch workflow](https://github.com/crabbuild/canopy/actions/runs/36921143618); this evidence update needs its own results | The eight provider gates cover SHA-256 round trips and native merge candidates, signed HTTP push options, signed SHA-256 SSH, stock SSH, bulk refs, partial clones @@ -201,8 +202,8 @@ At the same second, node 0 logged `repository directory operation failed` without an underlying error chain or request ID. That message does not establish the root cause or identify a Cellule bottleneck. -All eight finite gate/launcher processes subsequently exited. The original -three owners remained live, and the rejected post-load receipt contains an +All eight finite gate/launcher processes subsequently exited. At that checkpoint, +the original three owners remained live, and the rejected post-load receipt contains an empty signal list. There is no new recovery fleet or timed campaign directory. Their terminal receipts and logs have verified independent-filesystem copies; the provider retained its original start time, zero restarts and zero recorded @@ -234,8 +235,8 @@ nor a scheduled throughput measurement. **The root cause remains unresolved.** Temporary error-only classification lives in a separate diagnostic worktree, not this PR's production candidate. It records bounded request IDs and static error categories without raw headers, private error payloads or provider URLs. -It does not change routing, retries, deadlines or HTTP responses, and has not -yet been exercised against the original three-node corpus. +It does not change routing, retries, deadlines or HTTP responses. The later +same-corpus replay is recorded below; the original failed receipts stay intact. Its first release workspace suite failed the existing `failed_spawn_releases_parent_fence_before_cache_cleanup` test. Twenty full @@ -262,6 +263,60 @@ attempts. Every attempt log and digest is retained. This verifies the local test correction without erasing the original failure; it is not runtime recovery or proof that an intermittent Directory error is fixed. +### Same corpus authentication diagnosis + +The original fleet drained gracefully at **20:29:37 UTC**, after signaling only +its launcher. This was not an owner-loss test. Three diagnostic nodes then +started against the unchanged RustFS provider and existing corpus, with the same +admission limits. The UI preview was not changed. + +Three full metadata sweeps closed at **20:53:47 UTC**, with concurrency 16, +unchanged 30-second HTTP deadlines and no retries: + +| Metadata result | Observations | +| --- | ---: | +| HTTP 200 with matching identity | 29,996 | +| HTTP 503 | 4 | +| Total | 30,000 | + +All four failing request IDs matched authentication invocations classified as +`not_started` with a runtime `capacity` error. The bounded classifier reported +the reason as `other`; it did not identify the owner's specific budget. +These observations narrow the investigation, but do not prove that all earlier +503s share a cause. This replay excludes Git/LFS bodies, scheduled throughput +and crash recovery. A failed metadata check remains a failed correctness gate. + +### Bounded authentication candidate remains separate + +A regression at the actual `DirectoryCell.authenticate` call site held its SQL +worker while polling 16 valid authentication requests. The original generic +query contract admitted 15 and refused the sixteenth with `Cell mailbox bytes`. +All ten pre-fix repetitions reproduced that result. Its declared one-MiB result +bound over-reserves credit for an authentication result containing at most one +small principal row. + +The separate candidate adds a typed authentication query with a 36-byte input +bound and 256-byte output bound. It preserves existing query contracts, Cellule +`c51dd121`, runtime budgets, execution-time token expiry and minimum receipts. +The regression admitted all 16 requests without refusal, retaining 4,673 bytes; +all 11 Directory integration tests passed. This is an admission-contract result, +not a measured throughput improvement or a completed live-fleet fix. + +| Candidate verification | Closed result | +| --- | --- | +| Locked release build | Passed; executable retained outside Cargo targets | +| Authentication regression | Passed; 16 admitted, zero refused | +| Directory integration suite | 11 passed | +| Release workspace | **Failed**, closed at 21:09:16 UTC; multi-server suite had 95 passed, one failed and nine ignored | +| Failed case | `paused_cold_repository_does_not_serialize_other_cold_activations` returned HTTP 503 during repository creation; exact failing call and cause remain unresolved | +| Runtime upgrade | Not attempted; predecessor compatibility and old-code restore remain open | + +The original predecessor release descriptor was read from RustFS and its BLAKE3 +digest verified without changing selected state. Reading it is not rolling +compatibility or restore proof. The candidate is **not deployed or included in +this PR**. Its failure cannot be dismissed as test timing, and its passing +focused checks cannot replace full correctness verification. + New closed files are under `/Users/haipingfu/.codex/canopy-three-node-evidence-BwYz7P`, separately from the rebuildable Cargo target: @@ -311,6 +366,22 @@ The separate local diagnostic evidence is under | `canopy-c51-diagnostic-library-tests` | `795b6ff0abc2d1eee01c83d927b8e8b0a2b0827293b2f42acaee8f7e1496980f` | | `test-artifact.json` | `fbdd719e2dd5e0d12dfab935ff470eff5b5f87206354462fd5df97e844ad8ae9` | | `cleanup-reproduction/reproduction.json` | `7e71923d85025177a7b0a273656a4daccded11ba6c9f0e174814cfc4aaec50ed` | +| `instrumented-metadata-replay.json` | `75047e541aeb877966a1e52797eebcf1f762b4c2626bb3b2e6dd0d779cb998c9` | +| `instrumented-metadata-replay.samples.jsonl` | `1bdce950788e730011648c6614f633ee3d1e2f3a435029f37ea36388cd3a2e73` | +| `classification-correlations.json` | `4fd685ec69d3d55f67fc7ca20ee66ec7c74b4f1b7bbac67245f6f72bfed0e382` | + +The separate bounded candidate evidence is under +`/Users/haipingfu/.codex/canopy-authentication-bounded-query-PvyQdK`: + +| Closed candidate artifact | SHA-256 | +| --- | --- | +| `build-tests.json` | `81cb8ecb5be43ac4311ac3ec1356805f87cb8a81c4165626048af6fcfcc23099` | +| `authentication-regression.log` | `79c3114be7d3dd15a4faa56cefda884d749758f803b6afec50e28f12f507dbf1` | +| `directory-correctness.log` | `5210b95ffeec1196be875a18c5f469e738f17e146a7d58295705175610309af4` | +| `release-workspace.log` | `2aea6bf39c8ed8480070a8175b5af45711850d0742f3b9630d26f647b82ca07a` | +| `canopy-bounded-authentication-c51` | `e52ac937bb34d154185b5b6d49268d0ce6b7b39fab6e220478937f7a42472b2d` | +| `failed-workspace-multi-server-tests` | `8a79042dc26364dfd8844624d6ab0b448f65350c04e48cb3a53647aa57baf399` | +| `closed-failure-backup.json` | `9be839c98998ef970c1f25dda53e98e8ddd16958c09150ebf56962dd92f93081` | ### Independent evidence backup @@ -332,12 +403,20 @@ This local preservation does not recover the old missing ledgers or constitute an off-machine backup. Private fleet configs and changing verifier/load receipts are excluded. Their existence or intermediate counts are not passing results. +The closed 30,000-observation replay and correlations also have verified copies +at `/Volumes/Workspace/CrabData/canopy-closed-directory-replay-fk36kinj`. +The failed bounded candidate, its source bindings, exact test executables and +predecessor inputs have 278 verified file copies at +`/Volumes/Workspace/CrabData/canopy-bounded-auth-failure-8v6zmudg`. +Neither backup overwrites original failed receipts or certifies recovery. + | Required gate | Current status | | --- | --- | | Exact new artifact and source provenance | Passed locked build and current binding checks | | Rust workspace correctness and RustFS Git/LFS compatibility | Passed release suite, eight provider gates and initial 17-step proxy check | | New-pin non-sparse 5-GiB transfer | Open; excluded by `--provider-only` | | Full 10,000 identities/100 populated fixtures | Full seed closed; remote verification failed HTTP 503; cause unresolved | +| Bounded authentication candidate | Focused regression and Directory suite passed; full workspace failed; old-code compatibility and live verification open; not in PR | | Original matrix and every new ACK after owner loss | Downstream gates stopped without owner signals or timed load; recovery open | | Critical concurrent Git load and fault coverage | Open; separate from the initial 17-step functional check | | Higher admission profiles and matched comparisons | Open | From e2649bca80ad3b00a05b0d5cf938afd28cdfb2ce Mon Sep 17 00:00:00 2001 From: forhappy Date: Thu, 1 Oct 2026 13:49:12 -0700 Subject: [PATCH 30/60] test: reproduce directory authentication mailbox credit refusal --- .../tests/directory_cell/capacity.rs | 131 ++++++++++++++++++ .../tests/directory_cell/main.rs | 1 + 2 files changed, 132 insertions(+) create mode 100644 crates/canopy-server/tests/directory_cell/capacity.rs diff --git a/crates/canopy-server/tests/directory_cell/capacity.rs b/crates/canopy-server/tests/directory_cell/capacity.rs new file mode 100644 index 0000000..ed2da84 --- /dev/null +++ b/crates/canopy-server/tests/directory_cell/capacity.rs @@ -0,0 +1,131 @@ +//! Deterministic feedback loop at the actual authentication query call site. +//! +//! The held SQL worker controls overlap only. This is not a RustFS throughput +//! benchmark or proof that every observed peer capacity error has this cause. + +use super::*; +use std::{future::Future, task::Poll, time::Duration}; + +struct ReleaseOnDrop(Option>); + +impl ReleaseOnDrop { + fn release(&mut self) { + if let Some(sender) = self.0.take() { + let _ = sender.send(()); + } + } +} + +impl Drop for ReleaseOnDrop { + fn drop(&mut self) { + self.release(); + } +} + +#[tokio::test(flavor = "multi_thread")] +async fn sixteen_overlapping_authentication_queries_fit_the_directory_contract() +-> Result<(), Box> { + let application = Arc::new(CanopyApplication::compile(build_descriptor( + include_bytes!("../../../../Cargo.lock"), + "directory-authentication-capacity-test", + ))?); + let tenant = TenantId::from_bytes([71; 16]); + let application_id = ApplicationId::from_bytes([72; 16]); + let target = directory::directory_target(tenant, application_id)?; + let session = SessionId::from_bytes([73; 16]); + let runtime = runtime(session)?; + let layout = CellStorageLayout::new( + Store::new(Arc::new(InMemory::new())), + StorePath::from("authentication-capacity-test"), + *application_id.as_bytes(), + ); + let files = tempfile::TempDir::new()?; + let handle = bootstrap( + &runtime, + &application.registry(), + &layout, + &target, + (DirectoryModule::NAME, directory::SCHEMA), + session, + &files.path().join("directory.sqlite"), + ) + .await?; + let directory = DirectoryCell::new( + &app_handle(&application, tenant, application_id, handle.clone())?, + target, + )?; + directory + .create_account(random_identity()?, "owner", [1; 32], TokenScope::Admin) + .await?; + assert!( + directory + .authenticate([1; 32], None) + .await? + .output + .is_some() + ); + + let (entered, observed) = tokio::sync::oneshot::channel(); + let (release, released) = std::sync::mpsc::channel(); + let mut release = ReleaseOnDrop(Some(release)); + let blocker = tokio::spawn(async move { + handle + .query(0, 1, move |_| { + let _ = entered.send(()); + released + .recv_timeout(Duration::from_secs(4)) + .map_err(|source| cellule_runtime::Error::Facility { + name: "capacity test release", + source: Box::new(source), + })?; + Ok(Vec::new()) + }) + .await + }); + tokio::time::timeout(Duration::from_secs(3), observed).await??; + + // Poll each actual typed authentication query while the worker is held. + // Immediate capacity refusals stay in the result; there is no retry, + // lowered concurrency, synthetic oversized input or changed runtime limit. + let mut queries = (0..16) + .map(|_| Box::pin(directory.authenticate([1; 32], None))) + .collect::>(); + let mut pending = Vec::new(); + let mut refused = Vec::new(); + let mut premature = 0; + for query in &mut queries { + let polled = std::future::poll_fn(|cx| Poll::Ready(query.as_mut().poll(cx))).await; + match polled { + Poll::Pending => pending.push(query), + Poll::Ready(Err(error)) => refused.push(error), + Poll::Ready(Ok(_)) => premature += 1, + } + } + let admitted = pending.len(); + let retained = runtime.stats().retained_bytes(); + println!( + "authentication overlap: pending={admitted}, premature={premature}, \ + refused={refused:?}, retained_bytes={retained}" + ); + release.release(); + tokio::time::timeout(Duration::from_secs(5), blocker).await???; + for query in pending { + let principal = tokio::time::timeout(Duration::from_secs(5), query) + .await?? + .output + .ok_or("queued authentication lost its valid principal")?; + assert_eq!(principal.account, "owner"); + assert_eq!(principal.scope, TokenScope::Admin); + } + runtime.shutdown().await?; + assert_eq!(premature, 0, "a query bypassed the held FIFO worker"); + assert_eq!( + admitted, 16, + "authentication used generic SQL result credit" + ); + assert!( + refused.is_empty(), + "valid authentication queries were refused" + ); + Ok(()) +} diff --git a/crates/canopy-server/tests/directory_cell/main.rs b/crates/canopy-server/tests/directory_cell/main.rs index a7095ef..1309a57 100644 --- a/crates/canopy-server/tests/directory_cell/main.rs +++ b/crates/canopy-server/tests/directory_cell/main.rs @@ -1,4 +1,5 @@ mod accounts; +mod capacity; mod expiry; #[path = "../support/objects.rs"] mod objects; From 0746e406ebf86e269da6ae02b8354f39968732ac Mon Sep 17 00:00:00 2001 From: forhappy Date: Thu, 1 Oct 2026 13:50:34 -0700 Subject: [PATCH 31/60] perf: bound directory authentication query credit to its result --- crates/canopy-server/src/directory/mod.rs | 19 +++-- .../canopy-server/src/directory/timed_sql.rs | 84 +++++++++++++++++++ 2 files changed, 95 insertions(+), 8 deletions(-) diff --git a/crates/canopy-server/src/directory/mod.rs b/crates/canopy-server/src/directory/mod.rs index e9a782b..167c448 100644 --- a/crates/canopy-server/src/directory/mod.rs +++ b/crates/canopy-server/src/directory/mod.rs @@ -31,7 +31,11 @@ pub const SCHEMA: &str = include_str!("../directory_schema.sql"); pub const REPOSITORY_PAGE_SIZE: usize = 32; const COMMANDS: [OperationDescriptor; 2] = [operation(1), operation(3)]; -const QUERIES: [OperationDescriptor; 2] = [operation(2), operation(4)]; +const QUERIES: [OperationDescriptor; 3] = [ + operation(2), + operation(4), + timed_sql::AUTHENTICATE_OPERATION, +]; const fn operation(id: u32) -> OperationDescriptor { OperationDescriptor { @@ -100,7 +104,8 @@ impl CellModule for DirectoryModule { fn register(self, registry: &mut RegistryBuilder) -> cellule_runtime::Result<()> { register_sql::(registry)?; registry.bind_command::()?; - registry.bind_query::() + registry.bind_query::()?; + registry.bind_query::() } } @@ -305,12 +310,10 @@ impl DirectoryCell { token_digest: [u8; 32], minimum: Option, ) -> Result>, InvocationError>> { - let result = self.credential_query(minimum, SqlBatch { - statements: vec![SqlStatement { - sql: "SELECT a.name, t.scope, t.id FROM access_tokens AS t JOIN accounts AS a ON a.name = t.account WHERE t.digest = ?2 AND t.enabled = 1 AND (t.expires_ms IS NULL OR t.expires_ms > ?1) AND a.enabled = 1".into(), - parameters: vec![SqlValue::Blob(token_digest.to_vec())], - }], - }).await?; + let result = self + .application + .query::(&self.target, minimum, token_digest.to_vec()) + .await?; let principal = result .output .first() diff --git a/crates/canopy-server/src/directory/timed_sql.rs b/crates/canopy-server/src/directory/timed_sql.rs index 47ae6d6..f47bbb7 100644 --- a/crates/canopy-server/src/directory/timed_sql.rs +++ b/crates/canopy-server/src/directory/timed_sql.rs @@ -3,6 +3,48 @@ use cellule_runtime::{ Command, Query, registry::CommandContext, registry::CommandResult, registry::QueryContext, }; +// Authentication returns at most one row: a validated 64-byte account name, +// a five-byte scope and a 16-byte token ID. Do not reserve the generic 1-MiB +// SQL result ceiling for every small credential decision. Generic credential +// pages retain their existing operation IDs and bounds. +pub(super) const AUTHENTICATE_OPERATION: OperationDescriptor = OperationDescriptor { + id: 5, + codec_version: 1, + schema_min: 1, + schema_max: 1, + input_limit: 36, // Canonical Vec: four-byte length plus SHA-256 digest. + output_limit: 256, +}; + +pub(super) struct AuthenticateQuery; + +impl Query for AuthenticateQuery { + const MODULE: &'static str = DirectoryModule::NAME; + const ID: u32 = AUTHENTICATE_OPERATION.id; + const CODEC_VERSION: u32 = 1; + type Input = Vec; + type Output = Vec; + + fn execute( + context: &mut QueryContext<'_>, + input: Self::Input, + ) -> cellule_runtime::Result { + if input.len() != 32 { + return Err(Error::Command("invalid authentication digest length")); + } + let batch = bind_time( + SqlBatch { + statements: vec![SqlStatement { + sql: "SELECT a.name, t.scope, t.id FROM access_tokens AS t JOIN accounts AS a ON a.name = t.account WHERE t.digest = ?2 AND t.enabled = 1 AND (t.expires_ms IS NULL OR t.expires_ms > ?1) AND a.enabled = 1".into(), + parameters: vec![SqlValue::Blob(input)], + }], + }, + context.now_ms(), + )?; + context.sql(&batch) + } +} + // Credential decisions use one owner timestamp for the whole transaction. // Cellule samples context time before queueing; refresh it to fence expired // requests without racing separate decision/update statements. @@ -79,3 +121,45 @@ impl DirectoryCell { .await } } + +#[cfg(test)] +mod tests { + use super::*; + use cellule_runtime::codec::{BoundedDecoder, BoundedEncoder, WireValue}; + + #[test] + fn authentication_bounds_cover_the_largest_valid_principal() + -> Result<(), Box> { + let digest = vec![7; 32]; + let mut input = BoundedEncoder::new(AUTHENTICATE_OPERATION.input_limit)?; + digest.encode(&mut input)?; + let encoded = input.finish(); + assert_eq!(encoded.len(), 36); + assert_eq!( + Vec::::decode(&mut BoundedDecoder::new(&encoded, 36)?)?, + digest + ); + assert!(vec![7_u8; 33].encode(&mut BoundedEncoder::new(36)?).is_err()); + + let name = "a".repeat(64); + validate_component(&name)?; + let result = vec![SqlResultSet { + columns: vec!["name".into(), "scope".into(), "id".into()], + rows: vec![vec![ + SqlValue::Text(name), + SqlValue::Text(TokenScope::Admin.as_str().into()), + SqlValue::Blob(vec![8; 16]), + ]], + rows_affected: 0, + }]; + let mut output = BoundedEncoder::new(AUTHENTICATE_OPERATION.output_limit)?; + result.encode(&mut output)?; + let encoded = output.finish(); + assert!(encoded.len() <= 256); + assert_eq!( + Vec::::decode(&mut BoundedDecoder::new(&encoded, 256)?)?, + result + ); + Ok(()) + } +} From cc7130f894a8162d5f1b808fb970197080f5568d Mon Sep 17 00:00:00 2001 From: forhappy Date: Thu, 1 Oct 2026 13:01:33 -0700 Subject: [PATCH 32/60] test: isolate parent fence checks and cover inherited fork safety --- .../src/git_http/stream_tests.rs | 133 ++++++++++++++++++ 1 file changed, 133 insertions(+) diff --git a/crates/canopy-server/src/git_http/stream_tests.rs b/crates/canopy-server/src/git_http/stream_tests.rs index 165fb50..c0619b4 100644 --- a/crates/canopy-server/src/git_http/stream_tests.rs +++ b/crates/canopy-server/src/git_http/stream_tests.rs @@ -211,6 +211,24 @@ async fn completed_worker_releases_cache_before_headers_are_polled() #[tokio::test] async fn failed_spawn_releases_parent_fence_before_cache_cleanup() -> Result<(), Box> { + const ISOLATED: &str = "CANOPY_TEST_ISOLATED_FAILED_SPAWN"; + if std::env::var(ISOLATED).as_deref() != Ok("1") { + // This checks the command's parent fence, not unrelated forks that can + // briefly inherit its CLOEXEC descriptor. An inherited live fence must + // prevent cleanup; exercise that case separately below. + let status = Command::new(std::env::current_exe()?) + .args([ + "--exact", + "git_http::stream_tests::failed_spawn_releases_parent_fence_before_cache_cleanup", + "--test-threads=1", + "--nocapture", + ]) + .env(ISOLATED, "1") + .status() + .await?; + assert!(status.success()); + return Ok(()); + } let files = tempfile::TempDir::new()?; let budget = DiskBudget::new(1 << 20); let cache = GitCache::create( @@ -229,3 +247,118 @@ async fn failed_spawn_releases_parent_fence_before_cache_cleanup() assert_eq!(budget.used(), 0); Ok(()) } + +// A concurrent fork can inherit the cache fence until it execs, even though +// CLOEXEC remains set in the parent. Failed-spawn cleanup must not undercount +// or delete such a generation; conservative quarantine lasts until restart. +#[cfg(unix)] +#[tokio::test] +async fn inherited_fork_fence_keeps_failed_spawn_cache_charged() +-> Result<(), Box> { + use std::{ + io::{Read, Write}, + os::unix::{io::AsRawFd, net::UnixStream, process::CommandExt}, + thread::JoinHandle, + }; + + struct ForkBarrier { + control: UnixStream, + thread: Option>>, + } + impl ForkBarrier { + fn finish(&mut self) -> Result<(), Box> { + self.control.write_all(b"X")?; + let status = self + .thread + .take() + .ok_or("missing helper thread")? + .join() + .map_err(|_| "helper thread panicked")??; + assert!(status.success()); + Ok(()) + } + } + impl Drop for ForkBarrier { + fn drop(&mut self) { + if let Some(thread) = self.thread.take() { + let _ = self.control.write_all(b"X"); + let _ = thread.join(); + } + } + } + + let files = tempfile::TempDir::new()?; + let budget = DiskBudget::new(1 << 20); + let cache = GitCache::create( + files.path().into(), + budget.clone(), + "refs/heads/main", + crate::ObjectFormat::Sha1, + ) + .await?; + let charged = budget.used(); + assert!(charged > 0); + let git_dir = cache.git_dir(); + let mut command = crate::native_git::command(&git_dir)?; + command.current_dir(files.path().join("missing")); + + let (control, child_control) = UnixStream::pair()?; + control.set_read_timeout(Some(Duration::from_secs(5)))?; + let executable = std::env::current_exe()?; + let thread = std::thread::spawn(move || { + let mut unrelated = std::process::Command::new(executable); + unrelated + .arg("--help") + .stdin(std::process::Stdio::null()) + .stdout(std::process::Stdio::null()) + .stderr(std::process::Stdio::null()); + // SAFETY: the child uses only async-signal-safe read/write, with an + // owned live socket and stack/static byte buffers, before exec. Holding + // this barrier models a concurrent fork's inherited CLOEXEC descriptors. + unsafe { + unrelated.pre_exec(move || { + let fd = child_control.as_raw_fd(); + let mut release = 0_u8; + if libc::write(fd, b"R".as_ptr().cast(), 1) != 1 + || libc::read(fd, (&mut release as *mut u8).cast(), 1) != 1 + || release != b'X' + { + return Err(std::io::Error::from_raw_os_error(libc::EIO)); + } + Ok(()) + }); + } + unrelated.spawn()?.wait() + }); + let mut barrier = ForkBarrier { + control, + thread: Some(thread), + }; + let mut ready = [0_u8]; + barrier.control.read_exact(&mut ready)?; + assert_eq!(ready, *b"R"); + assert!(matches!( + GitProcess::spawn(command, cache), + Err(GitHttpError::Io(_)) + )); + assert_eq!(budget.used(), charged); + assert!( + git_dir.exists(), + "live inherited fence must prevent cache deletion" + ); + let busy = + crate::native_git::idle_fence(&git_dir).expect_err("inherited fence should still be held"); + assert_eq!(busy.kind(), std::io::ErrorKind::WouldBlock); + barrier.finish()?; + drop(crate::native_git::idle_fence(&git_dir)?); + // Drop quarantined the generation while its fence was busy. Closing the + // inherited descriptor later must not silently release its disk charge + // while files remain. Startup recovery reclaims quarantined generations. + assert_eq!( + budget.used(), + charged, + "quarantined files must stay charged until they are reclaimed" + ); + assert!(git_dir.exists()); + Ok(()) +} From 87889f4ccde5f6a30d7fb146b2e6871c027eb500 Mon Sep 17 00:00:00 2001 From: forhappy Date: Thu, 1 Oct 2026 15:17:05 -0700 Subject: [PATCH 33/60] fix: retain exact directory predecessor for bounded authentication --- crates/canopy-server/src/directory/mod.rs | 28 +++++++-- .../canopy-server/src/directory/timed_sql.rs | 6 +- .../tests/directory_cell/compatibility.rs | 61 +++++++++++++++++++ .../fixtures/c51-selected-release.json | 1 + .../tests/directory_cell/main.rs | 1 + 5 files changed, 91 insertions(+), 6 deletions(-) create mode 100644 crates/canopy-server/tests/directory_cell/compatibility.rs create mode 100644 crates/canopy-server/tests/directory_cell/fixtures/c51-selected-release.json diff --git a/crates/canopy-server/src/directory/mod.rs b/crates/canopy-server/src/directory/mod.rs index 167c448..bdf057a 100644 --- a/crates/canopy-server/src/directory/mod.rs +++ b/crates/canopy-server/src/directory/mod.rs @@ -18,10 +18,15 @@ use cellule_app::{ApplicationHandle, CellType}; use cellule_runtime::{ ApplicationId, CellModule, CellTarget, Committed, Digest, Error, InvocationError, MigrationDescriptor, ModuleDescriptor, MutationIdentity, NamespaceDescriptor, NamespaceId, - Observed, Receipt, RegistryBuilder, SqlCell, SqlModule, TenantId, cell::catalog::CatalogRole, - partition_for_shard, primitives::sql::SqlBatch, primitives::sql::SqlResultSet, - primitives::sql::SqlStatement, primitives::sql::SqlValue, primitives::sql::register_sql, - registry::OperationDescriptor, + Observed, Receipt, RegistryBuilder, SqlCell, SqlModule, TenantId, + cell::catalog::CatalogRole, + partition_for_shard, + primitives::sql::SqlBatch, + primitives::sql::SqlResultSet, + primitives::sql::SqlStatement, + primitives::sql::SqlValue, + primitives::sql::register_sql, + registry::{OperationDescriptor, RetainedCodeDescriptor}, }; use crate::{CanopyApplication, ReadIdentity, validate_repository_id}; @@ -37,6 +42,19 @@ const QUERIES: [OperationDescriptor; 3] = [ timed_sql::AUTHENTICATE_OPERATION, ]; +// The selected c51 release used the same schema and credential contracts, +// before the separately bounded authentication query was added. Keep that +// exact code executable while persisted Cells roll to the new descriptor. +const RETAINED_CODES: [RetainedCodeDescriptor; 1] = [RetainedCodeDescriptor { + code: Digest::from_bytes([ + 0xf7, 0x25, 0x4e, 0xda, 0x9d, 0x5d, 0x33, 0x95, 0x66, 0xf4, 0x54, 0x57, 0x50, 0x26, 0x18, + 0xad, 0x13, 0xcb, 0xbf, 0x6e, 0x5a, 0x74, 0x59, 0x5f, 0x5b, 0x3c, 0xe4, 0x66, 0x53, 0xea, + 0x12, 0xf1, + ]), + schema_min: 1, + schema_max: 1, +}]; + const fn operation(id: u32) -> OperationDescriptor { OperationDescriptor { id, @@ -76,7 +94,7 @@ impl CellModule for DirectoryModule { source.update(SCHEMA.as_bytes()); Digest::from_bytes(*source.finalize().as_bytes()) }, - retained_codes: &[], + retained_codes: &RETAINED_CODES, schema_min: 1, schema_max: 1, migrations: MIGRATIONS.get_or_init(|| { diff --git a/crates/canopy-server/src/directory/timed_sql.rs b/crates/canopy-server/src/directory/timed_sql.rs index f47bbb7..609c29b 100644 --- a/crates/canopy-server/src/directory/timed_sql.rs +++ b/crates/canopy-server/src/directory/timed_sql.rs @@ -139,7 +139,11 @@ mod tests { Vec::::decode(&mut BoundedDecoder::new(&encoded, 36)?)?, digest ); - assert!(vec![7_u8; 33].encode(&mut BoundedEncoder::new(36)?).is_err()); + assert!( + vec![7_u8; 33] + .encode(&mut BoundedEncoder::new(36)?) + .is_err() + ); let name = "a".repeat(64); validate_component(&name)?; diff --git a/crates/canopy-server/tests/directory_cell/compatibility.rs b/crates/canopy-server/tests/directory_cell/compatibility.rs new file mode 100644 index 0000000..c8d37db --- /dev/null +++ b/crates/canopy-server/tests/directory_cell/compatibility.rs @@ -0,0 +1,61 @@ +use super::*; +use cellule_runtime::Digest; + +// Captured read-only from the unchanged RustFS deployment. The canonical +// descriptor's SHA-256 is 8b85c842995e4a2b0bcc6d1f3d361ccdbd2c33a1c52592241b50ad8cfe9aeb88. +// This is contract compatibility, not proof that old Cells have been restored. +#[test] +fn bounded_authentication_retains_the_exact_selected_predecessor() +-> Result<(), Box> { + let previous = include_bytes!("fixtures/c51-selected-release.json").trim_ascii_end(); + assert_eq!( + blake3::hash(previous).to_hex().as_str(), + "e31bf1a951e2fa19d91e9f964b2ddeade1a81b05a20ad628362819a1487c16b1" + ); + let application = CanopyApplication::compile(build_descriptor( + include_bytes!("../../../../Cargo.lock"), + "directory-compatibility-test", + ))?; + let registry = application.registry(); + registry.verify_rolling_from(previous)?; + let predecessor: serde_json::Value = serde_json::from_slice(previous)?; + let old_directory = predecessor["modules"] + .as_array() + .ok_or("modules missing")? + .iter() + .find(|module| module["name"] == "directory") + .ok_or("directory missing")?; + let bytes: [u8; 32] = hex::decode(old_directory["code"].as_str().ok_or("code missing")?)? + .try_into() + .map_err(|_| "invalid predecessor code length")?; + let old_code = Digest::from_bytes(bytes); + assert_ne!(Some(old_code), registry.module_code("directory")); + assert!(registry.supports_cell(directory::DIRECTORY, CatalogRole::Sql, old_code, 1)); + for schema in [0, 2] { + assert!(!registry.supports_cell(directory::DIRECTORY, CatalogRole::Sql, old_code, schema)); + } + assert!(!registry.supports_cell( + directory::DIRECTORY, + CatalogRole::Sql, + Digest::from_bytes([1; 32]), + 1 + )); + let current: serde_json::Value = serde_json::from_slice(registry.release_bytes())?; + let module = current["modules"] + .as_array() + .ok_or("modules missing")? + .iter() + .find(|module| module["name"] == "directory") + .ok_or("directory missing")?; + let queries = module["queries"].as_array().ok_or("queries missing")?; + assert_eq!( + queries + .iter() + .map(|operation| operation["id"].as_u64().unwrap()) + .collect::>(), + vec![2, 4, 5] + ); + assert_eq!(queries[2]["input_limit"], 36); + assert_eq!(queries[2]["output_limit"], 256); + Ok(()) +} diff --git a/crates/canopy-server/tests/directory_cell/fixtures/c51-selected-release.json b/crates/canopy-server/tests/directory_cell/fixtures/c51-selected-release.json new file mode 100644 index 0000000..da34d20 --- /dev/null +++ b/crates/canopy-server/tests/directory_cell/fixtures/c51-selected-release.json @@ -0,0 +1 @@ +{"build":{"cargo_lock_digest":"275fecc12d547cafb0285ac39a92c20c1ee169fce125dfc131da6288f7aecc1c","source_revision":"0.1.0"},"modules":[{"activities":[],"code":"f7254eda9d5d339566f45457502618ad13cbbf6e5a74595f5b3ce46653ea12f1","commands":[{"codec":1,"id":1,"input_limit":1048576,"output_limit":1048576,"schema_max":1,"schema_min":1},{"codec":1,"id":3,"input_limit":1048576,"output_limit":1048576,"schema_max":1,"schema_min":1}],"migrations":[{"digest":"46bc0b8cd751045d6f02bdeacbe5d5d0454bd779868821f8ef2f310b86b6532a","version":1}],"name":"directory","queries":[{"codec":1,"id":2,"input_limit":1048576,"output_limit":1048576,"schema_max":1,"schema_min":1},{"codec":1,"id":4,"input_limit":1048576,"output_limit":1048576,"schema_max":1,"schema_min":1}],"retained_codes":[],"schema_max":1,"schema_min":1,"source_digest":"eb226223e5c2dd0ad47bb36cb608e5b998d7553486b661e78d1840312d1d544c","workflows":[]},{"activities":[],"code":"9e2427d83565ff9729af3fe116e70616c3d949c29ff87f8804232c14debc5c03","commands":[{"codec":1,"id":1,"input_limit":1048576,"output_limit":1048576,"schema_max":1,"schema_min":1},{"codec":2,"id":6,"input_limit":1048576,"output_limit":1048576,"schema_max":1,"schema_min":1},{"codec":2,"id":7,"input_limit":1048576,"output_limit":1048576,"schema_max":1,"schema_min":1},{"codec":2,"id":8,"input_limit":1048576,"output_limit":1048576,"schema_max":1,"schema_min":1},{"codec":2,"id":10,"input_limit":1048576,"output_limit":1048576,"schema_max":1,"schema_min":1},{"codec":3,"id":5,"input_limit":4194304,"output_limit":1048576,"schema_max":1,"schema_min":1},{"codec":4,"id":3,"input_limit":1048576,"output_limit":1048576,"schema_max":1,"schema_min":1},{"codec":4,"id":9,"input_limit":1048576,"output_limit":1048576,"schema_max":1,"schema_min":1},{"codec":6,"id":4,"input_limit":1048576,"output_limit":1048576,"schema_max":1,"schema_min":1}],"migrations":[{"digest":"ef03f54a27c9b1f483dde597ce61bcaa2aef0a575c3cee5c75f1eea5c6bd8927","version":1}],"name":"repository","queries":[{"codec":1,"id":2,"input_limit":1048576,"output_limit":1048576,"schema_max":1,"schema_min":1}],"retained_codes":[],"schema_max":1,"schema_min":1,"source_digest":"4eba2d3c793ecc1acf52412e9ef0d8a4c071187b5a1a45233b7c0576017757b2","workflows":[]}],"namespaces":[{"dead_letter":null,"effect_targets":[],"id":"47474747474747474747474747474747","module":"repository","name":"repository","role":"sql","shards":1},{"dead_letter":null,"effect_targets":[],"id":"48484848484848484848484848484848","module":"directory","name":"directory","role":"sql","shards":1}],"peer_versions":[1],"runtime":"cellule","version":1} diff --git a/crates/canopy-server/tests/directory_cell/main.rs b/crates/canopy-server/tests/directory_cell/main.rs index 1309a57..2b04f8c 100644 --- a/crates/canopy-server/tests/directory_cell/main.rs +++ b/crates/canopy-server/tests/directory_cell/main.rs @@ -1,5 +1,6 @@ mod accounts; mod capacity; +mod compatibility; mod expiry; #[path = "../support/objects.rs"] mod objects; From b32afeae492cc21553fff73730372239132462d0 Mon Sep 17 00:00:00 2001 From: forhappy Date: Thu, 1 Oct 2026 14:55:56 -0700 Subject: [PATCH 34/60] perf: rescan idle inventory briefly instead of refusing settling residents --- .../canopy-server/src/server/residency/mod.rs | 40 +++++++++++++++---- 1 file changed, 32 insertions(+), 8 deletions(-) diff --git a/crates/canopy-server/src/server/residency/mod.rs b/crates/canopy-server/src/server/residency/mod.rs index 13dfb68..7a50ca4 100644 --- a/crates/canopy-server/src/server/residency/mod.rs +++ b/crates/canopy-server/src/server/residency/mod.rs @@ -354,6 +354,8 @@ impl RepositoryManager { // request instead of turning one transient preflight race into HTTP 503. // Bound rescans so a whole busy working set still backpressures callers. let mut rejected = HashSet::new(); + let settle_deadline = Instant::now() + std::time::Duration::from_millis(200); + let mut settle_rescans = 0; loop { let candidates: HashMap<_, _> = self .node @@ -362,7 +364,7 @@ impl RepositoryManager { .into_iter() .map(|(cell, generation, _, _)| (cell, generation)) .collect(); - let (id, action, _transition) = { + let (chosen, may_settle) = { let mut loaded = self.loaded.lock().await; let mut eligible = Vec::new(); for (id, repository) in loaded.iter() { @@ -406,13 +408,35 @@ impl RepositoryManager { chosen = Some((id, action, transition)); break; } - chosen.ok_or_else(|| { - if rejected.is_empty() { - Error::Capacity("repository residency") - } else { - Error::CellDraining - } - })? + let may_settle = loaded.iter().any(|(id, repository)| { + !rejected.contains(id) + && repository.local + && Arc::strong_count(&repository.pin) == 1 + && matches!( + repository.state, + ResidencyState::Serving | ResidencyState::RefreshHandle + ) + }); + (chosen, may_settle) + }; + let Some((id, action, _transition)) = chosen else { + // Publication and renewal may still be settling just after an + // acknowledged request. Reobserve inventory only: never replay + // a mutation or release a Cell the runtime considers busy. + // Keep every request/residency permit charged while waiting, + // and do not wait at all when every resident is request-pinned. + let remaining = settle_deadline.saturating_duration_since(Instant::now()); + if may_settle && settle_rescans < 8 && !remaining.is_zero() { + settle_rescans += 1; + tokio::time::sleep(remaining.min(std::time::Duration::from_millis(25))).await; + continue; + } + return Err(if rejected.is_empty() { + Error::Capacity("repository residency") + } else { + Error::CellDraining + } + .into()); }; let (cell, generation) = match action { EvictionAction::DropRemote => { From eaddeef4c40c36778b829b6a46fe7c42cc5d5993 Mon Sep 17 00:00:00 2001 From: forhappy Date: Thu, 1 Oct 2026 15:26:12 -0700 Subject: [PATCH 35/60] docs: publish verified authentication and residency candidate checkpoint --- docs/contracts.md | 24 ++- docs/delivery-plan.md | 5 +- docs/performance-plan.md | 15 +- ...01-bounded-authentication-and-residency.md | 161 ++++++++++++++++++ .../2026-10-01-evidence-loss-and-c51dd121.md | 23 ++- 5 files changed, 212 insertions(+), 16 deletions(-) create mode 100644 docs/performance/2026-10-01-bounded-authentication-and-residency.md diff --git a/docs/contracts.md b/docs/contracts.md index 35dab29..f94e99f 100644 --- a/docs/contracts.md +++ b/docs/contracts.md @@ -448,6 +448,13 @@ starts. Requests for that candidate wait for the transition and recheck its serving state. Initialization must succeed before the ready route is exposed. Remote routes retain authoritative owner checks on the transition path. +If inventory has no safe eviction candidate but an unrejected, unpinned local +resident may settle, cold admission can reobserve inventory up to eight times +with at most 200 ms of settling waits. The registry lock is released while +waiting and request/residency permits stay charged. It never replays mutations +or releases runtime-busy Cells. A fully request-pinned working set refuses +immediately; the settling bound is not an HTTP or storage-I/O latency guarantee. + At most 32 transition operations may be executing or waiting on that path. One authenticated account may hold at most 16 of those slots, across all its tokens and repository endpoints. Anonymous readers share a separate 16-slot @@ -1187,7 +1194,7 @@ require a fresh development storage prefix; there is no upgrade reader for older development databases. The module descriptor and object paths will become compatibility boundaries at the first persistent preview. The current source pins Cellule revision -`a4500add51764fa0415791aefbfa561db6ada203`; running qualification artifacts +`c51dd121284ecc8878b75d32717a4dfbe2c406c2`; qualification artifacts remain bound to their recorded revisions, not silently replaced by this pin. The earlier entity-partition cutover also made pre-cutover prefixes incompatible; no migration is available. @@ -1474,7 +1481,7 @@ storage. Historical credentials and runtime command outcomes remain retained; account-count admission, request-rate controls, receipt retention and deployment storage quotas still need their own policies. -Directory credential command 3 and query 4 bind parameter 1 to one execution +Directory credential command 3 and queries 4/5 bind parameter 1 to one execution timestamp for the entire SQL operation. They use the greater of owner wall time and Cellule admission time. The pinned runtime captures context time before queuing, so the handler samples the clock again after queued work. Commands @@ -1482,6 +1489,14 @@ still persist their replay outcome through Cellule; replay does not rerun a successful mutation or reactivate an expired record. Normal SQL operations 1/2 remain for Directory operations that do not make credential-time decisions. +Authentication uses typed query 5, codec 1, with a 36-byte canonical input limit +and a 256-byte result limit. Its input is exactly one 32-byte token digest; the +result contains at most one validated principal. Existing commands 1/3 and +queries 2/4 keep their one-MiB contracts. Admission budgets and minimum receipts +are unchanged. The Directory descriptor retains the exact preceding `f7254eda` +code at schema 1; this is descriptor compatibility, not an automatic upgrade or +proof of restoring an existing catalog under a new release. + Authentication requires `expires_ms IS NULL OR expires_ms > now_ms`; it does not wait for a cleanup job. Token listing/issuance/revocation, account creation and account disablement use the same expiry rule for their exact authorizing digest. @@ -2480,8 +2495,9 @@ Stock `git lfs pull` without credentials verifies this behavior. Both Directory and Repository initialization schemas changed in this unreleased build. Use a fresh development prefix; existing data is not migrated or removed. -Source digest validation rejects old modules; mixed-build rolling upgrades are -not supported. Production upgrade migration remains a delivery gate. +Source digest validation rejects unsupported module codes. Exact Directory +predecessor retention is a contract check, not a migration. Mixed-build rolling +upgrades are not supported. Production upgrade migration remains a delivery gate. ## Deployment release enrollment and maintenance diff --git a/docs/delivery-plan.md b/docs/delivery-plan.md index ee2c60d..6e6cb78 100644 --- a/docs/delivery-plan.md +++ b/docs/delivery-plan.md @@ -72,12 +72,15 @@ These are the next implementation slices. Each slice must update its contract, b 1. Expand the S3-compatible process smoke into a node/lease fault matrix and test the target production object store. Run the checked-in CI workflow on - a Canopy remote. The current source pins Cellule revision `a4500add` and + a Canopy remote. The current source pins Cellule revision `c51dd121` and owns a startup storage probe with cleanup on failure. Historical dependency qualification runs below remain evidence for their recorded revisions only. The [live three-node trial](performance/2026-10-01-native-filesystem.md) retains its separately bound `e07670e` executable; latest source is not a claim that the running artifact was rebuilt or performance-qualified. + The [bounded authentication and residency candidate](performance/2026-10-01-bounded-authentication-and-residency.md) + passed release correctness and eight RustFS compatibility gates. Existing + catalog upgrade, full-corpus recovery and performance remain open. 2. Execute the [repository-density performance plan](performance-plan.md), then qualify the bounded Linux deployment and residency under faults and larger hot sets. Each node reserves one SQL slot for Directory takeover and admits diff --git a/docs/performance-plan.md b/docs/performance-plan.md index 7427dd7..fa415eb 100644 --- a/docs/performance-plan.md +++ b/docs/performance-plan.md @@ -6,7 +6,7 @@ dependency pin is `c51dd121284ecc8878b75d32717a4dfbe2c406c2`, observed upstream main at build preparation on October 1. Unlike the earlier `a4500add` documentation advance, this changes routing and resource-permit implementation. Locked metadata resolves all six packages; the lockfile changes only their sources. -The new locked release build passed. Its release workspace suite passed 227 +The original rebuilt artifact passed its locked release build and 227 release tests with 9 ignored gates; all eight real-RustFS compatibility gates subsequently passed, as did the production binary's initial 17-step check through three nodes and a proxy. The 84 Python harness tests cover accounting and guards. Full-corpus @@ -15,13 +15,21 @@ of these functional checks establishes reference capacity. Upstream subsequently advanced to `0dc04a6` with atomic API and test changes. That revision needs a separate artifact and runtime qualification; the current experiment and its frozen source remain on `c51dd121`. +The [combined authentication and residency candidate](performance/2026-10-01-bounded-authentication-and-residency.md) +is now in PR #18: 231 release tests and all eight RustFS gates passed, as did +20 unchanged cold-activation repetitions and all 15 residency tests. It has not +been deployed on the existing corpus. Descriptor compatibility passed, but +existing-catalog admission, old-code restore and the diagnostic fleet's terminal +lease-fencing failure remain open. These results do not qualify newer upstream +or replace the original complete performance schedule. The [three-node proxy campaign](performance/2026-09-30-three-node-proxy.md) retains the earlier fixture race and its regression, plus the candidate's failed seed and successful recovery of every recorded ACK. | Current gate | Evidence / status | | --- | --- | -| Latest Cellule candidate | `c51dd121` build, 227 Rust tests, eight RustFS gates and initial 17-step proxy check passed; full 10K/100 seed closed, but remote verification failed HTTP 503. A separate 10K metadata diagnostic returned 9,999 matching identities and one 503; downstream fault/load gates stopped without owner signals. Cause unresolved | +| Original rebuilt Cellule artifact | `c51dd121` build, 227 Rust tests, eight RustFS gates and initial 17-step proxy check passed; full 10K/100 seed closed, but remote verification failed HTTP 503. A separate 10K metadata diagnostic returned 9,999 matching identities and one 503; downstream fault/load gates stopped without owner signals. Cause unresolved | +| Combined authentication and residency candidate | 231 release tests, eight RustFS gates, 20 original cold-activation repetitions and 15 residency tests passed. Exact predecessor descriptor is retained; existing-catalog upgrade and old-code restore remain open. Not deployed or performance-qualified | | Candidate Git behavior | Bound `e07670e` production artifact passed all 17 critical stock-Git steps through three nodes/proxy against RustFS | | Mac-backed candidate seed | Failed: 3,561/10,000 identities and 39 Git/LFS fixtures; incomplete manifest retained | | Every candidate seed ACK | All 3,561 identities and 39 populated fixtures passed fresh-owner recovery; independent ledger/binding and 78-clone audit passed | @@ -43,7 +51,8 @@ The separate diagnostic build passed 230 release workspace tests with 9 ignored after isolating a concurrent-fork cleanup test and adding fence-safety coverage. All 20 subsequent full-library repetitions passed at four test threads. That local test finding does not explain the Directory 503, change production -cleanup or qualify the PR's frozen artifact; temporary diagnostics remain outside +cleanup or qualify its original frozen artifact. The cleanup-test correction is +included in the combined candidate; temporary diagnostic logging stays outside the PR. No Cellule bottleneck or matched performance improvement is proven. The 4-GiB diagnostic is a new resource envelope, not a passing 2-GiB result. Offline plans diff --git a/docs/performance/2026-10-01-bounded-authentication-and-residency.md b/docs/performance/2026-10-01-bounded-authentication-and-residency.md new file mode 100644 index 0000000..73cf8e7 --- /dev/null +++ b/docs/performance/2026-10-01-bounded-authentication-and-residency.md @@ -0,0 +1,161 @@ +# Bounded authentication and cold repository admission + +The combined candidate passes release correctness, all eight isolated RustFS +compatibility gates and repeated cold-activation regressions. It is now included +in PR #18 without temporary diagnostic logging. It has **not** been deployed +against the existing 10,000-repository corpus. Upgrade, full recovery and +performance qualification remain open; no matched speedup is claimed. + +This follows the [failed authentication candidate and corpus verification](2026-10-01-evidence-loss-and-c51dd121.md). +Their failures remain intact. The target is still three nodes behind a proxy, +10,000 identities, 100 populated Git/LFS fixtures and all 108 load windows: +114,960 arrivals over 8,640 scheduled seconds. + +## Authentication credit matches the result + +The actual `DirectoryCell.authenticate` regression holds the SQL worker while +polling 16 valid requests. All ten pre-fix repetitions admitted 15 and refused +the sixteenth with `Cell mailbox bytes`. This reproduces one admission +bottleneck; it does not prove every historical or live HTTP 503 has that cause. + +| Contract | Before | Combined candidate | +| --- | --- | --- | +| Authentication query | Generic credential query 4 | Typed query 5, codec 1 | +| Canonical input limit | 1 MiB | 36 bytes: length plus 32-byte token digest | +| Output limit | 1 MiB | 256 bytes: at most one validated principal row | +| Admitted in held-worker regression | 15 of 16 | 16 of 16, zero refusals | +| Retained runtime bytes in regression | 15,732,331 | 4,673 | + +Commands 1/3 and queries 2/4 retain their existing contracts. Runtime budgets, +FIFO execution, minimum receipts and principal validation are unchanged. Query 5 +binds expiry at SQL execution time, so queued work cannot extend token authority. +The codec test covers the largest valid account name, scope and token ID. +This is a memory-admission result, not an end-to-end latency measurement. + +## Cold admission can reobserve settling residents + +Replaying the exact failed executable reproduced five failures in 20 attempts +of `paused_cold_repository_does_not_serialize_other_cold_activations`: two HTTP +503s and three waits that never reached the expected store pause. Refusal probes +found local, serving residents without request pins but no runtime-eligible +eviction candidate. A timed-out request had already returned HTTP 503 before +reaching the injected pause. The admission seam is established; every internal +reason for ineligibility is not. + +The fix reobserves idle inventory when an unrejected, unpinned local resident +may settle. It sleeps at most eight times, up to 25 ms each, under a 200-ms +settling deadline. Admission permits remain charged and the registry lock is +dropped before waiting. Request-pinned working sets still refuse immediately. + +```mermaid +flowchart TD + A[Cold request needs a resident slot] --> B{Safe candidate available?} + B -->|Yes| C[Existing generation and release checks] + B -->|No| D{Unpinned local resident may settle?} + D -->|No| E[Capacity refusal] + D -->|Yes, within budget| F[Wait up to 25 ms without registry lock] + F --> B + D -->|Budget exhausted| E + C --> G[Release or refresh through existing transition] +``` + +No mutation is replayed or runtime-busy Cell forcibly released. Pin, generation, +movement and capacity checks remain in force. The 200-ms bound is for settling +waits, not an HTTP latency guarantee or a bound on storage I/O. + +## Predecessor compatibility is not a completed upgrade + +Before retention, the compiled registry rejected the stored predecessor with +`rolling release does not retain predecessor module code`. The candidate retains +only Directory code +`f7254eda9d5d339566f45457502618ad13cbbf6e5a74595f5b3ce46653ea12f1` +at schema 1. The fixture is the exact canonical descriptor read from RustFS, +with BLAKE3 digest +`e31bf1a951e2fa19d91e9f964b2ddeade1a81b05a20ad628362819a1487c16b1`. +Its test verifies rolling descriptor admission and rejects unsupported schemas +and unknown code. It does not open old persisted Cells. + +Startup still requires the exact selected release. There is no general upgrade +controller; no selected release or catalog was changed during these checks. +Source inspection identifies another unresolved upgrade risk: `acquire_sql_cell` +provisions the current module code, whereas an existing catalog's initial code +is immutable. A call-site reproduction and actual old-code restore are required. +Do not bypass release selection, rewrite catalog identity or treat a low-level +activation CAS as compatibility proof. + +## Closed verification results + +The frozen source is local commit `392a343b83a492d49c08a6f3e83ff4dd2917437f`. +The PR production source, tests, scripts and Cargo files are byte-identical to +that candidate. All six Cellule packages remain locked to +`c51dd121284ecc8878b75d32717a4dfbe2c406c2`. Upstream main was rechecked as +`0dc04a658bd99668936f7ec58032d054f6fbc141`; separate qualification is required. + +| Check | Closed result | Scope limit | +| --- | --- | --- | +| Locked release build | Passed; executable retained outside Cargo targets | Not live deployment | +| Authentication regression | 16 admitted, zero refused, 4,673 retained bytes | Held-worker overlap, not throughput | +| Directory integration | 12 passed | Includes descriptor admission, not old-code restore | +| Release workspace | 231 top-level tests passed, zero failed, nine ignored; closed 22:14:20 UTC | Nested subprocess tests counted once; ignored gates are separate | +| Real RustFS compatibility | All eight exact gates passed; closed 22:16:45 UTC | Disposable fixture, not matched capacity or five-GiB transfer | +| Original cold-activation regression | All 20 independent repetitions passed | Exact production test artifact; no HTTP retries or relaxed assertions | +| Complete residency suite | All 15 tests passed at four threads; closed 22:17:45 UTC | Pinning, cancellation, release and restore faults; not live corpus recovery | + +The cleanup-test correction isolates the parent-fence assertion and adds an +inherited-fork safety regression. Production cleanup remains conservative: +busy fenced files stay present and charged until recovery. Temporary refusal +classification and inventory traces are not in the PR. + +With Rust, stock Git, Git LFS, AWS CLI and a reachable Docker daemon: + +```sh +cargo test --release --locked --workspace -- --test-threads=4 +python3 scripts/qualify_size.py --provider-only --release +``` + +The second command creates and removes only its own disposable RustFS fixture. +It excludes the non-sparse five-GiB test and is not the three-node load driver. + +## Evidence and remaining gates + +Release evidence is under +`/Users/haipingfu/.codex/canopy-bounded-auth-residency-release-gwk4TA`. +It binds 260 tracked source/harness files plus the build helper and retains every +workspace executable. All 282 preserved files passed copy and independent reread +checks at `/Volumes/Workspace/CrabData/canopy-combined-correctness-evidence-y1l45kek`. +Provider/repetition evidence is under +`/Users/haipingfu/.codex/canopy-combined-provider-gates-HZBUvh`; its 303 preserved +files passed the same checks at +`/Volumes/Workspace/CrabData/canopy-combined-provider-residency-evidence-jjp_u395`. +These different-filesystem local copies are not off-machine backups. + +| Closed artifact | SHA-256 | +| --- | --- | +| Combined executable | `66adab6c8a2b8d801e4ff9d053a4edb22c483e407dc9627dc3ee57c805d29704` | +| `build-tests.json` | `8eedaa338a21db3dd8f5982f37ec57c34f2706a685d1a7362eea9c10d637f1f6` | +| `release-workspace.log` | `1a99446f42216732d63ab2227ecda9ab39060becf7a0bf04e9ee8d0a3bd3b2cc` | +| `provider-tests.json` | `cee298f5314b819b4ee704da6ac4f69e48910a99cdd4dcffc94a8b843ffd9dad` | +| `provider-tests.log` | `cc41133268681850e16105c172ce16db4ab7429afbd479723b765aa07550e344` | +| `residency-repetitions.json` | `db707b06c7af57ab21c569e13cb1878cc62a00d0060c58d18255b3208372ea05` | + +The earlier diagnostic fleet subsequently failed: all three nodes exited 1 +without forced shutdown. Logs record Cell fencing around 20:58 UTC, then +`node lease bounds are invalid` around 21:00:40 UTC and unconfirmed-drain +quarantine. Terminal logs, outcome and stored node tombstones are preserved. +The unchanged RustFS provider had zero restarts or OOM kills; observed host logs +did not establish a sleep or clock-jump cause. The stall cause remains unresolved. +This candidate has not been deployed and cannot be credited with fixing it. + +Required next proofs remain: + +- Safe existing-catalog admission, old-code restore and actual same-corpus upgrade. +- Full 10,000-identity/100-fixture Git/LFS verification and fresh-owner recovery. +- The complete schedule, critical concurrent workflows and faults, and every new + acknowledged write after owner loss. +- Newer Cellule qualification, higher admission profiles, matched comparisons, + non-sparse five-GiB transfer and complete CPU/cost accounting. +- Isolated Linux reference capacity; the shared Mac/Colima environment does not + establish it. Missing historical ledgers still prevent old every-ACK proof. + +The [performance plan](../performance-plan.md) remains the full scope. PR #18 +stays draft while these gates are open. diff --git a/docs/performance/2026-10-01-evidence-loss-and-c51dd121.md b/docs/performance/2026-10-01-evidence-loss-and-c51dd121.md index 773fffa..ade2fcc 100644 --- a/docs/performance/2026-10-01-evidence-loss-and-c51dd121.md +++ b/docs/performance/2026-10-01-evidence-loss-and-c51dd121.md @@ -5,6 +5,11 @@ incomplete. The new Cellule candidate passed release tests and initial Git verification. Its full corpus seed completed, but remote verification failed with HTTP 503. Recovery and timed load did not start.** +The [bounded authentication and residency follow-up](2026-10-01-bounded-authentication-and-residency.md) +now records the combined production changes included in PR #18, their closed +correctness checks and remaining upgrade/recovery gates. The failed attempts +below remain historical results, not passing retries. + This checkpoint supersedes earlier running/retained-artifact statements in the [native-filesystem trial](2026-10-01-native-filesystem.md). It keeps the original 10,000-repository target, all 108 load windows and separate correctness gates. @@ -120,9 +125,10 @@ upgrading the pin or changing the live experiment. The locked release build completed successfully. Its retained executable has SHA-256 `32b114119960608c0a91d1c783bb69eafec432831bfa452d54d8950b09bc0e99`. -The build receipt binds 256 production-source and harness files; those bindings -still match the published candidate. Evidence and the executable are outside the -rebuildable Cargo target directory. +The build receipt binds 256 production-source and harness files. Those bindings +matched the original PR source at `4651aaa`; that baseline is preserved separately +and does not bind the later combined candidate. Evidence and the executable are +outside the rebuildable Cargo target directory. | Verification | Result and scope | | --- | --- | @@ -286,7 +292,7 @@ These observations narrow the investigation, but do not prove that all earlier 503s share a cause. This replay excludes Git/LFS bodies, scheduled throughput and crash recovery. A failed metadata check remains a failed correctness gate. -### Bounded authentication candidate remains separate +### First bounded authentication candidate failed A regression at the actual `DirectoryCell.authenticate` call site held its SQL worker while polling 16 valid authentication requests. The original generic @@ -313,9 +319,10 @@ not a measured throughput improvement or a completed live-fleet fix. The original predecessor release descriptor was read from RustFS and its BLAKE3 digest verified without changing selected state. Reading it is not rolling -compatibility or restore proof. The candidate is **not deployed or included in -this PR**. Its failure cannot be dismissed as test timing, and its passing -focused checks cannot replace full correctness verification. +compatibility or restore proof. At the 21:09 UTC checkpoint, this first candidate +was **not deployed or included in the PR**. Its failure cannot be dismissed as +test timing, and its passing focused checks cannot replace full correctness +verification. New closed files are under `/Users/haipingfu/.codex/canopy-three-node-evidence-BwYz7P`, separately from the @@ -416,7 +423,7 @@ Neither backup overwrites original failed receipts or certifies recovery. | Rust workspace correctness and RustFS Git/LFS compatibility | Passed release suite, eight provider gates and initial 17-step proxy check | | New-pin non-sparse 5-GiB transfer | Open; excluded by `--provider-only` | | Full 10,000 identities/100 populated fixtures | Full seed closed; remote verification failed HTTP 503; cause unresolved | -| Bounded authentication candidate | Focused regression and Directory suite passed; full workspace failed; old-code compatibility and live verification open; not in PR | +| First bounded authentication candidate | Focused regression and Directory suite passed; full workspace failed. The [combined follow-up](2026-10-01-bounded-authentication-and-residency.md) is now in the PR with passing release/provider checks; existing-catalog upgrade and live verification stay open | | Original matrix and every new ACK after owner loss | Downstream gates stopped without owner signals or timed load; recovery open | | Critical concurrent Git load and fault coverage | Open; separate from the initial 17-step functional check | | Higher admission profiles and matched comparisons | Open | From 6136dfd552179e8e14ed61bf6deea451b8fc1189 Mon Sep 17 00:00:00 2001 From: forhappy Date: Thu, 1 Oct 2026 15:32:40 -0700 Subject: [PATCH 36/60] test: reproduce real startup against a retained predecessor catalog --- .../canopy-server/tests/multi_server/main.rs | 1 + .../tests/multi_server/retained_catalog.rs | 249 ++++++++++++++++++ 2 files changed, 250 insertions(+) create mode 100644 crates/canopy-server/tests/multi_server/retained_catalog.rs diff --git a/crates/canopy-server/tests/multi_server/main.rs b/crates/canopy-server/tests/multi_server/main.rs index 4ab06b8..bc6fe92 100644 --- a/crates/canopy-server/tests/multi_server/main.rs +++ b/crates/canopy-server/tests/multi_server/main.rs @@ -38,6 +38,7 @@ mod lfs_locks; mod lifecycle; mod peers; mod residency; +mod retained_catalog; mod tokens; mod transfers; mod visibility; diff --git a/crates/canopy-server/tests/multi_server/retained_catalog.rs b/crates/canopy-server/tests/multi_server/retained_catalog.rs new file mode 100644 index 0000000..7bda449 --- /dev/null +++ b/crates/canopy-server/tests/multi_server/retained_catalog.rs @@ -0,0 +1,249 @@ +//! Real startup seam for an immutable predecessor Directory catalog entry. +//! This owned in-memory fixture is not an upgrade of the live RustFS corpus +//! or proof that an old executable produced the stored bytes. + +use super::*; +use bytes::Bytes; +use canopy_server::{ + CanopyApplication, build_descriptor, + directory::{self, DirectoryCell, DirectoryModule, TokenScope}, +}; +use cellule_app::{ApplicationHandle, CellApplication}; +use cellule_ltx::{CellReplica, DiskBudget, Host, Limits}; +use cellule_runtime::{ + CellClient, CellModule, CellRuntime, MutationIdentity, SessionId, + cell::application::{ApplicationIdentity, ApplicationIdentityStore}, + cell::catalog::{CatalogEntry, CatalogRole, CellCatalog}, + cell::worker::SqlWorkerPool, + control::{ControlState, Owner, authority::CellAuthority}, + identity::{IncarnationId, RequestId}, + ltx::CellStorageLayout, + node::NodeDirectory, + recovery::release::{ReleaseState, ReleaseStore}, +}; +use cellule_store::Store; +use sha2::{Digest as _, Sha256}; +use std::time::{SystemTime, UNIX_EPOCH}; + +type Result = std::result::Result>; + +fn now() -> Result { + Ok(i64::try_from( + SystemTime::now().duration_since(UNIX_EPOCH)?.as_millis(), + )?) +} + +#[tokio::test(flavor = "multi_thread")] +async fn server_reopens_a_retained_directory_catalog_after_explicit_fixture_activation() -> Result { + let files = tempfile::TempDir::new()?; + let address = available_address().await?; + let configuration = config(address, files.path().join("server")); + let store: Arc = Arc::new(InMemory::new()); + let storage = Store::new(Arc::clone(&store)); + let identity = ApplicationIdentity::new(configuration.tenant, configuration.application); + storage + .create_strict( + &configuration + .store_prefix + .clone() + .join("canopy-root-v1.json"), + Bytes::from_static(br#"{"kind":"service"}"#), + ) + .await?; + ApplicationIdentityStore::new(storage.clone(), configuration.store_prefix.clone()) + .initialize(identity) + .await?; + let layout = CellStorageLayout::new( + storage, + configuration.store_prefix.clone(), + *configuration.application.as_bytes(), + ); + let application = Arc::new(CanopyApplication::compile(build_descriptor( + include_bytes!("../../../../Cargo.lock"), + env!("CARGO_PKG_VERSION"), + ))?); + let registry = application.registry(); + let previous = + include_bytes!("../directory_cell/fixtures/c51-selected-release.json").trim_ascii_end(); + assert_eq!( + blake3::hash(previous).to_hex().as_str(), + "e31bf1a951e2fa19d91e9f964b2ddeade1a81b05a20ad628362819a1487c16b1" + ); + registry.verify_rolling_from(previous)?; + let descriptor: serde_json::Value = serde_json::from_slice(previous)?; + let module = descriptor["modules"] + .as_array() + .ok_or("modules missing")? + .iter() + .find(|module| module["name"] == DirectoryModule::NAME) + .ok_or("Directory predecessor missing")?; + let old_code = Digest::from_bytes( + hex::decode(module["code"].as_str().ok_or("code missing")?)? + .try_into() + .map_err(|_| "invalid predecessor code")?, + ); + let releases = ReleaseStore::new(layout.clone(), identity)?; + let image = format!("sha256:{}", hex::encode(configuration.image.as_bytes())); + let old_operation = RequestId::from_bytes(uuid::Uuid::new_v4().into_bytes()); + let prepared = releases + .prepare( + previous, + Digest::from_bytes(*blake3::hash(previous).as_bytes()), + 0, + &image, + old_operation, + ) + .await?; + // The fixture has no Cells or advertised writers at first activation. + let activating = releases + .start_activation(prepared.revision(), old_operation) + .await?; + let old_ready = releases + .complete_activation(activating.revision(), old_operation) + .await?; + assert_eq!(old_ready.state(), ReleaseState::Ready); + + let target = directory::directory_target(configuration.tenant, configuration.application)?; + let catalog = CellCatalog::new(layout.clone(), configuration.tenant); + let entry = CatalogEntry::new(&target, CatalogRole::Sql, old_code, 1)?; + let proof = catalog.provision(entry.clone()).await?; + let authority = CellAuthority::new(layout.clone()); + let session = SessionId::from_bytes(uuid::Uuid::new_v4().into_bytes()); + let observed = authority + .create_initial( + &proof, + IncarnationId::from_bytes(uuid::Uuid::new_v4().into_bytes()), + Owner { + session, + endpoint: "https://canopy.test".into(), + }, + ) + .await?; + let runtime = CellRuntime::new_with_replica_host( + SqlWorkerPool::new(1, 4)?, + 64 * 1024 * 1024, + session, + Host::default().with_local_disk_budget(DiskBudget::new(1 << 30)), + )?; + let handle = runtime + .bootstrap( + proof, + CellReplica::new( + layout.clone(), + *target.cell_id().as_bytes(), + *observed.value().incarnation.as_bytes(), + Limits::default(), + )?, + authority.clone(), + observed, + files.path().join("predecessor.sqlite"), + |transaction| { + transaction.execute_batch(directory::SCHEMA)?; + Ok(()) + }, + ) + .await?; + let client = ApplicationHandle::::new( + CellClient::local(registry.clone(), handle), + application, + configuration.tenant, + configuration.application, + )?; + let directory = DirectoryCell::new(&client, target.clone())?; + let issued = now()?; + directory + .create_account( + MutationIdentity { + request_id: RequestId::from_bytes(uuid::Uuid::new_v4().into_bytes()), + issued_at_ms: issued, + expires_at_ms: issued + 60_000, + }, + "canopy", + Sha256::digest(configuration.token.as_bytes()).into(), + TokenScope::Admin, + ) + .await?; + runtime.shutdown().await?; + let old_control = authority + .load(target.cell_id()) + .await? + .ok_or("control missing")?; + assert_eq!(old_control.value().code, old_code); + assert_eq!(old_control.value().schema, 1); + assert_eq!(old_control.value().state, ControlState::Idle); + assert!(old_control.value().root.is_some()); + assert!(old_control.value().owner.is_none()); + + // Explicit fixture-only admission, before completing the new release. + // The low-level phase CAS does not itself prove compatibility or drain. + let nodes = NodeDirectory::new( + layout.clone(), + configuration.fleet, + configuration.image, + registry.release_digest(), + ); + assert!(nodes.advertised_sessions(now()?, 4096).await?.is_empty()); + let mut cells = 0; + for shard in 0..=u8::MAX { + let mut scan = catalog.scan_shard(shard).await?; + while let Some(page) = scan.next_page().await? { + for proof in page.entries() { + assert_eq!(proof.entry(), &entry); + assert!(registry.supports_cell( + proof.entry().namespace(), + proof.entry().role(), + proof.entry().initial_code(), + proof.entry().initial_schema(), + )); + cells += 1; + } + } + } + assert_eq!(cells, 1); + let operation = RequestId::from_bytes(uuid::Uuid::new_v4().into_bytes()); + let prepared = releases + .prepare( + registry.release_bytes(), + registry.release_digest(), + old_ready.revision(), + &image, + operation, + ) + .await?; + let activating = releases + .start_activation(prepared.revision(), operation) + .await?; + releases + .complete_activation(activating.revision(), operation) + .await?; + + // Exercise real Canopy startup, not a direct low-level client workaround. + let server = CanopyServer::start(configuration, Arc::clone(&store)).await?; + let url = create_repository(address, "after-activation").await?; + let listing: serde_json::Value = reqwest::Client::new() + .get(format!("http://{address}/api/repositories")) + .bearer_auth("local-test-token") + .send() + .await? + .error_for_status()? + .json() + .await?; + assert_eq!( + listing["repositories"] + .as_array() + .ok_or("list missing")? + .len(), + 1 + ); + assert_eq!( + catalog + .lookup(target.cell_id()) + .await? + .ok_or("catalog missing")? + .entry(), + &entry + ); + assert!(url.ends_with("/canopy/after-activation.git")); + server.shutdown().await?; + Ok(()) +} From 06834cc6eaace4a5db4bc835f1abb9b47e7dbf93 Mon Sep 17 00:00:00 2001 From: forhappy Date: Thu, 1 Oct 2026 15:42:02 -0700 Subject: [PATCH 37/60] fix: admit retained SQL catalog proofs without rewriting their identity --- .../src/server/catalog_admission.rs | 57 +++++++++++++++++++ crates/canopy-server/src/server/mod.rs | 22 ++++--- 2 files changed, 72 insertions(+), 7 deletions(-) create mode 100644 crates/canopy-server/src/server/catalog_admission.rs diff --git a/crates/canopy-server/src/server/catalog_admission.rs b/crates/canopy-server/src/server/catalog_admission.rs new file mode 100644 index 0000000..303b980 --- /dev/null +++ b/crates/canopy-server/src/server/catalog_admission.rs @@ -0,0 +1,57 @@ +//! Read-only admission of an existing immutable SQL catalog identity. + +use cellule_runtime::{ + CellTarget, Error, Registry, Result, + cell::catalog::{CatalogProof, CatalogRole}, + recovery::release::{ReleaseState, ReleaseStore}, +}; + +pub(super) async fn existing_sql_proof( + releases: &ReleaseStore, + registry: &Registry, + target: &CellTarget, + proof: CatalogProof, +) -> Result { + // New entries must still use ReleaseStore::provision and current code. + // An existing entry pins its original code/schema forever; admission must + // verify that identity instead of attempting to rewrite it during restore. + let before = releases.load().await?.ok_or(Error::Release( + "release is absent during existing Cell admission", + ))?; + let digest = registry.release_digest(); + let record = before.record(); + if record.state() != ReleaseState::Ready + || record.current() != Some(digest) + || record.desired() != Some(digest) + || releases.descriptor(digest).await? != registry.release_bytes() + { + return Err(Error::Release( + "existing Cell requires the exact ready release", + )); + } + let entry = proof.entry(); + if entry.cell() != target.cell_id() + || entry.namespace() != target.namespace() + || entry.partition() != target.partition() + || entry.role() != CatalogRole::Sql + || !registry.supports_cell( + entry.namespace(), + entry.role(), + entry.initial_code(), + entry.initial_schema(), + ) + { + return Err(Error::Release( + "existing SQL catalog identity is unsupported", + )); + } + let after = releases.load().await?.ok_or(Error::Release( + "release disappeared during existing Cell admission", + ))?; + if after.record() != record { + return Err(Error::Release( + "release changed during existing Cell admission", + )); + } + Ok(proof) +} diff --git a/crates/canopy-server/src/server/mod.rs b/crates/canopy-server/src/server/mod.rs index 2d9b0da..6e4eeae 100644 --- a/crates/canopy-server/src/server/mod.rs +++ b/crates/canopy-server/src/server/mod.rs @@ -46,6 +46,7 @@ use crate::{ repository_http::RepositoryHttp, }; +mod catalog_admission; mod discovery; mod lifecycle; pub(crate) mod peer; @@ -762,13 +763,20 @@ async fn acquire_sql_cell( layout.clone(), ApplicationIdentity::new(target.tenant(), target.application()), )?; - let proof = releases - .provision( - &catalog, - ®istry, - CatalogEntry::new(target, CatalogRole::Sql, code, 1)?, - ) - .await?; + let proof = match catalog.lookup(target.cell_id()).await? { + Some(proof) => { + catalog_admission::existing_sql_proof(&releases, ®istry, target, proof).await? + } + None => { + releases + .provision( + &catalog, + ®istry, + CatalogEntry::new(target, CatalogRole::Sql, code, 1)?, + ) + .await? + } + }; acquire_provisioned_sql_cell(node, layout, directory, spec, session, endpoint, proof).await } From ae4b1962a38f8180f8190e146144111ffe64b5ee Mon Sep 17 00:00:00 2001 From: forhappy Date: Thu, 1 Oct 2026 15:49:18 -0700 Subject: [PATCH 38/60] test: require unsupported persisted controls to stay unchanged on startup refusal --- .../tests/multi_server/retained_catalog.rs | 97 ++++++++++++++++++- 1 file changed, 95 insertions(+), 2 deletions(-) diff --git a/crates/canopy-server/tests/multi_server/retained_catalog.rs b/crates/canopy-server/tests/multi_server/retained_catalog.rs index 7bda449..d296fbd 100644 --- a/crates/canopy-server/tests/multi_server/retained_catalog.rs +++ b/crates/canopy-server/tests/multi_server/retained_catalog.rs @@ -33,8 +33,18 @@ fn now() -> Result { )?) } -#[tokio::test(flavor = "multi_thread")] -async fn server_reopens_a_retained_directory_catalog_after_explicit_fixture_activation() -> Result { +struct RetainedFixture { + _files: tempfile::TempDir, + configuration: ServerConfig, + store: Arc, + layout: CellStorageLayout, + catalog: CellCatalog, + authority: CellAuthority, + target: cellule_runtime::CellTarget, + entry: CatalogEntry, +} + +async fn retained_fixture() -> Result { let files = tempfile::TempDir::new()?; let address = available_address().await?; let configuration = config(address, files.path().join("server")); @@ -217,6 +227,32 @@ async fn server_reopens_a_retained_directory_catalog_after_explicit_fixture_acti .complete_activation(activating.revision(), operation) .await?; + Ok(RetainedFixture { + _files: files, + configuration, + store, + layout, + catalog, + authority, + target, + entry, + }) +} + +#[tokio::test(flavor = "multi_thread")] +async fn server_reopens_a_retained_directory_catalog_after_explicit_fixture_activation() -> Result { + let fixture = retained_fixture().await?; + let address = fixture.configuration.listen; + let RetainedFixture { + _files, + configuration, + store, + catalog, + target, + entry, + .. + } = fixture; + // Exercise real Canopy startup, not a direct low-level client workaround. let server = CanopyServer::start(configuration, Arc::clone(&store)).await?; let url = create_repository(address, "after-activation").await?; @@ -247,3 +283,60 @@ async fn server_reopens_a_retained_directory_catalog_after_explicit_fixture_acti server.shutdown().await?; Ok(()) } + +async fn unsupported_control_stays_unchanged(code: Option, schema: Option) -> Result { + let fixture = retained_fixture().await?; + let observed = fixture + .authority + .load(fixture.target.cell_id()) + .await? + .ok_or("control missing")?; + let mut unsupported = observed.value().clone(); + if let Some(code) = code { + unsupported.code = code; + } + if let Some(schema) = schema { + unsupported.schema = schema; + } + unsupported.revision += 1; + unsupported.progress += 1; + let path = fixture + .layout + .control_path(fixture.target.cell_id().as_bytes()); + let (_, token) = fixture.layout.store().get_with_etag(&path).await?; + let encoded = unsupported.encode()?; + // Inject unsupported persisted metadata only in this owned in-memory fixture. + // Its catalog remains supported. A failed startup must reject the control + // before acquisition, not claim it and rely on later SQL/client rejection. + fixture + .layout + .store() + .update(&path, Bytes::from(encoded.clone()), token) + .await?; + assert!( + CanopyServer::start(fixture.configuration, Arc::clone(&fixture.store)) + .await + .is_err() + ); + let after = fixture + .authority + .load(fixture.target.cell_id()) + .await? + .ok_or("control disappeared")?; + assert_eq!( + after.value().encode()?, + encoded, + "unsupported persisted control was acquired or rewritten" + ); + Ok(()) +} + +#[tokio::test(flavor = "multi_thread")] +async fn unknown_control_code_is_rejected_before_cell_acquisition() -> Result { + unsupported_control_stays_unchanged(Some(Digest::from_bytes([1; 32])), None).await +} + +#[tokio::test(flavor = "multi_thread")] +async fn unknown_control_schema_is_rejected_before_cell_acquisition() -> Result { + unsupported_control_stays_unchanged(None, Some(2)).await +} From 395111b188d5f4b84ae3a8ccebe0857c2b236244 Mon Sep 17 00:00:00 2001 From: forhappy Date: Thu, 1 Oct 2026 15:57:23 -0700 Subject: [PATCH 39/60] Reject unsupported persisted SQL controls before acquisition --- crates/canopy-server/src/server/mod.rs | 11 +++++++++++ 1 file changed, 11 insertions(+) diff --git a/crates/canopy-server/src/server/mod.rs b/crates/canopy-server/src/server/mod.rs index 6e4eeae..12e7f70 100644 --- a/crates/canopy-server/src/server/mod.rs +++ b/crates/canopy-server/src/server/mod.rs @@ -812,6 +812,17 @@ pub(crate) async fn acquire_provisioned_sql_cell( .await? } }; + // Catalog identity describes initial code/schema, not necessarily the + // current persisted Control. Reject unsupported Control metadata before + // bootstrap, ownership takeover, or restoration can mutate authority. + if !node.application().registry().supports_cell( + target.namespace(), + CatalogRole::Sql, + observed.value().code, + observed.value().schema, + ) { + return Err(Error::Control("persisted SQL Cell code/schema is unsupported").into()); + } let cell_type = node .application() .cell_types() From a435e6f4f7818079241e7d50010e29beba68dc75 Mon Sep 17 00:00:00 2001 From: forhappy Date: Thu, 1 Oct 2026 16:02:00 -0700 Subject: [PATCH 40/60] Test read-only catalog admission across release transitions --- .../src/server/catalog_admission.rs | 3 + .../src/server/catalog_admission/tests.rs | 350 ++++++++++++++++++ 2 files changed, 353 insertions(+) create mode 100644 crates/canopy-server/src/server/catalog_admission/tests.rs diff --git a/crates/canopy-server/src/server/catalog_admission.rs b/crates/canopy-server/src/server/catalog_admission.rs index 303b980..9f95819 100644 --- a/crates/canopy-server/src/server/catalog_admission.rs +++ b/crates/canopy-server/src/server/catalog_admission.rs @@ -1,5 +1,8 @@ //! Read-only admission of an existing immutable SQL catalog identity. +#[cfg(test)] +mod tests; + use cellule_runtime::{ CellTarget, Error, Registry, Result, cell::catalog::{CatalogProof, CatalogRole}, diff --git a/crates/canopy-server/src/server/catalog_admission/tests.rs b/crates/canopy-server/src/server/catalog_admission/tests.rs new file mode 100644 index 0000000..6c14a41 --- /dev/null +++ b/crates/canopy-server/src/server/catalog_admission/tests.rs @@ -0,0 +1,350 @@ +use std::{ + fmt, + pin::Pin, + sync::{Arc, Mutex}, + time::Duration, +}; + +use bytes::Bytes; +use cellule_app::CellApplication; +use cellule_runtime::{ + ApplicationId, CellModule, Digest, TenantId, + cell::application::ApplicationIdentity, + cell::catalog::{CatalogEntry, CellCatalog}, + identity::RequestId, + ltx::CellStorageLayout, +}; +use cellule_store::Store; +use futures_core::Stream; +use object_store::{ + CopyOptions, GetOptions, GetResult, ListResult, MultipartUpload, ObjectMeta, ObjectStore, + ObjectStoreExt, PutMultipartOptions, PutOptions, PutPayload, PutResult, memory::InMemory, + path::Path, +}; +use tokio::{sync::Notify, time::timeout}; + +use super::*; +use crate::{ + CanopyApplication, build_descriptor, + directory::{DirectoryModule, directory_target}, +}; + +type TestResult = std::result::Result>; +type StoreStream = Pin> + Send + 'static>>; + +// Fault injection belongs to this owned fixture, not a shared deployment. +#[derive(Debug, Default)] +struct PausedDescriptor { + inner: InMemory, + pause: Mutex>, + entered: Notify, + proceed: Notify, +} + +impl fmt::Display for PausedDescriptor { + fn fmt(&self, formatter: &mut fmt::Formatter<'_>) -> fmt::Result { + formatter.write_str("catalog-admission-fixture") + } +} + +#[async_trait::async_trait] +impl ObjectStore for PausedDescriptor { + async fn put_opts( + &self, + path: &Path, + payload: PutPayload, + options: PutOptions, + ) -> object_store::Result { + self.inner.put_opts(path, payload, options).await + } + async fn put_multipart_opts( + &self, + path: &Path, + options: PutMultipartOptions, + ) -> object_store::Result> { + self.inner.put_multipart_opts(path, options).await + } + async fn get_opts(&self, path: &Path, options: GetOptions) -> object_store::Result { + let pause = { + let mut armed = self.pause.lock().unwrap(); + if armed.as_ref() == Some(path) { + armed.take(); + true + } else { + false + } + }; + if pause { + self.entered.notify_one(); + self.proceed.notified().await; + } + self.inner.get_opts(path, options).await + } + fn delete_stream(&self, paths: StoreStream) -> StoreStream { + self.inner.delete_stream(paths) + } + fn list(&self, prefix: Option<&Path>) -> StoreStream { + self.inner.list(prefix) + } + async fn list_with_delimiter(&self, prefix: Option<&Path>) -> object_store::Result { + self.inner.list_with_delimiter(prefix).await + } + async fn copy_opts( + &self, + from: &Path, + to: &Path, + options: CopyOptions, + ) -> object_store::Result<()> { + self.inner.copy_opts(from, to, options).await + } +} + +struct Fixture { + store: Arc, + layout: CellStorageLayout, + registry: Arc, + target: CellTarget, + catalog: CellCatalog, + releases: ReleaseStore, +} + +impl Fixture { + fn new() -> TestResult { + let store = Arc::new(PausedDescriptor::default()); + let tenant = TenantId::from_bytes([1; 16]); + let application = ApplicationId::from_bytes([2; 16]); + let layout = CellStorageLayout::new( + Store::new(store.clone()), + Path::from("admission"), + *application.as_bytes(), + ); + let registry = CanopyApplication::compile(build_descriptor( + include_bytes!("../../../../../../Cargo.lock"), + env!("CARGO_PKG_VERSION"), + ))? + .registry(); + let target = directory_target(tenant, application)?; + let catalog = CellCatalog::new(layout.clone(), tenant); + let releases = ReleaseStore::new( + layout.clone(), + ApplicationIdentity::new(tenant, application), + )?; + Ok(Self { + store, + layout, + registry, + target, + catalog, + releases, + }) + } + + fn code(&self) -> Digest { + self.registry.module_code(DirectoryModule::NAME).unwrap() + } + + async fn prepare(&self) -> TestResult<(u64, RequestId)> { + let revision = self + .releases + .load() + .await? + .map_or(0, |value| value.record().revision()); + let operation = RequestId::from_bytes(uuid::Uuid::new_v4().into_bytes()); + let prepared = self + .releases + .prepare( + self.registry.release_bytes(), + self.registry.release_digest(), + revision, + &format!("sha256:{}", "a".repeat(64)), + operation, + ) + .await?; + Ok((prepared.revision(), operation)) + } + + async fn ready(&self) -> TestResult { + let (revision, operation) = self.prepare().await?; + let activating = self.releases.start_activation(revision, operation).await?; + self.releases + .complete_activation(activating.revision(), operation) + .await?; + Ok(()) + } + + async fn proof( + &self, + role: CatalogRole, + code: Digest, + schema: u32, + ) -> TestResult { + Ok(self + .catalog + .provision(CatalogEntry::new(&self.target, role, code, schema)?) + .await?) + } + + async fn admit(&self, proof: CatalogProof) -> cellule_runtime::Result { + existing_sql_proof(&self.releases, &self.registry, &self.target, proof).await + } +} + +#[tokio::test] +async fn supported_current_and_retained_proofs_are_returned_unchanged() -> TestResult { + let retained = Digest::from_bytes( + hex::decode("f7254eda9d5d339566f45457502618ad13cbbf6e5a74595f5b3ce46653ea12f1")? + .try_into() + .map_err(|_| "invalid retained code")?, + ); + for code in [None, Some(retained)] { + let fixture = Fixture::new()?; + fixture.ready().await?; + let proof = fixture + .proof(CatalogRole::Sql, code.unwrap_or(fixture.code()), 1) + .await?; + let entry = proof.entry().clone(); + let revision = proof.revision(); + let release = fixture.releases.load().await?.unwrap().record().clone(); + let admitted = fixture.admit(proof).await?; + assert_eq!(admitted.entry(), &entry); + assert_eq!(admitted.revision(), revision); + assert_eq!(fixture.releases.load().await?.unwrap().record(), &release); + } + Ok(()) +} + +#[tokio::test] +async fn absent_and_non_ready_releases_close_existing_admission() -> TestResult { + for state in [ + None, + Some(ReleaseState::Prepared), + Some(ReleaseState::Activating), + Some(ReleaseState::Maintenance), + ] { + let fixture = Fixture::new()?; + let proof = fixture.proof(CatalogRole::Sql, fixture.code(), 1).await?; + if let Some(state) = state { + let (revision, operation) = fixture.prepare().await?; + match state { + ReleaseState::Activating => { + fixture + .releases + .start_activation(revision, operation) + .await?; + } + ReleaseState::Maintenance => { + fixture + .releases + .start_maintenance(revision, operation) + .await?; + } + ReleaseState::Prepared => {} + _ => unreachable!(), + } + } + let before = fixture + .releases + .load() + .await? + .map(|value| value.record().clone()); + assert!(matches!(fixture.admit(proof).await, Err(Error::Release(_)))); + let after = fixture + .releases + .load() + .await? + .map(|value| value.record().clone()); + assert_eq!(before, after); + } + Ok(()) +} + +#[tokio::test] +async fn corrupted_selected_descriptor_closes_existing_admission() -> TestResult { + let fixture = Fixture::new()?; + fixture.ready().await?; + let proof = fixture.proof(CatalogRole::Sql, fixture.code(), 1).await?; + fixture + .store + .put( + &fixture + .layout + .release_descriptor_path(fixture.registry.release_digest().as_bytes()), + Bytes::from_static(b"corrupt").into(), + ) + .await?; + assert!(matches!(fixture.admit(proof).await, Err(Error::Release(_)))); + Ok(()) +} + +#[tokio::test] +async fn unsupported_catalog_identity_and_wrong_target_are_not_adopted() -> TestResult { + for case in 0..4 { + let fixture = Fixture::new()?; + fixture.ready().await?; + let role = if case == 0 { + CatalogRole::Kv + } else { + CatalogRole::Sql + }; + let code = if case == 1 { + Digest::from_bytes([1; 32]) + } else { + fixture.code() + }; + let schema = if case == 2 { 2 } else { 1 }; + let proof = fixture.proof(role, code, schema).await?; + let entry = proof.entry().clone(); + let other = directory_target(fixture.target.tenant(), ApplicationId::from_bytes([3; 16]))?; + let target = if case == 3 { &other } else { &fixture.target }; + assert!(matches!( + existing_sql_proof(&fixture.releases, &fixture.registry, target, proof).await, + Err(Error::Release( + "existing SQL catalog identity is unsupported" + )), + )); + assert_eq!( + fixture + .catalog + .lookup(fixture.target.cell_id()) + .await? + .unwrap() + .entry(), + &entry + ); + } + Ok(()) +} + +#[tokio::test] +async fn release_round_trip_to_ready_during_admission_is_rejected() -> TestResult { + let fixture = Fixture::new()?; + fixture.ready().await?; + let proof = fixture.proof(CatalogRole::Sql, fixture.code(), 1).await?; + let before = fixture.releases.load().await?.unwrap().record().clone(); + *fixture.store.pause.lock().unwrap() = Some( + fixture + .layout + .release_descriptor_path(fixture.registry.release_digest().as_bytes()), + ); + let change = async { + timeout(Duration::from_secs(5), fixture.store.entered.notified()).await?; + // This fixture has no owners or Controls; it is not a live upgrade controller. + let result = fixture.ready().await; + fixture.store.proceed.notify_one(); + result + }; + let (admission, changed) = tokio::join!(fixture.admit(proof), change); + changed?; + let after = fixture.releases.load().await?.unwrap().record().clone(); + assert_eq!(after.state(), ReleaseState::Ready); + assert_eq!(before.current(), after.current()); + assert_eq!(before.desired(), after.desired()); + assert_ne!(before.revision(), after.revision()); + assert!(matches!( + admission, + Err(Error::Release( + "release changed during existing Cell admission" + )) + )); + Ok(()) +} From 33e1dc897c337121bb92030bb9e7b48465d43591 Mon Sep 17 00:00:00 2001 From: forhappy Date: Thu, 1 Oct 2026 16:02:29 -0700 Subject: [PATCH 41/60] Correct catalog admission fixture lockfile path --- crates/canopy-server/src/server/catalog_admission/tests.rs | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/crates/canopy-server/src/server/catalog_admission/tests.rs b/crates/canopy-server/src/server/catalog_admission/tests.rs index 6c14a41..a9c8d33 100644 --- a/crates/canopy-server/src/server/catalog_admission/tests.rs +++ b/crates/canopy-server/src/server/catalog_admission/tests.rs @@ -119,7 +119,7 @@ impl Fixture { *application.as_bytes(), ); let registry = CanopyApplication::compile(build_descriptor( - include_bytes!("../../../../../../Cargo.lock"), + include_bytes!("../../../../../Cargo.lock"), env!("CARGO_PKG_VERSION"), ))? .registry(); From 5e9ca0ad1e79b9082ca4f5ab55d08b514678ef36 Mon Sep 17 00:00:00 2001 From: forhappy Date: Thu, 1 Oct 2026 16:24:32 -0700 Subject: [PATCH 42/60] Document verified retained catalog startup and safety boundaries --- docs/contracts.md | 16 +- docs/delivery-plan.md | 8 +- docs/performance-plan.md | 13 +- ...01-bounded-authentication-and-residency.md | 20 ++- .../2026-10-01-retained-catalog-admission.md | 151 ++++++++++++++++++ 5 files changed, 195 insertions(+), 13 deletions(-) create mode 100644 docs/performance/2026-10-01-retained-catalog-admission.md diff --git a/docs/contracts.md b/docs/contracts.md index f94e99f..9278a2a 100644 --- a/docs/contracts.md +++ b/docs/contracts.md @@ -1494,8 +1494,10 @@ and a 256-byte result limit. Its input is exactly one 32-byte token digest; the result contains at most one validated principal. Existing commands 1/3 and queries 2/4 keep their one-MiB contracts. Admission budgets and minimum receipts are unchanged. The Directory descriptor retains the exact preceding `f7254eda` -code at schema 1; this is descriptor compatibility, not an automatic upgrade or -proof of restoring an existing catalog under a new release. +code at schema 1. Descriptor compatibility alone is not an automatic upgrade or +old-binary restore proof. The [catalog admission regression](performance/2026-10-01-retained-catalog-admission.md) +exercises real startup against an owned in-memory predecessor fixture, not an +old executable or the existing RustFS corpus. Authentication requires `expires_ms IS NULL OR expires_ms > now_ms`; it does not wait for a cleanup job. Token listing/issuance/revocation, account creation and @@ -2523,6 +2525,16 @@ advertisement withdrawal follows the shutdown attempt. Unsettled controls still prevent maintenance completion if that attempt fails. The executable observes supervisor completion as well as OS stop signals. +New SQL catalog entries still use strict current-code provisioning. Existing SQL +entries are admitted read-only: their verified proof must match the derived +target, namespace, partition and SQL role, and the registry must support their +immutable initial code/schema. The selected release must be exact and Ready +before and after admission, with the complete release record unchanged. +The actual persisted Control code/schema is checked separately before bootstrap, +takeover or restoration can claim ownership. Unsupported Control metadata is +rejected without rewriting its canonical bytes. These checks do not activate a +release, migrate schema or replace runtime CAS/fencing. + Maintenance administration prepares the same compiled descriptor and enters Maintenance with a caller-supplied UUID. It does not select new code or modify Cell contents. Calls with the same UUID resume/replay the operation; another UUID is diff --git a/docs/delivery-plan.md b/docs/delivery-plan.md index 6e6cb78..db125c8 100644 --- a/docs/delivery-plan.md +++ b/docs/delivery-plan.md @@ -79,8 +79,12 @@ These are the next implementation slices. Each slice must update its contract, b retains its separately bound `e07670e` executable; latest source is not a claim that the running artifact was rebuilt or performance-qualified. The [bounded authentication and residency candidate](performance/2026-10-01-bounded-authentication-and-residency.md) - passed release correctness and eight RustFS compatibility gates. Existing - catalog upgrade, full-corpus recovery and performance remain open. + passed release correctness and eight RustFS compatibility gates. The + [retained catalog follow-up](performance/2026-10-01-retained-catalog-admission.md) + adds read-only existing-proof admission and rejects unsupported persisted + Control metadata before acquisition. Its 239-test release suite, eight RustFS + gates and repeated startup/residency regressions passed. Actual old-binary + RustFS upgrade, full-corpus recovery and performance remain open. 2. Execute the [repository-density performance plan](performance-plan.md), then qualify the bounded Linux deployment and residency under faults and larger hot sets. Each node reserves one SQL slot for Directory takeover and admits diff --git a/docs/performance-plan.md b/docs/performance-plan.md index fa415eb..7fe78c1 100644 --- a/docs/performance-plan.md +++ b/docs/performance-plan.md @@ -19,8 +19,14 @@ The [combined authentication and residency candidate](performance/2026-10-01-bou is now in PR #18: 231 release tests and all eight RustFS gates passed, as did 20 unchanged cold-activation repetitions and all 15 residency tests. It has not been deployed on the existing corpus. Descriptor compatibility passed, but -existing-catalog admission, old-code restore and the diagnostic fleet's terminal -lease-fencing failure remain open. These results do not qualify newer upstream +actual old-binary restore and the diagnostic fleet's terminal lease-fencing +failure remain open. The [catalog admission follow-up](performance/2026-10-01-retained-catalog-admission.md) +reproduces and fixes immutable-identity reprovisioning and late rejection of +unsupported Control metadata. Its combined candidate passed 239 release tests, +all eight RustFS gates, 20 cold-activation repetitions and 20 three-case retained +startup repetitions, plus all 15 residency tests. Retained startup uses an owned +in-memory fixture; actual old-binary RustFS upgrade remains open. +These results do not qualify newer upstream or replace the original complete performance schedule. The [three-node proxy campaign](performance/2026-09-30-three-node-proxy.md) retains the earlier fixture race and its regression, plus the candidate's failed @@ -29,7 +35,8 @@ seed and successful recovery of every recorded ACK. | Current gate | Evidence / status | | --- | --- | | Original rebuilt Cellule artifact | `c51dd121` build, 227 Rust tests, eight RustFS gates and initial 17-step proxy check passed; full 10K/100 seed closed, but remote verification failed HTTP 503. A separate 10K metadata diagnostic returned 9,999 matching identities and one 503; downstream fault/load gates stopped without owner signals. Cause unresolved | -| Combined authentication and residency candidate | 231 release tests, eight RustFS gates, 20 original cold-activation repetitions and 15 residency tests passed. Exact predecessor descriptor is retained; existing-catalog upgrade and old-code restore remain open. Not deployed or performance-qualified | +| Combined authentication and residency candidate | First publication passed 231 release tests, eight RustFS gates, 20 cold-activation repetitions and 15 residency tests. Exact predecessor descriptor is retained; see the separately qualified catalog follow-up | +| Retained catalog admission follow-up | 239 release tests, eight RustFS gates, 20 cold-activation and 20 three-case retained-startup repetitions, 15 residency tests, lints and 84 Python tests passed. Existing supported identities restore without reprovisioning; unsupported Controls stay unchanged on refusal. Owned in-memory retained fixture, not old-binary RustFS upgrade, live deployment or performance qualification | | Candidate Git behavior | Bound `e07670e` production artifact passed all 17 critical stock-Git steps through three nodes/proxy against RustFS | | Mac-backed candidate seed | Failed: 3,561/10,000 identities and 39 Git/LFS fixtures; incomplete manifest retained | | Every candidate seed ACK | All 3,561 identities and 39 populated fixtures passed fresh-owner recovery; independent ledger/binding and 78-clone audit passed | diff --git a/docs/performance/2026-10-01-bounded-authentication-and-residency.md b/docs/performance/2026-10-01-bounded-authentication-and-residency.md index 73cf8e7..84aacd8 100644 --- a/docs/performance/2026-10-01-bounded-authentication-and-residency.md +++ b/docs/performance/2026-10-01-bounded-authentication-and-residency.md @@ -6,6 +6,10 @@ in PR #18 without temporary diagnostic logging. It has **not** been deployed against the existing 10,000-repository corpus. Upgrade, full recovery and performance qualification remain open; no matched speedup is claimed. +This records the first published authentication/residency candidate. The +[retained catalog follow-up](2026-10-01-retained-catalog-admission.md) documents +the later startup fix, safety regressions and separately qualified combined source. + This follows the [failed authentication candidate and corpus verification](2026-10-01-evidence-loss-and-c51dd121.md). Their failures remain intact. The target is still three nodes behind a proxy, 10,000 identities, 100 populated Git/LFS fixtures and all 108 load windows: @@ -77,17 +81,20 @@ and unknown code. It does not open old persisted Cells. Startup still requires the exact selected release. There is no general upgrade controller; no selected release or catalog was changed during these checks. -Source inspection identifies another unresolved upgrade risk: `acquire_sql_cell` -provisions the current module code, whereas an existing catalog's initial code -is immutable. A call-site reproduction and actual old-code restore are required. +Source inspection identified another upgrade risk: `acquire_sql_cell` provisioned +the current module code, whereas an existing catalog's initial code is immutable. +The [follow-up](2026-10-01-retained-catalog-admission.md) reproduces and fixes that +call site, and tests retained startup in an owned in-memory fixture. Actual +old-executable RustFS restore and fully admitted same-corpus upgrade remain required. Do not bypass release selection, rewrite catalog identity or treat a low-level activation CAS as compatibility proof. ## Closed verification results The frozen source is local commit `392a343b83a492d49c08a6f3e83ff4dd2917437f`. -The PR production source, tests, scripts and Cargo files are byte-identical to -that candidate. All six Cellule packages remain locked to +At the first publication, the PR production source, tests, scripts and Cargo files +were byte-identical to that candidate. The follow-up records the current qualified +source separately. All six Cellule packages remain locked to `c51dd121284ecc8878b75d32717a4dfbe2c406c2`. Upstream main was rechecked as `0dc04a658bd99668936f7ec58032d054f6fbc141`; separate qualification is required. @@ -148,7 +155,8 @@ This candidate has not been deployed and cannot be credited with fixing it. Required next proofs remain: -- Safe existing-catalog admission, old-code restore and actual same-corpus upgrade. +- Actual old-executable RustFS restore and fully admitted same-corpus upgrade; + the later catalog admission regression is not that proof. - Full 10,000-identity/100-fixture Git/LFS verification and fresh-owner recovery. - The complete schedule, critical concurrent workflows and faults, and every new acknowledged write after owner loss. diff --git a/docs/performance/2026-10-01-retained-catalog-admission.md b/docs/performance/2026-10-01-retained-catalog-admission.md new file mode 100644 index 0000000..ebcf560 --- /dev/null +++ b/docs/performance/2026-10-01-retained-catalog-admission.md @@ -0,0 +1,151 @@ +# Retained catalog admission and startup safety + +The candidate reopens a supported predecessor Directory Cell after explicit fixture +activation without rewriting its immutable catalog identity. It also rejects +unsupported persisted Control code/schema before ownership changes. PR #18 +still does not provide an automatic upgrade controller or establish old-binary +RustFS upgrade, full-corpus recovery or reference performance. + +This extends the [authentication and residency checkpoint](2026-10-01-bounded-authentication-and-residency.md). +The Cellule pin remains `c51dd121284ecc8878b75d32717a4dfbe2c406c2`. +Upstream main was rechecked as `0dc04a658bd99668936f7ec58032d054f6fbc141`; +qualifying it remains a separate gate. + +## Existing identity is admitted without reprovisioning + +The real startup regression failed with `CatalogCollision` in all three +pre-fix attempts. Startup tried to provision the current Directory module code +against an existing entry whose initial code is immutable. Descriptor retention +alone could not repair this call-site behavior. + +New entries still use Cellule's strict `ReleaseStore::provision` path and current +code. Existing entries instead require a verified proof, the exact compiled +descriptor and a Ready release with matching current/desired digests. The target, +partition, namespace, SQL role and initial code/schema must be supported. +Admission rereads the complete release record and rejects any change, including +a transition back to Ready with the same digest. + +```mermaid +sequenceDiagram + participant S as Canopy startup + participant C as Catalog + participant R as Selected release + participant A as Cell authority + participant X as Runtime + S->>C: Lookup derived Cell identity + C-->>S: Existing verified immutable proof + S->>R: Require exact Ready release and descriptor + S->>S: Validate target and supported catalog code/schema + S->>R: Recheck complete release record + R-->>S: Unchanged record + S->>A: Load persisted Control + A-->>S: Observed code, schema and ownership + alt Control code/schema supported + S->>X: Acquire using existing proof and observed Control + X-->>S: Restored Cell + else Unsupported Control + S-->>S: Reject without claiming ownership + end +``` + +The checks do not select a release, migrate schema, rewrite catalog identity or +remove the runtime's ownership CAS and fencing requirements. Release enrollment +and readiness checks remain in force. + +## Refusal must not change unsupported Control metadata + +The initial read-only catalog fix passed retained startup but failed a stronger +safety test. Both unsupported-code and unsupported-schema cases failed in all +three attempts: startup returned an error only after runtime acquisition had +changed Control epoch/revision. + +The guard now checks the actual persisted Control against the registry before +bootstrap, takeover or restoration. The unchanged three-case regression then +passed in all three attempts. Its negative cases compare the complete canonical +Control bytes before and after refused startup, not merely the returned error. + +| Regression | Before correction | After correction | +| --- | --- | --- | +| Supported retained Directory startup | CatalogCollision in all three attempts | Restored persisted token, authenticated it and created/listed a repository | +| Unknown persisted Control code | Startup refused but Control changed in all three attempts | Refused with identical Control bytes in all three attempts | +| Unknown persisted Control schema | Startup refused but Control changed in all three attempts | Refused with identical Control bytes in all three attempts | +| Release and catalog admission guards | New targeted tests | All five passed: current/retained proof, missing/non-Ready release, corrupt descriptor, unsupported identity/wrong target and Ready round trip | + +These tests use owned in-memory stores. The predecessor descriptor is the exact +recorded fixture, but the fixture writes its data with the current test runtime. +It is not evidence that an old executable produced those bytes. Its explicit +activation occurs only after fixture-specific admission and with no advertised +writers; it must not be copied as a live upgrade procedure. + +## Combined qualification + +The frozen tested source is `402e9f93e21c576fc430cd20bbac16a4489d1558`. +PR production files are checked byte-for-byte against that source before publication. + +| Check | Closed result | Scope limit | +| --- | --- | --- | +| Locked release build and lints | Passed; all targets checked with warnings denied | Not a live deployment | +| Release workspace | 239 top-level tests passed, zero failed, nine ignored | Nested subprocess tests counted once; ignored provider/size gates are separate | +| Directory integration | All 12 passed | Includes bounded authentication and descriptor admission | +| Read-only catalog admission | All five unit regressions passed | Owned in-memory release/catalog faults | +| RustFS compatibility | All eight exact gates passed; closed 23:20:36 UTC | Disposable fresh fixture, not retained old-binary upgrade | +| Original cold activation | All 20 independent repetitions passed | Unchanged test against retained production artifact | +| Retained startup and refusal | All three cases passed in all 20 independent repetitions | Current test runtime wrote the predecessor fixture | +| Complete residency suite | All 15 passed at four threads; repetitions closed 23:22:14 UTC | Not live corpus recovery or capacity measurement | +| Python harness | All 84 passed | Accounting and guard coverage | + +Reproduce the focused checks and provider gates with: + +```sh +cargo test --release --locked -p canopy-server --lib server::catalog_admission::tests:: +cargo test --release --locked -p canopy-server --test multi_server retained_catalog:: +cargo test --release --locked --workspace -- --test-threads=4 +cargo clippy --release --workspace --all-targets --locked -- -D warnings +python3 -B scripts/qualify_size.py --provider-only --release +``` + +The RustFS command owns its disposable fixture. It does not qualify the existing +10,000-repository store, sustained load or the non-sparse five-GiB transfer. + +## Preserved evidence and remaining gates + +The failed catalog reproduction is retained at +`/Users/haipingfu/.codex/canopy-retained-catalog-before-ZOZqDQ`. +The failed unsupported-Control tests are at +`/Users/haipingfu/.codex/canopy-retained-control-before-kRqUeP`; +the corrected three repetitions are at +`/Users/haipingfu/.codex/canopy-retained-control-after-G4eQeG`. +Closed artifacts and exact source bindings were copied and independently reread +on a different local filesystem before subsequent source changes. These are +local backups, not off-machine copies. + +Combined release evidence is at +`/Users/haipingfu/.codex/canopy-catalog-admission-qualification-fsvZZ8`. +Its 290 files passed copy and independent reread checks at +`/Volumes/Workspace/CrabData/canopy-catalog-correctness-evidence-w_ouq9tq`. +Provider/repetition evidence is at +`/Users/haipingfu/.codex/canopy-catalog-provider-gates-p7WF68`; +its 328 files passed the same checks at +`/Volumes/Workspace/CrabData/canopy-catalog-provider-residency-evidence-sx7hfpq4`. + +| Closed artifact | SHA-256 | +| --- | --- | +| Combined executable | `04ac334406120d82162c6d2c47e693eaaf4b88215318e8b9b9cf14acd8a2da91` | +| `build-tests.json` | `021f9348bc4377a08339809fe37e696c486802ef0257ec3ca657f6c36f8e299c` | +| `release-workspace.log` | `40e0a801dfafe30d779953ae65d5bc6987a54a060660feaf8fc557d30c9f3098` | +| `provider-tests.json` | `d417e4743ac37bbbb6bd11e3b5441edeeaa76cae58411fa78ecefce7ebd34f5f` | +| `provider-tests.log` | `a0c6a05e14d47ef2362948ed65b154625e8b9b68418241ea2918f1d13ce4a20e` | +| `residency-repetitions.json` | `0d9531a6b1100cbe0423fa14c923b479965e17a257849ea4c1a8e100a32a50a9` | + +The existing corpus and UI preview were not upgraded by these checks. Required +gates remain actual old-executable to new-executable restore against RustFS, +fully admitted same-corpus upgrade, full Git/LFS recovery, and an explanation of +the separate diagnostic fleet's terminal lease-fencing failure. No owner restart +or release change is credited as resolving that failure. + +The full [performance plan](../performance-plan.md) is unchanged: three nodes +behind a proxy, 10,000 identities, 100 populated Git/LFS fixtures, all 108 windows, +114,960 arrivals and 8,640 scheduled seconds, critical concurrent workflows and +faults, every acknowledged write after owner loss, higher admission profiles, +matched comparisons and isolated Linux qualification. No matched speedup is +claimed; PR #18 remains draft. From 5f737092e4c3ac74012fa3d78e76a2799848b874 Mon Sep 17 00:00:00 2001 From: forhappy Date: Thu, 1 Oct 2026 16:29:25 -0700 Subject: [PATCH 43/60] Upgrade Cellule to origin main 0dc04a658 --- Cargo.lock | 12 ++++++------ crates/canopy-server/Cargo.toml | 10 +++++----- 2 files changed, 11 insertions(+), 11 deletions(-) diff --git a/Cargo.lock b/Cargo.lock index 40b95af..4d07109 100644 --- a/Cargo.lock +++ b/Cargo.lock @@ -448,7 +448,7 @@ dependencies = [ [[package]] name = "cellule-app" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=c51dd121284ecc8878b75d32717a4dfbe2c406c2#c51dd121284ecc8878b75d32717a4dfbe2c406c2" +source = "git+https://github.com/crabbuild/cellule.git?rev=0dc04a658bd99668936f7ec58032d054f6fbc141#0dc04a658bd99668936f7ec58032d054f6fbc141" dependencies = [ "blake3", "cellule-runtime", @@ -457,7 +457,7 @@ dependencies = [ [[package]] name = "cellule-host" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=c51dd121284ecc8878b75d32717a4dfbe2c406c2#c51dd121284ecc8878b75d32717a4dfbe2c406c2" +source = "git+https://github.com/crabbuild/cellule.git?rev=0dc04a658bd99668936f7ec58032d054f6fbc141#0dc04a658bd99668936f7ec58032d054f6fbc141" dependencies = [ "cellule-app", "cellule-runtime", @@ -471,7 +471,7 @@ dependencies = [ [[package]] name = "cellule-ltx" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=c51dd121284ecc8878b75d32717a4dfbe2c406c2#c51dd121284ecc8878b75d32717a4dfbe2c406c2" +source = "git+https://github.com/crabbuild/cellule.git?rev=0dc04a658bd99668936f7ec58032d054f6fbc141#0dc04a658bd99668936f7ec58032d054f6fbc141" dependencies = [ "async-trait", "blake3", @@ -493,7 +493,7 @@ dependencies = [ [[package]] name = "cellule-runtime" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=c51dd121284ecc8878b75d32717a4dfbe2c406c2#c51dd121284ecc8878b75d32717a4dfbe2c406c2" +source = "git+https://github.com/crabbuild/cellule.git?rev=0dc04a658bd99668936f7ec58032d054f6fbc141#0dc04a658bd99668936f7ec58032d054f6fbc141" dependencies = [ "blake3", "bytes", @@ -519,7 +519,7 @@ dependencies = [ [[package]] name = "cellule-store" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=c51dd121284ecc8878b75d32717a4dfbe2c406c2#c51dd121284ecc8878b75d32717a4dfbe2c406c2" +source = "git+https://github.com/crabbuild/cellule.git?rev=0dc04a658bd99668936f7ec58032d054f6fbc141#0dc04a658bd99668936f7ec58032d054f6fbc141" dependencies = [ "async-trait", "blake3", @@ -542,7 +542,7 @@ dependencies = [ [[package]] name = "cellule-types" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=c51dd121284ecc8878b75d32717a4dfbe2c406c2#c51dd121284ecc8878b75d32717a4dfbe2c406c2" +source = "git+https://github.com/crabbuild/cellule.git?rev=0dc04a658bd99668936f7ec58032d054f6fbc141#0dc04a658bd99668936f7ec58032d054f6fbc141" dependencies = [ "schemars", "serde", diff --git a/crates/canopy-server/Cargo.toml b/crates/canopy-server/Cargo.toml index 439cbbf..66fe98c 100644 --- a/crates/canopy-server/Cargo.toml +++ b/crates/canopy-server/Cargo.toml @@ -15,11 +15,11 @@ axum = "0.8.9" base64 = "0.22" blake3 = "1.8" bytes = "1.11" -cellule-app = { git = "https://github.com/crabbuild/cellule.git", rev = "c51dd121284ecc8878b75d32717a4dfbe2c406c2" } -cellule-host = { git = "https://github.com/crabbuild/cellule.git", rev = "c51dd121284ecc8878b75d32717a4dfbe2c406c2" } -cellule-ltx = { git = "https://github.com/crabbuild/cellule.git", rev = "c51dd121284ecc8878b75d32717a4dfbe2c406c2", features = ["replica"] } -cellule-runtime = { git = "https://github.com/crabbuild/cellule.git", rev = "c51dd121284ecc8878b75d32717a4dfbe2c406c2" } -cellule-store = { git = "https://github.com/crabbuild/cellule.git", rev = "c51dd121284ecc8878b75d32717a4dfbe2c406c2" } +cellule-app = { git = "https://github.com/crabbuild/cellule.git", rev = "0dc04a658bd99668936f7ec58032d054f6fbc141" } +cellule-host = { git = "https://github.com/crabbuild/cellule.git", rev = "0dc04a658bd99668936f7ec58032d054f6fbc141" } +cellule-ltx = { git = "https://github.com/crabbuild/cellule.git", rev = "0dc04a658bd99668936f7ec58032d054f6fbc141", features = ["replica"] } +cellule-runtime = { git = "https://github.com/crabbuild/cellule.git", rev = "0dc04a658bd99668936f7ec58032d054f6fbc141" } +cellule-store = { git = "https://github.com/crabbuild/cellule.git", rev = "0dc04a658bd99668936f7ec58032d054f6fbc141" } ed25519-dalek = "2" flate2 = "1.1" futures-core = "0.3" From 240203a7a2441db8b03d59e00de0b09b77e8c4a2 Mon Sep 17 00:00:00 2001 From: forhappy Date: Thu, 1 Oct 2026 17:10:01 -0700 Subject: [PATCH 44/60] Document latest Cellule qualification and open upgrade gates --- docs/contracts.md | 7 +- docs/delivery-plan.md | 7 +- docs/performance-plan.md | 16 +-- ...01-bounded-authentication-and-residency.md | 9 +- .../2026-10-01-cellule-main-qualification.md | 101 ++++++++++++++++++ .../2026-10-01-retained-catalog-admission.md | 10 +- 6 files changed, 132 insertions(+), 18 deletions(-) create mode 100644 docs/performance/2026-10-01-cellule-main-qualification.md diff --git a/docs/contracts.md b/docs/contracts.md index 9278a2a..61baa01 100644 --- a/docs/contracts.md +++ b/docs/contracts.md @@ -1194,7 +1194,7 @@ require a fresh development storage prefix; there is no upgrade reader for older development databases. The module descriptor and object paths will become compatibility boundaries at the first persistent preview. The current source pins Cellule revision -`c51dd121284ecc8878b75d32717a4dfbe2c406c2`; qualification artifacts +`0dc04a658bd99668936f7ec58032d054f6fbc141`; qualification artifacts remain bound to their recorded revisions, not silently replaced by this pin. The earlier entity-partition cutover also made pre-cutover prefixes incompatible; no migration is available. @@ -2794,10 +2794,13 @@ remain required. No new configuration surface was added for diagnostic output. Canopy directly uses `cellule-app`, `cellule-host`, `cellule-runtime`, `cellule-ltx` and `cellule-store`, pinned to Cellule commit -`a28de7bc09ce36d87e642adc4f4b6be50d6fcb69`. `cellule-types` is transitive. +`0dc04a658bd99668936f7ec58032d054f6fbc141`. `cellule-types` is transitive. The lockfile contains no Crab Cell, Crab product/server or Xet packages. Historical qualification runs in the delivery/performance logs retain their original dependency revisions; they are not performance evidence for this build. +The [current revision checkpoint](performance/2026-10-01-cellule-main-qualification.md) +records release correctness and fresh RustFS compatibility separately from the +still-open retained-store upgrade and performance gates. The application declares one entity-partitioned SQL Cell type for repositories. `repository_target` validates the canonical UUID, then calls the app crate's diff --git a/docs/delivery-plan.md b/docs/delivery-plan.md index db125c8..eb88c9d 100644 --- a/docs/delivery-plan.md +++ b/docs/delivery-plan.md @@ -72,7 +72,7 @@ These are the next implementation slices. Each slice must update its contract, b 1. Expand the S3-compatible process smoke into a node/lease fault matrix and test the target production object store. Run the checked-in CI workflow on - a Canopy remote. The current source pins Cellule revision `c51dd121` and + a Canopy remote. The current source pins Cellule revision `0dc04a6` and owns a startup storage probe with cleanup on failure. Historical dependency qualification runs below remain evidence for their recorded revisions only. The [live three-node trial](performance/2026-10-01-native-filesystem.md) @@ -83,7 +83,10 @@ These are the next implementation slices. Each slice must update its contract, b [retained catalog follow-up](performance/2026-10-01-retained-catalog-admission.md) adds read-only existing-proof admission and rejects unsupported persisted Control metadata before acquisition. Its 239-test release suite, eight RustFS - gates and repeated startup/residency regressions passed. Actual old-binary + gates and repeated startup/residency regressions passed. The + [latest Cellule qualification](performance/2026-10-01-cellule-main-qualification.md) + reruns these checks on the new immutable revision without changing runtime + budgets. Actual old-binary RustFS upgrade, full-corpus recovery and performance remain open. 2. Execute the [repository-density performance plan](performance-plan.md), then qualify the bounded Linux deployment and residency under faults and larger diff --git a/docs/performance-plan.md b/docs/performance-plan.md index 7fe78c1..8ce40d5 100644 --- a/docs/performance-plan.md +++ b/docs/performance-plan.md @@ -2,8 +2,12 @@ Use this plan to design capacity work and interpret Canopy benchmark results. It separates demonstrated behavior from proposed targets. The current Cellule -dependency pin is `c51dd121284ecc8878b75d32717a4dfbe2c406c2`, observed upstream -main at build preparation on October 1. Unlike the earlier `a4500add` documentation +dependency pin is `0dc04a658bd99668936f7ec58032d054f6fbc141`, verified against +upstream main on October 1. Its [qualification checkpoint](performance/2026-10-01-cellule-main-qualification.md) +records 239 release Rust tests, 84 Python tests, lints and all eight fresh RustFS +gates, with separately retained artifacts. Retained-store upgrade and reference +performance remain open. The preceding `c51dd121` build is historical evidence. +Unlike the earlier `a4500add` documentation advance, this changes routing and resource-permit implementation. Locked metadata resolves all six packages; the lockfile changes only their sources. The original rebuilt artifact passed its locked release build and 227 release @@ -13,8 +17,8 @@ and a proxy. The 84 Python harness tests cover accounting and guards. Full-corpu recovery, scheduled load and every-ACK verification remain separate gates; none of these functional checks establishes reference capacity. Upstream subsequently advanced to `0dc04a6` with atomic API and test changes. -That revision needs a separate artifact and runtime qualification; the current -experiment and its frozen source remain on `c51dd121`. +The new candidate is independently qualified; the original corpus experiment +and its frozen source remain on `c51dd121`. The [combined authentication and residency candidate](performance/2026-10-01-bounded-authentication-and-residency.md) is now in PR #18: 231 release tests and all eight RustFS gates passed, as did 20 unchanged cold-activation repetitions and all 15 residency tests. It has not @@ -26,14 +30,14 @@ unsupported Control metadata. Its combined candidate passed 239 release tests, all eight RustFS gates, 20 cold-activation repetitions and 20 three-case retained startup repetitions, plus all 15 residency tests. Retained startup uses an owned in-memory fixture; actual old-binary RustFS upgrade remains open. -These results do not qualify newer upstream -or replace the original complete performance schedule. +Neither checkpoint replaces the original complete performance schedule. The [three-node proxy campaign](performance/2026-09-30-three-node-proxy.md) retains the earlier fixture race and its regression, plus the candidate's failed seed and successful recovery of every recorded ACK. | Current gate | Evidence / status | | --- | --- | +| Current Cellule revision | `0dc04a6`: independently rebuilt release correctness, lints, Python harness and eight fresh RustFS gates passed. See the latest checkpoint for repeated regressions and exact bindings. Not retained-store upgrade, live deployment or measured performance | | Original rebuilt Cellule artifact | `c51dd121` build, 227 Rust tests, eight RustFS gates and initial 17-step proxy check passed; full 10K/100 seed closed, but remote verification failed HTTP 503. A separate 10K metadata diagnostic returned 9,999 matching identities and one 503; downstream fault/load gates stopped without owner signals. Cause unresolved | | Combined authentication and residency candidate | First publication passed 231 release tests, eight RustFS gates, 20 cold-activation repetitions and 15 residency tests. Exact predecessor descriptor is retained; see the separately qualified catalog follow-up | | Retained catalog admission follow-up | 239 release tests, eight RustFS gates, 20 cold-activation and 20 three-case retained-startup repetitions, 15 residency tests, lints and 84 Python tests passed. Existing supported identities restore without reprovisioning; unsupported Controls stay unchanged on refusal. Owned in-memory retained fixture, not old-binary RustFS upgrade, live deployment or performance qualification | diff --git a/docs/performance/2026-10-01-bounded-authentication-and-residency.md b/docs/performance/2026-10-01-bounded-authentication-and-residency.md index 84aacd8..0006e02 100644 --- a/docs/performance/2026-10-01-bounded-authentication-and-residency.md +++ b/docs/performance/2026-10-01-bounded-authentication-and-residency.md @@ -94,9 +94,10 @@ activation CAS as compatibility proof. The frozen source is local commit `392a343b83a492d49c08a6f3e83ff4dd2917437f`. At the first publication, the PR production source, tests, scripts and Cargo files were byte-identical to that candidate. The follow-up records the current qualified -source separately. All six Cellule packages remain locked to -`c51dd121284ecc8878b75d32717a4dfbe2c406c2`. Upstream main was rechecked as -`0dc04a658bd99668936f7ec58032d054f6fbc141`; separate qualification is required. +source separately. All six Cellule packages in this checkpoint were locked to +`c51dd121284ecc8878b75d32717a4dfbe2c406c2`. The +[later Cellule qualification](2026-10-01-cellule-main-qualification.md) records +`0dc04a658bd99668936f7ec58032d054f6fbc141` with its own artifacts and results. | Check | Closed result | Scope limit | | --- | --- | --- | @@ -160,7 +161,7 @@ Required next proofs remain: - Full 10,000-identity/100-fixture Git/LFS verification and fresh-owner recovery. - The complete schedule, critical concurrent workflows and faults, and every new acknowledged write after owner loss. -- Newer Cellule qualification, higher admission profiles, matched comparisons, +- Higher admission profiles, matched comparisons, non-sparse five-GiB transfer and complete CPU/cost accounting. - Isolated Linux reference capacity; the shared Mac/Colima environment does not establish it. Missing historical ledgers still prevent old every-ACK proof. diff --git a/docs/performance/2026-10-01-cellule-main-qualification.md b/docs/performance/2026-10-01-cellule-main-qualification.md new file mode 100644 index 0000000..0b34db3 --- /dev/null +++ b/docs/performance/2026-10-01-cellule-main-qualification.md @@ -0,0 +1,101 @@ +# Cellule main revision qualification + +Canopy now pins all six Cellule packages to +`0dc04a658bd99668936f7ec58032d054f6fbc141`, verified against upstream main on +October 1. Release correctness, fresh RustFS compatibility and repeated +regressions passed. This does not establish a retained-store upgrade, recovery +of the existing corpus, reference capacity or a performance improvement. + +This follows the [retained catalog admission checkpoint](2026-10-01-retained-catalog-admission.md). +That checkpoint and the [authentication and residency results](2026-10-01-bounded-authentication-and-residency.md) +retain their original Cellule revision and artifacts. + +## Dependency change and source equivalence + +The tested source is `bbd784a40c3867646044a8c716b7ee517f9aca49`. The dependency +change replaces five direct manifest pins and six lockfile source revisions; +it changes no other dependency versions, Canopy implementation or runtime budgets. +Locked metadata resolves all six Cellule packages to the same immutable commit. +Before publication, all 263 production, test, script and Cargo files on the PR +branch are checked against this frozen source. + +The new revision includes atomic API and upstream test changes. Passing +functional tests is not evidence of a throughput or latency improvement. + +## Closed verification results + +| Check | Result | Verification boundary | +| --- | --- | --- | +| Locked release build | Passed | Executable retained outside Cargo targets | +| Release workspace | 239 top-level tests passed; zero failed, nine ignored | Nested subprocess tests counted once; ignored provider and size gates are separate | +| Release lints | All targets passed with warnings denied | No deployment implied | +| Authentication overlap | 16 admitted, zero refused, 4,673 retained bytes | Held-worker regression, not throughput | +| Directory integration | All 12 passed | Includes bounded authentication and predecessor descriptor admission | +| Catalog admission guards | All five passed | Release and catalog fault coverage | +| Real RustFS compatibility | All eight exact gates passed | Disposable fresh fixture, not retained-store upgrade or five-GiB transfer | +| Cold activation | All 20 independent repetitions passed | Unchanged regression and exact retained test executable | +| Retained startup and refusal | All three cases passed in each of 20 independent repetitions | Owned in-memory predecessor fixture written by the current test runtime | +| Complete residency suite | All 15 passed at four threads | Not live corpus recovery or capacity | +| Python harness | All 84 passed on the same source | Accounting and qualification guards | + +Release verification closed October 1 at 23:52:15 UTC; provider gates closed at +23:59:32 UTC. Repeated regressions closed October 2 at 00:06:16 UTC, still +October 1 in the local Pacific timezone. No assertion was weakened, no HTTP +retry was added, and no runtime budget was raised for these checks. + +The eight provider gates cover SHA-256 round trips and native merge candidates, +signed push options, signed SHA-256 SSH, stock SSH, bulk mirror refs, filtered +clones and stock Git LFS over SSH. The existing 10,000-repository RustFS provider +had identical provider snapshots before and after these gates. Neither its +corpus nor the UI preview was upgraded. + +Reproduce the checks from a checkout with this lockfile: + +```sh +cargo build --release --locked --workspace +cargo test --release --locked --workspace -- --test-threads=4 +cargo clippy --release --locked --workspace --all-targets -- -D warnings +cargo test --release --locked -p canopy-server --lib server::catalog_admission::tests:: +cargo test --release --locked -p canopy-server --test multi_server retained_catalog:: +python3 -B scripts/qualify_size.py --provider-only --release +``` + +The provider command owns a disposable fixture. It is not an upgrade command. + +## Preserved artifacts + +Release and harness evidence is retained at +`/Users/haipingfu/.codex/canopy-cellule-main-qualification-WKtopc`. +Its 290 files were copied and independently reread on a different local filesystem +at `/Volumes/Workspace/CrabData/canopy-cellule-main-correctness-evidence-gxxlwgeb`. +Provider and repetition evidence is retained at +`/Users/haipingfu/.codex/canopy-cellule-main-provider-gates-4480qw`. Its 328 files +were copied and independently reread at +`/Volumes/Workspace/CrabData/canopy-cellule-main-provider-evidence-14_2ewqe`. +These are local, not off-machine backups. + +| Closed artifact | SHA-256 | +| --- | --- | +| Release executable | `a541cef36acae0b21a651316fda2e9b3204ef7ea3f02173b65a7c217cd6026e9` | +| `build-tests.json` | `b8a7b2e8005dac79aa19a10423095555c0b2feba9faf8efb6614be22a489ceb4` | +| `release-workspace.log` | `67fbad383f5a6c2b678bcd9d24144cc945dd52f25092d062cd4f110617be8686` | +| `provider-tests.json` | `7e3e87acd46506561d9051d5fe638d4fd30732cb79fd0583337955943192071e` | +| `provider-tests.log` | `f3b384520a7d926f55e82b5aadf435671ce0403bbf4088a28972cd93c473e902` | +| `residency-repetitions.json` | `ae1bc03205fdb26cb892a589e34452185bcede24f42255fd462335364274b9d1` | + +## Remaining upgrade and performance gates + +PR #18 remains draft. The actual old-executable RustFS fixture has not reached +upgrade admission: an attempted reserved-ref push was correctly rejected, and +a subsequent fresh clone failed the exact remote Git/LFS byte comparison. +Partial acknowledged writes and failed attempts remain retained. The byte +mismatch is unresolved; no old-to-new RustFS restore pass is claimed. + +The [full performance plan](../performance-plan.md) still requires three nodes +behind a proxy, 10,000 identities, 100 populated Git/LFS fixtures, all 108 windows, +114,960 arrivals and 8,640 scheduled seconds. Full remote verification, +same-corpus upgrade admission, concurrent critical operations and faults, +every acknowledged write after owner loss, higher admission profiles, matched +comparisons, non-sparse five-GiB transfers and isolated Linux capacity remain open. +The separate diagnostic fleet's terminal lease-fencing failure is also unresolved. +New passing suites do not explain it or certify missing historical recovery ledgers. diff --git a/docs/performance/2026-10-01-retained-catalog-admission.md b/docs/performance/2026-10-01-retained-catalog-admission.md index ebcf560..eca24c1 100644 --- a/docs/performance/2026-10-01-retained-catalog-admission.md +++ b/docs/performance/2026-10-01-retained-catalog-admission.md @@ -7,9 +7,10 @@ still does not provide an automatic upgrade controller or establish old-binary RustFS upgrade, full-corpus recovery or reference performance. This extends the [authentication and residency checkpoint](2026-10-01-bounded-authentication-and-residency.md). -The Cellule pin remains `c51dd121284ecc8878b75d32717a4dfbe2c406c2`. -Upstream main was rechecked as `0dc04a658bd99668936f7ec58032d054f6fbc141`; -qualifying it remains a separate gate. +The results below use Cellule `c51dd121284ecc8878b75d32717a4dfbe2c406c2`. +The [subsequent Cellule qualification](2026-10-01-cellule-main-qualification.md) +records the new revision and its independently rebuilt artifacts; it does not +replace this historical evidence. ## Existing identity is admitted without reprovisioning @@ -80,7 +81,8 @@ writers; it must not be copied as a live upgrade procedure. ## Combined qualification The frozen tested source is `402e9f93e21c576fc430cd20bbac16a4489d1558`. -PR production files are checked byte-for-byte against that source before publication. +At this checkpoint's publication, PR production files were checked byte-for-byte +against that source. Later dependency qualification has separate source bindings. | Check | Closed result | Scope limit | | --- | --- | --- | From 0498e22d05651de4c645fe247c494192ad7e1f43 Mon Sep 17 00:00:00 2001 From: forhappy Date: Thu, 1 Oct 2026 17:44:06 -0700 Subject: [PATCH 45/60] Document old-binary RustFS upgrade and critical workflow recovery --- docs/contracts.md | 6 +- docs/delivery-plan.md | 6 +- docs/performance-plan.md | 3 +- .../2026-10-01-cellule-main-qualification.md | 11 +- .../2026-10-01-old-binary-rustfs-upgrade.md | 154 ++++++++++++++++++ .../2026-10-01-retained-catalog-admission.md | 8 +- 6 files changed, 175 insertions(+), 13 deletions(-) create mode 100644 docs/performance/2026-10-01-old-binary-rustfs-upgrade.md diff --git a/docs/contracts.md b/docs/contracts.md index 61baa01..2de4324 100644 --- a/docs/contracts.md +++ b/docs/contracts.md @@ -2799,8 +2799,10 @@ The lockfile contains no Crab Cell, Crab product/server or Xet packages. Historical qualification runs in the delivery/performance logs retain their original dependency revisions; they are not performance evidence for this build. The [current revision checkpoint](performance/2026-10-01-cellule-main-qualification.md) -records release correctness and fresh RustFS compatibility separately from the -still-open retained-store upgrade and performance gates. +records release correctness and fresh RustFS compatibility separately from +performance. The [owned retained-store qualification](performance/2026-10-01-old-binary-rustfs-upgrade.md) +records actual old-executable restore and fresh-owner recovery; it does not +provide a general upgrade controller or qualify the existing 10K corpus. The application declares one entity-partitioned SQL Cell type for repositories. `repository_target` validates the canonical UUID, then calls the app crate's diff --git a/docs/delivery-plan.md b/docs/delivery-plan.md index eb88c9d..005c6d1 100644 --- a/docs/delivery-plan.md +++ b/docs/delivery-plan.md @@ -86,8 +86,10 @@ These are the next implementation slices. Each slice must update its contract, b gates and repeated startup/residency regressions passed. The [latest Cellule qualification](performance/2026-10-01-cellule-main-qualification.md) reruns these checks on the new immutable revision without changing runtime - budgets. Actual old-binary - RustFS upgrade, full-corpus recovery and performance remain open. + budgets. The [owned old-binary RustFS qualification](performance/2026-10-01-old-binary-rustfs-upgrade.md) + passed controlled upgrade, twenty scheduled critical workflows and complete + recovery of that fixture after three-owner loss. Original-corpus upgrade, + full reference performance and the CI port-handoff failure remain open. 2. Execute the [repository-density performance plan](performance-plan.md), then qualify the bounded Linux deployment and residency under faults and larger hot sets. Each node reserves one SQL slot for Directory takeover and admits diff --git a/docs/performance-plan.md b/docs/performance-plan.md index 8ce40d5..ecec26b 100644 --- a/docs/performance-plan.md +++ b/docs/performance-plan.md @@ -5,7 +5,7 @@ It separates demonstrated behavior from proposed targets. The current Cellule dependency pin is `0dc04a658bd99668936f7ec58032d054f6fbc141`, verified against upstream main on October 1. Its [qualification checkpoint](performance/2026-10-01-cellule-main-qualification.md) records 239 release Rust tests, 84 Python tests, lints and all eight fresh RustFS -gates, with separately retained artifacts. Retained-store upgrade and reference +gates, with separately retained artifacts. Original-corpus upgrade and reference performance remain open. The preceding `c51dd121` build is historical evidence. Unlike the earlier `a4500add` documentation advance, this changes routing and resource-permit implementation. Locked @@ -38,6 +38,7 @@ seed and successful recovery of every recorded ACK. | Current gate | Evidence / status | | --- | --- | | Current Cellule revision | `0dc04a6`: independently rebuilt release correctness, lints, Python harness and eight fresh RustFS gates passed. See the latest checkpoint for repeated regressions and exact bindings. Not retained-store upgrade, live deployment or measured performance | +| Owned old-binary RustFS upgrade | [Controlled fixture qualification](performance/2026-10-01-old-binary-rustfs-upgrade.md) passed actual old-byte restore through three fresh gateways/proxy, both new Git/LFS writes, twenty scheduled 17-step workflows and recovery of all forty critical repositories after all three owners were killed. Two retained repositories; not the original 10K corpus or reference capacity | | Original rebuilt Cellule artifact | `c51dd121` build, 227 Rust tests, eight RustFS gates and initial 17-step proxy check passed; full 10K/100 seed closed, but remote verification failed HTTP 503. A separate 10K metadata diagnostic returned 9,999 matching identities and one 503; downstream fault/load gates stopped without owner signals. Cause unresolved | | Combined authentication and residency candidate | First publication passed 231 release tests, eight RustFS gates, 20 cold-activation repetitions and 15 residency tests. Exact predecessor descriptor is retained; see the separately qualified catalog follow-up | | Retained catalog admission follow-up | 239 release tests, eight RustFS gates, 20 cold-activation and 20 three-case retained-startup repetitions, 15 residency tests, lints and 84 Python tests passed. Existing supported identities restore without reprovisioning; unsupported Controls stay unchanged on refusal. Owned in-memory retained fixture, not old-binary RustFS upgrade, live deployment or performance qualification | diff --git a/docs/performance/2026-10-01-cellule-main-qualification.md b/docs/performance/2026-10-01-cellule-main-qualification.md index 0b34db3..624aa52 100644 --- a/docs/performance/2026-10-01-cellule-main-qualification.md +++ b/docs/performance/2026-10-01-cellule-main-qualification.md @@ -85,11 +85,12 @@ These are local, not off-machine backups. ## Remaining upgrade and performance gates -PR #18 remains draft. The actual old-executable RustFS fixture has not reached -upgrade admission: an attempted reserved-ref push was correctly rejected, and -a subsequent fresh clone failed the exact remote Git/LFS byte comparison. -Partial acknowledged writes and failed attempts remain retained. The byte -mismatch is unresolved; no old-to-new RustFS restore pass is claimed. +PR #18 remains draft. The [subsequent owned-fixture upgrade and recovery](2026-10-01-old-binary-rustfs-upgrade.md) +verified actual old-executable RustFS data, controlled admission, new-executable +restore, twenty scheduled critical workflows and their complete fresh-owner +recovery. Its earlier clone-byte mismatch was traced to missing client LFS +filters; the full-byte assertion was retained. This closes a small prerequisite, +not full same-corpus upgrade or reference performance. The [full performance plan](../performance-plan.md) still requires three nodes behind a proxy, 10,000 identities, 100 populated Git/LFS fixtures, all 108 windows, diff --git a/docs/performance/2026-10-01-old-binary-rustfs-upgrade.md b/docs/performance/2026-10-01-old-binary-rustfs-upgrade.md new file mode 100644 index 0000000..ccb1466 --- /dev/null +++ b/docs/performance/2026-10-01-old-binary-rustfs-upgrade.md @@ -0,0 +1,154 @@ +# Old executable RustFS upgrade and critical workflow recovery + +The qualified Cellule `0dc04a6` executable restored data actually produced by the +old `c51dd121` executable on an owned RustFS fixture. Twenty scheduled critical +Git workflows passed through three nodes and a proxy, and all their acknowledged +state survived loss of all three owners and restoration from fresh local disks. + +This closes the small retained-store upgrade prerequisite. It does not qualify +the existing 10,000-repository corpus, its complete load matrix, isolated Linux +capacity or a speedup. The [latest Cellule checkpoint](2026-10-01-cellule-main-qualification.md) +records the separate release build, tests and provider gates. + +## Upgrade and recovery sequence + +```text +Old executable writes two Git/LFS repositories to owned RustFS + -> full-byte verification -> graceful drain -> same-code maintenance + -> full catalog and Control admission -> controlled new release activation + -> three fresh gateways + proxy -> restore old bytes -> verify new writes + -> 20 scheduled critical workflows -> preserve every receipt and arrival + -> SIGKILL all three gateways -> 32.008 seconds confirmed owner absence + -> three new gateways with fresh disks -> verify all retained and new state + -> graceful drain; all three recovery gateways exit zero +``` + +The producer used the exact old executable, not a current-runtime predecessor +simulation. Before activation, admission verified the actual old descriptor, +rolling compatibility, all 256 catalog shards, exactly the three expected Cells, +supported initial catalog identities and actual persisted Control code/schema. +Every Cell was durably idle and unowned; there were no advertised writers, +including expired advertisements. Repeated observations had identical catalog +revisions, immutable page digests, canonical Controls, identity and root purpose. + +The old executable ended its exact maintenance operation after confirmed drain. +The fixture-only tool then performed prepare, activation and completion, rescanning +the unchanged state at each boundary. It never rewrote immutable catalog identity +or bypassed runtime ownership CAS. This restricted tool is not a general upgrade +command and must not be applied blindly to another deployment. + +## Verification results + +| Gate | Closed result | Scope | +| --- | --- | --- | +| Old executable producer | Two repositories, two commits each, main/tag/custom refs and full LFS payloads verified | Actual retained old executable; each LFS body is 1,048,593 bytes | +| Old writer drain | Exit zero; zero advertised sessions and zero unsettled Cells | Same-code maintenance before upgrade | +| Upgrade admission | All 256 shards and three expected Cells passed; invalid mode and wrong endpoint refused | Exact owned fixture only | +| New executable restore | Git v0/v2 exact refs, ordinary clones, mirrors, full fsck and full LFS bytes passed | Three fresh gateways behind the proxy | +| New writes | Both fresh commits and LFS payloads acknowledged and freshly verified | New branches preserve all old main/tag/custom refs | +| Critical workflow schedule | 20/20 workflows and 340/340 steps passed; no failed or dropped arrivals | 300-second schedule, 15-second interval, concurrency cap four; observed peak overlap three | +| Owner loss | Exact three native gateway PIDs received SIGKILL; all absent for 32.007527 seconds | Both RustFS providers retained the same identity, start time and resource settings | +| Fresh-owner recovery | Both retained repositories, both new Git/LFS writes and all 40 critical repositories passed | Exact Git v0/v2 inventories, full fsck, payloads, notes and full old/new LFS bytes | +| Final drain | All three fresh recovery gateways exited zero without force | No UI preview or other fleet signalled | + +The initial byte-check failure was a client setup error: the producer installed +Git LFS only in its local source repository. A fresh clone without LFS filters +returned the correct 132-byte pointer and README, not the full payload. A real +stock-Git regression failed with user-global configuration disabled, then passed +with explicit LFS filters. The original full-byte assertion was retained. +The earlier reserved-ref attempt also remains recorded: main/tag were accepted, +while `refs/canopy/retained` was correctly rejected as server-owned. No partial +acknowledged state was discarded to make the fixture pass. + +## Throughput and latency + +This shared Mac/Colima run declared one complete workflow every 15 seconds. +All 20 completed inside the 300-second schedule: **0.066667 workflows/s**. +Including drain, elapsed time was 300.026007 seconds and throughput was +0.066661 workflows/s. This is the delivered rate at the declared arrival clock, +not a saturation capacity measurement. + +Each workflow includes two repository creations, atomic publication, mixed and +atomic refusal, correct/stale lease pushes, shallow/deepen/unshallow, filtered +lazy fetch under v0/v2, incremental push, fast-forward pull, branch deletion, +fetch-prune, mirror push, credential refusal and exact mirror/fsck verification. + +| Measured client wall time | p50 ms | p95 ms | p99 ms | +| --- | ---: | ---: | ---: | +| Complete scheduled workflow | 18,413.501 | 30,629.526 | 31,908.751 | +| Create source repository | 383.934 | 755.265 | 1,869.551 | +| Create mirror repository | 321.574 | 605.454 | 2,463.179 | +| Atomic multiple-ref publication | 942.296 | 1,993.050 | 2,479.839 | +| Incremental push | 720.945 | 1,541.330 | 1,726.267 | +| Fast-forward pull | 523.296 | 1,163.776 | 1,194.283 | +| Fetch-prune | 248.192 | 602.466 | 666.105 | +| Exact v0/v2 mirror and fsck checks | 3,683.990 | 5,051.374 | 5,529.689 | + +These include stock-Git processes, client work and validation. Step timings may +group multiple commands; they are not isolated RPC rates or server-only latency. +There are only 20 samples per step. This run supplies no matched baseline, +server CPU attribution, provider cost or isolated Linux capacity result. + +## Reproduction and retained evidence + +The checked-in workflow driver can reproduce the declared schedule on a +separately admitted, caller-owned fixture: + +```sh +python3 -B scripts/benchmark_critical_git.py \ + --fleet-dir /path/to/admitted-three-node-fleet --node-active-limit 100 \ + run --output-dir /path/to/new-critical-results \ + --duration 300 --interval 15 --concurrency 4 +``` + +Recovery additionally requires recorded owner loss, absence/expiry, unchanged +provider and fresh gateway workspaces. The verifier checks receipts; it does not +perform or prove those caller-controlled fault steps by itself. + +Evidence is retained at +`/Users/haipingfu/.codex/canopy-old-binary-rustfs-upgrade-n6W1G1`. Before owner loss, +7,388 closed files were copied and independently reread on another local filesystem +at `/Volumes/Workspace/CrabData/canopy-retained-upgrade-load-v9fs4p1c`. The final +13,292-file copy at +`/Volumes/Workspace/CrabData/canopy-retained-upgrade-recovery-4a3q8dyg` includes +closed recovery receipts, terminal fleets, exact source and failed attempts. +These are local, not off-machine backups. + +| Closed artifact | SHA-256 | +| --- | --- | +| Old producer receipt | `a861a56aa31aed2f9110fc26069c9c35c4a7a59179d687ceddd3f5d471f70cd4` | +| Upgrade admission | `37a152bbd70714974e1479e64dd73d7311c74b9d145c12bae83cee7d513978b1` | +| Upgrade activation | `aff121a9ab4eadcefe3c6ce04f46b8a15e770beee260738ec429a5c645566656` | +| Initial new-executable restore | `287a7519598ba3ab5bbeba2d5a48df95f6e582527ba69be18d939726df85e6f7` | +| Critical load report | `7857f31b1d52c88fe0335c7b84fa2120649abf2e92c6d53843bef3332fe83fe2` | +| Arrival ledger | `e880e348c62001480944b4f7462a3b14f19c93cd2c3189ccfe173a9fdf14cb57` | +| Owner-loss receipt | `4faf82d22f6f7e0efd43c2911b6def0f947523dc4ba96ecde8457d967846d973` | +| Post-loss verification | `010bb840514dc798b8bc83c69ef4c0399fb572c338380801dedc1585233ecd75` | + +The old executable SHA-256 is +`32b114119960608c0a91d1c783bb69eafec432831bfa452d54d8950b09bc0e99`; +the new qualified executable is +`a541cef36acae0b21a651316fda2e9b3204ef7ea3f02173b65a7c217cd6026e9`. +Production source still matches `bbd784a40c3867646044a8c716b7ee517f9aca49`. +The producer, upgrade, load and recovery ran October 2 at 00:13–00:36 UTC, +October 1 in the local Pacific timezone. + +## Remaining full campaign and CI work + +The original 10,000-repository provider was not upgraded or restarted. Its +read-only status observation reports the old Ready release, zero advertised +writers and **301 unsettled Cells**. It requires supported offline recovery and +full same-corpus admission before any new release or gateway launch. Restarting +the small fixture does not explain the original fleet's terminal lease failure. + +At PR head `240203a7`, both Python CI jobs and one complete Rust job passed. +The other Rust job failed with `AddrInUse` in the SHA-256 merge-candidate test. +The failure log is retained; a passing parallel job does not erase it. Port +reservation/handoff needs reproduction and correction before merge readiness. + +The [full performance plan](../performance-plan.md) remains unchanged: +10,000 identities, 100 populated Git/LFS fixtures, all 108 windows, 114,960 arrivals +and 8,640 scheduled seconds, independent operation rates/concurrency, uniform and +skewed active sets, ref-only/fresh-object pushes, higher admission profiles, +concurrent faults, full original-corpus/every-ACK recovery, matched comparisons, +non-sparse five-GiB transfers and isolated Linux qualification. PR #18 remains draft. diff --git a/docs/performance/2026-10-01-retained-catalog-admission.md b/docs/performance/2026-10-01-retained-catalog-admission.md index eca24c1..1cffe97 100644 --- a/docs/performance/2026-10-01-retained-catalog-admission.md +++ b/docs/performance/2026-10-01-retained-catalog-admission.md @@ -2,9 +2,11 @@ The candidate reopens a supported predecessor Directory Cell after explicit fixture activation without rewriting its immutable catalog identity. It also rejects -unsupported persisted Control code/schema before ownership changes. PR #18 -still does not provide an automatic upgrade controller or establish old-binary -RustFS upgrade, full-corpus recovery or reference performance. +unsupported persisted Control code/schema before ownership changes. This +checkpoint does not establish old-binary RustFS upgrade, full-corpus recovery or +reference performance. The [subsequent owned-fixture run](2026-10-01-old-binary-rustfs-upgrade.md) +records actual old-binary RustFS restore and recovery separately. PR #18 still +does not provide an automatic upgrade controller. This extends the [authentication and residency checkpoint](2026-10-01-bounded-authentication-and-residency.md). The results below use Cellule `c51dd121284ecc8878b75d32717a4dfbe2c406c2`. From eaebfc4cdd7ead54a8d62978aebcd8656f0de0b8 Mon Sep 17 00:00:00 2001 From: forhappy Date: Thu, 1 Oct 2026 18:20:18 -0700 Subject: [PATCH 46/60] fix: retain HTTP listener reservations through supervised startup The SHA-256 test helper dropped its listener before Canopy rebound the advertised address. Preserve ownership across startup, reject address mismatch before writes, and keep the existing supervisor, cancellation and drain semantics. Retain the deterministic pre-fix failure and separate new artifact verification from historical performance measurements. --- crates/canopy-server/src/server/lifecycle.rs | 22 +++- crates/canopy-server/src/server/mod.rs | 13 +- .../tests/multi_server/lifecycle/mod.rs | 71 ++++++++-- .../tests/multi_server/listener_handoff.rs | 108 +++++++++++++++ .../canopy-server/tests/multi_server/main.rs | 1 + .../tests/multi_server/sha256.rs | 36 +++-- docs/contracts.md | 5 + docs/delivery-plan.md | 8 +- docs/performance-plan.md | 6 + .../2026-10-01-cellule-main-qualification.md | 6 +- .../2026-10-01-listener-handoff.md | 123 ++++++++++++++++++ .../2026-10-01-old-binary-rustfs-upgrade.md | 10 +- 12 files changed, 376 insertions(+), 33 deletions(-) create mode 100644 crates/canopy-server/tests/multi_server/listener_handoff.rs create mode 100644 docs/performance/2026-10-01-listener-handoff.md diff --git a/crates/canopy-server/src/server/lifecycle.rs b/crates/canopy-server/src/server/lifecycle.rs index f3f985d..ec74dad 100644 --- a/crates/canopy-server/src/server/lifecycle.rs +++ b/crates/canopy-server/src/server/lifecycle.rs @@ -6,13 +6,33 @@ impl CanopyServer { pub async fn start( config: ServerConfig, store: Arc, + ) -> Result { + Self::start_supervised(config, store, None).await + } + + /// Starts a node using an already-bound HTTP listener, without rebinding it. + /// The listener's local address must exactly match `config.listen`; the + /// public URL may still name a proxy. All readiness and drain checks apply. + /// Cancelling startup requests cleanup after admitted initialization settles. + pub async fn start_with_listener( + config: ServerConfig, + store: Arc, + listener: TcpListener, + ) -> Result { + Self::start_supervised(config, store, Some(listener)).await + } + + async fn start_supervised( + config: ServerConfig, + store: Arc, + listener: Option, ) -> Result { let (ready, receive_ready) = oneshot::channel(); let (shutdown, receive_shutdown) = oneshot::channel(); // The task owns startup, drain and the workspace together. Dropping any // caller future only closes a channel; it cannot abandon admitted work. let finished = tokio::spawn(async move { - let server = match RunningServer::start(config, store).await { + let server = match RunningServer::start(config, store, listener).await { Ok(server) => server, Err(error) => { let _ = ready.send(Err(error)); diff --git a/crates/canopy-server/src/server/mod.rs b/crates/canopy-server/src/server/mod.rs index 12e7f70..4bb8f44 100644 --- a/crates/canopy-server/src/server/mod.rs +++ b/crates/canopy-server/src/server/mod.rs @@ -403,7 +403,15 @@ impl RunningServer { async fn start( config: ServerConfig, raw_store: Arc, + listener: Option, ) -> Result { + if let Some(listener) = &listener + && listener.local_addr()? != config.listen + { + return Err(ServerError::Http( + "HTTP listener address differs from listen configuration", + )); + } // Cellule permits 10,000 active Cells; reserve one for Directory takeover. if !(1..10_000).contains(&config.max_active_repositories) { return Err(ServerError::Http( @@ -462,7 +470,10 @@ impl RunningServer { identity.image, identity.release, ); - let listener = TcpListener::bind(config.listen).await?; + let listener = match listener { + Some(listener) => listener, + None => TcpListener::bind(config.listen).await?, + }; let mut config = config; let address = listener.local_addr()?; let ssh_listener = if let Some(ssh) = &config.ssh { diff --git a/crates/canopy-server/tests/multi_server/lifecycle/mod.rs b/crates/canopy-server/tests/multi_server/lifecycle/mod.rs index 2858cb9..e22e4c4 100644 --- a/crates/canopy-server/tests/multi_server/lifecycle/mod.rs +++ b/crates/canopy-server/tests/multi_server/lifecycle/mod.rs @@ -202,6 +202,41 @@ async fn wait_for_cleanup(data: &Path) -> Result { Ok(()) } +#[tokio::test(flavor = "multi_thread")] +async fn cancelled_prebound_startup_keeps_listener_and_workspace_until_cleanup() -> Result { + let files = tempfile::TempDir::new()?; + let data = files.path().join("node"); + let store = Arc::new(PausedStore::default()); + store.arm(ControlState::Serving); + let listener = TcpListener::bind("127.0.0.1:0").await?; + let address = listener.local_addr()?; + let start = tokio::spawn(CanopyServer::start_with_listener( + config(address, data.clone()), + store.clone(), + listener, + )); + store.wait().await?; + start.abort(); + assert!(start.await.is_err_and(|error| error.is_cancelled())); + assert!(matches!( + workspace_lock(&data)?.try_lock(), + Err(std::fs::TryLockError::WouldBlock) + )); + assert_eq!( + TcpListener::bind(address).await.unwrap_err().kind(), + std::io::ErrorKind::AddrInUse + ); + store.proceed.notify_one(); + wait_for_cleanup(&data).await?; + let listener = TcpListener::bind(address).await?; + let server = CanopyServer::start_with_listener(config(address, data), store, listener).await?; + create_repository(address, "after-cancellation").await?; + server.shutdown().await?; + let rebound = TcpListener::bind(address).await?; + assert_eq!(rebound.local_addr()?, address); + Ok(()) +} + #[tokio::test(flavor = "multi_thread")] async fn cancelled_startup_keeps_workspace_until_publication_and_cleanup_settle() -> Result { let files = tempfile::TempDir::new()?; @@ -564,17 +599,29 @@ mod lease; #[tokio::test(flavor = "multi_thread")] async fn startup_rejects_ignored_conditional_writes_before_enrollment() -> Result { - let store = Arc::new(PausedStore::default()); - store.ignore_conditions.store(true, Ordering::SeqCst); - let files = tempfile::TempDir::new()?; - let settings = config(available_address().await?, files.path().join("server")); - assert!(matches!( - CanopyServer::start(settings, store.clone()).await, - Err(canopy_server::server::ServerError::Repository( - "storage conditional create failed" - )) - )); - let remaining = store.list_with_delimiter(None).await?; - assert!(remaining.objects.is_empty() && remaining.common_prefixes.is_empty()); + for prebound in [false, true] { + let store = Arc::new(PausedStore::default()); + store.ignore_conditions.store(true, Ordering::SeqCst); + let files = tempfile::TempDir::new()?; + let listener = TcpListener::bind("127.0.0.1:0").await?; + let address = listener.local_addr()?; + let settings = config(address, files.path().join("server")); + let result = if prebound { + CanopyServer::start_with_listener(settings, store.clone(), listener).await + } else { + drop(listener); + CanopyServer::start(settings, store.clone()).await + }; + assert!(matches!( + result, + Err(canopy_server::server::ServerError::Repository( + "storage conditional create failed" + )) + )); + let remaining = store.list_with_delimiter(None).await?; + assert!(remaining.objects.is_empty() && remaining.common_prefixes.is_empty()); + let rebound = TcpListener::bind(address).await?; + assert_eq!(rebound.local_addr()?, address); + } Ok(()) } diff --git a/crates/canopy-server/tests/multi_server/listener_handoff.rs b/crates/canopy-server/tests/multi_server/listener_handoff.rs new file mode 100644 index 0000000..24fa370 --- /dev/null +++ b/crates/canopy-server/tests/multi_server/listener_handoff.rs @@ -0,0 +1,108 @@ +use super::*; +use canopy_server::server::ServerError; + +type Result = std::result::Result>; + +#[tokio::test(flavor = "multi_thread")] +async fn reserved_listener_survives_startup_and_serves_advertised_clone_url() -> Result { + let files = tempfile::TempDir::new()?; + let listener = TcpListener::bind("127.0.0.1:0").await?; + let address = listener.local_addr()?; + assert_eq!( + TcpListener::bind(address).await.unwrap_err().kind(), + std::io::ErrorKind::AddrInUse + ); + // A port chosen for an advertised clone URL must stay reserved until the + // serving task owns it, rather than being released and rebound at startup. + let server = CanopyServer::start_with_listener( + config(address, files.path().join("node")), + Arc::new(InMemory::new()), + listener, + ) + .await?; + assert_eq!(server.local_addr(), address); + let url = create_repository(address, "listener-handoff").await?; + assert_eq!(url, format!("http://{address}/canopy/listener-handoff.git")); + run_git( + None, + &[ + "-c", + "http.extraHeader=Authorization: Bearer local-test-token", + "ls-remote", + &url, + ], + ) + .await?; + server.shutdown().await?; + let rebound = TcpListener::bind(address).await?; + assert_eq!(rebound.local_addr()?, address); + Ok(()) +} + +#[tokio::test(flavor = "multi_thread")] +async fn mismatched_listener_is_rejected_before_workspace_and_storage_writes() -> Result { + let files = tempfile::TempDir::new()?; + let store = Arc::new(InMemory::new()); + let listener = TcpListener::bind("127.0.0.1:0").await?; + let address = listener.local_addr()?; + let other = TcpListener::bind("127.0.0.1:0").await?; + let configured = other.local_addr()?; + assert_ne!(address, configured); + let data = files.path().join("node"); + let result = CanopyServer::start_with_listener( + config(configured, data.clone()), + store.clone(), + listener, + ) + .await; + assert!(matches!( + result, + Err(ServerError::Http( + "HTTP listener address differs from listen configuration" + )) + )); + assert!(!data.exists(), "mismatch created a workspace"); + let remaining = store.list_with_delimiter(None).await?; + assert!(remaining.objects.is_empty() && remaining.common_prefixes.is_empty()); + let rebound = TcpListener::bind(address).await?; + assert_eq!(rebound.local_addr()?, address); + assert_eq!(other.local_addr()?, configured); + Ok(()) +} + +#[tokio::test(flavor = "multi_thread")] +async fn unpolled_startup_releases_reserved_listener() -> Result { + let files = tempfile::TempDir::new()?; + let store = Arc::new(InMemory::new()); + let listener = TcpListener::bind("127.0.0.1:0").await?; + let address = listener.local_addr()?; + let data = files.path().join("node"); + let startup = + CanopyServer::start_with_listener(config(address, data.clone()), store.clone(), listener); + drop(startup); + assert!(!data.exists()); + let remaining = store.list_with_delimiter(None).await?; + assert!(remaining.objects.is_empty() && remaining.common_prefixes.is_empty()); + let rebound = TcpListener::bind(address).await?; + assert_eq!(rebound.local_addr()?, address); + Ok(()) +} + +#[tokio::test(flavor = "multi_thread")] +async fn released_port_can_be_taken_before_server_startup() -> Result { + let files = tempfile::TempDir::new()?; + let address = available_address().await?; + // Deterministically insert the other binder into the test helper's gap. + let competing = TcpListener::bind(address).await?; + let result = CanopyServer::start( + config(address, files.path().join("node")), + Arc::new(InMemory::new()), + ) + .await; + assert!(matches!( + result, + Err(ServerError::Io(error)) if error.kind() == std::io::ErrorKind::AddrInUse + )); + assert_eq!(competing.local_addr()?, address); + Ok(()) +} diff --git a/crates/canopy-server/tests/multi_server/main.rs b/crates/canopy-server/tests/multi_server/main.rs index bc6fe92..ba39708 100644 --- a/crates/canopy-server/tests/multi_server/main.rs +++ b/crates/canopy-server/tests/multi_server/main.rs @@ -36,6 +36,7 @@ mod issues; mod large_objects; mod lfs_locks; mod lifecycle; +mod listener_handoff; mod peers; mod residency; mod retained_catalog; diff --git a/crates/canopy-server/tests/multi_server/sha256.rs b/crates/canopy-server/tests/multi_server/sha256.rs index dcce1ab..cf8442e 100644 --- a/crates/canopy-server/tests/multi_server/sha256.rs +++ b/crates/canopy-server/tests/multi_server/sha256.rs @@ -15,10 +15,12 @@ async fn sha256_real_provider_round_trip() -> Result { async fn sha256_repository_round_trip(store: Arc) -> Result { let workspace = tempfile::TempDir::new()?; - let first_address = available_address().await?; - let first = CanopyServer::start( + let first_listener = TcpListener::bind("127.0.0.1:0").await?; + let first_address = first_listener.local_addr()?; + let first = CanopyServer::start_with_listener( config(first_address, workspace.path().join("first")), Arc::clone(&store), + first_listener, ) .await?; let response: serde_json::Value = reqwest::Client::new() @@ -134,10 +136,12 @@ async fn sha256_repository_round_trip(store: Arc) -> Result { assert_eq!(expected.trim_ascii().len(), 64); first.shutdown().await?; - let second_address = available_address().await?; - let second = CanopyServer::start( + let second_listener = TcpListener::bind("127.0.0.1:0").await?; + let second_address = second_listener.local_addr()?; + let second = CanopyServer::start_with_listener( config(second_address, workspace.path().join("second")), Arc::clone(&store), + second_listener, ) .await?; let second_url = format!("http://{second_address}/canopy/sha256.git"); @@ -263,10 +267,12 @@ async fn sha256_checks_reviews_and_merge_survive_restore() -> Result { let store: Arc = Arc::new(InMemory::new()); let workspace = tempfile::TempDir::new()?; - let address = available_address().await?; - let server = CanopyServer::start( + let listener = TcpListener::bind("127.0.0.1:0").await?; + let address = listener.local_addr()?; + let server = CanopyServer::start_with_listener( config(address, workspace.path().join("first")), Arc::clone(&store), + listener, ) .await?; let client = Client::new(); @@ -421,10 +427,12 @@ async fn sha256_checks_reviews_and_merge_survive_restore() -> Result { assert_eq!(merged["merge"]["oid"], source); server.shutdown().await?; - let restored_address = available_address().await?; - let restored = CanopyServer::start( + let restored_listener = TcpListener::bind("127.0.0.1:0").await?; + let restored_address = restored_listener.local_addr()?; + let restored = CanopyServer::start_with_listener( config(restored_address, workspace.path().join("restored")), store, + restored_listener, ) .await?; let restored_repo = format!("http://{restored_address}/api/repositories/sha256-review"); @@ -479,10 +487,12 @@ async fn sha256_native_merge_candidates(store: Arc) -> Result { use serde_json::json; let workspace = tempfile::TempDir::new()?; - let address = available_address().await?; - let server = CanopyServer::start( + let listener = TcpListener::bind("127.0.0.1:0").await?; + let address = listener.local_addr()?; + let server = CanopyServer::start_with_listener( config(address, workspace.path().join("first")), Arc::clone(&store), + listener, ) .await?; let client = Client::new(); @@ -596,10 +606,12 @@ async fn sha256_native_merge_candidates(store: Arc) -> Result { } server.shutdown().await?; - let restored_address = available_address().await?; - let restored = CanopyServer::start( + let restored_listener = TcpListener::bind("127.0.0.1:0").await?; + let restored_address = restored_listener.local_addr()?; + let restored = CanopyServer::start_with_listener( config(restored_address, workspace.path().join("restored")), store, + restored_listener, ) .await?; for (name, repository, request, candidate) in prepared { diff --git a/docs/contracts.md b/docs/contracts.md index 2de4324..747f2d3 100644 --- a/docs/contracts.md +++ b/docs/contracts.md @@ -2803,6 +2803,11 @@ records release correctness and fresh RustFS compatibility separately from performance. The [owned retained-store qualification](performance/2026-10-01-old-binary-rustfs-upgrade.md) records actual old-executable restore and fresh-owner recovery; it does not provide a general upgrade controller or qualify the existing 10K corpus. +The [listener handoff checkpoint](performance/2026-10-01-listener-handoff.md) +qualifies the optional prebound HTTP listener through the same startup +supervisor. Address mismatch is rejected before workspace/storage writes; +readiness, fencing, cancellation cleanup and drain checks are unchanged. +Its executable has separate bindings; earlier measurements do not certify it. The application declares one entity-partitioned SQL Cell type for repositories. `repository_target` validates the canonical UUID, then calls the app crate's diff --git a/docs/delivery-plan.md b/docs/delivery-plan.md index 005c6d1..c23df8b 100644 --- a/docs/delivery-plan.md +++ b/docs/delivery-plan.md @@ -88,8 +88,12 @@ These are the next implementation slices. Each slice must update its contract, b reruns these checks on the new immutable revision without changing runtime budgets. The [owned old-binary RustFS qualification](performance/2026-10-01-old-binary-rustfs-upgrade.md) passed controlled upgrade, twenty scheduled critical workflows and complete - recovery of that fixture after three-owner loss. Original-corpus upgrade, - full reference performance and the CI port-handoff failure remain open. + recovery of that fixture after three-owner loss. The + [listener handoff follow-up](performance/2026-10-01-listener-handoff.md) + removes the affected test's reservation gap and passes 244 release tests, + lints, 84 Python tests and eight fresh RustFS gates. Its changed head needs + fresh Linux CI. Original-corpus upgrade and full reference performance + remain open. 2. Execute the [repository-density performance plan](performance-plan.md), then qualify the bounded Linux deployment and residency under faults and larger hot sets. Each node reserves one SQL slot for Directory takeover and admits diff --git a/docs/performance-plan.md b/docs/performance-plan.md index ecec26b..af2ff0c 100644 --- a/docs/performance-plan.md +++ b/docs/performance-plan.md @@ -19,6 +19,11 @@ of these functional checks establishes reference capacity. Upstream subsequently advanced to `0dc04a6` with atomic API and test changes. The new candidate is independently qualified; the original corpus experiment and its frozen source remain on `c51dd121`. +The subsequent [listener handoff correction](performance/2026-10-01-listener-handoff.md) +passes 244 release Rust tests, lints, 84 Python tests and eight fresh RustFS gates +on its own frozen source. It changes no runtime budget and supplies no new +throughput or latency measurement; the earlier measured executable remains +separately bound. The [combined authentication and residency candidate](performance/2026-10-01-bounded-authentication-and-residency.md) is now in PR #18: 231 release tests and all eight RustFS gates passed, as did 20 unchanged cold-activation repetitions and all 15 residency tests. It has not @@ -38,6 +43,7 @@ seed and successful recovery of every recorded ACK. | Current gate | Evidence / status | | --- | --- | | Current Cellule revision | `0dc04a6`: independently rebuilt release correctness, lints, Python harness and eight fresh RustFS gates passed. See the latest checkpoint for repeated regressions and exact bindings. Not retained-store upgrade, live deployment or measured performance | +| HTTP listener handoff | [Separate startup qualification](performance/2026-10-01-listener-handoff.md): the affected SHA-256 tests retain their bound listeners through supervised startup; 244 release tests, 84 Python tests, eight fresh RustFS gates and repeated cancellation/handoff checks passed. New PR head requires fresh Linux CI; not measured performance | | Owned old-binary RustFS upgrade | [Controlled fixture qualification](performance/2026-10-01-old-binary-rustfs-upgrade.md) passed actual old-byte restore through three fresh gateways/proxy, both new Git/LFS writes, twenty scheduled 17-step workflows and recovery of all forty critical repositories after all three owners were killed. Two retained repositories; not the original 10K corpus or reference capacity | | Original rebuilt Cellule artifact | `c51dd121` build, 227 Rust tests, eight RustFS gates and initial 17-step proxy check passed; full 10K/100 seed closed, but remote verification failed HTTP 503. A separate 10K metadata diagnostic returned 9,999 matching identities and one 503; downstream fault/load gates stopped without owner signals. Cause unresolved | | Combined authentication and residency candidate | First publication passed 231 release tests, eight RustFS gates, 20 cold-activation repetitions and 15 residency tests. Exact predecessor descriptor is retained; see the separately qualified catalog follow-up | diff --git a/docs/performance/2026-10-01-cellule-main-qualification.md b/docs/performance/2026-10-01-cellule-main-qualification.md index 624aa52..363b541 100644 --- a/docs/performance/2026-10-01-cellule-main-qualification.md +++ b/docs/performance/2026-10-01-cellule-main-qualification.md @@ -16,8 +16,10 @@ The tested source is `bbd784a40c3867646044a8c716b7ee517f9aca49`. The dependency change replaces five direct manifest pins and six lockfile source revisions; it changes no other dependency versions, Canopy implementation or runtime budgets. Locked metadata resolves all six Cellule packages to the same immutable commit. -Before publication, all 263 production, test, script and Cargo files on the PR -branch are checked against this frozen source. +At first publication (`240203a7`), all 263 production, test, script and Cargo +files on the PR branch matched this frozen source. The subsequent +[listener handoff verification](2026-10-01-listener-handoff.md) has its own +source, executable and checks; the artifacts below remain historical bindings. The new revision includes atomic API and upstream test changes. Passing functional tests is not evidence of a throughput or latency improvement. diff --git a/docs/performance/2026-10-01-listener-handoff.md b/docs/performance/2026-10-01-listener-handoff.md new file mode 100644 index 0000000..ba19e18 --- /dev/null +++ b/docs/performance/2026-10-01-listener-handoff.md @@ -0,0 +1,123 @@ +# HTTP listener handoff verification + +The SHA-256 verification tests now transfer an already-bound HTTP listener into +Canopy startup. This removes their port reservation gap without retries, +serialized tests, relaxed Git assertions or changed runtime budgets. Release +correctness and all eight fresh-RustFS gates passed on the separately frozen +source. This is not a throughput improvement or an upgrade of the +original 10,000-repository corpus. + +The [earlier three-node measurements](2026-10-01-old-binary-rustfs-upgrade.md) +remain bound to their original executable. They are not performance results for +this new build. + +## Reproduction and correction + +At PR head `240203a7`, one Linux Rust job failed with `AddrInUse` in +`sha256_native_merge_candidates_survive_restore_and_publish`; the parallel job +passed. Both later Rust jobs at `0498e22` passed before this correction. +Those passes do not erase the original failure. + +The affected helper returned the address of a temporary listener and dropped +the listener before Canopy bound that address. A competing binder inserted in +that gap reproduces the same error. A regression that retains the reservation +failed in all three pre-fix runs. The exact competing binder in the original +Linux job was not captured. + +```text +Before: select port -> drop listener -> unreserved gap -> bind again + another test may take the port + +After: bind listener -> transfer ownership -> supervised startup -> serve + listener remains bound throughout +``` + +`CanopyServer::start_with_listener` uses the existing startup and shutdown +supervisor. It requires the listener's actual address to equal `config.listen` +before creating a workspace or writing storage. The public URL can still name +a proxy. Ordinary `CanopyServer::start` retains its existing bind path, readiness +checks and drain behavior; SSH configuration is unchanged. + +An embedding caller can reserve its HTTP address as follows: + +```rust +let listener = tokio::net::TcpListener::bind("127.0.0.1:0").await?; +config.listen = listener.local_addr()?; +config.public_url = format!("http://{}", config.listen); +let server = CanopyServer::start_with_listener(config, store, listener).await?; +// Use the advertised clone URL only after startup returns readiness. +server.shutdown().await?; +``` + +All six startup addresses in the SHA-256 integration tests use owned listeners, +including both provider tests. Other tests still use the legacy address helper; +this correction does not claim to eliminate every test-suite port race. + +## Closed checks + +The tested source is `3dda2b47b1cba105642a62b4ad27d7c84d0940d5`, with the same +Cellule revision `0dc04a658bd99668936f7ec58032d054f6fbc141` and unchanged lockfile. + +| Check | Result and boundary | +| --- | --- | +| Pre-fix reservation regression | Three failures with `AddrInUse`; failed executable and exact source retained | +| Handoff regressions | All four passed: real advertised clone URL, competing binder refusal, mismatch before writes and unpolled cancellation | +| Admitted startup cancellation | Workspace and listener remain held until cleanup settles; subsequent rebind and repository creation passed | +| Repetitions | All four handoff tests and admitted cancellation passed in each of 20 independent runs; 100 checks, not throughput samples | +| Lifecycle suite | All 14 passed, including both bind paths rejecting ignored conditional writes before enrollment | +| SHA-256 suite | All three ordinary tests passed; includes all three native merge strategies, exact candidate restoration, stock Git clones and strict fsck | +| Release workspace | 244 top-level tests passed, zero failed, nine ignored; nested subprocess tests counted once | +| Release build and lints | Locked release executable built; all-target Clippy passed with warnings denied | +| Python harness | All 84 tests passed on the same frozen source | +| Fresh RustFS compatibility | All eight exact gates passed; disposable fresh fixture, not a retained-store or five-GiB gate | + +The full workspace preserved the SSH-LFS grant's real 300-second expiry check. +Release verification closed October 2 at 01:15:37 UTC, October 1 in the local +Pacific timezone. No runtime diagnostic logging was added. +Provider verification closed at 01:18:09 UTC. The original performance +provider's container, volume, image, configuration, start time and restart count +were identical before and after these gates. The disposable gate fixture was +removed by its normal cleanup. + +Reproduce the local checks: + +```sh +cargo test --release --locked --test multi_server listener_handoff:: -- --test-threads=4 +cargo test --release --locked --test multi_server lifecycle:: -- --test-threads=4 +cargo test --release --locked --test multi_server sha256:: -- --test-threads=4 +cargo test --release --locked --workspace -- --test-threads=4 +cargo clippy --release --locked --workspace --all-targets -- -D warnings +python3 -B -m unittest discover -s scripts -p 'test_*.py' +python3 -B scripts/qualify_size.py --provider-only --release +``` + +## Artifact bindings and remaining work + +Evidence is retained at +`/Users/haipingfu/.codex/canopy-listener-handoff-I4OJLd`. The failed source, +executable and original CI log were copied and independently reread at +`/Volumes/Workspace/CrabData/canopy-listener-red-8_2n504v`; the 330-file release +copy is at `/Volumes/Workspace/CrabData/canopy-listener-green-hreldv3h`. +All 695 closed files, including provider, Python and repetition receipts, were +copied and independently reread at +`/Volumes/Workspace/CrabData/canopy-listener-closed-omgsp1wu` before publication. +These are different local filesystems, not off-machine backups. + +| Artifact | SHA-256 | +| --- | --- | +| Pre-fix receipt | `fe23bd6d351a847f29b43ff8ed00c6902b86374b976c8cda52e72c31482f0f76` | +| New release executable | `a61ef0f2cb977348e4e4fc45330334a1f8fd68bf8c44146cd9343b0e6676a6c8` | +| Release verification receipt | `85845cf11b2977136904b6095d5be55ec628818ab220e5386e69efc1de26470b` | +| Release workspace log | `6b78ae474033ba44a7aa9a8662662364ca9ca1733bd9af2b3de7dd01959299dc` | +| Repetition receipt | `72b82bf3aaf667c7206c013fe8252f27941c0b1f5c0c60c66c867eee56e2dfb7` | +| Python receipt | `2f519343688f3e9055ad48bc0060d7e2088551935ef706e22b287617a397a23e` | +| RustFS receipt | `b5a1baa082d8178a2ceedcb589edb0c2aa542d0cea95ea550989549ab23216b2` | +| RustFS log | `3ed9009bb25bd7fb348eec0bec10c729cae3722976479cac24329f5ae0c2ba3b` | + +The changed PR head needs fresh Linux CI. PR #18 remains draft. The +[full campaign](../performance-plan.md) still requires original-corpus offline +recovery and upgrade admission, full remote Git/LFS verification, all 108 load +windows, every acknowledged write after owner loss, higher admission profiles, +matched comparisons, non-sparse five-GiB transfers and isolated Linux capacity. +The original provider's last read-only status had zero advertised writers and +301 unsettled Cells; neither its release nor its data was changed here. diff --git a/docs/performance/2026-10-01-old-binary-rustfs-upgrade.md b/docs/performance/2026-10-01-old-binary-rustfs-upgrade.md index ccb1466..8360a7b 100644 --- a/docs/performance/2026-10-01-old-binary-rustfs-upgrade.md +++ b/docs/performance/2026-10-01-old-binary-rustfs-upgrade.md @@ -129,7 +129,9 @@ The old executable SHA-256 is `32b114119960608c0a91d1c783bb69eafec432831bfa452d54d8950b09bc0e99`; the new qualified executable is `a541cef36acae0b21a651316fda2e9b3204ef7ea3f02173b65a7c217cd6026e9`. -Production source still matches `bbd784a40c3867646044a8c716b7ee517f9aca49`. +The measured production source matches `bbd784a40c3867646044a8c716b7ee517f9aca49`. +The subsequent [listener handoff fix](2026-10-01-listener-handoff.md) is a +separately verified build, not a new performance binding for this run. The producer, upgrade, load and recovery ran October 2 at 00:13–00:36 UTC, October 1 in the local Pacific timezone. @@ -143,8 +145,10 @@ the small fixture does not explain the original fleet's terminal lease failure. At PR head `240203a7`, both Python CI jobs and one complete Rust job passed. The other Rust job failed with `AddrInUse` in the SHA-256 merge-candidate test. -The failure log is retained; a passing parallel job does not erase it. Port -reservation/handoff needs reproduction and correction before merge readiness. +The failure log is retained; a passing parallel job does not erase it. The +[listener handoff follow-up](2026-10-01-listener-handoff.md) reproduces the +reservation gap, removes it from the affected SHA-256 tests and passes release +and fresh-RustFS checks. The changed PR head still needs its own Linux CI. The [full performance plan](../performance-plan.md) remains unchanged: 10,000 identities, 100 populated Git/LFS fixtures, all 108 windows, 114,960 arrivals From 4ddc74241d878908363cb351bac133358b998923 Mon Sep 17 00:00:00 2001 From: forhappy Date: Thu, 1 Oct 2026 19:03:25 -0700 Subject: [PATCH 47/60] docs: record full original-corpus maintenance recovery --- docs/delivery-plan.md | 9 +- docs/performance-plan.md | 14 ++- .../2026-10-01-listener-handoff.md | 16 ++- .../2026-10-01-old-binary-rustfs-upgrade.md | 16 ++- .../2026-10-01-original-corpus-recovery.md | 117 ++++++++++++++++++ 5 files changed, 154 insertions(+), 18 deletions(-) create mode 100644 docs/performance/2026-10-01-original-corpus-recovery.md diff --git a/docs/delivery-plan.md b/docs/delivery-plan.md index c23df8b..bac7969 100644 --- a/docs/delivery-plan.md +++ b/docs/delivery-plan.md @@ -91,9 +91,12 @@ These are the next implementation slices. Each slice must update its contract, b recovery of that fixture after three-owner loss. The [listener handoff follow-up](performance/2026-10-01-listener-handoff.md) removes the affected test's reservation gap and passes 244 release tests, - lints, 84 Python tests and eight fresh RustFS gates. Its changed head needs - fresh Linux CI. Original-corpus upgrade and full reference performance - remain open. + lints, 84 Python tests and eight fresh RustFS gates. Both Linux CI workflows + at `eaebfc4` passed. The + [original corpus maintenance recovery](performance/2026-10-01-original-corpus-recovery.md) + settled all 301 remaining Cells and verified 10,003 idle Controls with + unchanged catalog and published roots. Old Maintenance still closes serving + admission; new-release verification and full reference performance remain open. 2. Execute the [repository-density performance plan](performance-plan.md), then qualify the bounded Linux deployment and residency under faults and larger hot sets. Each node reserves one SQL slot for Directory takeover and admits diff --git a/docs/performance-plan.md b/docs/performance-plan.md index af2ff0c..75f0188 100644 --- a/docs/performance-plan.md +++ b/docs/performance-plan.md @@ -28,13 +28,18 @@ The [combined authentication and residency candidate](performance/2026-10-01-bou is now in PR #18: 231 release tests and all eight RustFS gates passed, as did 20 unchanged cold-activation repetitions and all 15 residency tests. It has not been deployed on the existing corpus. Descriptor compatibility passed, but -actual old-binary restore and the diagnostic fleet's terminal lease-fencing -failure remain open. The [catalog admission follow-up](performance/2026-10-01-retained-catalog-admission.md) +the original corpus's new-release restore and the diagnostic fleet's terminal +lease-fencing failure remain open. The [catalog admission follow-up](performance/2026-10-01-retained-catalog-admission.md) reproduces and fixes immutable-identity reprovisioning and late rejection of unsupported Control metadata. Its combined candidate passed 239 release tests, all eight RustFS gates, 20 cold-activation repetitions and 20 three-case retained startup repetitions, plus all 15 residency tests. Retained startup uses an owned -in-memory fixture; actual old-binary RustFS upgrade remains open. +in-memory fixture; the separate owned old-binary RustFS upgrade passed, while +original-corpus new-release verification remains open. The +[original corpus maintenance recovery](performance/2026-10-01-original-corpus-recovery.md) +now leaves all 10,003 Controls durably idle under the old Maintenance release. +Its catalog, roots and 9,702 previously idle Controls stayed unchanged; no new +gateway, remote payload verification or performance result is established. Neither checkpoint replaces the original complete performance schedule. The [three-node proxy campaign](performance/2026-09-30-three-node-proxy.md) retains the earlier fixture race and its regression, plus the candidate's failed @@ -43,7 +48,8 @@ seed and successful recovery of every recorded ACK. | Current gate | Evidence / status | | --- | --- | | Current Cellule revision | `0dc04a6`: independently rebuilt release correctness, lints, Python harness and eight fresh RustFS gates passed. See the latest checkpoint for repeated regressions and exact bindings. Not retained-store upgrade, live deployment or measured performance | -| HTTP listener handoff | [Separate startup qualification](performance/2026-10-01-listener-handoff.md): the affected SHA-256 tests retain their bound listeners through supervised startup; 244 release tests, 84 Python tests, eight fresh RustFS gates and repeated cancellation/handoff checks passed. New PR head requires fresh Linux CI; not measured performance | +| HTTP listener handoff | [Separate startup qualification](performance/2026-10-01-listener-handoff.md): the affected SHA-256 tests retain their bound listeners through supervised startup; 244 release tests, 84 Python tests, eight fresh RustFS gates and repeated cancellation/handoff checks passed. Both Linux CI workflows at `eaebfc4` passed; not measured performance | +| Original corpus maintenance recovery | [Complete metadata and fenced-recovery checkpoint](performance/2026-10-01-original-corpus-recovery.md): all 256 shards and 10,003 Controls checked twice before and after; 301 unsettled Cells drained, zero unsettled Cells, catalog and all published roots unchanged. Old Maintenance remains closed to serving; new-release admission and remote Git/LFS verification remain open | | Owned old-binary RustFS upgrade | [Controlled fixture qualification](performance/2026-10-01-old-binary-rustfs-upgrade.md) passed actual old-byte restore through three fresh gateways/proxy, both new Git/LFS writes, twenty scheduled 17-step workflows and recovery of all forty critical repositories after all three owners were killed. Two retained repositories; not the original 10K corpus or reference capacity | | Original rebuilt Cellule artifact | `c51dd121` build, 227 Rust tests, eight RustFS gates and initial 17-step proxy check passed; full 10K/100 seed closed, but remote verification failed HTTP 503. A separate 10K metadata diagnostic returned 9,999 matching identities and one 503; downstream fault/load gates stopped without owner signals. Cause unresolved | | Combined authentication and residency candidate | First publication passed 231 release tests, eight RustFS gates, 20 cold-activation repetitions and 15 residency tests. Exact predecessor descriptor is retained; see the separately qualified catalog follow-up | diff --git a/docs/performance/2026-10-01-listener-handoff.md b/docs/performance/2026-10-01-listener-handoff.md index ba19e18..7503773 100644 --- a/docs/performance/2026-10-01-listener-handoff.md +++ b/docs/performance/2026-10-01-listener-handoff.md @@ -114,10 +114,16 @@ These are different local filesystems, not off-machine backups. | RustFS receipt | `b5a1baa082d8178a2ceedcb589edb0c2aa542d0cea95ea550989549ab23216b2` | | RustFS log | `3ed9009bb25bd7fb348eec0bec10c729cae3722976479cac24329f5ae0c2ba3b` | -The changed PR head needs fresh Linux CI. PR #18 remains draft. The -[full campaign](../performance-plan.md) still requires original-corpus offline -recovery and upgrade admission, full remote Git/LFS verification, all 108 load +Both [Linux PR CI](https://github.com/crabbuild/canopy/actions/runs/36950769630) +and [push CI](https://github.com/crabbuild/canopy/actions/runs/36950765768) at +`eaebfc4cdd7ead54a8d62978aebcd8656f0de0b8` passed Rust, real RustFS compatibility, +server build and Python harness checks. PR #18 remains draft. The +[full campaign](../performance-plan.md) still requires original-corpus new-release +admission, full remote Git/LFS verification, all 108 load windows, every acknowledged write after owner loss, higher admission profiles, matched comparisons, non-sparse five-GiB transfers and isolated Linux capacity. -The original provider's last read-only status had zero advertised writers and -301 unsettled Cells; neither its release nor its data was changed here. +The original provider's read-only status at this qualification had zero advertised +writers and 301 unsettled Cells; neither its release nor its data was changed here. +The later [maintenance recovery](2026-10-01-original-corpus-recovery.md) settled +those Cells without changing catalog identities or published roots. It remains +in old Maintenance; that later action is not a new listener-build performance result. diff --git a/docs/performance/2026-10-01-old-binary-rustfs-upgrade.md b/docs/performance/2026-10-01-old-binary-rustfs-upgrade.md index 8360a7b..0bd81e6 100644 --- a/docs/performance/2026-10-01-old-binary-rustfs-upgrade.md +++ b/docs/performance/2026-10-01-old-binary-rustfs-upgrade.md @@ -137,18 +137,22 @@ October 1 in the local Pacific timezone. ## Remaining full campaign and CI work -The original 10,000-repository provider was not upgraded or restarted. Its -read-only status observation reports the old Ready release, zero advertised -writers and **301 unsettled Cells**. It requires supported offline recovery and -full same-corpus admission before any new release or gateway launch. Restarting -the small fixture does not explain the original fleet's terminal lease failure. +The original 10,000-repository provider was not upgraded or restarted during +this small-fixture qualification. The later +[original corpus maintenance recovery](2026-10-01-original-corpus-recovery.md) +settled its 301 remaining Cells using the exact old executable. All 10,003 +Controls are now idle and unowned; catalog identities and published roots stayed +unchanged. It remains in old Maintenance pending full new-release admission and +remote Git/LFS verification. Recovery does not explain the earlier fleet's +terminal lease failure. At PR head `240203a7`, both Python CI jobs and one complete Rust job passed. The other Rust job failed with `AddrInUse` in the SHA-256 merge-candidate test. The failure log is retained; a passing parallel job does not erase it. The [listener handoff follow-up](2026-10-01-listener-handoff.md) reproduces the reservation gap, removes it from the affected SHA-256 tests and passes release -and fresh-RustFS checks. The changed PR head still needs its own Linux CI. +and fresh-RustFS checks. Both Linux PR and push workflows at `eaebfc4` then passed +Rust, real RustFS compatibility, server build and Python harness checks. The [full performance plan](../performance-plan.md) remains unchanged: 10,000 identities, 100 populated Git/LFS fixtures, all 108 windows, 114,960 arrivals diff --git a/docs/performance/2026-10-01-original-corpus-recovery.md b/docs/performance/2026-10-01-original-corpus-recovery.md new file mode 100644 index 0000000..5f2a0b5 --- /dev/null +++ b/docs/performance/2026-10-01-original-corpus-recovery.md @@ -0,0 +1,117 @@ +# Original corpus maintenance recovery + +The original RustFS corpus now has **10,003 durably idle, unowned Cells**: +10,000 repository identities, two critical Git repositories and Directory. +Normal fenced recovery with the exact old executable settled all 301 remaining +Cells. Two complete post-recovery snapshots and an independent receipt audit +passed. All published roots, catalog identities and Cell incarnations stayed +unchanged. + +The deployment remains in **old-release Maintenance**, revision 5. No new +release or serving gateway was admitted. This closes the offline metadata and +maintenance-recovery prerequisite, not remote Git/LFS content verification, +same-corpus upgrade, performance or an explanation of the earlier lease failure. + +## Recovery sequence + +```text +Old Ready release, revision 3 + -> read-only scan of all 256 shards and 10,003 Controls + -> exact old descriptor supports every catalog and Control code/schema + -> three recorded owners canonically retired; zero advertised writers + -> identical second snapshot; preserve closed evidence + -> old executable begins exact maintenance operation, revision 5 + -> new recovery identity and scratch disk; normal fenced acquire and drain + -> old executable reports zero unsettled Cells; worker exits zero + -> two full read-only scans; independent canonical JSON and receipt audit + -> remain in old Maintenance; upgrade and serving admission still closed +``` + +The readers use a transport wrapper that refuses every mutation before it +reaches RustFS. Its test rejects put, create, multipart, copy and both delete +paths while preserving the original bytes. Unsupported command mode and wrong +endpoint are rejected before provider access. Both tools passed locked offline +release builds, unit tests and all-target Clippy with warnings denied; dependency +versions match the qualified Cellule `0dc04a6` metadata. + +Recovery used the retained **old** executable, its exact selected descriptor and +image setting, a new NodeID, an owned test signing key and a fresh workspace. +It did not erase ownership, force a Control CAS, widen leases, rewrite catalog +identity, change runtime budgets, restart RustFS or reseed the corpus. The worker +drained acquired Cells through the existing runtime before removing only its +closed restore scratch. + +## Verification results + +| Check | Closed result | +| --- | --- | +| Before recovery | 9,702 idle and 301 serving Controls; all 10,003 expected Cells present | +| Old release support | Exact persisted descriptor, initial catalog code/schema and actual Control code/schema verified across all 256 shards | +| Previous owners | All three canonically retired and not live; zero advertisements, including expired advertisements | +| Stable pre-recovery observation | Two identical complete snapshots; zero attempted provider writes | +| Old maintenance recovery | Exit zero; zero advertisements and zero unsettled Cells; recovery command took 234.461 seconds | +| Post-recovery observation | Two identical complete snapshots; all 10,003 Controls idle, unowned and rooted | +| Previously idle Cells | All 9,702 canonical Controls and their ETags unchanged | +| Recovered Cells | All 301 retain incarnation and code/schema; fencing epoch and revision advance | +| Durable metadata | All published roots, catalog shard revisions/page digests and service-root bytes/ETag unchanged | +| Independent audit | Receipt/source bindings verified; exact counts and root transaction/commit sequence checks passed; all owned workers and children absent | +| RustFS | Same container, volume, image, configuration, start time, resource envelope and restart count; no OOM | + +The durations are maintenance and complete-scan wall times on the shared +Mac/Colima host. They are not Git request latency, throughput, a matched +improvement or isolated Linux capacity. The corpus manifests still describe +100 populated Git/LFS fixtures and two critical repositories; this checkpoint +does not claim to have cloned or checked their payloads remotely. + +## Commands and admission boundaries + +On a separately verified, caller-owned deployment, the existing old-binary CLI +provides the maintenance operations: + +```sh +canopy maintenance /path/to/exact-old-config.json status +canopy maintenance /path/to/exact-old-config.json begin +canopy maintenance /path/to/exact-old-worker-config.json recover +canopy maintenance /path/to/exact-old-config.json status +``` + +These commands do not replace full catalog admission. Verify the exact old +executable, descriptor and supported persisted metadata first; use a fresh +worker identity and workspace. Do not end maintenance or activate another +release until its separate admission requirements are proven. The retained +readers and wrapper are hard-scoped qualification tools, not general upgrade +commands. + +## Evidence and remaining work + +Closed evidence is retained at +`/Users/haipingfu/.codex/canopy-original-corpus-admission-YxGodu`. +All 73 closed files and bound inputs were copied and independently reread at +`/Volumes/Workspace/CrabData/canopy-original-recovery-closed-181o8fa6`. +This is another local filesystem, not an off-machine or provider-data backup. +The first offline-resolution failure, map-entry lint failure and output-name +collision are preserved separately; existing receipts were never overwritten. + +| Artifact | SHA-256 | +| --- | --- | +| Exact old executable | `32b114119960608c0a91d1c783bb69eafec432831bfa452d54d8950b09bc0e99` | +| Pre-recovery snapshot, both copies | `3c53357ec9568ba878a7d21e347038fd09a366559b7bd06f24fc98d4ef89fc49` | +| Same-code recovery receipt | `28e6e6899b946b62bad8c4312daca884a4377d93c159aabb588f8feb8d027e38` | +| Post-recovery snapshot, both copies | `de62eebeaa609937759186126ff6540e823e3ddbb209df8becb6128a21fc96ce` | +| Post-recovery admission | `2b4b691d47d8910106593046dc6efbb3553e8496ee2c9cedfe516601db9a77bd` | +| Independent closed audit | `219331a18c0dcb0e9be65885f22f73bea610befefb66ec3d70e4b018824fd95a` | + +The checks ran October 2 at 01:46–01:57 UTC, October 1 Pacific. +The selected old release remains +`e31bf1a951e2fa19d91e9f964b2ddeade1a81b05a20ad628362819a1487c16b1`; +maintenance operation is `74f10df1-417c-4c63-83b0-b554fdb996e4`. +The UI preview and unrelated providers were not changed or signalled. + +Next gates are rolling-descriptor and full actual-Control admission of the new +executable on this same corpus, controlled activation, three fresh gateways and +a proxy, full remote Git/LFS verification and fresh-owner/every-ACK recovery. +The [full campaign](../performance-plan.md) remains **108 windows, 114,960 +arrivals and 8,640 scheduled seconds**, with independent rates/concurrency, +uniform/skewed access, higher residency profiles, concurrent critical workflows +and faults, matched comparisons, large transfers and isolated Linux capacity. +No new serving or performance result is claimed here. PR #18 remains draft. From 2c897c7eae1afa36049685dd35f2214082e03e23 Mon Sep 17 00:00:00 2001 From: forhappy Date: Thu, 1 Oct 2026 19:45:18 -0700 Subject: [PATCH 48/60] docs: record controlled original-corpus release activation --- docs/README.md | 1 + docs/performance-plan.md | 24 ++-- .../2026-10-01-cellule-main-qualification.md | 11 +- .../2026-10-01-listener-handoff.md | 10 +- .../2026-10-01-original-corpus-activation.md | 116 ++++++++++++++++++ .../2026-10-01-original-corpus-recovery.md | 17 ++- 6 files changed, 159 insertions(+), 20 deletions(-) create mode 100644 docs/performance/2026-10-01-original-corpus-activation.md diff --git a/docs/README.md b/docs/README.md index c8c2444..a51f268 100644 --- a/docs/README.md +++ b/docs/README.md @@ -14,6 +14,7 @@ Use this page to choose a document by task. Canopy's hosting core supports stock | Find the right Rust crate | [Rust workspace](workspace.md) | Crate ownership, dependencies and build commands | | Decide whether a release gate is closed | [Delivery plan](delivery-plan.md) | Required proof, current state and chronological implementation evidence | | Plan or evaluate capacity | [Repository density and latency](performance-plan.md) | Workloads, targets, measured results and limits of each result | +| Inspect the original RustFS corpus upgrade | [Full-corpus release activation](performance/2026-10-01-original-corpus-activation.md) | Admission of 10,003 Cells, exact release transitions, unchanged metadata and open remote-content gates | | Track three nodes behind a proxy and the latest dependency candidate | [Three-node proxy qualification](performance/2026-09-30-three-node-proxy.md) | Baseline/candidate pins, complete-corpus recovery, failing load windows and open gates | | Inspect the merged workspace and RustFS verification | [Workspace and RustFS verification](performance/2026-09-30-workspace-rustfs.md) | The `70bd25f` revision, conflict resolution and end-to-end gates | | Inspect the earlier Cellule dependency update | [Initial September 30 requalification](performance/2026-09-30-cellule-main.md) | Earlier revision, passing checks, failed suites and storage-blocked performance gates | diff --git a/docs/performance-plan.md b/docs/performance-plan.md index 75f0188..03c4085 100644 --- a/docs/performance-plan.md +++ b/docs/performance-plan.md @@ -5,8 +5,10 @@ It separates demonstrated behavior from proposed targets. The current Cellule dependency pin is `0dc04a658bd99668936f7ec58032d054f6fbc141`, verified against upstream main on October 1. Its [qualification checkpoint](performance/2026-10-01-cellule-main-qualification.md) records 239 release Rust tests, 84 Python tests, lints and all eight fresh RustFS -gates, with separately retained artifacts. Original-corpus upgrade and reference -performance remain open. The preceding `c51dd121` build is historical evidence. +gates, with separately retained artifacts. Original-corpus remote-content +verification and reference performance remain open. Upstream later advanced to +`191409685b001a82bd02780def45102b4fc2f164`; that revision is not yet qualified +by this PR. The preceding `c51dd121` build is historical evidence. Unlike the earlier `a4500add` documentation advance, this changes routing and resource-permit implementation. Locked metadata resolves all six packages; the lockfile changes only their sources. @@ -18,7 +20,8 @@ recovery, scheduled load and every-ACK verification remain separate gates; none of these functional checks establishes reference capacity. Upstream subsequently advanced to `0dc04a6` with atomic API and test changes. The new candidate is independently qualified; the original corpus experiment -and its frozen source remain on `c51dd121`. +and its original frozen source used `c51dd121`; the later activation below uses +the separately qualified `0dc04a6` executable. The subsequent [listener handoff correction](performance/2026-10-01-listener-handoff.md) passes 244 release Rust tests, lints, 84 Python tests and eight fresh RustFS gates on its own frozen source. It changes no runtime budget and supplies no new @@ -27,7 +30,7 @@ separately bound. The [combined authentication and residency candidate](performance/2026-10-01-bounded-authentication-and-residency.md) is now in PR #18: 231 release tests and all eight RustFS gates passed, as did 20 unchanged cold-activation repetitions and all 15 residency tests. It has not -been deployed on the existing corpus. Descriptor compatibility passed, but +been deployed at that checkpoint. Descriptor compatibility passed, but the original corpus's new-release restore and the diagnostic fleet's terminal lease-fencing failure remain open. The [catalog admission follow-up](performance/2026-10-01-retained-catalog-admission.md) reproduces and fixes immutable-identity reprovisioning and late rejection of @@ -37,19 +40,24 @@ startup repetitions, plus all 15 residency tests. Retained startup uses an owned in-memory fixture; the separate owned old-binary RustFS upgrade passed, while original-corpus new-release verification remains open. The [original corpus maintenance recovery](performance/2026-10-01-original-corpus-recovery.md) -now leaves all 10,003 Controls durably idle under the old Maintenance release. +left all 10,003 Controls durably idle under the old Maintenance release. Its catalog, roots and 9,702 previously idle Controls stayed unchanged; no new gateway, remote payload verification or performance result is established. -Neither checkpoint replaces the original complete performance schedule. +The subsequent [full-corpus activation](performance/2026-10-01-original-corpus-activation.md) +passed new-release admission and controlled activation to Ready revision 9. +All 10,003 Controls and published roots remained unchanged through activation; +three upgraded gateways are ready, but remote-content verification is still open. +None of these checkpoints replaces the original complete performance schedule. The [three-node proxy campaign](performance/2026-09-30-three-node-proxy.md) retains the earlier fixture race and its regression, plus the candidate's failed seed and successful recovery of every recorded ACK. | Current gate | Evidence / status | | --- | --- | -| Current Cellule revision | `0dc04a6`: independently rebuilt release correctness, lints, Python harness and eight fresh RustFS gates passed. See the latest checkpoint for repeated regressions and exact bindings. Not retained-store upgrade, live deployment or measured performance | +| Qualified Cellule revision | `0dc04a6`: release correctness, lints, Python harness, eight fresh RustFS gates and subsequent original-corpus activation passed. Upstream `1914096` remains unqualified; no measured improvement is claimed | | HTTP listener handoff | [Separate startup qualification](performance/2026-10-01-listener-handoff.md): the affected SHA-256 tests retain their bound listeners through supervised startup; 244 release tests, 84 Python tests, eight fresh RustFS gates and repeated cancellation/handoff checks passed. Both Linux CI workflows at `eaebfc4` passed; not measured performance | -| Original corpus maintenance recovery | [Complete metadata and fenced-recovery checkpoint](performance/2026-10-01-original-corpus-recovery.md): all 256 shards and 10,003 Controls checked twice before and after; 301 unsettled Cells drained, zero unsettled Cells, catalog and all published roots unchanged. Old Maintenance remains closed to serving; new-release admission and remote Git/LFS verification remain open | +| Original corpus maintenance recovery | [Complete metadata and fenced-recovery checkpoint](performance/2026-10-01-original-corpus-recovery.md): all 256 shards and 10,003 Controls checked twice before and after; 301 unsettled Cells drained, zero unsettled Cells, catalog and all published roots unchanged | +| Original corpus activation | [Controlled full-corpus activation](performance/2026-10-01-original-corpus-activation.md): old Ready 6 → Prepared 7 → Activating 8 → new Ready 9; four allowed release/descriptor writes, eight complete unchanged scans. Three upgraded gateways are ready; remote Git/LFS verification and performance remain open | | Owned old-binary RustFS upgrade | [Controlled fixture qualification](performance/2026-10-01-old-binary-rustfs-upgrade.md) passed actual old-byte restore through three fresh gateways/proxy, both new Git/LFS writes, twenty scheduled 17-step workflows and recovery of all forty critical repositories after all three owners were killed. Two retained repositories; not the original 10K corpus or reference capacity | | Original rebuilt Cellule artifact | `c51dd121` build, 227 Rust tests, eight RustFS gates and initial 17-step proxy check passed; full 10K/100 seed closed, but remote verification failed HTTP 503. A separate 10K metadata diagnostic returned 9,999 matching identities and one 503; downstream fault/load gates stopped without owner signals. Cause unresolved | | Combined authentication and residency candidate | First publication passed 231 release tests, eight RustFS gates, 20 cold-activation repetitions and 15 residency tests. Exact predecessor descriptor is retained; see the separately qualified catalog follow-up | diff --git a/docs/performance/2026-10-01-cellule-main-qualification.md b/docs/performance/2026-10-01-cellule-main-qualification.md index 363b541..57d291f 100644 --- a/docs/performance/2026-10-01-cellule-main-qualification.md +++ b/docs/performance/2026-10-01-cellule-main-qualification.md @@ -10,6 +10,10 @@ This follows the [retained catalog admission checkpoint](2026-10-01-retained-cat That checkpoint and the [authentication and residency results](2026-10-01-bounded-authentication-and-residency.md) retain their original Cellule revision and artifacts. +Upstream later advanced to `191409685b001a82bd02780def45102b4fc2f164`. +The PR still pins the independently qualified `0dc04a6`; this checkpoint does +not qualify the newer runtime or peer HTTP CI changes. + ## Dependency change and source equivalence The tested source is `bbd784a40c3867646044a8c716b7ee517f9aca49`. The dependency @@ -92,12 +96,15 @@ verified actual old-executable RustFS data, controlled admission, new-executable restore, twenty scheduled critical workflows and their complete fresh-owner recovery. Its earlier clone-byte mismatch was traced to missing client LFS filters; the full-byte assertion was retained. This closes a small prerequisite, -not full same-corpus upgrade or reference performance. +not reference performance. The later +[original corpus activation](2026-10-01-original-corpus-activation.md) passed +full admission and controlled activation of all 10,003 Cells. Remote Git/LFS +content verification remains a separate open gate. The [full performance plan](../performance-plan.md) still requires three nodes behind a proxy, 10,000 identities, 100 populated Git/LFS fixtures, all 108 windows, 114,960 arrivals and 8,640 scheduled seconds. Full remote verification, -same-corpus upgrade admission, concurrent critical operations and faults, +concurrent critical operations and faults, every acknowledged write after owner loss, higher admission profiles, matched comparisons, non-sparse five-GiB transfers and isolated Linux capacity remain open. The separate diagnostic fleet's terminal lease-fencing failure is also unresolved. diff --git a/docs/performance/2026-10-01-listener-handoff.md b/docs/performance/2026-10-01-listener-handoff.md index 7503773..2e07647 100644 --- a/docs/performance/2026-10-01-listener-handoff.md +++ b/docs/performance/2026-10-01-listener-handoff.md @@ -118,12 +118,14 @@ Both [Linux PR CI](https://github.com/crabbuild/canopy/actions/runs/36950769630) and [push CI](https://github.com/crabbuild/canopy/actions/runs/36950765768) at `eaebfc4cdd7ead54a8d62978aebcd8656f0de0b8` passed Rust, real RustFS compatibility, server build and Python harness checks. PR #18 remains draft. The -[full campaign](../performance-plan.md) still requires original-corpus new-release -admission, full remote Git/LFS verification, all 108 load +[full campaign](../performance-plan.md) still requires full remote Git/LFS verification, all 108 load windows, every acknowledged write after owner loss, higher admission profiles, matched comparisons, non-sparse five-GiB transfers and isolated Linux capacity. The original provider's read-only status at this qualification had zero advertised writers and 301 unsettled Cells; neither its release nor its data was changed here. The later [maintenance recovery](2026-10-01-original-corpus-recovery.md) settled -those Cells without changing catalog identities or published roots. It remains -in old Maintenance; that later action is not a new listener-build performance result. +those Cells without changing catalog identities or published roots. Subsequent +[full-corpus activation](2026-10-01-original-corpus-activation.md) reached new +Ready revision 9 after full admission and unchanged snapshots. Neither later +action is a listener-build performance result. Both PR and push CI at the +documentation-only head `4ddc742` also passed both Rust and harness jobs. diff --git a/docs/performance/2026-10-01-original-corpus-activation.md b/docs/performance/2026-10-01-original-corpus-activation.md new file mode 100644 index 0000000..0764243 --- /dev/null +++ b/docs/performance/2026-10-01-original-corpus-activation.md @@ -0,0 +1,116 @@ +# Original corpus release activation + +The original RustFS corpus passed full new-release admission and controlled +activation without changing any of its **10,003 Cell Controls, catalog entries +or published roots**. The qualified executable uses Cellule `0dc04a6`; its +selected release reached **Ready revision 9**. This closes the release-admission +gate after [same-code maintenance recovery](2026-10-01-original-corpus-recovery.md). +It does not establish remote Git/LFS content recovery, throughput or capacity. + +## Activation sequence + +```text +Old Maintenance 5: all 10,003 Cells idle and unowned + | + +-- read-only admission: actual old descriptor + new rolling compatibility + | all 256 shards and actual Control code/schema + | two unchanged complete snapshots + | + +-- exact old CLI ends maintenance --> Old Ready 6 + | + +-- new SDK prepares release ------> Prepared 7 + +-- new SDK begins activation -----> Activating 8 + +-- new SDK completes activation --> New Ready 9 + | + +-- preserve and reread closed evidence + +-- three fresh gateways and proxy ready + +-- remote Git/LFS verification still open +``` + +The activation controller repeats two complete scans at each of Ready 6, +Prepared 7, Activating 8 and Ready 9. Every snapshot matches the admitted catalog, +canonical Controls, ETags and service root. No serving owner is admitted during +these scans. The later gateway launch is a separate step, not part of this +unchanged-metadata claim. + +## Write boundary + +The fixture-scoped controller uses the existing Cellule release APIs, wrapped +by a transport gate. It arms only after checking the exact old Ready record and +two stable complete snapshots. While armed, it permits only: + +- Creation of the exact new descriptor with its canonical bytes. +- Conditional updates of the exact release key to the three expected records: + Prepared 7, Activating 8 and Ready 9. + +All Control, catalog, node, data, delete, copy and multipart writes remain +refused. The gate disarms after activation. It forwarded **four writes**, with +zero refused attempts during activation. The exact old CLI's preceding +maintenance-end transition is separate from that four-write count. + +Three unit tests cover read-only refusal, exact armed paths/bytes/write modes, +and real ReleaseStore transitions with canonical readback. Locked offline +metadata, release build and all-target Clippy passed. The earlier test attempt +that incorrectly listed an exact leaf as a directory prefix is preserved; its +SDK transitions had passed, and the corrected test verifies the exact record. +The controller is a retained qualification tool, not a general upgrade command. + +## Closed results + +| Check | Result | +| --- | --- | +| Corpus admission | All 256 shards and 10,003 Cells; actual predecessor descriptor and every catalog/Control code/schema supported | +| Admission snapshots | Two identical scans; zero owners, advertisements, unsettled Cells or attempted provider writes | +| Old maintenance end | Exact retained executable exited zero; old Ready revision 6, zero advertisements and unsettled Cells | +| Activation | SDK transitions 7 → 8 → 9; exact new descriptor and selected image verified | +| Activation snapshots | Eight complete scans; all canonical Controls, ETags, catalog and published roots unchanged | +| Provider | Container, volume, image, configuration, start time, resource envelope and restart count unchanged | +| Evidence | 369 closed files and exact inputs copied and independently reread on another local filesystem | + +Admission closed October 2 at 02:25:04 UTC; activation closed at 02:32:05 UTC +(October 1 Pacific). The controller took 225.749 seconds including complete +rescans. That is activation wall time, not Git request latency. + +## Artifact bindings + +| Artifact | SHA-256 | +| --- | --- | +| Qualified Canopy executable | `a61ef0f2cb977348e4e4fc45330334a1f8fd68bf8c44146cd9343b0e6676a6c8` | +| Scoped activation controller | `d5ebef17d975a349bb6db250e2154198dbda53a273ad1078db1958e8b819373f` | +| Admission receipt | `39cdcfce11fdc4e2261beaa51b4640a2321756b014d04887082d2a57ca614fe5` | +| Activation receipt | `316ef6b8449b3a41c09d41dd2578474ad77b1f17147d06ab14be45f035d66981` | +| Backup manifest | `03ca94a54b7683b32c717bc0fbd9b5839f1d0f9f705e696431b79c7d387294e8` | + +The Canopy source is `3dda2b47b1cba105642a62b4ad27d7c84d0940d5`, independently +qualified in the [listener handoff checkpoint](2026-10-01-listener-handoff.md). +The selected release is +`9a8df7ae5feba1f1760a843bc88d7870af4ebf515b8bc48e45bb7433569cd5d0`; +activation operation is `d36626b9-361b-47dc-a43a-3e14f9bd1a7d`. + +Receipts and exact tools are retained at +`/Users/haipingfu/.codex/canopy-original-corpus-admission-YxGodu`. +The verified copy is +`/Volumes/Workspace/CrabData/canopy-original-upgrade-activation-7ze24cer`. +This is not an off-machine or provider-data backup. No runtime budget was raised, +owner erased, Control force-CAS applied, catalog rewritten, corpus reseeded or +provider restarted. The UI preview and unrelated providers remain unchanged. + +## Remaining verification + +Three fresh gateways reported ready behind the proxy at 02:35:59 UTC with the +exact qualified executable and 100 active repositories per node. Readiness is +not a full remote-content result. Full verification of 10,000 identities, +100 populated Git/LFS fixtures and the two critical repositories remains open, +as does fresh-owner recovery of every acknowledged write. + +Cellule upstream subsequently advanced by two commits to +`191409685b001a82bd02780def45102b4fc2f164`, observed when publishing this checkpoint. +Those commits change runtime forwarding/compaction and peer HTTP CI gates. +They are **not** the dependency revision tested here and require independent +qualification; no performance result transfers to them. + +PR #18 remains draft. The [full campaign](../performance-plan.md) still requires +108 windows, 114,960 arrivals and 8,640 scheduled seconds, concurrent critical +operations and faults, higher admission profiles, matched comparisons, large +transfers and isolated Linux capacity. The earlier diagnostic lease-fencing +failure remains unexplained; successful activation does not establish its cause. diff --git a/docs/performance/2026-10-01-original-corpus-recovery.md b/docs/performance/2026-10-01-original-corpus-recovery.md index 5f2a0b5..39a1ad4 100644 --- a/docs/performance/2026-10-01-original-corpus-recovery.md +++ b/docs/performance/2026-10-01-original-corpus-recovery.md @@ -1,17 +1,22 @@ # Original corpus maintenance recovery -The original RustFS corpus now has **10,003 durably idle, unowned Cells**: +At this recovery checkpoint, the original RustFS corpus had +**10,003 durably idle, unowned Cells**: 10,000 repository identities, two critical Git repositories and Directory. Normal fenced recovery with the exact old executable settled all 301 remaining Cells. Two complete post-recovery snapshots and an independent receipt audit passed. All published roots, catalog identities and Cell incarnations stayed unchanged. -The deployment remains in **old-release Maintenance**, revision 5. No new +The deployment remained in **old-release Maintenance**, revision 5. No new release or serving gateway was admitted. This closes the offline metadata and maintenance-recovery prerequisite, not remote Git/LFS content verification, same-corpus upgrade, performance or an explanation of the earlier lease failure. +The subsequent [full-corpus release activation](2026-10-01-original-corpus-activation.md) +passed separate new-release admission and reached Ready revision 9 without +changing these Controls or roots. Remote Git/LFS verification remains open. + ## Recovery sequence ```text @@ -102,14 +107,14 @@ collision are preserved separately; existing receipts were never overwritten. | Independent closed audit | `219331a18c0dcb0e9be65885f22f73bea610befefb66ec3d70e4b018824fd95a` | The checks ran October 2 at 01:46–01:57 UTC, October 1 Pacific. -The selected old release remains +The selected old release at this checkpoint was `e31bf1a951e2fa19d91e9f964b2ddeade1a81b05a20ad628362819a1487c16b1`; maintenance operation is `74f10df1-417c-4c63-83b0-b554fdb996e4`. The UI preview and unrelated providers were not changed or signalled. -Next gates are rolling-descriptor and full actual-Control admission of the new -executable on this same corpus, controlled activation, three fresh gateways and -a proxy, full remote Git/LFS verification and fresh-owner/every-ACK recovery. +The subsequent activation checkpoint closes rolling-descriptor and full +actual-Control admission, controlled activation and three-gateway readiness. +Full remote Git/LFS verification and fresh-owner/every-ACK recovery remain open. The [full campaign](../performance-plan.md) remains **108 windows, 114,960 arrivals and 8,640 scheduled seconds**, with independent rates/concurrency, uniform/skewed access, higher residency profiles, concurrent critical workflows From 085a278cfad4ff136b9d056685f2fa21c275eaff Mon Sep 17 00:00:00 2001 From: forhappy Date: Thu, 1 Oct 2026 20:39:32 -0700 Subject: [PATCH 49/60] docs: record full corpus verification and failed load arrivals --- docs/README.md | 2 +- docs/performance-plan.md | 21 ++- .../2026-10-01-cellule-main-qualification.md | 10 +- .../2026-10-01-listener-handoff.md | 9 +- .../2026-10-01-original-corpus-activation.md | 122 ++++++++++++++++-- .../2026-10-01-original-corpus-recovery.md | 4 +- 6 files changed, 141 insertions(+), 27 deletions(-) diff --git a/docs/README.md b/docs/README.md index a51f268..9254f07 100644 --- a/docs/README.md +++ b/docs/README.md @@ -14,7 +14,7 @@ Use this page to choose a document by task. Canopy's hosting core supports stock | Find the right Rust crate | [Rust workspace](workspace.md) | Crate ownership, dependencies and build commands | | Decide whether a release gate is closed | [Delivery plan](delivery-plan.md) | Required proof, current state and chronological implementation evidence | | Plan or evaluate capacity | [Repository density and latency](performance-plan.md) | Workloads, targets, measured results and limits of each result | -| Inspect the original RustFS corpus upgrade | [Full-corpus release activation](performance/2026-10-01-original-corpus-activation.md) | Admission of 10,003 Cells, exact release transitions, unchanged metadata and open remote-content gates | +| Inspect the original RustFS corpus upgrade | [Full-corpus activation and remote verification](performance/2026-10-01-original-corpus-activation.md) | Admission of 10,003 Cells, full Git/LFS verification, failed diagnostic load windows and open owner-loss gates | | Track three nodes behind a proxy and the latest dependency candidate | [Three-node proxy qualification](performance/2026-09-30-three-node-proxy.md) | Baseline/candidate pins, complete-corpus recovery, failing load windows and open gates | | Inspect the merged workspace and RustFS verification | [Workspace and RustFS verification](performance/2026-09-30-workspace-rustfs.md) | The `70bd25f` revision, conflict resolution and end-to-end gates | | Inspect the earlier Cellule dependency update | [Initial September 30 requalification](performance/2026-09-30-cellule-main.md) | Earlier revision, passing checks, failed suites and storage-blocked performance gates | diff --git a/docs/performance-plan.md b/docs/performance-plan.md index 03c4085..33223a1 100644 --- a/docs/performance-plan.md +++ b/docs/performance-plan.md @@ -5,8 +5,9 @@ It separates demonstrated behavior from proposed targets. The current Cellule dependency pin is `0dc04a658bd99668936f7ec58032d054f6fbc141`, verified against upstream main on October 1. Its [qualification checkpoint](performance/2026-10-01-cellule-main-qualification.md) records 239 release Rust tests, 84 Python tests, lints and all eight fresh RustFS -gates, with separately retained artifacts. Original-corpus remote-content -verification and reference performance remain open. Upstream later advanced to +gates, with separately retained artifacts. Subsequent original-corpus remote +verification passed; scheduled load includes failed arrivals, and post-load +owner-loss recovery and reference capacity remain open. Upstream later advanced to `191409685b001a82bd02780def45102b4fc2f164`; that revision is not yet qualified by this PR. The preceding `c51dd121` build is historical evidence. Unlike the earlier `a4500add` documentation @@ -38,7 +39,7 @@ unsupported Control metadata. Its combined candidate passed 239 release tests, all eight RustFS gates, 20 cold-activation repetitions and 20 three-case retained startup repetitions, plus all 15 residency tests. Retained startup uses an owned in-memory fixture; the separate owned old-binary RustFS upgrade passed, while -original-corpus new-release verification remains open. The +original-corpus verification had not run at that checkpoint. The [original corpus maintenance recovery](performance/2026-10-01-original-corpus-recovery.md) left all 10,003 Controls durably idle under the old Maintenance release. Its catalog, roots and 9,702 previously idle Controls stayed unchanged; no new @@ -46,7 +47,12 @@ gateway, remote payload verification or performance result is established. The subsequent [full-corpus activation](performance/2026-10-01-original-corpus-activation.md) passed new-release admission and controlled activation to Ready revision 9. All 10,003 Controls and published roots remained unchanged through activation; -three upgraded gateways are ready, but remote-content verification is still open. +three upgraded gateways subsequently passed full remote verification of 10,000 +identities, 100 full LFS bodies, 200 v0/v2 clones and both critical fixtures. +The first ten scheduled metadata windows recorded 23,779 OK / 24,000 arrivals, +206 busy drops and 15 HTTP 503s. Concurrent critical load recorded 19/20 OK, +one busy drop and 323 successful steps. Both are failed arrival gates, not +capacity passes; the full schedule continues unchanged. None of these checkpoints replaces the original complete performance schedule. The [three-node proxy campaign](performance/2026-09-30-three-node-proxy.md) retains the earlier fixture race and its regression, plus the candidate's failed @@ -57,7 +63,8 @@ seed and successful recovery of every recorded ACK. | Qualified Cellule revision | `0dc04a6`: release correctness, lints, Python harness, eight fresh RustFS gates and subsequent original-corpus activation passed. Upstream `1914096` remains unqualified; no measured improvement is claimed | | HTTP listener handoff | [Separate startup qualification](performance/2026-10-01-listener-handoff.md): the affected SHA-256 tests retain their bound listeners through supervised startup; 244 release tests, 84 Python tests, eight fresh RustFS gates and repeated cancellation/handoff checks passed. Both Linux CI workflows at `eaebfc4` passed; not measured performance | | Original corpus maintenance recovery | [Complete metadata and fenced-recovery checkpoint](performance/2026-10-01-original-corpus-recovery.md): all 256 shards and 10,003 Controls checked twice before and after; 301 unsettled Cells drained, zero unsettled Cells, catalog and all published roots unchanged | -| Original corpus activation | [Controlled full-corpus activation](performance/2026-10-01-original-corpus-activation.md): old Ready 6 → Prepared 7 → Activating 8 → new Ready 9; four allowed release/descriptor writes, eight complete unchanged scans. Three upgraded gateways are ready; remote Git/LFS verification and performance remain open | +| Original corpus activation and remote verification | [Full-corpus checkpoint](performance/2026-10-01-original-corpus-activation.md): Ready 9 after four allowed writes and eight unchanged activation scans; subsequent 10K identities, 100 LFS bodies, 200 clones and critical-2 verification passed. Independent audit and 7,061-file verified copy closed; post-load owner-loss recovery remains open | +| Current original-corpus load snapshot | First ten metadata windows: 23,779/24,000 OK, 206 busy drops, 15 HTTP 503s. Concurrent critical schedule: 19/20 OK, one busy drop, 323 successful steps, 38 UUIDs. Sealed ledgers audited and copied; unchanged full schedule still running, not reference capacity | | Owned old-binary RustFS upgrade | [Controlled fixture qualification](performance/2026-10-01-old-binary-rustfs-upgrade.md) passed actual old-byte restore through three fresh gateways/proxy, both new Git/LFS writes, twenty scheduled 17-step workflows and recovery of all forty critical repositories after all three owners were killed. Two retained repositories; not the original 10K corpus or reference capacity | | Original rebuilt Cellule artifact | `c51dd121` build, 227 Rust tests, eight RustFS gates and initial 17-step proxy check passed; full 10K/100 seed closed, but remote verification failed HTTP 503. A separate 10K metadata diagnostic returned 9,999 matching identities and one 503; downstream fault/load gates stopped without owner signals. Cause unresolved | | Combined authentication and residency candidate | First publication passed 231 release tests, eight RustFS gates, 20 cold-activation repetitions and 15 residency tests. Exact predecessor descriptor is retained; see the separately qualified catalog follow-up | @@ -70,9 +77,9 @@ seed and successful recovery of every recorded ACK. | Post-OOM diagnostic | Three old owners killed as a verified batch; 32.033-s absence wait recorded. Same provider/data restarted with memory raised from 2 to 4 GiB; full 10K/100 and critical-2 recovery passed, with independent 200-clone/four-mirror audits | | Complete 4-GiB load attempt | All 108 windows finished: 60,197 OK and 54,763 failed arrivals. The original attempt remains failed; local raw evidence is now unavailable | | Post-load ACK recovery | Recorded three-owner loss and 32.249-s wait; critical fixtures passed. Fresh fleet then recorded ENOSPC and shut down; full corpus and every-ACK verification are incomplete | -| Concurrent critical Git workflows | [Separate driver](../scripts/benchmark_critical_git.py) retains the 17-step suite. The versioned gate separates failed performance from mandatory correctness recovery; no live critical-load/fault result is established | +| Concurrent critical Git workflows | [Separate driver](../scripts/benchmark_critical_git.py) retains the 17-step suite. Original-corpus concurrent load closed at 19/20 OK with one busy drop; all 19 attempted receipts complete. Fresh-owner verification of all 38 UUIDs and concurrent fault coverage remain open; the small-fixture 20/20 recovery result is separate | | Evidence availability | Old benchmark directories and executable disappeared; old receipts remain unavailable. New closed artifacts, all build bindings and the completed seed have verified backups on a separate filesystem outside Cargo targets; changing outputs are excluded | -| Performance qualification | Latest full-corpus recovery, scheduled load, every post-load ACK, critical-operation load/fault coverage, higher admission profiles and matched comparisons remain open | +| Performance qualification | Pre-load original-corpus remote verification passed. Complete scheduled load/audit, original-corpus and every-ACK recovery after owner loss, concurrent faults, higher admission profiles and matched comparisons remain open | The [native-filesystem trial](performance/2026-10-01-native-filesystem.md) records the storage hypothesis, graceful diagnostic drain and new bindings. diff --git a/docs/performance/2026-10-01-cellule-main-qualification.md b/docs/performance/2026-10-01-cellule-main-qualification.md index 57d291f..807ec76 100644 --- a/docs/performance/2026-10-01-cellule-main-qualification.md +++ b/docs/performance/2026-10-01-cellule-main-qualification.md @@ -98,13 +98,15 @@ recovery. Its earlier clone-byte mismatch was traced to missing client LFS filters; the full-byte assertion was retained. This closes a small prerequisite, not reference performance. The later [original corpus activation](2026-10-01-original-corpus-activation.md) passed -full admission and controlled activation of all 10,003 Cells. Remote Git/LFS -content verification remains a separate open gate. +full admission and controlled activation of all 10,003 Cells. Its subsequent +full remote Git/LFS verification passed with an independent audit and verified +7,061-file copy. The first ten scheduled metadata windows and concurrent +critical schedule include failed arrivals; see that checkpoint's result tables. The [full performance plan](../performance-plan.md) still requires three nodes behind a proxy, 10,000 identities, 100 populated Git/LFS fixtures, all 108 windows, -114,960 arrivals and 8,640 scheduled seconds. Full remote verification, -concurrent critical operations and faults, +114,960 arrivals and 8,640 scheduled seconds. Complete scheduled load/audit, +concurrent faults, original-corpus fresh-owner recovery, every acknowledged write after owner loss, higher admission profiles, matched comparisons, non-sparse five-GiB transfers and isolated Linux capacity remain open. The separate diagnostic fleet's terminal lease-fencing failure is also unresolved. diff --git a/docs/performance/2026-10-01-listener-handoff.md b/docs/performance/2026-10-01-listener-handoff.md index 2e07647..8e92812 100644 --- a/docs/performance/2026-10-01-listener-handoff.md +++ b/docs/performance/2026-10-01-listener-handoff.md @@ -118,7 +118,7 @@ Both [Linux PR CI](https://github.com/crabbuild/canopy/actions/runs/36950769630) and [push CI](https://github.com/crabbuild/canopy/actions/runs/36950765768) at `eaebfc4cdd7ead54a8d62978aebcd8656f0de0b8` passed Rust, real RustFS compatibility, server build and Python harness checks. PR #18 remains draft. The -[full campaign](../performance-plan.md) still requires full remote Git/LFS verification, all 108 load +[full campaign](../performance-plan.md) still requires completion of all 108 load windows, every acknowledged write after owner loss, higher admission profiles, matched comparisons, non-sparse five-GiB transfers and isolated Linux capacity. The original provider's read-only status at this qualification had zero advertised @@ -127,5 +127,8 @@ The later [maintenance recovery](2026-10-01-original-corpus-recovery.md) settled those Cells without changing catalog identities or published roots. Subsequent [full-corpus activation](2026-10-01-original-corpus-activation.md) reached new Ready revision 9 after full admission and unchanged snapshots. Neither later -action is a listener-build performance result. Both PR and push CI at the -documentation-only head `4ddc742` also passed both Rust and harness jobs. +action is a listener-build performance result. The activation checkpoint's +subsequent full remote Git/LFS verification passed; its first ten metadata +windows and concurrent critical load include failed arrivals. Both PR and push +CI at documentation-only heads `4ddc742` and `2c897c7` passed both Rust and +harness jobs. Later documentation commits require their own fresh CI. diff --git a/docs/performance/2026-10-01-original-corpus-activation.md b/docs/performance/2026-10-01-original-corpus-activation.md index 0764243..3950ae8 100644 --- a/docs/performance/2026-10-01-original-corpus-activation.md +++ b/docs/performance/2026-10-01-original-corpus-activation.md @@ -5,7 +5,9 @@ activation without changing any of its **10,003 Cell Controls, catalog entries or published roots**. The qualified executable uses Cellule `0dc04a6`; its selected release reached **Ready revision 9**. This closes the release-admission gate after [same-code maintenance recovery](2026-10-01-original-corpus-recovery.md). -It does not establish remote Git/LFS content recovery, throughput or capacity. +Activation alone does not establish remote Git/LFS content recovery, throughput +or capacity. The subsequent remote verification passed; the first ten load +windows below include failed arrivals and do not establish capacity. ## Activation sequence @@ -24,7 +26,9 @@ Old Maintenance 5: all 10,003 Cells idle and unowned | +-- preserve and reread closed evidence +-- three fresh gateways and proxy ready - +-- remote Git/LFS verification still open + +-- remote Git/LFS verification passed + +-- full load running: failed arrivals retained + +-- fresh-owner/every-ACK recovery still open ``` The activation controller repeats two complete scans at each of Ready 6, @@ -95,22 +99,118 @@ This is not an off-machine or provider-data backup. No runtime budget was raised owner erased, Control force-CAS applied, catalog rewritten, corpus reseeded or provider restarted. The UI preview and unrelated providers remain unchanged. -## Remaining verification +## Full remote-content verification Three fresh gateways reported ready behind the proxy at 02:35:59 UTC with the -exact qualified executable and 100 active repositories per node. Readiness is -not a full remote-content result. Full verification of 10,000 identities, -100 populated Git/LFS fixtures and the two critical repositories remains open, -as does fresh-owner recovery of every acknowledged write. +exact qualified executable and 100 active repositories per node. The subsequent +read-only remote verifier closed successfully at **03:02:09 UTC**; an independent +offline audit closed at **03:02:40 UTC** on October 2 (October 1 Pacific). + +| Check | Closed result | +| --- | --- | +| Original repository identities | All 10,000 names matched their exact UUIDs | +| Original LFS objects | All 100 complete bodies matched their expected sizes and SHA-256 digests | +| Stock Git | 200 clones across protocols v0/v2; exact HEAD, base commit, fixture files and full fsck | +| Original critical fixtures | Both UUIDs, four exact v0/v2 ref inventories and mirrors, payloads, notes and full fsck | +| Independent audit | 10,002 identities, 100 LFS digests, 824 Git commands, all 200 local clones and four mirrors reconciled | +| Evidence preservation | 7,061 closed files and exact inputs copied and independently reread on a different local filesystem | + +HTTP checks used concurrency 16 and a 30-second timeout; Git commands retained +their 120-second timeout. No retries, new seed or provider restart were used. +The original fixtures upload LFS objects directly without committing LFS pointer +files: this verifies complete download bytes, **not clone-side LFS hydration**. +The 756.929-second corpus stage is verification wall time, not scheduled request +latency. This check does not establish post-load owner-loss recovery. + +| Remote-verification artifact | SHA-256 | +| --- | --- | +| Verification receipt | `1447486483a253c3fad18000a73eaa449dbaa26a1223903c074497d2b431bd4a` | +| Independent audit | `2b1a256a40cec3bcf24359a21559da8b35927d697ec0e25ba5a42982cdb030d0` | +| Verified backup manifest | `783902ec100a41e7882c73478a98b912b06b6ec0293907146a60a7dd7d8b8078` | + +Closed receipts are in the qualification root's `upgraded-original-verification` +directory; the verified copy is +`/Volumes/Workspace/CrabData/canopy-original-upgraded-remote-dooxekr0`. + +## Scheduled load: failed diagnostic baseline + +The unchanged **108-window / 114,960-arrival / 8,640-second** campaign is running +on the same qualified executable and original RustFS provider. This snapshot +covers only its first **ten sealed, audited and preserved metadata windows**, +observed October 2 at 03:36 UTC. All offer 20 requests/s for 120 seconds, with +client concurrency 32 and node residency capped at 100. An active set of 500 +does not raise the per-node residency limit. + +| Active set / distribution / repetition | OK / 2,400 | Busy drops | HTTP 503 | Delivered RPS in window | Scheduled p95 / p99 (ms) | +| --- | --- | --- | --- | --- | --- | +| 100 / uniform / 1 | 2,344 | 56 | 0 | 19.500 | 1,447.299 / 3,198.533 | +| 100 / uniform / 2 | 2,392 | 8 | 0 | 19.900 | 732.404 / 1,556.094 | +| 100 / uniform / 3 | 2,400 | 0 | 0 | 20.000 | 387.751 / 714.107 | +| 100 / skewed / 1 | 2,400 | 0 | 0 | 20.000 | 169.310 / 388.415 | +| 100 / skewed / 2 | 2,400 | 0 | 0 | 20.000 | 206.813 / 395.493 | +| 100 / skewed / 3 | 2,400 | 0 | 0 | 20.000 | 297.194 / 611.530 | +| 500 / uniform / 1 | 2,349 | 50 | 1 | 19.467 | 1,360.391 / 2,787.405 | +| 500 / uniform / 2 | 2,315 | 76 | 9 | 19.267 | 1,557.115 / 3,131.934 | +| 500 / uniform / 3 | 2,379 | 16 | 5 | 19.783 | 684.082 / 1,993.517 | +| 500 / skewed / 1 | 2,400 | 0 | 0 | 19.983 | 1,012.309 / 1,710.166 | + +Total: **23,779 OK / 24,000 scheduled, 206 busy drops and 15 HTTP 503s**. +Busy drops have no completed-request latency and are not silently omitted from +arrival counts. Percentiles cover completed attempts, including HTTP errors; +RPS counts only successful completions inside the schedule window. Successful +requests completed during drain remain in OK counts but not in-window RPS. +No rate, cap, assertion or timeout was relaxed to obtain a passing result. + +The concurrent 300-second critical schedule also failed its arrival gate: + +| Check | Closed result | +| --- | --- | +| Scheduled workflows | 19 OK / 20 scheduled; one driver-busy drop, no writes for that dropped arrival | +| Attempted workflows | All 19 receipts complete: 323 successful steps and 38 acknowledged repository UUIDs | +| Delivered workflow throughput | 0.060000/s in-window; 0.060223/s including drain | +| Whole-workflow p50 / p95 / p99 | 37.089 / 60.551 / 60.551 seconds, including stock-Git work and validation | +| Correctness boundary | Ledger and receipt integrity passed; recovery of these ACKs after owner loss remains open | + +The read-only watcher audits each sealed window's exact sequence, deterministic +selection, outcomes, latency, throughput and resource bindings before copying +four finalized files. It also preserves each attempted workflow's receipt and +local Git data. Changing campaign indexes, logs and live provider data are not +copied. Observation and file-copy work are separate overhead; process self-CPU +and proxy counters are not full-host CPU or cost measurements. + +Campaign outputs are at +`/Volumes/Workspace/CrabData/canopy-original-full-0dc04a6-j6promvd`; +per-window copies and manifests are at +`/Users/haipingfu/.codex/canopy-original-full-window-copies-qknece2i`. +The critical audit is `closed-critical-load-audit.json` in the qualification root; +its verified final report/source copy is +`/Users/haipingfu/.codex/canopy-original-closed-critical-dved6eyf`. +These are local evidence copies, not off-machine or provider-data backups. + +| Closed critical artifact | SHA-256 | +| --- | --- | +| Report | `66a1468b8e51b2b4254813d06af12e04e802fa5598512f6bb0081cc484d16285` | +| Sample ledger | `1ae9c44266cb49eeffaec80406ac74f3e3a342e6ab5b07fc42761a99fdb9988a` | +| Verified backup manifest | `b8effa1957630581b09d1eec6ec7d25d8f7813449c956b8f203b564db50c8ea3` | + +Diagnosis remains open. In the first two windows, dispatch p99 was +12.760/20.340 ms versus service p99 3,190.791/1,550.308 ms. Proxy connection +deltas were balanced (85/85/84 and 106/106/105), with no proxy errors or +rejections. This weakens timer scheduling and connection-count imbalance as +dominant explanations; it does not prove a Directory or RustFS bottleneck. + +## Remaining verification Cellule upstream subsequently advanced by two commits to `191409685b001a82bd02780def45102b4fc2f164`, observed when publishing this checkpoint. Those commits change runtime forwarding/compaction and peer HTTP CI gates. -They are **not** the dependency revision tested here and require independent -qualification; no performance result transfers to them. +They are **not** the dependency revision tested here. An isolated candidate +passed locked metadata and formatting checks only; release and RustFS +qualification remain open, and no performance result transfers to it. PR #18 remains draft. The [full campaign](../performance-plan.md) still requires -108 windows, 114,960 arrivals and 8,640 scheduled seconds, concurrent critical -operations and faults, higher admission profiles, matched comparisons, large +the complete schedule and audit, concurrent faults, fresh-owner verification +of the original corpus and every newly acknowledged write, higher admission +profiles, matched comparisons, large transfers and isolated Linux capacity. The earlier diagnostic lease-fencing failure remains unexplained; successful activation does not establish its cause. diff --git a/docs/performance/2026-10-01-original-corpus-recovery.md b/docs/performance/2026-10-01-original-corpus-recovery.md index 39a1ad4..78e79db 100644 --- a/docs/performance/2026-10-01-original-corpus-recovery.md +++ b/docs/performance/2026-10-01-original-corpus-recovery.md @@ -114,7 +114,9 @@ The UI preview and unrelated providers were not changed or signalled. The subsequent activation checkpoint closes rolling-descriptor and full actual-Control admission, controlled activation and three-gateway readiness. -Full remote Git/LFS verification and fresh-owner/every-ACK recovery remain open. +Its subsequent full remote Git/LFS verification passed with an independent +audit and verified evidence copy. Fresh-owner/every-ACK recovery after the +current load remains open; that checkpoint records failed diagnostic arrivals. The [full campaign](../performance-plan.md) remains **108 windows, 114,960 arrivals and 8,640 scheduled seconds**, with independent rates/concurrency, uniform/skewed access, higher residency profiles, concurrent critical workflows From ec816507cb4cf853b48e2f64a807eb735b8a1c19 Mon Sep 17 00:00:00 2001 From: forhappy Date: Thu, 1 Oct 2026 21:35:43 -0700 Subject: [PATCH 50/60] docs: record completed original-corpus benchmark phases --- docs/performance-plan.md | 10 +-- .../2026-10-01-cellule-main-qualification.md | 5 +- .../2026-10-01-original-corpus-activation.md | 67 +++++++++++++++++-- 3 files changed, 72 insertions(+), 10 deletions(-) diff --git a/docs/performance-plan.md b/docs/performance-plan.md index 33223a1..4a38314 100644 --- a/docs/performance-plan.md +++ b/docs/performance-plan.md @@ -49,9 +49,11 @@ passed new-release admission and controlled activation to Ready revision 9. All 10,003 Controls and published roots remained unchanged through activation; three upgraded gateways subsequently passed full remote verification of 10,000 identities, 100 full LFS bodies, 200 v0/v2 clones and both critical fixtures. -The first ten scheduled metadata windows recorded 23,779 OK / 24,000 arrivals, -206 busy drops and 15 HTTP 503s. Concurrent critical load recorded 19/20 OK, -one busy drop and 323 successful steps. Both are failed arrival gates, not +The first four completed load phases recorded 41,485/43,200 metadata OK, +1,464/2,160 creation OK, 10,309/57,600 HTTP v2 discovery OK and 632/1,200 +stock-Git ls-remote OK across 40 windows. Busy drops and protocol/client errors +remain in the result tables. Concurrent critical load recorded 19/20 OK, +one busy drop and 323 successful steps. These are failed arrival gates, not capacity passes; the full schedule continues unchanged. None of these checkpoints replaces the original complete performance schedule. The [three-node proxy campaign](performance/2026-09-30-three-node-proxy.md) @@ -64,7 +66,7 @@ seed and successful recovery of every recorded ACK. | HTTP listener handoff | [Separate startup qualification](performance/2026-10-01-listener-handoff.md): the affected SHA-256 tests retain their bound listeners through supervised startup; 244 release tests, 84 Python tests, eight fresh RustFS gates and repeated cancellation/handoff checks passed. Both Linux CI workflows at `eaebfc4` passed; not measured performance | | Original corpus maintenance recovery | [Complete metadata and fenced-recovery checkpoint](performance/2026-10-01-original-corpus-recovery.md): all 256 shards and 10,003 Controls checked twice before and after; 301 unsettled Cells drained, zero unsettled Cells, catalog and all published roots unchanged | | Original corpus activation and remote verification | [Full-corpus checkpoint](performance/2026-10-01-original-corpus-activation.md): Ready 9 after four allowed writes and eight unchanged activation scans; subsequent 10K identities, 100 LFS bodies, 200 clones and critical-2 verification passed. Independent audit and 7,061-file verified copy closed; post-load owner-loss recovery remains open | -| Current original-corpus load snapshot | First ten metadata windows: 23,779/24,000 OK, 206 busy drops, 15 HTTP 503s. Concurrent critical schedule: 19/20 OK, one busy drop, 323 successful steps, 38 UUIDs. Sealed ledgers audited and copied; unchanged full schedule still running, not reference capacity | +| Current original-corpus load snapshot | Four completed phases / 40 windows: metadata 41,485/43,200 OK; creation 1,464/2,160; HTTP v2 discovery 10,309/57,600; stock-Git ls-remote 632/1,200. Failed arrivals retained. Concurrent critical schedule: 19/20 OK, 323 successful steps, 38 UUIDs. Sealed ledgers audited and copied; full schedule still running, not reference capacity | | Owned old-binary RustFS upgrade | [Controlled fixture qualification](performance/2026-10-01-old-binary-rustfs-upgrade.md) passed actual old-byte restore through three fresh gateways/proxy, both new Git/LFS writes, twenty scheduled 17-step workflows and recovery of all forty critical repositories after all three owners were killed. Two retained repositories; not the original 10K corpus or reference capacity | | Original rebuilt Cellule artifact | `c51dd121` build, 227 Rust tests, eight RustFS gates and initial 17-step proxy check passed; full 10K/100 seed closed, but remote verification failed HTTP 503. A separate 10K metadata diagnostic returned 9,999 matching identities and one 503; downstream fault/load gates stopped without owner signals. Cause unresolved | | Combined authentication and residency candidate | First publication passed 231 release tests, eight RustFS gates, 20 cold-activation repetitions and 15 residency tests. Exact predecessor descriptor is retained; see the separately qualified catalog follow-up | diff --git a/docs/performance/2026-10-01-cellule-main-qualification.md b/docs/performance/2026-10-01-cellule-main-qualification.md index 807ec76..8ca212d 100644 --- a/docs/performance/2026-10-01-cellule-main-qualification.md +++ b/docs/performance/2026-10-01-cellule-main-qualification.md @@ -100,8 +100,9 @@ not reference performance. The later [original corpus activation](2026-10-01-original-corpus-activation.md) passed full admission and controlled activation of all 10,003 Cells. Its subsequent full remote Git/LFS verification passed with an independent audit and verified -7,061-file copy. The first ten scheduled metadata windows and concurrent -critical schedule include failed arrivals; see that checkpoint's result tables. +7,061-file copy. The four completed load phases (40 windows) and concurrent +critical schedule include failed arrivals; see that checkpoint's result tables +for metadata, creation, HTTP v2 discovery and stock-Git ls-remote. The [full performance plan](../performance-plan.md) still requires three nodes behind a proxy, 10,000 identities, 100 populated Git/LFS fixtures, all 108 windows, diff --git a/docs/performance/2026-10-01-original-corpus-activation.md b/docs/performance/2026-10-01-original-corpus-activation.md index 3950ae8..e3e83b3 100644 --- a/docs/performance/2026-10-01-original-corpus-activation.md +++ b/docs/performance/2026-10-01-original-corpus-activation.md @@ -6,8 +6,8 @@ or published roots**. The qualified executable uses Cellule `0dc04a6`; its selected release reached **Ready revision 9**. This closes the release-admission gate after [same-code maintenance recovery](2026-10-01-original-corpus-recovery.md). Activation alone does not establish remote Git/LFS content recovery, throughput -or capacity. The subsequent remote verification passed; the first ten load -windows below include failed arrivals and do not establish capacity. +or capacity. The subsequent remote verification passed; the completed load +phases below include failed arrivals and do not establish capacity. ## Activation sequence @@ -132,10 +132,69 @@ Closed receipts are in the qualification root's `upgraded-original-verification` directory; the verified copy is `/Volumes/Workspace/CrabData/canopy-original-upgraded-remote-dooxekr0`. -## Scheduled load: failed diagnostic baseline +## Scheduled load results + +The unchanged **108-window / 114,960-arrival / 8,640-second** campaign remains +running on the qualified `0dc04a6` executable and original RustFS provider. +At October 2, 04:33 UTC (October 1 Pacific), its first four operation phases +were complete: **40 sealed, audited and preserved windows**. Later clone +windows had also begun; they are excluded from this completed-phase summary. + +| Completed phase | Windows | Scheduled | OK | Busy drops | HTTP 503 | Other failures | +| --- | --- | --- | --- | --- | --- | --- | +| Metadata | 18 | 43,200 | 41,485 | 1,669 | 46 | 0 | +| Repository creation | 6 | 2,160 | 1,464 | 690 | 6 | 0 | +| HTTP Git v2 discovery | 8 | 57,600 | 10,309 | 9,581 | 37,706 | 4 transport errors | +| Stock Git ls-remote | 8 | 1,200 | 632 | 514 | Not separately classified | 54 Git errors | + +These are failed arrival gates. The HTTP discovery probe checks the v2 +capability response, not the full stock-Git ref exchange; `ls-remote` runs the +actual Git client and checks the expected main ref. Busy arrivals do not start +requests and have no completed-request latency. Git errors remain client-level +failures rather than being silently classified as HTTP 503s. + +Fast refusal can make aggregate latency misleading. In the first uniform +100-RPS discovery window, completed attempts had p50 **7.824 ms**, but successful +attempts had p50/p95/p99 **1,406.853/4,472.563/6,349.508 ms**. Only 1,343 of +12,000 arrivals succeeded, delivering **10.992 successful requests/s** inside +the window. This is not a latency improvement or evidence of 100-RPS capacity. + +### Repository creation throughput and latency + +Each window lasts 120 seconds with client concurrency 16. Successful completions +during drain count as OK but not as delivered RPS inside the window. Percentiles +below cover completed attempts, including HTTP errors; they exclude busy drops. + +| Offered RPS / repetition | OK / scheduled | Busy drops | HTTP 503 | Delivered RPS | Scheduled p95 / p99 (ms) | +| --- | --- | --- | --- | --- | --- | +| 1 / 1 | 120 / 120 | 0 | 0 | 1.000 | 3,205.203 / 5,940.919 | +| 1 / 2 | 120 / 120 | 0 | 0 | 1.000 | 5,220.908 / 6,795.877 | +| 1 / 3 | 120 / 120 | 0 | 0 | 1.000 | 5,282.604 / 7,891.099 | +| 5 / 1 | 394 / 600 | 204 | 2 | 3.150 | 11,488.210 / 13,053.231 | +| 5 / 2 | 196 / 600 | 400 | 4 | 1.558 | 18,279.797 / 20,765.950 | +| 5 / 3 | 514 / 600 | 86 | 0 | 4.250 | 5,185.795 / 12,490.197 | + +All 1,464 positive creation ACKs have unique canonical UUIDs and their exact +scheduled names. They do not collide with the original 10,000 repositories, +the two original critical fixtures or the 38 current critical UUIDs. This checks +ledger integrity, not recovery after owner loss. Failed writes are not assumed +to have rolled back. + +### Stock Git ref listing + +The eight 60-second `ls-remote` windows vary offered rate (1 or 4/s) independently +of client concurrency (1 or 16), with two repetitions per combination. The +first concurrency-16, 1-RPS repetition completed all 60 arrivals; the second +returned **41 OK, 16 Git errors and three busy drops**, with completed-attempt +p95 **34,429.751 ms**. The two concurrency-16, 4-RPS repetitions returned +**236/240** and **111/240** OK, delivering **3.833** and **1.750** successful +operations/s respectively. The variability is retained; no matched speedup or +root cause is established. + +### Earlier ten-window metadata snapshot The unchanged **108-window / 114,960-arrival / 8,640-second** campaign is running -on the same qualified executable and original RustFS provider. This snapshot +on the same qualified executable and original RustFS provider. This earlier snapshot covers only its first **ten sealed, audited and preserved metadata windows**, observed October 2 at 03:36 UTC. All offer 20 requests/s for 120 seconds, with client concurrency 32 and node residency capped at 100. An active set of 500 From 842680aa4cbd4cf9a5aebba9dc5d10b35452e337 Mon Sep 17 00:00:00 2001 From: forhappy Date: Thu, 1 Oct 2026 19:54:33 -0700 Subject: [PATCH 51/60] build: pin Cellule to main revision 1914096 --- Cargo.lock | 12 ++++++------ crates/canopy-server/Cargo.toml | 10 +++++----- 2 files changed, 11 insertions(+), 11 deletions(-) diff --git a/Cargo.lock b/Cargo.lock index 4d07109..9c72ff9 100644 --- a/Cargo.lock +++ b/Cargo.lock @@ -448,7 +448,7 @@ dependencies = [ [[package]] name = "cellule-app" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=0dc04a658bd99668936f7ec58032d054f6fbc141#0dc04a658bd99668936f7ec58032d054f6fbc141" +source = "git+https://github.com/crabbuild/cellule.git?rev=191409685b001a82bd02780def45102b4fc2f164#191409685b001a82bd02780def45102b4fc2f164" dependencies = [ "blake3", "cellule-runtime", @@ -457,7 +457,7 @@ dependencies = [ [[package]] name = "cellule-host" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=0dc04a658bd99668936f7ec58032d054f6fbc141#0dc04a658bd99668936f7ec58032d054f6fbc141" +source = "git+https://github.com/crabbuild/cellule.git?rev=191409685b001a82bd02780def45102b4fc2f164#191409685b001a82bd02780def45102b4fc2f164" dependencies = [ "cellule-app", "cellule-runtime", @@ -471,7 +471,7 @@ dependencies = [ [[package]] name = "cellule-ltx" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=0dc04a658bd99668936f7ec58032d054f6fbc141#0dc04a658bd99668936f7ec58032d054f6fbc141" +source = "git+https://github.com/crabbuild/cellule.git?rev=191409685b001a82bd02780def45102b4fc2f164#191409685b001a82bd02780def45102b4fc2f164" dependencies = [ "async-trait", "blake3", @@ -493,7 +493,7 @@ dependencies = [ [[package]] name = "cellule-runtime" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=0dc04a658bd99668936f7ec58032d054f6fbc141#0dc04a658bd99668936f7ec58032d054f6fbc141" +source = "git+https://github.com/crabbuild/cellule.git?rev=191409685b001a82bd02780def45102b4fc2f164#191409685b001a82bd02780def45102b4fc2f164" dependencies = [ "blake3", "bytes", @@ -519,7 +519,7 @@ dependencies = [ [[package]] name = "cellule-store" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=0dc04a658bd99668936f7ec58032d054f6fbc141#0dc04a658bd99668936f7ec58032d054f6fbc141" +source = "git+https://github.com/crabbuild/cellule.git?rev=191409685b001a82bd02780def45102b4fc2f164#191409685b001a82bd02780def45102b4fc2f164" dependencies = [ "async-trait", "blake3", @@ -542,7 +542,7 @@ dependencies = [ [[package]] name = "cellule-types" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=0dc04a658bd99668936f7ec58032d054f6fbc141#0dc04a658bd99668936f7ec58032d054f6fbc141" +source = "git+https://github.com/crabbuild/cellule.git?rev=191409685b001a82bd02780def45102b4fc2f164#191409685b001a82bd02780def45102b4fc2f164" dependencies = [ "schemars", "serde", diff --git a/crates/canopy-server/Cargo.toml b/crates/canopy-server/Cargo.toml index 66fe98c..40e9086 100644 --- a/crates/canopy-server/Cargo.toml +++ b/crates/canopy-server/Cargo.toml @@ -15,11 +15,11 @@ axum = "0.8.9" base64 = "0.22" blake3 = "1.8" bytes = "1.11" -cellule-app = { git = "https://github.com/crabbuild/cellule.git", rev = "0dc04a658bd99668936f7ec58032d054f6fbc141" } -cellule-host = { git = "https://github.com/crabbuild/cellule.git", rev = "0dc04a658bd99668936f7ec58032d054f6fbc141" } -cellule-ltx = { git = "https://github.com/crabbuild/cellule.git", rev = "0dc04a658bd99668936f7ec58032d054f6fbc141", features = ["replica"] } -cellule-runtime = { git = "https://github.com/crabbuild/cellule.git", rev = "0dc04a658bd99668936f7ec58032d054f6fbc141" } -cellule-store = { git = "https://github.com/crabbuild/cellule.git", rev = "0dc04a658bd99668936f7ec58032d054f6fbc141" } +cellule-app = { git = "https://github.com/crabbuild/cellule.git", rev = "191409685b001a82bd02780def45102b4fc2f164" } +cellule-host = { git = "https://github.com/crabbuild/cellule.git", rev = "191409685b001a82bd02780def45102b4fc2f164" } +cellule-ltx = { git = "https://github.com/crabbuild/cellule.git", rev = "191409685b001a82bd02780def45102b4fc2f164", features = ["replica"] } +cellule-runtime = { git = "https://github.com/crabbuild/cellule.git", rev = "191409685b001a82bd02780def45102b4fc2f164" } +cellule-store = { git = "https://github.com/crabbuild/cellule.git", rev = "191409685b001a82bd02780def45102b4fc2f164" } ed25519-dalek = "2" flate2 = "1.1" futures-core = "0.3" From e175238992318b3c61725a6aa6c1ba9ae7c11d4d Mon Sep 17 00:00:00 2001 From: forhappy Date: Thu, 1 Oct 2026 22:15:57 -0700 Subject: [PATCH 52/60] ci: qualify release correctness against RustFS with retained artifacts --- .github/workflows/qualify-release.yml | 109 ++++++++++++++++++++++++++ 1 file changed, 109 insertions(+) create mode 100644 .github/workflows/qualify-release.yml diff --git a/.github/workflows/qualify-release.yml b/.github/workflows/qualify-release.yml new file mode 100644 index 0000000..3f45388 --- /dev/null +++ b/.github/workflows/qualify-release.yml @@ -0,0 +1,109 @@ +name: Qualify release correctness against RustFS + +on: + workflow_dispatch: + push: + branches: + - 'codex/cellule-release-ci-*' + +permissions: + contents: read + +defaults: + run: + shell: bash + +jobs: + release-correctness: + runs-on: ubuntu-latest + timeout-minutes: 60 + steps: + - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 + - name: Install tools + run: | + rustup component add clippy rustfmt + sudo apt-get update + sudo apt-get install -y awscli git-lfs openssh-client + git lfs install + - name: Bind exact source and dependency inputs + run: | + mkdir qualification + cargo metadata --locked --format-version 1 > qualification/metadata.json + docker pull rustfs/rustfs:1.0.0-beta.8-glibc + python3 - <<'PY' + import hashlib + import json + from pathlib import Path + import subprocess + import tomllib + + metadata = json.loads(Path('qualification/metadata.json').read_text()) + manifest = tomllib.loads(Path('crates/canopy-server/Cargo.toml').read_text()) + revisions = {value['rev'] for value in manifest['dependencies'].values() + if isinstance(value, dict) and value.get('git') == 'https://github.com/crabbuild/cellule.git'} + assert len(revisions) == 1 + revision = revisions.pop() + packages = {package['name']: package['source'] for package in metadata['packages'] + if package['name'].startswith('cellule-')} + assert set(packages) == {'cellule-app', 'cellule-host', 'cellule-ltx', + 'cellule-runtime', 'cellule-store', 'cellule-types'} + assert all(source == f'git+https://github.com/crabbuild/cellule.git?rev={revision}#{revision}' + for source in packages.values()) + files = subprocess.check_output(['git', 'ls-files', 'crates', 'scripts', 'Cargo.toml', + 'Cargo.lock', '.github/workflows'], text=True).splitlines() + image = json.loads(subprocess.check_output(['docker', 'image', 'inspect', + 'rustfs/rustfs:1.0.0-beta.8-glibc'], text=True))[0] + result = { + 'head': subprocess.check_output(['git', 'rev-parse', 'HEAD'], text=True).strip(), + 'cellule_revision': revision, + 'packages': packages, + 'source_sha256': {name: hashlib.sha256(Path(name).read_bytes()).hexdigest() for name in files}, + 'rustc': subprocess.check_output(['rustc', '-Vv'], text=True), + 'cargo': subprocess.check_output(['cargo', '-V'], text=True).strip(), + 'git': subprocess.check_output(['git', '--version'], text=True).strip(), + 'rustfs_image': {'id': image['Id'], 'repo_digests': image.get('RepoDigests', [])}, + 'scope': 'Linux release correctness and fresh RustFS compatibility only; ' + 'not native Mac qualification, retained-store upgrade, recovery or reference capacity.', + } + Path('qualification/inputs.json').write_text(json.dumps(result, indent=2) + '\n') + PY + cp Cargo.lock qualification/Cargo.lock + cp .github/workflows/qualify-release.yml qualification/qualify-release.yml + - name: Check release formatting and all-target lints + run: | + cargo fmt --all -- --check + cargo clippy --release --workspace --all-targets --locked -- -D warnings 2>&1 | tee qualification/release-clippy.log + - name: Test release workspace + run: cargo test --release --workspace --locked -- --test-threads=4 2>&1 | tee qualification/release-workspace.log + - name: Check Python qualification harness + run: python3 -B -m unittest discover -s scripts -p 'test_*.py' 2>&1 | tee qualification/python-harness.log + - name: Qualify release Git compatibility against RustFS + run: python3 scripts/qualify_size.py --provider-only --release 2>&1 | tee qualification/release-rustfs.log + - name: Build and retain release executable + run: | + cargo build --release --locked --bin canopy 2>&1 | tee qualification/release-build.log + python3 - <<'PY' + import hashlib + import json + from pathlib import Path + + inputs = json.loads(Path('qualification/inputs.json').read_text()) + assert all(hashlib.sha256(Path(name).read_bytes()).hexdigest() == digest + for name, digest in inputs['source_sha256'].items()) + binary = Path('target/release/canopy') + digest = hashlib.sha256(binary.read_bytes()).hexdigest() + result = {'head': inputs['head'], 'cellule_revision': inputs['cellule_revision'], + 'binary_sha256': digest, 'binary_size_bytes': binary.stat().st_size, + 'bound_source_unchanged': True, 'scope': inputs['scope']} + Path('qualification/release-output.json').write_text(json.dumps(result, indent=2) + '\n') + PY + sha256sum target/release/canopy > qualification/canopy.sha256 + tar -czf qualification/canopy-release-linux.tar.gz -C target/release canopy + - name: Retain qualification evidence including failed attempts + if: always() + uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4 + with: + name: release-correctness-${{ github.sha }}-${{ github.run_attempt }} + path: qualification/ + if-no-files-found: error + retention-days: 30 From 51fa42e81527697ba5ac7a5d0fbd198e0dbe0d03 Mon Sep 17 00:00:00 2001 From: forhappy Date: Thu, 1 Oct 2026 22:20:13 -0700 Subject: [PATCH 53/60] ci: use runner AWS CLI and retain evidence before setup --- .github/workflows/qualify-release.yml | 9 +++++++-- 1 file changed, 7 insertions(+), 2 deletions(-) diff --git a/.github/workflows/qualify-release.yml b/.github/workflows/qualify-release.yml index 3f45388..248a07e 100644 --- a/.github/workflows/qualify-release.yml +++ b/.github/workflows/qualify-release.yml @@ -19,15 +19,20 @@ jobs: timeout-minutes: 60 steps: - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 + - name: Initialize qualification evidence before tool setup + run: | + mkdir qualification + git rev-parse HEAD > qualification/head.txt + cp .github/workflows/qualify-release.yml qualification/qualify-release.yml - name: Install tools run: | rustup component add clippy rustfmt sudo apt-get update - sudo apt-get install -y awscli git-lfs openssh-client + sudo apt-get install -y git-lfs openssh-client + aws --version git lfs install - name: Bind exact source and dependency inputs run: | - mkdir qualification cargo metadata --locked --format-version 1 > qualification/metadata.json docker pull rustfs/rustfs:1.0.0-beta.8-glibc python3 - <<'PY' From 3a02d591b5ba651a274ebe0099789096257cc95d Mon Sep 17 00:00:00 2001 From: forhappy Date: Thu, 1 Oct 2026 22:38:18 -0700 Subject: [PATCH 54/60] docs: distinguish latest Cellule checks from frozen campaign evidence --- .github/workflows/qualify-release.yml | 1 + docs/performance-plan.md | 17 +++---- .../2026-10-01-cellule-main-qualification.md | 47 +++++++++++++++---- .../2026-10-01-original-corpus-activation.md | 9 ++-- 4 files changed, 54 insertions(+), 20 deletions(-) diff --git a/.github/workflows/qualify-release.yml b/.github/workflows/qualify-release.yml index 248a07e..3236233 100644 --- a/.github/workflows/qualify-release.yml +++ b/.github/workflows/qualify-release.yml @@ -5,6 +5,7 @@ on: push: branches: - 'codex/cellule-release-ci-*' + - 'codex/three-node-recovery-evidence' permissions: contents: read diff --git a/docs/performance-plan.md b/docs/performance-plan.md index 4a38314..a782cc4 100644 --- a/docs/performance-plan.md +++ b/docs/performance-plan.md @@ -2,14 +2,15 @@ Use this plan to design capacity work and interpret Canopy benchmark results. It separates demonstrated behavior from proposed targets. The current Cellule -dependency pin is `0dc04a658bd99668936f7ec58032d054f6fbc141`, verified against -upstream main on October 1. Its [qualification checkpoint](performance/2026-10-01-cellule-main-qualification.md) -records 239 release Rust tests, 84 Python tests, lints and all eight fresh RustFS -gates, with separately retained artifacts. Subsequent original-corpus remote +dependency pin is `191409685b001a82bd02780def45102b4fc2f164`, rechecked against +upstream main at publication. Its [qualification checkpoint](performance/2026-10-01-cellule-main-qualification.md) +records 244 Linux debug Rust tests, 84 Python tests, lints and all eight fresh +RustFS gates; release verification is running. The frozen `0dc04a6` executable +retains its separate release evidence and original-corpus campaign. Remote verification passed; scheduled load includes failed arrivals, and post-load -owner-loss recovery and reference capacity remain open. Upstream later advanced to -`191409685b001a82bd02780def45102b4fc2f164`; that revision is not yet qualified -by this PR. The preceding `c51dd121` build is historical evidence. +owner-loss recovery and reference capacity remain open. Neither that campaign +nor earlier release evidence qualifies `1914096` performance or recovery. +The preceding `c51dd121` build is historical evidence. Unlike the earlier `a4500add` documentation advance, this changes routing and resource-permit implementation. Locked metadata resolves all six packages; the lockfile changes only their sources. @@ -62,7 +63,7 @@ seed and successful recovery of every recorded ACK. | Current gate | Evidence / status | | --- | --- | -| Qualified Cellule revision | `0dc04a6`: release correctness, lints, Python harness, eight fresh RustFS gates and subsequent original-corpus activation passed. Upstream `1914096` remains unqualified; no measured improvement is claimed | +| Current Cellule pin and historical runtime | PR pins `1914096`: Linux debug tests, lints, Python harness and eight fresh RustFS gates passed; release verification is running. Frozen `0dc04a6`: release qualification and original-corpus activation passed. No recovery or performance result transfers between revisions | | HTTP listener handoff | [Separate startup qualification](performance/2026-10-01-listener-handoff.md): the affected SHA-256 tests retain their bound listeners through supervised startup; 244 release tests, 84 Python tests, eight fresh RustFS gates and repeated cancellation/handoff checks passed. Both Linux CI workflows at `eaebfc4` passed; not measured performance | | Original corpus maintenance recovery | [Complete metadata and fenced-recovery checkpoint](performance/2026-10-01-original-corpus-recovery.md): all 256 shards and 10,003 Controls checked twice before and after; 301 unsettled Cells drained, zero unsettled Cells, catalog and all published roots unchanged | | Original corpus activation and remote verification | [Full-corpus checkpoint](performance/2026-10-01-original-corpus-activation.md): Ready 9 after four allowed writes and eight unchanged activation scans; subsequent 10K identities, 100 LFS bodies, 200 clones and critical-2 verification passed. Independent audit and 7,061-file verified copy closed; post-load owner-loss recovery remains open | diff --git a/docs/performance/2026-10-01-cellule-main-qualification.md b/docs/performance/2026-10-01-cellule-main-qualification.md index 8ca212d..0405029 100644 --- a/docs/performance/2026-10-01-cellule-main-qualification.md +++ b/docs/performance/2026-10-01-cellule-main-qualification.md @@ -1,18 +1,46 @@ # Cellule main revision qualification -Canopy now pins all six Cellule packages to -`0dc04a658bd99668936f7ec58032d054f6fbc141`, verified against upstream main on -October 1. Release correctness, fresh RustFS compatibility and repeated -regressions passed. This does not establish a retained-store upgrade, recovery -of the existing corpus, reference capacity or a performance improvement. +PR #18 now pins all six Cellule packages to +`191409685b001a82bd02780def45102b4fc2f164`, rechecked against upstream main +when publishing this update. Linux debug correctness and fresh RustFS gates +passed; release verification is running. The original-corpus campaign still +uses the frozen, release-qualified `0dc04a6` executable. Results from that +executable do not transfer to the new dependency revision. This follows the [retained catalog admission checkpoint](2026-10-01-retained-catalog-admission.md). That checkpoint and the [authentication and residency results](2026-10-01-bounded-authentication-and-residency.md) retain their original Cellule revision and artifacts. -Upstream later advanced to `191409685b001a82bd02780def45102b4fc2f164`. -The PR still pins the independently qualified `0dc04a6`; this checkpoint does -not qualify the newer runtime or peer HTTP CI changes. +## Current dependency update + +The pin-only candidate is `63c93b0f3fadf12893c9c45832f7367ace04ff24`. +Its [Linux debug CI](https://github.com/crabbuild/canopy/actions/runs/36965371936) +passed 244 top-level Rust tests (zero failed, nine ignored), all 84 Python +tests, formatting, all-target lints, the debug server build and all eight +fresh RustFS compatibility gates. Only five direct pins and six lockfile +source entries changed; other dependency versions and runtime budgets did not. + +The new [release workflow](../../.github/workflows/qualify-release.yml) repeats +release workspace tests, all-target release lints, the Python harness and the +eight provider gates, then retains the Linux executable, source/dependency +provenance and logs. Its first attempt failed during tool setup because the +runner's package index had no `awscli` candidate; no tests ran. The corrected +workflow checks the existing runner CLI and initializes evidence before setup. +The [corrected release run](https://github.com/crabbuild/canopy/actions/runs/36968478780) +has passed setup, source binding, formatting, release lints, workspace tests +and the Python harness. RustFS gates and executable retention were still +running at publication. PR-head CI is a separate check. + +Native release qualification, retained-store recovery, every-ACK verification, +matched performance and reference capacity remain open for `1914096`. +Upstream peer HTTP CI changes are not evidence of Canopy latency equivalence. + +## Historical release qualification for 0dc04a6 + +The following results bind `0dc04a658bd99668936f7ec58032d054f6fbc141`, +not the current PR pin. Release correctness, fresh RustFS compatibility and +repeated regressions passed on that revision. These checks alone did not +establish retained-store recovery, reference capacity or a performance gain. ## Dependency change and source equivalence @@ -55,7 +83,8 @@ clones and stock Git LFS over SSH. The existing 10,000-repository RustFS provide had identical provider snapshots before and after these gates. Neither its corpus nor the UI preview was upgraded. -Reproduce the checks from a checkout with this lockfile: +Reproduce these historical checks from the frozen `bbd784a4` checkout and its +lockfile (the current PR lockfile instead tests `1914096`): ```sh cargo build --release --locked --workspace diff --git a/docs/performance/2026-10-01-original-corpus-activation.md b/docs/performance/2026-10-01-original-corpus-activation.md index e3e83b3..026b079 100644 --- a/docs/performance/2026-10-01-original-corpus-activation.md +++ b/docs/performance/2026-10-01-original-corpus-activation.md @@ -263,9 +263,12 @@ dominant explanations; it does not prove a Directory or RustFS bottleneck. Cellule upstream subsequently advanced by two commits to `191409685b001a82bd02780def45102b4fc2f164`, observed when publishing this checkpoint. Those commits change runtime forwarding/compaction and peer HTTP CI gates. -They are **not** the dependency revision tested here. An isolated candidate -passed locked metadata and formatting checks only; release and RustFS -qualification remain open, and no performance result transfers to it. +They are **not** the dependency revision tested here. PR #18 now pins that +revision: its Linux debug CI passed 244 Rust tests, 84 Python tests and all eight +fresh RustFS gates. Release verification is running; native qualification, +retained-store recovery and performance remain open. See the +[current dependency checkpoint](2026-10-01-cellule-main-qualification.md). +No result from this activation or frozen campaign transfers to the new pin. PR #18 remains draft. The [full campaign](../performance-plan.md) still requires the complete schedule and audit, concurrent faults, fresh-owner verification From 5581d5c61722609111f4fe5550a2fe722c3d2918 Mon Sep 17 00:00:00 2001 From: forhappy Date: Thu, 1 Oct 2026 22:40:00 -0700 Subject: [PATCH 55/60] docs: describe remaining gates independently of PR draft status --- docs/performance/2026-10-01-cellule-main-qualification.md | 2 +- docs/performance/2026-10-01-original-corpus-activation.md | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/performance/2026-10-01-cellule-main-qualification.md b/docs/performance/2026-10-01-cellule-main-qualification.md index 0405029..f68e602 100644 --- a/docs/performance/2026-10-01-cellule-main-qualification.md +++ b/docs/performance/2026-10-01-cellule-main-qualification.md @@ -120,7 +120,7 @@ These are local, not off-machine backups. ## Remaining upgrade and performance gates -PR #18 remains draft. The [subsequent owned-fixture upgrade and recovery](2026-10-01-old-binary-rustfs-upgrade.md) +The [subsequent owned-fixture upgrade and recovery](2026-10-01-old-binary-rustfs-upgrade.md) verified actual old-executable RustFS data, controlled admission, new-executable restore, twenty scheduled critical workflows and their complete fresh-owner recovery. Its earlier clone-byte mismatch was traced to missing client LFS diff --git a/docs/performance/2026-10-01-original-corpus-activation.md b/docs/performance/2026-10-01-original-corpus-activation.md index 026b079..018fea5 100644 --- a/docs/performance/2026-10-01-original-corpus-activation.md +++ b/docs/performance/2026-10-01-original-corpus-activation.md @@ -270,7 +270,7 @@ retained-store recovery and performance remain open. See the [current dependency checkpoint](2026-10-01-cellule-main-qualification.md). No result from this activation or frozen campaign transfers to the new pin. -PR #18 remains draft. The [full campaign](../performance-plan.md) still requires +The [full campaign](../performance-plan.md) still requires the complete schedule and audit, concurrent faults, fresh-owner verification of the original corpus and every newly acknowledged write, higher admission profiles, matched comparisons, large From 0ba777536ac390e58d9d6b12bce3d7826f16792c Mon Sep 17 00:00:00 2001 From: forhappy Date: Thu, 1 Oct 2026 23:43:52 -0700 Subject: [PATCH 56/60] docs: record release failure and audited residency diagnostic --- docs/performance-plan.md | 7 ++- .../2026-10-01-cellule-main-qualification.md | 47 +++++++++++++++++-- .../2026-10-01-original-corpus-activation.md | 7 ++- 3 files changed, 52 insertions(+), 9 deletions(-) diff --git a/docs/performance-plan.md b/docs/performance-plan.md index a782cc4..adc435a 100644 --- a/docs/performance-plan.md +++ b/docs/performance-plan.md @@ -5,7 +5,10 @@ It separates demonstrated behavior from proposed targets. The current Cellule dependency pin is `191409685b001a82bd02780def45102b4fc2f164`, rechecked against upstream main at publication. Its [qualification checkpoint](performance/2026-10-01-cellule-main-qualification.md) records 244 Linux debug Rust tests, 84 Python tests, lints and all eight fresh -RustFS gates; release verification is running. The frozen `0dc04a6` executable +RustFS gates. The release candidate passed, but PR-head release CI failed a +disconnected-admission test. All 100 isolated and two full-target diagnostic +runs passed without changing the source; the original failure remains unresolved. +The frozen `0dc04a6` executable retains its separate release evidence and original-corpus campaign. Remote verification passed; scheduled load includes failed arrivals, and post-load owner-loss recovery and reference capacity remain open. Neither that campaign @@ -63,7 +66,7 @@ seed and successful recovery of every recorded ACK. | Current gate | Evidence / status | | --- | --- | -| Current Cellule pin and historical runtime | PR pins `1914096`: Linux debug tests, lints, Python harness and eight fresh RustFS gates passed; release verification is running. Frozen `0dc04a6`: release qualification and original-corpus activation passed. No recovery or performance result transfers between revisions | +| Current Cellule pin and historical runtime | PR pins `1914096`: debug and candidate release correctness passed. PR-head release CI failed HTTP 503 in the disconnected-admission test; 100 isolated and two full-target diagnostic passes do not resolve it. Frozen `0dc04a6`: release qualification and original-corpus activation passed. No recovery or performance result transfers between revisions | | HTTP listener handoff | [Separate startup qualification](performance/2026-10-01-listener-handoff.md): the affected SHA-256 tests retain their bound listeners through supervised startup; 244 release tests, 84 Python tests, eight fresh RustFS gates and repeated cancellation/handoff checks passed. Both Linux CI workflows at `eaebfc4` passed; not measured performance | | Original corpus maintenance recovery | [Complete metadata and fenced-recovery checkpoint](performance/2026-10-01-original-corpus-recovery.md): all 256 shards and 10,003 Controls checked twice before and after; 301 unsettled Cells drained, zero unsettled Cells, catalog and all published roots unchanged | | Original corpus activation and remote verification | [Full-corpus checkpoint](performance/2026-10-01-original-corpus-activation.md): Ready 9 after four allowed writes and eight unchanged activation scans; subsequent 10K identities, 100 LFS bodies, 200 clones and critical-2 verification passed. Independent audit and 7,061-file verified copy closed; post-load owner-loss recovery remains open | diff --git a/docs/performance/2026-10-01-cellule-main-qualification.md b/docs/performance/2026-10-01-cellule-main-qualification.md index f68e602..c346f53 100644 --- a/docs/performance/2026-10-01-cellule-main-qualification.md +++ b/docs/performance/2026-10-01-cellule-main-qualification.md @@ -2,8 +2,10 @@ PR #18 now pins all six Cellule packages to `191409685b001a82bd02780def45102b4fc2f164`, rechecked against upstream main -when publishing this update. Linux debug correctness and fresh RustFS gates -passed; release verification is running. The original-corpus campaign still +when publishing the dependency update. Linux debug correctness and fresh RustFS +gates passed. The release candidate passed, but PR-head release CI subsequently +failed a disconnected-admission test. An unchanged-source diagnostic did not +reproduce that failure; its cause and resolution remain open. The original-corpus campaign still uses the frozen, release-qualified `0dc04a6` executable. Results from that executable do not transfer to the new dependency revision. @@ -27,9 +29,44 @@ provenance and logs. Its first attempt failed during tool setup because the runner's package index had no `awscli` candidate; no tests ran. The corrected workflow checks the existing runner CLI and initializes evidence before setup. The [corrected release run](https://github.com/crabbuild/canopy/actions/runs/36968478780) -has passed setup, source binding, formatting, release lints, workspace tests -and the Python harness. RustFS gates and executable retention were still -running at publication. PR-head CI is a separate check. +passed 244 top-level Rust tests (zero failed, nine ignored), all 84 Python tests, +the eight release RustFS gates, formatting, release lints and executable retention. +Its archive, 267 source/workflow hashes, locked metadata and executable checksum +were independently audited. PR-head CI is a separate check. + +## PR-head failure and unchanged-source diagnostic + +| Check | Exact source / result | Interpretation | +| --- | --- | --- | +| [PR-head release CI](https://github.com/crabbuild/canopy/actions/runs/36969961214) | `5581d5c`: disconnected-admission test failed HTTP 503; 240 top-level tests passed, one failed, nine ignored before Cargo aborted | Python, provider gates and executable retention were skipped; not a release pass | +| [PR debug CI](https://github.com/crabbuild/canopy/actions/runs/36969965999) and [push debug CI](https://github.com/crabbuild/canopy/actions/runs/36969961213) | `5581d5c`: each passed 244 top-level Rust tests, 84 Python tests and eight fresh RustFS gates | Neither run clears the release failure | +| [Unchanged-source release diagnostic](https://github.com/crabbuild/canopy/actions/runs/36972048946) | `49268ba`: all 100 isolated repetitions passed; both full-target runs passed 104 tests, zero failed, nine ignored | Nonreproduction, not a fix, full-workspace qualification or performance evidence | + +The failing test is +`residency::faults::disconnected_admission_finishes_release_and_allows_a_later_restore`. +The diagnostic's 264 production, test, script and Cargo inputs exactly match +`5581d5c`. It executes the same retained release test binary for every declared +case, including four-thread full-target runs. No test assertion, HTTP retry, +residency budget or production code changed. Failed repetitions would remain +failed; the workflow does not retry until green. Its first setup-only failure +is retained separately. + +The diagnostic closed October 2 at 06:27:34 UTC (October 1 Pacific). An independent +audit verified the GitHub archive digest, exact 111-file inventory, source hashes, +all six Cellule metadata entries, retained ELF checksum and all 102 result/log +bindings. The evidence and auditor were copied and independently reread as +115 files on another local filesystem; this is not an off-machine backup. + +| Diagnostic artifact | SHA-256 | +| --- | --- | +| GitHub artifact archive | `80522d2d0d33d6bbcf3d48c8b40a0583f32953bde75ac39218aac038735ff7e9` | +| Release `multi_server` executable | `b9d2fe7dc845b8c8f0791296208ed2b99c848820a498e6a14ae35821d99fda10` | +| Independent diagnostic audit | `ccd848c30d08b75e0958ff3b973e9416df027d1ebb81cd2fa685eca7f6ce018f` | +| Verified local copy manifest | `0735c73eaaa8ba60bb4bf34f02bbd437e10027d1592c30a597a10bb52a817ebe` | + +The earlier successful release candidate remains historical evidence. The later +failure remains unresolved despite the passing diagnostic. Establish its cause +with a reproducible cancellation/release probe before claiming a correction. Native release qualification, retained-store recovery, every-ACK verification, matched performance and reference capacity remain open for `1914096`. diff --git a/docs/performance/2026-10-01-original-corpus-activation.md b/docs/performance/2026-10-01-original-corpus-activation.md index 018fea5..57db5b8 100644 --- a/docs/performance/2026-10-01-original-corpus-activation.md +++ b/docs/performance/2026-10-01-original-corpus-activation.md @@ -265,8 +265,11 @@ Cellule upstream subsequently advanced by two commits to Those commits change runtime forwarding/compaction and peer HTTP CI gates. They are **not** the dependency revision tested here. PR #18 now pins that revision: its Linux debug CI passed 244 Rust tests, 84 Python tests and all eight -fresh RustFS gates. Release verification is running; native qualification, -retained-store recovery and performance remain open. See the +fresh RustFS gates. Its release candidate passed, but later PR-head release CI +failed the disconnected-admission test with HTTP 503. All 100 isolated and two +full-target diagnostic runs passed on unchanged source; that does not resolve +the failure. Native qualification, retained-store recovery and performance +remain open. See the [current dependency checkpoint](2026-10-01-cellule-main-qualification.md). No result from this activation or frozen campaign transfers to the new pin. From cb78793874b96f958e0a104dce7184b086ef8271 Mon Sep 17 00:00:00 2001 From: forhappy Date: Fri, 2 Oct 2026 07:55:43 -0700 Subject: [PATCH 57/60] perf: retain bounded Git failures and report closed recovery evidence --- docs/performance-plan.md | 19 ++- .../2026-10-01-original-corpus-activation.md | 104 ++++++++++++-- scripts/audit_three_node_campaign.py | 34 +++++ scripts/benchmark_repositories.py | 39 ++++- scripts/test_benchmark_git_failures.py | 133 ++++++++++++++++++ 5 files changed, 310 insertions(+), 19 deletions(-) create mode 100644 scripts/test_benchmark_git_failures.py diff --git a/docs/performance-plan.md b/docs/performance-plan.md index adc435a..e48a455 100644 --- a/docs/performance-plan.md +++ b/docs/performance-plan.md @@ -70,7 +70,8 @@ seed and successful recovery of every recorded ACK. | HTTP listener handoff | [Separate startup qualification](performance/2026-10-01-listener-handoff.md): the affected SHA-256 tests retain their bound listeners through supervised startup; 244 release tests, 84 Python tests, eight fresh RustFS gates and repeated cancellation/handoff checks passed. Both Linux CI workflows at `eaebfc4` passed; not measured performance | | Original corpus maintenance recovery | [Complete metadata and fenced-recovery checkpoint](performance/2026-10-01-original-corpus-recovery.md): all 256 shards and 10,003 Controls checked twice before and after; 301 unsettled Cells drained, zero unsettled Cells, catalog and all published roots unchanged | | Original corpus activation and remote verification | [Full-corpus checkpoint](performance/2026-10-01-original-corpus-activation.md): Ready 9 after four allowed writes and eight unchanged activation scans; subsequent 10K identities, 100 LFS bodies, 200 clones and critical-2 verification passed. Independent audit and 7,061-file verified copy closed; post-load owner-loss recovery remains open | -| Current original-corpus load snapshot | Four completed phases / 40 windows: metadata 41,485/43,200 OK; creation 1,464/2,160; HTTP v2 discovery 10,309/57,600; stock-Git ls-remote 632/1,200. Failed arrivals retained. Concurrent critical schedule: 19/20 OK, 323 successful steps, 38 UUIDs. Sealed ledgers audited and copied; full schedule still running, not reference capacity | +| Complete original-corpus load | All 108 windows / 114,960 arrivals / 8,640 scheduled seconds closed and independently audited: 59,554 OK, 55,406 failed arrivals. Concurrent critical schedule: 19/20 OK, 323 successful steps, 38 UUIDs. Every positive ACK inventoried; not reference capacity | +| Post-load fresh-owner recovery | Exact original owners removed; 32.008191 seconds of confirmed absence, unchanged RustFS. Separate fresh fleet reached readiness after a preserved helper failure. Original-corpus verification then failed a Git v2 clone with exit 128; no complete stage or every-ACK recovery pass. All 3,684 attempt files copied and hash-checked | | Owned old-binary RustFS upgrade | [Controlled fixture qualification](performance/2026-10-01-old-binary-rustfs-upgrade.md) passed actual old-byte restore through three fresh gateways/proxy, both new Git/LFS writes, twenty scheduled 17-step workflows and recovery of all forty critical repositories after all three owners were killed. Two retained repositories; not the original 10K corpus or reference capacity | | Original rebuilt Cellule artifact | `c51dd121` build, 227 Rust tests, eight RustFS gates and initial 17-step proxy check passed; full 10K/100 seed closed, but remote verification failed HTTP 503. A separate 10K metadata diagnostic returned 9,999 matching identities and one 503; downstream fault/load gates stopped without owner signals. Cause unresolved | | Combined authentication and residency candidate | First publication passed 231 release tests, eight RustFS gates, 20 cold-activation repetitions and 15 residency tests. Exact predecessor descriptor is retained; see the separately qualified catalog follow-up | @@ -85,7 +86,7 @@ seed and successful recovery of every recorded ACK. | Post-load ACK recovery | Recorded three-owner loss and 32.249-s wait; critical fixtures passed. Fresh fleet then recorded ENOSPC and shut down; full corpus and every-ACK verification are incomplete | | Concurrent critical Git workflows | [Separate driver](../scripts/benchmark_critical_git.py) retains the 17-step suite. Original-corpus concurrent load closed at 19/20 OK with one busy drop; all 19 attempted receipts complete. Fresh-owner verification of all 38 UUIDs and concurrent fault coverage remain open; the small-fixture 20/20 recovery result is separate | | Evidence availability | Old benchmark directories and executable disappeared; old receipts remain unavailable. New closed artifacts, all build bindings and the completed seed have verified backups on a separate filesystem outside Cargo targets; changing outputs are excluded | -| Performance qualification | Pre-load original-corpus remote verification passed. Complete scheduled load/audit, original-corpus and every-ACK recovery after owner loss, concurrent faults, higher admission profiles and matched comparisons remain open | +| Performance qualification | Pre-load remote verification and complete schedule/ledger audit passed their integrity checks. Arrival gate failed. Original-corpus and every-ACK recovery after owner loss, concurrent faults, higher admission profiles and matched comparisons remain open | The [native-filesystem trial](performance/2026-10-01-native-filesystem.md) records the storage hypothesis, graceful diagnostic drain and new bindings. @@ -859,6 +860,20 @@ acknowledgements. It does not kill a server or establish that an owner restarted record old/new node identities, process exit and fresh local-state evidence alongside it. Run the separate corpus `verify` for seeded identities and data. +New benchmark arrival samples retain optional `git_failure` details: the Git +command stage, nonzero exit or unchanged timeout, and up to **2,048 UTF-8 bytes** +of redacted stderr. Configured tokens, authorization headers, URL userinfo and +common secret query parameters are removed before truncation and hashing. +`redacted_stderr_sha256` and `redacted_stderr_bytes` describe the full redacted +text, not the original stderr. Review excerpts before sharing: this is not a +general-purpose secret detector. Raw command arguments are not added to these +details. Failures remain failed arrivals; no retry or timeout widening is added. + +The closed-ledger auditor validates the optional fields, their outcome/deadline +consistency, excerpt bounds and non-truncated digest. It still accepts historical +samples without these fields; their missing stderr cannot be reconstructed. +For matched comparisons, use the same instrumented driver for both binaries. + Regression fixtures cover SHA-1/SHA-256 pushes, wrong tips, exact-body differences, corrupt LFS, lost acknowledgements, tampered samples and inconsistent accounting. These driver checks do not themselves qualify production recovery diff --git a/docs/performance/2026-10-01-original-corpus-activation.md b/docs/performance/2026-10-01-original-corpus-activation.md index 57db5b8..588c78b 100644 --- a/docs/performance/2026-10-01-original-corpus-activation.md +++ b/docs/performance/2026-10-01-original-corpus-activation.md @@ -6,8 +6,9 @@ or published roots**. The qualified executable uses Cellule `0dc04a6`; its selected release reached **Ready revision 9**. This closes the release-admission gate after [same-code maintenance recovery](2026-10-01-original-corpus-recovery.md). Activation alone does not establish remote Git/LFS content recovery, throughput -or capacity. The subsequent remote verification passed; the completed load -phases below include failed arrivals and do not establish capacity. +or capacity. Pre-load remote verification passed. The complete load campaign +failed its arrival gate, and post-load fresh-owner verification failed during +an original-corpus Git v2 clone. Full recovery remains unproven. ## Activation sequence @@ -27,8 +28,9 @@ Old Maintenance 5: all 10,003 Cells idle and unowned +-- preserve and reread closed evidence +-- three fresh gateways and proxy ready +-- remote Git/LFS verification passed - +-- full load running: failed arrivals retained - +-- fresh-owner/every-ACK recovery still open + +-- full load closed: failed arrivals retained + +-- post-loss original-corpus verification failed + +-- every-ACK recovery not reached ``` The activation controller repeats two complete scans at each of Ready 6, @@ -134,11 +136,11 @@ directory; the verified copy is ## Scheduled load results -The unchanged **108-window / 114,960-arrival / 8,640-second** campaign remains -running on the qualified `0dc04a6` executable and original RustFS provider. -At October 2, 04:33 UTC (October 1 Pacific), its first four operation phases -were complete: **40 sealed, audited and preserved windows**. Later clone -windows had also begun; they are excluded from this completed-phase summary. +The unchanged **108-window / 114,960-arrival / 8,640-scheduled-second** campaign +closed October 2 at **07:04:46 UTC** on the qualified `0dc04a6` executable and +original RustFS provider. All 108 arrival ledgers and resource boundaries were +independently audited. It recorded **59,554 OK and 55,406 failed arrivals**; +completing the schedule is not a capacity pass. | Completed phase | Windows | Scheduled | OK | Busy drops | HTTP 503 | Other failures | | --- | --- | --- | --- | --- | --- | --- | @@ -146,12 +148,30 @@ windows had also begun; they are excluded from this completed-phase summary. | Repository creation | 6 | 2,160 | 1,464 | 690 | 6 | 0 | | HTTP Git v2 discovery | 8 | 57,600 | 10,309 | 9,581 | 37,706 | 4 transport errors | | Stock Git ls-remote | 8 | 1,200 | 632 | 514 | Not separately classified | 54 Git errors | +| Clone | 8 | 1,200 | 704 | 495 | Not separately classified | 1 Git error | +| Fresh-client fetch | 8 | 1,200 | 814 | 386 | Not separately classified | 0 | +| Incremental fetch | 8 | 1,200 | 771 | 418 | Not separately classified | 11 Git errors | +| Incremental pull | 8 | 1,200 | 743 | 457 | Not separately classified | 0 | +| Ref-only push | 4 | 1,200 | 1,023 | 85 | Not separately classified | 92 Git errors | +| Fresh-object push, 256 KiB and 1 MiB | 16 | 2,400 | 467 | 1,591 | Not separately classified | 333 Git errors, 9 timeouts | +| LFS download | 8 | 1,200 | 900 | 297 | 3 | 0 | +| LFS upload | 8 | 1,200 | 242 | 626 | 157 | 174 transport errors, 1 HTTP 500 | These are failed arrival gates. The HTTP discovery probe checks the v2 capability response, not the full stock-Git ref exchange; `ls-remote` runs the actual Git client and checks the expected main ref. Busy arrivals do not start requests and have no completed-request latency. Git errors remain client-level failures rather than being silently classified as HTTP 503s. +Fresh-client fetch does not guarantee a cold server cache. The frozen driver +did not retain Git stderr excerpts, so those errors cannot be assigned a cause +from the arrival ledger alone. Failed writes are not assumed to have rolled back. + +The eight 1-MiB push windows recorded **150/1,200 OK**, 855 busy drops, 186 Git +errors and nine timeouts. All positive ACKs and **157,286,400 declared payload +bytes** were reconciled; these are not wire bytes. The two 16-client, 4-push/s +windows delivered **0.117 / 0.050 successful in-window pushes/s**, with +successful-only scheduled p95 latencies of **70.764 / 100.998 seconds**. +Completion counts include drain; in-window rates exclude it. Fast refusal can make aggregate latency misleading. In the first uniform 100-RPS discovery window, completed attempts had p50 **7.824 ms**, but successful @@ -193,8 +213,8 @@ root cause is established. ### Earlier ten-window metadata snapshot -The unchanged **108-window / 114,960-arrival / 8,640-second** campaign is running -on the same qualified executable and original RustFS provider. This earlier snapshot +The now-closed campaign used the same qualified executable and original +RustFS provider. This earlier snapshot covers only its first **ten sealed, audited and preserved metadata windows**, observed October 2 at 03:36 UTC. All offer 20 requests/s for 120 seconds, with client concurrency 32 and node residency capped at 100. An active set of 500 @@ -258,6 +278,64 @@ deltas were balanced (85/85/84 and 106/106/105), with no proxy errors or rejections. This weakens timer scheduling and connection-count imbalance as dominant explanations; it does not prove a Directory or RustFS bottleneck. +## Post-load owner loss and failed recovery + +The closed campaign's positive ACK inventory contains **1,464 creations, +1,023 ref-only pushes, 467 fresh-object pushes, 242 LFS uploads and 19 critical +workflows**. The original and acknowledged namespace has 11,504 distinct +repository identities. Inventory integrity does not prove remote recovery. + +```text +108 closed windows + every positive ACK inventoried + -> exact three original gateway owners removed + -> 32.008191 seconds of confirmed owner absence + -> first fresh launch: helper metrics/readiness race; normally drained + -> separate fresh launch: three new owners and proxy ready + -> original-corpus verification: Git v2 clone exited 128 + -> remaining corpus, critical fixtures and every-ACK stages not completed +``` + +RustFS identity, start time, configuration and resource envelope were unchanged +across owner loss. The first fresh launch's helper failure remains preserved. +The separate launch waits for metrics from the same live child within the +original startup deadline; nine guard tests passed. It neither respawns a child +nor widens the deadline or runtime budgets. + +The actual read-only recovery attempt terminated at **08:13:33 UTC on October 2** +during the original-corpus stage. Its 3,998 request rows contain 3,632 identity +checks, 41 LFS downloads and 325 Git commands. One Git command failed: + +```sh +git -c protocol.version=2 clone \ + "$PROXY_URL/canopy/density-3a92b05e1d80-03629.git" \ + "$NEW_WORK_DIR/density-3a92b05e1d80-03629-v2" +``` + +It exited **128 after 0.570 seconds**, not a client timeout. The frozen verifier +retained stderr's SHA-256 but not its text; the root cause is unresolved. No +complete stage was recorded, and the every-ACK verification stages were not +reached. Partial successful requests do not establish full-corpus recovery. +Any later diagnostic success must remain separate from this failed attempt. + +All **3,684 closed attempt files**, including the request ledger and all remaining +cloned data, were copied and independently hash-checked on a +different local filesystem. All **651 input bindings** were checked before and +after copying. This is evidence preservation, not a provider or off-machine backup. + +| Artifact | SHA-256 | +| --- | --- | +| Full campaign audit | `23836f6a437826a58e5157b534a17346b939e9c4e7213168b1c872c19f387f59` | +| Complete ACK inventory | `9853c5f32551980165d62e06281e163a26d435fde070984883fe8c388758a0de` | +| Owner-loss receipt | `40cb5aa9068b8f32cffb51433682af6c8fa4ce36b4a2a856543c753018f0addb` | +| Failed recovery receipt | `54c3e058871c7c57ef65f9b8f32f3f94cdd93151915d757af49cfcce333a3b54` | +| Failed recovery request ledger | `a331fe7b71dcd546abee9fc46b623dcafaec0b4374e5de64941a9b8d0afa0348` | +| Failed recovery preservation manifest | `c9b314798f1071dcb625f95e9cd4da8b9d64618051d8f7d61b100dfe85a601fc` | + +The failed attempt is at +`/Volumes/Workspace/CrabData/canopy-full-post-owner-loss-5zhLBG`; +its verified copy and manifest are at +`/Users/haipingfu/.codex/canopy-full-recovery-failed-y871llh9`. + ## Remaining verification Cellule upstream subsequently advanced by two commits to @@ -273,8 +351,8 @@ remain open. See the [current dependency checkpoint](2026-10-01-cellule-main-qualification.md). No result from this activation or frozen campaign transfers to the new pin. -The [full campaign](../performance-plan.md) still requires -the complete schedule and audit, concurrent faults, fresh-owner verification +The [performance plan](../performance-plan.md) still requires +concurrent faults, fresh-owner verification of the original corpus and every newly acknowledged write, higher admission profiles, matched comparisons, large transfers and isolated Linux capacity. The earlier diagnostic lease-fencing diff --git a/scripts/audit_three_node_campaign.py b/scripts/audit_three_node_campaign.py index a6e9d30..1ed6ec1 100644 --- a/scripts/audit_three_node_campaign.py +++ b/scripts/audit_three_node_campaign.py @@ -27,6 +27,39 @@ def finite(value): return isinstance(value, (int, float)) and not isinstance(value, bool) and math.isfinite(value) and value >= 0 +def audit_git_failure(row, report): + """Validate optional evidence without inventing details for older ledgers.""" + details = row.get("git_failure") + if details is None: + return + keys = {"kind", "command", "exit_code", "timeout_seconds", "stderr_excerpt", + "redacted_stderr_sha256", "redacted_stderr_bytes", "stderr_truncated"} + require(isinstance(details, dict) and set(details) == keys, "Git failure fields differ") + commands = {"init", "clone", "fetch", "pull", "push", "ls-remote", "rev-parse", "fsck", + "show", "notes", "for-each-ref", "hash-object", "cat-file", "config", "add", "commit", "unknown"} + require(isinstance(details["command"], str) and details["command"] in commands, "invalid Git command stage") + if details["kind"] == "exit": + require(row["result"] == "git_error" and type(details["exit_code"]) is int + and details["exit_code"] != 0 and details["timeout_seconds"] is None, "invalid Git exit evidence") + else: + require(details["kind"] == "timeout" and row["result"] == "client_timeout" + and details["exit_code"] is None and finite(details["timeout_seconds"]) + and details["timeout_seconds"] > 0 + and details["timeout_seconds"] == report["git_timeout_seconds"], "invalid Git timeout evidence") + require(isinstance(details["stderr_excerpt"], str), "invalid Git stderr excerpt") + excerpt = details["stderr_excerpt"].encode("utf-8") + require(len(excerpt) <= 2048 and type(details["redacted_stderr_bytes"]) is int + and details["redacted_stderr_bytes"] >= len(excerpt) + and type(details["stderr_truncated"]) is bool + and isinstance(details["redacted_stderr_sha256"], str) + and re.fullmatch(r"[0-9a-f]{64}", details["redacted_stderr_sha256"]), "invalid Git stderr bounds/digest") + if details["stderr_truncated"]: + require(details["redacted_stderr_bytes"] > len(excerpt), "Git truncation flag differs") + else: + require(details["redacted_stderr_bytes"] == len(excerpt) + and hashlib.sha256(excerpt).hexdigest() == details["redacted_stderr_sha256"], "Git stderr digest differs") + + def quantiles(values): ordered = sorted(values) return {name: None if not ordered else round(ordered[math.ceil(q * len(ordered)) - 1], 3) @@ -89,6 +122,7 @@ def audit_window(path, window, manifest_digest, driver_digest, repository_ids, e seen.add(sequence) require(type(row["ingress_index"]) is int and row["ingress_index"] == 0 and isinstance(row["result"], str), "unexpected ingress/outcome") + audit_git_failure(row, report) if expected_selection is not None: require(row["repository_id"] == expected_selection[sequence], "deterministic active-set/distribution differs") if window["operation"] != "create": diff --git a/scripts/benchmark_repositories.py b/scripts/benchmark_repositories.py index 53f0149..0485389 100644 --- a/scripts/benchmark_repositories.py +++ b/scripts/benchmark_repositories.py @@ -132,10 +132,37 @@ def git_result(*args, cwd, token, timeout=120, request_id=None): capture_output=True, timeout=timeout, check=False) +def git_failure_details(args, token, stderr, *, kind, exit_code=None, timeout=None): + """Bound redacted stderr evidence; never retain command arguments or credentials.""" + values = list(args) + while len(values) >= 2 and values[0] == "-c": + values = values[2:] + commands = {"init", "clone", "fetch", "pull", "push", "ls-remote", "rev-parse", "fsck", + "show", "notes", "for-each-ref", "hash-object", "cat-file", "config", "add", "commit"} + command = values[0] if values and values[0] in commands else "unknown" + text = (stderr or b"").decode("utf-8", errors="replace") + if token: + text = text.replace(token, "[redacted]") + text = re.sub(r"(?im)\b(?:proxy-)?authorization\s*:[^\r\n]*", "[redacted authorization]", text) + text = re.sub(r"(?i)((?:https?|ssh)://)[^\s/@]+@", r"\1[redacted]@", text) + text = re.sub(r"(?i)([?&](?:access_token|token|password|secret|signature)=)[^&#\s]+", r"\1[redacted]", text) + redacted = text.encode("utf-8") + excerpt = redacted[:2048].decode("utf-8", errors="ignore") + return {"kind": kind, "command": command, "exit_code": exit_code, "timeout_seconds": timeout, + "stderr_excerpt": excerpt, "redacted_stderr_sha256": hashlib.sha256(redacted).hexdigest(), + "redacted_stderr_bytes": len(redacted), "stderr_truncated": len(redacted) > len(excerpt.encode("utf-8"))} + + def git(*args, cwd, token, timeout=120, request_id=None): - result = git_result(*args, cwd=cwd, token=token, timeout=timeout, request_id=request_id) + try: + result = git_result(*args, cwd=cwd, token=token, timeout=timeout, request_id=request_id) + except subprocess.TimeoutExpired as error: + error.git_failure = git_failure_details(args, token, error.stderr, kind="timeout", timeout=error.timeout) + raise if result.returncode: - raise RuntimeError(f"Git {args[0]} failed (exit {result.returncode})") + error = RuntimeError(f"Git {args[0]} failed (exit {result.returncode})") + error.git_failure = git_failure_details(args, token, result.stderr, kind="exit", exit_code=result.returncode) + raise error return result.stdout.strip().decode() @@ -828,6 +855,7 @@ def execute(sequence, entry, scheduled): uploaded_oid = None created_id = None push_receipt = {} + git_failure = None created_name = f"create-{create_run_id}-{sequence:07d}" if creation else None try: if creation: @@ -882,10 +910,12 @@ def execute(sequence, entry, scheduled): raise ValueError("unsupported benchmark operation") result = ("ok" if valid else (f"http_{status}" if not git_operation and status != 200 else "invalid_response")) - except subprocess.TimeoutExpired: + except subprocess.TimeoutExpired as error: result = "client_timeout" - except RuntimeError: + git_failure = getattr(error, "git_failure", None) + except RuntimeError as error: result = "git_error" + git_failure = getattr(error, "git_failure", None) except (OSError, ValueError, KeyError, http.client.HTTPException): pass finally: @@ -896,6 +926,7 @@ def execute(sequence, entry, scheduled): "push_commit": push_receipt.get("push_commit"), "git_push_command_ms": push_receipt.get("git_push_command_ms"), "git_client_preparation_ms": push_receipt.get("git_client_preparation_ms"), + "git_failure": git_failure, "completion_offset_seconds": finished - started, "elapsed_ms": (finished - scheduled) * 1000, "service_ms": (finished - dispatched) * 1000, diff --git a/scripts/test_benchmark_git_failures.py b/scripts/test_benchmark_git_failures.py new file mode 100644 index 0000000..148b815 --- /dev/null +++ b/scripts/test_benchmark_git_failures.py @@ -0,0 +1,133 @@ +"""Bounded Git failure evidence; local fixtures do not measure server capacity.""" +import hashlib +from copy import deepcopy +import json +from pathlib import Path +import subprocess +import tempfile +from types import SimpleNamespace +import unittest +from unittest.mock import patch +import uuid + +import benchmark_repositories as benchmark +import audit_three_node_campaign as auditor + + +class GitFailureEvidenceTests(unittest.TestCase): + def test_nonzero_exit_retains_redacted_stage_and_stderr(self): + token = 'fixture-private-token' + stderr = (f'Authorization: Bearer {token}\n' + 'Proxy-Authorization: Basic Y3JlZGVudGlhbHM=\n' + 'fatal: https://user:password@example.invalid/remote denied\n' + 'https://example.invalid/?TOKEN=query-private&secret=other-private\n' + f'{token}\n').encode() + result = subprocess.CompletedProcess(['git'], 128, stdout=b'', stderr=stderr) + with patch.object(benchmark, 'git_result', return_value=result), self.assertRaises(RuntimeError) as raised: + benchmark.git('-c', 'protocol.version=2', 'fetch', 'remote', cwd=Path('.'), token=token) + details = getattr(raised.exception, 'git_failure', None) + self.assertIsInstance(details, dict) + self.assertEqual((details['kind'], details['command'], details['exit_code']), ('exit', 'fetch', 128)) + encoded = json.dumps(details) + for secret in [token, 'user:password', 'Y3JlZGVudGlhbHM=', 'query-private', 'other-private']: + self.assertNotIn(secret, encoded) + self.assertIn('denied', details['stderr_excerpt']) + + def test_excerpt_bound_and_secret_crossing_truncation_boundary(self): + token = 'fixture-private-token' + result = subprocess.CompletedProcess(['git'], 1, stdout=b'', stderr=b'x' * 2040 + token.encode() + b'\x01' * 8000) + with patch.object(benchmark, 'git_result', return_value=result), self.assertRaises(RuntimeError) as raised: + benchmark.git('push', 'remote', cwd=Path('.'), token=token) + details = getattr(raised.exception, 'git_failure', None) + self.assertIsInstance(details, dict) + self.assertLessEqual(len(details['stderr_excerpt'].encode()), 2048) + self.assertTrue(details['stderr_truncated']) + self.assertNotIn(token[:8], details['stderr_excerpt']) + self.assertLess(len(json.dumps(details)), 13 * 1024) + auditor.audit_git_failure({'result': 'git_error', 'git_failure': details}, {'git_timeout_seconds': 120}) + + def test_invalid_utf8_and_empty_stderr_are_safe(self): + for stderr in [b'', b'\xff\xfe fatal: denied']: + with self.subTest(stderr=stderr), patch.object(benchmark, 'git_result', return_value= + subprocess.CompletedProcess(['git'], 1, stdout=b'', stderr=stderr)), self.assertRaises(RuntimeError) as raised: + benchmark.git('clone', 'remote', cwd=Path('.'), token='fixture-token') + details = getattr(raised.exception, 'git_failure', None) + self.assertIsInstance(details, dict) + if not stderr: + self.assertEqual(details['redacted_stderr_sha256'], hashlib.sha256(b'').hexdigest()) + self.assertLessEqual(len(details['stderr_excerpt'].encode()), 2048) + + def test_timeout_remains_timeout_with_same_deadline(self): + error = subprocess.TimeoutExpired(['git', 'fetch'], 120, stderr=b'fatal: fixture timeout') + with patch.object(benchmark, 'git_result', side_effect=error), self.assertRaises(subprocess.TimeoutExpired) as raised: + benchmark.git('fetch', 'remote', cwd=Path('.'), token='fixture-token', timeout=120) + self.assertIs(raised.exception, error) + details = getattr(error, 'git_failure', None) + self.assertIsInstance(details, dict) + self.assertEqual((details['kind'], details['command'], details['timeout_seconds']), ('timeout', 'fetch', 120)) + self.assertIsNone(details['exit_code']) + + def fixture(self, root): + manifest = root / 'corpus.json' + benchmark.save(manifest, {'version': 1, 'complete': True, 'requested_repositories': 1, + 'repositories': [{'name': 'absent', 'owner': 'canopy', 'repository_id': str(uuid.uuid4()), + 'commit': 'a' * 40, 'readme_sha256': 'b' * 64}]}) + args = SimpleNamespace(operation='ls_remote', manifest=manifest, seed=20260926, + active_repositories=1, distribution='uniform', rate=1, duration=1, concurrency=1, + timeout=30, git_timeout=120, work_dir=root / 'clients', output=root / 'report.json') + return args, SimpleNamespace(base_url=(root / 'nonexistent-backend').as_uri()) + + def test_real_local_git_failure_remains_failed_and_retains_evidence(self): + with tempfile.TemporaryDirectory() as directory: + root = Path(directory); args, client = self.fixture(root) + report = benchmark.measure(args, client, 'fixture-token') + self.assertEqual(report['outcomes'], {'git_error': 1}) + self.assertEqual(report['failed_arrivals'], 1) + sample = json.loads(args.output.with_suffix('.samples.jsonl').read_text()) + self.assertEqual(sample['git_failure']['command'], 'ls-remote') + self.assertNotEqual(sample['git_failure']['exit_code'], 0) + self.assertIn('fatal:', sample['git_failure']['stderr_excerpt']) + self.assertEqual(benchmark.file_sha256(args.output.with_suffix('.samples.jsonl')), report['samples_sha256']) + + def test_measured_timeout_is_not_relabelled_as_git_error(self): + with tempfile.TemporaryDirectory() as directory: + root = Path(directory); args, client = self.fixture(root) + with patch.object(benchmark, 'git_result', side_effect=subprocess.TimeoutExpired(['git'], 120, stderr=b'fixture-timeout')): + report = benchmark.measure(args, client, 'fixture-token') + self.assertEqual(report['outcomes'], {'client_timeout': 1}) + sample = json.loads(args.output.with_suffix('.samples.jsonl').read_text()) + self.assertEqual(sample['git_failure']['timeout_seconds'], 120) + self.assertEqual(sample['result'], 'client_timeout') + auditor.audit_git_failure(sample, report) + + def test_independent_auditor_rejects_corrupted_failure_details(self): + with tempfile.TemporaryDirectory() as directory: + root = Path(directory); args, client = self.fixture(root) + report = benchmark.measure(args, client, 'fixture-token') + path = args.output.with_suffix('.samples.jsonl') + sample = json.loads(path.read_text()) + window = {'operation': 'ls_remote', 'distribution': 'uniform', 'active_repositories': 1, + 'rate': 1, 'duration': 1, 'concurrency': 1} + def audit(): + return auditor.audit_window(args.output, window, benchmark.file_sha256(args.manifest), + benchmark.file_sha256(Path(benchmark.__file__)), {sample['repository_id']}, {0: sample['repository_id']}) + audit() + for key, value in [('exit_code', 0), ('command', 'secret-command'), + ('stderr_excerpt', 'x' * 2049), ('redacted_stderr_sha256', 'bad'), + ('arguments', ['secret']), ('timeout_seconds', 120), + ('exit_code', True), ('redacted_stderr_bytes', False), + ('stderr_truncated', 'yes'), ('kind', 'unknown'), + ('redacted_stderr_sha256', 'a' * 64)]: + changed = deepcopy(sample); changed['git_failure'][key] = value + path.write_text(json.dumps(changed) + '\n') + report['samples_sha256'] = benchmark.file_sha256(path); benchmark.save(args.output, report) + with self.subTest(key=key), self.assertRaises(ValueError): + audit() + legacy = deepcopy(sample); del legacy['git_failure'] + path.write_text(json.dumps(legacy) + '\n') + report['samples_sha256'] = benchmark.file_sha256(path); benchmark.save(args.output, report) + audit() + + +if __name__ == '__main__': + unittest.main() From 7497fc720e68a440e27614a83a1442576511cfd2 Mon Sep 17 00:00:00 2001 From: forhappy Date: Fri, 2 Oct 2026 08:51:07 -0700 Subject: [PATCH 58/60] deps: update Cellule main pin and qualification status --- Cargo.lock | 12 ++--- crates/canopy-server/Cargo.toml | 10 ++-- .../2026-10-01-cellule-main-qualification.md | 51 ++++++++++++++----- 3 files changed, 49 insertions(+), 24 deletions(-) diff --git a/Cargo.lock b/Cargo.lock index 9c72ff9..9e435fa 100644 --- a/Cargo.lock +++ b/Cargo.lock @@ -448,7 +448,7 @@ dependencies = [ [[package]] name = "cellule-app" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=191409685b001a82bd02780def45102b4fc2f164#191409685b001a82bd02780def45102b4fc2f164" +source = "git+https://github.com/crabbuild/cellule.git?rev=0f4ca0919b0dfe20a3dcd964d21da03135e42eed#0f4ca0919b0dfe20a3dcd964d21da03135e42eed" dependencies = [ "blake3", "cellule-runtime", @@ -457,7 +457,7 @@ dependencies = [ [[package]] name = "cellule-host" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=191409685b001a82bd02780def45102b4fc2f164#191409685b001a82bd02780def45102b4fc2f164" +source = "git+https://github.com/crabbuild/cellule.git?rev=0f4ca0919b0dfe20a3dcd964d21da03135e42eed#0f4ca0919b0dfe20a3dcd964d21da03135e42eed" dependencies = [ "cellule-app", "cellule-runtime", @@ -471,7 +471,7 @@ dependencies = [ [[package]] name = "cellule-ltx" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=191409685b001a82bd02780def45102b4fc2f164#191409685b001a82bd02780def45102b4fc2f164" +source = "git+https://github.com/crabbuild/cellule.git?rev=0f4ca0919b0dfe20a3dcd964d21da03135e42eed#0f4ca0919b0dfe20a3dcd964d21da03135e42eed" dependencies = [ "async-trait", "blake3", @@ -493,7 +493,7 @@ dependencies = [ [[package]] name = "cellule-runtime" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=191409685b001a82bd02780def45102b4fc2f164#191409685b001a82bd02780def45102b4fc2f164" +source = "git+https://github.com/crabbuild/cellule.git?rev=0f4ca0919b0dfe20a3dcd964d21da03135e42eed#0f4ca0919b0dfe20a3dcd964d21da03135e42eed" dependencies = [ "blake3", "bytes", @@ -519,7 +519,7 @@ dependencies = [ [[package]] name = "cellule-store" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=191409685b001a82bd02780def45102b4fc2f164#191409685b001a82bd02780def45102b4fc2f164" +source = "git+https://github.com/crabbuild/cellule.git?rev=0f4ca0919b0dfe20a3dcd964d21da03135e42eed#0f4ca0919b0dfe20a3dcd964d21da03135e42eed" dependencies = [ "async-trait", "blake3", @@ -542,7 +542,7 @@ dependencies = [ [[package]] name = "cellule-types" version = "0.1.0" -source = "git+https://github.com/crabbuild/cellule.git?rev=191409685b001a82bd02780def45102b4fc2f164#191409685b001a82bd02780def45102b4fc2f164" +source = "git+https://github.com/crabbuild/cellule.git?rev=0f4ca0919b0dfe20a3dcd964d21da03135e42eed#0f4ca0919b0dfe20a3dcd964d21da03135e42eed" dependencies = [ "schemars", "serde", diff --git a/crates/canopy-server/Cargo.toml b/crates/canopy-server/Cargo.toml index 40e9086..0741f9a 100644 --- a/crates/canopy-server/Cargo.toml +++ b/crates/canopy-server/Cargo.toml @@ -15,11 +15,11 @@ axum = "0.8.9" base64 = "0.22" blake3 = "1.8" bytes = "1.11" -cellule-app = { git = "https://github.com/crabbuild/cellule.git", rev = "191409685b001a82bd02780def45102b4fc2f164" } -cellule-host = { git = "https://github.com/crabbuild/cellule.git", rev = "191409685b001a82bd02780def45102b4fc2f164" } -cellule-ltx = { git = "https://github.com/crabbuild/cellule.git", rev = "191409685b001a82bd02780def45102b4fc2f164", features = ["replica"] } -cellule-runtime = { git = "https://github.com/crabbuild/cellule.git", rev = "191409685b001a82bd02780def45102b4fc2f164" } -cellule-store = { git = "https://github.com/crabbuild/cellule.git", rev = "191409685b001a82bd02780def45102b4fc2f164" } +cellule-app = { git = "https://github.com/crabbuild/cellule.git", rev = "0f4ca0919b0dfe20a3dcd964d21da03135e42eed" } +cellule-host = { git = "https://github.com/crabbuild/cellule.git", rev = "0f4ca0919b0dfe20a3dcd964d21da03135e42eed" } +cellule-ltx = { git = "https://github.com/crabbuild/cellule.git", rev = "0f4ca0919b0dfe20a3dcd964d21da03135e42eed", features = ["replica"] } +cellule-runtime = { git = "https://github.com/crabbuild/cellule.git", rev = "0f4ca0919b0dfe20a3dcd964d21da03135e42eed" } +cellule-store = { git = "https://github.com/crabbuild/cellule.git", rev = "0f4ca0919b0dfe20a3dcd964d21da03135e42eed" } ed25519-dalek = "2" flate2 = "1.1" futures-core = "0.3" diff --git a/docs/performance/2026-10-01-cellule-main-qualification.md b/docs/performance/2026-10-01-cellule-main-qualification.md index c346f53..166a1f5 100644 --- a/docs/performance/2026-10-01-cellule-main-qualification.md +++ b/docs/performance/2026-10-01-cellule-main-qualification.md @@ -1,13 +1,15 @@ # Cellule main revision qualification -PR #18 now pins all six Cellule packages to -`191409685b001a82bd02780def45102b4fc2f164`, rechecked against upstream main -when publishing the dependency update. Linux debug correctness and fresh RustFS -gates passed. The release candidate passed, but PR-head release CI subsequently -failed a disconnected-admission test. An unchanged-source diagnostic did not -reproduce that failure; its cause and resolution remain open. The original-corpus campaign still -uses the frozen, release-qualified `0dc04a6` executable. Results from that -executable do not transfer to the new dependency revision. +PR #18 pins all six Cellule packages to +`0f4ca0919b0dfe20a3dcd964d21da03135e42eed`, checked against upstream main on +October 2. Locked metadata resolves the new revision; its debug, release and +fresh RustFS checks are pending. Passing checks for the previous `1914096` +pin do not qualify this update. + +The original-corpus campaign still uses the frozen, release-qualified +`0dc04a6` executable. Its results do not transfer to either newer revision. +An earlier disconnected-admission failure also remains unexplained despite +subsequent passing checks. This follows the [retained catalog admission checkpoint](2026-10-01-retained-catalog-admission.md). That checkpoint and the [authentication and residency results](2026-10-01-bounded-authentication-and-residency.md) @@ -15,6 +17,19 @@ retain their original Cellule revision and artifacts. ## Current dependency update +Only five direct manifest pins and six lockfile source revisions changed from +`1914096` to `0f4ca09`. Other dependency versions, features, Canopy implementation +and runtime budgets are unchanged. Local `cargo metadata --locked` passed; +no native compilation or running-fleet upgrade was performed for publication. + +The upstream delta adds admitted-owner fences for application handlers and +node-log rotation after member expiry. These changes are not evidence of a +Canopy performance improvement. The existing release workflow will check the +new pin against an isolated fresh RustFS fixture; retained-store upgrade, +full-corpus recovery and matched performance require separate verification. + +## Previous pin: 1914096 + The pin-only candidate is `63c93b0f3fadf12893c9c45832f7367ace04ff24`. Its [Linux debug CI](https://github.com/crabbuild/canopy/actions/runs/36965371936) passed 244 top-level Rust tests (zero failed, nine ignored), all 84 Python @@ -34,6 +49,14 @@ the eight release RustFS gates, formatting, release lints and executable retenti Its archive, 267 source/workflow hashes, locked metadata and executable checksum were independently audited. PR-head CI is a separate check. +The later [PR-head release run at `cb78793`](https://github.com/crabbuild/canopy/actions/runs/37023329269) +passed 244 top-level Rust tests (zero failed, nine ignored), 91 Python tests, +all eight fresh RustFS gates, release lints and executable retention. Both +PR and push debug verification also passed. The release archive, 268 exact +source/workflow inputs, locked metadata and retained executable were independently +audited. These results apply to `1914096`, not the new `0f4ca09` pin, and do +not establish the cause or resolution of the earlier failure below. + ## PR-head failure and unchanged-source diagnostic | Check | Exact source / result | Interpretation | @@ -121,7 +144,7 @@ had identical provider snapshots before and after these gates. Neither its corpus nor the UI preview was upgraded. Reproduce these historical checks from the frozen `bbd784a4` checkout and its -lockfile (the current PR lockfile instead tests `1914096`): +lockfile (the current PR lockfile instead tests `0f4ca09`): ```sh cargo build --release --locked --workspace @@ -170,10 +193,12 @@ full remote Git/LFS verification passed with an independent audit and verified critical schedule include failed arrivals; see that checkpoint's result tables for metadata, creation, HTTP v2 discovery and stock-Git ls-remote. -The [full performance plan](../performance-plan.md) still requires three nodes -behind a proxy, 10,000 identities, 100 populated Git/LFS fixtures, all 108 windows, -114,960 arrivals and 8,640 scheduled seconds. Complete scheduled load/audit, -concurrent faults, original-corpus fresh-owner recovery, +The [full performance plan](../performance-plan.md) uses three nodes behind a +proxy, 10,000 identities and 100 populated Git/LFS fixtures. All 108 windows, +114,960 arrivals and 8,640 scheduled seconds completed and were audited on the +frozen `0dc04a6` runtime, with failed arrivals retained. The subsequent full-corpus +recovery attempt failed; see the [complete campaign and recovery evidence](2026-10-01-original-corpus-activation.md). +Concurrent faults, complete original-corpus fresh-owner recovery, every acknowledged write after owner loss, higher admission profiles, matched comparisons, non-sparse five-GiB transfers and isolated Linux capacity remain open. The separate diagnostic fleet's terminal lease-fencing failure is also unresolved. From e8eef7702f09f1fb582ad2517ab3996e367ff71b Mon Sep 17 00:00:00 2001 From: forhappy Date: Fri, 2 Oct 2026 09:46:18 -0700 Subject: [PATCH 59/60] docs: refresh Cellule qualification and recovery checkpoint --- docs/performance-plan.md | 31 ++++++++-------- .../2026-10-01-cellule-main-qualification.md | 37 ++++++++++++++++--- 2 files changed, 46 insertions(+), 22 deletions(-) diff --git a/docs/performance-plan.md b/docs/performance-plan.md index e48a455..76f8e1a 100644 --- a/docs/performance-plan.md +++ b/docs/performance-plan.md @@ -1,18 +1,17 @@ # Measure repository density and latency -Use this plan to design capacity work and interpret Canopy benchmark results. -It separates demonstrated behavior from proposed targets. The current Cellule -dependency pin is `191409685b001a82bd02780def45102b4fc2f164`, rechecked against -upstream main at publication. Its [qualification checkpoint](performance/2026-10-01-cellule-main-qualification.md) -records 244 Linux debug Rust tests, 84 Python tests, lints and all eight fresh -RustFS gates. The release candidate passed, but PR-head release CI failed a -disconnected-admission test. All 100 isolated and two full-target diagnostic -runs passed without changing the source; the original failure remains unresolved. -The frozen `0dc04a6` executable -retains its separate release evidence and original-corpus campaign. Remote -verification passed; scheduled load includes failed arrivals, and post-load -owner-loss recovery and reference capacity remain open. Neither that campaign -nor earlier release evidence qualifies `1914096` performance or recovery. +Use this plan to design capacity work and interpret Canopy benchmark results. Fresh-fixture correctness passes, but the complete original-corpus load failed its arrival gates and post-load recovery remains incomplete. No reference capacity or matched performance improvement is established. + +## Current verification status + +The current PR pins Cellule to `0f4ca0919b0dfe20a3dcd964d21da03135e42eed`. Its [qualification checkpoint](performance/2026-10-01-cellule-main-qualification.md) records five passing Linux CI checks at `7497fc7`: 244 top-level release Rust tests, 91 Python tests, release lints and all eight fresh RustFS gates. An independent audit binds the executable, dependencies and 268 source/workflow inputs. These results do not qualify native retained-store recovery or performance. + +The original-corpus campaign uses the separate, frozen `0dc04a6` executable. Its [full-corpus checkpoint](performance/2026-10-01-original-corpus-activation.md) records pre-load remote verification and all 108 scheduled windows. Of 114,960 arrivals, 59,554 succeeded and 55,406 failed. Concurrent critical load completed 19/20 workflows, retaining one busy drop. Two subsequent full recovery attempts failed; neither completed original-corpus and every-ACK verification. + +The earlier disconnected-admission failure and the diagnostic fleet's terminal lease-fencing failure remain unresolved despite later passing checks. Results do not transfer between dependency revisions or binaries. + +## Historical qualification checkpoints + The preceding `c51dd121` build is historical evidence. Unlike the earlier `a4500add` documentation advance, this changes routing and resource-permit implementation. Locked @@ -58,7 +57,7 @@ The first four completed load phases recorded 41,485/43,200 metadata OK, stock-Git ls-remote OK across 40 windows. Busy drops and protocol/client errors remain in the result tables. Concurrent critical load recorded 19/20 OK, one busy drop and 323 successful steps. These are failed arrival gates, not -capacity passes; the full schedule continues unchanged. +capacity passes. The full schedule subsequently closed with the totals above. None of these checkpoints replaces the original complete performance schedule. The [three-node proxy campaign](performance/2026-09-30-three-node-proxy.md) retains the earlier fixture race and its regression, plus the candidate's failed @@ -66,12 +65,12 @@ seed and successful recovery of every recorded ACK. | Current gate | Evidence / status | | --- | --- | -| Current Cellule pin and historical runtime | PR pins `1914096`: debug and candidate release correctness passed. PR-head release CI failed HTTP 503 in the disconnected-admission test; 100 isolated and two full-target diagnostic passes do not resolve it. Frozen `0dc04a6`: release qualification and original-corpus activation passed. No recovery or performance result transfers between revisions | +| Current Cellule pin and historical runtime | PR pins `0f4ca09`: all five Linux CI checks passed at `7497fc7`, including 244 release Rust tests, 91 Python tests and eight fresh RustFS gates. Frozen `0dc04a6`: release qualification and original-corpus activation passed. Earlier failures remain unresolved; recovery and performance results do not transfer between revisions | | HTTP listener handoff | [Separate startup qualification](performance/2026-10-01-listener-handoff.md): the affected SHA-256 tests retain their bound listeners through supervised startup; 244 release tests, 84 Python tests, eight fresh RustFS gates and repeated cancellation/handoff checks passed. Both Linux CI workflows at `eaebfc4` passed; not measured performance | | Original corpus maintenance recovery | [Complete metadata and fenced-recovery checkpoint](performance/2026-10-01-original-corpus-recovery.md): all 256 shards and 10,003 Controls checked twice before and after; 301 unsettled Cells drained, zero unsettled Cells, catalog and all published roots unchanged | | Original corpus activation and remote verification | [Full-corpus checkpoint](performance/2026-10-01-original-corpus-activation.md): Ready 9 after four allowed writes and eight unchanged activation scans; subsequent 10K identities, 100 LFS bodies, 200 clones and critical-2 verification passed. Independent audit and 7,061-file verified copy closed; post-load owner-loss recovery remains open | | Complete original-corpus load | All 108 windows / 114,960 arrivals / 8,640 scheduled seconds closed and independently audited: 59,554 OK, 55,406 failed arrivals. Concurrent critical schedule: 19/20 OK, 323 successful steps, 38 UUIDs. Every positive ACK inventoried; not reference capacity | -| Post-load fresh-owner recovery | Exact original owners removed; 32.008191 seconds of confirmed absence, unchanged RustFS. Separate fresh fleet reached readiness after a preserved helper failure. Original-corpus verification then failed a Git v2 clone with exit 128; no complete stage or every-ACK recovery pass. All 3,684 attempt files copied and hash-checked | +| Post-load fresh-owner recovery | Exact original owners removed; 32.008191 seconds of confirmed absence, unchanged RustFS. First full recovery attempt failed a Git v2 clone with exit 128; 3,684 attempt files preserved. A second attempt recorded nine metadata timeouts, one metadata 503 and one LFS 500; all 64 Git commands passed. All 1,454 attempt/input files and 653 bound inputs preserved. Neither attempt completed a stage or every-ACK recovery | | Owned old-binary RustFS upgrade | [Controlled fixture qualification](performance/2026-10-01-old-binary-rustfs-upgrade.md) passed actual old-byte restore through three fresh gateways/proxy, both new Git/LFS writes, twenty scheduled 17-step workflows and recovery of all forty critical repositories after all three owners were killed. Two retained repositories; not the original 10K corpus or reference capacity | | Original rebuilt Cellule artifact | `c51dd121` build, 227 Rust tests, eight RustFS gates and initial 17-step proxy check passed; full 10K/100 seed closed, but remote verification failed HTTP 503. A separate 10K metadata diagnostic returned 9,999 matching identities and one 503; downstream fault/load gates stopped without owner signals. Cause unresolved | | Combined authentication and residency candidate | First publication passed 231 release tests, eight RustFS gates, 20 cold-activation repetitions and 15 residency tests. Exact predecessor descriptor is retained; see the separately qualified catalog follow-up | diff --git a/docs/performance/2026-10-01-cellule-main-qualification.md b/docs/performance/2026-10-01-cellule-main-qualification.md index 166a1f5..48d0229 100644 --- a/docs/performance/2026-10-01-cellule-main-qualification.md +++ b/docs/performance/2026-10-01-cellule-main-qualification.md @@ -1,10 +1,6 @@ # Cellule main revision qualification -PR #18 pins all six Cellule packages to -`0f4ca0919b0dfe20a3dcd964d21da03135e42eed`, checked against upstream main on -October 2. Locked metadata resolves the new revision; its debug, release and -fresh RustFS checks are pending. Passing checks for the previous `1914096` -pin do not qualify this update. +PR #18 pins all six Cellule packages to `0f4ca0919b0dfe20a3dcd964d21da03135e42eed`, checked against upstream main on October 2. Linux debug, release and all eight fresh RustFS gates passed at `7497fc7`. Native qualification, retained-store recovery and performance remain separate, open gates. The original-corpus campaign still uses the frozen, release-qualified `0dc04a6` executable. Its results do not transfer to either newer revision. @@ -24,10 +20,39 @@ no native compilation or running-fleet upgrade was performed for publication. The upstream delta adds admitted-owner fences for application handlers and node-log rotation after member expiry. These changes are not evidence of a -Canopy performance improvement. The existing release workflow will check the +Canopy performance improvement. The release workflow checked the new pin against an isolated fresh RustFS fixture; retained-store upgrade, full-corpus recovery and matched performance require separate verification. +## Closed verification for the current pin + +All five PR and push checks passed at `7497fc720e68a440e27614a83a1442576511cfd2`. These results qualify that source and lockfile, not later diagnostic branches or the running original-corpus fleet. + +| Check | Result | Boundary | +| --- | --- | --- | +| [Linux release correctness](https://github.com/crabbuild/canopy/actions/runs/37029875331) | 244 top-level Rust tests passed, zero failed, nine ignored; 91 Python tests passed; all eight fresh RustFS gates passed | Release lints and executable retention also passed; nested subprocess tests counted once | +| [PR debug verification](https://github.com/crabbuild/canopy/actions/runs/37029883642) and [push debug verification](https://github.com/crabbuild/canopy/actions/runs/37029875188) | Rust and harness jobs passed in both workflows | Separate from native retained-store or performance qualification | +| Independent release audit | Archive digest, 268 source/workflow inputs, six locked Cellule packages and ELF checksum verified | Nineteen evidence files copied and reread on another local filesystem; not an off-machine backup | + +Release verification closed October 2 at 16:07:54 UTC. No runtime budget changed, and the frozen native fleet was not upgraded or restarted. + +| Audited artifact | SHA-256 | +| --- | --- | +| Release archive | `d2cb4060a1f6dc68dc122a53a833f2be04e0e3f812df21b88147cde48008329b` | +| Linux executable | `381ec918026d7d6421a86b0ed916ff40b847c83dccdc881f82624c104c88b33b` | +| Independent audit | `7c28852384d53d5265721af4daa58b37b898b08e5525ff49a61dbce29642abc6` | +| Verified local copy manifest | `16c7775c6525aa3fa7cd42b177a9ff7dba98d4f4118ffafb68f58d2994ac5276` | + +The audited evidence is retained at `/Volumes/Workspace/CrabData/canopy-latest-pin-release-0f4ca09-a9l7xj07`. Passing fresh-fixture checks does not resolve the earlier disconnected-admission failure or the original-corpus recovery failures. + +## Separate query deadline regression + +A [test-only diagnostic at `d1a0f1e`](https://github.com/crabbuild/canopy/actions/runs/37034626141), excluded from PR #18, reproduces a metadata recovery failure with the same Cellule pin. All three isolated release cases failed, as did the four-thread full-library run: 119 passed and one failed. The [debug suite](https://github.com/crabbuild/canopy/actions/runs/37034626111) failed the same assertion. + +The test triggers the existing SQL deadline, verifies that the old handle is fenced, and waits for Control to become idle with no owner. It verifies that the published root is unchanged and another repository remains readable. One subsequent metadata request then returns HTTP 503, not the required HTTP 200. The test drains its owned node before asserting that failure; it adds no HTTP retry or production change. + +This establishes a reproducible recovery defect on an owned in-memory fixture. It does not establish the cause of the retained RustFS failures or provide a qualified correction. The first diagnostic attempt failed compilation before any test ran; moving the test into lifecycle scope corrected that harness issue without changing production visibility. + ## Previous pin: 1914096 The pin-only candidate is `63c93b0f3fadf12893c9c45832f7367ace04ff24`. From bbd3c54589e5dca8f6df1cd49c5d712f5cce68eb Mon Sep 17 00:00:00 2001 From: forhappy Date: Fri, 2 Oct 2026 13:10:23 -0700 Subject: [PATCH 60/60] docs: record native qualification and retained maintenance blocker --- .../2026-10-01-cellule-main-qualification.md | 31 ++++++++++++++ .../2026-10-01-original-corpus-activation.md | 42 ++++++++++++++----- 2 files changed, 63 insertions(+), 10 deletions(-) diff --git a/docs/performance/2026-10-01-cellule-main-qualification.md b/docs/performance/2026-10-01-cellule-main-qualification.md index 48d0229..2902315 100644 --- a/docs/performance/2026-10-01-cellule-main-qualification.md +++ b/docs/performance/2026-10-01-cellule-main-qualification.md @@ -45,6 +45,37 @@ Release verification closed October 2 at 16:07:54 UTC. No runtime budget changed The audited evidence is retained at `/Volumes/Workspace/CrabData/canopy-latest-pin-release-0f4ca09-a9l7xj07`. Passing fresh-fixture checks does not resolve the earlier disconnected-admission failure or the original-corpus recovery failures. +## Native checks for the separate active-owner candidate + +Candidate `d1f9250026bd517576793af44bdb4fc849cedce7` remains excluded from PR #18. It depends on `49921181f1dfe78cc9442931ce6dfb36349c0525` from draft [Cellule #44](https://github.com/crabbuild/cellule/pull/44), not upstream main. Upstream main still points to `0f4ca0919b0dfe20a3dcd964d21da03135e42eed` at this checkpoint. + +Cellule #44's [follower-proof capacity check](https://github.com/crabbuild/cellule/actions/runs/37041888146/job/110953738398) failed. Passing Canopy checks do not clear that dependency failure. + +Native Mac checks closed on October 2. Independent audits verified all 270 source/workflow bindings and six Cellule package revisions: + +| Check | Result | Qualification boundary | +| --- | --- | --- | +| Formatting and release all-target lints | Passed with warnings denied | Exact candidate source and lockfile | +| Release workspace tests | 247 top-level tests passed, zero failed, nine ignored | Nested subprocess tests counted once | +| Python harness | 91 tests passed | Unchanged deadlines, retries and runtime budgets | +| Standalone release CLI build | Passed; executable retained outside Cargo targets | Not deployed against the retained corpus | +| Eight fresh RustFS gates | All eight passed | Test executable, not the standalone CLI; disposable fresh fixture | + +The provider gates cover SHA-256 round trips, SHA-256 merge candidates, signed push options, signed SHA-256 SSH, stock SSH, 4,096 mirror refs, filtered clones and Git LFS over SSH. The helper removed its disposable RustFS fixture after testing. The original retained provider remained unchanged. + +Two evidence collectors failed after their underlying checks passed. The native collector combined Cargo diagnostics with metadata JSON. The provider collector expected Cargo's mutable CLI path to retain the standalone build, but Cargo selected another cached dependency graph. Both original failures remain preserved. Separate collectors and independent audits verified the artifacts without rerunning the tests or builds. + +| Audited artifact | SHA-256 | +| --- | --- | +| Retained native standalone CLI | `a6e827989d850e34101276d7ddbf3698b1a2171e73183c298f5d21b8620111ef` | +| Fresh-provider test executable | `99640caa91b0eeb198f6498d0dc8ea26ba680b5a6a467465caa6399b63aa67be` | +| Independent native audit | `7f6c7b90fa0d04c48b83ad0771365c7e5f3f7a25842f42e332c35c3d0cf6529a` | +| Independent fresh-provider audit | `178064b07b05c31d4082499ccca8231834119ad7515703d3cb3343d103a7dd73` | + +Native evidence is retained at `/Volumes/Workspace/CrabData/canopy-active-owner-native-pp8ks07t`. Fresh-provider evidence is retained at `/Volumes/Workspace/CrabData/canopy-active-owner-native-rustfs-an77ivv2`. Verified copies reside on another local filesystem; they are not off-machine backups. + +The [retained-corpus inspection](2026-10-01-original-corpus-activation.md#read-only-retained-corpus-inspection) found a maintenance compatibility defect. These passing suites do not establish complete recovery, explain historical failures or demonstrate a performance improvement. No retained-store upgrade or activation occurred. + ## Separate query deadline regression A [test-only diagnostic at `d1a0f1e`](https://github.com/crabbuild/canopy/actions/runs/37034626141), excluded from PR #18, reproduces a metadata recovery failure with the same Cellule pin. All three isolated release cases failed, as did the four-thread full-library run: 119 passed and one failed. The [debug suite](https://github.com/crabbuild/canopy/actions/runs/37034626111) failed the same assertion. diff --git a/docs/performance/2026-10-01-original-corpus-activation.md b/docs/performance/2026-10-01-original-corpus-activation.md index 588c78b..4b2ec77 100644 --- a/docs/performance/2026-10-01-original-corpus-activation.md +++ b/docs/performance/2026-10-01-original-corpus-activation.md @@ -336,20 +336,42 @@ The failed attempt is at its verified copy and manifest are at `/Users/haipingfu/.codex/canopy-full-recovery-failed-y871llh9`. +## Read-only retained-corpus inspection + +Two complete catalog and Control scans matched on October 2 after all recorded gateway owners had exited. Control stores each Cell's ownership and lifecycle state. The inspector refused writes; it performed no enrollment, recovery, maintenance transition or activation. + +| Observation | Count | Meaning | +| --- | --- | --- | +| Actual catalog Cells | 11,509 | Includes every required identity and four additional Cells | +| Required Cells | 11,505 | 11,504 original/every-ACK repository identities plus Directory | +| Idle Cells | 10,993 | Catalog inspection only, not Git/LFS content verification | +| Serving Cells | 516 | Unsettled despite the owners' absence | +| Recorded owners | Six retired, zero live | No advertised writers | +| Provider write attempts | Zero | Read-only transport guard | + +The inspector rejected Directory's catalog and Control metadata because its previous-release predicate accepted only the current module code. Directory uses the declared retained code `f7254eda9d5d339566f45457502618ad13cbbf6e5a74595f5b3ce46653ea12f1`, schema 1. A separate read fetched the selected stored release descriptor and verified its BLAKE3 digest, `9a8df7ae5feba1f1760a843bc88d7870af4ebf515b8bc48e45bb7433569cd5d0`. That descriptor explicitly supports this retained code and schema. + +Source inspection found the same restriction in the frozen maintenance worker: `recover_maintenance` calls `Registry::is_current_cell` before restoring an unsettled Cell. This rejects the retained Directory code even though the descriptor declares it supported. The worker was not executed against this corpus, so this is a confirmed source-level compatibility defect, not a maintenance-run result. + +```text +Stored descriptor: current Directory code + retained code, schema 1 + -> retained catalog and Control both reference the retained code + -> inspector's current-code-only check rejects Directory + -> frozen maintenance worker has the same current-code-only check + -> maintenance and activation remain unattempted +``` + +Changing only the inspector would not qualify the frozen maintenance executable. A correction still needs regression coverage, a qualified executable and a supported release transition. Unknown codes, roles, namespaces and schema versions must continue to fail admission. No Control root, ownership, lease bound or deadline changed. + +The closed negative attempt and all bound inputs were preserved as 938 files, totaling 333,374,129 bytes. The verified copy is `/Users/haipingfu/.codex/canopy-retained-inspection-negative-ujryo2ur`; its manifest SHA-256 is `f22500a28ec405a7d3bc42390afbbdfb5e61e1c58d950c4d8a66a1ba277440e2`. Inspection outputs remain at `/Volumes/Workspace/CrabData/canopy-retained-upgrade-inspection-7pyhs2gd`. These are local evidence copies, not provider-data backups or complete remote recovery proof. + ## Remaining verification Cellule upstream subsequently advanced by two commits to `191409685b001a82bd02780def45102b4fc2f164`, observed when publishing this checkpoint. -Those commits change runtime forwarding/compaction and peer HTTP CI gates. -They are **not** the dependency revision tested here. PR #18 now pins that -revision: its Linux debug CI passed 244 Rust tests, 84 Python tests and all eight -fresh RustFS gates. Its release candidate passed, but later PR-head release CI -failed the disconnected-admission test with HTTP 503. All 100 isolated and two -full-target diagnostic runs passed on unchanged source; that does not resolve -the failure. Native qualification, retained-store recovery and performance -remain open. See the -[current dependency checkpoint](2026-10-01-cellule-main-qualification.md). -No result from this activation or frozen campaign transfers to the new pin. +Those commits change runtime forwarding/compaction and peer HTTP CI gates. They are **not** the dependency revision tested here. PR #18 subsequently advanced to upstream `0f4ca0919b0dfe20a3dcd964d21da03135e42eed`; its Linux qualification passed 244 top-level Rust tests, 91 Python tests and all eight fresh RustFS gates. The separate active-owner candidate passed native tests but remains excluded. See the [current dependency checkpoint](2026-10-01-cellule-main-qualification.md) for exact source and artifact boundaries. + +The earlier disconnected-admission release failure remains unexplained. Passing diagnostics and newer suites do not resolve it. Complete retained-store recovery and performance remain open; no result from this frozen activation or campaign transfers to either newer dependency revision. The [performance plan](../performance-plan.md) still requires concurrent faults, fresh-owner verification