Conversation
1962719 to
9850e91
Compare
|
Follow-up qualification and hardening (commit 9393caf):
The Kubernetes 5,000-commit stress artifact remains explicitly negative at the 500-commit fetch/repack boundary because the interrupted pre-fix repository state still requires a large historical pack; it is not being reported as parity proof. Hosted-provider, multipart, replica/tiering, mount/browser, S3 gateway, migration/recovery, and backup/restore rows remain release gates. |
|
Parity follow-up (commit d6431f8): removed the stale client rejection for protected pushes carrying a v2 mirror-plan ID. The plan ID is already authenticated in |
|
Added |
|
Documentation follow-up (commit 0233c25): the main capsule publication design now records the authenticated external thin-base rule and protected mirror-plan receipt path alongside the parity inventory and RustFS evidence. |
0233c25 to
6ad3b0b
Compare
V1 parity passImplemented and pushed in
Proof after rebase:
Remaining release blockers are intentionally explicit: hosted-provider checksum/multipart and 5,000-commit current-format replay; managed replica failover/repair; tier/archive restore; mount range/cancellation/unmount; browser/HTTP load and fault matrix; S3 gateway operation/concurrency/restart matrix; lifecycle/workflow/admin inventory; backup/restore export inventory; and migration fault/resume/provider plus older-Git/interrupted/adversarial qualification. Full v1 production parity is not claimed until those Level-3 gates pass. |
|
Parity closure update (c2d87d8):
Local proof after this change:
This closes the local wiring gap, but is not a claim of complete v1 production parity. Release gates remain: live S3/GCS/Azure lifecycle and restore behavior; replica readiness/failover/repair; restored-content verification; the full FUSE/NFS range/cache/cancellation matrix; browser and smart-HTTP load/fault coverage; S3 gateway restart/concurrency/request-count coverage; and migration, backup inventory, and delete/restore qualification on populated v1/v2 repositories. |
|
Follow-up test hardening: the focused mount module now passes 122/122. I also serialized the unmount test's HOME override through the existing test guard; this removes a process-global HOME race that could make |
|
Final local rerun after the test-only race fix: |
4b94ade to
cc8700f
Compare
|
Layered-pack qualification update (commit 1640af9):
The 5,000-commit replay with 500-commit fetch/repack checkpoints and hosted-provider matrix remain explicit release gates; v1 is not being retired until those pass. |
|
Follow-up pushed in The first fresh PR-208 5,000-commit run found a correctness blocker before replay could continue: an append-only in-memory visibility dictionary could be non-canonical when serialized into the layered capsule ( Proof for the fix:
I am rebuilding the exact release binary from this commit and rerunning the fresh local-RustFS qualification. The earlier run remains recorded as a failed negative qualification at seed + replay 1; no 5,000-commit success is claimed until all replay, checkpoint fetch/repack, clone, and fsck gates pass. |
|
Layered-pack follow-up pushed in
The PR is updated with the fixes and evidence, but v1 is not retired and the 5,000-commit/multi-pack cold-clone gates remain open until the authenticated multi-member install or equivalent whole-member union path is implemented and requalified. |
|
Post-push test completion for
The branch remains clean apart from the pre-existing untracked local |
|
Updated in commit 1cfdd63 (perf(fetch): preload layered locators for cold clones).\n\nWhat changed:\n- Complete layered views now coalesce and authenticate index, reverse-index, and kind-bearing locator sidecars, then expose inline locators to the planner.\n- Ordinary incremental haves stay on the footer/tip-bound path; cold, filtered, shallow, and tag requests promote only when complete visibility is required.\n- Multi-member cold clones no longer fall back to per-object visibility reads; direct one-pack installation remains fail-closed.\n- Design history and qualification evidence are recorded in crab/docs/design/capsule-layered-packs.md.\n\nVerification:\n- crab-read: 200 passed.\n- remote-helper: 139 passed.\n- upload-pack wire: 39 passed.\n- Local RustFS, Kubernetes-derived fixture: blob:none clone 14.50 s with 173 ms planning; cache-miss shallow blob:none clone 6.12 s with 1,095 planning reads and 372 terminal response-pack reads; unfiltered multi-member clone 85.94 s, exact source tip, native git fsck --full clean.\n- Full, filtered, and shallow clone tips all match the source.\n\nThe fresh 5,000-push/fetch/repack matrix, hosted-provider latency, and v1-retirement gates remain open; this update does not claim those are complete. |
|
CI follow-up: GitHub did not emit a synchronize run for the new head, so I manually dispatched the current commit (1cfdd63) against the repository workflows:\n- CI: https://github.com/crabbuild/crab/actions/runs/35675208869\n- Git protocol v2 partial-clone qualification: https://github.com/crabbuild/crab/actions/runs/35675210814\n- Large repository RustFS qualification: https://github.com/crabbuild/crab/actions/runs/35675212409\n\nAt this update they are queued, not yet green; the prior completed run was for an older head. |
|
Pushed dba2cb4 (fix(protocol-v2): retain negotiated fetch haves). What changed:
Proof:
Qualification status remains honest: the replay later reached an existing 503,980,520-byte Crab/Xet pointer commit and stopped with CRAB-E0086 because the replay harness had not staged its local chunks. The 5,000-push/xorb qualification gate is still open; this is a staging-contract failure, not evidence that the haves fix is incorrect. |
|
Pushed follow-up commit 1696f53 to PR 208. The fresh full-history GitHub-origin Kubernetes RustFS qualification (pr208-v2-fresh-github-smoke2-20260921, binary dba2cb4) passed seed publication, protocol-v2 incremental fetches at pushes 1/5/10, exact tip checks, cold and warm full clones, and native fsck. Measured incremental fetches were 33.457s / 15,494 storage range reads / 50.1MB response at push 1 and 24.227s / 16,999 reads / 55.5MB response at push 5. Cold full clone was 217.5s with 10 store requests; warm clone was 99.8s with zero store requests. The run stopped only at the blobless-clone qualification assertion blob-none-ordinal-metadata-lookup: the exact ordinal-metadata lookup branch emitted no locator_lookup_mode event, so the harness observed zero metadata events even though the request completed through the catalog-filter plan. The new commit adds that trace at the crab-metadata reader boundary; no data-path or authorization behavior changes. Focused crab-metadata tests pass (212/212). Fresh workflows have been dispatched at this new head:
I am not marking the PR green until those runs complete. |
|
Pushed This fixes the v2 incremental-fetch regression caused by generation-owner checkpoint compaction: the control-only reader previously saw no live capsule transitions after compaction and fell back to a visibility traversal. Layered checkpoints now carry a bounded, authenticated recent per-ref transition suffix in the footer. Control-only fetches consume that suffix; older/incomplete have chains still fail closed to the existing catalog/traversal planner. The complete ordinal visibility body remains authoritative for cold/strict paths. Proof:
Fresh manual workflows for this commit:
|
|
Qualification update for
|
|
Update: pushed
Known local gate: workspace |
|
Qualification update from the exact |
|
Focused regression gate from |
|
Qualification update: the 1,500-boundary checkpoint completed with 1,501/1,501 pushes successful. Owner maintenance was 511.8 s (peak child RSS ~1.69 GiB, two active packs). The incremental fetch then completed in 15.4 s with 146 storage requests and 17,445 logical objects; visibility planning was 1 ms. This reinforces the outstanding protocol-v2 response-pack scaling gap; correctness remains intact and the 5,000 replay is continuing. |
|
Qualification update: the 2,000-boundary checkpoint completed with 2,001/2,001 pushes successful. Owner maintenance was 659.3 s (peak child RSS ~0.79 GiB, two active packs). Incremental fetch completed in 5.88 s with 126 storage requests and 20,279 logical objects; visibility planning was 3 ms. Fetch wall time varied versus the 1,500 sample, but the request count remains dominated by protocol-v2 response-pack reads. The 5,000 replay continues with no correctness failure. |
|
Follow-up at 00b7177: source-matched release binary rebuilt and the full isolated RustFS 1.0.0 GA protocol-v2 partial-clone smoke passed all 92 checks. Actual filtered incremental fetch advertised version 2 / command=fetch; the preserved filtered-transfer-smaller gate now measures delivered Git pack bytes, 19,902 filtered (including initial lazy fetches) versus 188,716 full, rather than remote-reader counters that are zero on direct cold clones. Strict Git and lifecycle checks passed. Retained report SHA-256: 6d7dbc98a998c4233ca7d47e4da498d8a7326ce3c32c85626663d61fa62c8c55. Cell RustFS workflow now supplies the public-test environment scope; the new CI run is still underway. Kubernetes 5,000-push correctness remains passed but fetch p95 is 34 requests versus the unchanged 10-request gate. |
|
Qualification update (head 1a0165b): the corrected layered-checkpoint reader contract passes locally, 38/38 crab-remote checkpoint tests. CI has restarted and is still in progress. A fresh local RustFS GA Xet run exercised 40 GiB of large files across three versions (120 GiB logical history). It passed 96 checks through seed push, two incremental pushes, three layered repacks, retained refs/history, cross-repository chunk reuse, and byte-identical consumer hydration. Retained xorb bytes were 21.54 GB (16.7% of logical history); incremental pushes were 72 object-store requests each (large-file/xorb traffic, not the small Git-commit path). Seed upload was 749 requests. This is partial qualification, not a 100 GiB pass: I interrupted the full cold-clone hydrate when the shared workspace reached its 20 GiB safety floor. The final report records that deliberate interruption as a failed run; no product corruption was observed before the stop. Full cold-clone, 100 GiB, v1 comparison, and the remaining red CI jobs are still open gates. I preserved run artifacts and remote objects. |
|
Follow-up at head 158e104: the Xet qualification harness now releases only its own caches after each verified phase, budgets the modeled peak checkout/origin/cache footprint, and rechecks free space before every hydrate. A regression test reproduces the prior 40 GiB run being admitted at 153 GiB free; 31 local harness/meter tests pass. This is a safety and qualification-harness fix, not a completed 100 GiB run. The retained 5,000-commit trace also confirms the final warm fetch is already one complete GET per each of 24 capsule sources plus eight control/admission requests. The unchanged ten-request gate still fails; reader range coalescing alone cannot solve that physical source floor. CI has restarted on this head. |
|
Qualification update (c172bf1): the 500 MiB × 5-version RustFS 1.0.0 GA rehearsal passed all pushes/repack, independent consumer byte-identity, and first cold-clone hydrate; the second hydrate was stopped by the harness headroom guard (22,398,337,024 free < 22,523,412,480 required), not by data corruption. The harness now releases verified consumer worktrees, dehydrates the published source after copying the consumer fixture, and records available/required bytes. Focused Python tests: 33 passed. The source-dehydration E2E repeat and 100 GiB gate remain open; current mounted-volume free space is below the 24.6 GB minimum rehearsal preflight. Reports/logs retained, isolated disposable bucket/checkouts cleaned. |
|
Qualification follow-up: the fresh |
|
Qualification update (HEAD 9b91d0b): two fresh traced 1 GiB × five-version RustFS GA Xet runs passed 159 checks/174 commands each with zero proxy errors or 5xx responses. The first 1 GiB rehearsal had one proxy TimeoutError despite a passed user-visible report; this push adds a zero-proxy-error gate so that condition now fails qualification. Incremental Xet pushes in the latest clean repeat were 369–452 ms / 34 object-store requests. This is bounded lifecycle evidence, not completion of the 100 GiB, paired v1/v2, provider, or release gates. CI currently has a known mirror smoke fixture issue (a second repack does no work) and a prior public ECR rate-limit failure; both are being tracked explicitly. |
|
Current-head local RustFS 1.0.0 GA diagnostic (binary SHA-256 733823bf5f68b22e5c24b9fea9127959425c1551ec5209c4c4a40fc673f4f59a): the mirror smoke reproducibly reached 117 checks / 449 commands, then failed its source-ahead stale-plan assertion. The preceding second repack was a no-op (packs 1→1, bytes read/written 0), so applying the still-valid plan succeeded. A focused mirror unit test confirms plan identity changes when the destination snapshot actually changes. This is fixture setup, not evidence that stale plans are accepted after a real metadata mutation. Separately, the in-memory repack request assertion reproduces 17 operations for the current two-phase logical+physical publication against its <=12 expectation; the publication/recovery tradeoff remains open. CI and the 100 GiB Xet gate are not yet complete. |
Status
Protocol-v2 capsule roots, per-ref publication, external Xet xorbs/shards, and stable layered packs are implemented. This PR is not production-qualified and v1 must remain available. Current head:
c7c88bfd57d. Correctness and performance gates have not been relaxed.Local RustFS 1.0.0 GA evidence
A fresh GitHub Kubernetes clone supplied a 5,000-commit first-parent replay. All 5,000 pushes, ten exact-tip fetch-before-repack intervals, strict Git fsck, seed/final remote Crab fsck, independent cold/warm clones, and sampled blob digests passed. Incremental push mean/p95 was 228/422 ms with 7.012 mean origin requests and no growth across 500-commit windows. Fetch p95 was 5.996 s but 34 origin requests versus the unchanged ≤10 gate. Full data and retained artifact hashes:
crab/docs/benchmarks/capsule-v2-kubernetes-5000-rustfs-ga.md.A matched, isolated 500-commit diagnostic compared the retained 32-way per-ref capsule compaction policy with an uncommitted four-way experiment on the same upstream commit range. Four-way reduced fetch requests from 32 to 14 and interval-repack requests from 63 to 27, but raised mean push requests from 7.012 to 7.488; observed local fetch time was 6.079 s versus 4.423 s. Both runs passed exact tips, clones, strict Git/Crab fsck, and sampled blob bytes, but both failed the ≤10 fetch-request gate. The fan-in experiment was reverted; single-pair latency differences are not treated as causal proof. Details and report hashes are in the same benchmark document.
A separate default 100 GiB Xet workload exercised three file versions, historical hydration, cross-repo dedup, restore, republish, exact bytes, and remote fsck. All 3,850 file comparisons passed. The harness still failed because the metered seed push saw three 60-second proxy timeouts on duplicate xorb PUTs. Path tracing reproduced those requests; direct and later metered conditional-PUT probes returned 412 promptly, but the original stall cause is not established. Cold 100 GiB hydrate took 37.9 minutes, so performance is not qualified. Full evidence:
crab/docs/benchmarks/capsule-v2-xet-100g-rustfs-ga.md.Open gates
Design and release gates:
crab/docs/design/capsule-layered-packs.mdandcrab/docs/design/capsule-xorbs-shards.md.