Skip to content

Expose admitted owner fences to application command handlers - #38

Merged
forhappy merged 8 commits into
mainfrom
codex/canopy-admitted-owner-fence
Oct 2, 2026
Merged

forhappy merged 8 commits into
mainfrom
codex/canopy-admitted-owner-fence

Conversation

@forhappy

@forhappy forhappy commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Prepared application operations need to reject work created under a previous Cell owner. Expose OwnerFence { incarnation, epoch } on the activation-specific admission capability and CommandContext, so handlers can compare stored operation tokens with the fence of the execution that was admitted.

Local commands and authenticated inbox effects obtain this value from the validated destination handle. Renewal, root publication, migration, and failed transfer admission replacement preserve the fence; a new ownership claim advances it. Recorded outcomes replay without re-running a handler or substituting the successor fence. This adds no per-command authority read and changes no persisted SQL, request digest, or peer wire format. Manual CommandInvocation construction must now supply the fence. Application tokens must also bind their Cell or target.

Coverage includes stale epoch and wrong-incarnation rejection, exact typed outcome replay after takeover, renewal and migration stability, failed transfer admission replacement, follower-proven outcome recovery after owner loss, and authenticated inbox replay plus fresh delivery under the successor fence.

CI readiness: refreshed against current main, including the website Rust example linker fix and current routing qualification workflow. Reader publication tests now hold the recruitment clock fixed so the five-second periodic repair cannot coalesce their expected publication passes. They assert that repair never supplies a hint, retain a bounded watchdog, and resume real time to drain the host on success or failure. Updated the fence test helper for Rust 1.99 Clippy.

The strict reader-loss probe exposed a reference TCP transport starvation path: a killed container could consume the full request deadline during connection setup, leaving a receipt-bound lane without progress during replacement. Connection setup now has a 250 ms cap within the original deadline and reports a known pre-dispatch failure. Connected requests retain the overall reply deadline and unknown-outcome semantics. Regression coverage models a stalled connection and a dispatched request awaiting its reply. Qualification thresholds, lane progress requirements, receipt minima, and workloads are unchanged.

Validation:

  • Rust workspace and MSRV: all-feature tests, Clippy with warnings denied, API documentation, website Rust examples, executable examples, local LTX, and architecture/documentation checks.
  • Compose smoke and routing: real object storage, 3/5/10/20-node reader and writer scaling, and four balanced routing pairs per mode with retained raw evidence. The reader-loss window committed all 300 scheduled writes; every reader lane made progress during replacement, including 1,020 receipt-covering reads in the lane that previously stalled.
  • Object and follower write-capacity qualification, qualification contracts, and website.
  • 30 repeated complete application integration runs on the fixed recruitment tests. A local negative control deliberately omitted the third publication: the test rejected it before periodic repair and drained the host cleanly. The temporary diagnostic workflow was removed.

Routing evidence retains the nonblocking >10% latency alerts for review. The final leased comparison reports forwarded-query expired-burst p99 at +18.0% and p95 at +8.7%; all blocking gates pass in both modes, and object-only routing has no nonblocking alerts. An independent raw-evidence audit reconciled 484,384 latency samples, 65,536 publication records, and 98,304 recovered commands across the 16 benchmark runs.

This is the runtime prerequisite for Canopy's prepared-operation publication fence. The embedding application owns token creation and binding, publication integration, and its own capacity qualification.

Pause the recruitment clock and hold its auto-advance guard while external SQL workers reply. Slow or competing fixtures can otherwise cross the five-second periodic scan, which coalesces hints and invalidates per-publication counts. Keep the two-second virtual deadline and assert that periodic repair cannot supply a hint. Drain the host and release the guard on completion.

Update the owner-fence test helper for the Rust 1.99 chunks_exact_to_as_chunks lint, and exercise the full competing integration fixtures in the temporary CI diagnostic.
The missing-publication negative control correctly expires the virtual deadline. Resume time before shutdown so SQL close replies drain on the failure path even after the watchdog releases auto-advance. Remove the temporary diagnostic workflow after all 30 full application integration repeats passed on GitHub.
@forhappy
forhappy marked this pull request as ready for review October 2, 2026 09:17
@forhappy
forhappy merged commit 91e2aaf into main Oct 2, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant