Skip to content

File protocol version 2: per-write append, client-chosen handle IDs and bind-ordered stream succession - #359

Merged
SaladDay merged 4 commits into
feature/agent-outside-sandboxfrom
aos/file-protocol
Oct 1, 2026
Merged

SaladDay merged 4 commits into
feature/agent-outside-sandboxfrom
aos/file-protocol

Conversation

@SaladDay

@SaladDay SaladDay commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

File protocol version 2 (plan row 6a, lane L1c). It fixes three defects that the blind review of the world frontend (#349) found, and the frontend can't fix them within version 1.

It uses the Link bind sequence from #358 (merged).

Protocol (internal/sandboxfs/protocol.go, docs/file-access-protocol.md)

The wire version changes outright to 2, with no fallback.

Per-write append

  • OpenAppend is removed and the remaining OpenFlags are renumbered.
  • WriteRequest.Append replaces it. An append write ignores Offset and writes atomically at the end of the file.
  • Any handle that may write accepts both kinds of write, as a native descriptor does when fcntl toggles O_APPEND.
  • AtomicAppend declares support for this.

Client-chosen handle IDs (the 9P fid model)

  • Open, Create and OpenDir carry a HandleID that the client chooses, and their responses no longer return one. sandboxfs.HandleIDs allocates IDs per attachment.
  • The service reserves the ID before any file-system effect. Reserved IDs count toward MaxOpenHandles, and a duplicate ID is refused with InvalidArgument and EffectNone.
  • A Release or ReleaseDir of a reserved ID runs only after that ID's acquisition finishes. That settles a lost or cancelled acquisition: success or StaleHandle proves that no handle with that ID remains.
  • Declared residual: when a node reference's reply is lost, the reference stays held until Detach or lease end. At most MaxInFlight are leaked per transport loss. A Detach to recover them would break every open handle of the world.

Stream succession fence

  • For each (ServerInstanceID, AttachmentID), the server admits one stream at a time, in Link bind order.
  • Before it dispatches the successor's first request, the server stops the predecessor, cancels its requests and waits for every handler and state publication to finish. No timeout bypasses that wait.
  • A stream that was bound earlier but arrives later is refused with sandboxfs.ErrSuperseded.
  • So a Detach cleans up an uncertain Attach, and a Release cleans up an uncertain acquisition.

ErrorCode.Retryable()

ResourceExhausted is the retryable code. A failure with a retryable code and EffectNone may be resent unchanged.

Linux service (apps/sandboxio/internal/fileservice)

  • Files are opened without O_APPEND, and a handle's writes run one at a time. Each write first sets or clears the descriptor's O_APPEND to match Append; an append is then one write call.
  • pwritev2(RWF_APPEND) is not used. On this host's 5.15 kernel, overlayfs ignores it and writes at offset 0, and the container suites on overlayfs caught this.

World frontend (apps/daemon/internal/worldfs)

  • Append: WRITE sends Append from the FUSE write's O_APPEND.
  • Handle IDs and uncertain acquisitions: handle IDs come from HandleIDs. An uncertain Open, Create or OpenDir is never resent; it is released, on the new stream after a redial.
  • Resending: Release and ReleaseDir are resent after a transport failure, since IDs are never reused. Forget is never resent once it may have been applied.
  • Uncertain Attach: it is cleaned up with a Detach on a successor stream, which the fence orders after the failed stream.
  • ErrAttachmentDirty: it now remains only when that Detach can't be sent or answered. doc.go lists those cases.
  • Retry decision: the drain retry uses ErrorCode.Retryable().

Tests

  • Golden frames: regenerated, with an append Write frame added.

  • Server:

    • fence tests for Attach then Detach, and for Open then Release;
    • a Release that races a cancelled Open that is still running;
    • duplicate IDs, both pending and live.
  • Service:

    • per-write append toggling;
    • a lost Open reply released on the next stream.
  • worldfs:

    • TestAppendFollowsFcntl (privileged);
    • TestLostOpenIsReleased;
    • TestUncertainAttachIsDetached.
  • Checks:

    • go build ./...;
    • focused go vet, and go test -race -count=1 on sandboxfs, sandboxlink, sandboxio, worldfs, sessionview and gateway;
    • cross-OS builds;
    • FuzzDecode for 30 seconds.
  • Privileged suites in the container:

    Suite Pass Skip Fail
    worldfs 46 5 (existing documented gaps) 0
    sessionview 6 0 0
    gateway 11 0 0
    fileservice on overlayfs 20 0 0

Review fixes (8b666e2)

From the blind review:

  • An uncertain acquisition no longer blocks its FUSE operation. The drainer now releases it, and the operation returns EIO at once.
  • Superseded waiting streams end at once. The fence lives in shared per-stream state. The newest successor still waits for every earlier stream that ran a request.
  • The abort bound now covers teardown. That covers the stream close, the Detach and the drainer wait.
  • The node-reference residual is restated without a numeric bound. The doc names the client usage that avoids abandoned replies, and worldfs follows it.
  • New tests: TestSupersededWaiterEnds, TestAbortBoundedWhenTransportBlocks, and a rewritten TestLostOpenIsReleased.
  • Later review rounds (b68e740, 480bda2):
    • Pruning keeps a fence until every earlier stream that ran requests has drained.
    • A call cancelled mid-write no longer waits on the transport, so the File client returns at once and fails the stream.
    • A stream whose lease ended before admission is refused with sandboxfs.ErrLeaseEnded.
    • The cancellation doc names each outcome.

…nd bind-ordered stream succession

- sandboxfs.Version is 2, so Link refuses a File stream between builds of different layouts with VersionMismatch.
- WriteRequest.Append replaces OpenAppend. The Linux service opens files without O_APPEND, runs a handle's writes one at a time and sets or clears the descriptor's O_APPEND to match each write; overlayfs on some kernels, including 5.15, drops pwritev2's RWF_APPEND and writes at the offset.
- Open, Create and OpenDir carry a client-chosen HandleID that the service reserves before any filesystem effect; a duplicate is InvalidArgument with EffectNone, and a Release of a pending ID waits for its acquisition.
- sandboxfs.Server admits one stream per (ServerInstanceID, AttachmentID) in Link bind order: it drains the predecessor before a successor dispatches, and refuses a stream bound before one already admitted with ErrSuperseded, also after the newer stream ended, until the attachment's lease ends.
- ErrorCode.Retryable marks ResourceExhausted as the only code a client resends after EffectNone.
- The world frontend passes O_APPEND from each WRITE's flags as Append, so fcntl(F_SETFL) toggling works, and takes handle IDs from sandboxfs.HandleIDs. An Open, Create or OpenDir that may have taken effect without a response is never resent; its ID is released, on a successor stream after a redial, and StaleHandle settles it. A failed Serve detaches on a new stream fenced behind every earlier request, so ErrAttachmentDirty remains only when that Detach cannot be sent or answered.
- fileservicetest serves through sandboxfs.Server with a Dial-order bind sequence.
The world frontend queues the release of an acquisition in doubt for the
drainer, so the kernel request returns EIO without waiting for an
unreachable service. Abort, Stop and shutdown close streams without waiting
for the transport and send Detach on its own goroutine, so their bounds hold
when outgoing traffic blocks; no stream opens after shutdown.

The File server keeps its succession fence in shared per-attachment state:
a successor's requests wait before they run, and a successor that has run
nothing ends at once when it is superseded or its stream or context ends,
passing its wait on to the next one.

The File access protocol states the node-reference residual without a
numeric bound and names the client usage that avoids abandoned replies.
…ransport

The File server forgets an attachment's stream entry after its lease ends
only once that stream has drained and so has every earlier stream that ran
a request, so a later stream can no longer bypass a handler still running.

The File client no longer waits for a transport that holds a write or a
close: a call whose context ends while its request is being written fails
the stream and returns at once, and the client closes the transport in the
background after it fails the calls in flight. The world frontend relies on
this for its Detach and stream drops instead of its own goroutine wrapper,
and its abort test completes a request after Attach before it jams the
transport, so the cancellation finds Attach in doubt.
@SaladDay
SaladDay merged commit a079a70 into feature/agent-outside-sandbox Oct 1, 2026
@SaladDay
SaladDay deleted the aos/file-protocol branch October 1, 2026 04:14
Serve refuses a stream whose attachment lease ended before admission with
sandboxfs.ErrLeaseEnded, checked under the lock that forgets settled
streams, so a stream bound before a forgotten one is always refused. The
fence test now renews the lease for the later stream, as Link does for an
attachment ID bound again after its attachment closed.

The File access protocol states that a call cancelled while its request is
being written returns the transport failure, Unknown with EffectPossible.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant