Skip to content

New Runtime - #1

Merged
aruokhai merged 70 commits into
mainfrom
merkle-block-store
Sep 26, 2026
Merged

aruokhai merged 70 commits into
mainfrom
merkle-block-store

Conversation

@aruokhai

Copy link
Copy Markdown
Contributor

No description provided.

aruokhai and others added 30 commits August 8, 2026 18:59
The previous engine mapped POSIX paths onto S3 keys directly. For an enclave,
where the host is outside the trust boundary and storage must be assumed
hostile, it had three fatal properties: no integrity checking of any kind, no
anchor and so no way to detect a rollback, and mutation in place, which is
incompatible with S3 Object Lock.

This replaces the storage layer with a ZFS-style copy-on-write block store.
Content lives in immutable AEAD-encrypted blocks packed into slab objects, and
the whole filesystem hangs off one signed, hash-chained root record written
with a conditional PUT under Object Lock COMPLIANCE retention. Verifying that
record transitively verifies every byte beneath it.

Against an adversary holding full write access to the buckets: they can make
the filesystem unreadable; they cannot make it read wrong, and they cannot
make it read old.

Storage layer (crates/s3fs-core/src/store, crypto):
- BlkPtr addresses a byte range inside a slab rather than a content hash, so a
  commit costs a couple of PUTs however many blocks changed. Content addressing
  would have cost one PUT per 128 KiB.
- Indirect blocks are variable length with trailing holes trimmed, so a
  two-block file pays a 256-byte indirect block rather than a full record.
- Directories are a separator-indexed B+tree. Extendible hashing was planned,
  but its cheap-doubling property needs many slots sharing one bucket block,
  and the AEAD binds every block to its index — deliberately, since that is
  what stops an attacker relocating blocks. Sorted iteration comes free.
- Transaction group numbers are never reused. The non-obvious threat is a crash
  between writing slabs and publishing a root: the tip still names txg T, so a
  naive remount resumes at T+1 and repeats every nonce already spent on the
  orphaned slabs. next_safe_txg probes past orphans before the first commit.
- Losing the conditional PUT for a sequence number poisons the mount rather
  than retrying, for the same reason.
- Root signatures verify against the key we derived, never the key the record
  names, and the record's sequence is checked against the key it was read from.

Filesystem layer:
- rename is atomic: one directory-entry move in one commit. The async
  copy-then-delete queue and its redirect machinery are gone, and a directory
  rename no longer costs O(entries) copies.
- stat cannot go stale; link counts and all three timestamps are real.
- Identity is (objid, gen), stable across mounts, so is-same-object and
  metadata-hash are exact rather than a hash of an ETag.
- Hard links, and unlink-while-open, both previously documented as impossible.
- Snapshots: every root record already is one.

Also fixes a credential leak in s3fs-runner, which forwarded the entire host
environment — including AWS_SECRET_ACCESS_KEY — into the guest.

Deletes mpu.rs, rename.rs, flusher.rs, buffer/ and inode/ (~3.4k lines): the
MPU state machine, UploadPartCopy planning, race-three lookup and tiered part
schedule were all compensating for the path-to-key model.

333 unit tests, plus MinIO integration coverage of Object Lock retention,
remount, slab packing economics, tamper detection and the rollback floor.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The storage engine is done; the enclave is not. Three things stand between
this and running for real, and each has a decision in it worth writing down
rather than rediscovering.

M8, attestation and key release. --master-key is a development seam: a secret
on the command line is visible to the parent instance, which is exactly the
party the enclave exists to exclude. Notes the two seams that do not exist yet
— AwsS3BackendConfig has no http_client field, so the SDK cannot be pointed at
a vsock proxy, and its static credentials carry no expiry, so an attested
session token dies mid-run.

M9, garbage collection. Deliberately not shipped rather than shipped
approximately: a slab is dead when no root you intend to keep references it,
and "intend to keep" is policy, not a fact about the data. Getting it wrong
deletes live data. Records the mark-and-sweep design and the cheaper stopgaps.

M10, the enclave-runtime crate. s3fs-runner stays the dev CLI so the local
loop keeps working without NSM or vsock; a new crate is the deployment target.
The guest ships inside the EIF so PCR0 covers it, which is what makes the key
policy attest to the code that will actually read the data — proving that some
code with the right image hash is running is not the same claim.

Flags add_wasi_minus_filesystem for extraction into a shared crate: it is a
hand-copied clone of a wasmtime-wasi function and will drift silently on any
bump, and a second binary would double that.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The thing that actually runs inside a Nitro Enclave did not exist. s3fs-runner
is shaped as a developer CLI — every setting a required flag, the guest's
environment empty unless each variable is named — and an enclave has no shell
to type flags at and no operator to enumerate variables. It gets whatever the
image baked in, and the guest is expected to read it.

enclave-runtime is the deployment target: environment-first configuration with
matching flags, the guest loaded from a known path (/enclave/guest.wasm), and
the guest inheriting this process's environment minus anything under AWS_ or
S3FS_. That one prefix rule covers the master key and the whole credential set
without anyone enumerating them, and it keeps working when a new one is added.
Explicit --guest-env still wins, because explicit is a decision and inheritance
is an accident waiting to happen.

s3fs-host holds what both binaries need. It absorbs s3fs-wasmtime, which after
the rewire had exactly one consumer and pinned wasmtime separately; mount sits
behind an `aws` feature so the wasi:filesystem bindings stay usable over any
Backend without dragging in the SDK. The hand-copied clone of
add_to_linker_with_options_async now exists once rather than being duplicated
into a second binary.

Exit codes are no longer collapsed. A guest calling exit(3) reached us as a
trap carrying I32Exit and became exit 1, losing the code; it now propagates,
and a trap (70) is distinguished from a runtime failure before the guest
started (71) — an integrity or rollback failure at mount is a security event,
not a bug in the guest.

Two bugs found by running it rather than by reasoning about it:

- A clap bool with `env` demands exactly true/false, so ENV S3FS_FORCE_PATH_STYLE=1
  — the obvious thing to write in a Dockerfile — hard-failed at startup. Inside
  an enclave that is expensive to diagnose. Both binaries now accept 1/0,
  true/false, yes/no, on/off.
- crates/s3fs-core/src/backend/aws.rs had `.copy_object()AwsS3Backend` committed
  in b11f95b. It was introduced between the last green test run and `git add -A`
  and swept in; the tree at b11f95b does not compile.

examples/guest-smoke is a new std-only guest: no wasi-sdk, so CI can run it on
every push, unlike guest-fsdemo with its bundled SQLite. It exercises write,
in-place patch, rename, read_dir, asserts no AWS_/S3FS_ variable reached it,
and counts its own runs from a file it wrote — so a second run printing "run 2"
is proof that commits reached the store and the anchor chain was picked back
up, not that the filesystem was quietly reformatted.

The new guest-e2e CI job runs exactly that, twice. It is also the only real
drift guard for the linker list, since wasmtime's Linker cannot be enumerated
and no unit test can prove the list complete. Also removes the minio-integration
service container, which was malformed and unused — the tests start their own
via testcontainers, so the job was passing vacuously.

Verified end to end against MinIO with Object Lock on the roots bucket: three
consecutive runs, root_seq 0 -> 9 -> 17, distinct Merkle roots, ledger
persisted across processes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A real C database driving the filesystem is a far harsher test than anything
hand-written. SQLite does page-granular random reads and writes, creates and
unlinks a rollback journal on every transaction, fsyncs where it genuinely
needs durability, rewrites the whole file during VACUUM — and then tells us
whether the bytes came back correct, via PRAGMA integrity_check.

examples/guest-sqlite covers DDL, bulk and batched inserts, point and indexed
selects, updates, upserts, cascading deletes, rollback and nested savepoints,
all five constraint kinds, joins cross-checked against correlated subqueries,
aggregates, recursive CTEs, window functions, blobs up to 2 MiB verified byte
for byte, unicode and NULL handling, triggers, views, ALTER TABLE including
RENAME COLUMN, REINDEX, VACUUM, ANALYZE, and integrity_check both inline and
after closing and reopening the database. Every phase is timed, so it doubles
as the benchmark.

Running it found a real trap, and it is not in this filesystem. SQLite locates
a directory for temporary databases by probing candidates with access(2), which
WASI does not provide — so every candidate is rejected and VACUUM fails with a
bare "disk I/O error" that says nothing about a missing syscall. Setting
temp_store_directory does not help: that pragma validates the path the same way
and reports "not a writable directory". The fix is temp_store=MEMORY, now
documented in COMPATIBILITY.md alongside the locking_mode=EXCLUSIVE that WASI's
lack of fcntl already required.

journal_mode is left at DELETE rather than the MEMORY the older demo used. The
rollback journal is then a real sidecar file with a full create/extend/truncate/
unlink lifecycle, which exercises parts of the filesystem a memory journal skips
entirely.

Verified against MinIO with Object Lock on the roots bucket, at 2 000 and
20 000 accounts (60 000 rows total). Everything passes; integrity_check is
clean before and after reopen.

CI's guest-build job only ever compiled a guest and never ran one. It is
replaced by guest-sqlite, which builds and runs this workload end to end.

Also fixes .gitignore: `/target` anchored only the workspace root, so the
example guests' build directories were tracked, and 13 build artifacts went
into the previous commit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Extends the workload to the operations the first pass skipped, chosen for the
filesystem behaviour they provoke rather than for SQL coverage:

- DROP index/view/trigger/table, and a check that a dropped trigger stops
  firing. Dropping frees pages, which is what gives the following VACUUM
  something to reclaim — so drops now run last.
- ATTACH: a second database file open alongside the first, a join spanning
  both, and integrity_check on the attached one.
- Incremental blob I/O via sqlite3_blob_open, writing 4 KiB chunks in reverse
  order so pages are dirtied out of sequence. The rusqlite `blob` feature was
  enabled in the first pass and never used, which was sloppy.
- WITHOUT ROWID tables, STORED and VIRTUAL generated columns, partial indexes
  and expression indexes — all of which have different on-disk shapes.
- INSERT OR REPLACE / OR IGNORE, BEGIN IMMEDIATE / EXCLUSIVE, date/time and
  scalar functions, TEMP tables.
- JSON, FTS5 and R-Tree, probed rather than assumed. All three are present in
  the bundled build and all three pass; FTS5 indexes 2 000 documents and then
  rebuilds every shadow table.

The first run of this found the most useful number in the whole exercise. A
phase that should have taken half a second took 38.9 seconds, because I had
left 500 inserts in autocommit. Each statement outside a transaction is its own
commit, and a commit here is a transaction group: slab PUTs plus a signed root
record published with a conditional PUT. So the phase was measuring commit
latency, not the table shape.

Rather than only fixing it, there is now a phase that measures it deliberately:
25 inserts in autocommit take 1 796 ms (14 rows/s); the same 25 in one
transaction take 69 ms (364 rows/s). 20 000 batched inserts land in 142 ms —
less time than 25 unbatched ones. That is the single most important thing to
know when sizing a workload against this store, and it now has a line in the
benchmark table instead of hiding inside an unrelated phase.

README gains a "Running SQLite" section covering the two required pragmas and
why, what cannot work (WAL needs a shared-memory index WASI cannot provide;
concurrent connections need fcntl locking it also lacks — neither restricting
anything real, since the store is single-writer by design), the optional
modules, the benchmark table, and the batching guidance.

Verified end to end against MinIO with Object Lock on the roots bucket at
20 000 accounts. Every phase passes; integrity_check clean before and after
reopen.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
An enclave's system clock is not its own. It is seeded by the hypervisor at
boot, has no NTP, and drifts — and the party that sets it is the parent
instance, which is exactly the party the enclave exists to distrust. AWS
addresses this by exposing the Nitro card's PTP hardware clock, synchronised to
the Amazon Time Sync Service, at /dev/ptp0.

The guest's wasi:clocks/wall-clock now reads that device. set-times with "now"
shares the same clock, so a guest cannot read one answer from wall-clock and
have the filesystem record a different one.

Reading uses the POSIX dynamic-clock mechanism — open the character device,
derive a clock id from the fd, clock_gettime on that — via
rustix::time::clock_gettime_dynamic. rustix is already compiled into the tree
through wasmtime, so this adds no build cost and no direct libc dependency.
Hand-coding the ((~fd) << 3) | 3 encoding would have been worse than useless:
getting it wrong yields a valid-looking clock id for some *other* clock rather
than an error.

--clock-source is auto | ptp | host. `ptp` refuses to start without the device
and is what an enclave image sets; `auto` prefers PTP and warns loudly when it
falls back, so a misconfigured enclave never looks like a correct one in the
logs. --clock-check prints readings and exits without mounting anything.

The monotonic clock deliberately stays on CLOCK_MONOTONIC. It backs WASI's
timer subscriptions, so it must be cheap, and it must never step backwards —
which a clock disciplined by an external source can, and a wall clock is
allowed to.

HostWallClock::now cannot fail, so the adapter had to decide what a failed
device read means mid-run. It serves the last good reading and logs. A guest
should not trap because a driver hiccuped, and a zero timestamp — expired
certificates, rejected tokens, dates in 1970 — is far more damaging downstream
than one a few milliseconds stale.

Verified against a real PTP hardware clock, not a mock. /dev/ptp0 is root-owned
here and there is no passwordless sudo, so deploy/ptp-check.sh passes the host's
PHC into a container via --device. Readings:

  vs CLOCK_REALTIME   -520.851 ms, stable across five samples
  read cost           10-15 us, against ~0.3 us for the system clock

The skew is the useful part: it proves the PHC is genuinely a different clock
rather than CLOCK_REALTIME under another name. The read cost is documented, and
the standard mitigation — sample periodically, track with CLOCK_MONOTONIC
between samples — is deliberately not built until a measurement asks for it.

End to end, guest-smoke now prints the wall clock it sees and rejects anything
before 2020; run under --clock-source ptp against MinIO it reports PTP time.

The script also has a qemu mode using ptp_kvm, written but not exercised here
(it needs cloud-image-utils). It is the weaker of the two: the clock behind
ptp_kvm *is* the host's, so skew is ~0 and it cannot distinguish reading the
PHC from reading CLOCK_REALTIME. Neither mode emulates Nitro; that needs
QEMU >= 9.1's nitro-enclave machine and an EIF, noted in the roadmap with
attestation.

Roadmap also records the follow-on this leaves open: the Object Lock retention
deadline still uses SystemTime::now(), and a COMPLIANCE deadline computed from
a wrong clock cannot be corrected afterwards by anyone.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
wasi:random/random is what a guest builds keys, nonces and session identifiers
from. It came from wasmtime-wasi's default, the kernel's getrandom(2). Inside
an enclave that pool is NSM-seeded — but nothing in the path says so and
nothing fails if it isn't, which is too thin for the one interface where a
silent downgrade cannot be detected afterwards.

Bytes now come from /dev/nsm: the same device that signs attestation
documents, reached with the same ioctl. Straight from the device on every call,
with no DRBG in between, so the claim is "the NSM produced these" and there is
no software generator to reason about. The cost is real and documented: the
device answers 256 bytes per call, so a larger request is a loop of ioctls.

crates/s3fs-host/src/nsm.rs is the device layer, and deliberately generic — only
`NsmDevice::request` knows about the ioctl, everything above deals in CBOR
payloads. Attestation in M8 is the same call with a different payload, so that
work now starts from a working device rather than from scratch.

Two details that would otherwise cost an afternoon each, both from reading the
kernel's uapi/linux/nsm.h rather than guessing: the raw ioctl requires
CAP_SYS_ADMIN, so an unprivileged process gets EPERM from a device that exists
and is readable; and response.len is in/out, capacity going in and bytes
written coming back.

GuestRandom panics when the source fails, which looks inconsistent beside the
clock serving a stale reading — so the module says why. A stale timestamp is
still a timestamp and a caller can notice. Predictable bytes handed to a guest
that believes them random are indistinguishable from good ones at the point of
use; the guest builds a key and nothing downstream can ever tell. Stopping is
the only safe answer. A fallback to kernel entropy is logged at error, not
warning, for the same reason.

wasi:random/insecure keeps wasmtime-wasi's generator — making a
deliberately-not-cryptographic interface cost a device round trip would be
perverse — but its seed is drawn from the source once so it is not
deterministic across runs.

--clock-check becomes --self-check, now covering entropy too: it draws bytes and
applies the crude checks that catch how emulators and misconfigured drivers
actually fail — all zeros, a constant byte, two identical draws, a stuck byte
histogram. It touches no storage, which is what makes it the right thing to run
inside an emulator.

Verified end to end against MinIO: guest-smoke draws two 32-byte values through
wasi:random and rejects them if identical or zero; consecutive runs reported
f3bd77e4a331b205 and 71212eed82253f1c.

Also adds deploy/qemu-nitro/Dockerfile, which builds QEMU 9.2 with the
nitro-enclave machine. This is needed because no distro package works: Debian
sid ships QEMU 11 with the machine type present but `-device help` has no
virtio-nsm, since the NSM device is only compiled when libcbor and gnutls are
found at configure time. The image asserts virtio-nsm exists at build time
rather than letting a boot fail later for unrelated-looking reasons. Confirmed
before building anything that QEMU's virtio-nsm implements GetRandom, not only
attestation — hw/virtio/virtio-nsm.c handle_get_random, returning
{"GetRandom":{"random":<256 bytes>}}, which is what this decoder expects.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`/dev/nsm` exists nowhere but an enclave, so until now the device layer was
covered only by a fake and by reading the driver's ioctl definition. That is
thin for the interface a guest builds keys from: a wrong opcode or a
misread response length would pass every test in the tree.

deploy/qemu-nitro/ closes it. build-eif.sh assembles a real EIF — AWS's own
enclave kernel and init, plus a static musl nsm-selftest in the ramdisk — and
run-selftest.sh boots it on QEMU's nitro-enclave machine and greps the console.
The guest draws from the emulated NSM and reports a sample, a byte histogram
and the per-call cost. It passes, at roughly 62us per 64-byte device call, and
the sample differs across boots.

Four things had to line up, none of which announces itself when missing:

  - QEMU with virtio-nsm, only compiled when libcbor and gnutls are present at
    configure time, so no distro package has it. The Dockerfile fails the build
    rather than the run if the device is absent.
  - A vhost-user vsock backend; the machine has no built-in one.
  - An answer to init's boot heartbeat. It writes 0xB7 to the parent on vsock
    port 9000 and waits, with no timeout, so an unanswered heartbeat looks
    exactly like a broken image. vhost-device-vsock's unix-socket backend
    silently drops it — init dials CID 3, the Nitro parent convention, and that
    backend serves only the host CID — hence --forward-cid and a real AF_VSOCK
    listener.
  - Mountpoints inside rootfs/, since init binds /rootfs onto itself and mounts
    the pseudo-filesystems there. Missing ones abort with
    `mount: /dev: No such file or directory`, which reads like a bootstrap fault.

The NSM device layer moves to its own crate, nitro-nsm, so it can link into a
small static binary for an enclave image without dragging in wasmtime and the
AWS SDK. M8's attestation reuses the same ioctl, and eif_build already produces
genuine PCR0/1/2, so the harness extends to attestation rather than waiting for
hardware.

Also fixes two clippy lints new in the 1.97 toolchain the harness needed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The guest could read a Merkle-anchored filesystem, a trusted clock and NSM
entropy, but nothing could reach it: the runtime ran a command guest to
completion and exited. This adds the other half — a `wasi:http/proxy` guest
that answers requests, with `--mode serve`.

The guest receives a parsed request and returns a response. It never sees a
socket, a connection or a TLS record. That split is the whole reason to
terminate TLS inside the enclave later rather than in front of it: a reverse
proxy on the parent instance would read every request in the clear, and the
parent is the party an enclave exists to exclude.

Egress is refused rather than merely unused. `wasmtime-wasi-http`'s
`default-send-request` feature is off, which turns `send_request` from a
defaulted method into one this crate is required to write — so `EgressPolicy`
denies it explicitly, with a test, instead of the answer depending on which
crate features happened to be enabled.

`GuestEnvironment` now builds the per-instance state for both modes. A command
guest needs one; an HTTP guest needs a fresh one per request, since a Store
cannot be reused and reusing one would leak a guest's resource table into the
next caller's request. Both must produce identical environments, so both go
through one place rather than each assembling a WasiCtxBuilder and drifting.

Concurrency defaults to one request at a time. `Fs` is safe to share — handles
under a parking_lot mutex, transaction state under a tokio one — but the guest
is not necessarily safe to run twice over the same data: SQLite on WASI has to
hold `locking_mode=EXCLUSIVE` because WASI has no fcntl, so two instances would
each believe they had the database alone. Raising it is safe for a guest with
no cross-request state, and the flag says so.

examples/guest-http is deliberately stateful. `/counter` reads, increments and
writes a file, so a second request only returns 2 if the first request's write
was committed and the next instance read it back — a stateless handler would
pass even if every request got an empty store.

Verified live against MinIO, not just in tests: the counter reached 3 over
three curl requests, a POST/GET round-tripped a file, the environment policy
delivered DEPLOYMENT while withholding all AWS_ and S3FS_ variables, and after
killing and restarting the process the counter resumed at 4 with the file
intact — commits reached the store and the anchor chain was picked back up.

CI gains a job that builds the component and runs the dispatch path over the
in-memory backend, so this is covered on every push without Docker.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The device layer could ask the NSM for random bytes. This adds the other
request that matters: Attestation, which returns a COSE_Sign1 signed by AWS
carrying the enclave's PCRs and up to three caller-supplied fields. It is what
lets a remote client establish something TLS cannot — that the code
terminating its connection is the code it expected.

Absent request fields are omitted rather than sent as null. The device parses
each present key by pulling a byte string out of it, so `"nonce": null` makes
it reject the whole request as InvalidOperation, naming no field. That cost is
paid once here and recorded in a test.

Verification is a separate crate. A client checking a document has no
/dev/nsm, is frequently not Linux, and must not need a device layer to check a
signature — so nitro-attestation depends on neither nitro-nsm nor anything
else in the workspace. It carries the AWS Nitro root, pinned by its published
SHA-256 so an edit to that file fails the build rather than quietly moving the
trust anchor.

The tests build genuinely signed documents rather than hand-written blobs: a
real P-384 chain from rcgen, a real ES384 COSE_Sign1. A verifier that skipped
the signature, the chain, the validity window or the root pin passes the happy
path and fails the rest, which is why the rest are there — a tampered payload,
a document from another root, a leaf borrowed from another chain, an expired
chain, an empty cabundle, a replayed nonce, and a certificate binding that
does not match.

`Trust` distinguishes chain-verified from self-signed because the QEMU harness
cannot produce the former: the emulator signs with a key it generated. A
document from it is structurally real and internally consistent while proving
nothing about AWS hardware, and the two must never be confused.

`nitro-attest` puts the whole check in one command. It keeps the certificate
the server presented, asks for a document quoting a fresh nonce, verifies it,
and then checks that user_data binds *that* certificate. Without the last
step a valid document proves only that some enclave exists somewhere — a proxy
could fetch a real one and serve it over its own TLS session.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The runtime now serves HTTPS itself, with a key generated inside the enclave,
and hashes that certificate into the attestation document. A client can then
establish something TLS alone cannot: the session it holds terminates in an
enclave running a specific image — not in the EC2 instance hosting it, and not
in a proxy the operator controls.

Terminating on the parent and forwarding plaintext would be far simpler and
would concede exactly the thing an enclave exists to prevent.

user_data follows nitriding's layout: two multihash-prefixed SHA-256 digests,
the TLS leaf and the guest component. The certificate hash ties the connection
to the document; the guest hash says which application was behind it. The
prefixes are what let the format grow later without every existing verifier
silently misreading it.

/enclave/ is reserved for the runtime and checked before the guest ever sees a
request. A guest able to answer under that prefix could serve any attestation
it liked. The nonce is required rather than optional — a document without one
cannot be shown to be fresh, and serving one on demand invites precisely the
replay the nonce exists to stop. Documents are sent no-store for the same
reason.

Attestation is requested once at startup, so an NSM that will not attest fails
the boot rather than the first client request, and asking for attestation
without TLS is refused outright instead of promising a binding it cannot have.

serve_tls.rs is the milestone's real test: a client opens a genuine TLS
connection, fetches a document quoting a nonce it chose, verifies it with the
production verifier against a pinned root, and checks that user_data names the
certificate from the handshake. It also checks the failures — a substituted
certificate breaks the binding, an earlier document does not satisfy a later
nonce, and /enclave/* never reaches the guest.

Two things the tests found rather than the other way round. Guest responses
carry no content-length, so hyper chunks them; nitro-attest refused chunked
bodies and would have failed against this very server, and now decodes them.
And the TLS key's randomness deserved an explicit argument rather than a
shrug: inside an enclave the kernel pool is NSM-seeded and nothing else, which
the boot log shows, and the image sets S3FS_RANDOM_SOURCE=nsm so that is
checked rather than assumed. The reasoning is in serve/tls.rs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
An enclave has no NIC. Everything above this — the TLS listener, the S3
client, the ACME client that comes next — assumes a socket works, and until
now none of them could have. Its only channel out is AF_VSOCK to the parent.

gvproxy turns that channel into an interface: it terminates the enclave's
ethernet frames on the parent and provides a gateway, DHCP and DNS at
192.168.127.1. This is what nitriding and ArkLabs both do, and the reason is
that ordinary networking code then works unmodified. The alternative — a vsock
connector threaded through the AWS SDK and an ACME client — is two bespoke
transports to write and maintain against libraries that do not expect them.

Networking comes up before the mount, because mounting reaches S3: a runtime
that mounted first would fail with an S3 error naming the wrong cause.
Readiness is a TCP connection to the gateway rather than a sleep, since DHCP
takes an unpredictable moment and a fixed delay is either too short on a slow
boot or wasted on a fast one. The timeout message names the gvproxy command
the parent is missing, because "connection refused" thirty seconds into a boot
with no shell is not a diagnosis.

gvforwarder ships inside the image rather than being fetched at boot, so PCR0
covers it. A forwarder streamed in later would be code sitting on every packet
that the attestation says nothing about.

deploy/parent/ is the other half, and it is written down because getting it
wrong looks exactly like a broken enclave: it boots, and then answers nothing.
Inbound :443 is forwarded through gvproxy's API, which also carries ACME's
TLS-ALPN-01 challenge, so no second port is needed.

What the parent can see is worth being precise about. It carries ciphertext:
TLS to S3 and KMS is established inside the enclave, and inbound HTTPS is
terminated inside it. What the parent does learn is metadata, and it can
refuse to carry anything — neither is new, since it already decides whether
the enclave runs at all. What is not safe is trusting its DNS, which is why
every outbound connection validates certificates.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A self-signed certificate satisfies any client that verifies attestation — the
binding proves more than a CA signature can. It does not satisfy a browser,
which cannot check attestations and simply refuses. So the enclave gets a real
certificate, and the private key still never leaves it.

TLS-ALPN-01, so the challenge arrives on port 443 — the port the parent
already forwards. HTTP-01 would need port 80 forwarded as well; DNS-01 would
put DNS credentials inside the enclave, a standing secret able to mint
certificates for the whole zone. The ClientHello decides per connection which
configuration to use, so validation and service share one port.

The cache is the security-relevant part. rustls-acme hands it the ACME account
key and the certificate's private key as PEM, so both are sealed with
AES-256-GCM under a new HKDF label beside the block and root-signing keys, and
written as ordinary objects. The parent stores ciphertext. They are
deliberately *not* in the filesystem: the guest's preopen is the filesystem
root, and a guest holding the TLS private key could impersonate the enclave to
every client. The object key is the AEAD's associated data, so a sealed blob
cannot be moved from the account slot to the certificate slot by anyone who
can write to the bucket, and the directory URL is part of the key so a staging
certificate can never be served in production.

Caching is not an optimisation. Let's Encrypt allows five duplicate
certificates per week; an enclave re-issuing on every boot would exhaust that
and be unable to serve.

The attestation binding had to become dynamic. A certificate issued
asynchronously and replaced on renewal is not the one captured at startup, and
attesting a certificate that is no longer being presented would break the
binding for every client at the moment it mattered. EnclaveEndpoints now reads
a shared slot the cache publishes to, and answers 503 before any certificate
exists rather than serving a document that promises a binding it does not
have.

Verified: 446 tests pass, including sealing round-trips, cross-slot and
cross-key rejection, tamper detection, and that the leaf is taken from after
the private key rather than from it. A live run against a deliberately dead
directory confirms the runtime mounts, starts ACME, binds, serves, and retries
orders with backoff instead of dying on a CA that is not answering.

Not verified: a successful issuance. That needs a CA — Pebble or Let's Encrypt
— and a domain the CA can reach, neither of which exists here. The code path
from a signed order to a served certificate is written and unexercised.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Covers what the guest gets and what it deliberately cannot reach, the
concurrency default and the SQLite reason behind it, the attestation binding
and the one check that makes it worth anything, the two certificate modes and
what Let's Encrypt does and does not buy, and what the parent instance must be
running.

Says plainly what is not verified: ACME issuance has no CA to run against here,
and chain validation to the AWS root stays hardware-only.

M8's entry now reflects that attestation is built; what it still needs from
this layer is the public_key field for KMS Decrypt with Recipient.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`ServerConfig::builder()` resolves its provider from rustls's compiled-in
features and panics when more than one is present. Both are, in any build with
the `aws` feature: the AWS SDK brings rustls with `ring` while this crate asks
for `aws-lc-rs`, so there is no unambiguous default and every call panics.

This was a real startup crash, not a test artefact. `enclave-runtime` defaults
to `--tls self-signed`, so the first thing it does in its default serving
configuration is build a ServerConfig — and it would have died there. It
surfaced as two flaky-looking unit tests, which passed under `-p s3fs-host`
and failed under `--workspace`, because feature unification is what decides
whether both providers end up linked.

Naming the provider fixes it at all three call sites and also states the
intent: the whole image stays on one implementation of these primitives rather
than shipping two.

Verified against the path that would have crashed: the runtime generates a
certificate, serves the guest over real TLS, and the SHA-256 openssl reports
for the presented certificate matches the one logged and bound into user_data,
byte for byte. Workspace suite: 455 passing, stable across repeated runs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three things land together because each was useless without the others: a Nix
build that produces the same PCR0 twice, a QEMU harness that boots that image
with a real network, and the Packer and OpenTofu to put it on an instance.

PCR0 is a digest of the enclave image, pinned by a KMS key policy and checked
by clients. Its whole value is that someone else can rebuild and get the same
number — and the Dockerfile it replaces ran `apt-get update` over
`debian:bookworm-slim`, so it produced a different one every week and attested
to nothing anybody could reproduce.

`nix build .#eif --rebuild` now passes. Getting there was four things, none of
which announced itself:

  - `cpio --reproducible`. The newc header stores each file's inode and device
    numbers, so even the two-file *bootstrap* ramdisk came out different every
    build, taking PCR1 with it.
  - fixed uid/gid/mtime, and members sorted with LC_ALL=C, since cpio records
    the order it is handed.
  - `faketime` around eif_build, which stamps wall-clock BuildTime into the
    image's metadata. No PCR covers metadata, so the measurements were already
    reproducible — but the file differed by 15 bytes, and people compare
    artifacts by hashing them.
  - a pinned Cargo.lock for eif_build, which upstream gitignores. Letting the
    dependency set of the tool that *computes PCR0* float would mean the
    measurement came from something slightly different each time.

The end-to-end asserts three things that had never been checked together: the
filesystem mounts over gvproxy and a counter advances across requests, so
writes reach MinIO and come back; user_data binds the certificate from the
connection's own handshake; and the attested PCR0 equals the one the build
printed. Getting a boot that far surfaced four real bugs, each invisible until
the one before it was fixed:

  - the ramdisk had no /etc, so writing resolv.conf failed;
  - gvforwarder's stderr went to /dev/null, so the next three failures were
    only visible as "the gateway never answered";
  - it shells out to a DHCP client, which the image did not have — and
    nixpkgs patches busybox to find its lease script inside its own store
    path, so copying out just bin/busybox got a lease and applied none of it;
  - no CA bundle, so the AWS SDK panicked with "no CA certificates found"
    before making a request. That one matters in production too.

One thing I had assumed and was wrong about: QEMU's emulated NSM does not sign
attestation documents. Its source says "we don't actually sign the data, so we
use -1 as the 'alg' value", and -1 is not a COSE algorithm identifier — which
is why a strict parser rejected it. So the harness runs `nitro-attest
--unsigned-emulator`, which checks the contents the runtime put there and says
plainly on every run that no signature and no chain were verified. The earlier
claim that QEMU "signs with a key it generated" is corrected in the docs.

The parent is built with Packer over Amazon Linux 2023 rather than Nix, and
that is not inconsistency: the parent is the party the enclave excludes, PCR0
covers the enclave and not its host, so reproducing the host buys no security
property. No Docker on it either — that is only needed by `nitro-cli
build-enclave`. gvproxy and the gvforwarder inside the EIF come from one
nixpkgs package built static, so both ends of the vsock share a pin.

Verified here: reproducible PCR0, the full e2e, the entropy self-test on the
Nix-built image, 455 tests, clippy clean, `packer validate` and `tofu
validate`. Not verified: there are no AWS credentials on this machine, so no
AMI was built and nothing was applied.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
It was carved out as the development CLI — explicit flags, and a guest
environment that stays empty unless asked — against enclave-runtime as the
deployment target. That split stopped holding some time ago:

  - nothing ran it. Zero references in CI; its only tests were six over clap
    parsing, so it compiled and nothing more.
  - it could not do half the product. No serve, TLS, network or attestation
    flags at all, against enclave-runtime's eleven.
  - it went unused through the feature it existed to support. Every local
    MinIO test in the HTTPS milestone — the counter, the certificate binding,
    the restart durability check — ran through enclave-runtime, because the
    runner could not serve. The designated development tool was not reached
    for once while developing.

enclave-runtime is a strict superset: --no-inherit-env --guest-path X is
exactly what the runner did. A second binary that is a worse copy of the
first, and untested, is how a codebase acquires a tool nobody trusts.

The one thing it contributed was a safe default for an uncurated environment.
Inheritance is right inside an enclave, where the environment is the attested
deployment configuration, and wrong on a developer's machine, where it
forwards whatever is in your shell — the AWS_/S3FS_ denylist withholds this
runtime's own credentials, not a GITHUB_TOKEN. So --no-inherit-env is now the
documented local path, --guest-env gained the S3FS_GUEST_ENV binding every
other setting already had, and both are tested.

Nothing was lost with the deleted tests: the credential regression guard lives
in s3fs-host's env.rs, where it belongs — it is a property of GuestEnvPolicy
rather than of any particular CLI.

Demonstrated rather than assumed, against MinIO with guest-smoke reporting
what it could see: the default inherits the shell, --no-inherit-env withholds
it, --guest-env NAME passes exactly that one through, an explicit value
overrides, and no AWS_/S3FS_ variable reaches the guest in any configuration.
The first version of that check probed a variable the guest never prints and
"passed" three cases while measuring nothing.

M9's `s3fs-runner gc --keep-roots N` now has no home, and deliberately does
not move into enclave-runtime: that binary ships inside the enclave image, so
every flag added to it changes PCR0. The roadmap records that garbage
collection needs a small unattested CLI first.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
s3fs-host began as `wasi:filesystem` over the block store — a library any host
could embed, deliberately knowing nothing about AWS or enclaves. That is no
longer what it was. It had accumulated the NSM, the vsock tap device, TLS
termination, ACME and the attestation endpoints, none of which mean anything
outside an enclave, and after s3fs-runner was deleted it had exactly one
consumer. The crate boundary had stopped describing a real separation and was
only charging rent for it.

The layer that genuinely is reusable is s3fs-core, which still has no wasmtime
dependency at all. That boundary stays.

It becomes a library beside the binary rather than collapsing into main.rs,
because integration tests cannot import a binary-only crate — and
tests/serve_tls.rs, where a client checks that the attestation document binds
the certificate from its own handshake, is the test this design exists to
pass. Both suites still pass from their new home, 7 and 7.

The `aws` and `serve` features are gone. They existed so a host embedding only
the filesystem bindings could skip the AWS SDK and hyper; with one consumer
that always wanted both, they bought nothing and cost
`--features aws,s3fs-host/serve` on every build, test and clippy invocation,
in CI and by hand. There is now nothing to select: `cargo test --workspace`
runs all 451 tests. Two CI jobs that differed only by those flags were
byte-identical afterwards, so one is gone.

Also renames the WIT world from `s3fs-host` to `enclave-runtime`, since it is
generated into that crate and the old name would have been the last place the
dead one survived by accident rather than as a note.

Verified beyond compilation: the flake still builds a reproducible EIF from
the merged crate, and both serve suites run against a real component.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- Implemented `describe_pcr` and `extend_pcr` methods in the `Nsm` trait.
- Updated `FakeNsm` to support PCR operations, allowing for testing of PCR extension logic.
- Introduced `Pcr` struct to represent PCR state, including locked status and value.
- Added utility functions for encoding and decoding PCR requests and responses.

refactor(nitro-attestation): transition from single PCR to multiple PCRs

- Changed `Expectations` struct to use a map for multiple PCR values instead of a single `pcr0`.
- Updated related methods and tests to accommodate the new PCR structure.
- Enhanced the `expect` method to validate multiple PCRs against the attestation document.

fix(s3fs-core): separate filesystem creation from mounting

- Modified `Fs::mount` to fail with `FsError::NoFilesystem` if the store is empty, preventing accidental filesystem creation.
- Introduced `Fs::create` method to explicitly create a filesystem in an empty store.
- Updated tests to reflect the new behavior of filesystem mounting and creation.

chore(deployment): add deployment configuration for S3FS

- Created `deployment.nix` to define the S3 bucket configuration for the enclave image.
- Ensured that the configuration is baked into the image to maintain PCR integrity.

test(e2e): validate state origin across restarts

- Enhanced end-to-end tests to verify that a second boot resumes the filesystem state rather than creating a new one.
- Added checks to ensure that the state root remains consistent across boots.

docs(roadmap): update roadmap with recent changes

- Clarified the status of binding attestation into the anchor and closing the cold-mount gap.
- Updated descriptions to reflect the current implementation and future goals.
`Store::open` no longer formats an empty store, so `Fs::mount` returns
`NoFilesystem` where it used to hand back a fresh filesystem. Four harnesses
were still calling it against empty backends and had been failing since:

  tests/serve_guest.rs       6 tests, including the TLS-adjacent dispatch path
  tests/serve_tls.rs         7 tests, one of which the crate docs call
                             "the test this whole design exists to pass"
  tests/minio_integration.rs the shared mount helper, so all 10
  benches/fs_hot_paths.rs    would have panicked on first run

None of it showed up in `cargo test --workspace`: every one of those tests is
`#[ignore]`d behind Docker or a built wasm component, so the default run stayed
green while the suites that actually exercise the serving path did not run at
all. That is the more useful finding than the fix.

Harnesses over a fresh backend now say `Fs::create`. The MinIO helper both
creates and mounts, because several of its tests remount a bucket to prove the
state survived.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Object Lock COMPLIANCE protects a *version*, not the name it lives under. A
`DeleteObject` with no version id writes a delete marker and **succeeds**; an
ordinary `GetObject` then answers `NoSuchKey` while the retained version sits
underneath, genuinely undeletable. Verified against MinIO:

    $ aws s3api delete-object --bucket lockprobe --key roots/0
    { "DeleteMarker": true, "VersionId": "71eff5c3-…" }
    $ aws s3api get-object --bucket lockprobe --key roots/0 -
    NoSuchKey: The specified key does not exist.
    $ aws s3api delete-object --bucket lockprobe --key roots/0 --version-id afc3c2fd-…
    InvalidRequest: Object is WORM protected and cannot be overwritten

So "hide the tip" needed no lie from S3, just a delete nobody was allowed to
refuse. Two consequences, both load-bearing:

  - The boot machine treats a missing state-origin receipt as authorisation to
    create a filesystem. Mark the receipt *and* the sealed key and it reads
    `(None, None)`, takes the genesis path, and serves a second filesystem
    beside the one it was hiding — the exact substitution receipts exist to
    refuse, with every signature along the way valid.
  - Worse, `If-None-Match: *` tests the current version, so a marker makes the
    genesis lease free again. The "am I the first here?" check answers yes.

`Backend::get_retained_blob` finds the version through `ListObjectVersions` and
reads it by `versionId`, which a marker cannot conceal because the version
cannot be removed. Used by `boot::maybe_get` and by `RootStore::exists`/`load`,
which also closes the rollback half — hiding the newest root was the same
trick. It needs `s3:ListBucketVersions`, and errors rather than answering
"absent" without it, so a missing permission refuses the mount instead of
silently starting over.

`MemoryBackend` described itself as modelling "a non-versioned bucket", which
S3 will not let you have with Object Lock at all — so the fake was *stronger*
than the real thing and three tests proved a property that did not hold. It now
models delete markers, and those tests assert what is true: retention makes an
object impossible to destroy, not impossible to hide.

Found because `object_lock_makes_a_root_record_undeletable` was failing. It was
asserting the comfortable claim.

Also fixes a dead error arm in `boot::genesis`: it matched `FsError::Conflict`
on a failed conditional PUT, but both backends return `AlreadyExists`, so a
genesis race reported the generic message instead of the intended one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`pwrite` and `set_size` both refuse a handle opened without write. `set_times`
did not, so `OpenFlags::write` meant "may change the contents" rather than "may
change the file" — and the exemption was silent, which is the part that would
have cost someone an afternoon.

It matters more here than in an ordinary filesystem because setting a timestamp
is a *commit*: `flush` publishes a new signed root record under Object Lock
retention. A read-only handle that can advance the anchor chain is not read-only
in any sense worth the name.

POSIX would settle this by ownership rather than by the descriptor's open mode —
`futimens` on an `O_RDONLY` fd is legal for the owner. There is no owner here:
`Attrs` carries `mode` but no uid, so "are you allowed?" has no answer other
than what the handle was opened for. `set_times_at` stays ungated for the same
reason, having no handle to ask.

Both new tests were confirmed to fail without the guard.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`ServeHandle::handle` acquires the concurrency permit, moves the `Store` into a
spawned task, and blocks on a oneshot the guest fulfils by calling
`response-outparam::set`. The sender lives inside that `Store`. So a guest that
neither returns nor sets a response never drops the sender, the receiver never
resolves, the permit is never released — and at the default of one request in
flight, one such request is the whole server for the life of the process.

Nothing bounded it: no fuel, no epoch interruption, no timeout, and both engines
built with a bare `Config::new()`.

Two mechanisms now, because neither covers the other:

  - Epoch interruption, with a 250ms ticker and a per-request deadline. This is
    the only thing that can stop wasm that never yields.
  - A timeout on the response head, which bounds the wait regardless of *why*
    the guest is quiet — including a host call that never returns, which the
    epoch cannot see because no wasm is executing.

**The deadline must fire before the timeout, and that ordering is the whole
mechanism.** The first version of this gave the deadline slack *past* the
timeout so the error would read "took too long" rather than an opaque trap.
That inverted it: `task.abort()` fired first against a guest nothing could
interrupt, `abort` needs an await point a spinning guest never reaches, and the
task span on after the request was abandoned — hanging runtime shutdown instead
of the request. The tidier message cost the entire guarantee, and the test
caught it by hanging.

`examples/guest-http` gains `GET /hang`, deliberately rather than as a test
fixture: a guest that never answers is the one behaviour a host cannot provoke
from outside, so without it the watchdog has nothing to be tested against.
Reverting the deadline ordering makes the new test hang, which is how it was
confirmed to test anything.

`S3FS_REQUEST_TIMEOUT_SECS` is baked into the image like every other setting, so
PCR0 covers how long a guest may take.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`S3FS_GUEST_LIFETIME=session` keeps one `Store` and one `Proxy` alive across
requests, so a guest is a process rather than a handler: it holds its linear
memory, and with it any cache, connection or state it wants. Default stays
`request`, because a single session shared by every caller is worse than
per-request isolation — it becomes the right default once a session belongs to
one authenticated client.

The store still moves into the spawned task, because the guest keeps streaming a
body after the response head has gone out and a borrowed store cannot outlive
the function. What changes is that the task hands it back afterwards. The lock
is an `OwnedMutexGuard` for exactly that reason: it is `'static`, so it can move
into the task and carry the session with it.

`Option` lives *inside* the lock rather than beside it, so "checked out" and
"absent" stay distinguishable. A request holding the lock and finding `None`
knows to instantiate; one that finds the lock held knows to wait. Two separate
pieces of state could not tell those apart, and a request that guessed wrong
would build a second instance over the same filesystem.

Three things a session must get right, all of which have tests:

  - **Instantiate once.** `Instance` is a `Copy` index with no `Drop` and
    `Store.instances` only grows, so instantiating per request on a reused store
    leaks a component instance and its linear memory every time — bounded only
    by wasmtime's default of 10,000. This is the one way to leak unboundedly
    here; the resource table is a slab with a free list and does not.
  - **Reset the epoch deadline per request**, or request N+1 inherits whatever
    N left of it.
  - **Discard rather than carry.** A trap leaves the instance in an unknown
    state, and resources the guest failed to drop stay addressable by the next
    request on that session. Either sets the slot to `None` so the next checkout
    rebuilds.

A session is serialised by construction: the lock is held for the whole call
including the body, and two concurrent requests cannot share one linear memory
without wasm threads. That is a property of the model, not a limitation of this
implementation.

`examples/guest-http` gains `GET /memory`, a counter that touches no storage.
`/counter` proves the filesystem carried state between requests; this proves the
*instance* did, and under a per-request guest it can only ever return 1 — which
is what makes the three session tests discriminate. Forcing them to
`GuestLifetime::Request` fails all three, which is how that was confirmed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… be lied to

Client certificates were never requested — `with_no_client_auth()` was the only
server config in the crate — and `ServeHandle::handle` had no parameter for a
connection identity even if one had existed. The service closure was also built
*before* the TLS handshake, so nothing per-connection could reach a request.

mTLS carries the identity. There is no CA and no chain to build: a client
presents a self-signed certificate and the identity *is* that certificate. What
makes it mean anything is not who vouched for it but that the handshake proved
possession of the matching private key, which is why `AnyClientCertificate`
accepts any certificate while **delegating** `verify_tls13_signature` to the
provider. Answering `Ok` there instead would let anyone claim anyone's identity
by copying a certificate, which is public.

Client auth is offered, not required. A browser, `curl` or a health check
presents nothing, completes the handshake, and arrives without an identity —
which the guest sees as anonymous. Requiring it would break every ordinary
client and the ACME challenge with them.

`serve_connection<S>` is extracted so the closure is built after the handshake,
which is the only point at which a peer certificate exists. The three arms —
fixed TLS, ACME TLS, plaintext — now share one body; all three stream types
satisfy the bound, so it monomorphises.

The header is where this fails silently, so: `insert`, never `append` —
`append` leaves a client's own copies and the guest's `fields.get()` returns a
list, so `[0]` would be theirs. The `None` arm must `remove()`, or an
unauthenticated connection passes its own value through. And it cannot be done
in `is_forbidden_header`, which runs *inside* `new_incoming_request` — after
injection — so it could only delete the header, silently.

`rustls-acme` hardcodes `with_no_client_auth()` in `default_rustls_config()`, so
the ACME path builds its own config from `state.resolver()`. Without that the
identity plumbing would be present, correct, and dead in the one mode that has
a real certificate. The challenge config keeps no client auth: the CA validating
TLS-ALPN-01 presents none, and the one handshake that must not fail is not the
place to ask.

`GET /whoami` on the example guest echoes what the runtime told it. Three tests
drive real handshakes rather than passing a value to `handle` directly: a
certificate becomes the expected digest, two clients are two identities, and a
client without one is still served.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The identity was SHA-256 of the whole certificate. A certificate carries a
validity period and a serial number that change on every renewal, so a client
who rotated a certificate around the *same key* became a different client:

    public key inside each certificate   9bc0f6efe5aba683…  9bc0f6efe5aba683…
    whole certificate                    713e0ef105a01b33…  172f092635c5bb6b…

Survivable while an identity only picks a session — a renewed client loses
in-memory state and no more. Not survivable once an identity derives a
filesystem, which is the plan's next phase but one: a routine, scheduled
renewal would point a client at empty storage while their data stayed encrypted
under keys nothing would ever derive again, in a bucket that retains objects for
ten years. Silent, irreversible, and triggered by doing the correct thing.

Now SHA-256 over the `SubjectPublicKeyInfo` — the construction certificate
pinning uses (RFC 7469) — so the identity is the key, not the paperwork around
it.

The reason for the original was that extracting the key means parsing DER an
attacker supplies, and there was no parser in the image. There is: rustls has
already run `webpki` over this exact certificate to verify the handshake
signature, and `webpki::EndEntityCert` exposes `subject_public_key_info()`.
Naming it in `Cargo.toml` adds no code to the build — `cargo tree` shows the
same two `rustls-webpki` entries as before — only the ability to ask rather
than to re-derive a parser we would have to trust.

`from_certificate` now returns `Option`, because a certificate that will not
parse must produce no identity rather than a fabricated one. It should be
unreachable: the signature check that admitted the client parsed it first.

Two tests guard the property, one at each level. The unit test builds two
certificates on one key with different serials and asserts one identity —
reverting to the whole-certificate digest fails it, which is how it was
confirmed to test anything. The TLS test does the same over a real handshake.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
f90b1e0 added a guest that keeps its linear memory across requests, on the
reasoning that a transaction cosigner is a long-running process. That was
the wrong model, and this undoes it.

A `Store` is what separates two wasm instances. Give two clients one
instance and the separation stops being a runtime guarantee and becomes
guest code: a bug that confuses two clients is no longer a leak of one, it
is a total compromise. A trap poisons state everyone shares. Leaked
resource handles stay addressable by whoever calls next. All three of those
go away when the instance dies with the request, and none of them can be
engineered out of a shared one.

So the boundary goes back between clients, where it belongs. Per-client
state — nonce ledger, policy, pending transactions — lives in the anchored
filesystem, which is the only place it can, and that is the right kind of
constraint: attested, hash-chained, and it survives a restart. The linear
memory it would otherwise have sat in was none of those things.

Removed: GuestLifetime and --guest-lifetime, in_session and the session
mutex, the leaked-resource rebuild policy, and State::resources_settled,
which only meant anything for a store that outlived a call.

What replaces the eager start is narrower and honest about it:
verify_instantiates builds one instance at boot and drops it, so a guest
that traps in its initialiser stops the enclave instead of answering 500 to
every request while looking healthy.

`no_two_requests_share_an_instance` is the invariant as a test — /memory
must answer 1 for every request, forever. It is what fails if anyone
reintroduces sharing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two properties that were true in the code and nowhere else.

The per-request invariant, which the README never stated: no two requests
share an instance, and that is where the boundary between two clients
actually lives. A `Store` separates two wasm instances; give two clients
one instance and the separation becomes guest code instead.

And the client identity from c84edde, which shipped undocumented. With a
fresh instance per request it is the only thing tying a request to its
client's state in the store — there is no session and nothing else to key
on.

The e2e's fourth leg asserted /memory *increments*, which was the check
for the resident model f90b1e0 added and 6906cf9 removed. Inverted: two
requests must both answer 1, which makes the leg a proof of the isolation
invariant over real TLS rather than a leftover.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Written to settle an argument with a number rather than an intuition:
pooling instances per client saves exactly one thing — the instantiation
in the middle of every request — and the trade is only worth making if
that is a large fraction of a request.

    instantiate           24 µs
    mount                 51 µs
    dispatch/trivial      88 µs
    dispatch/committing  753 µs

Against MemoryBackend, so these are engine costs with no network. Two
things fall out. Instantiation is ~3% of a request that does one real
transaction, and against S3 the commit grows by round trips while
instantiation does not — so that share only shrinks.

And `mount` is the larger number, which is the one that matters if
filesystems are ever per client: it derives key material, reads a signed
root record, verifies it, and opens the object set. The 51 µs here is the
CPU floor; in production it is that plus two or three S3 round trips.
Whatever a per-client design keeps warm, the mounted filesystem is the
thing worth keeping, not the wasm instance.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Until now the only MasterKeySource was StaticKey, which "seals" by writing
the secret into the blob behind a marker saying it did not. The secret
arrived as S3FS_MASTER_KEY, so the parent instance held it before S3 did —
and from those 32 bytes HKDF derives the block AEAD key, the Ed25519
root-signing seed and the runtime-seal key. Whoever had them could decrypt
the filesystem, forge root records, and decrypt the enclave's TLS private
key to impersonate it to every client.

KmsAttestedKey mints the secret with GenerateDataKey and recovers it with
Decrypt, both carrying a Recipient: a fresh RSA key generated for that one
call, named in an NSM attestation document. KMS checks the document
against the key policy, encrypts its answer to that key, and omits
Plaintext entirely. The parent still proxies the HTTPS request, because
the enclave has no network of its own; what it cannot do is read the
answer.

The enforcement is the key policy, not this code. With
kms:RecipientAttestation:PCR0 pinned to the approved image, a wrong
enclave does not get a refused mount — it gets no key at all. That is the
difference between an enclave checking itself and something outside it
doing the checking.

CiphertextForRecipient is a CMS EnvelopedData. The `cms` crate parses it;
keys::recipient decides what is acceptable, and accepts exactly one shape:
one RSA-OAEP-SHA256 recipient, AES-256-CBC content, id-data inside
id-envelopedData. A general-purpose parser is fine, a general-purpose
policy on the path that unwraps the master key is not. Refusing a second
recipient is the sharp one — it would mean somebody besides this enclave
could open the content, which is what Recipient exists to prevent.

Storage splits in two. SSM holds the ciphertext and nothing else; the
roots bucket holds a pointer naming the parameter, the CMK, the encryption
context, and sha256 of the ciphertext. That keeps boot.rs's state machine
intact — it still decides genesis from resume by what is present — and
earns something: the state-origin receipt already commits to
sha256(sealed), so it now attests which CMK and which context this
filesystem was created under. A repointed pointer is caught by the
receipt, a swapped parameter by the hash. A host can delete the parameter
and stop the enclave booting; it cannot make it boot wrong.

--master-key-source has no default, and the refusals are the point. A
plaintext key under kms is an error, not an ignored setting: a production
image that still carried S3FS_MASTER_KEY would work perfectly while
quietly using a key the parent holds, and no ordering of "both were given"
is safe to guess at. Absent CiphertextForRecipient is likewise an error
and never a fallback — its absence means the request reached KMS without
an attestation, which would have returned the key in the clear.

The QEMU image stays on static and always will: its emulated NSM does not
sign, and KMS will not accept an unsigned document. PCR0 differs between
the two images, so a client can tell which it is talking to.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
aruokhai and others added 29 commits September 6, 2026 09:42
An ordinary request is approved by an assertion bound to
`{method, path, query, sha256(body)}`, and the body hash is the part that makes
the approval mean *this* operation rather than any operation at that route. A
stream has no such hash to offer: its body is the message sequence, and none of
it exists when the channel opens.

The wrong resolution is to drop the binding for streams and call the result
authenticated. That would hand out an approval good for anything the client
later chose to send — precisely the substitution the hash exists to prevent,
reintroduced through a weaker sibling.

So a stream is opened by a different approval, issued by a different endpoint,
that authorizes opening a channel on one route and nothing else. Enrollment
tokens already have this shape: they create a tenant and do nothing else.
Anything inside the stream that asks the enclave to sign carries its own fresh,
single-use assertion bound to that message. The prototype carries the fields
and refuses without them; verifying them needs a runtime hook that does not
exist yet, and it should be designed before a real key depends on it.

Two independent things stop the weaker approval leaking into the strong path,
and one of them is not a check at all: `BodyBinding::Unbound` never equals
`BodyBinding::Exact(_)`, so no comparison can be satisfied by the wrong kind
and no `if` can be forgotten. The second is `x-enclave-stream`, a runtime-owned
header stripped before the guest sees it — the header alone cannot talk the
runtime out of hashing a body, and an unbound approval alone cannot be spent on
a request that arrived as an ordinary one. Deliberately not keyed on
`content-type: application/grpc` or a path prefix: the content type is
application data the guest also reads, and a prefix would bake a service name
into a runtime that knows nothing about the guest's routes.

`Gate::verify` now consumes the challenge before touching the body, because the
binding is what decides whether there is a body to read. One consequence worth
naming: an oversized body burns its challenge before being refused, where
before the refusal came first. That is the safer direction — otherwise a caller
can probe the size limit without ever spending an approval.

Streams are refused outright when no gate is configured. Anonymous callers
share one lock, so an anonymous stream would hold it for its whole life and
starve every other anonymous caller — one client silencing the runtime by
connecting. An unauthenticated deployment is a development arrangement and
should not learn to depend on a shape that only works once identities exist.

docs/STREAMING.md states what a stream-open approval buys, what a stolen one
buys, what an open stream costs the tenant holding it, and what the runtime
does not promise: no retries, no resumption, no deduplication, and no
exactly-once execution.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…his repo

The guest hand-writes gRPC framing because tonic does not build for
wasm32-wasip2. Testing it with framing written here too would only show the two
halves agree with each other, which is the one thing that was never in doubt.

So tonic drives it: a real bidirectional stream over TLS and ALPN-negotiated
HTTP/2, decoded by prost on the client side, with `Status` and `Code` reported
by tonic rather than by anything in this repository. A refusal arrives as
`PermissionDenied` with its message intact, which is worth asserting through a
real client precisely because the head is 200 — a client reading only the
status line would call a refusal a success.

The attestation is verified from the *initial metadata* before the client sends
its first message, and the order is the point rather than an artefact of how
the test reads. Verifying afterwards would establish that the enclave was
listening only after you had already told it something. The document binds the
nonce the client chose and the leaf from its own handshake, so tonic connecting
through an already-open connection — rather than through tonic's own transport
— is what lets the test hold that exact certificate.

`grpc-timeout` gets a test of its own for what the runtime does *not* do with
it. It is neither honoured nor stripped: the runtime's deadlines are its own so
a client cannot lengthen them by asking, and interpreting a client's deadline
belongs to the guest, which is the only party that knows what its work is
worth.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The two are the same shape — an approval that authorizes creating one thing and
nothing else — so the reader who has just understood one is the reader best
placed to understand the other. Says plainly that an open channel is not
standing permission to sign, next to the paragraph explaining why an approval
for one transaction must not authorize another.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Authorization had two shapes that had grown apart. An ordinary request was
approved by an assertion bound to `sha256(body)`; a stream was approved by a
deliberately weaker assertion bound to nothing, because a stream's body is its
message sequence and does not exist when the channel opens. The weaker path
existed only because the stronger one could not express a stream.

Both are replaced by one. A person authenticates with their passkey and the
runtime issues a short-lived, single-use bearer token for one **interaction** —
one request and its response, or one stream until it closes or reaches its
deadline. The token names a method, a path and a query.

## What this gives up

It does not name the body. A token issued for `POST /sign` authorizes whatever
bytes follow, so a client compromised between the approval and the request can
substitute the payload, and the runtime will not notice because it is no longer
looking. That is a deliberate policy change, not an oversight, and the
documentation now says so where it used to promise the opposite. What survives:
a person still authenticates per interaction, tenant isolation is unchanged,
challenges and tokens are single-use and die on restart, and nothing about the
rate limiting moved.

`TokenStore` mirrors `ChallengeStore`, which was already the right shape:
bounded, swept on issue, refusing rather than evicting when full, and removing
an entry before deciding anything about it — so two callers racing one token
have exactly one winner and a token offered for the wrong route is spent by the
attempt. Entries are keyed by `sha256(token)`, so holding the store is not
holding a token.

The exchange is necessarily three trips. A passkey is a challenge-response, so
the assertion cannot exist until the challenge has been answered, and the token
cannot exist until the assertion has been checked. There were only two before
because the assertion rode on the operation itself; separating them is what
lets one approval cover a stream.

## Attestation moves to where a client can act on it

The runtime attests `/auth/` exchanges and leaves the rest alone. The challenge
request is the one that goes first on a connection nothing has vouched for, and
it is the right thing to send there: it carries the route but no approval, so a
party that intercepted it learns what is intended and holds nothing it can act
on. Its response identifies the enclave, and the client pins that certificate
for the two trips that follow — TLS proves the peer holds its private key, which
a second document would not add to. That halves the NSM signatures an
interaction costs, on a device that is the runtime's throughput ceiling.

The header is now *stripped* from unattested responses rather than left alone.
It was only ever guest-proof because every response overwrote it, and a guest
that could set it would be handing clients a document under the runtime's name.

`/enclave/config` is gone with it, and `--attestation` with it: a flag that
turns verification off is one that can be turned off by whoever starts the
process, which inside an enclave is the party the enclave exists to exclude.

## The client verifies, in the right order

`passkey-client` checked a document by parsing it, which does no signature or
chain check at all — an interceptor could mint one naming its own certificate.
It now verifies the chain, the age, the nonce, the certificate binding and the
measurements, and it refuses to run without `--pcr0`.

`--guest` is not accepted in its place. PCR0 is measured by the hypervisor and
locked; the guest hash is part of `user_data`, which the runtime being attested
chose — so an attacker's own enclave signs a genuine document claiming whatever
guest hash was going to be checked for. `--guest` is a second check on a trusted
measurement, not a substitute for one.

Both trips that carry a credential are pinned before a byte goes out: the
assertion, and the credential being registered. Sending either to an unpinned
connection meant handing it to whoever answered, and checking the response
afterwards would only prove the theft succeeded.

## An interaction now has a deadline

Nothing watched a call after its head went out — the `JoinHandle` was dropped,
which detaches. The epoch watchdog cannot help: a guest parked in a host call
executes no wasm and never reaches an epoch check. So a peer could open a stream,
say nothing, and hold its tenant's only slot until the process ended.

A wall clock works precisely where the epoch does not, because such a guest *is*
at an await point and `abort` reaches it. `--max-interaction-secs` bounds how
long an interaction may run; `--interaction-token-ttl-secs` bounds how long an
approval may sit unspent. This also restores a bound the pool lost: eviction
skips busy slots, so long-lived streams were pinning tenants above `max_tenants`
indefinitely.

`--max-request-body-bytes` is deleted. The runtime stopped buffering bodies when
it stopped hashing them, and a cap that applied to ordinary requests but not
streams would need the distinction this change removes. The guest bounds
messages, and the docs say so rather than pointing at a flag that could not.

## Also

A gRPC stream ending mid-frame now reports `INVALID_ARGUMENT` rather than OK.
The half-close looked clean at the HTTP layer; the stream did not end cleanly,
and the trailers are the only place that difference can be said.

`ClientMsg` loses its `challenge_id` and `assertion` fields. They promised a
per-message check nothing performed — and nothing *can* perform yet: verifying
an assertion needs the credential store, the issued challenge and the
relying-party configuration, all of which live in the runtime, and a guest is
given standard WASI with no host function to ask through. Fields that look like
a credential and are never checked read as a guarantee. Closing that gap needs a
new import, and it should be designed before a real key depends on it.

`serve_auth` ran in no CI job at all, which was not a gap to leave standing now
that it is the suite for the whole authorization model. It is wired in, with the
client binary built ahead of it so the test that drives the real thing does not
silently skip.

Not verified: the QEMU end-to-end harness. `run-e2e.sh` changed in three places
and only its shell syntax has been checked.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The guest component shipped inside the enclave image, so PCR0 covered it and
every guest change was an image rebuild. It is now fetched at boot from a key
in the roots bucket, `S3FS_GUEST_OBJECT`, and measured before anything asks for
a key: PCR16 is extended once with `sha256(component)` and locked. The image
still names *where* the guest comes from, under PCR0; PCR16 says *what*
arrived. A key policy pinning both still releases the key to exactly this
runtime running exactly this guest.

The object does not have to be trusted. A parent that substitutes it gets an
enclave with a different PCR16, which a policy pinning the approved guest
releases nothing to and a client pinning that guest refuses.

## The order is the point

`measure_guest` checks each step against what it should have produced, because
each is a place a wrong answer would pass silently into every attestation: the
register must start unlocked and zero, extending must give exactly what a client
computes from the component, and the device — not the lock call's return value
— must then report it locked with that value. A register that is extended but
never locked appears in no attestation document, so no key policy could see it.

`Pair::read` is the first thing `boot` does, before the store is read or KMS is
asked, and an unlocked or never-extended PCR16 refuses to boot. The bytes
measured are the bytes compiled and served. `lock_pcr` is new on the `Nsm`
trait; a successful `LockPCR` answers with the bare string rather than a map, so
it has its own decoder.

## PCR16 means something only beside PCR0

The runtime writes PCR16, so a runtime someone else wrote can put any value
there. PCR0 says the runtime that wrote it is yours; PCR16 says which
application that runtime loaded, which PCR0 can no longer say. Clients now
require both. `passkey-client` refuses to run without `--pcr0` and one of
`--pcr16` or `--guest`, and refuses the last two if they name different guests.
`nitro-attest` enforces the same through clap, and gains `--measure`, which
prints a component's PCR16 using the same `guest_pcr` the verifier checks with.
`nix build .#guest-release` uses it to emit `guest-pcr16.json`, so the value in
a key policy and the value a client pins cannot disagree.

A document without PCR16 fails an expectation on it rather than skipping the
check, because that is what a runtime that never locked the register produces.

## The key policy decides; boot records

`--authorise-successor` is gone. It extended PCR31 to name the next image and
never locked it, and since documents carry only locked registers, no document
it produced on hardware could have carried the register the successor was
checked against. It passed here only because the test NSMs put unlocked
registers into every document. They now keep registers the way the device
does: 0–15 locked from boot, 16 free until measured, only locked ones attested.

Nothing replaces the handoff, because the boot machine stops having an opinion
about which code may hold the state. An enclave KMS released the key to was
allowed by the policy, and a second copy of that decision could only disagree
with the real one. `verify_origin` now checks that the receipt names the state
that was loaded, not who wrote it.

What boot adds is a record. The first time a runtime and guest hold a state, the
enclave attests a *pair record* carrying both registers against the
`state_root` and stores it under Object Lock at a key derived from
`sha256(PCR0 ‖ PCR16)`. That boot reports `Upgrade`, which replaces
`Migration`; later boots of the pair find the record and write nothing. An
existing record is verified, not merely found, since anyone who can write the
roots bucket can derive its key and plant something there first. Two first
boots racing leave one record and both boot; junk that wins the race is
refused.

It records which pairs have held the state, not in what order: returning to an
earlier pair finds that pair's record and adds nothing. The order approvals
were given in lives in CloudTrail's record of key-policy edits. Both conditions
belong on `kms:GenerateDataKey` as well as `kms:Decrypt` — an unconditioned
`GenerateDataKey` hands the parent a data key in the clear, which it can plant
as a new filesystem's key before an enclave ever runs genesis.

## Also

Stored receipts are verified against their chain as of the moment they were
signed, not now. An attestation chain's certificates are short-lived next to a
filesystem, so checking against the current time would make a filesystem
unmountable by the passage of time. The timestamp is inside the signed payload,
so a document stamped outside its chain's validity still fails.

The flake's source filter excludes `examples/`, so a guest-only edit no longer
changes the runtime's store path and with it the image's PCR0.

`boot_origin` ran in no CI job, and neither did the unit tests of the
`passkey-client` and `nitro-attest` binaries, which `--lib` does not reach. All
three are wired in.

The QEMU harness builds `.#guest-release`, uploads it to MinIO, and checks the
attested PCR16 against `guest-pcr16.json`. A seventh leg replaces the object
and restarts: the enclave measures the substitute, records it as an upgrade,
and is refused by a client pinning the approved guest. The emulator has no KMS,
so the half where a substituted guest gets no key still needs hardware. Leg 5b
and the tofu `verify` output now describe interaction tokens and `/auth/`,
which the previous commit introduced.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- Added `existing_tenant_root` function to ensure tenant directories exist before scheduling tasks.
- Introduced background task handling in the `guest-http` example, including task enqueueing, status checking, cancellation, and forgetting tasks.
- Updated the `ServeConfig` to support background tasks in various test files.
- Created a new `BACKGROUND_TASKS.md` documentation file detailing the background task functionality and configuration.
- Modified the Dockerfile and Nix deployment files to include options for enabling background tasks.
- Added WIT interface definitions for task management, including enqueue, status, cancel, and forget operations.
- Updated dependencies in `Cargo.lock` and `Cargo.toml` for compatibility with new features.
Three strands landed together because they overlap in the same files —
`routes.rs`, `run-e2e.sh`, `flake.nix` and the README each carry more than one —
and separating them would have meant hunk-level surgery rather than a clean
division.

## Registration is open

Registering a passkey no longer needs an invitation. `/auth/register/options`
takes no token, `auth/enrollment.rs` is gone along with `--enrollment-token` and
the image configuration that carried one, and both the client and the QEMU
harness enrol with nothing but a URL.

What registration grants is unchanged and deliberately narrow: a new, empty
tenant. `PendingRegistration::join` is never `Some`, so a registration cannot
reach an existing tenant or add a passkey to somebody else's. The WebAuthn
challenge and verification are untouched, the tenant is created only after
verification succeeds, and the client still checks the enclave's attestation
before it trusts the response. An older client still sending `enrollment_token`
is refused rather than quietly admitted, because the request now denies unknown
fields.

What this gives up is admission control over resource creation. A tenant is a
directory the enclave keeps and a slot against `--max-tenants`, and anyone who
can open a connection can consume both. The rate limiter on challenge requests
is removed in the same breath, so `/auth/request/options` is unbounded per
credential as well. Whatever needs to bound either now has to sit in front of
the enclave.

## CI runs what it ships

The recipes moved out of inline YAML into `scripts/`, because there was nowhere
else for them to live: `run-e2e.sh` had grown its own copy of the MinIO setup,
and no CI job could be exercised without pushing. Each job is now one script
that behaves the same on a laptop, and the harness shares `minio-up.sh` instead
of keeping a second recipe that had to agree with the first.

Running every one of them found three defects that only CI would otherwise have
shown. `cargo bench --workspace -- --output-format bencher` hands that flag to
each crate's libtest harness, which rejects it and aborts the run after eight
minutes of compiling. The SQLite job had been broken since bbbee4d — first for a
missing `--master-key-source`, then for an NSM requirement no hosted runner can
meet — so it is out of CI, and the script it left behind says why. And
`wasi-sdk.sh` fetched without `curl -f`, which writes an error page into a
tarball and fails two steps later complaining about the archive format.

## The enclave schedules work, and that is checked where it runs

`run-e2e.sh` gains two legs: eight files written, read, overwritten and re-read
through the gate, and work scheduled by one approved interaction that runs later
with nothing signed at the moment it does. Three tests cover the same ground
quickly — the real client binary enqueueing and polling to completion, a guest
that genuinely fails retrying until the occurrence is failed, and a recurring
task firing a second time through the running scheduler.

Verified: rustfmt, clippy with `-D warnings`, 705 tests passing and 80 skipped,
and the emulated end-to-end passing all eight legs with Alice and Bob
registering without tokens into tenants that cannot see each other.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…pened

A guest had no way to reach its user between requests. Background tasks made
that gap plain: work finishes inside the enclave and nobody is told. A guest
still cannot reach the network — `serve::EgressPolicy` refuses every outgoing
request it makes, deliberately — so the runtime holds the Firebase credential
and sends on its behalf, and the guest gets a host import instead of a socket.

## A wake signal carries nothing

There is no title, no body and no `notification` object, and that absence is
the feature. An FCM payload crosses the parent instance — the party an enclave
exists to exclude — and then Google, so anything placed in it is disclosed to
both. What travels is an opaque category, an optional tenant-local reference
and a schema version. The app wakes and fetches the detail over its own
attested connection, where the parent is shut out again.

That is not a setting. A flag would be something a deployment turns on and a
guest then uses, at which point it stops being a property.

It is also the right shape mechanically: a `notification` block is rendered by
the OS without the app running, so a wake carrying one would show text *and*
fail to wake anything.

## Enrolling is interactive-only; waking is not

The one place this departs from `enclave:tasks/queue`, where every mutation
requires an interactive invocation. Enrolling a device grants standing ability
to reach somebody, so background work must not be able to do it, nor to silence
its owner by un-enrolling. But raising a wake *from* background work is the
whole point — a task that finishes at three in the morning telling its owner to
come and look. It grants nothing; it spends an enrolment an interactive call
already made.

The import is registered unconditionally, like tasks': authority is denied at
call time by the `Option` in `State` being `None`, never by omitting the import,
which would turn a policy refusal into a boot failure.

## What it gives up

Delivery is best effort and unacknowledged. Repeat wakes for one category
coalesce, a full queue drops and counts rather than blocking, and queued wakes
do not survive a restart — in `tasks` the record *is* the work, whereas here it
would be a stale pointer to work that already happened. A guest is never told
which devices accepted, because that would hand it the user's device inventory
and their online pattern, and it must be resilient to non-delivery regardless.

A parent that steals the service account can ring doorbells: send wakes to
tokens it obtains elsewhere, and delay or drop ours, which it could already do
because it carries every packet. It cannot read any tenant's data or the device
tokens, which live under a key KMS releases only against a matching PCR0 and
PCR16 — and a forged wake means nothing, because a wake carries no state.

## Zero new crates in the measured image

hyper's client, `webpki-roots` and `aws-lc-rs`'s RSA signing are already in the
production closure. The manifest comment claiming the runtime "never calls out
over HTTP" and that `hyper/client` was dev-only is corrected in passing: it has
been false for as long as `aws-config`'s `rustls` feature has been enabled.

The client's roots are compiled in rather than read from a file, because
`net.rs` notes the parent answers DNS — certificate validation is the only
thing stopping it pointing fcm.googleapis.com at itself. There is deliberately
no custom verifier on this path.

## Not proven here

Nothing in the suite reaches Google. The wire is covered by a transport double,
and leg 6b/8 of the QEMU harness proves the whole path inside an emulated
enclave against a local stub: a finished background task wakes an enrolled
device with a message carrying no content. Whether Google accepts it is
untested, the same way `guest_io::cloudwatch` is honest about its own
credential path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`crates/enclave-runtime/` sat beside `s3fs-core`, `nitro-attestation` and
`nitro-nsm` as though it were a peer. It is not: those three are libraries this
thing is built from, and it is the thing. The layout said otherwise to anyone
opening the repository for the first time, and where a directory sits is the
first claim a repository makes about what matters in it.

So `crates/enclave-runtime/` becomes `runtime/`. Pure `git mv`, plus the path
edits a move forces:

- the workspace member in `Cargo.toml`
- path dependencies inside `runtime/Cargo.toml`, which now reach *into*
  `../crates/` rather than sideways within it
- one directory level fewer between the crate and the repository root, so every
  `.join("../../examples/...")` in a test or bench loses a level, and
  `wit-bindgen`'s `path: "../../wit"` becomes `path: "../wit"`
- two README links

## What was deliberately left alone

`runtime/tests/serve_guest.rs` still contains `"../../tenants/…"` and
`"../../../../../../etc/passwd"`. Those are not paths on this disk. They are
the strings a guest sends to see whether the runtime lets it climb out of its
tenant directory, and the test asserts it does not. Rewriting them along with
the real paths would have quietly weakened the test that exists to catch
exactly that — which is why they are called out here rather than left for
whoever runs the next sweep.

No behaviour changes. The crate keeps its name, so nothing that depends on
`enclave-runtime` notices.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ying

A client could not be developed against the emulator, only pointed at it. Every
client meeting a QEMU enclave had to pass `--unsigned-emulator`, and that flag
skips the signature, the certificate chain and the validity windows — most of
what a client does, and precisely the part that has to be right before it meets
hardware. What was exercised locally was the contents check; what shipped was
code nothing had run.

## The emulator now signs, and says so

QEMU's NSM implements the protocol but not the signing. Its source says as much
— *"we don't actually sign the data, so we use -1 as the 'alg' value"* — and -1
is not a COSE algorithm identifier, so the envelope is unverifiable while the
contents are perfectly good.

`testing::CosigningNsm` wraps the device rather than replacing it: `attest`
asks the emulator for a document, parses it, and re-signs that same payload with
a chain minted at boot. Contents stay the device's answers — the real PCR0 of
the image, the PCR16 the runtime measured its guest into, the caller's nonce.
Only `certificate` and `cabundle` are replaced, and they have to be, or the
document would name a key that did not sign it. Nothing else is intercepted:
entropy and every PCR operation go straight through, so a client comparing PCR0
against what `nix build` printed is still comparing against the device's report.

It does not make a document mean anything. The key is minted inside an image its
operator controls, so a verified document here says *"this image said so"* where
on hardware it says *"a Nitro enclave with this measurement said so"*. Two
things keep that from being mistaken for the real property: the root is fresh at
every boot, so there is nothing long-lived to paste into an app; and it is a
`testing` build flag, so the production binary has no `--cosign-attestations` to
be talked into signing its own attestations.

## The clients needed no changes, which exposed one bug

`--trust-root` *without* `--allow-untrusted-root` already yields
`Trust::ChainVerified` and already makes `--pcr0` and `--pcr16` mandatory. So
the local path is the production path, flag for flag.

That reachability made a latent bug visible: `nitro-attest` printed "chain
verified to the AWS Nitro root" for any pinned root. Harmless while the only way
to reach `ChainVerified` was the default root; a lie the moment a test root
could get there. It now names the file and says it is not AWS's.

## One bring-up, two consumers

`run-e2e.sh` was 930 lines of which two thirds was standing the stack up. That
half moves to `deploy/qemu-nitro/lib.sh`, and `dev-enclave.sh` is the second
consumer: same emulated machine, same ACME issuance from a real CA, same
attested boot, except it serves a component of your choosing and stays up.

    deploy/qemu-nitro/dev-enclave.sh --guest path/to/component.wasm

It prints the three values a client pins and waits. Because both scripts share
the bring-up, what a client is developed against is what CI checks.

All thirteen e2e legs now verify by signature and chain against a pinned root.
Legs 7 and 8 re-pin after each reboot, since every boot mints a new chain.

## Found by running it, not by reading it

The MinIO container carried no label, because `docker update` has no
`--label-add` — so cleanup left it running. `minio-up.sh` takes `MINIO_LABEL`
now. S3 bucket names followed the run's name, but they are baked into the image
and therefore into PCR0, so the enclave looked in `e2e-data` while the script
had made `dev-data`. A `grep | head` inside an assignment killed the harness
silently under `pipefail` before the line it wanted had been written.
Backticks in an unquoted heredoc executed. And `--name ""` would have pointed
`rm -rf` at `target/qemu-nitro` itself.

Preflight also refuses to start on untracked `.rs`, `.toml`, `.wit` or `.nix`
files: Nix flakes copy only what git tracks, so an unadded module is invisible
to the build and fails twenty minutes later as `file not found for module`,
from inside a sandbox. `nix build … | tail -3` used to hide that error behind
the derivation summary; failures now print the compiler's own words.

docs/DEV_ENCLAVE.md covers what is and is not real about it.
docs/CLIENT_INTEGRATION.md is for a team bringing an existing client across.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Every assertion from the Android app was refused. Credential Manager signs
`clientDataJSON` with the origin `android:apk-key-hash:<hash>`, never
`https://<rp id>`, and the relying party was built with that one web origin
and no way to add another.

`--webauthn-allowed-origin` (repeatable; `S3FS_WEBAUTHN_ALLOWED_ORIGINS`,
comma-separated) appends further origins through webauthn-rs's
`append_allowed_origin`. Each is still compared exactly. This does not widen
who can sign: Android lets an app claim that origin only after
`https://<rp id>/.well-known/assetlinks.json` lists it, so the origin names an
app the domain already vouched for.

An Android origin is checked at boot, not at the first refused login: the hash
must be 32 bytes of unpadded base64url. The likeliest wrong value is the
colon-separated hex fingerprint keytool and the Play Console print, which names
the right certificate in a form Android never sends, so the error says to
convert it. Empty entries are dropped so an image with none still boots, and
the flag without a relying party is refused rather than silently ignored.

`deployment.nix` gains `webauthnAllowedOrigins`. Like `rpId` it is measured by
PCR0: adding an app is a new image.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A platform authenticator creates a passkey only for an rp id whose domain publishes
assetlinks.json naming the app, and the app then claims `android:apk-key-hash:<hash>`. The emulator
image fixed its relying party at `enclave.test`, which publishes nothing, so no phone could register
against a dev enclave at all.

`eifQemu { rpId, allowedOrigins }` builds the emulator image for another relying party, exposed as
`lib.<system>.eifQemu` because flake outputs take no arguments, and `dev-enclave.sh --rp-id
--allowed-origin` builds and boots one. Both values are held to their shapes before they reach the
Nix expression. `passkey-client` is given the rp id, and the certificate is still `enclave.test`:
the two were only ever the same by default.

`packages.eif-qemu` is `eifQemu { }` and evaluates to the same image: the derivations differ only in
the flake's own source path, and the environment baked into the image is byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A failing `run-task` wrote its error only into the sealed task record, so the
console showed `status=Failed attempts=5` and nothing else. Finding the cause
meant making the guest print to stderr.

Every failed attempt is now logged at warn with the error it records, whether
it will retry or has reached `Failed`, and so is an attempt a restart
interrupted. The text is escaped (`?`) so a guest cannot forge log lines with
it; it is nothing the guest could not already print to its own output.

The recorded error also lost its causes: `to_string()` on an anyhow error is
its outermost context, so a trap or a deadline said what failed and not why. It
is now `{:#}`, still bounded to 512 characters.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The WIT said only that `task-id` "remains stable across retries". The runtime
passes `<id>:<generation>:<occurrence>` (`Task::run_id`), so a guest that
validates it as the id it enqueued refuses every run — which is how the
cosigner's watch failed on every wallet.

The format is now documented in both copies of `tasks.wit` and in
BACKGROUND_TASKS.md: what each part is, that the enqueued id is everything
before the first `:` (an id cannot contain one), and which to use for what.
Documented rather than split into separate arguments, which would break every
existing guest and need a new interface version.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`/auth/register/options` used webauthn-rs's defaults: `residentKey:
discouraged` and no `authenticatorAttachment`. On Android below 14, Play
services takes that literally and creates a non-discoverable security-key
credential outside Password Manager, which One Tap then reports as "Cannot
find a matching credential". Every client would hit it; the app was working
around it by rewriting the options itself.

Registration now asks for a discoverable platform passkey, through webauthn-rs's
constructor for exactly this (`residentKey: required`, `requireResidentKey`,
`authenticatorAttachment: platform`, user verification still required). Its
feature flag adds no code beyond that constructor. It also stops requesting
credProtect, which Android does not support.

Ruling out roaming security keys and cross-device registration costs nothing:
the clients are native apps, never a browser.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A guest has had no outbound network at all: `EgressPolicy::Denied` refused every `wasi:http`
request. That is still the default. But some guests cannot do their job from inside it — a wallet
cosigner that holds a pre-signed renewal has to reach the ASP it renews with, and the only
alternative was routing the round through a phone that may be off.

So a deployment may name origins — scheme, host and port, nothing else — and guests may send
requests to exactly those (`guestEgressOrigins`; `S3FS_GUEST_EGRESS_ORIGINS`,
`--guest-egress-origin`). Compared exactly: no subdomain, no other port, no plaintext to an HTTPS
origin, no wildcards or paths; a malformed entry refuses the boot. HTTPS is verified against the
same webpki roots the FCM client uses — now one shared `web_pki_client_config` rather than two
stores to audit — and requests go over HTTP/1.1. Anything else is refused as before, with the same
warning. Requests and background tasks get the policy alike, since renewing on a schedule is
exactly the work that happens with nobody connected.

The list is image environment, so PCR0 covers it: a client learns where a guest can send traffic
from the attestation that tells it what the guest is. It is set only when non-empty, so an image
that names no origins is the image it was.

Two settings beside it, for the same kind of guest: `backgroundTimeoutSecs`, because a round waits
on the ASP's schedule and the default is 30 seconds, and `guestEnv`, which reaches the guest because
it inherits the image environment minus `AWS_*` and `S3FS_*` — how a cosigner learns its ASP's
address. `eifQemu` takes all three.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ch a local service

Every dev-enclave start booted into genesis: MinIO kept its data inside a `--rm` container, so a
restart or a guest rebuild threw away every tenant, passkey and file a guest wrote — and a phone
app under test had to be wiped and onboarded again each time.

`--keep-store` puts MinIO's data in target/qemu-nitro/<name>-store, outside the run directory that
is cleared on every start, so the next start resumes it and a new guest boots as an upgrade of the
same store. Pebble mints a new CA every start while a kept store serves the certificate an earlier
one issued, so every root and intermediate seen is kept and `pebble-root.pem` carries all of them —
the served-chain check verifies whichever issued it. The store directory gets a stable `id` for
clients that keep state per store, since the trust root is new every boot. `--fresh` discards it,
behind the same guard RUNDIR has. Without the flag nothing changes.

`--guest-egress ORIGIN`, `--background-timeout SECS` and `--guest-env NAME=VALUE` build the image
with the new `eifQemu` settings, each held to a shape that cannot break out of the Nix expression.
The machine running the script is 192.168.127.254 from inside the enclave.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…speak first

A guest has no execution context between invocations, so it cannot own a socket
that outlives a call. And a counterparty with no passkey for a tenant can never
call in, because the gate is checked before the path and fails closed. Neither
party could reach the other by the means each already had.

`enclave:streams` closes that. A guest asks for a connection to an origin its
image allows, and every message the far side sends becomes one invocation of
`on-message` — exactly as a due task becomes one invocation of `run-task`. The
guest stays stateless; the runtime keeps the socket.

What reconnects after a drop, a timeout or a restart is a supervisor task in the
runtime, started from the records on disk before anything is served. Not sealed
guest state, which only says a connection should exist, and not a scheduler.

Three things that are easy to get wrong, and were:

  - The wire id names the tenant as well as the stream. A stream id is
    tenant-local by contract, so every tenant a guest serves opens one under the
    same name; a far side keeping one connection per id would have each customer
    close the last one's, and would answer down the wrong socket.

  - `send_direct` hands back the connection worker with the response. It is
    abort-on-drop, so returning the response alone ended the body the moment the
    call returned — invisible for a request/response exchange, fatal for a
    stream held open.

  - `open` wakes the supervisor with `notify_one`, not `notify_waiters`. The
    loop is not a registered waiter while it is collecting records, and a wakeup
    that landed in that window was simply lost.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… streams

Store: a session claims its transaction group in the root chain before it
seals a block, and again after any transaction that did not end in a root.
Nonce uniqueness no longer rests on probing host-deletable slabs, and of two
mounts at one tip only one ever encrypts. The orphan probe is gone.

Filesystem: a failed sync keeps its dirty buffers for the retry; a directory
can no longer be renamed into its own subtree.

Backend: retained reads follow every page of ListObjectVersions, so delete
markers cannot push the real version out of sight.

Deploy: the parent role may list versions, read by version and put
retention headers, which boot needs; the deny on PutObjectRetention bought
nothing under COMPLIANCE.

Runtime: dropping a guest state closes the files it left open; a stream
supervisor is replaced when its record changes under the same key; the SSE
framer accepts CRLF and CR line endings.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…tream reopens

Store: publishing a root reads the retained version back after the
conditional PUT, and loses unless that version is ours. S3 accepts
`If-None-Match: *` over a delete marker, so a host that hid the winner's
claim could hand the same transaction group — the same nonces, the same slab
names — to a second mount. Claims, commits and format all go through
publish, so all three are covered.

Backend: a retained read takes the oldest version by list order, not by
LastModified. S3 lists a key's versions newest first, and a timestamp tie
left the newer one in front.

Runtime: dropping a guest state releases the files it left open without
flushing them. The flush ran after the tenant's lock had passed on, and could
land on top of what the next request committed. A guest that did not finish
loses what it had not synced, as on a crash.

Streams: a record carries the generation of the open that made it, so a close
and reopen to the same origin between two supervisor passes gets a new
connection — not the old one, whose status the close had removed, leaving
every send refused. The stream suite now runs in the guest CI stage.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The README still described the filesystem this repository started as. It
now covers the runtime as built: what exists, a local quick start, the
architecture and trust boundaries, the encrypted filesystem, boot and key
release, serving and verification, passkeys and tenants, background work and
held connections, SQLite, configuration, testing, deployment, embedding the
storage engine directly, and the limits that remain.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Scheduler: a recurring task's second success is waited for rather than
inferred from its output file. The guest writes the file before it returns
and the queue records the success after, so stopping the worker between the
two left `attempts` at 1 — the failure that ended both CI runs of the e2e job
before MinIO or the emulated enclave had started.

MinIO: two mounts, a delete marker over the first one's claim, and the second
refused. On real S3 the damage without the read-back is worse than nonce
reuse: the loser's slabs overwrite the winner's committed ones, so the test's
proof is the winner's data surviving.

README: the CI table names the suites these land in.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…step

rustfmt had failed on egress.rs and notify/transport.rs since 909ebee and on
serve/http.rs since before it, which stopped the unit job at its first
command on every run of this branch. Behind it, clippy denied the redundant
parentheses in the stream registry's key type and `ServeHandle::streams`,
which nothing calls; it is gone rather than allowed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Docker Hub refuses `minio/minio` and `minio/mc` without a login and quay.io
has no tags, so the e2e job's storage stage failed before a test ran, and its
QEMU stage would have failed the same way. It had never got that far before.

scripts/minio-image.sh builds the pinned server release and the mc release
before it from their GitHub tags, into `enclave-runtime/minio` under the tag
the testcontainers module pins. The integration tests take it by name,
minio-up.sh runs server and mc from it, and the harness uploads guests with
it. It builds only when missing, in about five minutes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… lint job

rust-toolchain.toml named `stable`, which floats. CI denies every clippy
warning, and 1.98.0 arrived with `chunks_exact_to_as_chunks`: a branch that
passed on 1.97 failed with no change to it, while a workstation still on 1.97
saw nothing. Now CI and every checkout run the same clippy, and a newer
release is a commit of its own.

The one thing 1.98.0's clippy asks for: indirect blocks are split with
`as_chunks`. The length is checked first, so its remainder is always empty,
as `chunks_exact`'s was.

The Nix enclave build is untouched; flake.lock already pins its toolchain.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
It relays every store lookup through the Actions cache API. That API
rate-limited it, it answered Nix with HTTP 418, and Nix treated the refusals
as fatal: the enclave image build failed on "rate limit exceeded" with
nothing wrong in the build. It was also the source of the FlakeHub
authentication error in every run's log.

nixpkgs still comes from cache.nixos.org. What goes is only the reuse of this
repository's own derivations between runs, which are rebuilt each time.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ing it

`escapeShellArg` turns a path into a string with `toString`, which carries no
context: the EIF derivation named the flake source's copy of pebble/ca.pem
without depending on it. Nix upstream copies the whole source eagerly, so
the file happened to be there and the QEMU image built. Determinate Nix,
which CI installs, evaluates flakes lazily and never put it in the store, so
the build failed on `cp: cannot stat` — after warning that the derivation
referenced a store path "without a proper context".

Interpolated instead, the file is its own store path and an input. The
image's contents are unchanged: PCR0 is the same.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@aruokhai
aruokhai merged commit b8254f6 into main Sep 26, 2026
5 of 6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant