New Runtime - #1
Merged
Merged
Conversation
The previous engine mapped POSIX paths onto S3 keys directly. For an enclave, where the host is outside the trust boundary and storage must be assumed hostile, it had three fatal properties: no integrity checking of any kind, no anchor and so no way to detect a rollback, and mutation in place, which is incompatible with S3 Object Lock. This replaces the storage layer with a ZFS-style copy-on-write block store. Content lives in immutable AEAD-encrypted blocks packed into slab objects, and the whole filesystem hangs off one signed, hash-chained root record written with a conditional PUT under Object Lock COMPLIANCE retention. Verifying that record transitively verifies every byte beneath it. Against an adversary holding full write access to the buckets: they can make the filesystem unreadable; they cannot make it read wrong, and they cannot make it read old. Storage layer (crates/s3fs-core/src/store, crypto): - BlkPtr addresses a byte range inside a slab rather than a content hash, so a commit costs a couple of PUTs however many blocks changed. Content addressing would have cost one PUT per 128 KiB. - Indirect blocks are variable length with trailing holes trimmed, so a two-block file pays a 256-byte indirect block rather than a full record. - Directories are a separator-indexed B+tree. Extendible hashing was planned, but its cheap-doubling property needs many slots sharing one bucket block, and the AEAD binds every block to its index — deliberately, since that is what stops an attacker relocating blocks. Sorted iteration comes free. - Transaction group numbers are never reused. The non-obvious threat is a crash between writing slabs and publishing a root: the tip still names txg T, so a naive remount resumes at T+1 and repeats every nonce already spent on the orphaned slabs. next_safe_txg probes past orphans before the first commit. - Losing the conditional PUT for a sequence number poisons the mount rather than retrying, for the same reason. - Root signatures verify against the key we derived, never the key the record names, and the record's sequence is checked against the key it was read from. Filesystem layer: - rename is atomic: one directory-entry move in one commit. The async copy-then-delete queue and its redirect machinery are gone, and a directory rename no longer costs O(entries) copies. - stat cannot go stale; link counts and all three timestamps are real. - Identity is (objid, gen), stable across mounts, so is-same-object and metadata-hash are exact rather than a hash of an ETag. - Hard links, and unlink-while-open, both previously documented as impossible. - Snapshots: every root record already is one. Also fixes a credential leak in s3fs-runner, which forwarded the entire host environment — including AWS_SECRET_ACCESS_KEY — into the guest. Deletes mpu.rs, rename.rs, flusher.rs, buffer/ and inode/ (~3.4k lines): the MPU state machine, UploadPartCopy planning, race-three lookup and tiered part schedule were all compensating for the path-to-key model. 333 unit tests, plus MinIO integration coverage of Object Lock retention, remount, slab packing economics, tamper detection and the rollback floor. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The storage engine is done; the enclave is not. Three things stand between this and running for real, and each has a decision in it worth writing down rather than rediscovering. M8, attestation and key release. --master-key is a development seam: a secret on the command line is visible to the parent instance, which is exactly the party the enclave exists to exclude. Notes the two seams that do not exist yet — AwsS3BackendConfig has no http_client field, so the SDK cannot be pointed at a vsock proxy, and its static credentials carry no expiry, so an attested session token dies mid-run. M9, garbage collection. Deliberately not shipped rather than shipped approximately: a slab is dead when no root you intend to keep references it, and "intend to keep" is policy, not a fact about the data. Getting it wrong deletes live data. Records the mark-and-sweep design and the cheaper stopgaps. M10, the enclave-runtime crate. s3fs-runner stays the dev CLI so the local loop keeps working without NSM or vsock; a new crate is the deployment target. The guest ships inside the EIF so PCR0 covers it, which is what makes the key policy attest to the code that will actually read the data — proving that some code with the right image hash is running is not the same claim. Flags add_wasi_minus_filesystem for extraction into a shared crate: it is a hand-copied clone of a wasmtime-wasi function and will drift silently on any bump, and a second binary would double that. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The thing that actually runs inside a Nitro Enclave did not exist. s3fs-runner is shaped as a developer CLI — every setting a required flag, the guest's environment empty unless each variable is named — and an enclave has no shell to type flags at and no operator to enumerate variables. It gets whatever the image baked in, and the guest is expected to read it. enclave-runtime is the deployment target: environment-first configuration with matching flags, the guest loaded from a known path (/enclave/guest.wasm), and the guest inheriting this process's environment minus anything under AWS_ or S3FS_. That one prefix rule covers the master key and the whole credential set without anyone enumerating them, and it keeps working when a new one is added. Explicit --guest-env still wins, because explicit is a decision and inheritance is an accident waiting to happen. s3fs-host holds what both binaries need. It absorbs s3fs-wasmtime, which after the rewire had exactly one consumer and pinned wasmtime separately; mount sits behind an `aws` feature so the wasi:filesystem bindings stay usable over any Backend without dragging in the SDK. The hand-copied clone of add_to_linker_with_options_async now exists once rather than being duplicated into a second binary. Exit codes are no longer collapsed. A guest calling exit(3) reached us as a trap carrying I32Exit and became exit 1, losing the code; it now propagates, and a trap (70) is distinguished from a runtime failure before the guest started (71) — an integrity or rollback failure at mount is a security event, not a bug in the guest. Two bugs found by running it rather than by reasoning about it: - A clap bool with `env` demands exactly true/false, so ENV S3FS_FORCE_PATH_STYLE=1 — the obvious thing to write in a Dockerfile — hard-failed at startup. Inside an enclave that is expensive to diagnose. Both binaries now accept 1/0, true/false, yes/no, on/off. - crates/s3fs-core/src/backend/aws.rs had `.copy_object()AwsS3Backend` committed in b11f95b. It was introduced between the last green test run and `git add -A` and swept in; the tree at b11f95b does not compile. examples/guest-smoke is a new std-only guest: no wasi-sdk, so CI can run it on every push, unlike guest-fsdemo with its bundled SQLite. It exercises write, in-place patch, rename, read_dir, asserts no AWS_/S3FS_ variable reached it, and counts its own runs from a file it wrote — so a second run printing "run 2" is proof that commits reached the store and the anchor chain was picked back up, not that the filesystem was quietly reformatted. The new guest-e2e CI job runs exactly that, twice. It is also the only real drift guard for the linker list, since wasmtime's Linker cannot be enumerated and no unit test can prove the list complete. Also removes the minio-integration service container, which was malformed and unused — the tests start their own via testcontainers, so the job was passing vacuously. Verified end to end against MinIO with Object Lock on the roots bucket: three consecutive runs, root_seq 0 -> 9 -> 17, distinct Merkle roots, ledger persisted across processes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A real C database driving the filesystem is a far harsher test than anything hand-written. SQLite does page-granular random reads and writes, creates and unlinks a rollback journal on every transaction, fsyncs where it genuinely needs durability, rewrites the whole file during VACUUM — and then tells us whether the bytes came back correct, via PRAGMA integrity_check. examples/guest-sqlite covers DDL, bulk and batched inserts, point and indexed selects, updates, upserts, cascading deletes, rollback and nested savepoints, all five constraint kinds, joins cross-checked against correlated subqueries, aggregates, recursive CTEs, window functions, blobs up to 2 MiB verified byte for byte, unicode and NULL handling, triggers, views, ALTER TABLE including RENAME COLUMN, REINDEX, VACUUM, ANALYZE, and integrity_check both inline and after closing and reopening the database. Every phase is timed, so it doubles as the benchmark. Running it found a real trap, and it is not in this filesystem. SQLite locates a directory for temporary databases by probing candidates with access(2), which WASI does not provide — so every candidate is rejected and VACUUM fails with a bare "disk I/O error" that says nothing about a missing syscall. Setting temp_store_directory does not help: that pragma validates the path the same way and reports "not a writable directory". The fix is temp_store=MEMORY, now documented in COMPATIBILITY.md alongside the locking_mode=EXCLUSIVE that WASI's lack of fcntl already required. journal_mode is left at DELETE rather than the MEMORY the older demo used. The rollback journal is then a real sidecar file with a full create/extend/truncate/ unlink lifecycle, which exercises parts of the filesystem a memory journal skips entirely. Verified against MinIO with Object Lock on the roots bucket, at 2 000 and 20 000 accounts (60 000 rows total). Everything passes; integrity_check is clean before and after reopen. CI's guest-build job only ever compiled a guest and never ran one. It is replaced by guest-sqlite, which builds and runs this workload end to end. Also fixes .gitignore: `/target` anchored only the workspace root, so the example guests' build directories were tracked, and 13 build artifacts went into the previous commit. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Extends the workload to the operations the first pass skipped, chosen for the filesystem behaviour they provoke rather than for SQL coverage: - DROP index/view/trigger/table, and a check that a dropped trigger stops firing. Dropping frees pages, which is what gives the following VACUUM something to reclaim — so drops now run last. - ATTACH: a second database file open alongside the first, a join spanning both, and integrity_check on the attached one. - Incremental blob I/O via sqlite3_blob_open, writing 4 KiB chunks in reverse order so pages are dirtied out of sequence. The rusqlite `blob` feature was enabled in the first pass and never used, which was sloppy. - WITHOUT ROWID tables, STORED and VIRTUAL generated columns, partial indexes and expression indexes — all of which have different on-disk shapes. - INSERT OR REPLACE / OR IGNORE, BEGIN IMMEDIATE / EXCLUSIVE, date/time and scalar functions, TEMP tables. - JSON, FTS5 and R-Tree, probed rather than assumed. All three are present in the bundled build and all three pass; FTS5 indexes 2 000 documents and then rebuilds every shadow table. The first run of this found the most useful number in the whole exercise. A phase that should have taken half a second took 38.9 seconds, because I had left 500 inserts in autocommit. Each statement outside a transaction is its own commit, and a commit here is a transaction group: slab PUTs plus a signed root record published with a conditional PUT. So the phase was measuring commit latency, not the table shape. Rather than only fixing it, there is now a phase that measures it deliberately: 25 inserts in autocommit take 1 796 ms (14 rows/s); the same 25 in one transaction take 69 ms (364 rows/s). 20 000 batched inserts land in 142 ms — less time than 25 unbatched ones. That is the single most important thing to know when sizing a workload against this store, and it now has a line in the benchmark table instead of hiding inside an unrelated phase. README gains a "Running SQLite" section covering the two required pragmas and why, what cannot work (WAL needs a shared-memory index WASI cannot provide; concurrent connections need fcntl locking it also lacks — neither restricting anything real, since the store is single-writer by design), the optional modules, the benchmark table, and the batching guidance. Verified end to end against MinIO with Object Lock on the roots bucket at 20 000 accounts. Every phase passes; integrity_check clean before and after reopen. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
An enclave's system clock is not its own. It is seeded by the hypervisor at boot, has no NTP, and drifts — and the party that sets it is the parent instance, which is exactly the party the enclave exists to distrust. AWS addresses this by exposing the Nitro card's PTP hardware clock, synchronised to the Amazon Time Sync Service, at /dev/ptp0. The guest's wasi:clocks/wall-clock now reads that device. set-times with "now" shares the same clock, so a guest cannot read one answer from wall-clock and have the filesystem record a different one. Reading uses the POSIX dynamic-clock mechanism — open the character device, derive a clock id from the fd, clock_gettime on that — via rustix::time::clock_gettime_dynamic. rustix is already compiled into the tree through wasmtime, so this adds no build cost and no direct libc dependency. Hand-coding the ((~fd) << 3) | 3 encoding would have been worse than useless: getting it wrong yields a valid-looking clock id for some *other* clock rather than an error. --clock-source is auto | ptp | host. `ptp` refuses to start without the device and is what an enclave image sets; `auto` prefers PTP and warns loudly when it falls back, so a misconfigured enclave never looks like a correct one in the logs. --clock-check prints readings and exits without mounting anything. The monotonic clock deliberately stays on CLOCK_MONOTONIC. It backs WASI's timer subscriptions, so it must be cheap, and it must never step backwards — which a clock disciplined by an external source can, and a wall clock is allowed to. HostWallClock::now cannot fail, so the adapter had to decide what a failed device read means mid-run. It serves the last good reading and logs. A guest should not trap because a driver hiccuped, and a zero timestamp — expired certificates, rejected tokens, dates in 1970 — is far more damaging downstream than one a few milliseconds stale. Verified against a real PTP hardware clock, not a mock. /dev/ptp0 is root-owned here and there is no passwordless sudo, so deploy/ptp-check.sh passes the host's PHC into a container via --device. Readings: vs CLOCK_REALTIME -520.851 ms, stable across five samples read cost 10-15 us, against ~0.3 us for the system clock The skew is the useful part: it proves the PHC is genuinely a different clock rather than CLOCK_REALTIME under another name. The read cost is documented, and the standard mitigation — sample periodically, track with CLOCK_MONOTONIC between samples — is deliberately not built until a measurement asks for it. End to end, guest-smoke now prints the wall clock it sees and rejects anything before 2020; run under --clock-source ptp against MinIO it reports PTP time. The script also has a qemu mode using ptp_kvm, written but not exercised here (it needs cloud-image-utils). It is the weaker of the two: the clock behind ptp_kvm *is* the host's, so skew is ~0 and it cannot distinguish reading the PHC from reading CLOCK_REALTIME. Neither mode emulates Nitro; that needs QEMU >= 9.1's nitro-enclave machine and an EIF, noted in the roadmap with attestation. Roadmap also records the follow-on this leaves open: the Object Lock retention deadline still uses SystemTime::now(), and a COMPLIANCE deadline computed from a wrong clock cannot be corrected afterwards by anyone. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
wasi:random/random is what a guest builds keys, nonces and session identifiers
from. It came from wasmtime-wasi's default, the kernel's getrandom(2). Inside
an enclave that pool is NSM-seeded — but nothing in the path says so and
nothing fails if it isn't, which is too thin for the one interface where a
silent downgrade cannot be detected afterwards.
Bytes now come from /dev/nsm: the same device that signs attestation
documents, reached with the same ioctl. Straight from the device on every call,
with no DRBG in between, so the claim is "the NSM produced these" and there is
no software generator to reason about. The cost is real and documented: the
device answers 256 bytes per call, so a larger request is a loop of ioctls.
crates/s3fs-host/src/nsm.rs is the device layer, and deliberately generic — only
`NsmDevice::request` knows about the ioctl, everything above deals in CBOR
payloads. Attestation in M8 is the same call with a different payload, so that
work now starts from a working device rather than from scratch.
Two details that would otherwise cost an afternoon each, both from reading the
kernel's uapi/linux/nsm.h rather than guessing: the raw ioctl requires
CAP_SYS_ADMIN, so an unprivileged process gets EPERM from a device that exists
and is readable; and response.len is in/out, capacity going in and bytes
written coming back.
GuestRandom panics when the source fails, which looks inconsistent beside the
clock serving a stale reading — so the module says why. A stale timestamp is
still a timestamp and a caller can notice. Predictable bytes handed to a guest
that believes them random are indistinguishable from good ones at the point of
use; the guest builds a key and nothing downstream can ever tell. Stopping is
the only safe answer. A fallback to kernel entropy is logged at error, not
warning, for the same reason.
wasi:random/insecure keeps wasmtime-wasi's generator — making a
deliberately-not-cryptographic interface cost a device round trip would be
perverse — but its seed is drawn from the source once so it is not
deterministic across runs.
--clock-check becomes --self-check, now covering entropy too: it draws bytes and
applies the crude checks that catch how emulators and misconfigured drivers
actually fail — all zeros, a constant byte, two identical draws, a stuck byte
histogram. It touches no storage, which is what makes it the right thing to run
inside an emulator.
Verified end to end against MinIO: guest-smoke draws two 32-byte values through
wasi:random and rejects them if identical or zero; consecutive runs reported
f3bd77e4a331b205 and 71212eed82253f1c.
Also adds deploy/qemu-nitro/Dockerfile, which builds QEMU 9.2 with the
nitro-enclave machine. This is needed because no distro package works: Debian
sid ships QEMU 11 with the machine type present but `-device help` has no
virtio-nsm, since the NSM device is only compiled when libcbor and gnutls are
found at configure time. The image asserts virtio-nsm exists at build time
rather than letting a boot fail later for unrelated-looking reasons. Confirmed
before building anything that QEMU's virtio-nsm implements GetRandom, not only
attestation — hw/virtio/virtio-nsm.c handle_get_random, returning
{"GetRandom":{"random":<256 bytes>}}, which is what this decoder expects.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`/dev/nsm` exists nowhere but an enclave, so until now the device layer was
covered only by a fake and by reading the driver's ioctl definition. That is
thin for the interface a guest builds keys from: a wrong opcode or a
misread response length would pass every test in the tree.
deploy/qemu-nitro/ closes it. build-eif.sh assembles a real EIF — AWS's own
enclave kernel and init, plus a static musl nsm-selftest in the ramdisk — and
run-selftest.sh boots it on QEMU's nitro-enclave machine and greps the console.
The guest draws from the emulated NSM and reports a sample, a byte histogram
and the per-call cost. It passes, at roughly 62us per 64-byte device call, and
the sample differs across boots.
Four things had to line up, none of which announces itself when missing:
- QEMU with virtio-nsm, only compiled when libcbor and gnutls are present at
configure time, so no distro package has it. The Dockerfile fails the build
rather than the run if the device is absent.
- A vhost-user vsock backend; the machine has no built-in one.
- An answer to init's boot heartbeat. It writes 0xB7 to the parent on vsock
port 9000 and waits, with no timeout, so an unanswered heartbeat looks
exactly like a broken image. vhost-device-vsock's unix-socket backend
silently drops it — init dials CID 3, the Nitro parent convention, and that
backend serves only the host CID — hence --forward-cid and a real AF_VSOCK
listener.
- Mountpoints inside rootfs/, since init binds /rootfs onto itself and mounts
the pseudo-filesystems there. Missing ones abort with
`mount: /dev: No such file or directory`, which reads like a bootstrap fault.
The NSM device layer moves to its own crate, nitro-nsm, so it can link into a
small static binary for an enclave image without dragging in wasmtime and the
AWS SDK. M8's attestation reuses the same ioctl, and eif_build already produces
genuine PCR0/1/2, so the harness extends to attestation rather than waiting for
hardware.
Also fixes two clippy lints new in the 1.97 toolchain the harness needed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The guest could read a Merkle-anchored filesystem, a trusted clock and NSM entropy, but nothing could reach it: the runtime ran a command guest to completion and exited. This adds the other half — a `wasi:http/proxy` guest that answers requests, with `--mode serve`. The guest receives a parsed request and returns a response. It never sees a socket, a connection or a TLS record. That split is the whole reason to terminate TLS inside the enclave later rather than in front of it: a reverse proxy on the parent instance would read every request in the clear, and the parent is the party an enclave exists to exclude. Egress is refused rather than merely unused. `wasmtime-wasi-http`'s `default-send-request` feature is off, which turns `send_request` from a defaulted method into one this crate is required to write — so `EgressPolicy` denies it explicitly, with a test, instead of the answer depending on which crate features happened to be enabled. `GuestEnvironment` now builds the per-instance state for both modes. A command guest needs one; an HTTP guest needs a fresh one per request, since a Store cannot be reused and reusing one would leak a guest's resource table into the next caller's request. Both must produce identical environments, so both go through one place rather than each assembling a WasiCtxBuilder and drifting. Concurrency defaults to one request at a time. `Fs` is safe to share — handles under a parking_lot mutex, transaction state under a tokio one — but the guest is not necessarily safe to run twice over the same data: SQLite on WASI has to hold `locking_mode=EXCLUSIVE` because WASI has no fcntl, so two instances would each believe they had the database alone. Raising it is safe for a guest with no cross-request state, and the flag says so. examples/guest-http is deliberately stateful. `/counter` reads, increments and writes a file, so a second request only returns 2 if the first request's write was committed and the next instance read it back — a stateless handler would pass even if every request got an empty store. Verified live against MinIO, not just in tests: the counter reached 3 over three curl requests, a POST/GET round-tripped a file, the environment policy delivered DEPLOYMENT while withholding all AWS_ and S3FS_ variables, and after killing and restarting the process the counter resumed at 4 with the file intact — commits reached the store and the anchor chain was picked back up. CI gains a job that builds the component and runs the dispatch path over the in-memory backend, so this is covered on every push without Docker. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The device layer could ask the NSM for random bytes. This adds the other request that matters: Attestation, which returns a COSE_Sign1 signed by AWS carrying the enclave's PCRs and up to three caller-supplied fields. It is what lets a remote client establish something TLS cannot — that the code terminating its connection is the code it expected. Absent request fields are omitted rather than sent as null. The device parses each present key by pulling a byte string out of it, so `"nonce": null` makes it reject the whole request as InvalidOperation, naming no field. That cost is paid once here and recorded in a test. Verification is a separate crate. A client checking a document has no /dev/nsm, is frequently not Linux, and must not need a device layer to check a signature — so nitro-attestation depends on neither nitro-nsm nor anything else in the workspace. It carries the AWS Nitro root, pinned by its published SHA-256 so an edit to that file fails the build rather than quietly moving the trust anchor. The tests build genuinely signed documents rather than hand-written blobs: a real P-384 chain from rcgen, a real ES384 COSE_Sign1. A verifier that skipped the signature, the chain, the validity window or the root pin passes the happy path and fails the rest, which is why the rest are there — a tampered payload, a document from another root, a leaf borrowed from another chain, an expired chain, an empty cabundle, a replayed nonce, and a certificate binding that does not match. `Trust` distinguishes chain-verified from self-signed because the QEMU harness cannot produce the former: the emulator signs with a key it generated. A document from it is structurally real and internally consistent while proving nothing about AWS hardware, and the two must never be confused. `nitro-attest` puts the whole check in one command. It keeps the certificate the server presented, asks for a document quoting a fresh nonce, verifies it, and then checks that user_data binds *that* certificate. Without the last step a valid document proves only that some enclave exists somewhere — a proxy could fetch a real one and serve it over its own TLS session. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The runtime now serves HTTPS itself, with a key generated inside the enclave, and hashes that certificate into the attestation document. A client can then establish something TLS alone cannot: the session it holds terminates in an enclave running a specific image — not in the EC2 instance hosting it, and not in a proxy the operator controls. Terminating on the parent and forwarding plaintext would be far simpler and would concede exactly the thing an enclave exists to prevent. user_data follows nitriding's layout: two multihash-prefixed SHA-256 digests, the TLS leaf and the guest component. The certificate hash ties the connection to the document; the guest hash says which application was behind it. The prefixes are what let the format grow later without every existing verifier silently misreading it. /enclave/ is reserved for the runtime and checked before the guest ever sees a request. A guest able to answer under that prefix could serve any attestation it liked. The nonce is required rather than optional — a document without one cannot be shown to be fresh, and serving one on demand invites precisely the replay the nonce exists to stop. Documents are sent no-store for the same reason. Attestation is requested once at startup, so an NSM that will not attest fails the boot rather than the first client request, and asking for attestation without TLS is refused outright instead of promising a binding it cannot have. serve_tls.rs is the milestone's real test: a client opens a genuine TLS connection, fetches a document quoting a nonce it chose, verifies it with the production verifier against a pinned root, and checks that user_data names the certificate from the handshake. It also checks the failures — a substituted certificate breaks the binding, an earlier document does not satisfy a later nonce, and /enclave/* never reaches the guest. Two things the tests found rather than the other way round. Guest responses carry no content-length, so hyper chunks them; nitro-attest refused chunked bodies and would have failed against this very server, and now decodes them. And the TLS key's randomness deserved an explicit argument rather than a shrug: inside an enclave the kernel pool is NSM-seeded and nothing else, which the boot log shows, and the image sets S3FS_RANDOM_SOURCE=nsm so that is checked rather than assumed. The reasoning is in serve/tls.rs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
An enclave has no NIC. Everything above this — the TLS listener, the S3 client, the ACME client that comes next — assumes a socket works, and until now none of them could have. Its only channel out is AF_VSOCK to the parent. gvproxy turns that channel into an interface: it terminates the enclave's ethernet frames on the parent and provides a gateway, DHCP and DNS at 192.168.127.1. This is what nitriding and ArkLabs both do, and the reason is that ordinary networking code then works unmodified. The alternative — a vsock connector threaded through the AWS SDK and an ACME client — is two bespoke transports to write and maintain against libraries that do not expect them. Networking comes up before the mount, because mounting reaches S3: a runtime that mounted first would fail with an S3 error naming the wrong cause. Readiness is a TCP connection to the gateway rather than a sleep, since DHCP takes an unpredictable moment and a fixed delay is either too short on a slow boot or wasted on a fast one. The timeout message names the gvproxy command the parent is missing, because "connection refused" thirty seconds into a boot with no shell is not a diagnosis. gvforwarder ships inside the image rather than being fetched at boot, so PCR0 covers it. A forwarder streamed in later would be code sitting on every packet that the attestation says nothing about. deploy/parent/ is the other half, and it is written down because getting it wrong looks exactly like a broken enclave: it boots, and then answers nothing. Inbound :443 is forwarded through gvproxy's API, which also carries ACME's TLS-ALPN-01 challenge, so no second port is needed. What the parent can see is worth being precise about. It carries ciphertext: TLS to S3 and KMS is established inside the enclave, and inbound HTTPS is terminated inside it. What the parent does learn is metadata, and it can refuse to carry anything — neither is new, since it already decides whether the enclave runs at all. What is not safe is trusting its DNS, which is why every outbound connection validates certificates. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A self-signed certificate satisfies any client that verifies attestation — the binding proves more than a CA signature can. It does not satisfy a browser, which cannot check attestations and simply refuses. So the enclave gets a real certificate, and the private key still never leaves it. TLS-ALPN-01, so the challenge arrives on port 443 — the port the parent already forwards. HTTP-01 would need port 80 forwarded as well; DNS-01 would put DNS credentials inside the enclave, a standing secret able to mint certificates for the whole zone. The ClientHello decides per connection which configuration to use, so validation and service share one port. The cache is the security-relevant part. rustls-acme hands it the ACME account key and the certificate's private key as PEM, so both are sealed with AES-256-GCM under a new HKDF label beside the block and root-signing keys, and written as ordinary objects. The parent stores ciphertext. They are deliberately *not* in the filesystem: the guest's preopen is the filesystem root, and a guest holding the TLS private key could impersonate the enclave to every client. The object key is the AEAD's associated data, so a sealed blob cannot be moved from the account slot to the certificate slot by anyone who can write to the bucket, and the directory URL is part of the key so a staging certificate can never be served in production. Caching is not an optimisation. Let's Encrypt allows five duplicate certificates per week; an enclave re-issuing on every boot would exhaust that and be unable to serve. The attestation binding had to become dynamic. A certificate issued asynchronously and replaced on renewal is not the one captured at startup, and attesting a certificate that is no longer being presented would break the binding for every client at the moment it mattered. EnclaveEndpoints now reads a shared slot the cache publishes to, and answers 503 before any certificate exists rather than serving a document that promises a binding it does not have. Verified: 446 tests pass, including sealing round-trips, cross-slot and cross-key rejection, tamper detection, and that the leaf is taken from after the private key rather than from it. A live run against a deliberately dead directory confirms the runtime mounts, starts ACME, binds, serves, and retries orders with backoff instead of dying on a CA that is not answering. Not verified: a successful issuance. That needs a CA — Pebble or Let's Encrypt — and a domain the CA can reach, neither of which exists here. The code path from a signed order to a served certificate is written and unexercised. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Covers what the guest gets and what it deliberately cannot reach, the concurrency default and the SQLite reason behind it, the attestation binding and the one check that makes it worth anything, the two certificate modes and what Let's Encrypt does and does not buy, and what the parent instance must be running. Says plainly what is not verified: ACME issuance has no CA to run against here, and chain validation to the AWS root stays hardware-only. M8's entry now reflects that attestation is built; what it still needs from this layer is the public_key field for KMS Decrypt with Recipient. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`ServerConfig::builder()` resolves its provider from rustls's compiled-in features and panics when more than one is present. Both are, in any build with the `aws` feature: the AWS SDK brings rustls with `ring` while this crate asks for `aws-lc-rs`, so there is no unambiguous default and every call panics. This was a real startup crash, not a test artefact. `enclave-runtime` defaults to `--tls self-signed`, so the first thing it does in its default serving configuration is build a ServerConfig — and it would have died there. It surfaced as two flaky-looking unit tests, which passed under `-p s3fs-host` and failed under `--workspace`, because feature unification is what decides whether both providers end up linked. Naming the provider fixes it at all three call sites and also states the intent: the whole image stays on one implementation of these primitives rather than shipping two. Verified against the path that would have crashed: the runtime generates a certificate, serves the guest over real TLS, and the SHA-256 openssl reports for the presented certificate matches the one logged and bound into user_data, byte for byte. Workspace suite: 455 passing, stable across repeated runs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three things land together because each was useless without the others: a Nix
build that produces the same PCR0 twice, a QEMU harness that boots that image
with a real network, and the Packer and OpenTofu to put it on an instance.
PCR0 is a digest of the enclave image, pinned by a KMS key policy and checked
by clients. Its whole value is that someone else can rebuild and get the same
number — and the Dockerfile it replaces ran `apt-get update` over
`debian:bookworm-slim`, so it produced a different one every week and attested
to nothing anybody could reproduce.
`nix build .#eif --rebuild` now passes. Getting there was four things, none of
which announced itself:
- `cpio --reproducible`. The newc header stores each file's inode and device
numbers, so even the two-file *bootstrap* ramdisk came out different every
build, taking PCR1 with it.
- fixed uid/gid/mtime, and members sorted with LC_ALL=C, since cpio records
the order it is handed.
- `faketime` around eif_build, which stamps wall-clock BuildTime into the
image's metadata. No PCR covers metadata, so the measurements were already
reproducible — but the file differed by 15 bytes, and people compare
artifacts by hashing them.
- a pinned Cargo.lock for eif_build, which upstream gitignores. Letting the
dependency set of the tool that *computes PCR0* float would mean the
measurement came from something slightly different each time.
The end-to-end asserts three things that had never been checked together: the
filesystem mounts over gvproxy and a counter advances across requests, so
writes reach MinIO and come back; user_data binds the certificate from the
connection's own handshake; and the attested PCR0 equals the one the build
printed. Getting a boot that far surfaced four real bugs, each invisible until
the one before it was fixed:
- the ramdisk had no /etc, so writing resolv.conf failed;
- gvforwarder's stderr went to /dev/null, so the next three failures were
only visible as "the gateway never answered";
- it shells out to a DHCP client, which the image did not have — and
nixpkgs patches busybox to find its lease script inside its own store
path, so copying out just bin/busybox got a lease and applied none of it;
- no CA bundle, so the AWS SDK panicked with "no CA certificates found"
before making a request. That one matters in production too.
One thing I had assumed and was wrong about: QEMU's emulated NSM does not sign
attestation documents. Its source says "we don't actually sign the data, so we
use -1 as the 'alg' value", and -1 is not a COSE algorithm identifier — which
is why a strict parser rejected it. So the harness runs `nitro-attest
--unsigned-emulator`, which checks the contents the runtime put there and says
plainly on every run that no signature and no chain were verified. The earlier
claim that QEMU "signs with a key it generated" is corrected in the docs.
The parent is built with Packer over Amazon Linux 2023 rather than Nix, and
that is not inconsistency: the parent is the party the enclave excludes, PCR0
covers the enclave and not its host, so reproducing the host buys no security
property. No Docker on it either — that is only needed by `nitro-cli
build-enclave`. gvproxy and the gvforwarder inside the EIF come from one
nixpkgs package built static, so both ends of the vsock share a pin.
Verified here: reproducible PCR0, the full e2e, the entropy self-test on the
Nix-built image, 455 tests, clippy clean, `packer validate` and `tofu
validate`. Not verified: there are no AWS credentials on this machine, so no
AMI was built and nothing was applied.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
It was carved out as the development CLI — explicit flags, and a guest
environment that stays empty unless asked — against enclave-runtime as the
deployment target. That split stopped holding some time ago:
- nothing ran it. Zero references in CI; its only tests were six over clap
parsing, so it compiled and nothing more.
- it could not do half the product. No serve, TLS, network or attestation
flags at all, against enclave-runtime's eleven.
- it went unused through the feature it existed to support. Every local
MinIO test in the HTTPS milestone — the counter, the certificate binding,
the restart durability check — ran through enclave-runtime, because the
runner could not serve. The designated development tool was not reached
for once while developing.
enclave-runtime is a strict superset: --no-inherit-env --guest-path X is
exactly what the runner did. A second binary that is a worse copy of the
first, and untested, is how a codebase acquires a tool nobody trusts.
The one thing it contributed was a safe default for an uncurated environment.
Inheritance is right inside an enclave, where the environment is the attested
deployment configuration, and wrong on a developer's machine, where it
forwards whatever is in your shell — the AWS_/S3FS_ denylist withholds this
runtime's own credentials, not a GITHUB_TOKEN. So --no-inherit-env is now the
documented local path, --guest-env gained the S3FS_GUEST_ENV binding every
other setting already had, and both are tested.
Nothing was lost with the deleted tests: the credential regression guard lives
in s3fs-host's env.rs, where it belongs — it is a property of GuestEnvPolicy
rather than of any particular CLI.
Demonstrated rather than assumed, against MinIO with guest-smoke reporting
what it could see: the default inherits the shell, --no-inherit-env withholds
it, --guest-env NAME passes exactly that one through, an explicit value
overrides, and no AWS_/S3FS_ variable reaches the guest in any configuration.
The first version of that check probed a variable the guest never prints and
"passed" three cases while measuring nothing.
M9's `s3fs-runner gc --keep-roots N` now has no home, and deliberately does
not move into enclave-runtime: that binary ships inside the enclave image, so
every flag added to it changes PCR0. The roadmap records that garbage
collection needs a small unattested CLI first.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
s3fs-host began as `wasi:filesystem` over the block store — a library any host could embed, deliberately knowing nothing about AWS or enclaves. That is no longer what it was. It had accumulated the NSM, the vsock tap device, TLS termination, ACME and the attestation endpoints, none of which mean anything outside an enclave, and after s3fs-runner was deleted it had exactly one consumer. The crate boundary had stopped describing a real separation and was only charging rent for it. The layer that genuinely is reusable is s3fs-core, which still has no wasmtime dependency at all. That boundary stays. It becomes a library beside the binary rather than collapsing into main.rs, because integration tests cannot import a binary-only crate — and tests/serve_tls.rs, where a client checks that the attestation document binds the certificate from its own handshake, is the test this design exists to pass. Both suites still pass from their new home, 7 and 7. The `aws` and `serve` features are gone. They existed so a host embedding only the filesystem bindings could skip the AWS SDK and hyper; with one consumer that always wanted both, they bought nothing and cost `--features aws,s3fs-host/serve` on every build, test and clippy invocation, in CI and by hand. There is now nothing to select: `cargo test --workspace` runs all 451 tests. Two CI jobs that differed only by those flags were byte-identical afterwards, so one is gone. Also renames the WIT world from `s3fs-host` to `enclave-runtime`, since it is generated into that crate and the old name would have been the last place the dead one survived by accident rather than as a note. Verified beyond compilation: the flake still builds a reproducible EIF from the merged crate, and both serve suites run against a real component. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- Implemented `describe_pcr` and `extend_pcr` methods in the `Nsm` trait. - Updated `FakeNsm` to support PCR operations, allowing for testing of PCR extension logic. - Introduced `Pcr` struct to represent PCR state, including locked status and value. - Added utility functions for encoding and decoding PCR requests and responses. refactor(nitro-attestation): transition from single PCR to multiple PCRs - Changed `Expectations` struct to use a map for multiple PCR values instead of a single `pcr0`. - Updated related methods and tests to accommodate the new PCR structure. - Enhanced the `expect` method to validate multiple PCRs against the attestation document. fix(s3fs-core): separate filesystem creation from mounting - Modified `Fs::mount` to fail with `FsError::NoFilesystem` if the store is empty, preventing accidental filesystem creation. - Introduced `Fs::create` method to explicitly create a filesystem in an empty store. - Updated tests to reflect the new behavior of filesystem mounting and creation. chore(deployment): add deployment configuration for S3FS - Created `deployment.nix` to define the S3 bucket configuration for the enclave image. - Ensured that the configuration is baked into the image to maintain PCR integrity. test(e2e): validate state origin across restarts - Enhanced end-to-end tests to verify that a second boot resumes the filesystem state rather than creating a new one. - Added checks to ensure that the state root remains consistent across boots. docs(roadmap): update roadmap with recent changes - Clarified the status of binding attestation into the anchor and closing the cold-mount gap. - Updated descriptions to reflect the current implementation and future goals.
`Store::open` no longer formats an empty store, so `Fs::mount` returns
`NoFilesystem` where it used to hand back a fresh filesystem. Four harnesses
were still calling it against empty backends and had been failing since:
tests/serve_guest.rs 6 tests, including the TLS-adjacent dispatch path
tests/serve_tls.rs 7 tests, one of which the crate docs call
"the test this whole design exists to pass"
tests/minio_integration.rs the shared mount helper, so all 10
benches/fs_hot_paths.rs would have panicked on first run
None of it showed up in `cargo test --workspace`: every one of those tests is
`#[ignore]`d behind Docker or a built wasm component, so the default run stayed
green while the suites that actually exercise the serving path did not run at
all. That is the more useful finding than the fix.
Harnesses over a fresh backend now say `Fs::create`. The MinIO helper both
creates and mounts, because several of its tests remount a bucket to prove the
state survived.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Object Lock COMPLIANCE protects a *version*, not the name it lives under. A
`DeleteObject` with no version id writes a delete marker and **succeeds**; an
ordinary `GetObject` then answers `NoSuchKey` while the retained version sits
underneath, genuinely undeletable. Verified against MinIO:
$ aws s3api delete-object --bucket lockprobe --key roots/0
{ "DeleteMarker": true, "VersionId": "71eff5c3-…" }
$ aws s3api get-object --bucket lockprobe --key roots/0 -
NoSuchKey: The specified key does not exist.
$ aws s3api delete-object --bucket lockprobe --key roots/0 --version-id afc3c2fd-…
InvalidRequest: Object is WORM protected and cannot be overwritten
So "hide the tip" needed no lie from S3, just a delete nobody was allowed to
refuse. Two consequences, both load-bearing:
- The boot machine treats a missing state-origin receipt as authorisation to
create a filesystem. Mark the receipt *and* the sealed key and it reads
`(None, None)`, takes the genesis path, and serves a second filesystem
beside the one it was hiding — the exact substitution receipts exist to
refuse, with every signature along the way valid.
- Worse, `If-None-Match: *` tests the current version, so a marker makes the
genesis lease free again. The "am I the first here?" check answers yes.
`Backend::get_retained_blob` finds the version through `ListObjectVersions` and
reads it by `versionId`, which a marker cannot conceal because the version
cannot be removed. Used by `boot::maybe_get` and by `RootStore::exists`/`load`,
which also closes the rollback half — hiding the newest root was the same
trick. It needs `s3:ListBucketVersions`, and errors rather than answering
"absent" without it, so a missing permission refuses the mount instead of
silently starting over.
`MemoryBackend` described itself as modelling "a non-versioned bucket", which
S3 will not let you have with Object Lock at all — so the fake was *stronger*
than the real thing and three tests proved a property that did not hold. It now
models delete markers, and those tests assert what is true: retention makes an
object impossible to destroy, not impossible to hide.
Found because `object_lock_makes_a_root_record_undeletable` was failing. It was
asserting the comfortable claim.
Also fixes a dead error arm in `boot::genesis`: it matched `FsError::Conflict`
on a failed conditional PUT, but both backends return `AlreadyExists`, so a
genesis race reported the generic message instead of the intended one.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`pwrite` and `set_size` both refuse a handle opened without write. `set_times` did not, so `OpenFlags::write` meant "may change the contents" rather than "may change the file" — and the exemption was silent, which is the part that would have cost someone an afternoon. It matters more here than in an ordinary filesystem because setting a timestamp is a *commit*: `flush` publishes a new signed root record under Object Lock retention. A read-only handle that can advance the anchor chain is not read-only in any sense worth the name. POSIX would settle this by ownership rather than by the descriptor's open mode — `futimens` on an `O_RDONLY` fd is legal for the owner. There is no owner here: `Attrs` carries `mode` but no uid, so "are you allowed?" has no answer other than what the handle was opened for. `set_times_at` stays ungated for the same reason, having no handle to ask. Both new tests were confirmed to fail without the guard. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`ServeHandle::handle` acquires the concurrency permit, moves the `Store` into a
spawned task, and blocks on a oneshot the guest fulfils by calling
`response-outparam::set`. The sender lives inside that `Store`. So a guest that
neither returns nor sets a response never drops the sender, the receiver never
resolves, the permit is never released — and at the default of one request in
flight, one such request is the whole server for the life of the process.
Nothing bounded it: no fuel, no epoch interruption, no timeout, and both engines
built with a bare `Config::new()`.
Two mechanisms now, because neither covers the other:
- Epoch interruption, with a 250ms ticker and a per-request deadline. This is
the only thing that can stop wasm that never yields.
- A timeout on the response head, which bounds the wait regardless of *why*
the guest is quiet — including a host call that never returns, which the
epoch cannot see because no wasm is executing.
**The deadline must fire before the timeout, and that ordering is the whole
mechanism.** The first version of this gave the deadline slack *past* the
timeout so the error would read "took too long" rather than an opaque trap.
That inverted it: `task.abort()` fired first against a guest nothing could
interrupt, `abort` needs an await point a spinning guest never reaches, and the
task span on after the request was abandoned — hanging runtime shutdown instead
of the request. The tidier message cost the entire guarantee, and the test
caught it by hanging.
`examples/guest-http` gains `GET /hang`, deliberately rather than as a test
fixture: a guest that never answers is the one behaviour a host cannot provoke
from outside, so without it the watchdog has nothing to be tested against.
Reverting the deadline ordering makes the new test hang, which is how it was
confirmed to test anything.
`S3FS_REQUEST_TIMEOUT_SECS` is baked into the image like every other setting, so
PCR0 covers how long a guest may take.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`S3FS_GUEST_LIFETIME=session` keeps one `Store` and one `Proxy` alive across
requests, so a guest is a process rather than a handler: it holds its linear
memory, and with it any cache, connection or state it wants. Default stays
`request`, because a single session shared by every caller is worse than
per-request isolation — it becomes the right default once a session belongs to
one authenticated client.
The store still moves into the spawned task, because the guest keeps streaming a
body after the response head has gone out and a borrowed store cannot outlive
the function. What changes is that the task hands it back afterwards. The lock
is an `OwnedMutexGuard` for exactly that reason: it is `'static`, so it can move
into the task and carry the session with it.
`Option` lives *inside* the lock rather than beside it, so "checked out" and
"absent" stay distinguishable. A request holding the lock and finding `None`
knows to instantiate; one that finds the lock held knows to wait. Two separate
pieces of state could not tell those apart, and a request that guessed wrong
would build a second instance over the same filesystem.
Three things a session must get right, all of which have tests:
- **Instantiate once.** `Instance` is a `Copy` index with no `Drop` and
`Store.instances` only grows, so instantiating per request on a reused store
leaks a component instance and its linear memory every time — bounded only
by wasmtime's default of 10,000. This is the one way to leak unboundedly
here; the resource table is a slab with a free list and does not.
- **Reset the epoch deadline per request**, or request N+1 inherits whatever
N left of it.
- **Discard rather than carry.** A trap leaves the instance in an unknown
state, and resources the guest failed to drop stay addressable by the next
request on that session. Either sets the slot to `None` so the next checkout
rebuilds.
A session is serialised by construction: the lock is held for the whole call
including the body, and two concurrent requests cannot share one linear memory
without wasm threads. That is a property of the model, not a limitation of this
implementation.
`examples/guest-http` gains `GET /memory`, a counter that touches no storage.
`/counter` proves the filesystem carried state between requests; this proves the
*instance* did, and under a per-request guest it can only ever return 1 — which
is what makes the three session tests discriminate. Forcing them to
`GuestLifetime::Request` fails all three, which is how that was confirmed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… be lied to Client certificates were never requested — `with_no_client_auth()` was the only server config in the crate — and `ServeHandle::handle` had no parameter for a connection identity even if one had existed. The service closure was also built *before* the TLS handshake, so nothing per-connection could reach a request. mTLS carries the identity. There is no CA and no chain to build: a client presents a self-signed certificate and the identity *is* that certificate. What makes it mean anything is not who vouched for it but that the handshake proved possession of the matching private key, which is why `AnyClientCertificate` accepts any certificate while **delegating** `verify_tls13_signature` to the provider. Answering `Ok` there instead would let anyone claim anyone's identity by copying a certificate, which is public. Client auth is offered, not required. A browser, `curl` or a health check presents nothing, completes the handshake, and arrives without an identity — which the guest sees as anonymous. Requiring it would break every ordinary client and the ACME challenge with them. `serve_connection<S>` is extracted so the closure is built after the handshake, which is the only point at which a peer certificate exists. The three arms — fixed TLS, ACME TLS, plaintext — now share one body; all three stream types satisfy the bound, so it monomorphises. The header is where this fails silently, so: `insert`, never `append` — `append` leaves a client's own copies and the guest's `fields.get()` returns a list, so `[0]` would be theirs. The `None` arm must `remove()`, or an unauthenticated connection passes its own value through. And it cannot be done in `is_forbidden_header`, which runs *inside* `new_incoming_request` — after injection — so it could only delete the header, silently. `rustls-acme` hardcodes `with_no_client_auth()` in `default_rustls_config()`, so the ACME path builds its own config from `state.resolver()`. Without that the identity plumbing would be present, correct, and dead in the one mode that has a real certificate. The challenge config keeps no client auth: the CA validating TLS-ALPN-01 presents none, and the one handshake that must not fail is not the place to ask. `GET /whoami` on the example guest echoes what the runtime told it. Three tests drive real handshakes rather than passing a value to `handle` directly: a certificate becomes the expected digest, two clients are two identities, and a client without one is still served. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The identity was SHA-256 of the whole certificate. A certificate carries a
validity period and a serial number that change on every renewal, so a client
who rotated a certificate around the *same key* became a different client:
public key inside each certificate 9bc0f6efe5aba683… 9bc0f6efe5aba683…
whole certificate 713e0ef105a01b33… 172f092635c5bb6b…
Survivable while an identity only picks a session — a renewed client loses
in-memory state and no more. Not survivable once an identity derives a
filesystem, which is the plan's next phase but one: a routine, scheduled
renewal would point a client at empty storage while their data stayed encrypted
under keys nothing would ever derive again, in a bucket that retains objects for
ten years. Silent, irreversible, and triggered by doing the correct thing.
Now SHA-256 over the `SubjectPublicKeyInfo` — the construction certificate
pinning uses (RFC 7469) — so the identity is the key, not the paperwork around
it.
The reason for the original was that extracting the key means parsing DER an
attacker supplies, and there was no parser in the image. There is: rustls has
already run `webpki` over this exact certificate to verify the handshake
signature, and `webpki::EndEntityCert` exposes `subject_public_key_info()`.
Naming it in `Cargo.toml` adds no code to the build — `cargo tree` shows the
same two `rustls-webpki` entries as before — only the ability to ask rather
than to re-derive a parser we would have to trust.
`from_certificate` now returns `Option`, because a certificate that will not
parse must produce no identity rather than a fabricated one. It should be
unreachable: the signature check that admitted the client parsed it first.
Two tests guard the property, one at each level. The unit test builds two
certificates on one key with different serials and asserts one identity —
reverting to the whole-certificate digest fails it, which is how it was
confirmed to test anything. The TLS test does the same over a real handshake.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
f90b1e0 added a guest that keeps its linear memory across requests, on the reasoning that a transaction cosigner is a long-running process. That was the wrong model, and this undoes it. A `Store` is what separates two wasm instances. Give two clients one instance and the separation stops being a runtime guarantee and becomes guest code: a bug that confuses two clients is no longer a leak of one, it is a total compromise. A trap poisons state everyone shares. Leaked resource handles stay addressable by whoever calls next. All three of those go away when the instance dies with the request, and none of them can be engineered out of a shared one. So the boundary goes back between clients, where it belongs. Per-client state — nonce ledger, policy, pending transactions — lives in the anchored filesystem, which is the only place it can, and that is the right kind of constraint: attested, hash-chained, and it survives a restart. The linear memory it would otherwise have sat in was none of those things. Removed: GuestLifetime and --guest-lifetime, in_session and the session mutex, the leaked-resource rebuild policy, and State::resources_settled, which only meant anything for a store that outlived a call. What replaces the eager start is narrower and honest about it: verify_instantiates builds one instance at boot and drops it, so a guest that traps in its initialiser stops the enclave instead of answering 500 to every request while looking healthy. `no_two_requests_share_an_instance` is the invariant as a test — /memory must answer 1 for every request, forever. It is what fails if anyone reintroduces sharing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two properties that were true in the code and nowhere else. The per-request invariant, which the README never stated: no two requests share an instance, and that is where the boundary between two clients actually lives. A `Store` separates two wasm instances; give two clients one instance and the separation becomes guest code instead. And the client identity from c84edde, which shipped undocumented. With a fresh instance per request it is the only thing tying a request to its client's state in the store — there is no session and nothing else to key on. The e2e's fourth leg asserted /memory *increments*, which was the check for the resident model f90b1e0 added and 6906cf9 removed. Inverted: two requests must both answer 1, which makes the leg a proof of the isolation invariant over real TLS rather than a leftover. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Written to settle an argument with a number rather than an intuition:
pooling instances per client saves exactly one thing — the instantiation
in the middle of every request — and the trade is only worth making if
that is a large fraction of a request.
instantiate 24 µs
mount 51 µs
dispatch/trivial 88 µs
dispatch/committing 753 µs
Against MemoryBackend, so these are engine costs with no network. Two
things fall out. Instantiation is ~3% of a request that does one real
transaction, and against S3 the commit grows by round trips while
instantiation does not — so that share only shrinks.
And `mount` is the larger number, which is the one that matters if
filesystems are ever per client: it derives key material, reads a signed
root record, verifies it, and opens the object set. The 51 µs here is the
CPU floor; in production it is that plus two or three S3 round trips.
Whatever a per-client design keeps warm, the mounted filesystem is the
thing worth keeping, not the wasm instance.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Until now the only MasterKeySource was StaticKey, which "seals" by writing the secret into the blob behind a marker saying it did not. The secret arrived as S3FS_MASTER_KEY, so the parent instance held it before S3 did — and from those 32 bytes HKDF derives the block AEAD key, the Ed25519 root-signing seed and the runtime-seal key. Whoever had them could decrypt the filesystem, forge root records, and decrypt the enclave's TLS private key to impersonate it to every client. KmsAttestedKey mints the secret with GenerateDataKey and recovers it with Decrypt, both carrying a Recipient: a fresh RSA key generated for that one call, named in an NSM attestation document. KMS checks the document against the key policy, encrypts its answer to that key, and omits Plaintext entirely. The parent still proxies the HTTPS request, because the enclave has no network of its own; what it cannot do is read the answer. The enforcement is the key policy, not this code. With kms:RecipientAttestation:PCR0 pinned to the approved image, a wrong enclave does not get a refused mount — it gets no key at all. That is the difference between an enclave checking itself and something outside it doing the checking. CiphertextForRecipient is a CMS EnvelopedData. The `cms` crate parses it; keys::recipient decides what is acceptable, and accepts exactly one shape: one RSA-OAEP-SHA256 recipient, AES-256-CBC content, id-data inside id-envelopedData. A general-purpose parser is fine, a general-purpose policy on the path that unwraps the master key is not. Refusing a second recipient is the sharp one — it would mean somebody besides this enclave could open the content, which is what Recipient exists to prevent. Storage splits in two. SSM holds the ciphertext and nothing else; the roots bucket holds a pointer naming the parameter, the CMK, the encryption context, and sha256 of the ciphertext. That keeps boot.rs's state machine intact — it still decides genesis from resume by what is present — and earns something: the state-origin receipt already commits to sha256(sealed), so it now attests which CMK and which context this filesystem was created under. A repointed pointer is caught by the receipt, a swapped parameter by the hash. A host can delete the parameter and stop the enclave booting; it cannot make it boot wrong. --master-key-source has no default, and the refusals are the point. A plaintext key under kms is an error, not an ignored setting: a production image that still carried S3FS_MASTER_KEY would work perfectly while quietly using a key the parent holds, and no ordering of "both were given" is safe to guess at. Absent CiphertextForRecipient is likewise an error and never a fallback — its absence means the request reached KMS without an attestation, which would have returned the key in the clear. The QEMU image stays on static and always will: its emulated NSM does not sign, and KMS will not accept an unsigned document. PCR0 differs between the two images, so a client can tell which it is talking to. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
An ordinary request is approved by an assertion bound to
`{method, path, query, sha256(body)}`, and the body hash is the part that makes
the approval mean *this* operation rather than any operation at that route. A
stream has no such hash to offer: its body is the message sequence, and none of
it exists when the channel opens.
The wrong resolution is to drop the binding for streams and call the result
authenticated. That would hand out an approval good for anything the client
later chose to send — precisely the substitution the hash exists to prevent,
reintroduced through a weaker sibling.
So a stream is opened by a different approval, issued by a different endpoint,
that authorizes opening a channel on one route and nothing else. Enrollment
tokens already have this shape: they create a tenant and do nothing else.
Anything inside the stream that asks the enclave to sign carries its own fresh,
single-use assertion bound to that message. The prototype carries the fields
and refuses without them; verifying them needs a runtime hook that does not
exist yet, and it should be designed before a real key depends on it.
Two independent things stop the weaker approval leaking into the strong path,
and one of them is not a check at all: `BodyBinding::Unbound` never equals
`BodyBinding::Exact(_)`, so no comparison can be satisfied by the wrong kind
and no `if` can be forgotten. The second is `x-enclave-stream`, a runtime-owned
header stripped before the guest sees it — the header alone cannot talk the
runtime out of hashing a body, and an unbound approval alone cannot be spent on
a request that arrived as an ordinary one. Deliberately not keyed on
`content-type: application/grpc` or a path prefix: the content type is
application data the guest also reads, and a prefix would bake a service name
into a runtime that knows nothing about the guest's routes.
`Gate::verify` now consumes the challenge before touching the body, because the
binding is what decides whether there is a body to read. One consequence worth
naming: an oversized body burns its challenge before being refused, where
before the refusal came first. That is the safer direction — otherwise a caller
can probe the size limit without ever spending an approval.
Streams are refused outright when no gate is configured. Anonymous callers
share one lock, so an anonymous stream would hold it for its whole life and
starve every other anonymous caller — one client silencing the runtime by
connecting. An unauthenticated deployment is a development arrangement and
should not learn to depend on a shape that only works once identities exist.
docs/STREAMING.md states what a stream-open approval buys, what a stolen one
buys, what an open stream costs the tenant holding it, and what the runtime
does not promise: no retries, no resumption, no deduplication, and no
exactly-once execution.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…his repo The guest hand-writes gRPC framing because tonic does not build for wasm32-wasip2. Testing it with framing written here too would only show the two halves agree with each other, which is the one thing that was never in doubt. So tonic drives it: a real bidirectional stream over TLS and ALPN-negotiated HTTP/2, decoded by prost on the client side, with `Status` and `Code` reported by tonic rather than by anything in this repository. A refusal arrives as `PermissionDenied` with its message intact, which is worth asserting through a real client precisely because the head is 200 — a client reading only the status line would call a refusal a success. The attestation is verified from the *initial metadata* before the client sends its first message, and the order is the point rather than an artefact of how the test reads. Verifying afterwards would establish that the enclave was listening only after you had already told it something. The document binds the nonce the client chose and the leaf from its own handshake, so tonic connecting through an already-open connection — rather than through tonic's own transport — is what lets the test hold that exact certificate. `grpc-timeout` gets a test of its own for what the runtime does *not* do with it. It is neither honoured nor stripped: the runtime's deadlines are its own so a client cannot lengthen them by asking, and interpreting a client's deadline belongs to the guest, which is the only party that knows what its work is worth. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The two are the same shape — an approval that authorizes creating one thing and nothing else — so the reader who has just understood one is the reader best placed to understand the other. Says plainly that an open channel is not standing permission to sign, next to the paragraph explaining why an approval for one transaction must not authorize another. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Authorization had two shapes that had grown apart. An ordinary request was approved by an assertion bound to `sha256(body)`; a stream was approved by a deliberately weaker assertion bound to nothing, because a stream's body is its message sequence and does not exist when the channel opens. The weaker path existed only because the stronger one could not express a stream. Both are replaced by one. A person authenticates with their passkey and the runtime issues a short-lived, single-use bearer token for one **interaction** — one request and its response, or one stream until it closes or reaches its deadline. The token names a method, a path and a query. ## What this gives up It does not name the body. A token issued for `POST /sign` authorizes whatever bytes follow, so a client compromised between the approval and the request can substitute the payload, and the runtime will not notice because it is no longer looking. That is a deliberate policy change, not an oversight, and the documentation now says so where it used to promise the opposite. What survives: a person still authenticates per interaction, tenant isolation is unchanged, challenges and tokens are single-use and die on restart, and nothing about the rate limiting moved. `TokenStore` mirrors `ChallengeStore`, which was already the right shape: bounded, swept on issue, refusing rather than evicting when full, and removing an entry before deciding anything about it — so two callers racing one token have exactly one winner and a token offered for the wrong route is spent by the attempt. Entries are keyed by `sha256(token)`, so holding the store is not holding a token. The exchange is necessarily three trips. A passkey is a challenge-response, so the assertion cannot exist until the challenge has been answered, and the token cannot exist until the assertion has been checked. There were only two before because the assertion rode on the operation itself; separating them is what lets one approval cover a stream. ## Attestation moves to where a client can act on it The runtime attests `/auth/` exchanges and leaves the rest alone. The challenge request is the one that goes first on a connection nothing has vouched for, and it is the right thing to send there: it carries the route but no approval, so a party that intercepted it learns what is intended and holds nothing it can act on. Its response identifies the enclave, and the client pins that certificate for the two trips that follow — TLS proves the peer holds its private key, which a second document would not add to. That halves the NSM signatures an interaction costs, on a device that is the runtime's throughput ceiling. The header is now *stripped* from unattested responses rather than left alone. It was only ever guest-proof because every response overwrote it, and a guest that could set it would be handing clients a document under the runtime's name. `/enclave/config` is gone with it, and `--attestation` with it: a flag that turns verification off is one that can be turned off by whoever starts the process, which inside an enclave is the party the enclave exists to exclude. ## The client verifies, in the right order `passkey-client` checked a document by parsing it, which does no signature or chain check at all — an interceptor could mint one naming its own certificate. It now verifies the chain, the age, the nonce, the certificate binding and the measurements, and it refuses to run without `--pcr0`. `--guest` is not accepted in its place. PCR0 is measured by the hypervisor and locked; the guest hash is part of `user_data`, which the runtime being attested chose — so an attacker's own enclave signs a genuine document claiming whatever guest hash was going to be checked for. `--guest` is a second check on a trusted measurement, not a substitute for one. Both trips that carry a credential are pinned before a byte goes out: the assertion, and the credential being registered. Sending either to an unpinned connection meant handing it to whoever answered, and checking the response afterwards would only prove the theft succeeded. ## An interaction now has a deadline Nothing watched a call after its head went out — the `JoinHandle` was dropped, which detaches. The epoch watchdog cannot help: a guest parked in a host call executes no wasm and never reaches an epoch check. So a peer could open a stream, say nothing, and hold its tenant's only slot until the process ended. A wall clock works precisely where the epoch does not, because such a guest *is* at an await point and `abort` reaches it. `--max-interaction-secs` bounds how long an interaction may run; `--interaction-token-ttl-secs` bounds how long an approval may sit unspent. This also restores a bound the pool lost: eviction skips busy slots, so long-lived streams were pinning tenants above `max_tenants` indefinitely. `--max-request-body-bytes` is deleted. The runtime stopped buffering bodies when it stopped hashing them, and a cap that applied to ordinary requests but not streams would need the distinction this change removes. The guest bounds messages, and the docs say so rather than pointing at a flag that could not. ## Also A gRPC stream ending mid-frame now reports `INVALID_ARGUMENT` rather than OK. The half-close looked clean at the HTTP layer; the stream did not end cleanly, and the trailers are the only place that difference can be said. `ClientMsg` loses its `challenge_id` and `assertion` fields. They promised a per-message check nothing performed — and nothing *can* perform yet: verifying an assertion needs the credential store, the issued challenge and the relying-party configuration, all of which live in the runtime, and a guest is given standard WASI with no host function to ask through. Fields that look like a credential and are never checked read as a guarantee. Closing that gap needs a new import, and it should be designed before a real key depends on it. `serve_auth` ran in no CI job at all, which was not a gap to leave standing now that it is the suite for the whole authorization model. It is wired in, with the client binary built ahead of it so the test that drives the real thing does not silently skip. Not verified: the QEMU end-to-end harness. `run-e2e.sh` changed in three places and only its shell syntax has been checked. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The guest component shipped inside the enclave image, so PCR0 covered it and every guest change was an image rebuild. It is now fetched at boot from a key in the roots bucket, `S3FS_GUEST_OBJECT`, and measured before anything asks for a key: PCR16 is extended once with `sha256(component)` and locked. The image still names *where* the guest comes from, under PCR0; PCR16 says *what* arrived. A key policy pinning both still releases the key to exactly this runtime running exactly this guest. The object does not have to be trusted. A parent that substitutes it gets an enclave with a different PCR16, which a policy pinning the approved guest releases nothing to and a client pinning that guest refuses. ## The order is the point `measure_guest` checks each step against what it should have produced, because each is a place a wrong answer would pass silently into every attestation: the register must start unlocked and zero, extending must give exactly what a client computes from the component, and the device — not the lock call's return value — must then report it locked with that value. A register that is extended but never locked appears in no attestation document, so no key policy could see it. `Pair::read` is the first thing `boot` does, before the store is read or KMS is asked, and an unlocked or never-extended PCR16 refuses to boot. The bytes measured are the bytes compiled and served. `lock_pcr` is new on the `Nsm` trait; a successful `LockPCR` answers with the bare string rather than a map, so it has its own decoder. ## PCR16 means something only beside PCR0 The runtime writes PCR16, so a runtime someone else wrote can put any value there. PCR0 says the runtime that wrote it is yours; PCR16 says which application that runtime loaded, which PCR0 can no longer say. Clients now require both. `passkey-client` refuses to run without `--pcr0` and one of `--pcr16` or `--guest`, and refuses the last two if they name different guests. `nitro-attest` enforces the same through clap, and gains `--measure`, which prints a component's PCR16 using the same `guest_pcr` the verifier checks with. `nix build .#guest-release` uses it to emit `guest-pcr16.json`, so the value in a key policy and the value a client pins cannot disagree. A document without PCR16 fails an expectation on it rather than skipping the check, because that is what a runtime that never locked the register produces. ## The key policy decides; boot records `--authorise-successor` is gone. It extended PCR31 to name the next image and never locked it, and since documents carry only locked registers, no document it produced on hardware could have carried the register the successor was checked against. It passed here only because the test NSMs put unlocked registers into every document. They now keep registers the way the device does: 0–15 locked from boot, 16 free until measured, only locked ones attested. Nothing replaces the handoff, because the boot machine stops having an opinion about which code may hold the state. An enclave KMS released the key to was allowed by the policy, and a second copy of that decision could only disagree with the real one. `verify_origin` now checks that the receipt names the state that was loaded, not who wrote it. What boot adds is a record. The first time a runtime and guest hold a state, the enclave attests a *pair record* carrying both registers against the `state_root` and stores it under Object Lock at a key derived from `sha256(PCR0 ‖ PCR16)`. That boot reports `Upgrade`, which replaces `Migration`; later boots of the pair find the record and write nothing. An existing record is verified, not merely found, since anyone who can write the roots bucket can derive its key and plant something there first. Two first boots racing leave one record and both boot; junk that wins the race is refused. It records which pairs have held the state, not in what order: returning to an earlier pair finds that pair's record and adds nothing. The order approvals were given in lives in CloudTrail's record of key-policy edits. Both conditions belong on `kms:GenerateDataKey` as well as `kms:Decrypt` — an unconditioned `GenerateDataKey` hands the parent a data key in the clear, which it can plant as a new filesystem's key before an enclave ever runs genesis. ## Also Stored receipts are verified against their chain as of the moment they were signed, not now. An attestation chain's certificates are short-lived next to a filesystem, so checking against the current time would make a filesystem unmountable by the passage of time. The timestamp is inside the signed payload, so a document stamped outside its chain's validity still fails. The flake's source filter excludes `examples/`, so a guest-only edit no longer changes the runtime's store path and with it the image's PCR0. `boot_origin` ran in no CI job, and neither did the unit tests of the `passkey-client` and `nitro-attest` binaries, which `--lib` does not reach. All three are wired in. The QEMU harness builds `.#guest-release`, uploads it to MinIO, and checks the attested PCR16 against `guest-pcr16.json`. A seventh leg replaces the object and restarts: the enclave measures the substitute, records it as an upgrade, and is refused by a client pinning the approved guest. The emulator has no KMS, so the half where a substituted guest gets no key still needs hardware. Leg 5b and the tofu `verify` output now describe interaction tokens and `/auth/`, which the previous commit introduced. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- Added `existing_tenant_root` function to ensure tenant directories exist before scheduling tasks. - Introduced background task handling in the `guest-http` example, including task enqueueing, status checking, cancellation, and forgetting tasks. - Updated the `ServeConfig` to support background tasks in various test files. - Created a new `BACKGROUND_TASKS.md` documentation file detailing the background task functionality and configuration. - Modified the Dockerfile and Nix deployment files to include options for enabling background tasks. - Added WIT interface definitions for task management, including enqueue, status, cancel, and forget operations. - Updated dependencies in `Cargo.lock` and `Cargo.toml` for compatibility with new features.
Three strands landed together because they overlap in the same files — `routes.rs`, `run-e2e.sh`, `flake.nix` and the README each carry more than one — and separating them would have meant hunk-level surgery rather than a clean division. ## Registration is open Registering a passkey no longer needs an invitation. `/auth/register/options` takes no token, `auth/enrollment.rs` is gone along with `--enrollment-token` and the image configuration that carried one, and both the client and the QEMU harness enrol with nothing but a URL. What registration grants is unchanged and deliberately narrow: a new, empty tenant. `PendingRegistration::join` is never `Some`, so a registration cannot reach an existing tenant or add a passkey to somebody else's. The WebAuthn challenge and verification are untouched, the tenant is created only after verification succeeds, and the client still checks the enclave's attestation before it trusts the response. An older client still sending `enrollment_token` is refused rather than quietly admitted, because the request now denies unknown fields. What this gives up is admission control over resource creation. A tenant is a directory the enclave keeps and a slot against `--max-tenants`, and anyone who can open a connection can consume both. The rate limiter on challenge requests is removed in the same breath, so `/auth/request/options` is unbounded per credential as well. Whatever needs to bound either now has to sit in front of the enclave. ## CI runs what it ships The recipes moved out of inline YAML into `scripts/`, because there was nowhere else for them to live: `run-e2e.sh` had grown its own copy of the MinIO setup, and no CI job could be exercised without pushing. Each job is now one script that behaves the same on a laptop, and the harness shares `minio-up.sh` instead of keeping a second recipe that had to agree with the first. Running every one of them found three defects that only CI would otherwise have shown. `cargo bench --workspace -- --output-format bencher` hands that flag to each crate's libtest harness, which rejects it and aborts the run after eight minutes of compiling. The SQLite job had been broken since bbbee4d — first for a missing `--master-key-source`, then for an NSM requirement no hosted runner can meet — so it is out of CI, and the script it left behind says why. And `wasi-sdk.sh` fetched without `curl -f`, which writes an error page into a tarball and fails two steps later complaining about the archive format. ## The enclave schedules work, and that is checked where it runs `run-e2e.sh` gains two legs: eight files written, read, overwritten and re-read through the gate, and work scheduled by one approved interaction that runs later with nothing signed at the moment it does. Three tests cover the same ground quickly — the real client binary enqueueing and polling to completion, a guest that genuinely fails retrying until the occurrence is failed, and a recurring task firing a second time through the running scheduler. Verified: rustfmt, clippy with `-D warnings`, 705 tests passing and 80 skipped, and the emulated end-to-end passing all eight legs with Alice and Bob registering without tokens into tenants that cannot see each other. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…pened A guest had no way to reach its user between requests. Background tasks made that gap plain: work finishes inside the enclave and nobody is told. A guest still cannot reach the network — `serve::EgressPolicy` refuses every outgoing request it makes, deliberately — so the runtime holds the Firebase credential and sends on its behalf, and the guest gets a host import instead of a socket. ## A wake signal carries nothing There is no title, no body and no `notification` object, and that absence is the feature. An FCM payload crosses the parent instance — the party an enclave exists to exclude — and then Google, so anything placed in it is disclosed to both. What travels is an opaque category, an optional tenant-local reference and a schema version. The app wakes and fetches the detail over its own attested connection, where the parent is shut out again. That is not a setting. A flag would be something a deployment turns on and a guest then uses, at which point it stops being a property. It is also the right shape mechanically: a `notification` block is rendered by the OS without the app running, so a wake carrying one would show text *and* fail to wake anything. ## Enrolling is interactive-only; waking is not The one place this departs from `enclave:tasks/queue`, where every mutation requires an interactive invocation. Enrolling a device grants standing ability to reach somebody, so background work must not be able to do it, nor to silence its owner by un-enrolling. But raising a wake *from* background work is the whole point — a task that finishes at three in the morning telling its owner to come and look. It grants nothing; it spends an enrolment an interactive call already made. The import is registered unconditionally, like tasks': authority is denied at call time by the `Option` in `State` being `None`, never by omitting the import, which would turn a policy refusal into a boot failure. ## What it gives up Delivery is best effort and unacknowledged. Repeat wakes for one category coalesce, a full queue drops and counts rather than blocking, and queued wakes do not survive a restart — in `tasks` the record *is* the work, whereas here it would be a stale pointer to work that already happened. A guest is never told which devices accepted, because that would hand it the user's device inventory and their online pattern, and it must be resilient to non-delivery regardless. A parent that steals the service account can ring doorbells: send wakes to tokens it obtains elsewhere, and delay or drop ours, which it could already do because it carries every packet. It cannot read any tenant's data or the device tokens, which live under a key KMS releases only against a matching PCR0 and PCR16 — and a forged wake means nothing, because a wake carries no state. ## Zero new crates in the measured image hyper's client, `webpki-roots` and `aws-lc-rs`'s RSA signing are already in the production closure. The manifest comment claiming the runtime "never calls out over HTTP" and that `hyper/client` was dev-only is corrected in passing: it has been false for as long as `aws-config`'s `rustls` feature has been enabled. The client's roots are compiled in rather than read from a file, because `net.rs` notes the parent answers DNS — certificate validation is the only thing stopping it pointing fcm.googleapis.com at itself. There is deliberately no custom verifier on this path. ## Not proven here Nothing in the suite reaches Google. The wire is covered by a transport double, and leg 6b/8 of the QEMU harness proves the whole path inside an emulated enclave against a local stub: a finished background task wakes an enrolled device with a message carrying no content. Whether Google accepts it is untested, the same way `guest_io::cloudwatch` is honest about its own credential path. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`crates/enclave-runtime/` sat beside `s3fs-core`, `nitro-attestation` and
`nitro-nsm` as though it were a peer. It is not: those three are libraries this
thing is built from, and it is the thing. The layout said otherwise to anyone
opening the repository for the first time, and where a directory sits is the
first claim a repository makes about what matters in it.
So `crates/enclave-runtime/` becomes `runtime/`. Pure `git mv`, plus the path
edits a move forces:
- the workspace member in `Cargo.toml`
- path dependencies inside `runtime/Cargo.toml`, which now reach *into*
`../crates/` rather than sideways within it
- one directory level fewer between the crate and the repository root, so every
`.join("../../examples/...")` in a test or bench loses a level, and
`wit-bindgen`'s `path: "../../wit"` becomes `path: "../wit"`
- two README links
## What was deliberately left alone
`runtime/tests/serve_guest.rs` still contains `"../../tenants/…"` and
`"../../../../../../etc/passwd"`. Those are not paths on this disk. They are
the strings a guest sends to see whether the runtime lets it climb out of its
tenant directory, and the test asserts it does not. Rewriting them along with
the real paths would have quietly weakened the test that exists to catch
exactly that — which is why they are called out here rather than left for
whoever runs the next sweep.
No behaviour changes. The crate keeps its name, so nothing that depends on
`enclave-runtime` notices.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ying
A client could not be developed against the emulator, only pointed at it. Every
client meeting a QEMU enclave had to pass `--unsigned-emulator`, and that flag
skips the signature, the certificate chain and the validity windows — most of
what a client does, and precisely the part that has to be right before it meets
hardware. What was exercised locally was the contents check; what shipped was
code nothing had run.
## The emulator now signs, and says so
QEMU's NSM implements the protocol but not the signing. Its source says as much
— *"we don't actually sign the data, so we use -1 as the 'alg' value"* — and -1
is not a COSE algorithm identifier, so the envelope is unverifiable while the
contents are perfectly good.
`testing::CosigningNsm` wraps the device rather than replacing it: `attest`
asks the emulator for a document, parses it, and re-signs that same payload with
a chain minted at boot. Contents stay the device's answers — the real PCR0 of
the image, the PCR16 the runtime measured its guest into, the caller's nonce.
Only `certificate` and `cabundle` are replaced, and they have to be, or the
document would name a key that did not sign it. Nothing else is intercepted:
entropy and every PCR operation go straight through, so a client comparing PCR0
against what `nix build` printed is still comparing against the device's report.
It does not make a document mean anything. The key is minted inside an image its
operator controls, so a verified document here says *"this image said so"* where
on hardware it says *"a Nitro enclave with this measurement said so"*. Two
things keep that from being mistaken for the real property: the root is fresh at
every boot, so there is nothing long-lived to paste into an app; and it is a
`testing` build flag, so the production binary has no `--cosign-attestations` to
be talked into signing its own attestations.
## The clients needed no changes, which exposed one bug
`--trust-root` *without* `--allow-untrusted-root` already yields
`Trust::ChainVerified` and already makes `--pcr0` and `--pcr16` mandatory. So
the local path is the production path, flag for flag.
That reachability made a latent bug visible: `nitro-attest` printed "chain
verified to the AWS Nitro root" for any pinned root. Harmless while the only way
to reach `ChainVerified` was the default root; a lie the moment a test root
could get there. It now names the file and says it is not AWS's.
## One bring-up, two consumers
`run-e2e.sh` was 930 lines of which two thirds was standing the stack up. That
half moves to `deploy/qemu-nitro/lib.sh`, and `dev-enclave.sh` is the second
consumer: same emulated machine, same ACME issuance from a real CA, same
attested boot, except it serves a component of your choosing and stays up.
deploy/qemu-nitro/dev-enclave.sh --guest path/to/component.wasm
It prints the three values a client pins and waits. Because both scripts share
the bring-up, what a client is developed against is what CI checks.
All thirteen e2e legs now verify by signature and chain against a pinned root.
Legs 7 and 8 re-pin after each reboot, since every boot mints a new chain.
## Found by running it, not by reading it
The MinIO container carried no label, because `docker update` has no
`--label-add` — so cleanup left it running. `minio-up.sh` takes `MINIO_LABEL`
now. S3 bucket names followed the run's name, but they are baked into the image
and therefore into PCR0, so the enclave looked in `e2e-data` while the script
had made `dev-data`. A `grep | head` inside an assignment killed the harness
silently under `pipefail` before the line it wanted had been written.
Backticks in an unquoted heredoc executed. And `--name ""` would have pointed
`rm -rf` at `target/qemu-nitro` itself.
Preflight also refuses to start on untracked `.rs`, `.toml`, `.wit` or `.nix`
files: Nix flakes copy only what git tracks, so an unadded module is invisible
to the build and fails twenty minutes later as `file not found for module`,
from inside a sandbox. `nix build … | tail -3` used to hide that error behind
the derivation summary; failures now print the compiler's own words.
docs/DEV_ENCLAVE.md covers what is and is not real about it.
docs/CLIENT_INTEGRATION.md is for a team bringing an existing client across.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Every assertion from the Android app was refused. Credential Manager signs `clientDataJSON` with the origin `android:apk-key-hash:<hash>`, never `https://<rp id>`, and the relying party was built with that one web origin and no way to add another. `--webauthn-allowed-origin` (repeatable; `S3FS_WEBAUTHN_ALLOWED_ORIGINS`, comma-separated) appends further origins through webauthn-rs's `append_allowed_origin`. Each is still compared exactly. This does not widen who can sign: Android lets an app claim that origin only after `https://<rp id>/.well-known/assetlinks.json` lists it, so the origin names an app the domain already vouched for. An Android origin is checked at boot, not at the first refused login: the hash must be 32 bytes of unpadded base64url. The likeliest wrong value is the colon-separated hex fingerprint keytool and the Play Console print, which names the right certificate in a form Android never sends, so the error says to convert it. Empty entries are dropped so an image with none still boots, and the flag without a relying party is refused rather than silently ignored. `deployment.nix` gains `webauthnAllowedOrigins`. Like `rpId` it is measured by PCR0: adding an app is a new image. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A platform authenticator creates a passkey only for an rp id whose domain publishes
assetlinks.json naming the app, and the app then claims `android:apk-key-hash:<hash>`. The emulator
image fixed its relying party at `enclave.test`, which publishes nothing, so no phone could register
against a dev enclave at all.
`eifQemu { rpId, allowedOrigins }` builds the emulator image for another relying party, exposed as
`lib.<system>.eifQemu` because flake outputs take no arguments, and `dev-enclave.sh --rp-id
--allowed-origin` builds and boots one. Both values are held to their shapes before they reach the
Nix expression. `passkey-client` is given the rp id, and the certificate is still `enclave.test`:
the two were only ever the same by default.
`packages.eif-qemu` is `eifQemu { }` and evaluates to the same image: the derivations differ only in
the flake's own source path, and the environment baked into the image is byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A failing `run-task` wrote its error only into the sealed task record, so the
console showed `status=Failed attempts=5` and nothing else. Finding the cause
meant making the guest print to stderr.
Every failed attempt is now logged at warn with the error it records, whether
it will retry or has reached `Failed`, and so is an attempt a restart
interrupted. The text is escaped (`?`) so a guest cannot forge log lines with
it; it is nothing the guest could not already print to its own output.
The recorded error also lost its causes: `to_string()` on an anyhow error is
its outermost context, so a trap or a deadline said what failed and not why. It
is now `{:#}`, still bounded to 512 characters.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The WIT said only that `task-id` "remains stable across retries". The runtime passes `<id>:<generation>:<occurrence>` (`Task::run_id`), so a guest that validates it as the id it enqueued refuses every run — which is how the cosigner's watch failed on every wallet. The format is now documented in both copies of `tasks.wit` and in BACKGROUND_TASKS.md: what each part is, that the enqueued id is everything before the first `:` (an id cannot contain one), and which to use for what. Documented rather than split into separate arguments, which would break every existing guest and need a new interface version. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`/auth/register/options` used webauthn-rs's defaults: `residentKey: discouraged` and no `authenticatorAttachment`. On Android below 14, Play services takes that literally and creates a non-discoverable security-key credential outside Password Manager, which One Tap then reports as "Cannot find a matching credential". Every client would hit it; the app was working around it by rewriting the options itself. Registration now asks for a discoverable platform passkey, through webauthn-rs's constructor for exactly this (`residentKey: required`, `requireResidentKey`, `authenticatorAttachment: platform`, user verification still required). Its feature flag adds no code beyond that constructor. It also stops requesting credProtect, which Android does not support. Ruling out roaming security keys and cross-device registration costs nothing: the clients are native apps, never a browser. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A guest has had no outbound network at all: `EgressPolicy::Denied` refused every `wasi:http` request. That is still the default. But some guests cannot do their job from inside it — a wallet cosigner that holds a pre-signed renewal has to reach the ASP it renews with, and the only alternative was routing the round through a phone that may be off. So a deployment may name origins — scheme, host and port, nothing else — and guests may send requests to exactly those (`guestEgressOrigins`; `S3FS_GUEST_EGRESS_ORIGINS`, `--guest-egress-origin`). Compared exactly: no subdomain, no other port, no plaintext to an HTTPS origin, no wildcards or paths; a malformed entry refuses the boot. HTTPS is verified against the same webpki roots the FCM client uses — now one shared `web_pki_client_config` rather than two stores to audit — and requests go over HTTP/1.1. Anything else is refused as before, with the same warning. Requests and background tasks get the policy alike, since renewing on a schedule is exactly the work that happens with nobody connected. The list is image environment, so PCR0 covers it: a client learns where a guest can send traffic from the attestation that tells it what the guest is. It is set only when non-empty, so an image that names no origins is the image it was. Two settings beside it, for the same kind of guest: `backgroundTimeoutSecs`, because a round waits on the ASP's schedule and the default is 30 seconds, and `guestEnv`, which reaches the guest because it inherits the image environment minus `AWS_*` and `S3FS_*` — how a cosigner learns its ASP's address. `eifQemu` takes all three. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ch a local service Every dev-enclave start booted into genesis: MinIO kept its data inside a `--rm` container, so a restart or a guest rebuild threw away every tenant, passkey and file a guest wrote — and a phone app under test had to be wiped and onboarded again each time. `--keep-store` puts MinIO's data in target/qemu-nitro/<name>-store, outside the run directory that is cleared on every start, so the next start resumes it and a new guest boots as an upgrade of the same store. Pebble mints a new CA every start while a kept store serves the certificate an earlier one issued, so every root and intermediate seen is kept and `pebble-root.pem` carries all of them — the served-chain check verifies whichever issued it. The store directory gets a stable `id` for clients that keep state per store, since the trust root is new every boot. `--fresh` discards it, behind the same guard RUNDIR has. Without the flag nothing changes. `--guest-egress ORIGIN`, `--background-timeout SECS` and `--guest-env NAME=VALUE` build the image with the new `eifQemu` settings, each held to a shape that cannot break out of the Nix expression. The machine running the script is 192.168.127.254 from inside the enclave. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…speak first
A guest has no execution context between invocations, so it cannot own a socket
that outlives a call. And a counterparty with no passkey for a tenant can never
call in, because the gate is checked before the path and fails closed. Neither
party could reach the other by the means each already had.
`enclave:streams` closes that. A guest asks for a connection to an origin its
image allows, and every message the far side sends becomes one invocation of
`on-message` — exactly as a due task becomes one invocation of `run-task`. The
guest stays stateless; the runtime keeps the socket.
What reconnects after a drop, a timeout or a restart is a supervisor task in the
runtime, started from the records on disk before anything is served. Not sealed
guest state, which only says a connection should exist, and not a scheduler.
Three things that are easy to get wrong, and were:
- The wire id names the tenant as well as the stream. A stream id is
tenant-local by contract, so every tenant a guest serves opens one under the
same name; a far side keeping one connection per id would have each customer
close the last one's, and would answer down the wrong socket.
- `send_direct` hands back the connection worker with the response. It is
abort-on-drop, so returning the response alone ended the body the moment the
call returned — invisible for a request/response exchange, fatal for a
stream held open.
- `open` wakes the supervisor with `notify_one`, not `notify_waiters`. The
loop is not a registered waiter while it is collecting records, and a wakeup
that landed in that window was simply lost.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…Encrypt integration
… streams Store: a session claims its transaction group in the root chain before it seals a block, and again after any transaction that did not end in a root. Nonce uniqueness no longer rests on probing host-deletable slabs, and of two mounts at one tip only one ever encrypts. The orphan probe is gone. Filesystem: a failed sync keeps its dirty buffers for the retry; a directory can no longer be renamed into its own subtree. Backend: retained reads follow every page of ListObjectVersions, so delete markers cannot push the real version out of sight. Deploy: the parent role may list versions, read by version and put retention headers, which boot needs; the deny on PutObjectRetention bought nothing under COMPLIANCE. Runtime: dropping a guest state closes the files it left open; a stream supervisor is replaced when its record changes under the same key; the SSE framer accepts CRLF and CR line endings. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…tream reopens Store: publishing a root reads the retained version back after the conditional PUT, and loses unless that version is ours. S3 accepts `If-None-Match: *` over a delete marker, so a host that hid the winner's claim could hand the same transaction group — the same nonces, the same slab names — to a second mount. Claims, commits and format all go through publish, so all three are covered. Backend: a retained read takes the oldest version by list order, not by LastModified. S3 lists a key's versions newest first, and a timestamp tie left the newer one in front. Runtime: dropping a guest state releases the files it left open without flushing them. The flush ran after the tenant's lock had passed on, and could land on top of what the next request committed. A guest that did not finish loses what it had not synced, as on a crash. Streams: a record carries the generation of the open that made it, so a close and reopen to the same origin between two supervisor passes gets a new connection — not the old one, whose status the close had removed, leaving every send refused. The stream suite now runs in the guest CI stage. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The README still described the filesystem this repository started as. It now covers the runtime as built: what exists, a local quick start, the architecture and trust boundaries, the encrypted filesystem, boot and key release, serving and verification, passkeys and tenants, background work and held connections, SQLite, configuration, testing, deployment, embedding the storage engine directly, and the limits that remain. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Scheduler: a recurring task's second success is waited for rather than inferred from its output file. The guest writes the file before it returns and the queue records the success after, so stopping the worker between the two left `attempts` at 1 — the failure that ended both CI runs of the e2e job before MinIO or the emulated enclave had started. MinIO: two mounts, a delete marker over the first one's claim, and the second refused. On real S3 the damage without the read-back is worse than nonce reuse: the loser's slabs overwrite the winner's committed ones, so the test's proof is the winner's data surviving. README: the CI table names the suites these land in. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…step rustfmt had failed on egress.rs and notify/transport.rs since 909ebee and on serve/http.rs since before it, which stopped the unit job at its first command on every run of this branch. Behind it, clippy denied the redundant parentheses in the stream registry's key type and `ServeHandle::streams`, which nothing calls; it is gone rather than allowed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Docker Hub refuses `minio/minio` and `minio/mc` without a login and quay.io has no tags, so the e2e job's storage stage failed before a test ran, and its QEMU stage would have failed the same way. It had never got that far before. scripts/minio-image.sh builds the pinned server release and the mc release before it from their GitHub tags, into `enclave-runtime/minio` under the tag the testcontainers module pins. The integration tests take it by name, minio-up.sh runs server and mc from it, and the harness uploads guests with it. It builds only when missing, in about five minutes. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… lint job rust-toolchain.toml named `stable`, which floats. CI denies every clippy warning, and 1.98.0 arrived with `chunks_exact_to_as_chunks`: a branch that passed on 1.97 failed with no change to it, while a workstation still on 1.97 saw nothing. Now CI and every checkout run the same clippy, and a newer release is a commit of its own. The one thing 1.98.0's clippy asks for: indirect blocks are split with `as_chunks`. The length is checked first, so its remainder is always empty, as `chunks_exact`'s was. The Nix enclave build is untouched; flake.lock already pins its toolchain. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
It relays every store lookup through the Actions cache API. That API rate-limited it, it answered Nix with HTTP 418, and Nix treated the refusals as fatal: the enclave image build failed on "rate limit exceeded" with nothing wrong in the build. It was also the source of the FlakeHub authentication error in every run's log. nixpkgs still comes from cache.nixos.org. What goes is only the reuse of this repository's own derivations between runs, which are rebuilt each time. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ing it `escapeShellArg` turns a path into a string with `toString`, which carries no context: the EIF derivation named the flake source's copy of pebble/ca.pem without depending on it. Nix upstream copies the whole source eagerly, so the file happened to be there and the QEMU image built. Determinate Nix, which CI installs, evaluates flakes lazily and never put it in the store, so the build failed on `cp: cannot stat` — after warning that the derivation referenced a store path "without a proper context". Interpolated instead, the file is its own store path and an input. The image's contents are unchanged: PCR0 is the same. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.