You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Drop image-volume admission (split to a follow-up), the Docker and core
mount-validation refactor, redundant per-arm Podman mount checks, and
unrelated doc and test churn. Keep image VOLUME masking validation for
custom workdirs.
Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
Copy file name to clipboardExpand all lines: architecture/compute-runtimes.md
+20-31Lines changed: 20 additions & 31 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -9,12 +9,10 @@ Podman provisions a paired workload and supervisor container using its native
9
9
libpod API. The workload uses `network=none`; the external supervisor alone joins
10
10
the configured network. A per-sandbox named volume carries their mutually
11
11
authenticated gRPC Unix socket, with supervisor credentials kept in its separate
12
-
filesystem. The supervisor and final sandbox runtime run as the resolved non-root
13
-
identity with all capabilities dropped. Only the managed `/sandbox` fallback uses
14
-
a trusted root bootstrap to prepare its driver-owned workspace before dropping
15
-
irreversibly to that identity. The containers do not share PID, mount, or network
16
-
namespaces. Podman owns paired lifecycle and health; the common protocol owns
17
-
process, identity, TCP, DNS, and forwarding semantics.
12
+
filesystem. Both containers run as the resolved non-root identity with all
13
+
capabilities dropped. They share only a user namespace for volume ownership,
14
+
not PID, mount, or network namespaces. Podman owns paired lifecycle and health;
15
+
the common protocol owns process, identity, TCP, DNS, and forwarding semantics.
18
16
19
17
## Driver Contract
20
18
@@ -307,7 +305,7 @@ delete, reconciliation removes the row; otherwise it can remain `Deleting`.
307
305
| Runtime | Best fit | Sandbox boundary | Notes |
308
306
|---|---|---|---|
309
307
| Docker | Local development with Docker available. | Capability-free workload container. | Uses `network_mode=none`; a separate capability-free supervisor container mediates egress and access over a private daemon-local Unix socket volume. |
310
-
| Podman |Local development with Podman available. |Capability-free workload container. |Uses `network=none`; a separate capability-free supervisor container mediates egress and access over a private Unix socket volume. |
308
+
| Podman |Existing rootless driver. |Container. |Not converted by this isolation stack. |
311
309
| Kubernetes | Cluster deployment through Helm. | Capability-free sandbox Pod. | Always creates a namespace-wide empty-egress workload NetworkPolicy and a separate capability-free supervisor Pod over mutually authenticated TLS. It requires an enforcing CNI and trusted sandbox namespace; the Kubernetes API does not attest policy enforcement. |
312
310
| VM | Experimental microVM isolation. | Per-sandbox libkrun or QEMU VM. | The NIC-less guest runs `openshell-sandbox` as PID 1; host `openshell-supervisor` owns gateway networking and reaches the guest over vsock. |
313
311
| Extension | Out-of-tree drivers operated alongside the gateway. | Whatever boundary the driver implements. | Selected by a custom `compute_drivers = ["<name>"]` entry with `[openshell.drivers.<name>].socket_path`, or at launch time by pairing `--drivers <name>` with `--compute-driver-socket=<path>`. A launch-time endpoint may use a canonical built-in name to preserve its driver-config key while replacing in-process construction. The gateway connects to an operator-provisioned UDS, snapshots `GetCapabilities`, and dispatches all sandbox lifecycle calls through `compute_driver.proto`. The driver process and socket lifecycle are operator-owned; the gateway does not spawn, supervise, or remove unmanaged extension drivers. The trust boundary is the socket's filesystem permissions: the operator must ensure only the gateway uid can read/write it. |
@@ -392,7 +390,7 @@ Drivers deliver the two binaries to separate trust domains:
392
390
| Runtime | Delivery model |
393
391
|---|---|
394
392
| Docker | A digest-pinned daemon-local volume supplies `openshell-sandbox`; the companion image runs `openshell-supervisor`. |
395
-
| Podman |The driver pins `sandbox_runtime_image` and `supervisor_image` to image IDs. The former supplies `openshell-sandbox`; the latter is passed as the companion container's image and runs `openshell-supervisor`. |
393
+
| Podman |Existing driver behavior; not converted by this stack. |
396
394
| Kubernetes | A non-root init container stages `openshell-sandbox` into a memory volume; a directly managed Pod runs `openshell-supervisor`. |
397
395
| VM |`openshell-sandbox` is embedded in the guest rootfs; a separately digest-checked native `openshell-supervisor` runs on the host. |
398
396
| Extension | Defined by the out-of-tree driver. |
@@ -408,33 +406,24 @@ The gateway preserves whether each policy process field was omitted and passes
408
406
the admitted selectors to the driver. The driver resolves one exact UID, GID,
409
407
and supplementary-group set before creating the immutable workload:
410
408
411
-
- Docker pins the image ID, resolves policy selectors against the image's
412
-
`/etc/passwd` and `/etc/group`, and validates its OCI working directory.
413
-
- Podman pins the image ID, resolves policy selectors against the image's
414
-
`/etc/passwd` and `/etc/group`, and validates its OCI working directory.
409
+
- Docker and Podman pin the image ID, resolve policy selectors against the
410
+
image's `/etc/passwd` and `/etc/group`, and validate its OCI working
411
+
directory.
415
412
- Kubernetes uses platform-resolved numeric values, including OpenShift
416
413
namespace ranges.
417
414
- VM uses the configured numeric guest identity.
418
415
419
-
UID/GID zero and `u32::MAX` are invalid. Agent commands run as the resolved
420
-
non-root user. For Podman's managed `/sandbox` workspace, trusted setup briefly
421
-
starts as root to prepare the workspace, then switches to that user before
422
-
reading bootstrap material or accepting commands. Identity-changing policy
423
-
updates require sandbox recreation, while other policy updates remain live.
424
-
425
-
Docker uses an absolute OCI working directory as the workspace. Empty, root,
426
-
and explicit `/sandbox` values select `/sandbox`; other paths must already
427
-
exist without symlink or reserved-mount collisions and must be usable by the
428
-
resolved identity.
429
-
430
-
Podman reads the image user and working directory from one pinned image. Empty,
431
-
`/`, and explicit `/sandbox` values use the managed `/sandbox` workspace. A
432
-
custom path must be absolute, normalized, and outside system and OpenShell
433
-
reserved paths. Image and driver mounts cannot cover the workspace. Custom
434
-
paths use the image's container filesystem directly; the final non-root user
435
-
must be able to write the existing directory. The managed `/sandbox` fallback
436
-
uses a driver-owned workspace volume.
437
-
Kubernetes and VM use `/sandbox`.
416
+
UID/GID zero and `u32::MAX` are invalid. The sandbox and every child start with
417
+
the resolved identity and zero capability masks; neither process performs an
0 commit comments