Conversation
…sockets On kernels without SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV (RHEL 9 / RHCOS 5.14), the broker cannot safely write an accepted peer address into workload memory, so accept/accept4 with a peer-address buffer failed with EOPNOTSUPP. Static binaries and Go servers, which issue the raw syscall, could not accept connections at all. Move local acceptance into the kernel and replace per-accept inspection with standing kernel confinement: - Bind every broker-created TCP/UDP socket to the loopback device before injection and verify the binding. Accepted sockets inherit it, so they can neither receive routed ingress nor emit routed egress. - Stop notifying accept/accept4 and remove the accept workers, the SIGUSR2 accept-interrupt monitor, and the 64-accept ceiling. - Continue getpeername natively for every descriptor except relayed connections. - Deny interface-selection socket options and MSG_FASTOPEN sends from scalar syscall arguments, so confinement does not depend on the capability state of the user namespace that owns the network namespace and covers unregistered descriptors. - Drop loopback-interface ingress on the non-loopback TCP control listener before it listens, so a workload cannot reach it by reconnecting a natively accepted socket. - Mark descriptors above stdio close-on-exec in every workload pre_exec path. - Require an active confinement probe at qualification and carry it as required authenticated audit evidence. Document OpenShift 4.19 as the minimum release: RHCOS kernels for 4.16 through 4.18 are built without Landlock. Closes #4058 Signed-off-by: Drew Newberry <anewberry@nvidia.com>
- Reject a TCP boundary control listener on a loopback address, including IPv4-mapped loopback. Production drivers bind the unspecified address; a loopback listener would be reachable from workload sockets the broker does not track. - Extend the confinement probe to IPv6 and prove that an accepted socket keeps its loopback binding after an AF_UNSPEC disconnect. - Pin the accepted-peer contract with tests: a loopback-bound listener admits only clients in the sandbox network namespace, and workload sockets cannot bind a non-loopback source address. - Correct the support matrix: accepted connections are bounded by per-process descriptor limits, the runtime PID limit, and the sandbox memory limit, not by a broker limit. Signed-off-by: Drew Newberry <anewberry@nvidia.com>
|
🌿 Preview your docs: https://nvidia-preview-pr-4150.docs.buildwithfern.com/openshell |
|
Label |
PR Review StatusThe independent review found no blocking code defects in the loopback confinement and native-accept change. The support matrix and OpenShift documentation cover the changed behavior and kernel requirements. Action required: rerun all jobs in Branch E2E Checks after its current attempt finishes, so it picks up Blocking findings: None. Carried findings: None. Non-blocking suggestion: Consider extending the native-accept fixture to assert Gator metadata
|
Socket confinement now uses socket2 for device binding and the ingress filter, and rustix for interface lookup and AF_UNSPEC disconnect, so the module contains no unsafe code. The control-listener filter is no longer locked: no safe API exposes SO_LOCK_FILTER, and the listener descriptor never leaves the trusted sandbox process, which marks every descriptor above stdio close-on-exec before running workload code. Tests added for native accept use rustix and socket2 instead of raw libc calls. rustix issues accept4 and getpeername as raw syscalls, so the direct-syscall test keeps its meaning. The close-on-exec sweep test runs in a re-executed test process instead of a forked child. The close_range(CLOSE_RANGE_CLOEXEC) syscall in the pre_exec hook remains the only unsafe added by this branch; neither rustix nor nix wraps it. Signed-off-by: Drew Newberry <anewberry@nvidia.com>
PR Review StatusDrew, I checked your follow-up replacing raw socket calls with socket2 and rustix, including device binding, the control-listener filter, direct-syscall accept coverage, and the isolated descriptor-sweep test. The independent delta review found no blocking defects, and the existing support-matrix and OpenShift docs still cover the behavior. The earlier test-dispatch blocker is cleared: current-head Branch E2E Checks is running Docker, Podman, VM, and Kubernetes suites with Blocking findings: None. Carried findings: None. Gator metadata
|
The workload seccomp filter denied only AF_PACKET, AF_BLUETOOTH, and AF_VSOCK, and the broker continued socket() for every non-INET family. Several protocol families, including AF_RXRPC, AF_SMC, and AF_KCM, carry traffic over kernel-owned sockets that the broker never creates and that are not bound to loopback. Allow only AF_UNIX, AF_NETLINK (still limited to NETLINK_ROUTE), and the brokered AF_INET/AF_INET6 families, and restrict socketpair(2) to AF_UNIX. Both decisions use the scalar domain argument. The broker independently refuses non-INET families other than AF_UNIX and AF_NETLINK with EAFNOSUPPORT. Signed-off-by: Drew Newberry <anewberry@nvidia.com>
PR Review StatusDrew, I checked your follow-up restricting workload socket families in seccomp and independently in the broker, the socketpair restriction, and the new behavioral tests. The independent delta review found no blocking defects. The updated security docs describe the allowlist and retain the NETLINK_ROUTE exception. Current-head Branch Checks and Branch E2E Checks are running. Gator will inspect the results next cycle. Blocking findings: None. Carried findings: None. Gator metadata
|
After native local accept, only two broker paths still wrote into workload memory: getpeername on relayed connections and the per-message lengths of a first sendmmsg to the DNS relay. Legacy mode existed only to refuse those writes on kernels without WAIT_KILLABLE_RECV, and the getpeername refusal broke CPython TLS on RHEL 9. The broker now never writes workload memory: - getpeername is no longer mediated and reports the kernel peer on every kernel; relayed connections report the loopback relay address. - A first send to the DNS relay pins the broker's socket copy to the relay and continues the syscall, so the kernel performs the send and writes any per-message results. Remove the listener mode, the task-memory write path and probe, and the task_memory_write, cancellation, and task_memory_writes_disabled audit evidence fields. WAIT_KILLABLE_RECV is still used when available and is reported for diagnostics, but no mediation decision depends on it. Signed-off-by: Drew Newberry <anewberry@nvidia.com>
…safe The broker no longer writes workload memory, so killable notification waits only reduced how often a signal restarts a notified syscall. RHEL 9 kernels never had them, so the handlers must tolerate restarts anyway. Install the listener without the flag on every kernel instead of special-casing newer ones. Handlers now check that the notification is still live immediately before each side effect (bind, connect of the retained socket, listen, and the first DNS send), and answer a restarted operation the broker already completed as the kernel would: a repeated TCP connect returns EISCONN, repeating the same UDP association succeeds, a repeat of a completed bind succeeds, and a failed relay reports its errno. Signed-off-by: Drew Newberry <anewberry@nvidia.com>
… exited Teardown waited only for registered process groups to disappear. A root is unregistered once reaped, so a descendant that ignored SIGTERM could outlive it while termination reported success and never sent SIGKILL, violating the bounded-termination requirement. The sandbox now becomes a child subreaper when it is not PID 1, so orphaned descendants stay in its tree and are reaped. Termination waits until no live descendant remains, and every scanned process is signalled through a pidfd after confirming its start time, so a reused PID is never signalled. Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Freezing stops every workload process with SIGSTOP. tgkill and rt_tgsigqueueinfo were not mediated, and mediated kill, rt_sigqueueinfo, and tkill passed SIGCONT through, so a workload process that was not yet stopped could resume the others while the supervisor recovered. Mediate tgkill and rt_tgsigqueueinfo like tkill: refuse targets in the sandbox thread group, report a thread outside the named group as missing, and continue otherwise. The boundary marks the broker frozen before stopping the workload and clears it after resuming; while frozen, workload requests to send SIGCONT fail with EPERM. Signed-off-by: Drew Newberry <anewberry@nvidia.com>
PR Review StatusDrew, I checked your follow-up removing workload-memory writes and the legacy listener mode, including DNS send continuation and restart-safe bind, connect, and listen handling. The independent Critical-only delta review found no new Critical defects; no carried blocking findings remain. The support-matrix and OpenShift updates cover the changed kernel behavior. Required E2E is blocked by registry infrastructure: the gateway-image push failed with GHCR Action required: a maintainer should rerun all jobs in current-head Branch E2E Checks after GHCR access recovers. The earlier E2E Label Help instruction targets an older head; the current run already picked up Blocking findings: None. Carried findings: None. Gator metadata
|
The native UDP send path re-ran the workload's own sendmsg, so per-message ancillary data such as IP_PKTINFO rode along; only the loopback destination contained a routing override. The broker now reads the datagram and sends it from its retained, relay-connected socket, which it builds without ancillary data, so a per-message override cannot redirect the packet. It writes nothing back into workload memory. A send carrying control data is refused with EOPNOTSUPP rather than silently stripped. Signed-off-by: Drew Newberry <anewberry@nvidia.com>
…ronment The canonical process inherits the sandbox environment so the image's own ENV reaches the workload, then removed only a denylist of credential variables. Other variables in the reserved OPENSHELL_ namespace, such as the serialized user environment and the log level, still reached the workload. Remove every inherited OPENSHELL_ variable before applying the declared environment, and restore OPENSHELL_SANDBOX=1. The image's ordinary ENV and declared variables are unaffected; declared variables cannot use the reserved namespace. Signed-off-by: Drew Newberry <anewberry@nvidia.com>
GPU mode granted read-write access to all of /proc so CUDA's cuInit could write thread names to /proc/<pid>/task/<tid>/comm. Keep /proc read-only. Open mediation now serves a writable open of the caller's own thread comm file: the broker opens it and injects the descriptor, with no syscall continued. The kernel accepts a comm write only from the target's own thread group, so a substituted path or reused thread ID cannot rename another process's thread. Any other /proc write is left to Landlock, which denies it. Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Sending DNS datagrams from the broker's socket required the broker to keep its copy, but the connect paths to the relay and to loopback still released it. glibc connects the resolver socket and then sends A and AAAA together with sendmmsg, which is mediated, so resolution failed. Release the copy on those connect paths only for TCP. Add a regression test for connect followed by sendmmsg that fails fast rather than hanging. Signed-off-by: Drew Newberry <anewberry@nvidia.com>
The loopback connect path held the registry lock while polling a TCP connect for up to five seconds on the single notification dispatcher. A workload connecting to a busy local listener stalled every other mediated syscall, including opens and signals. Connect a duplicate of the retained socket without holding the lock. A nonblocking socket gets the native EINPROGRESS and the kernel completes the handshake on the shared socket; a blocking socket waits on a bounded worker thread. Signed-off-by: Drew Newberry <anewberry@nvidia.com>
A panic in a notification handler unwound the single broker thread while the health flag still read healthy, so every later blocked workload syscall hung until the sandbox was killed. Run each dispatch under catch_unwind: a panicking handler fails only that syscall with EIO and the broker keeps mediating. If the broker thread ever exits, mark it unhealthy so dependent operations fail closed instead of blocking. Signed-off-by: Drew Newberry <anewberry@nvidia.com>
…nding Sending mediated DNS datagrams from the broker's own socket required keeping a broker handle for each DNS socket's whole life, which regressed real name resolution through reclaim and resource accounting that the mock-based unit tests did not exercise. Revert to continuing the kernel send, but refuse a send that carries ancillary control data (msg_controllen != 0) at read time, so a per-message routing override such as IP_PKTINFO cannot ride a mediated DNS send. A loopback destination contains an override that races the check. Full broker-side UDP mediation is left to a separate change with deployment e2e. Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Fixes: - The blocking local-connect worker no longer toggles O_NONBLOCK on the open file it shares with the workload; it waits with a plain connect. - tgkill and rt_tgsigqueueinfo are notified only when they send SIGCONT, so ordinary thread signals such as Go preemption stay in the kernel. The frozen flag is re-read immediately before delivery. - A shell redirect to the caller's own thread comm file works again; O_CREAT and O_TRUNC are no-ops there and O_EXCL returns EEXIST. - The teardown scan is a linear walk and fails closed when /proc cannot be read. Simplifications: - Remove tests that need infrastructure outside the repo or prove nothing: the topology harness tests, the static-server benchmark, the direct-syscall accept test, the comm rename test without Landlock, a serde-default test, and the test-only panic hook in setsockopt. - Drop the unused Failed connect outcome, the redundant read-back after binding to loopback, redundant OPENSHELL_SANDBOX settings, and stale or duplicated comments; reuse the reserved environment prefix constant. - Tighten the support matrix and OpenShift wording. Signed-off-by: Drew Newberry <anewberry@nvidia.com>
|
Label |
PR Review StatusDrew, I checked your follow-up hardening process-tree teardown, frozen-workload signal handling, DNS sends, broker panic and slow-connect handling, inherited environment filtering, and GPU thread-name writes with read-only I added Action required: rerun all jobs after the current attempt finishes so GPU E2E is dispatched. Gator will retry the authorized rerun next cycle; a maintainer may also perform it. Blocking findings: None. Carried findings: None. Gator metadata
|
Processes that inherit the same stdin share one open file description, so any of them can make it nonblocking for all. `sandbox exec` then read no input yet, got EAGAIN, and failed with "Resource temporarily unavailable (os error 11)". Parallel e2e tests inherit the runner's stdin, which made credential_gating fail intermittently. The stdin reader now blocks in poll() until input or end of file arrives instead of treating EAGAIN as fatal, and it leaves the shared descriptor's flags alone. The e2e exec helper also stops inheriting the runner's stdin, since its commands take no input. Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Summary
Workload
acceptnow runs natively in the kernel instead of through the seccomp broker. This fixes servers on RHEL 9 / RHCOS 5.14, where peer-addressacceptfailed withEOPNOTSUPP. It also removes legacy read-only mode and hardens the sandbox runtime. Supersedes #4087.Related Issue
Closes #4058
Changes
Native local accept
lo(SO_BINDTODEVICE) before injection. Accepted sockets inherit the binding, soaccept/accept4are no longer mediated.SIGUSR2interrupt monitor, and the 64-accept limit.socket_loopback_confinement).One mode on every kernel
getpeernameis no longer mediated; relayed connections report the loopback relay address.WAIT_KILLABLE_RECVare removed, along with the related audit fields. Handlers check the notification is live before acting and answer repeatedconnect/bindthe way the kernel would.Hardening
AF_UNIX,AF_NETLINKand the brokeredAF_INET/AF_INET6.socketpairis limited toAF_UNIX.MSG_FASTOPENsends are denied, and DNS sends carrying ancillary data are refused.pre_execmarks descriptors above stdio close-on-exec.SIGCONTis refused while frozen, andtgkill/rt_tgsigqueueinfocarryingSIGCONTare mediated.OPENSHELL_*variables no longer reach the workload./procread-only; thread-name writes to the caller's owncommfile are served by open mediation.CLI:
sandbox execno longer fails withEAGAINwhen another process has made its inherited stdin nonblocking; it waits for input instead. The e2e exec helper no longer inherits the runner's stdin. Together these fix an intermittentcredential_gatingfailure.Docs: the support matrix adds the socket-binding requirement and drops legacy mode. The OpenShift page states that 4.19 is the minimum, because the RHCOS kernels in 4.16–4.18 lack Landlock. The security page describes the family allowlist.
Testing
mise run lint; unit suites foropenshell-isolation-interface,openshell-sandbox-backendandopenshell-sandbox. The only failure ismetadata_loopback_connect_is_relayed_to_supervisor, which also fails onmain(bug: sandbox unit test metadata_loopback_connect_is_relayed_to_supervisor times out when run with other broker tests #4149).SIGCONTrefusal, environment stripping, comm path parsing and the slow-connect stall.credential_gatingfailure (EAGAINfromsandbox exec) is fixed in this PR, with a deterministic regression test. Earlier commits also passed Podman CI (31/31), VM (10/10) and Kubernetes isolation (94/95; bug: Kubernetes isolation e2e canonical_main_connect_recovers_its_ssh_transport times out reattaching after transport loss #4148).mainwithaccept4: operation not supportedworks on this branch.Checklist
🤖 Generated with Claude Code