Skip to content

feat(sandbox): restore peer addresses on kernels without WAIT_KILLABLE_RECV - #4087

Open
russellb wants to merge 7 commits into
NVIDIA:mainfrom
russellb:fix/4058-legacy-accept-peer-address/russellb
Open

russellb wants to merge 7 commits into
NVIDIA:mainfrom
russellb:fix/4058-legacy-accept-peer-address/russellb

Conversation

@russellb

@russellb russellb commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Restore libc-based server workloads on kernels without SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV (introduced in Linux 5.19). In legacy read-only mode, the broker cannot safely write output buffers into workload memory. The compatibility shim rewrites address-bearing accept/accept4 into a null-address accept followed by kernel-handled getpeername.

Closes #4058.

Implementation

  • Embed a freestanding, per-architecture shared object. It resolves the already-loaded libc's accept4 wrapper to retain libc's cancellation handling, then uses raw syscalls for peer lookup and error-path close. Failed peer lookup closes the accepted descriptor; ENOTCONN becomes ECONNABORTED, and invalid pointers return EFAULT without userspace dereferences.
  • Probe executable memfd support only for legacy listeners. Each entrypoint or operator exec gets a fresh sealed anonymous inode rather than a file under /run. Descriptors remain close-on-exec in the boundary; only the intended forked child makes its descriptor inheritable. A workload cannot change another launch's shim contents or permissions. No shim-specific Landlock admission or user ruleset is added.
  • Compose LD_PRELOAD after the launch environment is assembled, preserving the effective provider, workload, or session value. Skip direct foreign-class/architecture ELF executables, and avoid blocking executable probes on FIFOs or devices. Installation failure is non-fatal and logged; the broker retains fail-closed behavior.
  • Reject unsupported raw address-bearing accepts before consuming a queued client.
  • Continue getpeername in the kernel for local, accepted-local, and connected DNS sockets. Legacy relayed outbound sockets report the kernel's loopback relay address and port, not the original destination. Modern listeners still substitute the original outbound destination. Unconnected registered sockets return ENOTCONN.
  • Resolve target compilers through cc, including Cargo's CC/CFLAGS rebuild tracking. Compile C test fixtures with that target compiler. Add automated ELF, peer-error, cancellation, descriptor-isolation, Landlock, and full forced-legacy broker/preload regressions.
  • Update the crate README and support documentation. Remove PR-specific reproducer files; the slide deck is not tracked.

Coverage and limits

The shim interposes dynamic libc symbols. CPython and Bun address-bearing accepts are covered; Node's null-address accept does not need the shim. Go, static binaries, secure-execution loaders, and direct syscalls do not gain address-bearing accept support from it. Kernel-handled peer queries do not depend on the preload.

Descendants must retain the shim descriptor and access to /proc/self/fd. Descendants that close it, hide /proc, or invoke a foreign-architecture loader must remove its LD_PRELOAD entry themselves; arbitrary descendant execs cannot be checked by the boundary and may otherwise emit loader warnings. Connection authorization continues to use the requested destination and broker-observed addresses, not shim output.

Validation

  • Local macOS Podman, non-root Linux account: 198 sandbox library tests passed. The shim's 13 behavioral/ELF tests and 7 unit tests passed; anonymous loading also passed with /tmp mounted noexec.
  • OpenShift/RHCOS kernel 5.14.0-570.141.1.el9_6.x86_64, randomized non-root uid, all capabilities dropped, RuntimeDefault seccomp: legacy broker/preloaded CPython, queued-client preservation after raw accept rejection, concurrent-session descriptor isolation, Landlock, environment, and all 20 shim tests passed.
  • Zig 0.14.1 release build for x86_64-unknown-linux-musl passed. Inspection of the actual release object confirmed no DT_NEEDED, no text relocations, a non-executable stack, and exactly accept/accept4 exports. Dynamic Alpine musl cancellation smoke matrix passed.
  • Scoped Linux and macOS Clippy with warnings denied, Rust formatting, whitespace checks, and Fern docs validation passed. Fern reported three existing warnings and zero errors.

The OpenShift tests ran the real broker and preload directly in a test pod. A separate full gateway deployment could not start because the cluster's storage provisioner was unavailable; this validation does not claim a completed gateway E2E run.

…ed sockets

getpeername(2) was always answered by writing the sockaddr into the
workload's address space from the broker. On kernels without
SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV the broker runs in LegacyReadOnly
mode, where that cross-process write fails closed with EOPNOTSUPP, so
getpeername failed for every socket.

For SocketState::Local and SocketState::AcceptedLocal the descriptor is
genuinely connected to the recorded peer, so the kernel's own answer is
identical to the broker's. Respond with SECCOMP_USER_NOTIF_FLAG_CONTINUE
for those states and let the kernel write the sockaddr in the workload's
own address space. That needs no cross-process task-memory write and
therefore works on kernels before 5.19.

SocketState::Connected keeps the existing substitution. Those descriptors
are connected to a loopback relay rather than to the destination the
workload requested, so the kernel's answer would leak the relay address.
A regression test pins that distinction.

This is the path libuv-based runtimes depend on: libuv passes a null peer
address to accept4(2) and resolves the peer lazily through getpeername(2)
when the application reads remoteAddress.

Refs NVIDIA#4058

Signed-off-by: Russell Bryant <rbryant@redhat.com>
…E_RECV

Kernels before Linux 5.19 lack SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV, so
the broker cannot hold a notified workload thread in a kill-only wait and
cannot safely write into workload memory. It fails closed with EOPNOTSUPP
on every mediated syscall whose result is an output buffer. accept(fd,
&addr, &len) has exactly that shape, so server workloads on RHEL 9.x and
RHCOS see EOPNOTSUPP where they expect a connection.

Add openshell-accept-shim, a freestanding preloadable library that rewrites
an address-bearing accept into accept4(fd, NULL, NULL, flags) followed by
getpeername on the accepted descriptor. The first call has no output buffer,
so the broker injects the descriptor without touching workload memory; the
second is answered with SECCOMP_USER_NOTIF_FLAG_CONTINUE, so the kernel
writes the address into the caller's own buffer. Because the workload writes
its own memory, the cross-process TOCTOU that WAIT_KILLABLE_RECV exists to
close does not apply.

The sandbox installs the object at /run/openshell-compat only when the
listener reports writes disabled, and composes LD_PRELOAD so a workload's
own value is preserved. Installation failure is non-fatal and logged as an
OCSF config state change; the sandbox keeps the existing fail-closed
behavior.

Landlock gates open independently of the file mode, so a world-readable
object is not by itself reachable: without an explicit admission the loader
reports "cannot open shared object file" and silently drops the shim.
Workload children are launched from more than one place -- the sandbox
entrypoint and sandbox exec through the boundary -- and each prepares
Landlock with its own runtime paths. Admit the object and its directory
inside prepare_child_sandbox, the single production entry point, so a
launch path cannot set LD_PRELOAD without also letting the loader open what
it points at.

Resolve the shim's C compiler through the cc crate so the standard
CC_<target>, TARGET_CC, and CC overrides apply, along with the per-target
wrappers cargo-zigbuild installs when release binaries are cross-compiled
from a non-Linux host. Verify the emitted object's ELF machine against
CARGO_CFG_TARGET_ARCH and fail the build on a mismatch: a misresolved
compiler otherwise produces a valid object for the wrong architecture,
which the loader reports as "cannot open shared object file" and is
indistinguishable from a policy denial.

This is a compatibility aid, not a security control. Nothing in OpenShell
trusts its output: loopback and authorization decisions continue to use the
kernel's peer address from the broker's own accept4. A workload that unsets
LD_PRELOAD, links statically, or issues raw syscalls gains nothing it did
not already have.

Coverage follows from LD_PRELOAD interposing symbols rather than syscalls:
Bun and CPython are covered, Node.js does not need it because libuv passes
a null address and resolves the peer lazily through getpeername, and Go and
statically linked binaries remain unsupported for address-bearing accept.
getpeername is fixed for every runtime because the kernel answers it.

The object is built freestanding with -nostdlib so one build per
architecture loads under both glibc and musl, carries no DT_NEEDED entry
and no text relocations, and leaves __errno_location as its sole undefined
symbol.

Add an e2e test that asserts what the workload observes rather than how the
sandbox arranges it: a listener accepts a connection from a client in the
same sandbox and must report the client's address, not EOPNOTSUPP and not
its own. It passes unchanged on both kernel generations.

Signed-off-by: Russell Bryant <rbryant@redhat.com>
@copy-pr-bot

copy-pr-bot Bot commented Oct 1, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

Signed-off-by: Russell Bryant <rbryant@redhat.com>
Signed-off-by: Russell Bryant <rbryant@redhat.com>
@sallyom

sallyom commented Oct 2, 2026

Copy link
Copy Markdown

I hit this same issue! Testing this now in my env, thank you 😁

Signed-off-by: Russell Bryant <rbryant@redhat.com>
Signed-off-by: Russell Bryant <rbryant@redhat.com>
Replace workload-owned shim paths with sealed anonymous objects isolated per
launch. Keep descriptors close-on-exec in the boundary and enable inheritance
only in the intended child. Preserve effective preload environment overrides
without introducing a Landlock user ruleset.

Delegate cancellation to libc's accept4 wrapper, close descriptors on peer
lookup failure, and return ECONNABORTED for reset peers or EFAULT for invalid
output pointers. Reject unsupported raw accepts before consuming clients and
allow kernel peer queries on connected DNS sockets. Legacy outbound peer
queries intentionally retain the documented loopback relay-address fallback.

Add target-compiler fixtures and ELF, peer-error, legacy broker, and concurrent
launch regressions. Update compatibility documentation and remove PR-specific
reproducer artifacts.

Signed-off-by: Russell Bryant <rbryant@redhat.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(sandbox): support peer-address accept() for server workloads on pre-5.19 kernels (legacy read-only mode)

2 participants