Conversation
…ed sockets getpeername(2) was always answered by writing the sockaddr into the workload's address space from the broker. On kernels without SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV the broker runs in LegacyReadOnly mode, where that cross-process write fails closed with EOPNOTSUPP, so getpeername failed for every socket. For SocketState::Local and SocketState::AcceptedLocal the descriptor is genuinely connected to the recorded peer, so the kernel's own answer is identical to the broker's. Respond with SECCOMP_USER_NOTIF_FLAG_CONTINUE for those states and let the kernel write the sockaddr in the workload's own address space. That needs no cross-process task-memory write and therefore works on kernels before 5.19. SocketState::Connected keeps the existing substitution. Those descriptors are connected to a loopback relay rather than to the destination the workload requested, so the kernel's answer would leak the relay address. A regression test pins that distinction. This is the path libuv-based runtimes depend on: libuv passes a null peer address to accept4(2) and resolves the peer lazily through getpeername(2) when the application reads remoteAddress. Refs NVIDIA#4058 Signed-off-by: Russell Bryant <rbryant@redhat.com>
…E_RECV Kernels before Linux 5.19 lack SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV, so the broker cannot hold a notified workload thread in a kill-only wait and cannot safely write into workload memory. It fails closed with EOPNOTSUPP on every mediated syscall whose result is an output buffer. accept(fd, &addr, &len) has exactly that shape, so server workloads on RHEL 9.x and RHCOS see EOPNOTSUPP where they expect a connection. Add openshell-accept-shim, a freestanding preloadable library that rewrites an address-bearing accept into accept4(fd, NULL, NULL, flags) followed by getpeername on the accepted descriptor. The first call has no output buffer, so the broker injects the descriptor without touching workload memory; the second is answered with SECCOMP_USER_NOTIF_FLAG_CONTINUE, so the kernel writes the address into the caller's own buffer. Because the workload writes its own memory, the cross-process TOCTOU that WAIT_KILLABLE_RECV exists to close does not apply. The sandbox installs the object at /run/openshell-compat only when the listener reports writes disabled, and composes LD_PRELOAD so a workload's own value is preserved. Installation failure is non-fatal and logged as an OCSF config state change; the sandbox keeps the existing fail-closed behavior. Landlock gates open independently of the file mode, so a world-readable object is not by itself reachable: without an explicit admission the loader reports "cannot open shared object file" and silently drops the shim. Workload children are launched from more than one place -- the sandbox entrypoint and sandbox exec through the boundary -- and each prepares Landlock with its own runtime paths. Admit the object and its directory inside prepare_child_sandbox, the single production entry point, so a launch path cannot set LD_PRELOAD without also letting the loader open what it points at. Resolve the shim's C compiler through the cc crate so the standard CC_<target>, TARGET_CC, and CC overrides apply, along with the per-target wrappers cargo-zigbuild installs when release binaries are cross-compiled from a non-Linux host. Verify the emitted object's ELF machine against CARGO_CFG_TARGET_ARCH and fail the build on a mismatch: a misresolved compiler otherwise produces a valid object for the wrong architecture, which the loader reports as "cannot open shared object file" and is indistinguishable from a policy denial. This is a compatibility aid, not a security control. Nothing in OpenShell trusts its output: loopback and authorization decisions continue to use the kernel's peer address from the broker's own accept4. A workload that unsets LD_PRELOAD, links statically, or issues raw syscalls gains nothing it did not already have. Coverage follows from LD_PRELOAD interposing symbols rather than syscalls: Bun and CPython are covered, Node.js does not need it because libuv passes a null address and resolves the peer lazily through getpeername, and Go and statically linked binaries remain unsupported for address-bearing accept. getpeername is fixed for every runtime because the kernel answers it. The object is built freestanding with -nostdlib so one build per architecture loads under both glibc and musl, carries no DT_NEEDED entry and no text relocations, and leaves __errno_location as its sole undefined symbol. Add an e2e test that asserts what the workload observes rather than how the sandbox arranges it: a listener accepts a connection from a client in the same sandbox and must report the client's address, not EOPNOTSUPP and not its own. It passes unchanged on both kernel generations. Signed-off-by: Russell Bryant <rbryant@redhat.com>
russellb
requested review from
a team,
derekwaynecarr,
mrunalp and
sjenning
as code owners
October 1, 2026 22:53
Signed-off-by: Russell Bryant <rbryant@redhat.com>
Signed-off-by: Russell Bryant <rbryant@redhat.com>
|
I hit this same issue! Testing this now in my env, thank you 😁 |
Signed-off-by: Russell Bryant <rbryant@redhat.com>
Signed-off-by: Russell Bryant <rbryant@redhat.com>
Replace workload-owned shim paths with sealed anonymous objects isolated per launch. Keep descriptors close-on-exec in the boundary and enable inheritance only in the intended child. Preserve effective preload environment overrides without introducing a Landlock user ruleset. Delegate cancellation to libc's accept4 wrapper, close descriptors on peer lookup failure, and return ECONNABORTED for reset peers or EFAULT for invalid output pointers. Reject unsupported raw accepts before consuming clients and allow kernel peer queries on connected DNS sockets. Legacy outbound peer queries intentionally retain the documented loopback relay-address fallback. Add target-compiler fixtures and ELF, peer-error, legacy broker, and concurrent launch regressions. Update compatibility documentation and remove PR-specific reproducer artifacts. Signed-off-by: Russell Bryant <rbryant@redhat.com>
2 of 3 tasks
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Restore libc-based server workloads on kernels without
SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV(introduced in Linux 5.19). In legacy read-only mode, the broker cannot safely write output buffers into workload memory. The compatibility shim rewrites address-bearingaccept/accept4into a null-address accept followed by kernel-handledgetpeername.Closes #4058.
Implementation
accept4wrapper to retain libc's cancellation handling, then uses raw syscalls for peer lookup and error-path close. Failed peer lookup closes the accepted descriptor;ENOTCONNbecomesECONNABORTED, and invalid pointers returnEFAULTwithout userspace dereferences./run. Descriptors remain close-on-exec in the boundary; only the intended forked child makes its descriptor inheritable. A workload cannot change another launch's shim contents or permissions. No shim-specific Landlock admission or user ruleset is added.LD_PRELOADafter the launch environment is assembled, preserving the effective provider, workload, or session value. Skip direct foreign-class/architecture ELF executables, and avoid blocking executable probes on FIFOs or devices. Installation failure is non-fatal and logged; the broker retains fail-closed behavior.getpeernamein the kernel for local, accepted-local, and connected DNS sockets. Legacy relayed outbound sockets report the kernel's loopback relay address and port, not the original destination. Modern listeners still substitute the original outbound destination. Unconnected registered sockets returnENOTCONN.cc, including Cargo's CC/CFLAGS rebuild tracking. Compile C test fixtures with that target compiler. Add automated ELF, peer-error, cancellation, descriptor-isolation, Landlock, and full forced-legacy broker/preload regressions.Coverage and limits
The shim interposes dynamic libc symbols. CPython and Bun address-bearing accepts are covered; Node's null-address accept does not need the shim. Go, static binaries, secure-execution loaders, and direct syscalls do not gain address-bearing accept support from it. Kernel-handled peer queries do not depend on the preload.
Descendants must retain the shim descriptor and access to
/proc/self/fd. Descendants that close it, hide/proc, or invoke a foreign-architecture loader must remove itsLD_PRELOADentry themselves; arbitrary descendant execs cannot be checked by the boundary and may otherwise emit loader warnings. Connection authorization continues to use the requested destination and broker-observed addresses, not shim output.Validation
/tmpmountednoexec.5.14.0-570.141.1.el9_6.x86_64, randomized non-root uid, all capabilities dropped,RuntimeDefaultseccomp: legacy broker/preloaded CPython, queued-client preservation after raw accept rejection, concurrent-session descriptor isolation, Landlock, environment, and all 20 shim tests passed.x86_64-unknown-linux-muslpassed. Inspection of the actual release object confirmed noDT_NEEDED, no text relocations, a non-executable stack, and exactlyaccept/accept4exports. Dynamic Alpine musl cancellation smoke matrix passed.The OpenShift tests ran the real broker and preload directly in a test pod. A separate full gateway deployment could not start because the cluster's storage provisioner was unavailable; this validation does not claim a completed gateway E2E run.