User Story
As an operator running OpenShell on OpenShift / RHEL 9.x nodes,
I want server-shaped workloads to accept local connections inside the sandbox,
so that agent harnesses that run a local HTTP control server work on my existing fleet without waiting for a kernel bump.
Problem Statement
On kernels before Linux 5.19 the sandbox runs in legacy read-only mode, where the broker refuses the mediated operations that write results back into workload memory. For accept/accept4 this means the call succeeds with a null peer-address argument and fails closed with EOPNOTSUPP with a non-null one.
This is documented and deliberate (support matrix). Without SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV a non-fatal signal can cancel the notification and restart the syscall between SECCOMP_IOCTL_NOTIF_ID_VALID and the broker's cross-process write, so the workload may free or reuse the buffer before the write lands. There is no atomic validate-and-write ioctl for memory the way SECCOMP_IOCTL_NOTIF_ADDFD provides one for descriptors.
The gap is that OpenShell offers no alternative path for these workloads. #3417 made the sandbox start on 5.14; this is the next layer of consequence, which that issue's acceptance criteria did not cover.
Impact / Why This Matters
Today, operators on RHCOS/RHEL 9.x must either run the workload outside the sandbox or move to a node kernel with WAIT_KILLABLE_RECV. Neither is available in-cluster: the kernel is fixed for the RHEL 9 lifecycle, and running outside the sandbox defeats the purpose.
Observed on a 5.14.0-570.103.1.el9_6 node: OpenCode fails its readiness timeout because its local server cannot accept connections. A direct in-sandbox test confirmed accept() with a peer-address argument returns EOPNOTSUPP. No network-policy denial is reported, because this is a mediation refusal rather than a policy decision.
The affected set is wider than one tool. The trigger is the accept wrapper of whatever runtime the workload uses, not the workload's own code. CPython's socket.accept() always passes a peer-address buffer; so does uSockets, which is what Bun-based harnesses go through. Notably, these workloads do not need the value — a loopback-only control server never reads the peer address.
RHCOS is the default OpenShift substrate, so this recurs for every server-shaped workload anyone tries to sandbox there.
Proposed Design
In legacy read-only mode, make accept/accept4 with a non-null peer-address argument succeed, with the address supplied by the workload process itself rather than written across process boundaries by the broker.
Externally observable behavior:
- On a pre-5.19 node, a workload calling
accept(fd, &addr, &len) receives a working connected descriptor and a populated sockaddr, instead of EOPNOTSUPP.
- The returned address reports the correct family and loopback IP. The peer port may be reported as
0 if the real value is not recoverable; this should be documented rather than silently implied to be accurate.
addrlen follows accept(2) semantics, including reporting the true size when the caller's buffer is too small.
- Behavior on 5.19+ nodes is unchanged.
- The isolation boundary is unchanged. The broker still mediates the accept, still enforces the existing loopback-only peer check, and still registers the accepted descriptor so subsequent operations on that connection remain mediated.
- Coverage limits are stated in the support matrix, including which workload classes are not covered.
The security property this must preserve: the write happens inside the workload's own address space with the workload's own privileges, so it cannot do anything the workload could not already do. The TOCTOU that forces today's fail-closed is specific to the broker's cross-process write; performing the write in-process removes it by construction rather than by narrowing the window.
One approach that satisfies this is a small preload library shipped in the sandbox image and enabled only when seccomp_listener_mode == legacy_read_only, which issues the null-address accept the broker already permits and fills in the caller's buffer locally. Its known limit is that it only covers callers reaching accept4 through a dynamic libc symbol — a statically linked Go binary issuing the syscall directly would bypass it. Implementers should weigh this against alternatives.
Explicitly out of scope: answering the notification with SECCOMP_USER_NOTIF_FLAG_CONTINUE. That would let the kernel perform the write safely, but the accepted descriptor would never pass through the broker's registry, losing the loopback check and all subsequent mediation of that connection.
Alternatives Considered
- Require a 5.19+ node kernel. Leaves current OpenShift and RHEL 9 users blocked for the RHEL 9 lifecycle. This is the status quo.
- Have the broker write anyway, relying on
NOTIF_ID_VALID alone. Reintroduces the race that legacy read-only mode exists to avoid. Not acceptable.
SECCOMP_USER_NOTIF_FLAG_CONTINUE for accept. Safe write, but surrenders mediation of the accepted socket. See above.
- Document the limitation only, and let users patch their own images. Already effectively the situation; it pushes a subtle, security-adjacent shim onto every affected user.
- Recover the true peer port for fidelity. Would need a broker side channel; disproportionate, since affected workloads do not read the port.
Acceptance Criteria
- On a pre-5.19 node, a Python workload calling
socket.accept() inside the sandbox receives a connected socket and a peer address rather than EOPNOTSUPP.
- On the same node, a Bun- or uSockets-based local HTTP server reaches readiness inside the sandbox.
- The accepted connection remains fully mediated: it is present in the socket registry, and the existing non-loopback peer rejection still applies.
addrlen truncation semantics match accept(2).
- Behavior and qualification output on 5.19+ nodes are unchanged.
- The support matrix and the OpenShift page state the new behavior, the fidelity caveat on the reported peer port, and which workload classes remain uncovered.
- Coverage limits are verified by test, not only asserted.
Related
Note on observability
While diagnosing this, the legacy-mode refusal was found to emit no log or OCSF event — the workload sees EOPNOTSUPP and the operator sees nothing, which is why the failure initially looked like a policy denial. That is arguably a separate defect worth its own issue, independent of whether this feature is accepted.
User Story
As an operator running OpenShell on OpenShift / RHEL 9.x nodes,
I want server-shaped workloads to accept local connections inside the sandbox,
so that agent harnesses that run a local HTTP control server work on my existing fleet without waiting for a kernel bump.
Problem Statement
On kernels before Linux 5.19 the sandbox runs in legacy read-only mode, where the broker refuses the mediated operations that write results back into workload memory. For
accept/accept4this means the call succeeds with a null peer-address argument and fails closed withEOPNOTSUPPwith a non-null one.This is documented and deliberate (support matrix). Without
SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECVa non-fatal signal can cancel the notification and restart the syscall betweenSECCOMP_IOCTL_NOTIF_ID_VALIDand the broker's cross-process write, so the workload may free or reuse the buffer before the write lands. There is no atomic validate-and-write ioctl for memory the waySECCOMP_IOCTL_NOTIF_ADDFDprovides one for descriptors.The gap is that OpenShell offers no alternative path for these workloads. #3417 made the sandbox start on 5.14; this is the next layer of consequence, which that issue's acceptance criteria did not cover.
Impact / Why This Matters
Today, operators on RHCOS/RHEL 9.x must either run the workload outside the sandbox or move to a node kernel with
WAIT_KILLABLE_RECV. Neither is available in-cluster: the kernel is fixed for the RHEL 9 lifecycle, and running outside the sandbox defeats the purpose.Observed on a 5.14.0-570.103.1.el9_6 node: OpenCode fails its readiness timeout because its local server cannot accept connections. A direct in-sandbox test confirmed
accept()with a peer-address argument returnsEOPNOTSUPP. No network-policy denial is reported, because this is a mediation refusal rather than a policy decision.The affected set is wider than one tool. The trigger is the accept wrapper of whatever runtime the workload uses, not the workload's own code. CPython's
socket.accept()always passes a peer-address buffer; so does uSockets, which is what Bun-based harnesses go through. Notably, these workloads do not need the value — a loopback-only control server never reads the peer address.RHCOS is the default OpenShift substrate, so this recurs for every server-shaped workload anyone tries to sandbox there.
Proposed Design
In legacy read-only mode, make
accept/accept4with a non-null peer-address argument succeed, with the address supplied by the workload process itself rather than written across process boundaries by the broker.Externally observable behavior:
accept(fd, &addr, &len)receives a working connected descriptor and a populatedsockaddr, instead ofEOPNOTSUPP.0if the real value is not recoverable; this should be documented rather than silently implied to be accurate.addrlenfollowsaccept(2)semantics, including reporting the true size when the caller's buffer is too small.The security property this must preserve: the write happens inside the workload's own address space with the workload's own privileges, so it cannot do anything the workload could not already do. The TOCTOU that forces today's fail-closed is specific to the broker's cross-process write; performing the write in-process removes it by construction rather than by narrowing the window.
One approach that satisfies this is a small preload library shipped in the sandbox image and enabled only when
seccomp_listener_mode == legacy_read_only, which issues the null-address accept the broker already permits and fills in the caller's buffer locally. Its known limit is that it only covers callers reachingaccept4through a dynamic libc symbol — a statically linked Go binary issuing the syscall directly would bypass it. Implementers should weigh this against alternatives.Explicitly out of scope: answering the notification with
SECCOMP_USER_NOTIF_FLAG_CONTINUE. That would let the kernel perform the write safely, but the accepted descriptor would never pass through the broker's registry, losing the loopback check and all subsequent mediation of that connection.Alternatives Considered
NOTIF_ID_VALIDalone. Reintroduces the race that legacy read-only mode exists to avoid. Not acceptable.SECCOMP_USER_NOTIF_FLAG_CONTINUEfor accept. Safe write, but surrenders mediation of the accepted socket. See above.Acceptance Criteria
socket.accept()inside the sandbox receives a connected socket and a peer address rather thanEOPNOTSUPP.addrlentruncation semantics matchaccept(2).Related
WAIT_KILLABLE_RECV). That issue made the sandbox start in legacy read-only mode; this one addresses the server-workload consequence left behind.Note on observability
While diagnosing this, the legacy-mode refusal was found to emit no log or OCSF event — the workload sees
EOPNOTSUPPand the operator sees nothing, which is why the failure initially looked like a policy denial. That is arguably a separate defect worth its own issue, independent of whether this feature is accepted.