Skip to content

Fix exit detection for stopped child processes - #371

Open
bitbemol wants to merge 1 commit into
swiftlang:mainfrom
bitbemol:codex/fix-stopped-child-waitid
Open

bitbemol wants to merge 1 commit into
swiftlang:mainfrom
bitbemol:codex/fix-stopped-child-waitid

Conversation

@bitbemol

@bitbemol bitbemol commented Sep 5, 2026 •

Copy link
Copy Markdown

Summary

This change makes _peekIfExited recognize only terminal child states: CLD_EXITED, CLD_KILLED, and CLD_DUMPED.

It addresses unexpected waitid behavior reproduced on macOS 27 beta and macOS 26.6.2: an exit-only query returned a stop notification for a child that was still alive and suspended.

This is proposed as defensive handling of the observed behavior, rather than a claim that the existing condition is incorrect on every platform.

Expected and observed behavior

For a child suspended with SIGSTOP, this call should not report a terminal exit:

waitid(P_PID, child, &info, WEXITED | WNOHANG | WNOWAIT);

With a freshly zero-initialized siginfo_t, the observed result was:

return value = 0
si_pid       = child PID
si_signo     = SIGCHLD
si_code      = CLD_STOPPED
si_status    = SIGSTOP

Because si_pid and si_signo are nonzero, the original _peekIfExited condition returns true, although the child is paused rather than terminated. This can bypass exit monitoring and lead to a nonterminal status reaching TerminationStatus, which traps on unexpected status codes.

The proposed condition returns false for this stop notification.

Reproduction and validation

macOS 27.0 beta — build 26A5425a

Environment:

  • Apple Silicon (arm64)
  • Xcode 27.0 beta, build 27A5252f
  • Swift 6.4 (swiftlang-6.4.0.33.1)

The PR includes two Swift regression tests:

  • testStoppedChildIsNotReportedAsExited: suspend a child and verify that exit observation returns false without consuming the pending stop notification.
  • testMonitoringStoppedChildWaitsForItsFinalExit: monitor a stopped child, resume it, and verify its final exit status is 42.

Both tests passed with the patch in the latest local run on this environment.

An independent C diagnostic reproduced the unexpected OS result in 80/80 trials: 20 trials for each combination of fork/posix_spawn and exit queries with/without WNOWAIT.

That diagnostic observed the paused state through proc_pidinfo, without any preceding waitid call. All 80 children subsequently resumed and exited with status 42. A live, non-paused child returned zero PID, signal, and code as a negative control.

Summary lines from that run:

fork reap: old false exits 20/20; stop reports 20/20; subsequent real exits 20/20
fork peek: old false exits 20/20; stop reports 20/20; subsequent real exits 20/20
spawn reap: old false exits 20/20; stop reports 20/20; subsequent real exits 20/20
spawn peek: old false exits 20/20; stop reports 20/20; subsequent real exits 20/20
running control: rc=0 pid=0 signo=0 code=0

Here, “old false exits” counts how often the original nonzero-PID/signal condition would classify a paused child as exited. “Subsequent real exits” counts the later CLD_EXITED results with status 42, after explicitly resuming the children.

macOS 26.6.2 — build 25G83

A smaller, single-case C diagnostic also reproduced the behavior using fork and WEXITED | WNOHANG | WNOWAIT.

Recorded output:

Paused child: nonzero PID=1, signal=20, code=5, status=17
REPRODUCED: exit-only query returned a pause notification.
Confirmed: after resuming, the child really exited with status 42.

nonzero PID=1 is a Boolean indicator, not the child's actual PID. The other values correspond to SIGCHLD = 20, CLD_STOPPED = 5, and SIGSTOP = 17.

This extends the reproduction evidence beyond macOS 27 beta. The 80-trial matrix, the query without WNOWAIT, and the Swift regression tests have not been run on macOS 26.6.2 as part of this validation.

Standalone reproduction

The source below is the smaller diagnostic used for the macOS 26.6.2 result, not the earlier 80-trial matrix.

It does not link against Swift or Swift Subprocess. It creates its own child, confirms the paused state through proc_pidinfo, and then makes the exit-only waitid query with freshly initialized output. It subsequently resumes the child and checks its actual exit status.

It includes cleanup for normal exits and interruption/timeout signals.

  1. Save the source as waitid-check.c.
  2. Compile and run from the directory containing the file:
xcrun clang -Wall -Wextra -Werror waitid-check.c -o waitid-check
./waitid-check
C diagnostic source
#include <errno.h>
#include <libproc.h>
#include <signal.h>
#include <stdio.h>
#include <stdlib.h>
#include <sys/proc.h>
#include <sys/wait.h>
#include <unistd.h>

/* No file/network access. Every signal targets an owned, unreaped child. */
static volatile sig_atomic_t owned_child = 0;
static sigset_t cleanup_signals;

static void cleanup(void) {
    sigset_t previous;
    sigprocmask(SIG_BLOCK, &cleanup_signals, &previous);
    pid_t pid = (pid_t)owned_child;
    if (pid > 0) {
        kill(pid, SIGKILL);
        while (waitpid(pid, NULL, 0) < 0 && errno == EINTR) {}
        owned_child = 0;
    }
    sigprocmask(SIG_SETMASK, &previous, NULL);
}

static void interrupted(int signal_number) {
    cleanup();
    const char message[] = "Interrupted or timed out; test-child cleanup completed.\n";
    (void)write(STDERR_FILENO, message, sizeof(message) - 1);
    _exit(128 + signal_number);
}

int main(void) {
    const int signals[] = {SIGALRM, SIGINT, SIGTERM, SIGHUP};
    sigemptyset(&cleanup_signals);
    for (unsigned i = 0; i < sizeof(signals) / sizeof(signals[0]); i++)
        sigaddset(&cleanup_signals, signals[i]);
    struct sigaction action = {0};
    action.sa_handler = interrupted;
    action.sa_mask = cleanup_signals;
    for (unsigned i = 0; i < sizeof(signals) / sizeof(signals[0]); i++)
        if (sigaction(signals[i], &action, NULL)) return 1;
    alarm(15);

    sigset_t previous;
    if (sigprocmask(SIG_BLOCK, &cleanup_signals, &previous)) return 1;
    pid_t pid = fork();
    if (pid < 0) { perror("fork"); return 1; }
    if (pid == 0) {
        action.sa_handler = SIG_DFL;
        for (unsigned i = 0; i < sizeof(signals) / sizeof(signals[0]); i++)
            if (sigaction(signals[i], &action, NULL)) _exit(1);
        if (sigprocmask(SIG_SETMASK, &previous, NULL)) _exit(1);
        if (raise(SIGSTOP)) _exit(1);
        _exit(42);
    }
    owned_child = pid;
    atexit(cleanup);
    sigprocmask(SIG_SETMASK, &previous, NULL);

    /* Observe the pause without making a preliminary wait/waitid call. */
    int paused = 0;
    for (int i = 0; i < 5000; i++) {
        struct proc_bsdinfo state = {0};
        if (proc_pidinfo(pid, PROC_PIDTBSDINFO, 0, &state, sizeof(state)) != sizeof(state)) break;
        if (state.pbi_status == SSTOP) { paused = 1; break; }
        usleep(1000);
    }
    if (!paused) { puts("INCONCLUSIVE: could not confirm the test child was paused."); return 2; }
    siginfo_t info = {0};
    if (waitid(P_PID, (id_t)pid, &info, WEXITED | WNOHANG | WNOWAIT)) {
        perror("waitid"); return 3;
    }
    printf("Paused child: nonzero PID=%d, signal=%d, code=%d, status=%d\n",
        info.si_pid != 0, info.si_signo, info.si_code, info.si_status);
    int reproduced = info.si_code == CLD_STOPPED && (info.si_pid != 0 || info.si_signo != 0);
    if (reproduced) puts("REPRODUCED: exit-only query returned a pause notification.");
    else if (info.si_pid == 0 && info.si_signo == 0)
        puts("NOT REPRODUCED: exit-only query did not report the paused child.");
    else puts("INCONCLUSIVE: a different result needs examination.");

    if (kill(pid, SIGCONT)) { perror("resume"); return 4; }
    for (int i = 0; i < 5000; i++) {
        siginfo_t final = {0};
        if (waitid(P_PID, (id_t)pid, &final, WEXITED | WNOHANG | WNOWAIT)) {
            perror("final waitid"); return 5;
        }
        if (final.si_code == CLD_EXITED && final.si_status == 42) {
            puts("Confirmed: after resuming, the child really exited with status 42.");
            cleanup();
            alarm(0);
            return 0;
        }
        usleep(1000);
    }
    puts("INCONCLUSIVE: final exit was not confirmed before the deadline.");
    return 6;
}

Interpretation and scope

The diagnostics demonstrate that an exit-only query can return CLD_STOPPED on the two tested macOS versions, and that the original library condition misclassifies that result as terminal.

Apple's published XNU implementation gates stop notifications on WSTOPPED. The observed behavior therefore warrants investigation at the OS boundary as well as consideration of defensive handling in the library.

The C diagnostics isolate the OS behavior; they do not themselves execute the Swift library's subsequent crash path.

The investigation began with application subprocess hangs. After adopting the patched library, I observed the application working normally again. However, I haven't established that this specific condition caused the original hangs, so I’m treating that improvement as supporting context rather than proof of the root cause.

Remaining review

  • Whether the behavior occurs on other macOS versions or non-macOS platforms.
  • Why the tested OS versions return a stop notification for an exit-only query.
  • Whether defensive handling in Swift Subprocess is the preferred solution.
  • Extending the handling and regression coverage to _reap, if that approach is agreed upon.

The adjacent _reap helper still uses the original nonempty-result check. The macOS 27 diagnostic reproduced the same stop notification without WNOWAIT, so that path also needs consideration and focused regression coverage. The current patch does not address it.

Feedback is welcome on whether this approach is appropriate or whether a different solution would be preferable.

@bitbemol
bitbemol requested a review from iCharlesHu as a code owner September 5, 2026 01:41
internal func _reap(pid: pid_t) throws(Errno) -> TerminationStatus? {
let siginfo = try _waitid(idtype: P_PID, id: id_t(pid), flags: WEXITED | WNOHANG)
// If si_pid and si_signo are both 0, the child is still running since we used WNOHANG
if siginfo.si_pid == 0 && siginfo.si_signo == 0 {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would this need adjustment as well?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we should consider the same handling here.

I also observed the stop notification without WNOWAIT, as described in my main reply. If this approach looks appropriate to you, I'd be happy to extend the fix and add a focused regression test.

@jakepetroules

Copy link
Copy Markdown
Contributor

Would you be able to provide some additional context on the change? e.g. what are some cases where si_pid and si_signo both being zero are not equivalent to the proposed condition?

@bitbemol

bitbemol commented Sep 5, 2026 •

Copy link
Copy Markdown
Author

Would you be able to provide some additional context on the change? e.g. what are some cases where si_pid and si_signo both being zero are not equivalent to the proposed condition?

Thanks for reviewing this, @jakepetroules, and sorry for leaving the description empty! I've updated it with the reproduction details and test results. 😅

On macOS 27 beta (26A5425a), I observed waitid with WEXITED | WNOHANG | WNOWAIT returning a stopped child's PID, SIGCHLD, and CLD_STOPPED. (Also reproduced with the C code I attached in 26.6.2)

That is where the checks differ: the original condition reports that the child has exited because the PID and signal are nonzero, while the proposed condition correctly rejects the nonterminal status. Reporting an exit prematurely can bypass monitoring and lead to the stop notification reaching TerminationStatus, which traps on that status.

I confirmed the unexpected waitid result in a standalone C diagnostic, independently of Swift Subprocess. I haven't verified it on stable macOS or other platforms, so this may be an OS-specific issue that the library could handle defensively.

Regarding your _reap comment: I also observed the stop notification without WNOWAIT. If this approach looks appropriate, I'd be happy to extend the fix there and add a focused regression test.

Thanks for your guidance!

@jakepetroules

Copy link
Copy Markdown
Contributor

I'd also like @FranzBusch and/or @weissi to take a look if possible.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants