Skip to content

pm test --run reports a passing name-filtered Bun command as failed empty_run (Bun's stderr summary is not accepted as an execution receipt) #1418

Description

@unbraind

Problem

pm test --run marks a passing, name-filtered Bun test command as failed with failure_category: "empty_run" and rewrites its exit code to 1, although Bun ran tests, all passed, and it exited 0. The same file without the filter is classified passed.

Reproduced on @unbrained/pm-cli 2026.10.6, Bun 1.3.5 (also reported on 2026.10.4 / Bun 1.3.14 in a package CI):

mkdir repro && cd repro && git init -q && pm init --yes
mkdir test && cat > test/x.test.ts <<'TS'
import { test, expect } from "bun:test";
test("alpha one", () => { expect(1).toBe(1); });
test("alpha two", () => { expect(2).toBe(2); });
test("beta", () => { expect(3).toBe(3); });
TS
id=$(pm create --title "Repro" --type Task --json | jq -r .item.id)
pm test "$id" --add command="bun test --test-name-pattern alpha test/x.test.ts"
pm test "$id" --add command="bun test test/x.test.ts"
pm test "$id" --run --json | jq -c '.run_results[]|{command,status,exit_code,failure_category}'

Result:

{"command":"bun test --test-name-pattern alpha test/x.test.ts","status":"failed","exit_code":1,"failure_category":"empty_run"}
{"command":"bun test test/x.test.ts","status":"passed","exit_code":0,"failure_category":null}

Bun's own output for the filtered command (stderr):

 2 pass
 1 filtered out
 0 fail
 2 expect() calls
Ran 2 tests across 1 file. [51.00ms]

The runner recognizes Bun's name filter (that is why it looks for a positive execution receipt at all), but its receipt patterns apparently don't accept Bun's summary ( N pass / Ran N tests across M files). When a filter is present, the missing receipt is read as "the filter matched nothing".

Why it matters

Agents link narrow, fast regression commands to items (pm test <id> --add command="bun test --test-name-pattern …"), which is exactly the TDD flow pm recommends. With this bug every filtered Bun receipt is a false failure. A --validate-close gate then blocks closing a correctly tested item. Agents learn to drop the filter and run whole suites, which costs time and tokens, or to stop linking Bun tests.

Expected

  • Treat Bun's summary as a positive execution receipt: ^\s*(\d+) pass with N > 0, or Ran (\d+) tests? across with N > 0, on stdout or stderr (Bun prints its summary to stderr).
  • empty_run only when the receipt shows 0 tests ran (e.g. Ran 0 tests, or Bun's "did not match any test names" message).
  • Never rewrite a zero exit code to 1 without a stated, matched reason. Include the matched (or missing) receipt pattern in the run result so the classification is explainable.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions