Skip to content

ci(podman): report e2e-podman archive selected-vs-eligible coverage - #4162

Open
politerealism wants to merge 3 commits into
NVIDIA:mainfrom
politerealism:3712-podman-e2e-coverage-summary
Open

politerealism wants to merge 3 commits into
NVIDIA:mainfrom
politerealism:3712-podman-e2e-coverage-summary

Conversation

@politerealism

Copy link
Copy Markdown
Contributor

Summary

Adds a job-summary report on the e2e-podman nextest archive's coverage: how many targets are eligible under the e2e-podman feature, how many are excluded, how many are selected into the archive, and how many actually execute — so a narrow Podman CI run cannot be mistaken for full coverage.

Related Issue

Addresses the "selected-vs-eligible reporting" acceptance criterion on #3712 (Job names and summaries report the selected test set or target count so a narrow Podman job cannot be mistaken for the complete suite.).

Why this is scoped so narrowly

This issue has a lot of adjacent, actively-moving work (the e2e-podman test migration, the RPM/DEB installation-profile qualification effort, RFC 0016's broader testing-strategy discussion). Given that, this PR deliberately does only one small, self-contained thing:

  • No new selection mechanism. It reuses the exact cargo nextest list invocation generate-podman-e2e-ci-tests already uses, and reads podmanE2eFollowUpBinaries — the same declarative exclusion list that already drives the archive filter — rather than introducing a second, parallel way to know what's "in scope." If that list changes, this report changes with it automatically; there's nothing to keep in sync by hand.
  • No new CI gating. This only prints a summary; it doesn't fail a build, block a merge, or add a new required check. A wrong or stale number here is a documentation problem, not a correctness problem — safe to land independently of the bigger migration work in flight, and safe to delete or supersede later without unwinding any enforcement behavior.
  • Output only, not a new data model. It's a handful of integers and names rendered as markdown into $GITHUB_STEP_SUMMARY. RFC 0016 (docs(rfc): propose OpenShell testing strategy #3460 / test: define and track OpenShell testing strategy #3954) is still working out "the meaning of a complete conformance result" as project-wide testing strategy; this is intentionally small enough to be cheap to redo or retire once that lands, rather than pre-empting it with a bespoke coverage model of its own.

Why precision mattered here specifically

The first version of this (see commit history) reported "selected" as eligible - excluded. That's wrong in one specific way: two of the eligible, non-excluded targets (internet_network_perf, live_internet_traffic_perf) are real #[ignore]'d manual-only benchmarks — included in the archive's scope, but never actually executed by nextest. Reporting them as "selected" without qualification would have reproduced the exact failure mode this change exists to prevent: a number that looks like full coverage but isn't, just one level removed. The final version reports eligible, excluded, selected, and actually-executing as four distinct numbers and names the manual-only benchmarks explicitly, so the report is honest about the one case where "in scope" and "runs automatically" diverge.

Changes

  • tests/artifacts.nix: new podmanE2eCoverageSummary Nix app. Computes the eligible set via cargo nextest list (full listing, not binaries-only, so per-test #[ignore] status is visible), compares it against podmanE2eFollowUpBinaries, and reports eligible/excluded/selected/executing counts, naming any selected-but-ignored benchmarks.
  • flake.nix: exposes it as nix run .#podman-e2e-coverage-summary.
  • .github/workflows/integration-runner.yml: new step in the shared integration job, scoped to matrix.testsuite == 'e2e-podman' only, if: always() so it reports regardless of pass/fail, writing to $GITHUB_STEP_SUMMARY.
  • TESTING.md / CI.md: documented in both, matching this repo's existing split (CI.md = current implemented behavior, TESTING.md = local entry points).

Testing

  • nix-instantiate --parse on both modified .nix files
  • YAML validity check on the modified workflow file
  • Manually re-ran the exact jq/nextest pipeline against the real e2e/rust crate with a real cargo-nextest install, confirming correct output end to end: 36 eligible, 20 excluded, 16 selected (2 manual-only: internet_network_perf, live_internet_traffic_perf), 14 actually execute
  • Have not run nix run .#podman-e2e-coverage-summary inside the real flake/tmachine/KVM CI environment — no working Nix daemon available in the environment this was developed in. This will get its first real exercise when this PR's CI runs.

Checklist

  • Follows Conventional Commits
  • Commits are signed off (DCO)
  • Architecture/CI documentation updated (TESTING.md, CI.md)

🤖 Generated with Claude Code

The e2e-podman nextest archive is a curated subset of everything
eligible under the e2e-podman feature: podmanE2eFollowUpBinaries in
tests/artifacts.nix excludes targets that still depend on wrapper-owned
gateway controls, host fixtures, or other unmigrated test
infrastructure. CI reported a per-binary pass/fail count but never the
selected-vs-eligible split, so a narrow run could be mistaken for full
Podman coverage.

Add podman-e2e-coverage-summary, a Nix app that computes the eligible
count via a real cargo nextest list (the same mechanism
generate-podman-e2e-ci-tests already uses) and compares it against
podmanE2eFollowUpBinaries -- the same declarative exclusion list that
drives the archive filter, so this can't drift from what actually
runs. Wire it into integration-runner.yml as a job-summary step for
every e2e-podman run, independent of pass/fail.

Addresses the selected-vs-eligible reporting acceptance criterion in
NVIDIA#3712.

Signed-off-by: politerealism <burdcat17@gmail.com>
Address review feedback on the previous commit:

- The coverage summary's "selected" count included targets whose only
  test is a manual-only #[ignore]'d benchmark, so it overstated what
  actually executes by the count of those benchmarks. Switch from
  --list-type binaries-only to the full nextest listing so per-test
  ignored status is visible, and report eligible, excluded, selected,
  and actually-executing as distinct numbers, naming any selected
  manual-only benchmarks explicitly instead of folding them into
  "selected" silently.
- Document the new job-summary step in CI.md (current implemented CI
  behavior) alongside the existing TESTING.md entry (local entry
  point), matching this repo's documented split between the two files.

Signed-off-by: politerealism <burdcat17@gmail.com>
@copy-pr-bot

copy-pr-bot Bot commented Oct 3, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

Address the rootless/rootful-scope acceptance criterion in NVIDIA#3712:
document why the e2e-podman archive runs rootless-only in branch CI
(unlike driver-podman, which runs both), and file a tracking issue for
the one confirmed gap.

- podman_host_gateway was checked directly rather than assumed: the
  outer container's host-gateway alias uses Podman's own native
  "host-gateway" value (resolved identically in both modes), and the
  supervisor-side resolution defaults to a hardcoded 127.0.0.1 on
  Linux regardless of rootful/rootless. No real coverage gap there.
- podman_resource_limits is the one confirmed exception: it reads real
  cgroup v2 state, and rootless cgroup delegation is the harder,
  already-covered case. Rootful is unverified. Tracked in NVIDIA#4163, which
  also notes the e2e-podman tmachine playbook hardcodes rootless-only
  values and would need the same tmachine_container_runtime role
  pattern NVIDIA#3663 already proved out for the driver-podman suite.
- CI.md and tests/artifacts.nix both say plainly that closing NVIDIA#4163
  means running the whole archive rootful, not just the one target
  that needs it -- nextest archive filters select whole binaries, so
  there's no cheaper way to isolate just podman_resource_limits.

Signed-off-by: politerealism <burdcat17@gmail.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant