Problem
tests/conftest.py:11-12 states the governing principle:
GPU tests never run by default: a silently-passing GPU test on a CPU-only box would be a false green.
The marker/DEVMM_GPU gating enforces that correctly. But inside the hardware suites, consumer libraries are resolved with pytest.importorskip, so once the operator has explicitly set DEVMM_GPU=cuda|rocm, a missing torch, cupy, rmm, numba, or numpy skips rather than fails.
The result is a weaker version of the same false green the conftest is designed to prevent: the operator asked for the hardware suite, the run reports success, and the fact that a chunk of it never executed is visible only in the skip count.
Why this matters more now
PR #1 removed torch from the gpu-test-cuda extra (correctly — the CUDA build has to come from PyTorch's own index, per docs/testing.md). That makes step 2 of the documented install recipe load-bearing. An operator who runs step 1 and skips step 2 gets a green make test-gpu-cuda with the torch DLPack round-trips silently absent.
ADR 0003's manual-run record pins the expected shape of a complete run at 41 passed, 1 skipped. Nothing currently checks a run against that number, so a partial install degrades quietly instead of failing loudly.
This is pre-existing on main — not introduced by #1 — but #1 raises the odds of hitting it.
Scope
22 importorskip sites across three files:
tests/test_cuda_gpu.py — lines 47, 149, 150, 165, 166, 181, 182, 218, 242, 249, 280
tests/test_rocm_gpu.py — lines 46, 150, 151, 166, 167, 182, 183
tests/test_integrations_gpu.py — lines 45, 58, 71, 84, 99, 100, 122
The ROCm half matters for the still-open T3 side of the ADR 0003 waiver: whenever AMD hardware appears, the first T3 run should not be able to report green off a partial install.
Suggested direction
Once DEVMM_GPU is set, a missing consumer library should error rather than skip. Options, roughly in increasing strictness:
- A conftest helper (
require_module(name)) that calls pytest.fail when the relevant DEVMM_GPU value is active and importorskip otherwise — a mechanical swap at the 22 sites.
- A session-scoped fixture that asserts the full consumer set is importable up front, so the failure is one clear message rather than N.
- Additionally assert the collected/passed count against the ADR-recorded baseline, so silent shrinkage of the suite is itself a failure.
Genuine judgement call worth settling first: torch.cuda.is_available() being false (tests/test_cuda_gpu.py:183) is arguably a legitimate skip — a CPU-only torch build is a different condition from torch being absent. Whatever lands should probably keep that one a skip while making absence an error.
Follow-up from #1. 🤖 Generated with Claude Code
Problem
tests/conftest.py:11-12states the governing principle:The marker/
DEVMM_GPUgating enforces that correctly. But inside the hardware suites, consumer libraries are resolved withpytest.importorskip, so once the operator has explicitly setDEVMM_GPU=cuda|rocm, a missingtorch,cupy,rmm,numba, ornumpyskips rather than fails.The result is a weaker version of the same false green the conftest is designed to prevent: the operator asked for the hardware suite, the run reports success, and the fact that a chunk of it never executed is visible only in the skip count.
Why this matters more now
PR #1 removed
torchfrom thegpu-test-cudaextra (correctly — the CUDA build has to come from PyTorch's own index, perdocs/testing.md). That makes step 2 of the documented install recipe load-bearing. An operator who runs step 1 and skips step 2 gets a greenmake test-gpu-cudawith the torch DLPack round-trips silently absent.ADR 0003's manual-run record pins the expected shape of a complete run at 41 passed, 1 skipped. Nothing currently checks a run against that number, so a partial install degrades quietly instead of failing loudly.
This is pre-existing on
main— not introduced by #1 — but #1 raises the odds of hitting it.Scope
22
importorskipsites across three files:tests/test_cuda_gpu.py— lines 47, 149, 150, 165, 166, 181, 182, 218, 242, 249, 280tests/test_rocm_gpu.py— lines 46, 150, 151, 166, 167, 182, 183tests/test_integrations_gpu.py— lines 45, 58, 71, 84, 99, 100, 122The ROCm half matters for the still-open T3 side of the ADR 0003 waiver: whenever AMD hardware appears, the first T3 run should not be able to report green off a partial install.
Suggested direction
Once
DEVMM_GPUis set, a missing consumer library should error rather than skip. Options, roughly in increasing strictness:require_module(name)) that callspytest.failwhen the relevantDEVMM_GPUvalue is active andimportorskipotherwise — a mechanical swap at the 22 sites.Genuine judgement call worth settling first:
torch.cuda.is_available()being false (tests/test_cuda_gpu.py:183) is arguably a legitimate skip — a CPU-only torch build is a different condition from torch being absent. Whatever lands should probably keep that one a skip while making absence an error.Follow-up from #1. 🤖 Generated with Claude Code