Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 12 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,18 @@ and this project adheres to

## [Unreleased]

### Added

- `gpu-test-cuda` extra and `make test-gpu-cuda` to run the CUDA (T2) hardware
suite, with the setup recipe in [`docs/testing.md`](docs/testing.md#running-the-cuda-gpu-suite-on-hardware).
First hardware run recorded in [ADR 0003](docs/adr/0003-gpu-suite-waiver-for-0.1.0.md).

### Fixed

- `integrations.numba`: `DevmmEMMPlugin` passes the device pointer as a
`ctypes.c_void_p`, the only type numba-cuda ≥ 0.30 converts to a driver
pointer.

## [0.1.0] - 2026-07-17

- Repository layout: the package lives at `src/devmm/` (src-layout); imports
Expand Down
16 changes: 13 additions & 3 deletions Makefile
Original file line number Diff line number Diff line change
@@ -1,9 +1,9 @@
.PHONY: help test test-all test-devmode lint typecheck fmt fmt-check coverage verify gate-all \
release-gate dev clean
.PHONY: help test test-all test-devmode test-gpu-cuda lint typecheck fmt fmt-check coverage \
verify gate-all release-gate dev clean

help: ## Show this help
@awk 'BEGIN {FS = ":.*##"; printf "Usage: make \033[36m<target>\033[0m\n\nTargets:\n"} \
/^[a-zA-Z0-9_-]+:.*?##/ { printf " \033[36m%-12s\033[0m %s\n", $$1, $$2 }' $(MAKEFILE_LIST)
/^[a-zA-Z0-9_-]+:.*?##/ { printf " \033[36m%-14s\033[0m %s\n", $$1, $$2 }' $(MAKEFILE_LIST)

# Tests run with the `test` extra (numpy, array-api-strict): the suite's
# differential oracles and DLPack round-trips need a consumer library (§9).
Expand All @@ -18,6 +18,16 @@ test-all: test ## Run the full suite (override to add integration/e2e)
test-devmode: ## Run the suite under PYTHONDEVMODE=1 with faulthandler
PYTHONDEVMODE=1 PYTHONFAULTHANDLER=1 uv run --extra test pytest -q

# The CUDA (T2) hardware suite (docs/adr/0003). Install the deps first —
# `uv pip install '.[test,gpu-test-cuda]'` plus the PyTorch CUDA wheel — then
# this runs against the live device without re-resolving (`--no-sync` keeps the
# hand-installed torch/toolchain wheels). CUDA_HOME is cleared so Numba links
# libnvvm/nvjitlink from the cu12 wheels, not a newer system CUDA. Full setup
# recipe and rationale: docs/testing.md.
test-gpu-cuda: ## Run the CUDA (T2) GPU suite on hardware (deps must be installed; see docs/testing.md)
env -u CUDA_HOME -u CUDA_PATH DEVMM_GPU=cuda uv run --no-sync \
pytest tests/test_cuda_gpu.py tests/test_integrations_gpu.py

# Coverage thresholds (see docs/testing.md): >= 90% overall, >= 95% on the
# core domain model + DLPack layer. Residual uncovered lines carry reasoned
# `# pragma: no cover` / exclusions (see pyproject.toml).
Expand Down
27 changes: 27 additions & 0 deletions docs/adr/0003-gpu-suite-waiver-for-0.1.0.md
Original file line number Diff line number Diff line change
Expand Up @@ -59,3 +59,30 @@ Release 0.1.0 with the T2 and T3 hardware suites **waived**, on these terms:
- This ADR is append-only per repository convention; running the suites on
hardware supersedes it with a new ADR (or an appended record) rather than
an edit to the decision.

## Manual-run record — CUDA (T2), 2026-07-18

First execution of the T2 suite on real hardware, per the "manual-run record"
term above. ROCm (T3) remains unexecuted.

- **Hardware / driver**: 4× NVIDIA GH200 120GB (aarch64), driver 590.48.01,
system CUDA 13.1.
- **Consumers**: `rmm-cu12` 26.06.00, `cupy-cuda12x` 14.1.1,
`torch` 2.11.0+cu128, `numba` 0.66.0 + `numba-cuda` 0.30.4, `filecheck` 1.0.3.
Numba's JIT toolchain aligned on the cu12 wheels (`nvidia-cuda-nvcc-cu12` /
`nvidia-nvjitlink-cu12` both 12.8.93, `CUDA_HOME` unset) — see
[`docs/testing.md`](../testing.md#running-the-cuda-gpu-suite-on-hardware).
- **Result**: `tests/test_cuda_gpu.py` + `tests/test_integrations_gpu.py` —
**41 passed, 1 skipped** (the stream-race misorder canary is best-effort and
did not manifest). Reproduce with `make test-gpu-cuda`.
- **Fixes required to get green** (both against dependency versions newer than
the suite was written for, written to work across the old and new APIs):
the Numba EMM plugin now hands `MemoryPointer` a `ctypes.c_void_p`
(numba-cuda ≥ 0.30 only converts that type to a driver `CUdeviceptr`); the
rmm statistics test normalises `allocation_counts`, a `Statistics` object in
rmm ≥ 26.06 and a dict before.

The Decision's release-blocking clause is met rather than deferred: both
failures were fixed before any tag ships, so no T2 failure is outstanding
against a release. T2 having now run on hardware, only the ROCm (T3) half of
the waiver still has work to do; a new ADR supersedes this one when T3 runs.
43 changes: 43 additions & 0 deletions docs/testing.md
Original file line number Diff line number Diff line change
Expand Up @@ -32,6 +32,49 @@ GPU** — that is the backbone of the suite (design §9).
per platform: rmm-pool + torch/cupy `from_dlpack` round trips and a
stream-race canary.

## Running the CUDA GPU suite on hardware

The T2 suites (`tests/test_cuda_gpu.py`, `tests/test_integrations_gpu.py`) are
skip-gated behind the `gpu_cuda` marker and opt in with `DEVMM_GPU=cuda`
(`tests/conftest.py`). On a CUDA-12 host:

1. Install the consumer stack — the `gpu-test-cuda` extra, then a PyTorch CUDA
wheel. `torch` is not in the extra: PyTorch's CUDA builds, including the
aarch64/sbsa ones, live on its own index, so resolving it from PyPI would
install a build the second command replaces, along with a conflicting set
of `nvidia-*` runtime wheels. The extra is Python 3.12–3.14 only (no
numba-cuda wheel exists for 3.15); on 3.15+ it installs nothing, silently.

```sh
uv pip install '.[test,gpu-test-cuda]'
uv pip install torch --index-url https://download.pytorch.org/whl/cu128
```

2. Run the suite (or `make test-gpu-cuda`, which wraps this):

```sh
env -u CUDA_HOME -u CUDA_PATH DEVMM_GPU=cuda uv run --no-sync \
pytest tests/test_cuda_gpu.py tests/test_integrations_gpu.py
```

Two hardware-only gotchas, both about Numba's JIT toolchain rather than devmm:

- **libnvvm ↔ nvjitlink alignment.** Numba compiles a kernel with `libnvvm`
and links it with `nvjitlink`; if `libnvvm` is newer than `nvjitlink`, the
link fails with `nvJitLinkError: ERROR 4 in nvvmAddNVVMContainerToProgram,
may need newer version of nvJitLink library`. The `gpu-test-cuda` extra keeps
`nvidia-cuda-nvcc-cu12` (libnvvm) in the 12.8 series so it stays minor-aligned
with the nvjitlink PyTorch's wheels pin. That pin only makes sense on a
CUDA-12 stack, which is the other half of why the extra is marked
`python_full_version < '3.15'`: from 3.15 on the resolver reaches for the
CUDA-13 `cuda-toolkit` wheels and the cu12 libnvvm becomes the mismatched
half.
- **A newer system CUDA shadows the wheels.** If `CUDA_HOME` points at a system
CUDA newer than the cu12 wheels (e.g. a 13.x module on an HPC box), Numba
picks up that `libnvvm` and the alignment above breaks again. Clearing
`CUDA_HOME`/`CUDA_PATH` (as the run command and `make test-gpu-cuda` do) makes
Numba use the cu12 wheels instead.

## Coverage targets

The behavioural bar comes first: every shipped `LayoutPolicy`, every CPU MR,
Expand Down
22 changes: 22 additions & 0 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -32,6 +32,28 @@ cupy = ["cupy"]
numba = ["numba"]
# Libraries needed to exercise DLPack round-trips in the test suite (§9).
test = ["numpy", "array-api-strict"]
# Consumer stack for the CUDA (T2) GPU suite on hardware
# (tests/test_cuda_gpu.py, tests/test_integrations_gpu.py), targeting a CUDA-12
# runtime. `filecheck` is pulled in by Numba's own EMM test protocol.
# `nvidia-cuda-nvcc-cu12` supplies libnvvm: keep it in the 12.8 series so it
# stays minor-aligned with the nvjitlink PyTorch's wheels pin, or Numba's JIT
# linker rejects the NVVM container. The `python_full_version < '3.15'` markers
# bound the extra to where this stack exists: numba-cuda publishes no cp315
# wheel, and from 3.15 on the resolver reaches for the unsuffixed (CUDA-13)
# `cuda-toolkit` stack, which the cu12 nvcc pin above would contradict.
# `torch` is deliberately absent — its CUDA builds live on PyTorch's own index,
# so installing it from here would pull a PyPI build and a conflicting nvidia-*
# set that step 2 of the recipe then replaces. Install it separately, and on a
# host with a newer system CUDA unset `CUDA_HOME` so Numba links these wheels —
# see docs/testing.md for the full recipe.
gpu-test-cuda = [
"devmm[cuda]; python_full_version < '3.15'",
"cupy-cuda12x; python_full_version < '3.15'",
"numba; python_full_version < '3.15'",
"numba-cuda; python_full_version < '3.15'",
"filecheck; python_full_version < '3.15'",
"nvidia-cuda-nvcc-cu12>=12.8,<12.9; python_full_version < '3.15'",
]

[project.urls]
Homepage = "https://github.com/grAItools/devmm"
Expand Down
11 changes: 10 additions & 1 deletion src/devmm/integrations/numba.py
Original file line number Diff line number Diff line change
Expand Up @@ -95,9 +95,18 @@ def memalloc(self, size: int) -> Any:
# makes.
buffer = DeviceBuffer(size, mr=mr, stream=ForeignHandleStream(device, 0))
self.allocations[buffer.ptr] = buffer
# The device pointer goes in as a `ctypes.c_void_p`, the type
# Numba's EMM plugin docs prescribe: numba-cuda >= 0.30 builds a
# driver `CUdeviceptr` only inside an
# `isinstance(pointer, ctypes.c_void_p)` branch and stores any
# other type raw, which the driver then rejects; older Numba read
# `.value`, which `c_void_p` also has. `c_void_p(0).value` is None
# rather than 0, but that NULL must not arise: MRs must return
# non-null pointers even for zero-byte allocations (contract
# pinned by `devmm.testing.mr_conformance`).
return numba_cuda.MemoryPointer(
self.context,
ctypes.c_uint64(buffer.ptr),
ctypes.c_void_p(buffer.ptr),
size,
finalizer=_make_finalizer(self.allocations, buffer.ptr),
)
Expand Down
29 changes: 26 additions & 3 deletions tests/_integration_fakes.py
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@

from __future__ import annotations

import ctypes
from collections.abc import Callable
from types import ModuleType, TracebackType
from typing import Any
Expand Down Expand Up @@ -140,9 +141,25 @@ def take_ptr(self, nbytes: int) -> int:
return ptr


class FakeNumbaCUdeviceptr:
"""`cuda.bindings.driver.CUdeviceptr` double: the driver pointer type
numba-cuda converts an incoming `ctypes.c_void_p` into."""

def __init__(self, value: int | None) -> None:
self.value = value


class FakeNumbaMemoryPointer:
"""`numba.cuda.MemoryPointer` double: frees exactly once through the
finalizer, as Numba's deallocation machinery does."""
"""`numba.cuda.MemoryPointer` double: enforces numba-cuda >= 0.30's
pointer contract and frees exactly once through the finalizer, as Numba's
deallocation machinery does.

Numba converts the incoming pointer to a driver `CUdeviceptr` only under
`isinstance(pointer, ctypes.c_void_p)`; any other type is stored raw and
fails later inside a driver call. This double raises on that path instead,
so passing the wrong ctypes type is caught off-hardware (a driver pointer
would also work upstream; devmm always passes `c_void_p`).
"""

def __init__(
self,
Expand All @@ -152,8 +169,14 @@ def __init__(
owner: Any = None,
finalizer: Callable[[], None] | None = None,
) -> None:
if not isinstance(pointer, ctypes.c_void_p):
raise TypeError(
"device pointer must be a ctypes.c_void_p, got "
f"{type(pointer).__name__}; Numba converts only that type to a "
"driver pointer and stores the rest raw"
)
self.context = context
self.pointer = pointer
self.pointer = FakeNumbaCUdeviceptr(pointer.value)
self.size = size
self.owner = owner
self._finalizer = finalizer
Expand Down
22 changes: 20 additions & 2 deletions tests/test_cuda_gpu.py
Original file line number Diff line number Diff line change
Expand Up @@ -256,6 +256,24 @@ def test_stream_race_canary_can_misorder_without_the_handoff(
pytest.skip("race did not manifest with the handoff disabled (best-effort)")


_STATISTICS_FIELDS = ("current_bytes", "peak_bytes", "total_bytes")


def _allocation_counts(adaptor: Any) -> dict[str, int]:
"""rmm >= 26.06 reports `allocation_counts` as a `Statistics` object;
older rmm returned a plain dict. Normalise to a dict either way, failing
loudly if a field is missing rather than reporting a subset — an upstream
rename must not read as a pass."""
counts = adaptor.allocation_counts
if not isinstance(counts, dict):
counts = {
name: getattr(counts, name) for name in _STATISTICS_FIELDS if hasattr(counts, name)
}
missing = [name for name in _STATISTICS_FIELDS if name not in counts]
assert not missing, f"rmm allocation_counts is missing {missing}"
return counts


def test_rmm_pool_statistics_agree_with_statistics_adaptor(
runtime: CudaRuntime, stream: Stream
) -> None:
Expand All @@ -266,11 +284,11 @@ def test_rmm_pool_statistics_agree_with_statistics_adaptor(
mr = StatisticsAdaptor(RmmMemoryResource(upstream, _DEVICE))
sizes = (256, 1024, 4096)
ptrs = [mr.allocate(nbytes, stream) for nbytes in sizes]
counts = upstream.allocation_counts
counts = _allocation_counts(upstream)
assert counts["current_bytes"] == mr.current_bytes == sum(sizes)
assert counts["peak_bytes"] == mr.peak_bytes == sum(sizes)
for ptr, nbytes in zip(ptrs, sizes, strict=True):
mr.deallocate(ptr, nbytes, stream)
counts = upstream.allocation_counts
counts = _allocation_counts(upstream)
assert counts["current_bytes"] == mr.current_bytes == 0
assert counts["total_bytes"] == mr.total_bytes == sum(sizes)
16 changes: 15 additions & 1 deletion tests/test_integrations_numba.py
Original file line number Diff line number Diff line change
Expand Up @@ -23,6 +23,7 @@
from devmm.testing import RecordingMemoryResource
from tests._integration_fakes import (
FakeNumbaCuda,
FakeNumbaCUdeviceptr,
FakeNumbaMemoryInfo,
FakeNumbaMemoryPointer,
)
Expand Down Expand Up @@ -91,7 +92,10 @@ def test_memalloc_allocates_through_the_current_devmm_mr(
assert isinstance(memory, FakeNumbaMemoryPointer)
name, ptr, nbytes, stream = mr.calls[0]
assert (name, nbytes) == ("allocate", 256)
assert isinstance(memory.pointer, ctypes.c_uint64)
# The pointer arrived as the `c_void_p` the EMM contract requires, so
# Numba converted it to a driver pointer (contract enforced by
# `FakeNumbaMemoryPointer`, rationale in `integrations.numba`).
assert isinstance(memory.pointer, FakeNumbaCUdeviceptr)
assert memory.pointer.value == ptr
assert memory.size == 256
# The EMM protocol carries no stream, so allocations ride the
Expand Down Expand Up @@ -151,6 +155,16 @@ def test_get_memory_info_without_mr_support_raises(self, fake_numba: FakeNumbaCu
manager.get_memory_info()


class TestFakeContract:
"""The double's own guarantees, on which the plugin tests above rest."""

def test_a_pointer_that_is_not_a_c_void_p_is_refused(self) -> None:
# Were the double to accept any ctypes integer, the memalloc test
# would stop guarding the pointer type Numba actually converts.
with pytest.raises(TypeError, match="c_void_p"):
FakeNumbaMemoryPointer(None, ctypes.c_uint64(0x1000), 256)


class TestInstall:
def test_install_sets_the_manager_and_uninstall_restores_none(
self, fake_numba: FakeNumbaCuda
Expand Down
Loading
Loading