diff --git a/CHANGELOG.md b/CHANGELOG.md index b43bc5f..2fe91ce 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -8,6 +8,18 @@ and this project adheres to ## [Unreleased] +### Added + +- `gpu-test-cuda` extra and `make test-gpu-cuda` to run the CUDA (T2) hardware + suite, with the setup recipe in [`docs/testing.md`](docs/testing.md#running-the-cuda-gpu-suite-on-hardware). + First hardware run recorded in [ADR 0003](docs/adr/0003-gpu-suite-waiver-for-0.1.0.md). + +### Fixed + +- `integrations.numba`: `DevmmEMMPlugin` passes the device pointer as a + `ctypes.c_void_p`, the only type numba-cuda ≥ 0.30 converts to a driver + pointer. + ## [0.1.0] - 2026-07-17 - Repository layout: the package lives at `src/devmm/` (src-layout); imports diff --git a/Makefile b/Makefile index 9715a21..8fde7c3 100644 --- a/Makefile +++ b/Makefile @@ -1,9 +1,9 @@ -.PHONY: help test test-all test-devmode lint typecheck fmt fmt-check coverage verify gate-all \ - release-gate dev clean +.PHONY: help test test-all test-devmode test-gpu-cuda lint typecheck fmt fmt-check coverage \ + verify gate-all release-gate dev clean help: ## Show this help @awk 'BEGIN {FS = ":.*##"; printf "Usage: make \033[36m\033[0m\n\nTargets:\n"} \ - /^[a-zA-Z0-9_-]+:.*?##/ { printf " \033[36m%-12s\033[0m %s\n", $$1, $$2 }' $(MAKEFILE_LIST) + /^[a-zA-Z0-9_-]+:.*?##/ { printf " \033[36m%-14s\033[0m %s\n", $$1, $$2 }' $(MAKEFILE_LIST) # Tests run with the `test` extra (numpy, array-api-strict): the suite's # differential oracles and DLPack round-trips need a consumer library (§9). @@ -18,6 +18,16 @@ test-all: test ## Run the full suite (override to add integration/e2e) test-devmode: ## Run the suite under PYTHONDEVMODE=1 with faulthandler PYTHONDEVMODE=1 PYTHONFAULTHANDLER=1 uv run --extra test pytest -q +# The CUDA (T2) hardware suite (docs/adr/0003). Install the deps first — +# `uv pip install '.[test,gpu-test-cuda]'` plus the PyTorch CUDA wheel — then +# this runs against the live device without re-resolving (`--no-sync` keeps the +# hand-installed torch/toolchain wheels). CUDA_HOME is cleared so Numba links +# libnvvm/nvjitlink from the cu12 wheels, not a newer system CUDA. Full setup +# recipe and rationale: docs/testing.md. +test-gpu-cuda: ## Run the CUDA (T2) GPU suite on hardware (deps must be installed; see docs/testing.md) + env -u CUDA_HOME -u CUDA_PATH DEVMM_GPU=cuda uv run --no-sync \ + pytest tests/test_cuda_gpu.py tests/test_integrations_gpu.py + # Coverage thresholds (see docs/testing.md): >= 90% overall, >= 95% on the # core domain model + DLPack layer. Residual uncovered lines carry reasoned # `# pragma: no cover` / exclusions (see pyproject.toml). diff --git a/docs/adr/0003-gpu-suite-waiver-for-0.1.0.md b/docs/adr/0003-gpu-suite-waiver-for-0.1.0.md index d330a44..392e3a5 100644 --- a/docs/adr/0003-gpu-suite-waiver-for-0.1.0.md +++ b/docs/adr/0003-gpu-suite-waiver-for-0.1.0.md @@ -59,3 +59,30 @@ Release 0.1.0 with the T2 and T3 hardware suites **waived**, on these terms: - This ADR is append-only per repository convention; running the suites on hardware supersedes it with a new ADR (or an appended record) rather than an edit to the decision. + +## Manual-run record — CUDA (T2), 2026-07-18 + +First execution of the T2 suite on real hardware, per the "manual-run record" +term above. ROCm (T3) remains unexecuted. + +- **Hardware / driver**: 4× NVIDIA GH200 120GB (aarch64), driver 590.48.01, + system CUDA 13.1. +- **Consumers**: `rmm-cu12` 26.06.00, `cupy-cuda12x` 14.1.1, + `torch` 2.11.0+cu128, `numba` 0.66.0 + `numba-cuda` 0.30.4, `filecheck` 1.0.3. + Numba's JIT toolchain aligned on the cu12 wheels (`nvidia-cuda-nvcc-cu12` / + `nvidia-nvjitlink-cu12` both 12.8.93, `CUDA_HOME` unset) — see + [`docs/testing.md`](../testing.md#running-the-cuda-gpu-suite-on-hardware). +- **Result**: `tests/test_cuda_gpu.py` + `tests/test_integrations_gpu.py` — + **41 passed, 1 skipped** (the stream-race misorder canary is best-effort and + did not manifest). Reproduce with `make test-gpu-cuda`. +- **Fixes required to get green** (both against dependency versions newer than + the suite was written for, written to work across the old and new APIs): + the Numba EMM plugin now hands `MemoryPointer` a `ctypes.c_void_p` + (numba-cuda ≥ 0.30 only converts that type to a driver `CUdeviceptr`); the + rmm statistics test normalises `allocation_counts`, a `Statistics` object in + rmm ≥ 26.06 and a dict before. + +The Decision's release-blocking clause is met rather than deferred: both +failures were fixed before any tag ships, so no T2 failure is outstanding +against a release. T2 having now run on hardware, only the ROCm (T3) half of +the waiver still has work to do; a new ADR supersedes this one when T3 runs. diff --git a/docs/testing.md b/docs/testing.md index 8fe673b..2da0c87 100644 --- a/docs/testing.md +++ b/docs/testing.md @@ -32,6 +32,49 @@ GPU** — that is the backbone of the suite (design §9). per platform: rmm-pool + torch/cupy `from_dlpack` round trips and a stream-race canary. +## Running the CUDA GPU suite on hardware + +The T2 suites (`tests/test_cuda_gpu.py`, `tests/test_integrations_gpu.py`) are +skip-gated behind the `gpu_cuda` marker and opt in with `DEVMM_GPU=cuda` +(`tests/conftest.py`). On a CUDA-12 host: + +1. Install the consumer stack — the `gpu-test-cuda` extra, then a PyTorch CUDA + wheel. `torch` is not in the extra: PyTorch's CUDA builds, including the + aarch64/sbsa ones, live on its own index, so resolving it from PyPI would + install a build the second command replaces, along with a conflicting set + of `nvidia-*` runtime wheels. The extra is Python 3.12–3.14 only (no + numba-cuda wheel exists for 3.15); on 3.15+ it installs nothing, silently. + + ```sh + uv pip install '.[test,gpu-test-cuda]' + uv pip install torch --index-url https://download.pytorch.org/whl/cu128 + ``` + +2. Run the suite (or `make test-gpu-cuda`, which wraps this): + + ```sh + env -u CUDA_HOME -u CUDA_PATH DEVMM_GPU=cuda uv run --no-sync \ + pytest tests/test_cuda_gpu.py tests/test_integrations_gpu.py + ``` + +Two hardware-only gotchas, both about Numba's JIT toolchain rather than devmm: + +- **libnvvm ↔ nvjitlink alignment.** Numba compiles a kernel with `libnvvm` + and links it with `nvjitlink`; if `libnvvm` is newer than `nvjitlink`, the + link fails with `nvJitLinkError: ERROR 4 in nvvmAddNVVMContainerToProgram, + may need newer version of nvJitLink library`. The `gpu-test-cuda` extra keeps + `nvidia-cuda-nvcc-cu12` (libnvvm) in the 12.8 series so it stays minor-aligned + with the nvjitlink PyTorch's wheels pin. That pin only makes sense on a + CUDA-12 stack, which is the other half of why the extra is marked + `python_full_version < '3.15'`: from 3.15 on the resolver reaches for the + CUDA-13 `cuda-toolkit` wheels and the cu12 libnvvm becomes the mismatched + half. +- **A newer system CUDA shadows the wheels.** If `CUDA_HOME` points at a system + CUDA newer than the cu12 wheels (e.g. a 13.x module on an HPC box), Numba + picks up that `libnvvm` and the alignment above breaks again. Clearing + `CUDA_HOME`/`CUDA_PATH` (as the run command and `make test-gpu-cuda` do) makes + Numba use the cu12 wheels instead. + ## Coverage targets The behavioural bar comes first: every shipped `LayoutPolicy`, every CPU MR, diff --git a/pyproject.toml b/pyproject.toml index 671ffbc..646ec11 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -32,6 +32,28 @@ cupy = ["cupy"] numba = ["numba"] # Libraries needed to exercise DLPack round-trips in the test suite (§9). test = ["numpy", "array-api-strict"] +# Consumer stack for the CUDA (T2) GPU suite on hardware +# (tests/test_cuda_gpu.py, tests/test_integrations_gpu.py), targeting a CUDA-12 +# runtime. `filecheck` is pulled in by Numba's own EMM test protocol. +# `nvidia-cuda-nvcc-cu12` supplies libnvvm: keep it in the 12.8 series so it +# stays minor-aligned with the nvjitlink PyTorch's wheels pin, or Numba's JIT +# linker rejects the NVVM container. The `python_full_version < '3.15'` markers +# bound the extra to where this stack exists: numba-cuda publishes no cp315 +# wheel, and from 3.15 on the resolver reaches for the unsuffixed (CUDA-13) +# `cuda-toolkit` stack, which the cu12 nvcc pin above would contradict. +# `torch` is deliberately absent — its CUDA builds live on PyTorch's own index, +# so installing it from here would pull a PyPI build and a conflicting nvidia-* +# set that step 2 of the recipe then replaces. Install it separately, and on a +# host with a newer system CUDA unset `CUDA_HOME` so Numba links these wheels — +# see docs/testing.md for the full recipe. +gpu-test-cuda = [ + "devmm[cuda]; python_full_version < '3.15'", + "cupy-cuda12x; python_full_version < '3.15'", + "numba; python_full_version < '3.15'", + "numba-cuda; python_full_version < '3.15'", + "filecheck; python_full_version < '3.15'", + "nvidia-cuda-nvcc-cu12>=12.8,<12.9; python_full_version < '3.15'", +] [project.urls] Homepage = "https://github.com/grAItools/devmm" diff --git a/src/devmm/integrations/numba.py b/src/devmm/integrations/numba.py index 55075b2..f6cb6b4 100644 --- a/src/devmm/integrations/numba.py +++ b/src/devmm/integrations/numba.py @@ -95,9 +95,18 @@ def memalloc(self, size: int) -> Any: # makes. buffer = DeviceBuffer(size, mr=mr, stream=ForeignHandleStream(device, 0)) self.allocations[buffer.ptr] = buffer + # The device pointer goes in as a `ctypes.c_void_p`, the type + # Numba's EMM plugin docs prescribe: numba-cuda >= 0.30 builds a + # driver `CUdeviceptr` only inside an + # `isinstance(pointer, ctypes.c_void_p)` branch and stores any + # other type raw, which the driver then rejects; older Numba read + # `.value`, which `c_void_p` also has. `c_void_p(0).value` is None + # rather than 0, but that NULL must not arise: MRs must return + # non-null pointers even for zero-byte allocations (contract + # pinned by `devmm.testing.mr_conformance`). return numba_cuda.MemoryPointer( self.context, - ctypes.c_uint64(buffer.ptr), + ctypes.c_void_p(buffer.ptr), size, finalizer=_make_finalizer(self.allocations, buffer.ptr), ) diff --git a/tests/_integration_fakes.py b/tests/_integration_fakes.py index cbc46dc..ac95616 100644 --- a/tests/_integration_fakes.py +++ b/tests/_integration_fakes.py @@ -8,6 +8,7 @@ from __future__ import annotations +import ctypes from collections.abc import Callable from types import ModuleType, TracebackType from typing import Any @@ -140,9 +141,25 @@ def take_ptr(self, nbytes: int) -> int: return ptr +class FakeNumbaCUdeviceptr: + """`cuda.bindings.driver.CUdeviceptr` double: the driver pointer type + numba-cuda converts an incoming `ctypes.c_void_p` into.""" + + def __init__(self, value: int | None) -> None: + self.value = value + + class FakeNumbaMemoryPointer: - """`numba.cuda.MemoryPointer` double: frees exactly once through the - finalizer, as Numba's deallocation machinery does.""" + """`numba.cuda.MemoryPointer` double: enforces numba-cuda >= 0.30's + pointer contract and frees exactly once through the finalizer, as Numba's + deallocation machinery does. + + Numba converts the incoming pointer to a driver `CUdeviceptr` only under + `isinstance(pointer, ctypes.c_void_p)`; any other type is stored raw and + fails later inside a driver call. This double raises on that path instead, + so passing the wrong ctypes type is caught off-hardware (a driver pointer + would also work upstream; devmm always passes `c_void_p`). + """ def __init__( self, @@ -152,8 +169,14 @@ def __init__( owner: Any = None, finalizer: Callable[[], None] | None = None, ) -> None: + if not isinstance(pointer, ctypes.c_void_p): + raise TypeError( + "device pointer must be a ctypes.c_void_p, got " + f"{type(pointer).__name__}; Numba converts only that type to a " + "driver pointer and stores the rest raw" + ) self.context = context - self.pointer = pointer + self.pointer = FakeNumbaCUdeviceptr(pointer.value) self.size = size self.owner = owner self._finalizer = finalizer diff --git a/tests/test_cuda_gpu.py b/tests/test_cuda_gpu.py index 5096be0..07eabf5 100644 --- a/tests/test_cuda_gpu.py +++ b/tests/test_cuda_gpu.py @@ -256,6 +256,24 @@ def test_stream_race_canary_can_misorder_without_the_handoff( pytest.skip("race did not manifest with the handoff disabled (best-effort)") +_STATISTICS_FIELDS = ("current_bytes", "peak_bytes", "total_bytes") + + +def _allocation_counts(adaptor: Any) -> dict[str, int]: + """rmm >= 26.06 reports `allocation_counts` as a `Statistics` object; + older rmm returned a plain dict. Normalise to a dict either way, failing + loudly if a field is missing rather than reporting a subset — an upstream + rename must not read as a pass.""" + counts = adaptor.allocation_counts + if not isinstance(counts, dict): + counts = { + name: getattr(counts, name) for name in _STATISTICS_FIELDS if hasattr(counts, name) + } + missing = [name for name in _STATISTICS_FIELDS if name not in counts] + assert not missing, f"rmm allocation_counts is missing {missing}" + return counts + + def test_rmm_pool_statistics_agree_with_statistics_adaptor( runtime: CudaRuntime, stream: Stream ) -> None: @@ -266,11 +284,11 @@ def test_rmm_pool_statistics_agree_with_statistics_adaptor( mr = StatisticsAdaptor(RmmMemoryResource(upstream, _DEVICE)) sizes = (256, 1024, 4096) ptrs = [mr.allocate(nbytes, stream) for nbytes in sizes] - counts = upstream.allocation_counts + counts = _allocation_counts(upstream) assert counts["current_bytes"] == mr.current_bytes == sum(sizes) assert counts["peak_bytes"] == mr.peak_bytes == sum(sizes) for ptr, nbytes in zip(ptrs, sizes, strict=True): mr.deallocate(ptr, nbytes, stream) - counts = upstream.allocation_counts + counts = _allocation_counts(upstream) assert counts["current_bytes"] == mr.current_bytes == 0 assert counts["total_bytes"] == mr.total_bytes == sum(sizes) diff --git a/tests/test_integrations_numba.py b/tests/test_integrations_numba.py index 7fcfe0c..577cf44 100644 --- a/tests/test_integrations_numba.py +++ b/tests/test_integrations_numba.py @@ -23,6 +23,7 @@ from devmm.testing import RecordingMemoryResource from tests._integration_fakes import ( FakeNumbaCuda, + FakeNumbaCUdeviceptr, FakeNumbaMemoryInfo, FakeNumbaMemoryPointer, ) @@ -91,7 +92,10 @@ def test_memalloc_allocates_through_the_current_devmm_mr( assert isinstance(memory, FakeNumbaMemoryPointer) name, ptr, nbytes, stream = mr.calls[0] assert (name, nbytes) == ("allocate", 256) - assert isinstance(memory.pointer, ctypes.c_uint64) + # The pointer arrived as the `c_void_p` the EMM contract requires, so + # Numba converted it to a driver pointer (contract enforced by + # `FakeNumbaMemoryPointer`, rationale in `integrations.numba`). + assert isinstance(memory.pointer, FakeNumbaCUdeviceptr) assert memory.pointer.value == ptr assert memory.size == 256 # The EMM protocol carries no stream, so allocations ride the @@ -151,6 +155,16 @@ def test_get_memory_info_without_mr_support_raises(self, fake_numba: FakeNumbaCu manager.get_memory_info() +class TestFakeContract: + """The double's own guarantees, on which the plugin tests above rest.""" + + def test_a_pointer_that_is_not_a_c_void_p_is_refused(self) -> None: + # Were the double to accept any ctypes integer, the memalloc test + # would stop guarding the pointer type Numba actually converts. + with pytest.raises(TypeError, match="c_void_p"): + FakeNumbaMemoryPointer(None, ctypes.c_uint64(0x1000), 256) + + class TestInstall: def test_install_sets_the_manager_and_uninstall_restores_none( self, fake_numba: FakeNumbaCuda diff --git a/uv.lock b/uv.lock index d3d5783..caa699d 100644 --- a/uv.lock +++ b/uv.lock @@ -170,6 +170,29 @@ wheels = [ { url = "https://files.pythonhosted.org/packages/33/10/c71a07cd2a1d4db119bada1848b4752a874ccfe4927d419bfdd05f250920/cuda_bindings-12.9.7-cp314-cp314t-win_amd64.whl", hash = "sha256:ece8dfbc22e6de96a26940ab9887eb3cfe1fc1bc3966169391cdb866bb82bb64", size = 8208198, upload-time = "2026-05-27T18:44:39.053Z" }, ] +[[package]] +name = "cuda-core" +version = "1.1.0" +source = { registry = "https://pypi.org/simple" } +dependencies = [ + { name = "cuda-pathfinder", marker = "python_full_version < '3.15'" }, + { name = "numpy", marker = "python_full_version < '3.15'" }, +] +wheels = [ + { url = "https://files.pythonhosted.org/packages/a2/fe/5a511a9f8990e6732af7ab4e88dbfcea262ef1dba4f5aeef10fd7138886d/cuda_core-1.1.0-cp312-cp312-manylinux_2_24_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:5bdbe75ef002e3ba1685f1f53c174674deec33d26961c6d91dd52571e8a4ab79", size = 5796917, upload-time = "2026-07-09T23:13:22.921Z" }, + { url = "https://files.pythonhosted.org/packages/8a/cc/d87a4a3e07d11a11f05992f99a8af11abc1e857110ed923b26bd29b32722/cuda_core-1.1.0-cp312-cp312-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:2cd8637074b450f9e99137cbcf4c7b84e786de1beb8671cabc34a9e885985f67", size = 6164535, upload-time = "2026-07-09T23:13:24.713Z" }, + { url = "https://files.pythonhosted.org/packages/cd/55/248c86654e36d3611e382072101b9346ae477f34d6233679c451e2ae5267/cuda_core-1.1.0-cp312-cp312-win_amd64.whl", hash = "sha256:6214debf3c2bbfa28fde67b346b6cd3d8fbf001127d2034fcd177c36336ab637", size = 5569720, upload-time = "2026-07-09T23:13:26.9Z" }, + { url = "https://files.pythonhosted.org/packages/17/87/8ac49f49629fd4ec8f8bf74a30fe17169d8105467efd8bdf97223677a601/cuda_core-1.1.0-cp313-cp313-manylinux_2_24_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:2fb1b994a6252a1fd3d14314e41e169307bb33a35c1d78333747f567a3201a3f", size = 5730557, upload-time = "2026-07-09T23:13:28.518Z" }, + { url = "https://files.pythonhosted.org/packages/ee/58/ad675f0b2eca221bdb8ddf6b5d4ab9d093e96b48397de3bb336514b8acc0/cuda_core-1.1.0-cp313-cp313-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:abe884a521804c10e8fd63bfcf52c3d010fc2f130c0e1141826050b2b2c7a477", size = 6105170, upload-time = "2026-07-09T23:13:30.369Z" }, + { url = "https://files.pythonhosted.org/packages/fb/d0/d7a60381c405708cdd1088a68f8e60ca63f6e593b2a7c793983ed6303d63/cuda_core-1.1.0-cp313-cp313-win_amd64.whl", hash = "sha256:1e16d74de7ffded40e6a1b0d1fecebdf60dfcd61571424f73971c5c62bb65a1e", size = 5541307, upload-time = "2026-07-09T23:13:32.13Z" }, + { url = "https://files.pythonhosted.org/packages/e0/14/db2256f3045fa17763726c09d519be4f876bfadeed85cfea539b42ecae25/cuda_core-1.1.0-cp314-cp314-manylinux_2_24_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:0a47d03b0bdb57f82ce3e8563b4fc4d7432be629ac07ad749ecd6821d683288d", size = 5804417, upload-time = "2026-07-09T23:13:33.97Z" }, + { url = "https://files.pythonhosted.org/packages/32/d4/fad332f1923d47a0eafac37fde5259aad8987969eed6cceb33cb04f035a4/cuda_core-1.1.0-cp314-cp314-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:dc8103c78b9c15ad6b1ea48de7763b429c8fde42bec43a9dd0550868994a0ce4", size = 6136554, upload-time = "2026-07-09T23:13:35.752Z" }, + { url = "https://files.pythonhosted.org/packages/6f/15/8d1f52cba966f7a130108132b197b13fbcacee715aa764baa79cd67f9169/cuda_core-1.1.0-cp314-cp314-win_amd64.whl", hash = "sha256:d0531306d2894523e261418491a68b8385481c9fa0552d78fd3ba6b746a20942", size = 5668580, upload-time = "2026-07-09T23:13:37.654Z" }, + { url = "https://files.pythonhosted.org/packages/b2/8e/6eb564dbda45b5672454078e06d9ba28d2579ef7ae71e78b6e16da0e2a11/cuda_core-1.1.0-cp314-cp314t-manylinux_2_24_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:656b7b755b3881a5752311b343ed20267a47ad630d25f625b9d7c89b1f90a295", size = 6002276, upload-time = "2026-07-09T23:13:39.427Z" }, + { url = "https://files.pythonhosted.org/packages/5a/e0/572f2179b4353a0a4ff1265b8708629192ff1245fabdc0576e5276731f24/cuda_core-1.1.0-cp314-cp314t-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:3982a67eb16a8b9e9881455aa3d123db049891fab48b028434323bbde935e006", size = 6255746, upload-time = "2026-07-09T23:13:41.559Z" }, + { url = "https://files.pythonhosted.org/packages/2a/4b/7cbf954bc3c756e79844589d18f8552845e0cf01132ca0b701181b45226e/cuda_core-1.1.0-cp314-cp314t-win_amd64.whl", hash = "sha256:e4817919d37253134e06765afe551d964c1dab3f900679f4467de60a94b1bb56", size = 6553516, upload-time = "2026-07-09T23:13:43.186Z" }, +] + [[package]] name = "cuda-pathfinder" version = "1.5.6" @@ -199,6 +222,28 @@ dependencies = [ ] sdist = { url = "https://files.pythonhosted.org/packages/85/86/e99f52f4a50e5faaf4932cf3574d64d870bce224314be7fca91a6d02cda7/cupy-14.1.1.tar.gz", hash = "sha256:3d4d193a5b028e1823d1d44fe6268364d8fb76c3793ad4e28e1427c56ceb2f80", size = 3953009, upload-time = "2026-06-01T04:55:30.379Z" } +[[package]] +name = "cupy-cuda12x" +version = "14.1.1" +source = { registry = "https://pypi.org/simple" } +dependencies = [ + { name = "cuda-pathfinder", marker = "python_full_version < '3.15'" }, + { name = "numpy", marker = "python_full_version < '3.15'" }, +] +wheels = [ + { url = "https://files.pythonhosted.org/packages/a3/6e/290ee2d7cc4ad63d66e67acfd7ff3026f2b648dd04449a1bf88ffaa36b1e/cupy_cuda12x-14.1.1-cp312-cp312-manylinux2014_aarch64.whl", hash = "sha256:7aae7d3bed37985e2aa39f0914b88ad90dbd3a6141d3e8198d73fce65859013c", size = 144383812, upload-time = "2026-06-01T04:52:23.799Z" }, + { url = "https://files.pythonhosted.org/packages/f6/6e/dc03c1ddc940f33b3d32803898e2fdae5c9538a2127a25f499494c84b183/cupy_cuda12x-14.1.1-cp312-cp312-manylinux2014_x86_64.whl", hash = "sha256:a1138f20080489a46209291498cd12f792226d0a57d50c64a586c162a875a069", size = 133516927, upload-time = "2026-06-01T04:52:35.765Z" }, + { url = "https://files.pythonhosted.org/packages/cc/da/d4a8045b533af634bc791572e8c87981065e4a27b5d3e09d0d4d285742fd/cupy_cuda12x-14.1.1-cp312-cp312-win_amd64.whl", hash = "sha256:85bebce86ffc25ecf31727b25da7b3793daf07b6fd9952704546af574d250988", size = 95238722, upload-time = "2026-06-01T04:52:46.296Z" }, + { url = "https://files.pythonhosted.org/packages/30/90/00fe874c47207b26c9b6ac950d0cecc533b4a145491641932df17e573f3c/cupy_cuda12x-14.1.1-cp313-cp313-manylinux2014_aarch64.whl", hash = "sha256:afbb3d1fa9484b0ae20d76372c5939a8c5da327e3fc8711b77b2354566cac355", size = 143920086, upload-time = "2026-06-01T04:52:51.726Z" }, + { url = "https://files.pythonhosted.org/packages/89/a4/c46ff91dba0dbe2a0a557974faf4c090a3159d6e7296431ca6846038d047/cupy_cuda12x-14.1.1-cp313-cp313-manylinux2014_x86_64.whl", hash = "sha256:76ea35469e2aa0a8332b88f72505ea2f7871a0bc8f9b0c87184f57e47c9aa3bf", size = 133071615, upload-time = "2026-06-01T04:52:57.428Z" }, + { url = "https://files.pythonhosted.org/packages/ec/a0/46778424035ad3fc920d49471f079687a054f74d179142e9520014c2514e/cupy_cuda12x-14.1.1-cp313-cp313-win_amd64.whl", hash = "sha256:64072f4139b44df38215f0519a6badc14138fa0e4bb5b2db44fe94d05f8b9c8b", size = 95219598, upload-time = "2026-06-01T04:53:02.774Z" }, + { url = "https://files.pythonhosted.org/packages/b6/99/d72336481264c3483b162ea128d58f80abb50009f1df82ca82905e0b8fd7/cupy_cuda12x-14.1.1-cp314-cp314-manylinux2014_aarch64.whl", hash = "sha256:22d0ff2755a7f29cb225d1d5fb979a73428c5534ea0bca91b0c02698e9948f84", size = 143788629, upload-time = "2026-06-01T04:53:09.404Z" }, + { url = "https://files.pythonhosted.org/packages/c7/77/c43a67e6809e03780d88caf690fa44a8b3152db2d8f848714bec327c9881/cupy_cuda12x-14.1.1-cp314-cp314-manylinux2014_x86_64.whl", hash = "sha256:1059581507343e7cf6231facce30932a195c7aad4fa7771d00e4a252683915a1", size = 132406367, upload-time = "2026-06-01T04:53:16.232Z" }, + { url = "https://files.pythonhosted.org/packages/7d/dc/96cd37de6da41239e02fc7f17e3364d60f99bd6816673d622916a06113ec/cupy_cuda12x-14.1.1-cp314-cp314-win_amd64.whl", hash = "sha256:e707e0eceee174d323be21652e87bb97be982e6966b5dc241756307df42842aa", size = 95793971, upload-time = "2026-06-01T04:53:21.478Z" }, + { url = "https://files.pythonhosted.org/packages/20/c6/0ddec1be851de546e883ae3da5f03c1ea69738628b38234dce4362b5e38b/cupy_cuda12x-14.1.1-cp314-cp314t-manylinux2014_aarch64.whl", hash = "sha256:e09897636b7468a90efa1152109f0b19ba49ebc9a423d5dbd4682ed589e57843", size = 144093057, upload-time = "2026-06-01T04:53:28.647Z" }, + { url = "https://files.pythonhosted.org/packages/a4/80/5e05de89ba61df072aab6f8a6ee3ffeec57db68a0a456825b3b4ce608426/cupy_cuda12x-14.1.1-cp314-cp314t-manylinux2014_x86_64.whl", hash = "sha256:238080487174268d0f09770fe518de7c5b206527bef5c6792aef7ba0626a1c48", size = 132635338, upload-time = "2026-06-01T04:53:35.229Z" }, +] + [[package]] name = "devmm" version = "0.1.0" @@ -211,6 +256,14 @@ cuda = [ cupy = [ { name = "cupy" }, ] +gpu-test-cuda = [ + { name = "cupy-cuda12x", marker = "python_full_version < '3.15'" }, + { name = "filecheck", marker = "python_full_version < '3.15'" }, + { name = "numba", marker = "python_full_version < '3.15'" }, + { name = "numba-cuda", marker = "python_full_version < '3.15'" }, + { name = "nvidia-cuda-nvcc-cu12", marker = "python_full_version < '3.15'" }, + { name = "rmm-cu12", marker = "python_full_version < '3.15'" }, +] numba = [ { name = "numba" }, ] @@ -236,11 +289,17 @@ requires-dist = [ { name = "amd-hipmm", marker = "extra == 'rocm'" }, { name = "array-api-strict", marker = "extra == 'test'" }, { name = "cupy", marker = "extra == 'cupy'" }, + { name = "cupy-cuda12x", marker = "python_full_version < '3.15' and extra == 'gpu-test-cuda'" }, + { name = "devmm", extras = ["cuda"], marker = "python_full_version < '3.15' and extra == 'gpu-test-cuda'" }, + { name = "filecheck", marker = "python_full_version < '3.15' and extra == 'gpu-test-cuda'" }, + { name = "numba", marker = "python_full_version < '3.15' and extra == 'gpu-test-cuda'" }, { name = "numba", marker = "extra == 'numba'" }, + { name = "numba-cuda", marker = "python_full_version < '3.15' and extra == 'gpu-test-cuda'" }, { name = "numpy", marker = "extra == 'test'" }, + { name = "nvidia-cuda-nvcc-cu12", marker = "python_full_version < '3.15' and extra == 'gpu-test-cuda'", specifier = ">=12.8,<12.9" }, { name = "rmm-cu12", marker = "extra == 'cuda'" }, ] -provides-extras = ["cuda", "rocm", "cupy", "numba", "test"] +provides-extras = ["cuda", "rocm", "cupy", "numba", "test", "gpu-test-cuda"] [package.metadata.requires-dev] dev = [ @@ -251,6 +310,15 @@ dev = [ { name = "ruff" }, ] +[[package]] +name = "filecheck" +version = "1.0.3" +source = { registry = "https://pypi.org/simple" } +sdist = { url = "https://files.pythonhosted.org/packages/1e/1c/10167ce597b4badb548cd8f70c7d8f36fdefafa41ba50e19f1a6a775708d/filecheck-1.0.3.tar.gz", hash = "sha256:ccb70500858e8f362f06d5c3e33c9c133785543ade50ddbeb9390681991f1b05", size = 20060, upload-time = "2025-08-19T10:00:06.256Z" } +wheels = [ + { url = "https://files.pythonhosted.org/packages/5c/40/69ca9ea803303e14301fff9d4931b6d080b9603e134df0419c55e9764df4/filecheck-1.0.3-py3-none-any.whl", hash = "sha256:1427d0e82d9c5209ec5cd9fb65745cae16d3003f800321f3c29f4b0729e68a19", size = 23943, upload-time = "2025-08-19T10:00:05.139Z" }, +] + [[package]] name = "hypothesis" version = "6.156.6" @@ -489,6 +557,29 @@ wheels = [ { url = "https://files.pythonhosted.org/packages/4c/f2/dca53d50b8f2289dd01954ace9da261e0487d5b74b188b4304e4ecc3492c/numba-0.66.0-cp314-cp314t-win_amd64.whl", hash = "sha256:d426178fb991a85714c43112a8ea7b9d9579ea856ad8dcdb9c1c3941903ba5be", size = 2804772, upload-time = "2026-07-01T23:12:44.399Z" }, ] +[[package]] +name = "numba-cuda" +version = "0.30.4" +source = { registry = "https://pypi.org/simple" } +dependencies = [ + { name = "cuda-bindings", marker = "python_full_version < '3.15'" }, + { name = "cuda-core", marker = "python_full_version < '3.15'" }, + { name = "cuda-pathfinder", marker = "python_full_version < '3.15'" }, + { name = "numba", marker = "python_full_version < '3.15'" }, + { name = "packaging", marker = "python_full_version < '3.15'" }, +] +wheels = [ + { url = "https://files.pythonhosted.org/packages/7c/14/47b329af0047e27e1f0fe2a2f5544f5c6660a01e1a1e0fe8445dddd5747f/numba_cuda-0.30.4-cp312-cp312-manylinux_2_24_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:d363e80f0e98d3e6616ced13f226b0e70295085ad97f682c2259eac3ecd1d052", size = 1902336, upload-time = "2026-07-03T03:50:05.839Z" }, + { url = "https://files.pythonhosted.org/packages/e3/b6/1db1a2f1164b82e5fd867e880dfd964516f1382fb7d35a916c0c2929fa55/numba_cuda-0.30.4-cp312-cp312-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:4f229a9314c98a6084e43c3a29ae1568a182bdcab7e87d8fc628a651ea2336ca", size = 1899900, upload-time = "2026-07-03T03:50:07.778Z" }, + { url = "https://files.pythonhosted.org/packages/6b/1e/d873bd274561eb1484be3bba59a2016c30b46edb9a0fa49014b2a5bc06f8/numba_cuda-0.30.4-cp312-cp312-win_amd64.whl", hash = "sha256:478ce786d562eb04e73087442967e0f49cbdc7f5d24c316eff3ea390e657ee38", size = 1879110, upload-time = "2026-07-03T03:50:09.493Z" }, + { url = "https://files.pythonhosted.org/packages/52/8f/dc09092e69115eca1b49a2edb1112b424fb2b105e2cb0b2d0497b7da09ec/numba_cuda-0.30.4-cp313-cp313-manylinux_2_24_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:c667006d2331b3dcdc98863cfe16d125bbc3852b9cbd5454cfd4d7fe866acce0", size = 1909491, upload-time = "2026-07-03T03:50:11.299Z" }, + { url = "https://files.pythonhosted.org/packages/5a/bd/8550fb0dd4ae4e8ef58ea6dc1465395f1f9bfa5220ef2ff857fb20a2aa52/numba_cuda-0.30.4-cp313-cp313-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:ed7f67afaf92dd9973c3985d8323ddd3f53a81aa9cf96b4f3275f3bbfb7182d3", size = 1907244, upload-time = "2026-07-03T03:50:13.082Z" }, + { url = "https://files.pythonhosted.org/packages/0a/ea/a2619a7c14688d0ed4fc11cf24bc71b66795619b7db9f45664dff79f5bd2/numba_cuda-0.30.4-cp313-cp313-win_amd64.whl", hash = "sha256:efebb025ef48475414b66c4f4973b11b8fdf6bd6eb8c73be32edcbd66503355e", size = 1879200, upload-time = "2026-07-03T03:50:14.984Z" }, + { url = "https://files.pythonhosted.org/packages/bc/35/3c2a05375b01159632ed90384a04a08725b83b49d2e511f44158c5ec2c43/numba_cuda-0.30.4-cp314-cp314-manylinux_2_24_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:b153d7f59e9a7834dc5866681b607317c72f0d49f6bbfed06eef82b396cbee84", size = 1878752, upload-time = "2026-07-03T03:50:17.092Z" }, + { url = "https://files.pythonhosted.org/packages/df/73/8934508d27efd9a930489059b980bc1e75336574be31bffb284a9e782fa7/numba_cuda-0.30.4-cp314-cp314-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:ec8e14ab7d4ecacaa1ff8d662efe5ab3c268e2f1cdad0330d28804b0d19a78c1", size = 1876463, upload-time = "2026-07-03T03:50:19.237Z" }, + { url = "https://files.pythonhosted.org/packages/3f/e3/613f6fc457f430775d124bea2e38033802fe97ab412a032191ac9b117d3a/numba_cuda-0.30.4-cp314-cp314-win_amd64.whl", hash = "sha256:a99227a6a7bc9f953a7abb61917d5ff767b9667f9c01e28d0b6b1def8f0fd991", size = 1884556, upload-time = "2026-07-03T03:50:21.483Z" }, +] + [[package]] name = "numpy" version = "2.4.6" @@ -550,6 +641,16 @@ wheels = [ { url = "https://files.pythonhosted.org/packages/43/bb/e1c71a4295b1b1d1393d50dbb4f2a36283c6859d9d3892e84f00ec5a91d5/numpy-2.4.6-cp314-cp314t-win_arm64.whl", hash = "sha256:0c9136e14ed34a9e343a31c533d78a9813a69a3148332bce5e9821cb2f996e66", size = 10565867, upload-time = "2026-05-18T23:36:47.114Z" }, ] +[[package]] +name = "nvidia-cuda-nvcc-cu12" +version = "12.8.93" +source = { registry = "https://pypi.org/simple" } +wheels = [ + { url = "https://files.pythonhosted.org/packages/86/0a/962e9c861541b08d62ee38a9e9776a46f75b979c75e7236fa7e09b024470/nvidia_cuda_nvcc_cu12-12.8.93-py3-none-manylinux2010_x86_64.manylinux_2_12_x86_64.whl", hash = "sha256:2d6dc36fb7cb5ac9c0b8825bc13d193c35487a315664007287d0126531238011", size = 40128098, upload-time = "2025-03-07T01:41:22.271Z" }, + { url = "https://files.pythonhosted.org/packages/df/26/38b9f7d19a1d4bc7bc5988aa07c8c7ab8fb20f5a2d92018fd037eaff4eb7/nvidia_cuda_nvcc_cu12-12.8.93-py3-none-manylinux2014_aarch64.manylinux_2_17_aarch64.whl", hash = "sha256:2b35ada240938d19398119ca596faaf626ba5e729511be9715842a29d112926d", size = 38021796, upload-time = "2025-03-07T01:41:01.162Z" }, + { url = "https://files.pythonhosted.org/packages/05/e9/0adfcf7b228413645b41b0fafaa0b93a71dfe8f23cb15591a0f7fe6c3f63/nvidia_cuda_nvcc_cu12-12.8.93-py3-none-win_amd64.whl", hash = "sha256:bb9dc7c291cd105d14f9b220bdcdf6ce6063c2d6ff01a505e358471530a8d226", size = 33424101, upload-time = "2025-03-07T01:51:40.555Z" }, +] + [[package]] name = "packaging" version = "26.2"