fix(integrations): state Numba's real pointer contract and enforce it - #1
Conversation
…a >=0.30 numba-cuda >=0.30 coerces MemoryPointer's device pointer with int(), which a ctypes.c_uint64 cannot satisfy; ctypes.c_void_p works on both the new and old numba (older numba read .value, shared by both types). Also green the CUDA (T2) hardware suite against rmm >=26.06, whose allocation_counts is now a Statistics object rather than a dict, and codify running the suite: a gpu-test-cuda extra, a make test-gpu-cuda target, a docs/testing.md recipe with the libnvvm/nvjitlink alignment gotcha, and the first hardware manual-run record appended to ADR 0003 (4x GH200, 41 passed). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The EMM plugin's comment described numba-cuda >= 0.30 as coercing the device pointer with int(); it does a type test instead, converting only inside an isinstance(pointer, ctypes.c_void_p) branch and storing any other type raw for the driver to reject. Restate that at the definition site, note that c_void_p(0).value is None but MRs must return non-null pointers, and drop the copy of the rationale from the test. FakeNumbaMemoryPointer now mirrors that contract — converting a c_void_p to a driver-pointer double and refusing anything else — so the T0 suite catches a regression to the wrong ctypes type off-hardware. Also: drop torch from gpu-test-cuda (its CUDA builds live on PyTorch's index) and mark the extra python_full_version < '3.15', where numba-cuda would otherwise resolve the CUDA-13 stack against the cu12 nvcc pin; resolve the ADR 0003 record against the Decision's release-blocking clause; reorder the changelog groups; assert the rmm statistics fields exist; widen the make help column. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Reviewer follow-ups on 0b92fb6, all wording and placement: - memalloc's NULL note said the case "cannot arise"; mr_conformance is opt-in and the MR comes from the registry, so a third-party MR can still return 0. It must not arise, and the contract is what pins it. - The gpu-test-cuda marker now leads with the reason it describes reality: numba-cuda publishes no cp315 wheel. Step 1 of the recipe says so too, since on 3.15+ the extra installs nothing silently. - FakeNumbaMemoryPointer's refusal read as if every non-c_void_p fails upstream; a real driver pointer works there. Say so. - The double's contract test moves to TestFakeContract: it exercises the fake, not the plugin. - ADR 0003's Status records that the T2 half has run. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
Pull request overview
This PR tightens the devmm ↔ Numba EMM integration by documenting and enforcing Numba’s real device-pointer type contract (ctypes.c_void_p), and adds a reproducible CUDA (T2) on-hardware test recipe (extra + Makefile target + docs) captured in ADR/changelog.
Changes:
- Update the Numba EMM plugin to pass device pointers as
ctypes.c_void_p(the only type numba-cuda ≥ 0.30 converts to a driver pointer). - Strengthen the Numba integration test doubles to enforce that contract off-hardware, and add a focused test for the fake’s guarantee.
- Add a
gpu-test-cudaextra plusmake test-gpu-cuda, and document/record a first T2 hardware run.
Reviewed changes
Copilot reviewed 9 out of 10 changed files in this pull request and generated 2 comments.
Show a summary per file
| File | Description |
|---|---|
src/devmm/integrations/numba.py |
Passes ctypes.c_void_p(buffer.ptr) into numba_cuda.MemoryPointer and documents the contract rationale. |
tests/_integration_fakes.py |
Enforces ctypes.c_void_p in FakeNumbaMemoryPointer and simulates conversion to a driver-pointer type. |
tests/test_integrations_numba.py |
Updates assertions to validate the contract and adds a regression test ensuring the fake refuses non-c_void_p pointers. |
tests/test_cuda_gpu.py |
Normalizes rmm statistics access across API versions via _allocation_counts. |
pyproject.toml |
Adds gpu-test-cuda optional-dependency extra (CUDA-12-focused, <3.15-guarded). |
uv.lock |
Locks new dependencies required by the CUDA hardware test extra (e.g., numba-cuda, cupy-cuda12x, nvidia-cuda-nvcc-cu12). |
Makefile |
Adds test-gpu-cuda target to run the CUDA T2 suite with the documented environment constraints. |
docs/testing.md |
Documents the CUDA hardware suite install + run recipe and the key Numba toolchain caveats. |
docs/adr/0003-gpu-suite-waiver-for-0.1.0.md |
Appends a CUDA (T2) manual-run record; also edits earlier ADR text (see comments). |
CHANGELOG.md |
Notes the new CUDA hardware test extra/target and the Numba pointer-contract fix. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| Accepted — signed off by the grAItools maintainer on 2026-07-17, explicitly | ||
| including the extension of the waiver to T2 (CUDA) beyond the p12 spec's | ||
| T3-only waiver clause. | ||
| T3-only waiver clause (T2 half exercised on hardware 2026-07-18, see the | ||
| manual-run record). |
There was a problem hiding this comment.
Fixed in 1a1c2f5. Reverted the Status header to its original wording; the appended manual-run record carries the update, as ADR 0003's own Consequences section requires ("an appended record ... rather than an edit to the decision").
| - The manual-run record (which tests ran, hardware, driver versions) is | ||
| appended to this file when hardware access happens; as of this writing no | ||
| manual GPU run has been performed. | ||
| manual GPU run has been performed. (One has since: see the CUDA (T2) | ||
| manual-run record at the end of this file.) |
There was a problem hiding this comment.
Fixed in 1a1c2f5. Reverted the Decision bullet. Worth noting the appended record itself is sanctioned by this ADR rather than being an exception to append-only: the bullet you flagged is the one that says the record "is appended to this file when hardware access happens". git diff main -- docs/adr/0003-*.md now shows additions only.
The Status header and a Decision bullet were edited to cross-reference the manual-run record. ADR 0003 itself prescribes the opposite: the record is "appended to this file" (Decision) "rather than an edit to the decision" (Consequences), and docs/adr/README.md makes the historical record the point. Revert both edits; the appended record already carries the update. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Lands the fix that came out of the first manual T2 (CUDA hardware) run of the
GPU suite, plus the packaging and documentation needed to reproduce that run.
What changed
src/devmm/integrations/numba.py— the EMM plugin hands Numba actypes.c_void_pdevice pointer. numba-cuda >= 0.30 builds a driverCUdeviceptronly inside anisinstance(pointer, ctypes.c_void_p)branch;anything else is stored raw and rejected downstream.
tests/_integration_fakes.py—FakeNumbaMemoryPointernow enforces thatcontract instead of accepting anything, so the T0 fake suite catches a
regression off-hardware. Verified by reverting the plugin to
c_uint64:three tests fail.
pyproject.toml/uv.lock— newgpu-test-cudaextra, markedpython_full_version < '3.15'(numba-cuda publishes no cp315 wheel) and withtorchdeliberately left out, since the CUDA build comes from PyTorch's ownindex. No existing dependency drifted.
docs/testing.md,docs/adr/0003,CHANGELOG.md,Makefile— theinstall recipe, the hardware-run record against the ADR 0003 waiver terms, and
a
test-gpu-cudatarget.Verification
make verifygreen: 896 passed, 79 skipped. On hardware, the T2 suite ran41 passed / 1 skipped (the skip is the pre-existing best-effort race canary).
No test was deleted,
xfailed, or unhooked from its marker.Reviewed over two rounds — 11 defects in the first pass, 5 nits in the second,
all addressed; the reviewer independently re-ran the gate and the revert
experiment.
Known follow-up (pre-existing on
main, not introduced here)tests/test_cuda_gpu.py:181andtests/test_integrations_gpu.py:99,122usepytest.importorskip, so withDEVMM_GPU=cudaset a missing torch or numbaskips instead of failing. Now that torch is correctly out of the
gpu-test-cudaextra, step 2 of the install recipe is load-bearing, and apartial install can report green with fewer tests than the 41 recorded in
ADR 0003. Once
DEVMM_GPUis set, missing consumer libraries should errorrather than skip.
🤖 Generated with Claude Code