Skip to content

fix(integrations): state Numba's real pointer contract and enforce it - #1

Merged
egparedes merged 4 commits into
mainfrom
cuda-t2-suite-green
Jul 18, 2026
Merged

egparedes merged 4 commits into
mainfrom
cuda-t2-suite-green

Conversation

@egparedes

Copy link
Copy Markdown
Contributor

Lands the fix that came out of the first manual T2 (CUDA hardware) run of the
GPU suite, plus the packaging and documentation needed to reproduce that run.

What changed

  • src/devmm/integrations/numba.py — the EMM plugin hands Numba a
    ctypes.c_void_p device pointer. numba-cuda >= 0.30 builds a driver
    CUdeviceptr only inside an isinstance(pointer, ctypes.c_void_p) branch;
    anything else is stored raw and rejected downstream.
  • tests/_integration_fakes.py — FakeNumbaMemoryPointer now enforces that
    contract instead of accepting anything, so the T0 fake suite catches a
    regression off-hardware. Verified by reverting the plugin to c_uint64:
    three tests fail.
  • pyproject.toml / uv.lock — new gpu-test-cuda extra, marked
    python_full_version < '3.15' (numba-cuda publishes no cp315 wheel) and with
    torch deliberately left out, since the CUDA build comes from PyTorch's own
    index. No existing dependency drifted.
  • docs/testing.md, docs/adr/0003, CHANGELOG.md, Makefile — the
    install recipe, the hardware-run record against the ADR 0003 waiver terms, and
    a test-gpu-cuda target.

Verification

make verify green: 896 passed, 79 skipped. On hardware, the T2 suite ran
41 passed / 1 skipped (the skip is the pre-existing best-effort race canary).
No test was deleted, xfailed, or unhooked from its marker.

Reviewed over two rounds — 11 defects in the first pass, 5 nits in the second,
all addressed; the reviewer independently re-ran the gate and the revert
experiment.

Known follow-up (pre-existing on main, not introduced here)

tests/test_cuda_gpu.py:181 and tests/test_integrations_gpu.py:99,122 use
pytest.importorskip, so with DEVMM_GPU=cuda set a missing torch or numba
skips instead of failing. Now that torch is correctly out of the
gpu-test-cuda extra, step 2 of the install recipe is load-bearing, and a
partial install can report green with fewer tests than the 41 recorded in
ADR 0003. Once DEVMM_GPU is set, missing consumer libraries should error
rather than skip.

🤖 Generated with Claude Code

enriqueg and others added 3 commits July 18, 2026 00:54
…a >=0.30

numba-cuda >=0.30 coerces MemoryPointer's device pointer with int(), which a
ctypes.c_uint64 cannot satisfy; ctypes.c_void_p works on both the new and old
numba (older numba read .value, shared by both types).

Also green the CUDA (T2) hardware suite against rmm >=26.06, whose
allocation_counts is now a Statistics object rather than a dict, and codify
running the suite: a gpu-test-cuda extra, a make test-gpu-cuda target, a
docs/testing.md recipe with the libnvvm/nvjitlink alignment gotcha, and the
first hardware manual-run record appended to ADR 0003 (4x GH200, 41 passed).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The EMM plugin's comment described numba-cuda >= 0.30 as coercing the
device pointer with int(); it does a type test instead, converting only
inside an isinstance(pointer, ctypes.c_void_p) branch and storing any
other type raw for the driver to reject. Restate that at the definition
site, note that c_void_p(0).value is None but MRs must return non-null
pointers, and drop the copy of the rationale from the test.

FakeNumbaMemoryPointer now mirrors that contract — converting a c_void_p
to a driver-pointer double and refusing anything else — so the T0 suite
catches a regression to the wrong ctypes type off-hardware.

Also: drop torch from gpu-test-cuda (its CUDA builds live on PyTorch's
index) and mark the extra python_full_version < '3.15', where numba-cuda
would otherwise resolve the CUDA-13 stack against the cu12 nvcc pin;
resolve the ADR 0003 record against the Decision's release-blocking
clause; reorder the changelog groups; assert the rmm statistics fields
exist; widen the make help column.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Reviewer follow-ups on 0b92fb6, all wording and placement:

- memalloc's NULL note said the case "cannot arise"; mr_conformance is
  opt-in and the MR comes from the registry, so a third-party MR can
  still return 0. It must not arise, and the contract is what pins it.
- The gpu-test-cuda marker now leads with the reason it describes
  reality: numba-cuda publishes no cp315 wheel. Step 1 of the recipe
  says so too, since on 3.15+ the extra installs nothing silently.
- FakeNumbaMemoryPointer's refusal read as if every non-c_void_p fails
  upstream; a real driver pointer works there. Say so.
- The double's contract test moves to TestFakeContract: it exercises
  the fake, not the plugin.
- ADR 0003's Status records that the T2 half has run.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR tightens the devmm ↔ Numba EMM integration by documenting and enforcing Numba’s real device-pointer type contract (ctypes.c_void_p), and adds a reproducible CUDA (T2) on-hardware test recipe (extra + Makefile target + docs) captured in ADR/changelog.

Changes:

  • Update the Numba EMM plugin to pass device pointers as ctypes.c_void_p (the only type numba-cuda ≥ 0.30 converts to a driver pointer).
  • Strengthen the Numba integration test doubles to enforce that contract off-hardware, and add a focused test for the fake’s guarantee.
  • Add a gpu-test-cuda extra plus make test-gpu-cuda, and document/record a first T2 hardware run.

Reviewed changes

Copilot reviewed 9 out of 10 changed files in this pull request and generated 2 comments.

Show a summary per file
File Description
src/devmm/integrations/numba.py Passes ctypes.c_void_p(buffer.ptr) into numba_cuda.MemoryPointer and documents the contract rationale.
tests/_integration_fakes.py Enforces ctypes.c_void_p in FakeNumbaMemoryPointer and simulates conversion to a driver-pointer type.
tests/test_integrations_numba.py Updates assertions to validate the contract and adds a regression test ensuring the fake refuses non-c_void_p pointers.
tests/test_cuda_gpu.py Normalizes rmm statistics access across API versions via _allocation_counts.
pyproject.toml Adds gpu-test-cuda optional-dependency extra (CUDA-12-focused, <3.15-guarded).
uv.lock Locks new dependencies required by the CUDA hardware test extra (e.g., numba-cuda, cupy-cuda12x, nvidia-cuda-nvcc-cu12).
Makefile Adds test-gpu-cuda target to run the CUDA T2 suite with the documented environment constraints.
docs/testing.md Documents the CUDA hardware suite install + run recipe and the key Numba toolchain caveats.
docs/adr/0003-gpu-suite-waiver-for-0.1.0.md Appends a CUDA (T2) manual-run record; also edits earlier ADR text (see comments).
CHANGELOG.md Notes the new CUDA hardware test extra/target and the Numba pointer-contract fix.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +5 to +8
Accepted — signed off by the grAItools maintainer on 2026-07-17, explicitly
including the extension of the waiver to T2 (CUDA) beyond the p12 spec's
T3-only waiver clause.
T3-only waiver clause (T2 half exercised on hardware 2026-07-18, see the
manual-run record).

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 1a1c2f5. Reverted the Status header to its original wording; the appended manual-run record carries the update, as ADR 0003's own Consequences section requires ("an appended record ... rather than an edit to the decision").

Comment on lines +47 to +50
- The manual-run record (which tests ran, hardware, driver versions) is
appended to this file when hardware access happens; as of this writing no
manual GPU run has been performed.
manual GPU run has been performed. (One has since: see the CUDA (T2)
manual-run record at the end of this file.)

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 1a1c2f5. Reverted the Decision bullet. Worth noting the appended record itself is sanctioned by this ADR rather than being an exception to append-only: the bullet you flagged is the one that says the record "is appended to this file when hardware access happens". git diff main -- docs/adr/0003-*.md now shows additions only.

The Status header and a Decision bullet were edited to cross-reference the
manual-run record. ADR 0003 itself prescribes the opposite: the record is
"appended to this file" (Decision) "rather than an edit to the decision"
(Consequences), and docs/adr/README.md makes the historical record the point.
Revert both edits; the appended record already carries the update.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@egparedes
egparedes merged commit 1beb326 into main Jul 18, 2026
9 checks passed
@egparedes
egparedes deleted the cuda-t2-suite-green branch July 18, 2026 20:04
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants