Skip to content

Support wrapping host memory as a oneArray (draft) - #651

Draft
maleadt wants to merge 3 commits into
mainfrom
tb/unsafe_wrap
Draft

maleadt wants to merge 3 commits into
mainfrom
tb/unsafe_wrap

Conversation

@maleadt

@maleadt maleadt commented Sep 29, 2026

Copy link
Copy Markdown
Member

Draft: this works, but I don't think it's ready to merge. See below. The first two commits are #650 and #649; only the last one is new here.

The goal is the same as for the other back-ends (JuliaGPU/CUDA.jl#3308, JuliaGPU/Metal.jl#991, JuliaGPU/OpenCL.jl#513, JuliaGPU/AMDGPU.jl#1116): let unsafe_wrap(oneArray, a::Array) give the GPU direct access to an existing CPU array, without copying.

a = rand(Float32, 10^8)
b = unsafe_wrap(oneArray, a)
b .= sin.(b) .* 2          # runs on the GPU, directly on `a`'s memory
synchronize()

On an Iris Xe that broadcast takes 57 ms, against 229 ms when copying through a regular oneArray and back (and 446 ms for plain Base).

How it works

GPUs like the Iris Xe can't access arbitrary host memory (no shared system USM), but the driver supports ZE_extension_external_memmap_sysmem, which imports an existing range of host memory as a host allocation at the same address. After making it resident, kernels can use it. The extension only works with whole pages, so we import the pages containing the array.

Why this isn't ready

The extension doesn't allow importing ranges that overlap, but Julia arrays often share pages with their neighbours. This implementation keeps track of the imported ranges. Memory that falls entirely within an existing import can be wrapped, but memory whose pages only partially overlap an existing import can't, and throws an ArgumentError. Whether a wrap succeeds therefore depends on how Julia happened to lay out memory, and on what else is wrapped at the time. Wrapping 100 arrays allocated back to back:

array size wrapped
16 B 100/100
256 B 95/100
2 KB 83/100
4 KB 96/100
40 KB 99/100
4 MB 20/20

Large arrays, which are the interesting ones to wrap, don't seem to be affected, and a failure is always an error rather than corruption. Still, an API that occasionally fails depending on memory layout doesn't feel right. For comparison, SYCL only uses this kind of host-memory import to speed up explicit copies (sycl_ext_oneapi_copy_optimize), and simply declares overlapping ranges undefined behaviour. Intel's OpenMP runtime only lets kernels use host pointers on devices with shared system USM.

That's probably the way forward: on devices that support shared system USM (recent GPUs with the Xe kernel driver), the GPU can access any host memory directly, like CUDA's HMM. Wrapping then needs no imports and has none of these limitations. I don't have such hardware to test on, so this PR only implements the import-based approach.

Tested on an Iris Xe: 2385 tests pass. The implementation was reviewed for lifetime issues: wrapped arrays stay alive for as long as any import covering them, and imports are released from a background task rather than from a finalizer.

maleadt and others added 3 commits September 29, 2026 21:01
The Array returned by unsafe_wrap(Array, ::oneArray) does not keep the oneArray alive,
which wasn't documented. Add a prominent warning, as the other back-ends do. Also allow
wrapping arrays backed by host buffers, which are just as accessible from the host as
shared buffers.
Scalar getindex and setindex! on shared and host buffers accessed the memory directly,
without waiting for queued work. Reading the result of a reduction therefore returned
stale data, e.g. sum() and maximum() of such arrays returned 0. Synchronize the current
task's stream first.
Add unsafe_wrap(oneArray, ::Array) and unsafe_wrap(oneArray, ::Ptr, dims), which
create a oneArray backed by existing host memory without copying it, so that kernels
and broadcasts operate directly on Julia arrays. This uses the Level Zero
ZE_extension_external_memmap_sysmem extension, which maps page-aligned system memory
as a host allocation; the mapping has to be made resident before kernels can use it.

Mappings cover whole pages and may not overlap, so wrappers of memory within an
existing mapping share it, and memory whose pages only partially overlap an existing
mapping is rejected with an ArgumentError. Mappings are never replaced while in use.
Wrapping an Array keeps it alive. Wrappers freed by a finalizer are queued and released
by a separate task, since releasing blocks until the device is done with the memory.
@codecov

codecov Bot commented Sep 30, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 92.12598% with 10 lines in your changes missing coverage. Please review.
✅ Project coverage is 81.03%. Comparing base (85dd4cf) to head (24ba399).

Files with missing lines Patch % Lines
src/array.jl 92.37% 9 Missing ⚠️
lib/level-zero/memory.jl 88.88% 1 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main     #651      +/-   ##
==========================================
+ Coverage   80.87%   81.03%   +0.16%     
==========================================
  Files          56       56              
  Lines        4088     4214     +126     
==========================================
+ Hits         3306     3415     +109     
- Misses        782      799      +17     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant