Conversation
The Array returned by unsafe_wrap(Array, ::oneArray) does not keep the oneArray alive, which wasn't documented. Add a prominent warning, as the other back-ends do. Also allow wrapping arrays backed by host buffers, which are just as accessible from the host as shared buffers.
Scalar getindex and setindex! on shared and host buffers accessed the memory directly, without waiting for queued work. Reading the result of a reduction therefore returned stale data, e.g. sum() and maximum() of such arrays returned 0. Synchronize the current task's stream first.
Add unsafe_wrap(oneArray, ::Array) and unsafe_wrap(oneArray, ::Ptr, dims), which create a oneArray backed by existing host memory without copying it, so that kernels and broadcasts operate directly on Julia arrays. This uses the Level Zero ZE_extension_external_memmap_sysmem extension, which maps page-aligned system memory as a host allocation; the mapping has to be made resident before kernels can use it. Mappings cover whole pages and may not overlap, so wrappers of memory within an existing mapping share it, and memory whose pages only partially overlap an existing mapping is rejected with an ArgumentError. Mappings are never replaced while in use. Wrapping an Array keeps it alive. Wrappers freed by a finalizer are queued and released by a separate task, since releasing blocks until the device is done with the memory.
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #651 +/- ##
==========================================
+ Coverage 80.87% 81.03% +0.16%
==========================================
Files 56 56
Lines 4088 4214 +126
==========================================
+ Hits 3306 3415 +109
- Misses 782 799 +17 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Draft: this works, but I don't think it's ready to merge. See below. The first two commits are #650 and #649; only the last one is new here.
The goal is the same as for the other back-ends (JuliaGPU/CUDA.jl#3308, JuliaGPU/Metal.jl#991, JuliaGPU/OpenCL.jl#513, JuliaGPU/AMDGPU.jl#1116): let
unsafe_wrap(oneArray, a::Array)give the GPU direct access to an existing CPU array, without copying.On an Iris Xe that broadcast takes 57 ms, against 229 ms when copying through a regular
oneArrayand back (and 446 ms for plain Base).How it works
GPUs like the Iris Xe can't access arbitrary host memory (no shared system USM), but the driver supports
ZE_extension_external_memmap_sysmem, which imports an existing range of host memory as a host allocation at the same address. After making it resident, kernels can use it. The extension only works with whole pages, so we import the pages containing the array.Why this isn't ready
The extension doesn't allow importing ranges that overlap, but Julia arrays often share pages with their neighbours. This implementation keeps track of the imported ranges. Memory that falls entirely within an existing import can be wrapped, but memory whose pages only partially overlap an existing import can't, and throws an
ArgumentError. Whether a wrap succeeds therefore depends on how Julia happened to lay out memory, and on what else is wrapped at the time. Wrapping 100 arrays allocated back to back:Large arrays, which are the interesting ones to wrap, don't seem to be affected, and a failure is always an error rather than corruption. Still, an API that occasionally fails depending on memory layout doesn't feel right. For comparison, SYCL only uses this kind of host-memory import to speed up explicit copies (
sycl_ext_oneapi_copy_optimize), and simply declares overlapping ranges undefined behaviour. Intel's OpenMP runtime only lets kernels use host pointers on devices with shared system USM.That's probably the way forward: on devices that support shared system USM (recent GPUs with the Xe kernel driver), the GPU can access any host memory directly, like CUDA's HMM. Wrapping then needs no imports and has none of these limitations. I don't have such hardware to test on, so this PR only implements the import-based approach.
Tested on an Iris Xe: 2385 tests pass. The implementation was reviewed for lifetime issues: wrapped arrays stay alive for as long as any import covering them, and imports are released from a background task rather than from a finalizer.