From 41bbf766e4a37d5844438087e1be918e8b341561 Mon Sep 17 00:00:00 2001 From: Tim Besard Date: Wed, 7 Oct 2026 17:41:31 +0200 Subject: [PATCH 1/6] Fix two broken sentences in the @spawn docs A sentence about the device of a new task was cut off, and the note on exceptions lacked punctuation and ran into the next paragraph. --- docs/src/implementations.md | 4 ++-- src/spawn.jl | 3 ++- 2 files changed, 4 insertions(+), 3 deletions(-) diff --git a/docs/src/implementations.md b/docs/src/implementations.md index 273b7beab..234fa95d8 100644 --- a/docs/src/implementations.md +++ b/docs/src/implementations.md @@ -41,8 +41,8 @@ which backends can support with two optional functions: A new Julia task does not inherit the device of the task that spawned it: backends keep the active device in task-local state, which Julia does not copy into a child task, so the task -Backends with more than one device -**must** implement the device interface ([`device`](@ref KernelAbstractions.device), +starts on the backend's default device. Backends with more than one device **must** +implement the device interface ([`device`](@ref KernelAbstractions.device), [`ndevices`](@ref KernelAbstractions.ndevices), [`device!`](@ref KernelAbstractions.device!)) for `@spawn` to run on the right device. diff --git a/src/spawn.jl b/src/spawn.jl index 12a5b1904..816ed8b1c 100644 --- a/src/spawn.jl +++ b/src/spawn.jl @@ -62,7 +62,8 @@ more than one device implement this with a cross-device [`wait_event`](@ref KernelAbstractions.wait_event) yourself. !!! note - If `expr` throws the state of the device and the internal queue is unspecified. + If `expr` throws, the state of the device and of the task's queue is unspecified. + Backend authors: see the [notes for backend implementations](@ref implementations_notes) for the protocol behind these guarantees, and for how to support it without a full [`synchronize`](@ref). From b9510637abfb46f2132573f68052250298cd712f Mon Sep 17 00:00:00 2001 From: Tim Besard Date: Wed, 7 Oct 2026 17:41:40 +0200 Subject: [PATCH 2/6] Document that @spawn doesn't wait for the parent's work on the host The spawned task's work is ordered after its parent's on the device, but code in the task that passes the parent's results to something outside the backend, e.g., an MPI call on a GPU buffer, has to synchronize first. Some backends currently do so implicitly when the buffer's pointer is taken, which portable code can't rely on. Also say that the child isn't ordered against the parent's later work, rather than that it can't rely on that work's data. --- src/spawn.jl | 15 +++++++++++++-- 1 file changed, 13 insertions(+), 2 deletions(-) diff --git a/src/spawn.jl b/src/spawn.jl index 816ed8b1c..ee1da5001 100644 --- a/src/spawn.jl +++ b/src/spawn.jl @@ -50,8 +50,19 @@ more than one device implement this with a cross-device [`wait_event`](@ref KernelAbstractions.wait_event). !!! note - `expr` should not rely on data that the spawning task queues *after* `@spawn` returns. - Order later work by waiting on the task, or by spawning again. + `expr` is not ordered against work that the spawning task queues *after* `@spawn` + returns. Order conflicting uses of shared data by waiting on the task, or by spawning + again. + +!!! note + The ordering is between work queued on `backend`; the host need not wait for the + spawning task's work. Before `expr` passes that work's results to something that + doesn't queue on `backend`, e.g., an MPI call on a GPU buffer, call + `synchronize(backend)` in `expr`, which also waits for the spawning task's work since + the task's queue is ordered after it. + Some backends synchronize implicitly when such code takes the buffer's pointer, but + portable code should not rely on that. The trailing `synchronize` of `@spawn` does not + complete asynchronous operations outside the backend, like `MPI.Isend`. !!! note Prefer `device=` over calling [`device!`](@ref KernelAbstractions.device!) inside From 9ade488fcb658ca2d51bcb918854716cfa6d7c92 Mon Sep 17 00:00:00 2001 From: Tim Besard Date: Wed, 7 Oct 2026 17:41:40 +0200 Subject: [PATCH 3/6] Ask backends not to wait for work that wait_event already ordered Backends that track which queue last used an array make another queue wait before using it. If that wait ignores the order set up by wait_event, or by a completed synchronize, the spawned task's first use of an array shared with its parent waits for everything the parent has queued, including work queued after @spawn (#817). CUDA.jl, AMDGPU.jl, Metal.jl and OpenCL.jl all track arrays this way. Recommend respecting that order, and likewise for what synchronize waits for. Like task-local queues, this is about overlap rather than correctness, so it is a recommendation. Waits for memory accessibility or lifetime are not affected. --- docs/src/implementations.md | 17 +++++++++++++++++ lib/KernelInterface/src/host.jl | 14 ++++++++++---- 2 files changed, 27 insertions(+), 4 deletions(-) diff --git a/docs/src/implementations.md b/docs/src/implementations.md index 234fa95d8..8f4ca72c8 100644 --- a/docs/src/implementations.md +++ b/docs/src/implementations.md @@ -39,6 +39,23 @@ which backends can support with two optional functions: `wait(task)` in any other task implies that all work queued by the spawned task has completed. +Backends that track which queue last used an array, and wait for that queue, on the host or +on the device, before using the array on another queue, **should** respect the points up to +which the current queue is already ordered after the previous one: an event recorded on the +previous queue that the current queue waited for, or a [`synchronize`](@ref) of the +previous queue that returned. If the array's last use precedes such a point, using it on +the current queue should not wait, on the host or on the device, for work queued on the +previous queue after that point. Waits needed to make memory accessible, e.g., from another +device, or to keep it alive are not affected, and uses through a pointer taken before such a +point may still be synchronized conservatively. + +For the same reason, `synchronize` **should not** wait for work on other queues that the +current queue is not ordered after. + +A backend that ignores this is still correct, but the spawned task then waits for work the +parent queued after `@spawn`, and the two tasks' work doesn't overlap. Following it doesn't +guarantee overlap either; it only rules out these waits. + A new Julia task does not inherit the device of the task that spawned it: backends keep the active device in task-local state, which Julia does not copy into a child task, so the task starts on the backend's default device. Backends with more than one device **must** diff --git a/lib/KernelInterface/src/host.jl b/lib/KernelInterface/src/host.jl index b769783cd..716176184 100644 --- a/lib/KernelInterface/src/host.jl +++ b/lib/KernelInterface/src/host.jl @@ -38,7 +38,8 @@ completed. !!! note Backend implementations **must** implement this function, and it **must** be cooperative: it may not block inside a driver call, but has to yield to the Julia - scheduler while waiting. See the + scheduler while waiting. It **should not** wait for other tasks' work that the calling + task's queue was not ordered after. See the [notes for backend implementations](@ref implementations_notes) for why. """ function synchronize end @@ -48,7 +49,9 @@ function synchronize end Capture the work the calling task has queued on `backend`'s currently active device so far, and return a handle that [`wait_event`](@ref) can use to order later work after it, -either from another task or from the same task after switching devices. +either from another task or from the same task after switching devices. Work queued after +`record_event` returns is not captured, and recording need not wait for the captured work +to complete. The handle is only meaningful for the pair `record_event`/`wait_event`; do not use it for anything else. @@ -69,7 +72,8 @@ end wait_event(backend::Backend, event) Order the work the calling task subsequently queues on `backend`'s currently active device -after the work captured by `event`, which was returned by [`record_event`](@ref). +after the work captured by `event`, which was returned by [`record_event`](@ref). This +orders work on the device; it need not wait for the captured work on the host. The dependency is queue-ordered rather than task-ordered: it applies to the device that is active when `wait_event` is called, and a later [`device!`](@ref) leaves the newly selected @@ -85,7 +89,9 @@ wait_event(backend, event) # device 2 now waits for that work `wait_event(::Backend, ::Nothing)` is a no-op, matching the default `record_event`. A backend that implements [`record_event`](@ref) **must** implement this for the event type it returns, either by enqueuing a dependency on the current task's queue, or by - waiting cooperatively as [`synchronize`](@ref) does. A backend with more than one + waiting cooperatively as [`synchronize`](@ref) does. A backend that tracks which + queue last used an array **should** take this ordering into account when the array moves + to the waiting queue. A backend with more than one device **must** also accept an `event` that was recorded on a different device, by enqueuing the cross-device dependency if the driver supports one (CUDA's `cuStreamWaitEvent` does) and by waiting cooperatively otherwise. See the From 4403553592c8691fd4bd693816673cb4dc4648b6 Mon Sep 17 00:00:00 2001 From: Tim Besard Date: Wed, 7 Oct 2026 18:41:40 +0200 Subject: [PATCH 4/6] Tighten the @spawn and event docs Device selection was explained three times and the implementation requirements of record_event and wait_event were repeated on the notes page. Keep each fact in one place: the user-facing behavior in the @spawn docstring, the correctness requirements in the KernelInterface docstrings, and the protocol and the performance recommendations on the notes page. Also drop the note on exceptions, which the guarantees at the top already cover, and stop claiming that switching devices leaves the new queue unordered, since it may already be ordered by other means. --- docs/src/implementations.md | 94 ++++++++++++--------------------- lib/KernelInterface/src/host.jl | 50 ++++++++---------- src/spawn.jl | 58 ++++++++------------ 3 files changed, 78 insertions(+), 124 deletions(-) diff --git a/docs/src/implementations.md b/docs/src/implementations.md index 8f4ca72c8..5349d7efc 100644 --- a/docs/src/implementations.md +++ b/docs/src/implementations.md @@ -16,66 +16,40 @@ thread instead of letting independent tasks run concurrently. ## Task-local queues and `KernelAbstractions.@spawn` -Backends should give each Julia task its own queue/stream, so that kernels -launched from different tasks can execute concurrently. This implies that work queued -from two tasks is not ordered with respect to each other. - -[`KernelAbstractions.@spawn`](@ref) hides this from users by following a fixed protocol, -which backends can support with two optional functions: - -- Before the new task is created, the spawning task calls - [`record_event`](@ref KernelAbstractions.record_event) on the backend. The default - implementation is a full [`synchronize`](@ref) returning `nothing`, which is always - correct. A backend with task-local queues **may** instead record an event on the - current task's queue and return it, so that the spawning task does not have to wait. -- The new task selects its device with [`device!`](@ref KernelAbstractions.device!) — the - spawning task's, or the one the user asked for with `@spawn backend device=id` — and then - calls [`wait_event`](@ref KernelAbstractions.wait_event) with the recorded handle. The - order matters: `wait_event` makes the queue of the *currently active* device wait, so the - device has to be selected first. A backend that overrides `record_event` **must** - implement `wait_event` for its event type, typically by making the current task's queue - wait on the event. -- After the user's code returns, the new task calls [`synchronize`](@ref), so that - `wait(task)` in any other task implies that all work queued by the spawned task has - completed. - -Backends that track which queue last used an array, and wait for that queue, on the host or -on the device, before using the array on another queue, **should** respect the points up to -which the current queue is already ordered after the previous one: an event recorded on the -previous queue that the current queue waited for, or a [`synchronize`](@ref) of the -previous queue that returned. If the array's last use precedes such a point, using it on -the current queue should not wait, on the host or on the device, for work queued on the -previous queue after that point. Waits needed to make memory accessible, e.g., from another -device, or to keep it alive are not affected, and uses through a pointer taken before such a -point may still be synchronized conservatively. - -For the same reason, `synchronize` **should not** wait for work on other queues that the -current queue is not ordered after. - -A backend that ignores this is still correct, but the spawned task then waits for work the -parent queued after `@spawn`, and the two tasks' work doesn't overlap. Following it doesn't -guarantee overlap either; it only rules out these waits. - -A new Julia task does not inherit the device of the task that spawned it: backends keep the -active device in task-local state, which Julia does not copy into a child task, so the task -starts on the backend's default device. Backends with more than one device **must** -implement the device interface ([`device`](@ref KernelAbstractions.device), -[`ndevices`](@ref KernelAbstractions.ndevices), [`device!`](@ref KernelAbstractions.device!)) -for `@spawn` to run on the right device. - -`@spawn backend device=id` records the event on the spawning task's device but waits on -`id`, so a multi-device backend **must** accept an event recorded on a device other than the -one active in `wait_event`. A backend whose driver cannot **must** fall back -to waiting cooperatively, as [`synchronize`](@ref) does. - -Because `device!` selects the queue that `wait_event` acts on, the same two functions are -what lets users order work across a device switch they make themselves: - -```julia -event = KernelAbstractions.record_event(backend) -KernelAbstractions.device!(backend, 2) -KernelAbstractions.wait_event(backend, event) -``` +Backends **should** give each Julia task its own queue, so that work from different tasks +can execute concurrently. Separate queues do not by themselves order work. + +[`KernelAbstractions.@spawn`](@ref) orders it with this protocol: + +1. The spawning task calls [`record_event`](@ref KernelAbstractions.record_event). +2. The new task selects its device with [`device!`](@ref KernelAbstractions.device!), then + calls [`wait_event`](@ref KernelAbstractions.wait_event) before running the user's code. +3. If that code returns normally, the new task calls [`synchronize`](@ref), so that a + successful `wait(task)` implies its queued work has completed. + +The default `record_event` synchronizes and returns `nothing`, for which `wait_event` does +nothing. Backends can return an event instead, so that the spawning task doesn't wait; see +the docstrings of both functions for what that requires. Backends with more than one device +**must** implement [`device`](@ref KernelAbstractions.device), +[`ndevices`](@ref KernelAbstractions.ndevices) and +[`device!`](@ref KernelAbstractions.device!), and `wait_event` **must** accept an event +recorded on another device, since `@spawn backend device=id` records on the spawning task's +device. + +Backends that track which queue last used an array, and wait for that queue before using the +array on another one, **should** skip that wait when the current queue is already ordered +after the array's last use: through an event recorded on the previous queue after that use +and waited for by the current queue, or through a [`synchronize`](@ref) of the previous +queue that completed that use. In particular, they should not wait, on the host or on the +device, for work queued on the previous queue after that event or synchronization. Waits +needed to make memory accessible or to keep it alive still apply, and uses through a pointer +taken before that event or synchronization may be synchronized conservatively. Likewise, +`synchronize` **should not** wait for work on other queues that the current queue is not +ordered after. + +Otherwise, a spawned task's first use of an array shared with its parent waits for work the +parent queued after `@spawn`, reducing overlap. Neither recommendation guarantees that the +work of different tasks runs concurrently. ## Moving data with `adapt` diff --git a/lib/KernelInterface/src/host.jl b/lib/KernelInterface/src/host.jl index 716176184..16151fe2b 100644 --- a/lib/KernelInterface/src/host.jl +++ b/lib/KernelInterface/src/host.jl @@ -36,25 +36,19 @@ Block the calling task until all work it has queued on the active device of `bac completed. !!! note - Backend implementations **must** implement this function, and it **must** be - cooperative: it may not block inside a driver call, but has to yield to the Julia - scheduler while waiting. It **should not** wait for other tasks' work that the calling - task's queue was not ordered after. See the - [notes for backend implementations](@ref implementations_notes) for why. + Backend implementations **must** implement this function cooperatively, yielding to + the Julia scheduler while waiting rather than blocking inside a driver call. See the + [notes for backend implementations](@ref implementations_notes) for why, and for what + it should not wait for. """ function synchronize end """ record_event(backend::Backend) -Capture the work the calling task has queued on `backend`'s currently active device so -far, and return a handle that [`wait_event`](@ref) can use to order later work after it, -either from another task or from the same task after switching devices. Work queued after -`record_event` returns is not captured, and recording need not wait for the captured work -to complete. - -The handle is only meaningful for the pair `record_event`/`wait_event`; do not use it for -anything else. +Capture the work the calling task has queued on `backend`'s active device so far, and +return a handle for [`wait_event`](@ref). Work queued later is not captured, and recording +need not wait for the captured work to complete. The handle is only meant for `wait_event`. !!! note The default implementation calls [`synchronize`](@ref) and returns `nothing`. @@ -71,31 +65,29 @@ end """ wait_event(backend::Backend, event) -Order the work the calling task subsequently queues on `backend`'s currently active device -after the work captured by `event`, which was returned by [`record_event`](@ref). This -orders work on the device; it need not wait for the captured work on the host. +Order the work the calling task subsequently queues on `backend`'s active device after the +work captured by `event`, which was returned by [`record_event`](@ref). Returning does not +mean that the captured work has completed. -The dependency is queue-ordered rather than task-ordered: it applies to the device that is -active when `wait_event` is called, and a later [`device!`](@ref) leaves the newly selected -device unordered with respect to `event`. Select the device first and wait afterwards: +The wait applies to the queue of the device that is active when `wait_event` is called; +switching devices adds no ordering. To order work across a device switch, select the device +first and wait afterwards: ```julia event = record_event(backend) # captures work on the current device device!(backend, 2) -wait_event(backend, event) # device 2 now waits for that work +wait_event(backend, event) # orders this task's work on device 2 after it ``` !!! note `wait_event(::Backend, ::Nothing)` is a no-op, matching the default `record_event`. - A backend that implements [`record_event`](@ref) **must** implement this for the event - type it returns, either by enqueuing a dependency on the current task's queue, or by - waiting cooperatively as [`synchronize`](@ref) does. A backend that tracks which - queue last used an array **should** take this ordering into account when the array moves - to the waiting queue. A backend with more than one - device **must** also accept an `event` that was recorded on a different device, by - enqueuing the cross-device dependency if the driver supports one (CUDA's - `cuStreamWaitEvent` does) and by waiting cooperatively otherwise. See the - [notes for backend implementations](@ref implementations_notes). + A backend that returns another event type **must** implement `wait_event` for it, + either by adding a dependency to the current task's queue or by waiting cooperatively + as [`synchronize`](@ref) does. A backend with more than one device **must** accept an + event recorded on another device, waiting cooperatively if the driver cannot add a + cross-device dependency. See the + [notes for backend implementations](@ref implementations_notes) for how this ordering + should interact with implicit synchronization. """ wait_event(::Backend, ::Nothing) = nothing diff --git a/src/spawn.jl b/src/spawn.jl index ee1da5001..e5d9c793e 100644 --- a/src/spawn.jl +++ b/src/spawn.jl @@ -8,10 +8,10 @@ place of `Threads.@spawn` to launch kernels from a task. It guarantees that that argument is given; - the work the task queues on `backend` runs after the work the spawning task had queued on `backend` before calling `@spawn`; -- once `wait(task)` or `fetch(task)` returns, all work the task queued on `backend` has - completed, so its results may be used from any task. `fetch(task)` returns the value of - `expr`. If `expr` throws, the task is not synchronized: its queued work may still be - running when the exception surfaces. +- once `wait(task)` or `fetch(task)` returns successfully, all work the task queued on + `backend` has completed, so its results may be used from any task. `fetch(task)` + returns the value of `expr`. If `expr` throws, the task is not synchronized: its queued + work may still be running when the exception surfaces. Everything else works as for `Threads.@spawn`: the optional `threadpool` argument (`:default` or `:interactive`) is forwarded, `\$x` captures the value of `x` at spawn time, @@ -32,11 +32,9 @@ fetch(task) == 4 * length(A) # Choosing the device -Backends keep the active device in task-local state, and Julia does not copy that state -into a child task. A task started with plain `Threads.@spawn` therefore runs on the -backend's *default* device, whichever device the spawning task was using. `@spawn` selects -the device explicitly instead: by default the one active in the spawning task, or the one -named by `device`, a 1-based index into `1:ndevices(backend)`: +A task started with plain `Threads.@spawn` runs on the backend's default device, not on +the device of the task that started it. `@spawn` selects the spawning task's device, or the +one given by `device`, an index into `1:ndevices(backend)`: ```julia task = KernelAbstractions.@spawn backend device=2 begin @@ -44,40 +42,30 @@ task = KernelAbstractions.@spawn backend device=2 begin end ``` -The ordering guarantee holds across that switch: the task's work on `device` is still -ordered after the work the spawning task had queued on *its* device. Backends that support -more than one device implement this with a cross-device -[`wait_event`](@ref KernelAbstractions.wait_event). +The task's work on that device still runs after the spawning task's earlier work on its own +device. !!! note - `expr` is not ordered against work that the spawning task queues *after* `@spawn` - returns. Order conflicting uses of shared data by waiting on the task, or by spawning - again. + `@spawn` does not order the task's work against work the spawning task queues + afterwards. Order conflicting uses of shared data by waiting for the task, or by spawning a new + task after that work. !!! note - The ordering is between work queued on `backend`; the host need not wait for the - spawning task's work. Before `expr` passes that work's results to something that - doesn't queue on `backend`, e.g., an MPI call on a GPU buffer, call - `synchronize(backend)` in `expr`, which also waits for the spawning task's work since - the task's queue is ordered after it. - Some backends synchronize implicitly when such code takes the buffer's pointer, but - portable code should not rely on that. The trailing `synchronize` of `@spawn` does not - complete asynchronous operations outside the backend, like `MPI.Isend`. + Queued work is ordered, but the spawning task's earlier work need not have completed + when `expr` starts. Before passing that work's results to a consumer outside that + ordering, e.g., an MPI call on a GPU buffer, call `synchronize(backend)` in `expr`, + which also waits for that work. Some backends synchronize implicitly when the buffer's + pointer is taken, but portable code should not rely on that. `@spawn` does not wait for + asynchronous operations outside the backend, like `MPI.Isend`. !!! note - Prefer `device=` over calling [`device!`](@ref KernelAbstractions.device!) inside - `expr`. A `device!` in the body carries no ordering of its own, so work queued after it - is ordered neither against the spawning task nor against what the body queued before - the switch; you would have to bracket it with - [`record_event`](@ref KernelAbstractions.record_event) and - [`wait_event`](@ref KernelAbstractions.wait_event) yourself. - -!!! note - If `expr` throws, the state of the device and of the task's queue is unspecified. + Prefer `device=` over calling [`device!`](@ref KernelAbstractions.device!) in `expr`. + To keep work ordered across a manual switch, call + [`record_event`](@ref KernelAbstractions.record_event) before it and + [`wait_event`](@ref KernelAbstractions.wait_event) after it, as shown for `wait_event`. Backend authors: see the [notes for backend implementations](@ref implementations_notes) -for the protocol behind these guarantees, and for how to support it without a full -[`synchronize`](@ref). +for the protocol behind these guarantees. """ macro spawn(args...) usage = "@spawn expects `@spawn [threadpool] backend [device=id] expr`" From 8e6f074bb1382fd0bb26a96f419f9c2302e0e431 Mon Sep 17 00:00:00 2001 From: Tim Besard Date: Thu, 8 Oct 2026 08:51:04 +0200 Subject: [PATCH 5/6] Require that synchronize not wait for unrelated queues A synchronize that also waits for work on other queues, e.g. work another task queued, costs overlap without making any program more correct. All GPU backends except OpenCL.jl already wait only for the current queue (OpenCL.jl waits for other queues to check its shared exception mailbox, JuliaGPU/OpenCL.jl#549), so make this a requirement in the synchronize docstring rather than a recommendation. Also state the purpose of the ownership recommendation as a goal instead of describing what happens without it. --- docs/src/implementations.md | 11 ++++------- lib/KernelInterface/src/host.jl | 7 ++++--- 2 files changed, 8 insertions(+), 10 deletions(-) diff --git a/docs/src/implementations.md b/docs/src/implementations.md index 5349d7efc..30f188e4a 100644 --- a/docs/src/implementations.md +++ b/docs/src/implementations.md @@ -43,13 +43,10 @@ and waited for by the current queue, or through a [`synchronize`](@ref) of the p queue that completed that use. In particular, they should not wait, on the host or on the device, for work queued on the previous queue after that event or synchronization. Waits needed to make memory accessible or to keep it alive still apply, and uses through a pointer -taken before that event or synchronization may be synchronized conservatively. Likewise, -`synchronize` **should not** wait for work on other queues that the current queue is not -ordered after. - -Otherwise, a spawned task's first use of an array shared with its parent waits for work the -parent queued after `@spawn`, reducing overlap. Neither recommendation guarantees that the -work of different tasks runs concurrently. +taken before that event or synchronization may be synchronized conservatively. This lets a +spawned task use arrays it shares with its parent while the parent keeps queuing work, which +is also why [`synchronize`](@ref) **must not** wait for work on other queues that the +current queue is not ordered after. ## Moving data with `adapt` diff --git a/lib/KernelInterface/src/host.jl b/lib/KernelInterface/src/host.jl index 16151fe2b..42d9751cb 100644 --- a/lib/KernelInterface/src/host.jl +++ b/lib/KernelInterface/src/host.jl @@ -37,9 +37,10 @@ completed. !!! note Backend implementations **must** implement this function cooperatively, yielding to - the Julia scheduler while waiting rather than blocking inside a driver call. See the - [notes for backend implementations](@ref implementations_notes) for why, and for what - it should not wait for. + the Julia scheduler while waiting rather than blocking inside a driver call. It + **must not** wait for work on other queues that the current queue is not ordered + after, e.g., by [`wait_event`](@ref). See the + [notes for backend implementations](@ref implementations_notes) for why. """ function synchronize end From ae556928ebf2d09691e307e67b168df50c45d461 Mon Sep 17 00:00:00 2001 From: Tim Besard Date: Thu, 8 Oct 2026 08:51:04 +0200 Subject: [PATCH 6/6] Don't claim which device a plain Threads.@spawn task uses CUDA.jl and AMDGPU.jl make device! set the default device for new tasks as well, while oneAPI.jl and OpenCL.jl keep it task-local, so a task started with Threads.@spawn may or may not run on its parent's device. Also say which work record_event captures without describing what it doesn't. --- lib/KernelInterface/src/host.jl | 6 +++--- src/spawn.jl | 6 +++--- 2 files changed, 6 insertions(+), 6 deletions(-) diff --git a/lib/KernelInterface/src/host.jl b/lib/KernelInterface/src/host.jl index 42d9751cb..1c3ff0573 100644 --- a/lib/KernelInterface/src/host.jl +++ b/lib/KernelInterface/src/host.jl @@ -47,9 +47,9 @@ function synchronize end """ record_event(backend::Backend) -Capture the work the calling task has queued on `backend`'s active device so far, and -return a handle for [`wait_event`](@ref). Work queued later is not captured, and recording -need not wait for the captured work to complete. The handle is only meant for `wait_event`. +Capture the work the calling task has queued on `backend`'s active device before this +call, and return a handle for [`wait_event`](@ref). Recording need not wait for that work to +complete. The handle is only meant for `wait_event`. !!! note The default implementation calls [`synchronize`](@ref) and returns `nothing`. diff --git a/src/spawn.jl b/src/spawn.jl index e5d9c793e..14417426e 100644 --- a/src/spawn.jl +++ b/src/spawn.jl @@ -32,9 +32,9 @@ fetch(task) == 4 * length(A) # Choosing the device -A task started with plain `Threads.@spawn` runs on the backend's default device, not on -the device of the task that started it. `@spawn` selects the spawning task's device, or the -one given by `device`, an index into `1:ndevices(backend)`: +Which device a task started with plain `Threads.@spawn` uses depends on the backend, and +need not be the spawning task's. `@spawn` selects the spawning task's device, or the one +given by `device`, an index into `1:ndevices(backend)`: ```julia task = KernelAbstractions.@spawn backend device=2 begin