Conversation
Contributor
Benchmark ResultsShow table
Benchmark PlotsA plot of the benchmark results have been uploaded as an artifact to the workflow run for this PR. |
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## vc/ki-subgroup-ops #830 +/- ##
======================================================
+ Coverage 74.15% 75.70% +1.55%
======================================================
Files 24 25 +1
Lines 2275 2437 +162
======================================================
+ Hits 1687 1845 +158
- Misses 588 592 +4 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
vchuravy
force-pushed
the
vc/groupreduce
branch
from
October 3, 2026 22:49
507c83a to
df037a3
Compare
vchuravy
force-pushed
the
vc/groupreduce
branch
from
October 4, 2026 06:44
df037a3 to
25c9cb5
Compare
vchuravy
added this pull request to stack #834
October 4, 2026 06:46
vchuravy
force-pushed
the
vc/groupreduce
branch
6 times, most recently
from
October 4, 2026 10:43
81e4230 to
06fdf57
Compare
vchuravy
force-pushed
the
vc/groupreduce
branch
from
October 4, 2026 10:59
06fdf57 to
280c4e1
Compare
vchuravy
force-pushed
the
vc/groupreduce
branch
3 times, most recently
from
October 4, 2026 12:31
5f5c328 to
a424d5d
Compare
Rework of #559 on top of KernelInterface: - `@groupreduce(op, val, neutral[, groupsize]; subgroups=false)` reduces over the workgroup and returns the result on every work-item. It uses a local-memory tree by default, or a two-level reduction based on `KI.shfl_down` with `subgroups=true` (gated on the host by `KI.supports_shuffle`). The local memory is sized by the static workgroup size or an explicit upper bound. - `@subgroupreduce(op, val, neutral)` reduces over the sub-group with shuffles; the result is defined on the first lane. - Both are collectives in `@kernel`: the split treats them like `@synchronize`, and padding work-items contribute `neutral` without evaluating `val`, so ndranges that are not a multiple of the workgroup size work. Co-authored-by: Anton Smirnov <tonysmn97@gmail.com> Assisted-by: Claude Code (Opus 5.5)
- `@groupscan(op, val, neutral[, groupsize]; inclusive = true)` scans over the workgroup in the order of the local linear index, with a Hillis-Steele scan in double-buffered local memory, sized like `@groupreduce`. - `@subgroupscan(op, val, neutral; inclusive = true)` scans over the lanes of a sub-group with `KI.shfl_up`. Both only need `op` to be associative, and are collectives in `@kernel` like the reductions: padding work-items contribute `neutral`. Assisted-by: Claude Code (Opus 5.5)
`@subgroupreduce` and `@subgroupscan`, and the sub-group stage of `@groupreduce`, now use `KI.sub_group_reduce` and `KI.sub_group_scan`, which backends can implement with native operations. `@subgroupreduce` returns the result on every work-item of the sub-group. Test reductions of (value, index) pairs, and don't assume that sub-groups are formed from consecutive work-items. Assisted-by: Claude Code (Opus 5.5)
KernelAbstractions' collectives are executed by the padding work-items of a partial workgroup, but direct calls of KernelInterface's sub-group functions aren't: kernels that use them need `unsafe_indices=true`. Assisted-by: Claude Code (Opus 5.5)
The type of the value passed to `@groupreduce`, `@subgroupreduce`,
`@groupscan` or `@subgroupscan` may differ between the work-items: in a
`@kernel` the padding work-items contribute `neutral` instead of `val`, so
`@groupreduce(+, x[i]::Float32, 0.0)` reduces a `Union{Float32, Float64}`,
and an accumulator that only some work-items add a `Float64` to is a `Union`
as well. Julia union-splits the call of the collective with such an argument
into one call per type, so the work-items of a workgroup executed different
copies of its barriers and shuffles. On PoCL this silently gave wrong
results (0 for the reduction of a padded workgroup, NaN for the energy of a
Float32 system with a Float64 Coulomb constant in Molly).
Convert the value to the type of `neutral` at the call site instead, so that
only the conversion is union-split, and test mixed types for all four
collectives.
Assisted-by: Claude Code (Opus 5.5)
vchuravy
force-pushed
the
vc/groupreduce
branch
from
October 4, 2026 16:35
8333aea to
9b20320
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Supersedes #559 by @pxl-th (credited as co-author), reworked on top of KernelInterface.
Stacked on #831:
@subgroupscanneedsKI.shfl_up(and the struct shuffles) from there, so this PR targets that branch for now. Merge #831 first; this one then retargets tomain.API
@groupreduce(op, val, neutral[, groupsize]; subgroups = false)reducesvalover the workgroup and returns the result on every work-item.subgroups = true: each sub-group reduces withKI.shfl_down, then the sub-group results are combined (constant number of barriers). Gate it on the host withKI.supports_shuffle(backend, T)and pass it in as a constant (e.g.::Val{S}).groupsizeas a compile-time upper bound for dynamic workgroup sizes.@subgroupreduce(op, val, neutral)reduces over the sub-group withKI.sub_group_reduce(backends may use native reductions); the result is defined on every lane. It only needsopto be associative.@groupscan(op, val, neutral[, groupsize]; inclusive = true)scans in the order of@index(Local, Linear), inclusive or (withinclusive = false) exclusive. It uses a Hillis-Steele scan in double-buffered local memory, sized like@groupreduce. There is no sub-group variant: how work-items form sub-groups is unspecified in KernelInterface, so a scan built from sub-group scans wouldn't follow the local index order.@subgroupscan(op, val, neutral; inclusive = true)scans over the lanes of a sub-group withKI.sub_group_scan.opto be associative. They are collectives like the reductions, so padding work-items contributeneutral. A use case is stream compaction: offsets from an exclusive@groupscanof the predicates (cf. Molly.jl's neighbor finder).Differences to #559
shfl_down/supports_warp_reductionare gone: they'reKI.shfl_down/KI.supports_shufflenow, and the sub-group width comes from KI instead of a hardcoded 32.@kernelsplit, like@synchronize. Padding work-items take part, contributeneutraland don't evaluateval. This fixes the uninitialized-local-memory issue raised in Implement groupreduce API #559, and is whyneutralis required.y[i] = @groupreduce(...)) is a macro-expansion error, since it would run on padding work-items.Values of different types
The macros convert
valto the type ofneutralat the call site. Otherwise, avalwhose type differs between work-items makes Julia union-split the call, and the work-items run different copies of its barriers and shuffles. Two cases trigger this:neutralhaving a different type thanval(e.g.@groupreduce(+, x[i]::Float32, 0.0));On POCL this silently returned wrong results. It was found through a NaN in the Molly.jl port.
Tests
test/groupreduce.jlruns as part of the backend testsuite and covers:+/maxon Int32, Int64 and Float32unsafe_indices@groupscan(inclusive/exclusive): partial and non-power-of-two workgroups, a dynamic workgroup size with a bound, Cartesian workgroups with padding in the middle, reuse in a loop@subgroupscan, checked against the lane order the kernel reports@groupreduceof (value, index) pairs (argmin)@subgroupreduceon every lane, without assuming that sub-groups are consecutive work-items, and the macro errorsThe full test suite passes locally on POCL.
🤖 Generated with Claude Code