Skip to content

KernelInterface: don't promise how work-items form sub-groups - #815

Merged
maleadt merged 1 commit into
mainfrom
tb/subgroup-contract
Oct 1, 2026
Merged

maleadt merged 1 commit into
mainfrom
tb/subgroup-contract

Conversation

@maleadt

@maleadt maleadt commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

KernelInterface promised that a work-group is divided into sub-groups of the sub-group width, the last of which can be partial, so that there are cld(prod(get_local_size()), get_max_sub_group_size()) of them. Porting OpenCL.jl (JuliaGPU/OpenCL.jl#519) showed that this doesn't hold everywhere: Intel's CPU OpenCL runtime forms sub-groups per row of a multi-dimensional work-group, so a 33×2 work-group consists of four sub-groups of 32 and 1 work-items, not three.

Recording get_sub_group_id() and get_sub_group_size() for every work-item of a 33×2 work-group, with OpenCL.jl and sub-groups of 32:

function record!(ids, sizes)
    l = KI.get_local_id()
    i = (l.y - 1) * KI.get_local_size().x + l.x
    @inbounds ids[i] = KI.get_sub_group_id()
    @inbounds sizes[i] = KI.get_sub_group_size()
    return
end
KI.@launch backend workgroupsize=(33, 2) record!(ids, sizes)
Portable Computing Language: 1 => 32, 2 => 32, 3 => 2
Intel(R) OpenCL: 1 => 32, 2 => 1, 3 => 32, 4 => 1

That isn't a quirk of one runtime. Only CUDA (Programming Guide §3.2.2.2), HSA (PRM §2.6, which AMD's hardware follows) and Metal (MSL §5.2.3.6) guarantee linear, x-fastest packing for multi-dimensional work-groups. The others leave it implementation-defined:

  • OpenCL 3.0 (§3.2.1): "The mapping of work-items to sub-groups is implementation-defined". Only the sub-group with the highest index may be smaller, which Intel's CPU runtime arguably violates, but the conformance tests only check that for 1-D work-groups.
  • SYCL 2020 leaves it implementation-defined. The next revision (SYCL-Docs#860) only specifies the case where the fastest dimension is a multiple of the sub-group size, and its own example allows per-row formation otherwise.
  • Vulkan, WebGPU and HLSL say there is no relation between sub-group ids and local ids. HLSL's GetGroupWaveCount is "at least ceil(...), but may be larger".

Implementations use that freedom. Besides Intel's CPU runtime, which has a setting for it (CL_CONFIG_CPU_SUB_GROUP_CONSTRUCTION, per row by default), Intel's GPU runtime chooses a hardware walk order that isn't x-fastest for some work-group shapes on recent GPUs (Arc, Ponte Vecchio, Meteor Lake and later), and Mesa's Intel and Adreno drivers tile some shaders. Software that relies on linear packing in multi-dimensional work-groups (CUB, rocPRIM, Kokkos' CUDA back end) only targets CUDA and HIP; portable libraries use 1-D work-groups, or only a sub-group's own ids. Nothing in the Julia ecosystem uses KernelInterface's sub-groups yet, so this is the time to fix the contract.

This PR makes KernelInterface promise only what every backend can ensure. Backends that report sub-group support have to guarantee that:

  • every work-item has a unique (get_sub_group_id(), get_sub_group_local_id()) pair in its work-group, which doesn't change during the kernel;
  • the sub-group ids are 1:get_num_sub_groups(), and the lanes of a sub-group are 1:get_sub_group_size();
  • a 1-D work-group of at most sub_group_size(backend) work-items is a single sub-group, which is what kernels that use one sub-group per work-group need.

Which work-items form a sub-group, and which sub-groups are partial, is unspecified. get_num_sub_groups() is at least the count that full sub-groups would need, but can be larger. The price is that storage for a value per sub-group, e.g. the partial results of a reduction, has to allow for one per work-item, and the code combining them has to use get_num_sub_groups() rather than compute the count. That's what the new test does:

function subgroup_combine_kernel(out, ::Val{N}) where {N}
    partial = KI.localmemory(Int32, N)          # N is the work-group size
    if KI.get_sub_group_local_id() == 1
        @inbounds partial[KI.get_sub_group_id()] = KI.get_sub_group_size()
    end
    KI.barrier()
    l = KI.get_local_id()
    if l.x == 1 && l.y == 1 && l.z == 1
        total = Int32(0)
        for i in 1:KI.get_num_sub_groups()
            @inbounds total += partial[i]
        end
        @inbounds out[KI.get_group_id().x] = total
    end
    return
end

The fixed width that sub_group_size(backend) promises is unchanged: kernels still execute with that maximum width. Splitting support for sub-group operations from that guarantee (#810, item 4) is left for another PR.

The testsuite now checks those guarantees over 1-D work-groups around the width, and over 2-D and 3-D ones whose first dimension is or isn't a multiple of the width, and combines a value per sub-group as above. The shuffle test no longer assumes that a work-group of twice the width consists of two full sub-groups. It all passes on PoCL, on CUDA, and on Intel's CPU OpenCL runtime with sub-group support reported. After this, OpenCL.jl#519 no longer has to hide sub-groups on that runtime.

KernelInterface promised that a work-group is divided into sub-groups of the
sub-group width with at most one partial sub-group, i.e.
`cld(prod(get_local_size()), get_max_sub_group_size())` of them. Only CUDA, HSA
(AMD) and Metal specify linear packing of multi-dimensional work-groups. OpenCL
requires that only the highest-numbered sub-group is smaller, but leaves the
mapping implementation-defined, and SYCL, Vulkan and WebGPU don't specify
either. Implementations differ: Intel's CPU OpenCL runtime forms sub-groups per
row of a work-group by default, so a 33x2 work-group has four sub-groups of 32
and 1 work-items, and Intel's GPU runtime picks a hardware walk order that
isn't x-fastest for some shapes.

Require only what backends can ensure everywhere: unique and invariant
(sub-group, lane) pairs, dense ids and lanes, and that a 1-D work-group of at
most the sub-group width is a single sub-group. `get_num_sub_groups` is at
least the count full sub-groups would need, but can be larger, so storage for a
value per sub-group has to allow for one per work-item.

The testsuite checks those invariants over 1-D, 2-D and 3-D work-groups whose
size is or isn't a multiple of the width, and combines a value per sub-group
through local memory, as reductions do. The shuffle test no longer assumes
that a work-group of twice the width consists of two full sub-groups.
maleadt added a commit to JuliaGPU/OpenCL.jl that referenced this pull request Sep 30, 2026
Implement the sub-group queries, `sub_group_barrier` and `shfl_down` with
OpenCL's sub-group builtins, whose semantics already match KernelInterface's:
the last sub-group of a work-group can be partial, and `get_sub_group_size`
counts the work-items that are present.

KernelInterface requires kernels to execute with the sub-group width that
`KI.sub_group_size` reports. `clfunction` requests that width on devices with
`cl_intel_required_subgroup_size`, so only those report sub-group support, and
shuffles for the types that `cl_khr_subgroup_shuffle` covers on the device. A
`sub_group_size` compiler option asking for another width is rejected.

How work-items form sub-groups differs between devices: Intel's CPU runtime
forms them per row of a multi-dimensional work-group. KernelInterface doesn't
promise more than that (JuliaGPU/KernelAbstractions.jl#815).

Co-authored-by: Christian Guinard <28689358+christiangnrd@users.noreply.github.com>
@github-actions

Copy link
Copy Markdown
Contributor

Benchmark Results

Show table
main 1e46c0a... main / 1e46c0a...
const/@Const/Float32/262144 0.278 ± 0.025 ms 0.281 ± 0.027 ms 0.99 ± 0.13
const/@Const/Float32/65536 0.111 ± 0.024 ms 0.113 ± 0.024 ms 0.98 ± 0.29
const/@Const/Float64/262144 0.422 ± 0.02 ms 0.426 ± 0.022 ms 0.992 ± 0.07
const/@Const/Float64/65536 0.162 ± 0.014 ms 0.163 ± 0.013 ms 0.994 ± 0.12
const/unmarked/Float32/262144 1.13 ± 0.034 ms 1.14 ± 0.037 ms 0.989 ± 0.044
const/unmarked/Float32/65536 0.331 ± 0.031 ms 0.32 ± 0.035 ms 1.03 ± 0.15
const/unmarked/Float64/262144 1.37 ± 0.039 ms 1.38 ± 0.042 ms 0.993 ± 0.041
const/unmarked/Float64/65536 0.394 ± 0.033 ms 0.396 ± 0.02 ms 0.996 ± 0.096
launch/3D static workgroup, dynamic ndrange 0.0722 ± 0.016 ms 0.0726 ± 0.016 ms 0.995 ± 0.31
launch/3D static workgroup, static ndrange 0.0717 ± 0.027 ms 0.0718 ± 0.024 ms 0.998 ± 0.5
launch/dynamic workgroup, dynamic ndrange 0.0697 ± 0.036 ms 0.0718 ± 0.028 ms 0.971 ± 0.63
launch/dynamic workgroup, dynamic ndrange, workgroupsize given 0.0687 ± 0.036 ms 0.0715 ± 0.032 ms 0.961 ± 0.66
launch/static workgroup, dynamic ndrange 0.0719 ± 0.023 ms 0.0726 ± 0.02 ms 0.991 ± 0.43
launch/static workgroup, static ndrange 0.0718 ± 0.018 ms 0.0725 ± 0.019 ms 0.99 ± 0.36
partition/dynamic workgroup, dynamic ndrange 0.0581 ± 0.0011 μs 0.0636 ± 0.0027 μs 0.913 ± 0.043
partition/static workgroup, dynamic ndrange 0.0692 ± 0.012 μs 0.0646 ± 0.012 μs 1.07 ± 0.27
partition/static workgroup, static ndrange 1.55 ± 0.009 ns 1.55 ± 0.01 ns 1 ± 0.0087
saxpy/default/Float16/1024 0.0729 ± 0.025 ms 0.0714 ± 0.028 ms 1.02 ± 0.54
saxpy/default/Float16/1048576 0.794 ± 0.043 ms 0.796 ± 0.045 ms 0.998 ± 0.078
saxpy/default/Float16/16384 0.0606 ± 0.031 ms 0.061 ± 0.031 ms 0.993 ± 0.72
saxpy/default/Float16/2048 0.0736 ± 0.028 ms 0.0745 ± 0.021 ms 0.988 ± 0.47
saxpy/default/Float16/256 0.0718 ± 0.035 ms 0.0719 ± 0.03 ms 0.999 ± 0.63
saxpy/default/Float16/262144 0.244 ± 0.035 ms 0.246 ± 0.035 ms 0.994 ± 0.2
saxpy/default/Float16/32768 0.0699 ± 0.03 ms 0.0707 ± 0.03 ms 0.988 ± 0.6
saxpy/default/Float16/4096 0.072 ± 0.032 ms 0.0746 ± 0.032 ms 0.965 ± 0.59
saxpy/default/Float16/512 0.0719 ± 0.034 ms 0.0714 ± 0.031 ms 1.01 ± 0.65
saxpy/default/Float16/64 0.0731 ± 0.013 ms 0.0724 ± 0.026 ms 1.01 ± 0.41
saxpy/default/Float16/65536 0.0913 ± 0.028 ms 0.0927 ± 0.03 ms 0.984 ± 0.44
saxpy/default/Float32/1024 0.0708 ± 0.032 ms 0.0714 ± 0.03 ms 0.991 ± 0.61
saxpy/default/Float32/1048576 0.402 ± 0.041 ms 0.388 ± 0.038 ms 1.04 ± 0.15
saxpy/default/Float32/16384 0.0527 ± 0.029 ms 0.0554 ± 0.029 ms 0.951 ± 0.73
saxpy/default/Float32/2048 0.0723 ± 0.028 ms 0.0722 ± 0.026 ms 1 ± 0.53
saxpy/default/Float32/256 0.0718 ± 0.035 ms 0.0727 ± 0.023 ms 0.987 ± 0.58
saxpy/default/Float32/262144 0.139 ± 0.03 ms 0.14 ± 0.031 ms 0.993 ± 0.3
saxpy/default/Float32/32768 0.0585 ± 0.031 ms 0.058 ± 0.031 ms 1.01 ± 0.75
saxpy/default/Float32/4096 0.0743 ± 0.026 ms 0.0746 ± 0.018 ms 0.996 ± 0.43
saxpy/default/Float32/512 0.0708 ± 0.033 ms 0.0718 ± 0.016 ms 0.986 ± 0.52
saxpy/default/Float32/64 0.0724 ± 0.033 ms 0.0723 ± 0.031 ms 1 ± 0.62
saxpy/default/Float32/65536 0.0712 ± 0.03 ms 0.0717 ± 0.03 ms 0.994 ± 0.59
saxpy/default/Float64/1024 0.0711 ± 0.031 ms 0.0711 ± 0.028 ms 1 ± 0.59
saxpy/default/Float64/1048576 0.639 ± 0.076 ms 0.619 ± 0.063 ms 1.03 ± 0.16
saxpy/default/Float64/16384 0.0532 ± 0.027 ms 0.0537 ± 0.028 ms 0.991 ± 0.73
saxpy/default/Float64/2048 0.0637 ± 0.033 ms 0.0634 ± 0.032 ms 1.01 ± 0.72
saxpy/default/Float64/256 0.0712 ± 0.032 ms 0.071 ± 0.031 ms 1 ± 0.63
saxpy/default/Float64/262144 0.196 ± 0.036 ms 0.198 ± 0.037 ms 0.992 ± 0.26
saxpy/default/Float64/32768 0.0654 ± 0.03 ms 0.0638 ± 0.031 ms 1.02 ± 0.69
saxpy/default/Float64/4096 0.0634 ± 0.028 ms 0.0653 ± 0.027 ms 0.971 ± 0.59
saxpy/default/Float64/512 0.0709 ± 0.033 ms 0.0711 ± 0.03 ms 0.998 ± 0.63
saxpy/default/Float64/64 0.0734 ± 0.014 ms 0.0724 ± 0.026 ms 1.01 ± 0.41
saxpy/default/Float64/65536 0.0852 ± 0.031 ms 0.0849 ± 0.032 ms 1 ± 0.52
saxpy/static workgroup=(1024,)/Float16/1024 0.0715 ± 0.03 ms 0.0713 ± 0.03 ms 1 ± 0.59
saxpy/static workgroup=(1024,)/Float16/1048576 0.793 ± 0.041 ms 0.794 ± 0.042 ms 0.998 ± 0.074
saxpy/static workgroup=(1024,)/Float16/16384 0.0597 ± 0.029 ms 0.0607 ± 0.031 ms 0.983 ± 0.69
saxpy/static workgroup=(1024,)/Float16/2048 0.074 ± 0.028 ms 0.0743 ± 0.027 ms 0.996 ± 0.53
saxpy/static workgroup=(1024,)/Float16/256 0.0713 ± 0.033 ms 0.0716 ± 0.03 ms 0.996 ± 0.63
saxpy/static workgroup=(1024,)/Float16/262144 0.241 ± 0.033 ms 0.243 ± 0.035 ms 0.991 ± 0.2
saxpy/static workgroup=(1024,)/Float16/32768 0.07 ± 0.029 ms 0.0708 ± 0.029 ms 0.989 ± 0.58
saxpy/static workgroup=(1024,)/Float16/4096 0.065 ± 0.032 ms 0.0692 ± 0.032 ms 0.939 ± 0.63
saxpy/static workgroup=(1024,)/Float16/512 0.0717 ± 0.03 ms 0.0726 ± 0.026 ms 0.989 ± 0.55
saxpy/static workgroup=(1024,)/Float16/64 0.072 ± 0.03 ms 0.0725 ± 0.023 ms 0.993 ± 0.51
saxpy/static workgroup=(1024,)/Float16/65536 0.0936 ± 0.028 ms 0.0924 ± 0.028 ms 1.01 ± 0.43
saxpy/static workgroup=(1024,)/Float32/1024 0.0719 ± 0.029 ms 0.0709 ± 0.034 ms 1.01 ± 0.63
saxpy/static workgroup=(1024,)/Float32/1048576 0.395 ± 0.034 ms 0.394 ± 0.035 ms 1 ± 0.12
saxpy/static workgroup=(1024,)/Float32/16384 0.054 ± 0.027 ms 0.0558 ± 0.027 ms 0.969 ± 0.67
saxpy/static workgroup=(1024,)/Float32/2048 0.0716 ± 0.03 ms 0.0715 ± 0.029 ms 1 ± 0.58
saxpy/static workgroup=(1024,)/Float32/256 0.072 ± 0.034 ms 0.0717 ± 0.031 ms 1 ± 0.64
saxpy/static workgroup=(1024,)/Float32/262144 0.148 ± 0.029 ms 0.148 ± 0.028 ms 1 ± 0.27
saxpy/static workgroup=(1024,)/Float32/32768 0.0584 ± 0.027 ms 0.0605 ± 0.028 ms 0.965 ± 0.63
saxpy/static workgroup=(1024,)/Float32/4096 0.0741 ± 0.029 ms 0.0734 ± 0.028 ms 1.01 ± 0.55
saxpy/static workgroup=(1024,)/Float32/512 0.0718 ± 0.023 ms 0.071 ± 0.028 ms 1.01 ± 0.52
saxpy/static workgroup=(1024,)/Float32/64 0.0726 ± 0.022 ms 0.0723 ± 0.024 ms 1 ± 0.45
saxpy/static workgroup=(1024,)/Float32/65536 0.0724 ± 0.029 ms 0.0729 ± 0.03 ms 0.993 ± 0.57
saxpy/static workgroup=(1024,)/Float64/1024 0.0714 ± 0.024 ms 0.0709 ± 0.026 ms 1.01 ± 0.5
saxpy/static workgroup=(1024,)/Float64/1048576 0.604 ± 0.064 ms 0.58 ± 0.054 ms 1.04 ± 0.15
saxpy/static workgroup=(1024,)/Float64/16384 0.0587 ± 0.028 ms 0.0595 ± 0.029 ms 0.987 ± 0.67
saxpy/static workgroup=(1024,)/Float64/2048 0.0601 ± 0.033 ms 0.0623 ± 0.032 ms 0.966 ± 0.72
saxpy/static workgroup=(1024,)/Float64/256 0.0706 ± 0.031 ms 0.0716 ± 0.025 ms 0.987 ± 0.55
saxpy/static workgroup=(1024,)/Float64/262144 0.196 ± 0.034 ms 0.198 ± 0.032 ms 0.989 ± 0.23
saxpy/static workgroup=(1024,)/Float64/32768 0.0673 ± 0.029 ms 0.0688 ± 0.028 ms 0.979 ± 0.58
saxpy/static workgroup=(1024,)/Float64/4096 0.0531 ± 0.029 ms 0.0556 ± 0.028 ms 0.954 ± 0.72
saxpy/static workgroup=(1024,)/Float64/512 0.0714 ± 0.024 ms 0.0711 ± 0.026 ms 1 ± 0.5
saxpy/static workgroup=(1024,)/Float64/64 0.0728 ± 0.015 ms 0.0723 ± 0.024 ms 1.01 ± 0.4
saxpy/static workgroup=(1024,)/Float64/65536 0.0898 ± 0.032 ms 0.0905 ± 0.032 ms 0.993 ± 0.5
time_to_load 0.76 ± 0.013 s 0.788 ± 0.013 s 0.964 ± 0.022
main 1e46c0a... main / 1e46c0a...
const/@Const/Float32/262144 2 allocs: 32 B 2 allocs: 32 B 1
const/@Const/Float32/65536 2 allocs: 32 B 2 allocs: 32 B 1
const/@Const/Float64/262144 2 allocs: 32 B 2 allocs: 32 B 1
const/@Const/Float64/65536 2 allocs: 32 B 2 allocs: 32 B 1
const/unmarked/Float32/262144 2 allocs: 32 B 2 allocs: 32 B 1
const/unmarked/Float32/65536 2 allocs: 32 B 2 allocs: 32 B 1
const/unmarked/Float64/262144 2 allocs: 32 B 2 allocs: 32 B 1
const/unmarked/Float64/65536 2 allocs: 32 B 2 allocs: 32 B 1
launch/3D static workgroup, dynamic ndrange 6 allocs: 0.156 kB 6 allocs: 0.156 kB 1
launch/3D static workgroup, static ndrange 6 allocs: 0.156 kB 6 allocs: 0.156 kB 1
launch/dynamic workgroup, dynamic ndrange 2 allocs: 32 B 2 allocs: 32 B 1
launch/dynamic workgroup, dynamic ndrange, workgroupsize given 2 allocs: 32 B 2 allocs: 32 B 1
launch/static workgroup, dynamic ndrange 2 allocs: 32 B 2 allocs: 32 B 1
launch/static workgroup, static ndrange 2 allocs: 32 B 2 allocs: 32 B 1
partition/dynamic workgroup, dynamic ndrange 2 allocs: 0.0625 kB 2 allocs: 0.0625 kB 1
partition/static workgroup, dynamic ndrange 2 allocs: 32 B 2 allocs: 32 B 1
partition/static workgroup, static ndrange 0 allocs: 0 B 0 allocs: 0 B
saxpy/default/Float16/1024 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float16/1048576 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float16/16384 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float16/2048 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float16/256 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/default/Float16/262144 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float16/32768 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float16/4096 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float16/512 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float16/64 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/default/Float16/65536 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float32/1024 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float32/1048576 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float32/16384 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float32/2048 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float32/256 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/default/Float32/262144 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float32/32768 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float32/4096 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float32/512 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float32/64 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/default/Float32/65536 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float64/1024 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float64/1048576 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float64/16384 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float64/2048 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float64/256 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/default/Float64/262144 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float64/32768 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float64/4096 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float64/512 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/default/Float64/64 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/default/Float64/65536 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float16/1024 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float16/1048576 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float16/16384 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float16/2048 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float16/256 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/static workgroup=(1024,)/Float16/262144 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float16/32768 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float16/4096 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float16/512 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float16/64 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/static workgroup=(1024,)/Float16/65536 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float32/1024 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float32/1048576 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float32/16384 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float32/2048 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float32/256 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/static workgroup=(1024,)/Float32/262144 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float32/32768 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float32/4096 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float32/512 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float32/64 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/static workgroup=(1024,)/Float32/65536 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float64/1024 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float64/1048576 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float64/16384 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float64/2048 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float64/256 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/static workgroup=(1024,)/Float64/262144 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float64/32768 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float64/4096 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float64/512 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
saxpy/static workgroup=(1024,)/Float64/64 2 allocs: 32 B 2 allocs: 32 B 1
saxpy/static workgroup=(1024,)/Float64/65536 5 allocs: 0.0781 kB 5 allocs: 0.0781 kB 1
time_to_load 0.2 k allocs: 11.8 kB 0.2 k allocs: 11.8 kB 1

Benchmark Plots

A plot of the benchmark results have been uploaded as an artifact to the workflow run for this PR.
Go to "Actions"->"Benchmark a pull request"->[the most recent run]->"Artifacts" (at the bottom).

@maleadt
maleadt merged commit d3ef28a into main Oct 1, 2026
71 of 73 checks passed
@maleadt
maleadt deleted the tb/subgroup-contract branch October 1, 2026 06:33
@codecov

codecov Bot commented Oct 1, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 69.54%. Comparing base (225a67a) to head (1e46c0a).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files
@@           Coverage Diff           @@
##             main     #815   +/-   ##
=======================================
  Coverage   69.54%   69.54%           
=======================================
  Files          26       26           
  Lines        2157     2157           
=======================================
  Hits         1500     1500           
  Misses        657      657           

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant