KernelInterface - #1047
Closed
christiangnrd wants to merge 13 commits into
Closed
KernelInterface#1047christiangnrd wants to merge 13 commits into
christiangnrd wants to merge 13 commits into
Conversation
christiangnrd
force-pushed
the
interface
branch
5 times, most recently
from
August 22, 2026 19:57
f95ea6d to
744d030
Compare
christiangnrd
force-pushed
the
interface
branch
3 times, most recently
from
September 5, 2026 22:00
8bc556f to
c991fed
Compare
Contributor
There was a problem hiding this comment.
AMDGPU.jl Benchmarks
Details
| Benchmark suite | Current: 21fc97c | Previous: 482cd05 | Ratio |
|---|---|---|---|
amdgpu/synchronization/context/device |
557.5 ns |
545 ns |
1.02 |
amdgpu/synchronization/stream/blocking |
227.5 ns |
225 ns |
1.01 |
amdgpu/synchronization/stream/nonblocking |
305 ns |
305 ns |
1 |
applications/bitonic_sort |
1102303.5 ns |
1190667 ns |
0.93 |
applications/convolution |
74421 ns |
97316.5 ns |
0.76 |
applications/floyd_warshall |
9021578 ns |
9117782.75 ns |
0.99 |
applications/histogram |
803131.75 ns |
812499.25 ns |
0.99 |
applications/prefix_sum |
190235.25 ns |
186780 ns |
1.02 |
array/accumulate/Float32/1d |
81428.5 ns |
81956.25 ns |
0.99 |
array/accumulate/Float32/dims=1 |
267966.5 ns |
270074 ns |
0.99 |
array/accumulate/Float32/dims=1L |
90609 ns |
91604 ns |
0.99 |
array/accumulate/Float32/dims=2 |
103084 ns |
89373.75 ns |
1.15 |
array/accumulate/Float32/dims=2L |
3030614.5 ns |
3029106.5 ns |
1.00 |
array/accumulate/Int64/1d |
83953.5 ns |
82571.25 ns |
1.02 |
array/accumulate/Int64/dims=1 |
249953.75 ns |
246918.5 ns |
1.01 |
array/accumulate/Int64/dims=1L |
88306.25 ns |
87951.25 ns |
1.00 |
array/accumulate/Int64/dims=2 |
90346.25 ns |
76646.25 ns |
1.18 |
array/accumulate/Int64/dims=2L |
3321223.75 ns |
3388611.75 ns |
0.98 |
array/broadcast |
74498.5 ns |
73561 ns |
1.01 |
array/construct |
2127.5 ns |
2072.75 ns |
1.03 |
array/copy |
39313 ns |
38445.5 ns |
1.02 |
array/copyto!/cpu_to_gpu |
89446.25 ns |
89271.25 ns |
1.00 |
array/copyto!/gpu_to_cpu |
89489 ns |
89381.25 ns |
1.00 |
array/copyto!/gpu_to_gpu |
33943 ns |
33858 ns |
1.00 |
array/iteration/findall/bool |
131377 ns |
127474.25 ns |
1.03 |
array/iteration/findall/int |
147089.5 ns |
140999.5 ns |
1.04 |
array/iteration/findfirst/bool |
142817 ns |
155232.25 ns |
0.92 |
array/iteration/findfirst/int |
164550 ns |
156597.25 ns |
1.05 |
array/iteration/findmin/1d |
103819 ns |
103336.5 ns |
1.00 |
array/iteration/findmin/2d |
90283.75 ns |
90036.25 ns |
1.00 |
array/iteration/logical |
226783.25 ns |
224733.25 ns |
1.01 |
array/iteration/scalar |
293996.75 ns |
295471.75 ns |
1.00 |
array/permutedims/2d |
69976 ns |
69488.5 ns |
1.01 |
array/permutedims/3d |
49540.75 ns |
55823.25 ns |
0.89 |
array/permutedims/4d |
61131 ns |
73458.75 ns |
0.83 |
array/random/rand/Float32 |
45098.25 ns |
44653.25 ns |
1.01 |
array/random/rand/Int64 |
54625.75 ns |
54085.75 ns |
1.01 |
array/random/rand!/Float32 |
41005.5 ns |
41163.25 ns |
1.00 |
array/random/rand!/Int64 |
49220.75 ns |
56780.75 ns |
0.87 |
array/random/randn/Float32 |
70828.5 ns |
69911 ns |
1.01 |
array/random/randn!/Float32 |
56520.75 ns |
55493.25 ns |
1.02 |
array/reductions/mapreduce/Float32/1d |
81476.25 ns |
81456.25 ns |
1.00 |
array/reductions/mapreduce/Float32/dims=1 |
86468.75 ns |
86601.25 ns |
1.00 |
array/reductions/mapreduce/Float32/dims=1L |
837757.25 ns |
841199.75 ns |
1.00 |
array/reductions/mapreduce/Float32/dims=2 |
84998.75 ns |
81728.5 ns |
1.04 |
array/reductions/mapreduce/Float32/dims=2L |
135447 ns |
136954.5 ns |
0.99 |
array/reductions/mapreduce/Int64/1d |
81181.25 ns |
77736 ns |
1.04 |
array/reductions/mapreduce/Int64/dims=1 |
86361.25 ns |
80831.25 ns |
1.07 |
array/reductions/mapreduce/Int64/dims=1L |
842372.25 ns |
844804.75 ns |
1.00 |
array/reductions/mapreduce/Int64/dims=2 |
72793.5 ns |
75918.5 ns |
0.96 |
array/reductions/mapreduce/Int64/dims=2L |
137424.5 ns |
137289.5 ns |
1.00 |
array/reductions/reduce/Float32/1d |
81683.75 ns |
77718.5 ns |
1.05 |
array/reductions/reduce/Float32/dims=1 |
86538.75 ns |
75563.5 ns |
1.15 |
array/reductions/reduce/Float32/dims=1L |
850340 ns |
841087 ns |
1.01 |
array/reductions/reduce/Float32/dims=2 |
82871.25 ns |
76336 ns |
1.09 |
array/reductions/reduce/Float32/dims=2L |
136382 ns |
136187 ns |
1.00 |
array/reductions/reduce/Int64/1d |
80938.75 ns |
81263.5 ns |
1.00 |
array/reductions/reduce/Int64/dims=1 |
69996 ns |
76798.5 ns |
0.91 |
array/reductions/reduce/Int64/dims=1L |
841349.75 ns |
842952.25 ns |
1.00 |
array/reductions/reduce/Int64/dims=2 |
73901.25 ns |
76308.75 ns |
0.97 |
array/reductions/reduce/Int64/dims=2L |
137589.5 ns |
137287 ns |
1.00 |
array/reverse/1d |
42835.75 ns |
35455.5 ns |
1.21 |
array/reverse/1dL |
70628.5 ns |
70423.5 ns |
1.00 |
array/reverse/1dL_inplace |
54688.5 ns |
54620.75 ns |
1.00 |
array/reverse/1d_inplace |
35825.5 ns |
35605.5 ns |
1.01 |
array/reverse/2d |
46635.75 ns |
46543.25 ns |
1.00 |
array/reverse/2dL |
88466.25 ns |
88903.75 ns |
1.00 |
array/reverse/2dL_inplace |
65548.5 ns |
65341 ns |
1.00 |
array/reverse/2d_inplace |
31928 ns |
36125.5 ns |
0.88 |
array/sorting/1d |
327152.25 ns |
327314.75 ns |
1.00 |
gemm/tiled |
1947713.25 ns |
1970010.75 ns |
0.99 |
gemm/tiled_unbounded |
2905917 ns |
1968730.75 ns |
1.48 |
integration/byval/reference |
39401 ns |
39211 ns |
1.00 |
integration/byval/slices=1 |
39810 ns |
40181 ns |
0.99 |
integration/byval/slices=2 |
160742 ns |
144092 ns |
1.12 |
integration/byval/slices=3 |
238863 ns |
234554 ns |
1.02 |
integration/volumerhs |
4879730 ns |
4897830 ns |
1.00 |
kernel/indexing |
34285.5 ns |
28278 ns |
1.21 |
kernel/indexing_checked |
50158.25 ns |
33610.5 ns |
1.49 |
kernel/launch |
1167.5 ns |
1162.5 ns |
1.00 |
kernel/rand |
71286 ns |
96961.25 ns |
0.74 |
latency/import |
1451933250 ns |
1444012666 ns |
1.01 |
latency/precompile |
22930534917 ns |
22844392889 ns |
1.00 |
latency/ttfp |
2224689624 ns |
2200398894 ns |
1.01 |
stencil/diffusion3d |
1603250.75 ns |
1621100.75 ns |
0.99 |
stencil/diffusion3d_checked |
1658281.5 ns |
1656716.25 ns |
1.00 |
This comment was automatically generated by workflow using github-action-benchmark.
christiangnrd
force-pushed
the
interface
branch
4 times, most recently
from
September 12, 2026 15:58
7fe45f5 to
4d75db5
Compare
christiangnrd
force-pushed
the
interface
branch
2 times, most recently
from
September 22, 2026 12:58
a3c00b8 to
9a165fe
Compare
christiangnrd
marked this pull request as ready for review
September 22, 2026 20:14
christiangnrd
force-pushed
the
interface
branch
3 times, most recently
from
September 26, 2026 18:13
a32063f to
db60d30
Compare
christiangnrd
force-pushed
the
interface
branch
from
September 28, 2026 03:24
db60d30 to
9aef27b
Compare
Implement KI.max_work_group_dims from the device's maxThreadsDim, cached in HIPDevice since it is queried on every automatically-sized launch, and KI.max_num_groups as the number of workgroups that fits the work-item grid limit for any valid workgroup size. Requires KernelInterface 0.2.3.
- subtype `KI.Backend`, and implement `KI.launch` instead of calling the kernel; - implement the four primitive index queries with `% T`, and let KernelInterface derive the global ones; - typed sub-group queries, `supports_subgroups` and `supports_shuffle`; - `max_work_group_size(kernel)` is the kernel's limit, `launch_configuration` the occupancy-based recommendation; - report `Float64` and atomics support, and the device of an array; - `kernel_function` keeps the backend it was given; - `copyto!` checks the lengths and returns the destination; - use the generic `zeros` and `ones`; - `_print` is documented as unsupported. KernelInterface 0.3 isn't registered yet, so get it from its branch, and develop it from a clone on Julia 1.10, which ignores `[sources]`.
…wavefronts - `get_num_sub_groups` counts a partial last wavefront (`cld` instead of `÷`); - `get_sub_group_size` is the number of work-items in the wavefront, which is smaller than the wavefront size for the partial one; - `get_sub_group_local_id` is the hardware lane (`mbcnt`), not the index among the active lanes (`activelane`), which changes in divergent code.
`llvm.amdgcn.wave.barrier` only keeps the compiler from moving code across it, so writes before it weren't guaranteed to be visible to the other lanes afterwards. Fence it at wavefront scope, like `sync_workgroup` does at workgroup scope. This makes `KernelInterface.sub_group_barrier` fence memory, as KernelInterface 0.3 requires.
`KI.sub_group_size` promises that kernels from `KI.kernel_function` execute with that sub-group width. `hipfunction` compiles for the device's wavefront size by default, but `wavefrontsize64` could override it; reject a conflicting value.
Julia doesn't specialize a method on `args...` that it only passes through, which made every launch through KernelInterface's generic launch dispatch dynamically (+1.2 µs and +1 kB per launch on CUDA).
KernelInterface 0.3 passes the number of work-items to launch as `nitems`, separately from the bound on the work-group size.
Member
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Do not merge until KernelInterface has been reviewed and interface fully decided