Skip to content

KernelInterface - #1047

Closed
christiangnrd wants to merge 13 commits into
JuliaGPU:mainfrom
christiangnrd:interface
Closed

christiangnrd wants to merge 13 commits into
JuliaGPU:mainfrom
christiangnrd:interface

Conversation

@christiangnrd

@christiangnrd christiangnrd commented Aug 22, 2026 •

Copy link
Copy Markdown
Member

Do not merge until KernelInterface has been reviewed and interface fully decided

@christiangnrd
christiangnrd force-pushed the interface branch 5 times, most recently from f95ea6d to 744d030 Compare August 22, 2026 19:57
@christiangnrd
christiangnrd force-pushed the interface branch 3 times, most recently from 8bc556f to c991fed Compare September 5, 2026 22:00

@github-actions github-actions Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AMDGPU.jl Benchmarks

Details
Benchmark suite Current: 21fc97c Previous: 482cd05 Ratio
amdgpu/synchronization/context/device 557.5 ns 545 ns 1.02
amdgpu/synchronization/stream/blocking 227.5 ns 225 ns 1.01
amdgpu/synchronization/stream/nonblocking 305 ns 305 ns 1
applications/bitonic_sort 1102303.5 ns 1190667 ns 0.93
applications/convolution 74421 ns 97316.5 ns 0.76
applications/floyd_warshall 9021578 ns 9117782.75 ns 0.99
applications/histogram 803131.75 ns 812499.25 ns 0.99
applications/prefix_sum 190235.25 ns 186780 ns 1.02
array/accumulate/Float32/1d 81428.5 ns 81956.25 ns 0.99
array/accumulate/Float32/dims=1 267966.5 ns 270074 ns 0.99
array/accumulate/Float32/dims=1L 90609 ns 91604 ns 0.99
array/accumulate/Float32/dims=2 103084 ns 89373.75 ns 1.15
array/accumulate/Float32/dims=2L 3030614.5 ns 3029106.5 ns 1.00
array/accumulate/Int64/1d 83953.5 ns 82571.25 ns 1.02
array/accumulate/Int64/dims=1 249953.75 ns 246918.5 ns 1.01
array/accumulate/Int64/dims=1L 88306.25 ns 87951.25 ns 1.00
array/accumulate/Int64/dims=2 90346.25 ns 76646.25 ns 1.18
array/accumulate/Int64/dims=2L 3321223.75 ns 3388611.75 ns 0.98
array/broadcast 74498.5 ns 73561 ns 1.01
array/construct 2127.5 ns 2072.75 ns 1.03
array/copy 39313 ns 38445.5 ns 1.02
array/copyto!/cpu_to_gpu 89446.25 ns 89271.25 ns 1.00
array/copyto!/gpu_to_cpu 89489 ns 89381.25 ns 1.00
array/copyto!/gpu_to_gpu 33943 ns 33858 ns 1.00
array/iteration/findall/bool 131377 ns 127474.25 ns 1.03
array/iteration/findall/int 147089.5 ns 140999.5 ns 1.04
array/iteration/findfirst/bool 142817 ns 155232.25 ns 0.92
array/iteration/findfirst/int 164550 ns 156597.25 ns 1.05
array/iteration/findmin/1d 103819 ns 103336.5 ns 1.00
array/iteration/findmin/2d 90283.75 ns 90036.25 ns 1.00
array/iteration/logical 226783.25 ns 224733.25 ns 1.01
array/iteration/scalar 293996.75 ns 295471.75 ns 1.00
array/permutedims/2d 69976 ns 69488.5 ns 1.01
array/permutedims/3d 49540.75 ns 55823.25 ns 0.89
array/permutedims/4d 61131 ns 73458.75 ns 0.83
array/random/rand/Float32 45098.25 ns 44653.25 ns 1.01
array/random/rand/Int64 54625.75 ns 54085.75 ns 1.01
array/random/rand!/Float32 41005.5 ns 41163.25 ns 1.00
array/random/rand!/Int64 49220.75 ns 56780.75 ns 0.87
array/random/randn/Float32 70828.5 ns 69911 ns 1.01
array/random/randn!/Float32 56520.75 ns 55493.25 ns 1.02
array/reductions/mapreduce/Float32/1d 81476.25 ns 81456.25 ns 1.00
array/reductions/mapreduce/Float32/dims=1 86468.75 ns 86601.25 ns 1.00
array/reductions/mapreduce/Float32/dims=1L 837757.25 ns 841199.75 ns 1.00
array/reductions/mapreduce/Float32/dims=2 84998.75 ns 81728.5 ns 1.04
array/reductions/mapreduce/Float32/dims=2L 135447 ns 136954.5 ns 0.99
array/reductions/mapreduce/Int64/1d 81181.25 ns 77736 ns 1.04
array/reductions/mapreduce/Int64/dims=1 86361.25 ns 80831.25 ns 1.07
array/reductions/mapreduce/Int64/dims=1L 842372.25 ns 844804.75 ns 1.00
array/reductions/mapreduce/Int64/dims=2 72793.5 ns 75918.5 ns 0.96
array/reductions/mapreduce/Int64/dims=2L 137424.5 ns 137289.5 ns 1.00
array/reductions/reduce/Float32/1d 81683.75 ns 77718.5 ns 1.05
array/reductions/reduce/Float32/dims=1 86538.75 ns 75563.5 ns 1.15
array/reductions/reduce/Float32/dims=1L 850340 ns 841087 ns 1.01
array/reductions/reduce/Float32/dims=2 82871.25 ns 76336 ns 1.09
array/reductions/reduce/Float32/dims=2L 136382 ns 136187 ns 1.00
array/reductions/reduce/Int64/1d 80938.75 ns 81263.5 ns 1.00
array/reductions/reduce/Int64/dims=1 69996 ns 76798.5 ns 0.91
array/reductions/reduce/Int64/dims=1L 841349.75 ns 842952.25 ns 1.00
array/reductions/reduce/Int64/dims=2 73901.25 ns 76308.75 ns 0.97
array/reductions/reduce/Int64/dims=2L 137589.5 ns 137287 ns 1.00
array/reverse/1d 42835.75 ns 35455.5 ns 1.21
array/reverse/1dL 70628.5 ns 70423.5 ns 1.00
array/reverse/1dL_inplace 54688.5 ns 54620.75 ns 1.00
array/reverse/1d_inplace 35825.5 ns 35605.5 ns 1.01
array/reverse/2d 46635.75 ns 46543.25 ns 1.00
array/reverse/2dL 88466.25 ns 88903.75 ns 1.00
array/reverse/2dL_inplace 65548.5 ns 65341 ns 1.00
array/reverse/2d_inplace 31928 ns 36125.5 ns 0.88
array/sorting/1d 327152.25 ns 327314.75 ns 1.00
gemm/tiled 1947713.25 ns 1970010.75 ns 0.99
gemm/tiled_unbounded 2905917 ns 1968730.75 ns 1.48
integration/byval/reference 39401 ns 39211 ns 1.00
integration/byval/slices=1 39810 ns 40181 ns 0.99
integration/byval/slices=2 160742 ns 144092 ns 1.12
integration/byval/slices=3 238863 ns 234554 ns 1.02
integration/volumerhs 4879730 ns 4897830 ns 1.00
kernel/indexing 34285.5 ns 28278 ns 1.21
kernel/indexing_checked 50158.25 ns 33610.5 ns 1.49
kernel/launch 1167.5 ns 1162.5 ns 1.00
kernel/rand 71286 ns 96961.25 ns 0.74
latency/import 1451933250 ns 1444012666 ns 1.01
latency/precompile 22930534917 ns 22844392889 ns 1.00
latency/ttfp 2224689624 ns 2200398894 ns 1.01
stencil/diffusion3d 1603250.75 ns 1621100.75 ns 0.99
stencil/diffusion3d_checked 1658281.5 ns 1656716.25 ns 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@christiangnrd
christiangnrd force-pushed the interface branch 4 times, most recently from 7fe45f5 to 4d75db5 Compare September 12, 2026 15:58
@christiangnrd
christiangnrd force-pushed the interface branch 2 times, most recently from a3c00b8 to 9a165fe Compare September 22, 2026 12:58
@christiangnrd
christiangnrd marked this pull request as ready for review September 22, 2026 20:14
@christiangnrd
christiangnrd force-pushed the interface branch 3 times, most recently from a32063f to db60d30 Compare September 26, 2026 18:13
Implement KI.max_work_group_dims from the device's maxThreadsDim, cached in
HIPDevice since it is queried on every automatically-sized launch, and
KI.max_num_groups as the number of workgroups that fits the work-item grid
limit for any valid workgroup size. Requires KernelInterface 0.2.3.
- subtype `KI.Backend`, and implement `KI.launch` instead of calling the kernel;
- implement the four primitive index queries with `% T`, and let KernelInterface derive
  the global ones;
- typed sub-group queries, `supports_subgroups` and `supports_shuffle`;
- `max_work_group_size(kernel)` is the kernel's limit, `launch_configuration` the
  occupancy-based recommendation;
- report `Float64` and atomics support, and the device of an array;
- `kernel_function` keeps the backend it was given;
- `copyto!` checks the lengths and returns the destination;
- use the generic `zeros` and `ones`;
- `_print` is documented as unsupported.

KernelInterface 0.3 isn't registered yet, so get it from its branch, and develop it from a
clone on Julia 1.10, which ignores `[sources]`.
…wavefronts

- `get_num_sub_groups` counts a partial last wavefront (`cld` instead of `÷`);
- `get_sub_group_size` is the number of work-items in the wavefront, which is smaller than
  the wavefront size for the partial one;
- `get_sub_group_local_id` is the hardware lane (`mbcnt`), not the index among the active
  lanes (`activelane`), which changes in divergent code.
`llvm.amdgcn.wave.barrier` only keeps the compiler from moving code across it, so writes
before it weren't guaranteed to be visible to the other lanes afterwards. Fence it at
wavefront scope, like `sync_workgroup` does at workgroup scope. This makes
`KernelInterface.sub_group_barrier` fence memory, as KernelInterface 0.3 requires.
`KI.sub_group_size` promises that kernels from `KI.kernel_function` execute with that
sub-group width. `hipfunction` compiles for the device's wavefront size by default, but
`wavefrontsize64` could override it; reject a conflicting value.
Julia doesn't specialize a method on `args...` that it only passes through, which
made every launch through KernelInterface's generic launch dispatch dynamically
(+1.2 µs and +1 kB per launch on CUDA).
KernelInterface 0.3 passes the number of work-items to launch as `nitems`,
separately from the bound on the work-group size.
@maleadt

maleadt commented Sep 30, 2026

Copy link
Copy Markdown
Member

#1117

@maleadt maleadt closed this Sep 30, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants