ggml : support F32 input for int8 convrot tensorwise mulmat in Vulkan - #8
Open
noctrex wants to merge 4 commits into
Open
ggml : support F32 input for int8 convrot tensorwise mulmat in Vulkan#8noctrex wants to merge 4 commits into
noctrex wants to merge 4 commits into
Conversation
Remove the blanket GGML_USE_HIP exclusion from the INT8 tensorwise convrot ops and scope the capability check to GPUs with working integer GEMM support (turing_mma_available, amd_wmma_available, amd_mfma_available). RDNA2 stays excluded: hipBLASLt INT8 is broken there. vendors/hip.h: map CUDA_R_8I, CUDA_R_32I, CUBLAS_COMPUTE_32I and __shfl_down_sync to their HIP equivalents. HIP's masked shuffle templates require a 64-bit mask, so route to the maskless intrinsic like the other __shfl_*_sync mappings. Verified on gfx1100 (Windows ROCm 7.1): ggml tests/test-int8-convrot passes, and CPU-vs-GPU output is bit-exact at real model shapes (K=3840/10240, N up to 11520, 8104 rows).
Template the INT8 convrot quantization kernels on group_size and instantiate 64 alongside 256. The Hadamard normalization is now derived from the group size (2^-log4(group_size): 1/8 for G=64, 1/16 for G=256). The small-kernel max reduction uses a shared-memory tree instead of warp shuffles so it stays correct for 16-thread blocks. Capability gates accept group size 64 and 256. Z-Image int8_tensorwise exports in the wild use convrot_groupsize 64; without this they fall back to CPU on every GPU backend. Verified on gfx1100 (Windows ROCm 7.1): exact-reference tests pass for H64 (single-group and multi-group packed GEMM), H256 unchanged, and CPU-vs-GPU output is bit-exact at real Z-Image shapes.
Parameterize the quantize shader with a GROUP_SIZE macro (default 256,
validated {64, 256}), register the G64 SPIR-V variant, relax the quantize
and tensorwise mulmat gates to accept group size 64, and select the
pipeline by the op param group size. The GEMM shaders are group-size
agnostic and unchanged. The G256 variant compiles to byte-identical
SPIR-V as before.
The tensorwise mulmat accepts src1 as either the packed QUANTIZE_I8_CONVROT output or raw F32, but the Vulkan backend only handled the packed form, so F32 inputs fell back to the CPU via the scheduler. Relax the supports_op gate to accept the F32 form, bounded by maxStorageBufferRange for the packed staging size. Quantize the F32 input on the fly into a grow-on-demand device prealloc buffer in the packed payload+scales layout, then run the existing tensorwise GEMM unchanged. Extend the Vulkan test with F32-input cases (GPU packed equivalence, exact CPU reference).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The known limitation from #7: mul_mat_i8_tensorwise accepts src1 as either the packed QUANTIZE_I8_CONVROT output or raw F32 (CUDA/HIP quantize inline), but Vulkan only handled the packed form, so F32 inputs were rejected in supports_op and those graphs fell back to CPU via the scheduler.
Consumers such as stable-diffusion.cpp emit the F32 form for Linears with a folded activation scale, which made the Vulkan path fall back on real models.
This closes that gap for group sizes 64 and 256.
Implementation (7a91b8f):
Tests: tests/test-int8-convrot.cpp gained F32-input cases - GPU packed-vs-F32 equivalence plus an exact CPU reference in CPU mode.
Verification (RX 7900 XTX, Windows, Adrenalin, KHR_coopmat + integer dot):
Remaining gap (unchanged, smaller scope): convrot_group_size 0 (plain rowwise) F32 input still falls back to CPU - there is no group-0 Vulkan quantize shader.
Not yet verified: NVIDIA hardware, Linux.
Related: #6, #7, ggml-org/llama.cpp#26914, leejet/stable-diffusion.cpp#1857.
LLM Disclose: YES, used GLM-5.3 for the grunt work