ggml-cpu: add Q4_0 8x8 gemv/gemm for riscv vlenb=16 case - #28642
Open
hongyang-7 wants to merge 1 commit into
Open
ggml-cpu: add Q4_0 8x8 gemv/gemm for riscv vlenb=16 case#28642hongyang-7 wants to merge 1 commit into
hongyang-7 wants to merge 1 commit into
Conversation
|
Hi @hongyang-7, thanks for your contribution! Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:
Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below. |
hongyang-7
marked this pull request as ready for review
September 9, 2026 11:21
Author
|
Add Requirements section to respect PR Template. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This PR extends existing riscv gemv/gemm kernels for q4_0_q8_0 in 8x8 layout for vlenb=16.
Key Changes
Perplexity Test
Performance Test
Performance is tested on a real high-performance RISC-V server.
numactl -C 0-15 -m 0 llama-batched-benchLlama-3.2-1B-Instruct-Q4_0.gguf (738M)
(master)
(opt)
(opt/master)
(master)
(opt)
(opt/master)
PP: 4.3x speedup
TG: 1.2x~3.0x speedup
Meta-Llama-3-8B-Instruct.Q4_0.gguf (4.4G)
(master)
(opt)
(opt/master)
(master)
(opt)
(opt/master)
PP: 4.7x speedup
TG: 1.4x~4.0x speedup
Future Work
#if defined __riscv_zvfhpath from top ggml-cpu/repack.cppRequirements