-
Notifications
You must be signed in to change notification settings - Fork 22.9k
Pull requests: ggml-org/llama.cpp
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
chat : preserve object alternatives in Qwen XML tool schemas
documentation
Improvements or additions to documentation
testing
Everything test related
#28651
opened Sep 9, 2026 by
arc-uri-el
Loading…
webui: stop re-probing disabled /tools endpoint on every message
server/ui
#28646
opened Sep 9, 2026 by
geckguy
Contributor
Loading…
1 task done
model: fix all granite family parameter counts
model
Model specific
#28643
opened Sep 9, 2026 by
taronaeo
Member
Loading…
ggml-cpu: add Q4_0 8x8 gemv/gemm for riscv vlenb=16 case
ggml
changes relating to the ggml tensor library for machine learning
#28642
opened Sep 9, 2026 by
hongyang-7
Loading…
common : reject non-positive batch size (#28525)
testing
Everything test related
#28639
opened Sep 9, 2026 by
goodruyas
Loading…
OpenVINO: optimize stateful decode and GPU MoE inference
documentation
Improvements or additions to documentation
ggml
changes relating to the ggml tensor library for machine learning
OpenVINO
#28638
opened Sep 9, 2026 by
wine99
Contributor
Loading…
model : fix MTP context kv cache allocation for deepseek2, glm4moe, c…
testing
Everything test related
#28630
opened Sep 9, 2026 by
LoganChu
Loading…
hexagon: rope updates
ggml
changes relating to the ggml tensor library for machine learning
Hexagon
testing
Everything test related
#28628
opened Sep 9, 2026 by
tboinovski1
Contributor
Loading…
Qwen4 next flash multi gpu & buffer size issues solved
CUDA
Related to the CUDA backend
ggml
changes relating to the ggml tensor library for machine learning
model
Model specific
testing
Everything test related
#28623
opened Sep 9, 2026 by
gopinath87607
•
Draft
cmake : Use x86-64 microarchitecture level changes relating to the ggml tensor library for machine learning
-march when sufficient features are selected
ggml
#28621
opened Sep 9, 2026 by
bberberov
Contributor
Loading…
vulkan: use CPU writes in ggml_backend_vk_cpy_tensor_async if the context is idle
ggml
changes relating to the ggml tensor library for machine learning
Vulkan
Issues specific to the Vulkan backend
#28618
opened Sep 8, 2026 by
jeffbolznv
Contributor
Loading…
HIP: branch-free SWAR for __vsub4 / __vcmpne4 / __vcmpeq4
CUDA
Related to the CUDA backend
ggml
changes relating to the ggml tensor library for machine learning
#28616
opened Sep 8, 2026 by
SimonTeixidor
Contributor
•
Draft
HIP: tune MMVQ batch thresholds on RDNA3.5
CUDA
Related to the CUDA backend
ggml
changes relating to the ggml tensor library for machine learning
#28613
opened Sep 8, 2026 by
SimonTeixidor
Contributor
Loading…
chat : keep DSML markup out of DeepSeek V3.2 string tool arguments
testing
Everything test related
#28612
opened Sep 8, 2026 by
iamganesha
•
Draft
ggml-vulkan : tune L-tile warp micro-dimension for RDNA3 iGPUs
ggml
changes relating to the ggml tensor library for machine learning
Vulkan
Issues specific to the Vulkan backend
#28611
opened Sep 8, 2026 by
Troncooooo
Loading…
mtmd: add KaniTTS-2 support
conversion
documentation
Improvements or additions to documentation
model
Model specific
mtmd
Related to multimodal functionality (video/image/audio)
mtmd : add Soprano audio generation
conversion
documentation
Improvements or additions to documentation
mtmd
Related to multimodal functionality (video/image/audio)
#28607
opened Sep 8, 2026 by
hans00
Loading…
ggml-cpu(s390x): add Q1_0 vector intrinsic support
documentation
Improvements or additions to documentation
ggml
changes relating to the ggml tensor library for machine learning
merge ready
A maintainer can use this label to indicate that they consider the changes final and ready to merge.
#28606
opened Sep 8, 2026 by
taronaeo
Member
Loading…
metal : fix MiniCPM3 crash with automatic flash attention
Apple Metal
https://en.wikipedia.org/wiki/Metal_(API)
ggml
changes relating to the ggml tensor library for machine learning
testing
Everything test related
#28599
opened Sep 8, 2026 by
wyanzhao
Contributor
Loading…
json-schema-to-grammar : report incomplete-conversion warnings only once
#28598
opened Sep 8, 2026 by
RIVOIRA
Loading…
quantize: filter per-layer metadata when pruning layers (#27976)
testing
Everything test related
#28597
opened Sep 8, 2026 by
Sumire-no-kai
Loading…
Previous Next
ProTip!
Exclude everything labeled
bug with -label:bug.