Skip to content

Pull requests: ggml-org/llama.cpp

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

chat : preserve object alternatives in Qwen XML tool schemas documentation Improvements or additions to documentation testing Everything test related
#28651 opened Sep 9, 2026 by arc-uri-el Loading…
py : bump numpy to 2.4.6 conversion server
#28649 opened Sep 9, 2026 by CISC Member Loading…
webui: stop re-probing disabled /tools endpoint on every message server/ui
#28646 opened Sep 9, 2026 by geckguy Contributor Loading…
1 task done
model: fix all granite family parameter counts model Model specific
#28643 opened Sep 9, 2026 by taronaeo Member Loading…
ggml-cpu: add Q4_0 8x8 gemv/gemm for riscv vlenb=16 case ggml changes relating to the ggml tensor library for machine learning
#28642 opened Sep 9, 2026 by hongyang-7 Loading…
common : reject non-positive batch size (#28525) testing Everything test related
#28639 opened Sep 9, 2026 by goodruyas Loading…
OpenVINO: optimize stateful decode and GPU MoE inference documentation Improvements or additions to documentation ggml changes relating to the ggml tensor library for machine learning OpenVINO
#28638 opened Sep 9, 2026 by wine99 Contributor Loading…
model : fix MTP context kv cache allocation for deepseek2, glm4moe, c… testing Everything test related
#28630 opened Sep 9, 2026 by LoganChu Loading…
hexagon: rope updates ggml changes relating to the ggml tensor library for machine learning Hexagon testing Everything test related
#28628 opened Sep 9, 2026 by tboinovski1 Contributor Loading…
Qwen4 next flash multi gpu & buffer size issues solved CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning model Model specific testing Everything test related
#28623 opened Sep 9, 2026 by gopinath87607 Draft
ui : Accept WEBM video files server/ui
#28622 opened Sep 9, 2026 by EpicEric Loading…
cmake : Use x86-64 microarchitecture level -march when sufficient features are selected ggml changes relating to the ggml tensor library for machine learning
#28621 opened Sep 9, 2026 by bberberov Contributor Loading…
vulkan: use CPU writes in ggml_backend_vk_cpy_tensor_async if the context is idle ggml changes relating to the ggml tensor library for machine learning Vulkan Issues specific to the Vulkan backend
#28618 opened Sep 8, 2026 by jeffbolznv Contributor Loading…
HIP: branch-free SWAR for __vsub4 / __vcmpne4 / __vcmpeq4 CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning
#28616 opened Sep 8, 2026 by SimonTeixidor Contributor Draft
HIP: tune MMVQ batch thresholds on RDNA3.5 CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning
#28613 opened Sep 8, 2026 by SimonTeixidor Contributor Loading…
ggml-vulkan : tune L-tile warp micro-dimension for RDNA3 iGPUs ggml changes relating to the ggml tensor library for machine learning Vulkan Issues specific to the Vulkan backend
#28611 opened Sep 8, 2026 by Troncooooo Loading…
mtmd: add KaniTTS-2 support conversion documentation Improvements or additions to documentation model Model specific mtmd Related to multimodal functionality (video/image/audio)
#28609 opened Sep 8, 2026 by hans00 Draft
mtmd : add Soprano audio generation conversion documentation Improvements or additions to documentation mtmd Related to multimodal functionality (video/image/audio)
#28607 opened Sep 8, 2026 by hans00 Loading…
ggml-cpu(s390x): add Q1_0 vector intrinsic support documentation Improvements or additions to documentation ggml changes relating to the ggml tensor library for machine learning merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge.
#28606 opened Sep 8, 2026 by taronaeo Member Loading…
metal : fix MiniCPM3 crash with automatic flash attention Apple Metal https://en.wikipedia.org/wiki/Metal_(API) ggml changes relating to the ggml tensor library for machine learning testing Everything test related
#28599 opened Sep 8, 2026 by wyanzhao Contributor Loading…
quantize: filter per-layer metadata when pruning layers (#27976) testing Everything test related
#28597 opened Sep 8, 2026 by Sumire-no-kai Loading…
ProTip! Exclude everything labeled bug with -label:bug.