forked from GenerelSchwerz/llama.cpp
-
Notifications
You must be signed in to change notification settings - Fork 0
Pull requests: Piggidragon/llama.cpp
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
ggml-meta, sched : allocate a meta transport ring at each device's share
documentation
Improvements or additions to documentation
ggml
#9
opened Sep 17, 2026 by
Piggidragon
Owner
Loading…
sched, ggml-meta : pipeline the host KV delivery under split mode tensor
documentation
Improvements or additions to documentation
ggml
testing
#8
opened Sep 17, 2026 by
Piggidragon
Owner
Loading…
llama : add an attention split separate from the tensor split
testing
#7
opened Sep 17, 2026 by
Piggidragon
Owner
Loading…
dflash : run DFlash2 under split mode tensor
ggml
model
#6
opened Sep 17, 2026 by
Piggidragon
Owner
Loading…
ggml-meta : split a host-resident KV cache by head
devops
documentation
Improvements or additions to documentation
ggml
testing
#5
opened Sep 17, 2026 by
Piggidragon
Owner
Loading…
server : allow cache reuse for text-only prompts with mmproj loaded
server
#4
opened Sep 17, 2026 by
Piggidragon
Owner
Loading…
llama : split a tied output projection under split mode tensor
testing
#3
opened Sep 17, 2026 by
Piggidragon
Owner
Loading…
ggml : report allocation failure from the meta buffer type
ggml
testing
#2
opened Sep 17, 2026 by
Piggidragon
Owner
Loading…
ProTip!
Filter pull requests by the default branch with base:llama-tensor.