forked from ggml-org/llama.cpp
-
Notifications
You must be signed in to change notification settings - Fork 17
Pull requests: GenerelSchwerz/llama.cpp
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
speculative : capped recurrent planes for DFlash (--spec-draft-rs-planes)
documentation
Improvements or additions to documentation
server
testing
#91
opened Sep 14, 2026 by
Piggidragon
Loading…
ggml-alloc : share meta compute buffers through a borrower alias
ggml
testing
#88
opened Sep 13, 2026 by
Piggidragon
Loading…
qwen4exp: prefetch requested lazy PLE pages (upstream reference)
model
#86
opened Sep 10, 2026 by
GenerelSchwerz
Owner
•
Draft
cuda: opt-in quantized-native MMA FlashAttention
CUDA
devops
documentation
Improvements or additions to documentation
ggml
testing
#85
opened Sep 10, 2026 by
Piggidragon
Loading…
cuda: Windows MoE pin budget - layer-subset routing + chunked pinned store
CUDA
ggml
testing
#84
opened Sep 10, 2026 by
jasonlnheath
Loading…
cuda/moe-cache: pinned-alloc fallback puts experts on the CPU backend
CUDA
ggml
#83
opened Sep 10, 2026 by
LokenSI
Loading…
cuda : add experimental bounded MoE host pinning
CUDA
documentation
Improvements or additions to documentation
ggml
server
testing
#76
opened Sep 8, 2026 by
GenerelSchwerz
Owner
•
Draft
kv-cache : spend the partial residency budget on the slowest link first
documentation
Improvements or additions to documentation
examples
testing
#68
opened Sep 3, 2026 by
Piggidragon
Loading…
kv-cache : resolve the partial KV residency set once for the model
documentation
Improvements or additions to documentation
examples
testing
#67
opened Sep 3, 2026 by
Piggidragon
Loading…
llama-bench placement flags, two host-KV fixes, and H1/H2/H3/H5/H13 measured
documentation
Improvements or additions to documentation
examples
ggml
#62
opened Sep 2, 2026 by
Piggidragon
•
Draft
ggml-cuda: coalesce contiguous staging copies
#46
opened Aug 26, 2026 by
GenerelSchwerz
Owner
•
Draft
ggml-cuda: prefetch overflow expert siblings
#45
opened Aug 26, 2026 by
GenerelSchwerz
Owner
•
Draft
ggml-cuda: reuse cached experts during overflow staging
#44
opened Aug 26, 2026 by
GenerelSchwerz
Owner
•
Draft
ggml-cuda: skip duplicate decode sibling prefetch
#43
opened Aug 26, 2026 by
GenerelSchwerz
Owner
•
Draft
ggml-cuda : reuse routing IDs across MoE siblings
#42
opened Aug 26, 2026 by
GenerelSchwerz
Owner
•
Draft
1 task done
sched: pipeline the delivery of a host-resident KV cache
documentation
Improvements or additions to documentation
examples
ggml
testing
#39
opened Aug 26, 2026 by
Piggidragon
Loading…
2 tasks
kv: replace eligible dense causal masks with compact prefixes
CUDA
documentation
Improvements or additions to documentation
ggml
server
testing
#7
opened Aug 21, 2026 by
GenerelSchwerz
Owner
Loading…
ProTip!
Follow long discussions with comments:>50.