Skip to content

Pull requests: GenerelSchwerz/llama.cpp

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

speculative : capped recurrent planes for DFlash (--spec-draft-rs-planes) documentation Improvements or additions to documentation server testing
#91 opened Sep 14, 2026 by Piggidragon Loading…
cuda: opt-in quantized-native MMA FlashAttention CUDA devops documentation Improvements or additions to documentation ggml testing
#85 opened Sep 10, 2026 by Piggidragon Loading…
kv-cache : resolve the partial KV residency set once for the model documentation Improvements or additions to documentation examples testing
#67 opened Sep 3, 2026 by Piggidragon Loading…
ggml-cuda : reuse routing IDs across MoE siblings
#42 opened Aug 26, 2026 by GenerelSchwerz Owner Draft
1 task done
sched: pipeline the delivery of a host-resident KV cache documentation Improvements or additions to documentation examples ggml testing
#39 opened Aug 26, 2026 by Piggidragon Loading…
2 tasks
kv: replace eligible dense causal masks with compact prefixes CUDA documentation Improvements or additions to documentation ggml server testing
#7 opened Aug 21, 2026 by GenerelSchwerz Owner Loading…
ProTip! Follow long discussions with comments:>50.