-
Notifications
You must be signed in to change notification settings - Fork 53
Pull requests: mudler/vllm.cpp
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
feat(BACKEND-ROCM): kKdaGatedDeltaRule ROCm kernel — the per-K-channel-decay recurrence
#3120
opened Sep 10, 2026 by
localai-org-maint-bot
Collaborator
Loading…
docs(ENG-MM-INPUT-PIPELINE): correct HTTP support claims
#3118
opened Sep 10, 2026 by
localai-org-maint-bot
Collaborator
Loading…
docs: clarify Qwen3.8 benchmark limits
#3104
opened Sep 9, 2026 by
localai-org-maint-bot
Collaborator
Loading…
fix(ENG-QWEN35-FULL-ATTN-STATE): validate state only for GDN consumers
#3101
opened Sep 9, 2026 by
VikashLoomba
Contributor
•
Draft
feat(BACKEND-ROCM-QUANT-GATHER): gather packed embeddings on ROCm
#3097
opened Sep 9, 2026 by
VikashLoomba
Contributor
•
Draft
4 tasks done
feat(BACKEND-ROCM-BF16-MOE): run BF16 grouped experts on ROCm
#3096
opened Sep 9, 2026 by
VikashLoomba
Contributor
•
Draft
3
feat(BACKEND-ROCM-F16-WEIGHTS): retain F16 dense weights on ROCm
#3095
opened Sep 9, 2026 by
VikashLoomba
Contributor
•
Draft
2
fix(BACKEND-GATE-ROCM-VLLM): validate the pinned Strix oracle securely
#3052
opened Sep 8, 2026 by
localai-org-maint-bot
Collaborator
•
Draft
6 tasks done
docs(KV-WARMUP-PROFILE): specify startup memory profiling
#3050
opened Sep 8, 2026 by
VikashLoomba
Contributor
•
Draft
perf(KERNEL-QUANT-CIQ-GEMM-ROCM): cooperative-tile WMMA kernels — Shared rejected, BigTile accepted
#3036
opened Sep 7, 2026 by
joral
Contributor
Loading…
feat(KERNEL-QUANT-CIQ-GEMM-ROCM-IQUANT): port IQ4_XS/IQ3_XXS to ROCm
#3029
opened Sep 6, 2026 by
joral
Contributor
Loading…
perf(BACKEND-ROCM): widen sampling block to 1024 and split-phase random sample
#3010
opened Sep 6, 2026 by
ghazni101
Contributor
Loading…
fix(MODEL-TEXT-GLM4-MOE-LITE-GATE-2839): the near-tie predicate cannot fire, so apply the bar the oracle capture licenses -- and it FAILS 69/128
#2906
opened Sep 4, 2026 by
localai-org-maint-bot
Collaborator
Loading…
perf(GFX1100-TG200): T14 row-split greedy argmax arm
#2876
opened Sep 4, 2026 by
ghazni101
Contributor
Loading…
perf(GFX1100-TG200): T9 cooperative gated norm arm
#2875
opened Sep 4, 2026 by
ghazni101
Contributor
Loading…
perf(GFX1100-TG200): T8 cooperative single-row rmsnorm arm
#2874
opened Sep 4, 2026 by
ghazni101
Contributor
Loading…
perf(GFX1100-TG200): T6b cooperative attn preamble arm
#2868
opened Sep 4, 2026 by
ghazni101
Contributor
Loading…
perf(GFX1100-TG200): T6a cooperative GDN scan arm
#2866
opened Sep 4, 2026 by
ghazni101
Contributor
Loading…
feat(GFX1100-TG200): T25 keep ssm_out as Q5_K with runtime input permutation
#2807
opened Sep 3, 2026 by
ghazni101
Contributor
Loading…
feat(GFX1100-TG200): T21 keep-quant for V-head row-permuted GDN projections
#2804
opened Sep 3, 2026 by
ghazni101
Contributor
Loading…
perf(GFX1100-TG200): T16 YTILE=4 default for wvSplitK decode-skinny GEMV
#2787
opened Sep 3, 2026 by
ghazni101
Contributor
Loading…
perf(GFX1100-TG200): T2b flips ROCm support_static_graph_mode
#2777
opened Sep 3, 2026 by
ghazni101
Contributor
Loading…
feat(MODEL-MM-GLM53-FLASH-KPOOL-CUDA): give GLM-5.3-Flash's k-pool indexer a device it can run on
#2432
opened Aug 31, 2026 by
localai-org-maint-bot
Collaborator
Loading…
docs: align GLM-5.3 public status
#2374
opened Aug 31, 2026 by
localai-org-maint-bot
Collaborator
Loading…
feat(BACKEND-VULKAN-TQ1_0): TQ1_0 ternary keep-quant matmul, MoE, and rope shaders for Vulkan
#2248
opened Aug 29, 2026 by
phantomic12
Loading…
Previous Next
ProTip!
Type g i on any issue or pull request to go back to the issue listing page.