Skip to content

llama: GPU-resident LRU cache for host-offloaded MoE expert weights - #27861

Draft
csantiago78 wants to merge 1 commit into
ggml-org:masterfrom
csantiago78:moe-expert-cache
Draft

llama: GPU-resident LRU cache for host-offloaded MoE expert weights#27861
csantiago78 wants to merge 1 commit into
ggml-org:masterfrom
csantiago78:moe-expert-cache

MoE expert cache: GPU-resident LRU cache for host-offloaded expert we…

bccbacd
Select commit
Loading
Failed to load commit list.
Sign in for the full log view