SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving
-
Updated
Sep 28, 2026 - Python
SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving
(SMG) Shepherd Model Gateway official documentation, designed by studio-noiich.
Orchestrate LLM inference across your entire fleet. Run vLLM, SGLang, TensorRT-LLM, llama.cpp, or MLX (Apple Silicon) as a coordinated multi-node cluster (KV-aware, disaggregation-ready, energy-aware) on bare metal, Kubernetes, or any major cloud, installed with a single curl line.
RPC distributed prefill + slot-state PD disaggregation PoC for llama.cpp
Prefill/decode disaggregation for LLM serving on RBLN-CA25 NPUs (vllm_rbln optimum path) + heterogeneous GPU→NPU KV bridge (dlfloat16, verified cos=0.9997). Deployable kits with paste-in runbooks.
To associate your repository with the prefill-decode-disaggregation topic, visit your repo's landing page and select "manage topics."