Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
-
Updated
Oct 2, 2026 - Go
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
Adaptive Disaggregated Inference on a Role-free Fleet!
A lightweight LLM inference runtime, evolving toward heterogeneous, stage-disaggregated serving.
Prefill/Decode-disaggregated LLM serving library in Go — KV-cache connectors + SLO-aware elastic GPU flipping; published benchmark suite (colocated TPOT p99 inflates 133x under burst)
RPC distributed prefill + slot-state PD disaggregation PoC for llama.cpp
To associate your repository with the pd-disaggregation topic, visit your repo's landing page and select "manage topics."