Orchestrate LLM inference across your entire fleet. Run vLLM, SGLang, TensorRT-LLM, llama.cpp, or MLX (Apple Silicon) as a coordinated multi-node cluster (KV-aware, disaggregation-ready, energy-aware) on bare metal, Kubernetes, or any major cloud, installed with a single curl line.
inference tensorrt kv-cache llm vllm sglang prefill-benchmarking llm-cluster prefill-decode-disaggregation
-
Updated
Sep 23, 2026 - Rust