Popular repositories Loading
-
volta-nvfp4
volta-nvfp4 PublicNVFP4 Mixture-of-Experts inference on Tesla V100 (SM70). Fork of dnv2003/v100-skinny extending its QPN2 kernels to fused MoE expert stacks.
-
omlx
omlx PublicForked from jundot/omlx
LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
Python
-
vllm-ultimate-dgx-spark
vllm-ultimate-dgx-spark PublicForked from AEON-7/vllm-ultimate-dgx-spark
AEON vLLM Ultimate — vLLM 0.24.0 built from source for DGX Spark / Blackwell (sm_121a/GB10). One image serves the whole AEON fleet (Gemma-4-26B-A4B, Qwen3.6-27B, Qwen3.6-35B-A3B) with DFlash specul…
Python
-
TensorFold
TensorFold PublicForked from ashhart/TensorFold
Fast, exact LLM decoding on Apple Silicon (MLX) behind an OpenAI-compatible endpoint
Python
If the problem persists, check the GitHub status page or contact support.
