Skip to content
View gilby's full-sized avatar

Block or report gilby

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Popular repositories Loading

  1. volta-nvfp4 volta-nvfp4 Public

    NVFP4 Mixture-of-Experts inference on Tesla V100 (SM70). Fork of dnv2003/v100-skinny extending its QPN2 kernels to fused MoE expert stacks.

    Python 7 1

  2. omlx omlx Public

    Forked from jundot/omlx

    LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar

    Python

  3. vllm-ultimate-dgx-spark vllm-ultimate-dgx-spark Public

    Forked from AEON-7/vllm-ultimate-dgx-spark

    AEON vLLM Ultimate — vLLM 0.24.0 built from source for DGX Spark / Blackwell (sm_121a/GB10). One image serves the whole AEON fleet (Gemma-4-26B-A4B, Qwen3.6-27B, Qwen3.6-35B-A3B) with DFlash specul…

    Python

  4. TensorFold TensorFold Public

    Forked from ashhart/TensorFold

    Fast, exact LLM decoding on Apple Silicon (MLX) behind an OpenAI-compatible endpoint

    Python