Popular repositories Loading
-
glm53-flash-apple-silicon
glm53-flash-apple-silicon PublicGLM-5.3-Flash on Apple Silicon (M5 Ultra w/256GB): 2089 tok/s Prefill and 55.2 tok/s decode. Fully Ledgered, all tests and attempts included so that it can be developed further.
Python 3
-
local-ai-speed-simulator
local-ai-speed-simulator PublicHow fast would AI run on your own hardware? Single-file simulator for local LLM inference — prefill, decode, KV cache, multi-GPU scaling, electricity vs cloud cost.
HTML
-
DeepSeek-V4-Flash-3x-DGX-Sparks
DeepSeek-V4-Flash-3x-DGX-Sparks PublicForked from FlyCockpit/DeepSeek-V4-Flash-3x-DGX-Sparks
3× NVIDIA DGX Spark runbook for DeepSeek-V4-Flash-0731 — expert parallel + NCCL mesh + DSpark MTP
Python
-
DGX-Mac-Prefill-Sim
DGX-Mac-Prefill-Sim PublicInteractive token speed simulator: DGX Spark clusters vs Mac Studio local LLM prefill and decode realtime sim, for a genuine feel of local AI use.
JavaScript
-
side-buttons
side-buttons PublicMouse side buttons (Back/Forward) for macOS - a tribute to (the no longer supported on modern ARM devices) Sensible Side Buttons
Swift
-
TensorFold
TensorFold PublicForked from chadhurley25075-png/TensorFold
Fast, exact LLM decoding on Apple Silicon (MLX) behind an OpenAI-compatible endpoint
Python
If the problem persists, check the GitHub status page or contact support.