DeepSeek V4 Flash 284B on AMD Strix Halo (gfx1151) — up to 32 tok/s decode & ~250 tok/s prefill via ROCmFPX, DSpark & ROCm 7.2
-
Updated
Jul 24, 2026 - C
DeepSeek V4 Flash 284B on AMD Strix Halo (gfx1151) — up to 32 tok/s decode & ~250 tok/s prefill via ROCmFPX, DSpark & ROCm 7.2
DeepSeek-V4-Flash + DSpark speculative decoding on a pair of NVIDIA DGX Sparks (vLLM TP=2 over RoCE) — tuned recipe, overlays that halve multi-turn TTFT, contamination-guarded benchmarks, ops runbook
Lossless inference speedup benchmarking suite for local LLMs using DeepSeek DSpark speculative draft heads.
Multi-Ensemble Memory-Elastic Token Prediction — a numpy implementation of DeepSeek's DSpark speculative decoding + the MEMTP elastic-ensemble extension, on AMD Strix Halo (gfx1151).
MiniMax M3 NVFP4 on 4x DGX Spark with NVIDIA DSpark, native multi-node vLLM TP=4, reasoning, and tool calling
Add a description, image, and links to the dspark topic page so that developers can more easily learn about it.
To associate your repository with the dspark topic, visit your repo's landing page and select "manage topics."