Run open LLMs such as Qwen and Gemma locally on 1, 2 or 3 NVIDIA DGX Sparks (GB10) with vLLM: tested recipes, one setup script, and multi-Spark tensor parallelism over direct QSFP cables.
What to do first, how to cable two or three Sparks, and which recipe to pick, with measured speeds.
- NVIDIA-DGX-Spark-LLM-Setup: the start-here guide
- Qwen3.8-Flash-Next-DGX-Spark-TP1-TP3: Qwen3.8-Flash-Next on 1–3 Sparks, up to 1M context
- Gemma-4-31B-IT-DGX-Spark-TP1-TP2: Gemma-4-31B-IT on 1–2 Sparks, with draft-model speculative decoding
- dgx-spark-recipe-kit: the shared setup, benchmark and recipe template inside every recipe