Serve GLM-5.2 469B (REAP-pruned, NVFP4) across 3× NVIDIA DGX Spark with vLLM pipeline parallelism — 256K context, production-ready config and patches
-
Updated
Jul 12, 2026 - Python
Serve GLM-5.2 469B (REAP-pruned, NVFP4) across 3× NVIDIA DGX Spark with vLLM pipeline parallelism — 256K context, production-ready config and patches
Qwen3.8-Flash-Next on 4x RTX 3090 with vLLM: 806,792 kv pool, 160-175tok/s single decode, 3 x 262K sessions resident, MTP, host-mapped PLE, pinned build and container recipes.
Swift 1.5 post-train of Qwen3.8-Flash-Next on 4x RTX 3090 with vLLM: W4A16 checkpoint on the unchanged Flash-Next v2.2.0 engine, TP2 x PP2 + EP, 806,792-token FP8 KV pool, 262K context, MTP K=3, recalibrated KV scales.
Qwen3.8 Flash Next FP8 on 4x CMP 170HX: PP3 transformer stages with the 51B PLE table on GPU 3
A ~800-line PyTorch implementation of Megatron-LM's TP + PP + DP + AMP. 1.6-2x faster than Megatron-Core on 125M models.
Xe2 dual-B70 kernel and 2x2 parallelism lab (TP=2, PP=2)
FSDP · DDP · Pipeline · Tensor parallel on 4× NVIDIA A30 — real throughput benchmarks for Qwen2.5-7B multi-GPU training
JAX/Flax NNX training infrastructure: device meshes and SPMD sharding, an Orbax checkpoint store, early stopping and callbacks, W&B and MLflow logging.
To associate your repository with the pipeline-parallel topic, visit your repo's landing page and select "manage topics."