GLM-5.3-Flash on 1x RTX 5090 + 2x DGX Spark: attention on the 5090, routed experts on the Sparks over MCDMA RoCE, on TensorFold v0.6.5 (+ 0.6.6's commits) with Mia's GLM work. 2.0 at tag v2.0, v1.0 (glm53f-afd) at tag v1.0. Deployment recipe, as-is.
-
Updated
Oct 8, 2026 - Shell