We run large open models on one small machine: a 128 GB AMD Ryzen AI Max+ 395 (Strix Halo). We publish the engine and EXL3 quantisations of GLM-5.3-Flash and MiMo-V2.6-Flash-MOPD, with quality measured against the official FP8 weights. Every number has a date and a protocol; what we did not measure, we say so. Only gfx1151 with ROCm is tested.
- Engine: kyojin, built on ExLlamaV3, with a ROCm path for gfx1151.
- Models on Hugging Face:
Not affiliated with Z.ai, Xiaomi or turboderp.
