Skip to content
@Yamz-Labs

Yamz Labs

Quantized LLMs and a fast inference engine for AMD Strix Halo

Yamz Labs

Yamz Labs

We run large open models on one small machine: a 128 GB AMD Ryzen AI Max+ 395 (Strix Halo). We publish the engine and EXL3 quantisations of GLM-5.3-Flash and MiMo-V2.6-Flash-MOPD, with quality measured against the official FP8 weights. Every number has a date and a protocol; what we did not measure, we say so. Only gfx1151 with ROCm is tested.

Not affiliated with Z.ai, Xiaomi or turboderp.

Popular repositories Loading

  1. kyojin kyojin Public

    Kyojin: the Yamz inference engine for AMD Strix Halo (ROCm, gfx1151), built on ExLlamaV3. Runs 300B-class MoE models on one 128 GB mini PC.

    Python 91 9

  2. yamz-presets yamz-presets Public

    Python 1 1

  3. .github .github Public

  4. yamzlabs.com yamzlabs.com Public

    HTML

Repositories

Showing 4 of 4 repositories
  • kyojin Public

    Kyojin: the Yamz inference engine for AMD Strix Halo (ROCm, gfx1151), built on ExLlamaV3. Runs 300B-class MoE models on one 128 GB mini PC.

    Yamz-Labs/kyojin's past year of commit activity
    Python 91 MIT 9 8 (2 issues need help) 0 Updated Oct 9, 2026
  • yamz-presets Public
    Yamz-Labs/yamz-presets's past year of commit activity
    Python 1 MIT 1 1 0 Updated Oct 7, 2026
  • yamzlabs.com Public
    Yamz-Labs/yamzlabs.com's past year of commit activity
    HTML 0 0 0 0 Updated Oct 3, 2026
  • .github Public
    Yamz-Labs/.github's past year of commit activity
    0 0 0 0 Updated Oct 2, 2026

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Python HTML

Most used topics

Loading…