Skip to content

feat(quantization): share the int8_convrot scheme and read it for Krea-2 and Z-Image - #194

Draft
Pfannkuchensack wants to merge 5 commits into
mainfrom
feat/int8-convrot-shared
Draft

feat(quantization): share the int8_convrot scheme and read it for Krea-2 and Z-Image#194
Pfannkuchensack wants to merge 5 commits into
mainfrom
feat/int8-convrot-shared

Conversation

@Pfannkuchensack

@Pfannkuchensack Pfannkuchensack commented Sep 1, 2026

Copy link
Copy Markdown
Member

Summary

Bug fix, plus a memory saving. int8_tensorwise is a ComfyUI-wide quantization scheme, not one architecture's format — but the implementation lived in backend/minimax_h3/, where no other loader could reach it. This moves it to backend/quantization/ beside gguf, sdnq and bnb, and teaches Krea-2 and Z-Image to read it. The weights stay int8-resident: each quantized nn.Linear becomes an Int8ConvrotLinear that dequantizes and derotates per forward, as MiniMax H3 already does.

Neither architecture could load such a checkpoint, and they failed differently:

  • Krea-2 failed silently. _dequantize_scaled_fp8 keys off .weight_scale alone, so it scaled the int8 weights and never un-rotated them. No error, no NaN, no warning — correlation against a correct decode is 0.06. It loads, and it generates noise.
  • Z-Image failed loudly. Its fused-QKV split carried the weight but dropped the scale and the marker under the fused name, so 408 orphaned keys reached a strict load_state_dict.

Z-Image is the model this format is useful for. Krea-2 is 12.6 GiB as int8 — still too big for the 12 GB cards that fall back to GGUF today. Z-Image is 5.8 GiB. Measured at 1024px: 11.74 GiB resident as bf16, 6.04 GiB as int8, at 6 % more time, with the same image.

What it is worth

dense decode int8-resident (this PR)
load time 61.3 s 0.2 s
resident 24 GB bf16, or 12.3 GB only with the fp8-storage opt-in 12.3 GB, always
needs fp8 storage? yes, or it does not fit a 24 GB card no
8 steps at 512px 8.4 s 9.5 s
peak allocated / reserved 13.38 / 13.69 GiB 13.41 / 13.68 GiB

The load time is not a trick: safetensors maps lazily, and with no decode at load there is nothing to read. The 13 % on the denoise is the per-forward dequantize and derotation — and it does not grow with resolution, because the derotation is a fixed [out, in/256, 256] @ [256, 256] matmul while the layer's own matmul scales with sequence length. At 512px it is ~12 % of the layer's matmul; at 1024px, ~3 %.

Why resident rather than a dense decode. Measured against enabled fp8 storage the two are equal, both at 12.0 GiB — but _should_use_fp8 requires default_settings.fp8_storage is True, which is off by default and unavailable outside CUDA/XPU. Against the state a user actually gets, it is 12.3 GB versus 24 GB — and on Apple Silicon there is no fp8 alternative at all.

Ordering is load-bearing three times over, and none of it fails loudly

  • Before the fp8 fold, which would scale an int8 weight without un-rotating it.
  • Before the native→diffusers key conversion, which renames .attn.wq.weight by substring — carrying .weight_scale along but orphaning .comfy_quant. Module paths are therefore resolved by following the weight's own rename, not the marker's key.
  • Before the encoder's fp8 detection, which answers yes to any .weight_scale and would otherwise keep an int8 encoder "fp8-resident" over weights that were never fp8.

Each is pinned by a test that shows the damage rather than asserting the order.

Three things the real checkpoints forced

A bug in the existing key converter. last.linear was renamed by exact match per suffix (k == "last.linear.weight") while everything else uses prefix slicing. last.linear.weight_scale fell through every rule and kept its old name. This was invisible until now because both the fp8 path and a dense int8 decode consume the scales before the rename. Now a prefix rule, with a test over all ten rename rules.

input_scale. One Qwen3-VL repack ships 337 activation scales for W8A8 inference. This code dequantizes the weight and computes in bf16, so there is nothing to apply them to — and the existing filter spelled it scale_input, which never matched. Both spellings are now dropped, in both load paths.

LoRA. model_is_quantized checked only GGUFQuantized. An int8 Krea-2 is a plain checkpoint as far as the config knows, so LayerPatcher would have written directly into int8 buffers. Extracted as requires_sidecar_patching() and tested on its own — it was previously inline in a 600-line method and unreachable by any test.

Z-Image: the fused QKV

attention.qkv.weight is split into to_q/to_k/to_v. The split already handled the weight; the scale and the marker were left behind under a module name the model does not have. The per-output-channel scale now splits with its weight and the marker is copied to all three.

That decision is made from the key suffix, not the tensor shape — a 72-byte JSON marker is also divisible by three, and a shape-based rule cuts it into three fragments of broken JSON. A test asserts the fixture reproduces that trap.

Z-Image's converter carries .comfy_quant onto the final module names, so its markers are read after the conversion and need no re-keying — the opposite of Krea-2, whose converter orphans them. Both arrangements are pinned by tests.

Also added: an int8 weight with no marker is refused at load. It would otherwise be handed to a float Linear and fail only at forward time, if at all.

Related Issues / Discussions

This is the storage side of int8: the checkpoint loads correctly and the weights stay int8-resident, with the math in bf16. Running the int8 tensor cores is a separate matter and is not part of this PR.

Note for whoever merges v6 into v7: invoke-ai#9478 touches model_loaders/krea2.py (+261/−69) and test_krea2_state_dict_utils.py (+87/−43) — the same two files this PR touches.

QA Instructions

Verified against four real checkpoints on RTX 4090, torch 2.7.1+cu128, Windows.

Key coverage, driven through the real loader pipeline on meta tensors (no weight data allocated):

file result
Krea2_Turbo_convrot_int8mixed.safetensors (12.0 GiB, 264 quantized layers) 430/430 parameters, 0 missing, 0 extra, 0 shape mismatch, 0 orphaned int8 tensors
Krea2_Turbo_fp8mixed.safetensors (12.0 GiB) same — the fp8 path is unaffected
qwen3vl_4b_int8.safetensors (no convrot, 199 unquantized bf16 weights) 357 swapped, 0 missing, 0 extra
qwen3vl_4b_uncensored_int8_convrot.safetensors (convrot, 219 unquantized) 337 swapped, 0 missing, 0 extra
z_image_turbo_int8_convrot_bf16emixed.safetensors (5.75 GiB, 209 unquantized bf16) 272 modules after the qkv split, 0 orphaned, 0 missing, 0 extra

Numerics. Z-Image ships an unquantized bf16 release, so its decode is checked against ground truth rather than a second quantization. Over 12 layers spanning depth and both refiner stacks:

  • decoded: corr 0.99978–1.00000
  • without the un-rotation (negative control): corr 0.054–0.065
  • relative error: 0.85–1 % — the int8 error alone

to_q, to_k and to_v all land at 0.9998+, which tests the fused split, the scale split and the rotation together.

For Krea-2 no unquantized release was to hand, so its decode was checked against the fp8 build of the same repack (confirmed first to be the same weights: 159 of 166 unquantized tensors bit-identical, the 7 that differ all biases of quantized layers). Over 14 layers: decoded corr 0.997–0.9999, negative control 0.06–0.10, relative error a uniform 2.8 % with a structureless residual (corr to the reference ±0.009). That 2.8 % is mostly the reference's own error — the fp8 build carries one scale for the whole matrix where int8 carries one per output channel, and against true bf16 ground truth the same decode measures 0.9 %.

Generation. Same prompt, same seed, 8 steps. Z-Image int8 against the bf16 release at 1024px: the same photograph, mean absolute difference 8.61/255, correlation 0.950 — against 81/255 and −0.0004 for random noise. Krea-2 int8 against fp8 at 512px: likewise the same photograph, 12.25/255, correlation 0.948. A wrong rotation does not fail, it produces noise, so this is the end-to-end form of the negative control.

The resident path agrees with a dense decode at corr ≥ 0.999883, the difference being bf16 rather than fp32 arithmetic.

To verify by hand: load a *_int8_convrot.safetensors Krea-2 and generate. Then load the fp8 build of the same model at the same seed — the images should be the same picture. Before this PR the int8 file produced noise.

Tests: ruff check / format --check clean. No openapi.json change. Eleven mutations were verified as caught — seven on the Krea-2 side (removing either swap, letting the fp8 fold eat int8 scales, restoring the last.linear exact match, neutralising the scale-layout guard, neutralising the sidecar detection, no longer dropping activation scales) and four on Z-Image (removing the swap, not carrying the marker through the qkv split, deciding the split by shape again, removing the orphan guard).

Not verified: a mixed-precision Krea-2 transformer specifically. The repack tested here quantizes all 266 weights despite "mixed" in its name. The Z-Image checkpoint does exercise that path on a real transformer (209 unmarked bf16 weights beside 204 quantized), as do both Qwen3-VL encoders — the decision is made per tensor rather than per file — but no Krea-2 file to hand covers it.

Merge Plan

No DB schema, no redux slice, no API change, no dependency change. Expect a conflict in krea2.py against the v6→v7 merge that brings invoke-ai#9478.

Checklist

  • The PR has a short but descriptive title, suitable for a changelog
  • Tests added / updated (if applicable)
  • ❗Changes to a redux slice have a corresponding migration — n/a
  • Documentation added / updated (if applicable) — module docstrings; the scheme's framing changed from "H3's format" to "a ComfyUI-wide scheme"
  • Updated What's New copy (if doing a release after this PR)

🤖 Generated with Claude Code

Pfannkuchensack and others added 2 commits September 1, 2026 03:31
… read it

`int8_tensorwise` is a ComfyUI-wide scheme, not one architecture's format, but the
implementation lived in `backend/minimax_h3/`. Krea-2's loaders could not reach it, so
`*_int8_convrot.safetensors` loaded as confident nonsense: `_dequantize_scaled_fp8`
keys off `.weight_scale` alone, scaled the int8 weights and never un-rotated them.
No error, no NaN -- correlation against a correct decode is 0.06.

Moves `int8_convrot.py` to `backend/quantization/` beside the other schemes (gguf,
sdnq, bnb) and wires both Krea-2 loaders to it. The weights stay int8-resident: each
quantized Linear becomes an `Int8ConvrotLinear` that dequantizes and derotates per
forward, as MiniMax H3 already does. That is 12.3 GB instead of 24 GB, on every
platform and without the fp8-storage opt-in -- which is off by default and
unavailable outside CUDA/XPU.

Ordering is load-bearing three times over, and none of it fails loudly:

- before the fp8 fold, which would scale an int8 weight without un-rotating it;
- before the native->diffusers key conversion, which renames `.attn.wq.weight` by
  substring, carrying `.weight_scale` along but orphaning `.comfy_quant`;
- before the encoder's fp8 detection, which answers yes to any `.weight_scale` and
  would keep an int8 encoder "fp8-resident" over weights that were never fp8.

Each is pinned by a test that shows the damage rather than asserting the order.

Three things the real checkpoints forced. `last.linear` was renamed by exact match
per suffix while everything else uses prefix slicing, so `last.linear.weight_scale`
kept its old name -- invisible until now, because both the fp8 path and a dense int8
decode consume the scales before the rename. One Qwen3-VL repack ships 337
`input_scale` activation scales, under a spelling the existing filter did not match.
And `model_is_quantized` checked only the config format, so LoRA would have been
written directly into int8 buffers; it is now `requires_sidecar_patching()`, which
consults the module tree and can be tested on its own.

Verified against four real checkpoints: 430/430 key coverage for the int8 and fp8
Krea-2 builds and 357/337 swapped layers for the two Qwen3-VL encoders, all with no
missing, extra or orphaned tensors; weights at corr 0.997-0.9999 against an
independent quantization of the same model (0.06-0.10 without the un-rotation); and
a generation that produces the same photograph as the fp8 build.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Pfannkuchensack and others added 3 commits September 1, 2026 14:14
Z-Image is the model this format is actually useful for. Krea-2 is 12.6 GiB as
int8 -- too big for the 12 GB cards that fall back to GGUF today -- while Z-Image
is 5.8 GiB, and measured here it drops a 1024px generation from 11.74 GiB resident
to 6.04 GiB at 6 % more time.

Moves the architecture-agnostic half of the Krea-2 work into
`backend/quantization/int8_convrot.py` (resolve_quantized_module_paths,
swap_in_int8_linears, cast_unquantized, drop_unconsumed_quantization_sidecars) so
Z-Image imports it rather than copying it. No behaviour change.

Z-Image's own hazard is the fused QKV: `attention.qkv.weight` is split into
to_q/to_k/to_v, and the split already handled the weight but dropped the scale and
the marker under the fused name -- 408 keys that then reached a strict
`load_state_dict`. The per-output-channel scale now splits with its weight and the
marker is copied to all three. That decision is made from the key suffix, not the
tensor shape: a 72-byte JSON marker is also divisible by three, and cutting it into
thirds yields three fragments of broken JSON.

Because this converter carries `.comfy_quant` onto the final module names, the
markers are read after the conversion and need no re-keying -- the opposite of
Krea-2, whose converter orphans them.

Adds a guard for an int8 weight with no marker: it would be handed to a float
Linear and fail only at forward time, if at all.

Verified against the real 5.75 GiB checkpoint and the bf16 diffusers release of the
same model: 272 quantized modules with no orphans, no missing and no extra keys;
weights at corr 0.99978-1.00000 against unquantized ground truth (0.054-0.065
without the un-rotation) at 0.85-1 % relative error; and a generation that produces
the same photograph.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@Pfannkuchensack Pfannkuchensack changed the title feat(quantization): share the int8_convrot scheme and teach Krea-2 to read it feat(quantization): share the int8_convrot scheme and read it for Krea-2 and Z-Image Sep 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant