Summary
The docs cover weight quantization well — TurboQuant KV-cache compression and the
Q4_0/Q4_K/Q6_K/Q8_0 SIMD matmul kernels — but there is no single article that explains the
end-to-end quantization process, and in particular activation post-training quantization
(PTQ) with calibration, which is what int8 NPU/accelerator targets need.
This issue tracks adding an explanation/ article that ties the picture together and documents the
calibration pattern.
Motivation
Two quantization regimes coexist in SKaiNET and are easy to conflate:
- Weight-only quantization (TurboQuant / GGUF
Q8_0/Q4_K, QuantizedMatmul): weights are
stored in low-bit blocks and dequant-fused against fp32 activations at matmul time. This is
eager and already documented.
- Activation PTQ (int8 accelerators): both operands of the matmul are int8, with an i32
accumulator and a requantize epilogue. This needs calibration — representative activation
ranges — and a different lowering. It is not documented and is not what QuantizedMatmul
does.
A contributor targeting an int8 NPU currently has no map of which piece does what, or why
QuantizedMatmul cannot be reused for the activation-int8 path.
What the article should cover
- The two regimes side by side (weight quant vs activation PTQ) and when each applies.
- The PTQ pipeline: calibrate → per-tensor symmetric scales → quantize / i8 matmul (i32) /
requantize → int8 codegen.
- The calibration mechanism: tapping the eager execution stream to capture per-tensor
activation ranges (max|x|) over representative inputs, via an ExecutionObserver /
TensorOps-decorator on matmul.
- The honest limitation finding:
QuantizedMatmul is eager, weight-only and
TensorOps.matmul is dtype-homogeneous (Tensor<T> × Tensor<T> → Tensor<T>, no i8×i8→i32
form, no round), so the int8 activation-matmul pattern cannot be expressed at the tape level
today — the int8 emission belongs in backend codegen, fed by SKaiNET-side calibration.
- Where each responsibility lives (calibration in the framework, int8 emission in the target
backend) and the per-matmul scale set the emitter consumes (Sa, Sb, So).
Out of scope (this issue)
Implementing a tape-level int8 matmul op or per-channel weight scales — those are follow-ups the
article can point at. This issue is the explanation page plus its nav entry.
Summary
The docs cover weight quantization well — TurboQuant KV-cache compression and the
Q4_0/Q4_K/Q6_K/Q8_0 SIMD matmul kernels — but there is no single article that explains the
end-to-end quantization process, and in particular activation post-training quantization
(PTQ) with calibration, which is what int8 NPU/accelerator targets need.
This issue tracks adding an
explanation/article that ties the picture together and documents thecalibration pattern.
Motivation
Two quantization regimes coexist in SKaiNET and are easy to conflate:
Q8_0/Q4_K,QuantizedMatmul): weights arestored in low-bit blocks and dequant-fused against fp32 activations at matmul time. This is
eager and already documented.
accumulator and a requantize epilogue. This needs calibration — representative activation
ranges — and a different lowering. It is not documented and is not what
QuantizedMatmuldoes.
A contributor targeting an int8 NPU currently has no map of which piece does what, or why
QuantizedMatmulcannot be reused for the activation-int8 path.What the article should cover
requantize → int8 codegen.
activation ranges (
max|x|) over representative inputs, via anExecutionObserver/TensorOps-decorator onmatmul.QuantizedMatmulis eager, weight-only andTensorOps.matmulis dtype-homogeneous (Tensor<T> × Tensor<T> → Tensor<T>, noi8×i8→i32form, no
round), so the int8 activation-matmul pattern cannot be expressed at the tape leveltoday — the int8 emission belongs in backend codegen, fed by SKaiNET-side calibration.
backend) and the per-matmul scale set the emitter consumes (
Sa,Sb,So).Out of scope (this issue)
Implementing a tape-level int8 matmul op or per-channel weight scales — those are follow-ups the
article can point at. This issue is the explanation page plus its nav entry.