Skip to content

docs: explain the end-to-end quantization process (weight quant + activation PTQ/calibration) #746

Description

@michalharakal

Summary

The docs cover weight quantization well — TurboQuant KV-cache compression and the
Q4_0/Q4_K/Q6_K/Q8_0 SIMD matmul kernels — but there is no single article that explains the
end-to-end quantization process, and in particular activation post-training quantization
(PTQ) with calibration, which is what int8 NPU/accelerator targets need.

This issue tracks adding an explanation/ article that ties the picture together and documents the
calibration pattern.

Motivation

Two quantization regimes coexist in SKaiNET and are easy to conflate:

  1. Weight-only quantization (TurboQuant / GGUF Q8_0/Q4_K, QuantizedMatmul): weights are
    stored in low-bit blocks and dequant-fused against fp32 activations at matmul time. This is
    eager and already documented.
  2. Activation PTQ (int8 accelerators): both operands of the matmul are int8, with an i32
    accumulator and a requantize epilogue. This needs calibration — representative activation
    ranges — and a different lowering. It is not documented and is not what QuantizedMatmul
    does.

A contributor targeting an int8 NPU currently has no map of which piece does what, or why
QuantizedMatmul cannot be reused for the activation-int8 path.

What the article should cover

  • The two regimes side by side (weight quant vs activation PTQ) and when each applies.
  • The PTQ pipeline: calibrate → per-tensor symmetric scales → quantize / i8 matmul (i32) /
    requantize → int8 codegen
    .
  • The calibration mechanism: tapping the eager execution stream to capture per-tensor
    activation ranges (max|x|) over representative inputs, via an ExecutionObserver /
    TensorOps-decorator on matmul.
  • The honest limitation finding: QuantizedMatmul is eager, weight-only and
    TensorOps.matmul is dtype-homogeneous (Tensor<T> × Tensor<T> → Tensor<T>, no i8×i8→i32
    form, no round), so the int8 activation-matmul pattern cannot be expressed at the tape level
    today — the int8 emission belongs in backend codegen, fed by SKaiNET-side calibration.
  • Where each responsibility lives (calibration in the framework, int8 emission in the target
    backend) and the per-matmul scale set the emitter consumes (Sa, Sb, So).

Out of scope (this issue)

Implementing a tape-level int8 matmul op or per-channel weight scales — those are follow-ups the
article can point at. This issue is the explanation page plus its nav entry.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    documentationImprovements or additions to documentation (DARC: D)

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions