Skip to content

docs(explanation): the quantization process — weights, activations, calibration - #747

Merged
michalharakal merged 1 commit into
developfrom
feature/746-quantization-process-docs
Jun 19, 2026
Merged

michalharakal merged 1 commit into
developfrom
feature/746-quantization-process-docs

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

What

Adds an explanation/ article documenting SKaiNET's quantization process end to end, and a nav
entry under Explanation.

The existing docs cover weight quantization well (TurboQuant KV-cache, quantized SIMD kernels),
but nothing explains activation post-training quantization (PTQ) with calibration — what an int8
NPU/accelerator target needs. This page fills that gap.

Contents

  • The two regimes side by side (weight-only block quant vs activation PTQ) and when each applies.
  • The int8 matmul shape — quantize → i8 matmul (i32) → requantize — and the symmetric per-tensor
    scales Sa / Sb / So.
  • Calibration: tapping the fp32 eager run with a TensorOps matmul decorator (Kotlin
    interface delegation) to capture per-tensor activation ranges over representative inputs, with the
    real API and a sample scales file.
  • The honest finding on why QuantizedMatmul can't carry the activation-int8 path (eager +
    weight-only; TensorOps.matmul is dtype-homogeneous — no i8×i8→i32, no round), and the
    resulting split: the framework calibrates, the target backend emits int8.
  • End-to-end pipeline sketch and follow-ups (tape-level i8 matmul, per-channel scales, observer
    notifyOp wiring).

Notes

  • Documentation only — no code changes.
  • Both xref: targets (quantized-simd-kernels, turboquant-kv-compression) verified to exist;
    uses the component-level {framework_name} attribute like the other pages.
  • The calibration example is written as a reusable framework pattern (a generic matmul tap), not
    tied to any downstream consumer.

Closes #746.

…tions, calibration)

Add an explanation article that separates SKaiNET's two quantization regimes —
weight-only block quant (TurboQuant / Q8_0/Q4_K, eager fp32 activations) vs
activation post-training quantization (int8 activations for NPU/accelerator
targets) — and walks the PTQ pipeline end to end:

- the int8 matmul shape (quantize / i8 matmul i32 / requantize) and the
  symmetric per-tensor scales Sa/Sb/So it needs;
- calibration: tapping the fp32 eager run with a TensorOps `matmul` decorator
  to capture per-tensor activation ranges over representative inputs;
- why QuantizedMatmul can't carry it (eager, weight-only; TensorOps.matmul is
  dtype-homogeneous with no i8xi8->i32 form and no round) — so the framework
  calibrates and the target backend emits int8;
- the responsibility split, an end-to-end pipeline sketch, and follow-ups
  (tape-level i8 matmul, per-channel scales, observer notifyOp wiring).

Nav entry under Explanation. Closes #746.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

📖 Documentation Preview

The documentation has been built successfully for this PR.

Generated Files:

  • Operator documentation: docs/modules/operators/_generated_/
  • JSON schema output: operators.json

Artifacts:

  • Download the documentation-preview-747 artifact to view the complete documentation locally.

This comment will be updated automatically when the PR is updated.

@michalharakal
michalharakal merged commit 72d561f into develop Jun 19, 2026
8 checks passed
@michalharakal
michalharakal deleted the feature/746-quantization-process-docs branch June 19, 2026 11:46
MacOS pushed a commit to MacOS/SKaiNET that referenced this pull request Jul 10, 2026
Patch release. Bumps VERSION_NAME 0.31.2 -> 0.32.1 and brings the release
metadata current: develop never received the 0.32.0 release back-merge (PR SKaiNET-developers#753
was tagged + published but blocked from merging), so this consolidates the 0.32.0
AND 0.32.1 CHANGELOG / README "What's New" entries (supersedes SKaiNET-developers#753).

0.32.1 fix: GroupNorm now emits real stablehlo.reduce instead of
@reduce_mean/@reduce_variance custom_calls, so a groupNorm module compiles on
stock iree-compile. Verified end-to-end via skainet-iree-conformance:
PASS max_abs_err=1.2e-7. (PR SKaiNET-developers#754)

0.32.0 (folded in): GroupNorm StableHLO converter (SKaiNET-developers#752), SKEEP proposals docs
(SKaiNET-developers#750), quantization-process doc (SKaiNET-developers#747), dependency bumps.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

docs: explain the end-to-end quantization process (weight quant + activation PTQ/calibration)

1 participant