docs(explanation): the quantization process — weights, activations, calibration - #747
Merged
Merged
Conversation
…tions, calibration) Add an explanation article that separates SKaiNET's two quantization regimes — weight-only block quant (TurboQuant / Q8_0/Q4_K, eager fp32 activations) vs activation post-training quantization (int8 activations for NPU/accelerator targets) — and walks the PTQ pipeline end to end: - the int8 matmul shape (quantize / i8 matmul i32 / requantize) and the symmetric per-tensor scales Sa/Sb/So it needs; - calibration: tapping the fp32 eager run with a TensorOps `matmul` decorator to capture per-tensor activation ranges over representative inputs; - why QuantizedMatmul can't carry it (eager, weight-only; TensorOps.matmul is dtype-homogeneous with no i8xi8->i32 form and no round) — so the framework calibrates and the target backend emits int8; - the responsibility split, an end-to-end pipeline sketch, and follow-ups (tape-level i8 matmul, per-channel scales, observer notifyOp wiring). Nav entry under Explanation. Closes #746. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
📖 Documentation Preview The documentation has been built successfully for this PR. Generated Files:
Artifacts:
This comment will be updated automatically when the PR is updated. |
This was referenced Jun 22, 2026
MacOS
pushed a commit
to MacOS/SKaiNET
that referenced
this pull request
Jul 10, 2026
Patch release. Bumps VERSION_NAME 0.31.2 -> 0.32.1 and brings the release metadata current: develop never received the 0.32.0 release back-merge (PR SKaiNET-developers#753 was tagged + published but blocked from merging), so this consolidates the 0.32.0 AND 0.32.1 CHANGELOG / README "What's New" entries (supersedes SKaiNET-developers#753). 0.32.1 fix: GroupNorm now emits real stablehlo.reduce instead of @reduce_mean/@reduce_variance custom_calls, so a groupNorm module compiles on stock iree-compile. Verified end-to-end via skainet-iree-conformance: PASS max_abs_err=1.2e-7. (PR SKaiNET-developers#754) 0.32.0 (folded in): GroupNorm StableHLO converter (SKaiNET-developers#752), SKEEP proposals docs (SKaiNET-developers#750), quantization-process doc (SKaiNET-developers#747), dependency bumps. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Adds an
explanation/article documenting SKaiNET's quantization process end to end, and a naventry under Explanation.
The existing docs cover weight quantization well (TurboQuant KV-cache, quantized SIMD kernels),
but nothing explains activation post-training quantization (PTQ) with calibration — what an int8
NPU/accelerator target needs. This page fills that gap.
Contents
quantize → i8 matmul (i32) → requantize— and the symmetric per-tensorscales
Sa/Sb/So.TensorOpsmatmuldecorator (Kotlininterface delegation) to capture per-tensor activation ranges over representative inputs, with the
real API and a sample scales file.
QuantizedMatmulcan't carry the activation-int8 path (eager +weight-only;
TensorOps.matmulis dtype-homogeneous — noi8×i8→i32, noround), and theresulting split: the framework calibrates, the target backend emits int8.
notifyOpwiring).Notes
xref:targets (quantized-simd-kernels,turboquant-kv-compression) verified to exist;uses the component-level
{framework_name}attribute like the other pages.matmultap), nottied to any downstream consumer.
Closes #746.