Skip to content

Row-packed sub-byte fully-connected weights (#23403) - #23403

Open
mcremon-meta wants to merge 1 commit into
mainfrom
export-D117289220
Open

mcremon-meta wants to merge 1 commit into
mainfrom
export-D117289220

Conversation

@mcremon-meta

@mcremon-meta mcremon-meta commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Summary:

Adds 4- and 6-bit row-packed weight storage for fully-connected layers, from
the packing utilities all the way through to a dispatchable operator. Nothing
is landed unreachable: the packer, the operator, the registration, the
lowering hook and the backend fallback rows all arrive together.

The operator is quantized_fully_connected_packed rather than a packed
quantized_linear because wakeword stage 2 lowers to fully-connected, so this
is the shape the model will actually hit.

Layout: a [out_dim, in_dim] int8 weight becomes [out_dim, ceil(in_dim*bits/8)].
Packing per row is what keeps out_dim as weight.size(0), so the existing
out_size[-1] = weight.size(0) meta computation still holds. in_dim and
weight_bits become explicit operator arguments, since neither can be read off
the packed shape any more. 6-bit follows the torchao UINT6 byte order; signed
values are stored as offset binary.

Packing is storage-only and never changes a value. 6-bit means "quantize with
quant_min/quant_max of -32/31, then store it in 6 bits", not "round an 8-bit
weight down". A layer whose weights do not fit the narrower range, or whose
in_dim is not a multiple of the group size, is skipped and stays 8-bit rather
than failing the compile.

Selection is an explicit compile-time option: a weight_bits field on
EdgePassesConfig, defaulting to 8 (off). The transform runs in
apply_exir_ops_passes after .transform(...) rather than as a pipeline pass,
because it changes a constant's shape and so needs the ExportedProgram, not
the GraphModule a pass is handed. Scalar qparams on a .per_tensor source are
lifted to length-1 constants, since the packed operator only has the
tensor-qparam form.

The reference implementation unpacks and delegates to quantized_linear_common
so the packed path cannot drift from the unpacked arithmetic. The C++ operator
wraps the same inner kernel op_quantized_fully_connected.cpp uses.

Backend wiring: only the generic row is added to operator_fallback.bzl. HiFi
and vision resolve to it through the existing fallback chain, which is the
point of the chain - a backend gets a row only when it has a specialized
implementation. Verified by removing the redundant HiFi row and confirming the
wakeword build still resolves the operator. HiFi's CMake build lists generic
fallback sources explicitly, so the new source is added there. Jarvis keeps its
own operator schema registry in min_runtime/custom_ops.yaml which
codegen.py reads instead of the cadence yamls; a cadence op missing from it
fails the runtime build with a bare KeyError, so it is declared there too.

The requant expression in the packed kernel is deliberately byte-for-byte
identical to quantized_linear.h, including the (1 << 31), because the tests
assert the packed and dense kernels agree exactly.

Differential Revision: D117289220

@pytorch-bot

pytorch-bot Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23403

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit 0f6b8c8 with merge base 5e21c13 (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Oct 3, 2026
@meta-codesync

meta-codesync Bot commented Oct 3, 2026

Copy link
Copy Markdown
Contributor

@mcremon-meta has exported this pull request. If you are a Meta employee, you can view the originating Diff in D117289220.

@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@meta-codesync meta-codesync Bot changed the title Row-packed sub-byte fully-connected weights Row-packed sub-byte fully-connected weights (#23403) Oct 4, 2026
meta-codesync Bot pushed a commit that referenced this pull request Oct 4, 2026
Summary:

Adds 4- and 6-bit row-packed weight storage for fully-connected layers, from
the packing utilities all the way through to a dispatchable operator. Nothing
is landed unreachable: the packer, the operator, the registration, the
lowering hook and the backend fallback rows all arrive together.

The operator is `quantized_fully_connected_packed` rather than a packed
`quantized_linear` because wakeword stage 2 lowers to fully-connected, so this
is the shape the model will actually hit.

Layout: a [out_dim, in_dim] int8 weight becomes [out_dim, ceil(in_dim*bits/8)].
Packing per row is what keeps `out_dim` as `weight.size(0)`, so the existing
`out_size[-1] = weight.size(0)` meta computation still holds. `in_dim` and
`weight_bits` become explicit operator arguments, since neither can be read off
the packed shape any more. 6-bit follows the torchao UINT6 byte order; signed
values are stored as offset binary.

Packing is storage-only and never changes a value. 6-bit means "quantize with
quant_min/quant_max of -32/31, then store it in 6 bits", not "round an 8-bit
weight down". A layer whose weights do not fit the narrower range, or whose
in_dim is not a multiple of the group size, is skipped and stays 8-bit rather
than failing the compile.

Selection is an explicit compile-time option: a `weight_bits` field on
`EdgePassesConfig`, defaulting to 8 (off). The transform runs in
`apply_exir_ops_passes` after `.transform(...)` rather than as a pipeline pass,
because it changes a constant's *shape* and so needs the ExportedProgram, not
the GraphModule a pass is handed. Scalar qparams on a `.per_tensor` source are
lifted to length-1 constants, since the packed operator only has the
tensor-qparam form.

The reference implementation unpacks and delegates to `quantized_linear_common`
so the packed path cannot drift from the unpacked arithmetic. The C++ operator
wraps the same inner kernel `op_quantized_fully_connected.cpp` uses.

Backend wiring: only the generic row is added to `operator_fallback.bzl`. HiFi
and vision resolve to it through the existing fallback chain, which is the
point of the chain - a backend gets a row only when it has a specialized
implementation. Verified by removing the redundant HiFi row and confirming the
wakeword build still resolves the operator. HiFi's CMake build lists generic
fallback sources explicitly, so the new source is added there. Jarvis keeps its
own operator schema registry in `min_runtime/custom_ops.yaml` which
`codegen.py` reads instead of the cadence yamls; a cadence op missing from it
fails the runtime build with a bare KeyError, so it is declared there too.

The requant expression in the packed kernel is deliberately byte-for-byte
identical to `quantized_linear.h`, including the `(1 << 31)`, because the tests
assert the packed and dense kernels agree exactly.

Differential Revision: D117289220
@meta-codesync
meta-codesync Bot force-pushed the export-D117289220 branch from 357e815 to d03c21f Compare October 4, 2026 04:07
meta-codesync Bot pushed a commit that referenced this pull request Oct 5, 2026
Summary:

Adds 4- and 6-bit row-packed weight storage for fully-connected layers, from
the packing utilities all the way through to a dispatchable operator. Nothing
is landed unreachable: the packer, the operator, the registration, the
lowering hook and the backend fallback rows all arrive together.

The operator is `quantized_fully_connected_packed` rather than a packed
`quantized_linear` because wakeword stage 2 lowers to fully-connected, so this
is the shape the model will actually hit.

Layout: a [out_dim, in_dim] int8 weight becomes [out_dim, ceil(in_dim*bits/8)].
Packing per row is what keeps `out_dim` as `weight.size(0)`, so the existing
`out_size[-1] = weight.size(0)` meta computation still holds. `in_dim` and
`weight_bits` become explicit operator arguments, since neither can be read off
the packed shape any more. 6-bit follows the torchao UINT6 byte order; signed
values are stored as offset binary.

Packing is storage-only and never changes a value. 6-bit means "quantize with
quant_min/quant_max of -32/31, then store it in 6 bits", not "round an 8-bit
weight down". A layer whose weights do not fit the narrower range, or whose
in_dim is not a multiple of the group size, is skipped and stays 8-bit rather
than failing the compile.

Selection is an explicit compile-time option: a `weight_bits` field on
`EdgePassesConfig`, defaulting to 8 (off). The transform runs in
`apply_exir_ops_passes` after `.transform(...)` rather than as a pipeline pass,
because it changes a constant's *shape* and so needs the ExportedProgram, not
the GraphModule a pass is handed. Scalar qparams on a `.per_tensor` source are
lifted to length-1 constants, since the packed operator only has the
tensor-qparam form.

The reference implementation unpacks and delegates to `quantized_linear_common`
so the packed path cannot drift from the unpacked arithmetic. The C++ operator
wraps the same inner kernel `op_quantized_fully_connected.cpp` uses.

Backend wiring: only the generic row is added to `operator_fallback.bzl`. HiFi
and vision resolve to it through the existing fallback chain, which is the
point of the chain - a backend gets a row only when it has a specialized
implementation. Verified by removing the redundant HiFi row and confirming the
wakeword build still resolves the operator. HiFi's CMake build lists generic
fallback sources explicitly, so the new source is added there. Jarvis keeps its
own operator schema registry in `min_runtime/custom_ops.yaml` which
`codegen.py` reads instead of the cadence yamls; a cadence op missing from it
fails the runtime build with a bare KeyError, so it is declared there too.

The requant expression in the packed kernel is deliberately byte-for-byte
identical to `quantized_linear.h`, including the `(1 << 31)`, because the tests
assert the packed and dense kernels agree exactly.

Differential Revision: D117289220
@meta-codesync
meta-codesync Bot force-pushed the export-D117289220 branch from d03c21f to ed5cffa Compare October 5, 2026 04:12
Summary:
Pull Request resolved: #23403

Adds 4- and 6-bit row-packed weight storage for fully-connected layers, from
the packing utilities all the way through to a dispatchable operator. Nothing
is landed unreachable: the packer, the operator, the registration, the
lowering hook and the backend fallback rows all arrive together.

The operator is `quantized_fully_connected_packed` rather than a packed
`quantized_linear` because wakeword stage 2 lowers to fully-connected, so this
is the shape the model will actually hit.

Layout: a [out_dim, in_dim] int8 weight becomes [out_dim, ceil(in_dim*bits/8)].
Packing per row is what keeps `out_dim` as `weight.size(0)`, so the existing
`out_size[-1] = weight.size(0)` meta computation still holds. `in_dim` and
`weight_bits` become explicit operator arguments, since neither can be read off
the packed shape any more. 6-bit follows the torchao UINT6 byte order; signed
values are stored as offset binary.

Packing is storage-only and never changes a value. 6-bit means "quantize with
quant_min/quant_max of -32/31, then store it in 6 bits", not "round an 8-bit
weight down". A layer whose weights do not fit the narrower range, or whose
in_dim is not a multiple of the group size, is skipped and stays 8-bit rather
than failing the compile.

Selection is an explicit compile-time option: a `weight_bits` field on
`EdgePassesConfig`, defaulting to 8 (off). The transform runs in
`apply_exir_ops_passes` after `.transform(...)` rather than as a pipeline pass,
because it changes a constant's *shape* and so needs the ExportedProgram, not
the GraphModule a pass is handed. Scalar qparams on a `.per_tensor` source are
lifted to length-1 constants, since the packed operator only has the
tensor-qparam form.

The reference implementation unpacks and delegates to `quantized_linear_common`
so the packed path cannot drift from the unpacked arithmetic. The C++ operator
wraps the same inner kernel `op_quantized_fully_connected.cpp` uses.

Backend wiring: only the generic row is added to `operator_fallback.bzl`. HiFi
and vision resolve to it through the existing fallback chain, which is the
point of the chain - a backend gets a row only when it has a specialized
implementation. Verified by removing the redundant HiFi row and confirming the
wakeword build still resolves the operator. HiFi's CMake build lists generic
fallback sources explicitly, so the new source is added there. Jarvis keeps its
own operator schema registry in `min_runtime/custom_ops.yaml` which
`codegen.py` reads instead of the cadence yamls; a cadence op missing from it
fails the runtime build with a bare KeyError, so it is declared there too.

The requant expression in the packed kernel is deliberately byte-for-byte
identical to `quantized_linear.h`, including the `(1 << 31)`, because the tests
assert the packed and dense kernels agree exactly.

Differential Revision: D117289220

This branch was successfully deployed

1 active deployment
cadence — 0f6b8c8b Deployed Oct 5, 2026 by meta-codesync[bot] via hifi-op-test / hifi4 #31518
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant