fix(compile-hlo): lower layerNorm/rmsNorm/batchNorm to real stablehlo.reduce (compile on stock IREE) - #769
Merged
Conversation
….reduce convertLayerNorm and convertRmsNorm emitted stablehlo.custom_call @reduce_mean / @reduce_variance placeholders that no MLIR toolchain can compile, so the native norm ops were export-only — downstream users got modules that fail in iree-compile. (RMSNorm's only working path was the NN-DSL primitive decomposition, not the converter op.) LayerNorm/RmsNorm also applied scale/offset without a broadcast_in_dim, a shape mismatch for the usual rank-1 affine params. BatchNorm emitted stablehlo.batch_norm_training, whose 3-tuple result the string emitter can't bind to a single SSA value. Lower all three to real, compilable StableHLO, matching convertGroupNorm (already fixed): - mean/variance via stablehlo.reduce (sum) + divide; variance = E[x²] - E[x]² (ddof=0). - broadcast_in_dim the reduced mean/std AND the scale/offset (shape [axisSize]/(C,)) back to the input shape before the elementwise affine. - batchNorm decomposed to the same elementwise form: 5-operand inference (running mean/var) or 3-operand training (batch stats reduced over the non-feature axes); drop the batch_norm_* emission and the now-unused buildBatchNormOperation helper. Validated end-to-end on stock IREE 3.11.0 (via the skainet-iree-conformance harness): native layerNorm/rmsNorm/batchNorm now iree-compile, run, and match numpy (max abs err 0 / 1.2e-7 / 6e-8). Converter unit tests updated to assert real stablehlo.reduce and the absence of @reduce_* custom_calls. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
MacOS
pushed a commit
to MacOS/SKaiNET
that referenced
this pull request
Jul 10, 2026
Bumps VERSION_NAME 0.32.4 -> 0.33.0. Bundles the develop changes since 0.32.4: GRU layer (SKaiNET-developers#772/SKaiNET-developers#217), upsample2d Bilinear + StableHLO export (SKaiNET-developers#771), the autodiff dispatch correctness fix + 7 newly-differentiable ops + KSP coverage guard (SKaiNET-developers#774), and norm converters lowering to real stablehlo.reduce (SKaiNET-developers#769). Minor bump (not patch): TensorOps.sin/cos/convTranspose1d became abstract, a source/binary-incompatible change for downstream TensorOps implementers. Validated: full conformance suite (12/12 models + 33/33 ops) green end-to-end on IREE llvm-cpu against this tree (via local-maven 0.32.5-localdev1). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
NeuralNetOperationsConverterlowered the normalization ops to MLIR that stock IREE cannot compile, so the native ops were effectively export-only:convertLayerNorm/convertRmsNormemittedstablehlo.custom_call @reduce_mean/@reduce_varianceplaceholders — no MLIR toolchain understands them, soiree-compilefails. (RMSNorm's only working end-to-end path was the NN-DSL primitive decomposition, not the converter op.) They also appliedscale/offsetwith a barestablehlo.multiply/addagainst a rank-1[axisSize]operand — a shape mismatch (nobroadcast_in_dim).convertBatchNormemittedstablehlo.batch_norm_training, whose result is a 3-tuple(output, mean, var)— the string emitter binds it to a single SSA value, which is invalid.convertGroupNormwas already fixed (#752/#754) to use realstablehlo.reduce; this brings the other three in line.Fix
Lower all three to real, compilable StableHLO, mirroring
convertGroupNorm:stablehlo.reduce(sum) +divide; variance =E[x²] − E[x]²(population, ddof=0).broadcast_in_dimthe reduced mean/std and thescale/offset(shape[axisSize]/(C,)) back to the input shape before the elementwise affine.batchNormdecomposed to the same elementwise form: 5-operand inference (running mean/var) or 3-operand training (batch stats reduced over the non-feature axes). Dropped thebatch_norm_*emission and the now-unusedbuildBatchNormOperationhelper.No
@reduce_*custom-call stubs remain.Validation
End-to-end on stock IREE 3.11.0 via the
skainet-iree-conformanceharness — the native ops nowiree-compile→ run → match numpy:rmsNormlayerNormbatchNormConverter unit tests (
LayerNormConverterTest,RmsNormConverterTest) updated to assert realstablehlo.reduceand the absence of@reduce_*custom calls; fullskainet-compile-hlo:jvmTestsuite green.🤖 Generated with Claude Code