Repository navigation
feat(compile): graph-output pruning for export + finite causal mask - #760
Merged
Merged
Conversation
A traced decoder marks every leaf tensor as a graph output, so a StableHLO/IREE export returns the logits plus dozens of dangling per-layer intermediates (extra func returns + dead op subgraphs). Add OutputDesignatedGraph (compile-dag) to override the output set by node id, and ComputeGraph.prunedToOutputs(ids) (compile-opt) = designate + DeadCodeElimination, so an exporter keeps only the logits. Verified on a real TinyLlama export: 45 outputs -> 1, 1467 -> 1247 nodes, attention intact. Adds GraphPruningTest. Also: SDPA causal-mask converter emits -1e30 instead of -inf (0xFF800000) for the masked-fill, matching buildSlidingCausalMask (numerically equivalent after softmax; avoids a -inf splat). Refs the iree-compile decoder-export crash (constant-fold null-deref); this removes the dangling-output contribution. A deeper trigger (K activations frozen as constants) is tracked separately. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…output-pruning # Conflicts: # CHANGELOG.md
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
StableHLO/IREE export hygiene for decoder graphs, toward unblocking the Llama→IREE path (#759).
OutputDesignatedGraph(compile-dag) +ComputeGraph.prunedToOutputs(outputNodeIds)(compile-opt, = designate outputs +DeadCodeEliminationPass). A traced decoder marks every leaf as a graph output, so the export returned the logits plus ~44 dangling per-layer intermediates (extrafuncreturns + dead subgraphs). Exporters can now keep only the logits. Verified on a real TinyLlama export: 45 outputs → 1, 1467 → 1247 nodes, attention intact. AddsGraphPruningTest.-1.000000e+30instead of-inf(0xFF800000) for the masked-fill, matchingMultiHeadAttention.buildSlidingCausalMask(numerically equivalent after softmax; avoids a-infsplat in the IR).Why
Part of fixing the
iree-compiledecoder-export crash (#759). This removes the dangling-output contribution and the-infsplat. It does not fully fix #759 — a deeper trigger remains (K activations are const-folded into parameters during theVoidTensorOps+embedConstantsexport, and IREE'sinsert_slice/extract_slicefold over those frozen constants null-derefs). That root cause is documented in #759 and is the next step.Test
:skainet-compile:skainet-compile-opt:jvmTest --tests "*GraphPruningTest"✅compile-dag,compile-opt,compile-hlo).Refs #759
🤖 Generated with Claude Code