Found while building the Gemma 3n StableHLO export in SKaiNET-transformers (the #377 DSL lane / mobile path). The DSL → tape → toComputeGraph → StableHLO pipeline works (proven at 135M–270M by SmolLM2/FunctionGemma) but cannot emit a 4.5B-parameter model on a 48 GB host:
-
VoidTensorOps allocates real dense zero buffers for every recorded op (ShapeOnlyDataFactory.zeros → DenseTensorDataFactory.createFloatTensorData). Worst case: matmulWeightTransposed calls transpose(weight) per projection, allocating a weight-sized zero tensor that the tape then retains — ~9 GB of retained zeros across a 30-layer trace of Gemma 3n E2B (incl. a 2.1 GB zero for the 262k-vocab lm_head transpose). Shape-only propagation shouldn't materialize data.
-
TraceToGraphBuilder.extractFloatArray copies every frozen weight into the graph while the originals stay referenced by the module tree — dense originals (~10 GB for the E2B trunk+embedding) and graph copies (~10 GB) are co-resident, plus the retained zeros from (1). Measured: OOM at a 46 GB heap on the E2B export, in TraceToGraphBuilder.finalize.
-
Packed tensors cannot become graph constants at all — a keep-packed load traces them into opaque function arguments (we measured a func @gemma3n with 190+ weight args and zero dot ops), silently producing an unservable module. A loud error (or packed→dense dequant-on-extract) would be better than the silent signature leak; transformers now guards this with an arg-count check on its side.
Suggested directions (any one of the first two unblocks billion-scale export):
- Shape-only tensor data for
VoidTensorOps (no backing array).
- Streaming/late constant extraction: dequant-or-copy per tensor at emission time, releasing each source afterwards, instead of building the full copied set up front.
- Reject packed frozen params in the tracer with a clear message.
Repro: SKaiNET-transformers :llm-inference:gemma3n:exportGemma3n on gemma-3n-E2B-it-Q4_K_M.gguf (48 GB host, -PexportMaxHeap=46g). The truncated GEMMA3N_LAYERS=4 export passes, isolating scale as the only variable.
Found while building the Gemma 3n StableHLO export in SKaiNET-transformers (the #377 DSL lane / mobile path). The DSL → tape →
toComputeGraph→ StableHLO pipeline works (proven at 135M–270M by SmolLM2/FunctionGemma) but cannot emit a 4.5B-parameter model on a 48 GB host:VoidTensorOpsallocates real dense zero buffers for every recorded op (ShapeOnlyDataFactory.zeros→DenseTensorDataFactory.createFloatTensorData). Worst case:matmulWeightTransposedcallstranspose(weight)per projection, allocating a weight-sized zero tensor that the tape then retains — ~9 GB of retained zeros across a 30-layer trace of Gemma 3n E2B (incl. a 2.1 GB zero for the 262k-vocab lm_head transpose). Shape-only propagation shouldn't materialize data.TraceToGraphBuilder.extractFloatArraycopies every frozen weight into the graph while the originals stay referenced by the module tree — dense originals (~10 GB for the E2B trunk+embedding) and graph copies (~10 GB) are co-resident, plus the retained zeros from (1). Measured: OOM at a 46 GB heap on the E2B export, inTraceToGraphBuilder.finalize.Packed tensors cannot become graph constants at all — a keep-packed load traces them into opaque function arguments (we measured a
func @gemma3nwith 190+ weight args and zero dot ops), silently producing an unservable module. A loud error (or packed→dense dequant-on-extract) would be better than the silent signature leak; transformers now guards this with an arg-count check on its side.Suggested directions (any one of the first two unblocks billion-scale export):
VoidTensorOps(no backing array).Repro: SKaiNET-transformers
:llm-inference:gemma3n:exportGemma3nongemma-3n-E2B-it-Q4_K_M.gguf(48 GB host,-PexportMaxHeap=46g). The truncatedGEMMA3N_LAYERS=4export passes, isolating scale as the only variable.