Skip to content

Trace/export pipeline cannot emit multi-billion-param models: Void ops materialize real zeros and constant extraction duplicates all weights #1247

Description

@michalharakal

Found while building the Gemma 3n StableHLO export in SKaiNET-transformers (the #377 DSL lane / mobile path). The DSL → tape → toComputeGraph → StableHLO pipeline works (proven at 135M–270M by SmolLM2/FunctionGemma) but cannot emit a 4.5B-parameter model on a 48 GB host:

  1. VoidTensorOps allocates real dense zero buffers for every recorded op (ShapeOnlyDataFactory.zeros → DenseTensorDataFactory.createFloatTensorData). Worst case: matmulWeightTransposed calls transpose(weight) per projection, allocating a weight-sized zero tensor that the tape then retains — ~9 GB of retained zeros across a 30-layer trace of Gemma 3n E2B (incl. a 2.1 GB zero for the 262k-vocab lm_head transpose). Shape-only propagation shouldn't materialize data.

  2. TraceToGraphBuilder.extractFloatArray copies every frozen weight into the graph while the originals stay referenced by the module tree — dense originals (~10 GB for the E2B trunk+embedding) and graph copies (~10 GB) are co-resident, plus the retained zeros from (1). Measured: OOM at a 46 GB heap on the E2B export, in TraceToGraphBuilder.finalize.

  3. Packed tensors cannot become graph constants at all — a keep-packed load traces them into opaque function arguments (we measured a func @gemma3n with 190+ weight args and zero dot ops), silently producing an unservable module. A loud error (or packed→dense dequant-on-extract) would be better than the silent signature leak; transformers now guards this with an arg-count check on its side.

Suggested directions (any one of the first two unblocks billion-scale export):

  • Shape-only tensor data for VoidTensorOps (no backing array).
  • Streaming/late constant extraction: dequant-or-copy per tensor at emission time, releasing each source afterwards, instead of building the full copied set up front.
  • Reject packed frozen params in the tracer with a clear message.

Repro: SKaiNET-transformers :llm-inference:gemma3n:exportGemma3n on gemma-3n-E2B-it-Q4_K_M.gguf (48 GB host, -PexportMaxHeap=46g). The truncated GEMMA3N_LAYERS=4 export passes, isolating scale as the only variable.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions