functiongemma: position-selected graphs gemma_at / gemma_prefill_at — LM head on one position (#406) - #415
Merged
Conversation
… LM head on one one-hot-selected position, single-token result (#406) GemmaModel.forwardAt / forwardPrefillAt multiply the final-normed hidden state by a one-hot [1, seq] row before lm_head (plain matmul, no dynamic-index op), so the compiled graphs return [1, vocab] logits and a 1xi32 token instead of seq x vocab logits plus a seq x vocab argmax scratch. Measured on a 4-core arm32 box at seq 64: 6.6 s/step vs 10.4 s; the 1024-position graph now fits a 32-bit process (the all-positions variant died on a 1 GB calloc). Contract additions are additive: FN_REDECODE_AT, FN_PREFILL_AT, selectArgs(), prefillAtOutputs(), manifest entries, qualified(); CLI GEMMA_GRAPH=redecode_at|prefill_at; dump test for both graphs. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
michalharakal
force-pushed
the
functiongemma-export-floats-cast
branch
from
September 3, 2026 21:43
15e1ef3 to
6686107
Compare
michalharakal
force-pushed
the
functiongemma-position-selected-head
branch
from
September 3, 2026 21:43
c2cb2e3 to
95f2d27
Compare
This was referenced Sep 4, 2026
This was referenced Sep 4, 2026
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Addresses #406 (the LM-head half). Stacked on #414 — merge that first.
Why
The exported redecode and prefill graphs run
lm_headand the argmax over all SEQ positions and then materialise three moreSEQ×262144tensors for the argmax, although the caller reads one position per step. Measured on an arm32 Android device (arm32, 4 threads, FP32): that is 36 % of every step, and at SEQ 1024 the1024×262144i32 scratch (1.07 GB) cannot be allocated in a 32-bit process — the catalog-sized prompt could not run at all.Change (additive; the existing graphs are untouched)
GemmaModel.forwardAt(input, selectAt, ctx)/forwardPrefillAt(...): a one-hot[1, seq]row multiplies the final-normed hidden state down to[1, hidden]beforelm_head— a plain matmul, so no dynamic-index op is needed and every backend lowers it.FunctionGemmaExportHarness.exportRedecodeAt/exportPrefillAt:gemma_at(tokens 1×SEQ i32, select 1×SEQ f32) → 1×i32andgemma_prefill_at(tokens SEQ i32, select) → per-layer K/V…, 1×i32(own archives, like the other graphs).FN_REDECODE_AT,FN_PREFILL_AT,selectArgs(),prefillAtOutputs(),qualified()(the runtime wantsmodule.<fn>; see iree-android: nativeStep returns null silently on any failure; FunctionGemmaContract.FN_REDECODE ("gemma") is not the qualified name the runtime needs ("module.gemma") #404), manifest entriesredecodeAt/prefillAt+ arg/result arrays.CONTRACT_VERSIONunchanged (nothing existing moves).GEMMA_GRAPH=redecode_at|prefill_at.FunctionGemmaExportDumpTest.positionSelectedGraphs_emitContractShapes.Measured
gemma_atat SEQ 64 FP32: host oracle1xi32=32691; on the device1xi32=32691in 6.06 s/step vs 10.40 s for the all-positions graph (iree-benchmark-module, 4 threads).gemma_prefill_atat SEQ 1024 with an 852-token prompt on the host: first token48, equal to the all-positions redecode graph at position P−1; its K/V drivegemma_with_pastto the truth token (6639) once the sliding layers' cache is windowed to 512 — details in iree-android: stateful prefill/step contract with a reusable prompt-prefix KV snapshot (and an embeddings-input variant for Vulkan) #410.Not in this PR
The per-SEQ parameter archives (the other half of #406) and the sliding-window contract for the with-past graph (#410).