Skip to content

Thread the packed matmul kernels — ~3× decode on Android, the biggest lever under the bandwidth ceiling #1195

Description

@michalharakal

Follow-up to #1189/#1190; sibling of #1193. The Q4_K/Q6_K (and other packed) C matmul kernels run single-threaded per call. On the Pixel 8a that leaves most of the SoC idle during decode:

ceiling (Qwen2.5-1.5B Q4_K_M, Pixel 8a) tok/s
measured today — single thread, incl. lm_head ~5.5 (180 ms/token; projections measured at 153 ms via the M2-A5 harness)
threaded, 4 threads @ ~2.7× scaling ~14–16
sustained DRAM bandwidth (~22 GB/s ÷ ~870 MB streamed/token) ~25
LPDDR5X peak (51 GB/s) ~59 (unreachable)

llama.cpp reaches 8–12 tok/s on this phone/model class with 4 threads — threading is the gap, not the memory path: mapped page-cache reads are byte-for-byte DRAM reads once resident (measured 0–1 major faults in steady state), so this lever is fully orthogonal to residency.

Design notes

  • Partition over output rows (o-ranges): works identically for the feed-order and the row-major (_rm, Packed-tensor mapped staging: serve Q4_K/Q6_K/ternary blocks straight from the mmap on Android #1189) variants — each thread owns a disjoint slice of out[], no accumulation races, no synchronization beyond join. The Q8 activation quantization can be done once before the fork (it is shared read-only).
  • Precedent in-tree: the vendored NeoGPU ternary kernel (hs_ml_ternary_neon.c) already threads with pthreads under both JNI and FFM — N_THREADS 4, THREAD_THRESHOLD 512 (skip threading for small outputDim, where fixed costs dominate — exactly what the SmolLM2-135M numbers showed). Same pattern applies; consider a shared pool rather than create/join per call once profiling says the per-call cost matters.
  • JNI: threads spawned inside the native call while the arrays stay pinned is the pattern the ternary kernel already uses; the direct-buffer entries (Packed-tensor mapped staging: serve Q4_K/Q6_K/ternary blocks straight from the mmap on Android #1189) have no pins on the weight at all.
  • big.LITTLE: benchmark 2 vs 4 threads and X3+A715 vs all-core — A510s likely hurt via straggler effect; a threshold on outputDim (like the ternary kernel's 512) keeps small projections single-threaded.
  • Measure with the M2-A5 harness (Pixel 8a, detached am instrument): expect ms/step to drop from ~153 toward ~60 on the 1.5B, and the 135M numbers to move little (its matmuls sit under any sane threshold).

Scope

Q4_K + Q6_K first (both orders — the _rm twins must thread too or mapped decode stays behind), then the remaining formats with #1192. Bit-identity per output row is preserved by row partitioning (each out[o] keeps its exact accumulation order), so the existing parity suites remain the oracle.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions