You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Follow-up to #1189/#1190; sibling of #1193. The Q4_K/Q6_K (and other packed) C matmul kernels run single-threaded per call. On the Pixel 8a that leaves most of the SoC idle during decode:
ceiling (Qwen2.5-1.5B Q4_K_M, Pixel 8a)
tok/s
measured today — single thread, incl. lm_head
~5.5 (180 ms/token; projections measured at 153 ms via the M2-A5 harness)
threaded, 4 threads @ ~2.7× scaling
~14–16
sustained DRAM bandwidth (~22 GB/s ÷ ~870 MB streamed/token)
~25
LPDDR5X peak (51 GB/s)
~59 (unreachable)
llama.cpp reaches 8–12 tok/s on this phone/model class with 4 threads — threading is the gap, not the memory path: mapped page-cache reads are byte-for-byte DRAM reads once resident (measured 0–1 major faults in steady state), so this lever is fully orthogonal to residency.
Design notes
Partition over output rows (o-ranges): works identically for the feed-order and the row-major (_rm, Packed-tensor mapped staging: serve Q4_K/Q6_K/ternary blocks straight from the mmap on Android #1189) variants — each thread owns a disjoint slice of out[], no accumulation races, no synchronization beyond join. The Q8 activation quantization can be done once before the fork (it is shared read-only).
Precedent in-tree: the vendored NeoGPU ternary kernel (hs_ml_ternary_neon.c) already threads with pthreads under both JNI and FFM — N_THREADS 4, THREAD_THRESHOLD 512 (skip threading for small outputDim, where fixed costs dominate — exactly what the SmolLM2-135M numbers showed). Same pattern applies; consider a shared pool rather than create/join per call once profiling says the per-call cost matters.
big.LITTLE: benchmark 2 vs 4 threads and X3+A715 vs all-core — A510s likely hurt via straggler effect; a threshold on outputDim (like the ternary kernel's 512) keeps small projections single-threaded.
Measure with the M2-A5 harness (Pixel 8a, detached am instrument): expect ms/step to drop from ~153 toward ~60 on the 1.5B, and the 135M numbers to move little (its matmuls sit under any sane threshold).
Scope
Q4_K + Q6_K first (both orders — the _rm twins must thread too or mapped decode stays behind), then the remaining formats with #1192. Bit-identity per output row is preserved by row partitioning (each out[o] keeps its exact accumulation order), so the existing parity suites remain the oracle.
Follow-up to #1189/#1190; sibling of #1193. The Q4_K/Q6_K (and other packed) C matmul kernels run single-threaded per call. On the Pixel 8a that leaves most of the SoC idle during decode:
llama.cpp reaches 8–12 tok/s on this phone/model class with 4 threads — threading is the gap, not the memory path: mapped page-cache reads are byte-for-byte DRAM reads once resident (measured 0–1 major faults in steady state), so this lever is fully orthogonal to residency.
Design notes
o-ranges): works identically for the feed-order and the row-major (_rm, Packed-tensor mapped staging: serve Q4_K/Q6_K/ternary blocks straight from the mmap on Android #1189) variants — each thread owns a disjoint slice ofout[], no accumulation races, no synchronization beyond join. The Q8 activation quantization can be done once before the fork (it is shared read-only).hs_ml_ternary_neon.c) already threads with pthreads under both JNI and FFM —N_THREADS 4,THREAD_THRESHOLD 512(skip threading for smalloutputDim, where fixed costs dominate — exactly what the SmolLM2-135M numbers showed). Same pattern applies; consider a shared pool rather than create/join per call once profiling says the per-call cost matters.outputDim(like the ternary kernel's 512) keeps small projections single-threaded.am instrument): expect ms/step to drop from ~153 toward ~60 on the 1.5B, and the 135M numbers to move little (its matmuls sit under any sane threshold).Scope
Q4_K + Q6_K first (both orders — the
_rmtwins must thread too or mapped decode stays behind), then the remaining formats with #1192. Bit-identity per output row is preserved by row partitioning (eachout[o]keeps its exact accumulation order), so the existing parity suites remain the oracle.