Skip to content

feat(bench): ternary f32 gemv scenario + loaded-weight memory evidence (#1141) - #1167

Merged
michalharakal merged 1 commit into
developfrom
feature/1141-ternary-benchmark
Aug 26, 2026
Merged

michalharakal merged 1 commit into
developfrom
feature/1141-ternary-benchmark

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Phase 5 of #1136 — closes #1141 (the docs half landed in #1163).

What

  • engine-ternary-f32-gemv in jvm-cpu-publish + ScenarioRegistry entry — the first ternary scenario. Benches the real KernelDispatch path (adapters included), not a raw kernel loop, at BitNet-2B FFN dims (k=2560, n=6912 — crosses the LUT kernel's internal pthread threshold). Providers: ffm-lut | int8 | f32-reference.
  • TernaryWeightMemoryTest (io-gguf) — the memory evidence artifact: same I2_S tensor loaded keep-packed vs FP32-widened, asserts ≥15× (measures 15.98×).
  • jvm-cpu-publish gains the backend-native-cpu dependency (bundled libskainet_kernels runs exactly as a consumer would).

Measured (Apple arm64 dev host, real dispatch, warmups 3 / measured 3)

provider GOP/s
ffm-lut (vendored NeoGPU kernel, exact) 51.8
int8 (requantize adapter + portable bitnet_gemv) 0.43
f32-reference (Kotlin oracle) 0.30

Two orders of magnitude over the portable path, with exact math — #1136 proof points 1 and 2 now have their evidence artifacts. (Pi-4/A72 numbers will differ; upstream measured 6.78 GOPS there — the qemu lane and an on-device run remain the arm-Linux checks.)

Optional follow-ups not included: OpenBenchmarking profile dir, JMH twin.

🤖 Generated with Claude Code

…dence

engine-ternary-f32-gemv is the first ternary scenario, and deliberately
benches the REAL dispatch path rather than a raw kernel loop — what a
decode step pays, adapters included. Three providers for the same
operands: ffm-lut (the vendored NeoGPU kernel behind the exact
FP32×b1.58 key), int8 (per-call I8-absmax requantize + the portable
bitnet_gemv reference), f32-reference (the Kotlin oracle on the exact
key). BitNet-2B FFN dims (k=2560, n=6912 — crosses the LUT kernel's
internal pthread threshold).

Measured on an Apple-arm64 dev host through real dispatch:
ffm-lut 51.8 GOPS vs int8 0.43 GOPS vs reference 0.30 GOPS — exact math,
two orders of magnitude over the portable path (#1136 proof point 1).

TernaryWeightMemoryTest is the evidence artifact for proof point 2: the
same I2_S GGUF tensor loaded keep-packed vs FP32-widened, asserting the
>=15x ratio (measured 15.98x: 0.25 B/weight + scale vs 4 B/weight).

jvm-cpu-publish gains the backend-native-cpu dependency so the benchmark
runs the bundled libskainet_kernels exactly as a consumer would.

Closes #1141 together with the tutorial that landed in #1163. The
OpenBenchmarking profile dir and a JMH twin stay optional follow-ups.

Refs #1141, #1136

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@michalharakal
michalharakal merged commit 6e806b6 into develop Aug 26, 2026
13 of 14 checks passed
@michalharakal
michalharakal deleted the feature/1141-ternary-benchmark branch August 26, 2026 12:30
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[ternary-f32] Phase 5: benchmark scenario + docs

1 participant