fix(jni): eager library init; emulator-validated parity suite; gate q4k/q6k on #944 (#920) - #945
Conversation
…#944 Emulator validation (x86_64 AVD) of the JNI bridge surfaced three things: 1. JniKernels loaded its library lazily via the variant property — a direct call to an external fun never triggered the load and threw UnsatisfiedLinkError. The loader now runs eagerly in object init, so first access to ANY member loads the library. 2. The parity test data now mirrors the NativeKn* generators exactly (Q4_K pins d AND dMin; Q6_K pins d at the block END, bytes 208-209 — the first version pinned bytes 0-1 for every format, which is wrong for the K-formats). 3. Validation uncovered #944: the C q4k/q6k kernels and the Kotlin scalar references disagree on the 6-bit scale decode — data-dependent AND compiler-dependent (the exact seed/shape combo that passes with the gcc-built archive fails with the NDK clang build, up to double-digit percent per row), while q8_0/q4_0 agree exactly. The q4k/q6k parity tests are @ignore'd referencing #944; bridge mechanics stay covered by the q8_0/q4_0/smoke parity and the loader/ provider tests. Emulator result: 8 tests — 4 passed (provider availability, q8_0 parity, q4_0 parity, smoke roundtrip), 4 skipped (#944). Host unit tests still 4/4. Refs #920 #944
…xel 8a) Add loader_selects_tier_matching_cpu_features: derive the expected library variant from the same /proc/cpuinfo signal the loader uses, then assert JniKernels.variant matches. Guards the core #920 runtime-dispatch contract — a broken feature probe or a mis-packaged .so would otherwise pass every parity test (they run on whatever loaded) while silently running the wrong tier. Validated on physical hardware: Pixel 8a (arm64-v8a, asimddp+asimdhp present) selects V82_DOTPROD; x86_64 emulator selects BASELINE — the test self-adjusts to both without hard-coding. Physical-device run summary (Pixel 8a, this commit): 9 connected tests, 0 failures, 4 skipped (#944). q8_0/q4_0 NEON parity and the smoke roundtrip pass on-device; tier selection confirmed dotprod. Refs #920
|
Validated on physical hardware — Pixel 8a (arm64-v8a, Tensor G3). CPU Device run (this branch): 9 connected tests, 0 failures, 4 skipped (#944).
Same suite on the x86_64 emulator: 0 failures, baseline tier selected. The bridge is proven end-to-end on real ARM hardware — packaging, feature-gated loader, JNI resolution, array pinning, ServiceLoader discovery, and the dotprod path. Still to come (separate): a decode-throughput (tok/s) measurement on the SmolLM2 reproducer once #944 unblocks the K-quant path that a real GGUF actually uses. |
…solved) #944 turned out not to be a bug: the C q4k/q6k kernels quantize the activation to int8 (ggml block_q8_K fast path, faithful to ggml_vec_dot_q4_K_q8_K), so they are deliberately lossy vs the exact-float scalar reference. Per-row relative parity is the wrong gate for a lossy kernel — on zero-mean random fixtures a near-zero row shows unbounded relative error from a tiny absolute one. Adopt the aggregate RMS(error)/RMS(signal) gate (AGG_REL_TOL = 0.03) the NativeKn* K-format parity tests already use, and drop the four @ignore's. q8_0/q4_0 keep the exact per-row tolerance — those kernels dequantize the weight and accumulate in FP32, no activation quant, so bit-level parity is right and catches any bug. Split the shared driver into assertExactParity / assertRmsParity accordingly. The RMS gate still catches every structural bridge failure (wrong offset, layout, or loaded library diverge by orders of magnitude); only the intended quantization loss passes under it. Verified on device AND emulator: 9/9 tests, 0 skipped, 0 failures on both the physical Pixel 8a (V82_DOTPROD tier) and the x86_64 emulator (BASELINE). Host unit tests 4/4. Refs #920 #944
|
Update: q4k/q6k parity re-enabled (was @ignore'd), now that #944 is understood. The bisection (full writeup in #944) showed the K-quant divergence is not a bug — the C kernels quantize the activation to int8 (ggml This branch now uses the aggregate RMS(error)/RMS(signal) gate ( Full green on real hardware and emulator now:
So the JNI bridge is fully validated end-to-end, K-quant path included — no remaining skips. |
Kernel-throughput projection of SmolLM2-135M Q8_0 autoregressive decode: runs the real JNI kernels at the model's actual weight shapes (30 layers of q/k/v/o + gate/up/down projections, plus lm_head) on the device CPU, sums one token's matmul wall-clock, reports projected tok/s vs the scalar floor. Decode of a 135M model is matmul-bound; attention/RoPE/sampling are negligible at this size, so timing the matmuls at real shapes captures the bottleneck. Not an end-to-end generation (that needs the transformers stack — transformers#272); labeled as a projection. Measured on Pixel 8a (Tensor G3, V82_DOTPROD tier): jni = 23.96 tok/s scalar = 3.76 tok/s (6.4x) usability gate (3 tok/s): PASS Answers the #920 field report (1.0 tok/s scalar on Android, AI feature disabled below 3 tok/s): the NEON JNI path clears the gate ~8x. Refs #920
Follow-up to the just-merged #943, carrying the results of the first real-device-class validation (x86_64 emulator run of the connected suite).
Fixes
JniKernelsloaded its.solazily via thevariantproperty; a consumer calling anyexternalfun before touchingvarianthitUnsatisfiedLinkError. The loader now runs in object init, so first access to ANY member loads the library. (Caught bysmoke_roundtriprunning before the availability probe.)NativeKn*generators exactly — the first version pinned bytes 0–1 as the scale slot for every format, which is wrong for the K-formats (Q4_K needsd+dMinat 0–3; Q6_K'sdsits at bytes 208–209).@Ignored with the issue reference; details and reproducers are in Q4_K/Q6_K parity tests compare the ggml-faithful int8 (block_q8_K) native kernels against exact-float scalar references — divergence is intended activation-quant loss, not a bug #944.Emulator result (Pixel 8a AVD, x86_64)
8 connected tests — 4 passed (provider availability + variant selection, q8_0 parity, q4_0 parity, smoke roundtrip), 4 skipped (#944). Host unit tests 4/4. This validates the bridge mechanics end-to-end: packaging, two-tier loader, JNI symbol resolution, array pinning, ServiceLoader discovery.
Physical arm64 run (NEON/dotprod tiers + tok/s measurement) still pending — next on the device bench.
Refs #920 #944