Skip to content

bench(spike): Phase-2 TensorView/Storage access-path spike — JMH benchmarks, flat-RSS loop, report (SKEEP-003 decision #6, S1.0) - #1060

Merged
michalharakal merged 1 commit into
developfrom
feature/1016-p2-spike
Aug 23, 2026
Merged

michalharakal merged 1 commit into
developfrom
feature/1016-p2-spike

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Summary

SKEEP-003 slice S1.0 (milestone M1 #1002, decision #6, #1016): the Phase-2 access-path spike — what does a TensorView layer between a kernel and its bytes cost on the JVM? Report: docs/design/memory/spike-p2-access-path-2026-08-23.md; benchmarks kept reproducible in the JMH module (sk.ainet.bench.spike, run with -PjmhIncludes='spike.*' and runFlatRssSpike).

Verdict: go for Phase 2, with the access-path rule written into the M1 slices — kernels take TensorViews and unwrap once per call (asHeapArray() / segment()); decoding get() is the reference path only; JVM Scope.Forward defaults to heap slabs; vector kernels keep explicit offsets.

Numbers (i7-9750H, JDK 25, JMH fork 1 / 3 warmup / 5 it):

  • GEMV 256×1024: view over heap, unwrap once +1 % · heap get per element 0 % · off-heap unwrap +3 % · off-heap get +4 % — within budget.
  • Elementwise add 1 M: raw 286 µs · raw with loop-invariant offsets (control, no view) 539 µs · view-heap unwrap 567 ± 182 · view-heap get 516 · view-segment unwrap 495 · view-segment get 1579. The 1.9× on the vectorizable loop comes from C2 not auto-vectorizing the offset form, not from the view (view-unwrap ≈ raw-offset). SKaiNET's vector kernels take offsets explicitly, so the ≤ 3 % elementwise budget is re-checked on them at S1.4 / S1.7 against the baseline.
  • Off-heap: scalar MemorySegment.getAtIndex loops are 1.7–5.5× slower than arrays → activations stay on the heap by default on the JVM; mapped/off-heap weights are fine (already consumed by ByteVector.fromMemorySegment).
  • Flat RSS (3 000 decode-shaped steps, -Xmx256m): recycled bump slab 62–86 MiB with zero slab bytes allocated after warm-up vs 214 MiB plateau with per-step heap arrays; the slab's residual creep is the one view object per allocate() (193/step) — cheap, not free.

Also in this PR: -PjmhIncludes=<regex> to run a subset of the JMH suite; runFlatRssSpike JavaExec task.

Open: the Android half (Cortex-A55-class device, direct ByteBuffer) is pending the reference-device choice (PRD §9); recorded on #1016 when available — the JVM result already sets the rule.

Test plan

Full local gate (scripts/pr-gate.sh, JDK 25); results in the first comment. Benchmarks themselves are not run by the gate (JMH source set); they were run for the report as described.

Closes #1016

🤖 Generated with Claude Code

…hmarks, flat-RSS loop, report (SKEEP-003 decision #6, #1016)

Throw-away model of the proposed Storage (heap FloatArray | off-heap
MemorySegment) and TensorView (offset/length over a storage) measured
against raw arrays: elementwise add over 1 M floats and gemv 256x1024,
each as raw / view-get / view-unwrap-once over heap and off-heap, plus a
raw-with-offsets control; and FlatRssSpike, a 3000-step decode-shaped
loop with a recycled Arena.ofShared bump slab vs fresh heap arrays.

Findings (docs/design/memory/spike-p2-access-path-2026-08-23.md): the
view indirection is within budget when unwrapped once per call (gemv
+0..4 %); offset-based array indexing, not the view, costs 1.9x on the
vectorizable elementwise loop (C2 SuperWord); per-element get() is the
slow reference path (+80 % heap, 5.5x segment); off-heap activations need
Vector-API-over-segment kernels, so Scope.Forward on the JVM should default
to heap slabs; the slab keeps RSS at 62-86 MiB over 3000 steps vs 214 MiB
with per-step arrays. Verdict: go for Phase 2 with the unwrap-once rule;
the <= 3 % elementwise budget is re-checked on the real vector kernels at
S1.4/S1.7. Android half pending the reference device.

Also: -PjmhIncludes=<regex> to run a subset of the JMH suite, and the
runFlatRssSpike JavaExec task.

Closes #1016

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@michalharakal

Copy link
Copy Markdown
Contributor Author

Local gate scripts/pr-gate.sh (JDK 25) on 62fd1c2:

=== pr-gate: JVM tests ===                                      BUILD SUCCESSFUL
=== pr-gate: apiCheck ===                                       BUILD SUCCESSFUL
=== pr-gate: verifyNpmPins jsTest wasmJsTest wasmWasiTest ===   BUILD SUCCESSFUL (first run hit the pre-existing timing test `SlicingTest.testPerformanceAccessPatterns` on JS/Chrome — unrelated to this change, passed on re-run)
=== pr-gate: linuxX64Test ===                                   BUILD SUCCESSFUL
=== pr-gate: assemble ===                                       BUILD SUCCESSFUL
=== pr-gate: :skainet-test:skainet-test-java:test ===           BUILD SUCCESSFUL
pr-gate: all legs passed.

@github-actions

Copy link
Copy Markdown

📖 Documentation Preview

The documentation has been built successfully for this PR.

Generated Files:

  • Operator documentation: docs/modules/operators/_generated_/
  • JSON schema output: operators.json

Artifacts:

  • Download the documentation-preview-1060 artifact to view the complete documentation locally.

This comment will be updated automatically when the PR is updated.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[S1.0] P2 spike (decision #6): prototype TensorView/Storage access-path benchmark on JVM + Android, flat-RSS loop — report only

1 participant