bench(spike): Phase-2 TensorView/Storage access-path spike — JMH benchmarks, flat-RSS loop, report (SKEEP-003 decision #6, S1.0) - #1060
Merged
Conversation
…hmarks, flat-RSS loop, report (SKEEP-003 decision #6, #1016) Throw-away model of the proposed Storage (heap FloatArray | off-heap MemorySegment) and TensorView (offset/length over a storage) measured against raw arrays: elementwise add over 1 M floats and gemv 256x1024, each as raw / view-get / view-unwrap-once over heap and off-heap, plus a raw-with-offsets control; and FlatRssSpike, a 3000-step decode-shaped loop with a recycled Arena.ofShared bump slab vs fresh heap arrays. Findings (docs/design/memory/spike-p2-access-path-2026-08-23.md): the view indirection is within budget when unwrapped once per call (gemv +0..4 %); offset-based array indexing, not the view, costs 1.9x on the vectorizable elementwise loop (C2 SuperWord); per-element get() is the slow reference path (+80 % heap, 5.5x segment); off-heap activations need Vector-API-over-segment kernels, so Scope.Forward on the JVM should default to heap slabs; the slab keeps RSS at 62-86 MiB over 3000 steps vs 214 MiB with per-step arrays. Verdict: go for Phase 2 with the unwrap-once rule; the <= 3 % elementwise budget is re-checked on the real vector kernels at S1.4/S1.7. Android half pending the reference device. Also: -PjmhIncludes=<regex> to run a subset of the JMH suite, and the runFlatRssSpike JavaExec task. Closes #1016 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Contributor
Author
|
Local gate |
|
📖 Documentation Preview The documentation has been built successfully for this PR. Generated Files:
Artifacts:
This comment will be updated automatically when the PR is updated. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
SKEEP-003 slice S1.0 (milestone M1 #1002, decision #6, #1016): the Phase-2 access-path spike — what does a
TensorViewlayer between a kernel and its bytes cost on the JVM? Report:docs/design/memory/spike-p2-access-path-2026-08-23.md; benchmarks kept reproducible in the JMH module (sk.ainet.bench.spike, run with-PjmhIncludes='spike.*'andrunFlatRssSpike).Verdict: go for Phase 2, with the access-path rule written into the M1 slices — kernels take
TensorViews and unwrap once per call (asHeapArray()/segment()); decodingget()is the reference path only; JVMScope.Forwarddefaults to heap slabs; vector kernels keep explicit offsets.Numbers (i7-9750H, JDK 25, JMH fork 1 / 3 warmup / 5 it):
getper element 0 % · off-heap unwrap +3 % · off-heapget+4 % — within budget.get516 · view-segment unwrap 495 · view-segmentget1579. The 1.9× on the vectorizable loop comes from C2 not auto-vectorizing the offset form, not from the view (view-unwrap ≈ raw-offset). SKaiNET's vector kernels take offsets explicitly, so the ≤ 3 % elementwise budget is re-checked on them at S1.4 / S1.7 against the baseline.MemorySegment.getAtIndexloops are 1.7–5.5× slower than arrays → activations stay on the heap by default on the JVM; mapped/off-heap weights are fine (already consumed byByteVector.fromMemorySegment).-Xmx256m): recycled bump slab 62–86 MiB with zero slab bytes allocated after warm-up vs 214 MiB plateau with per-step heap arrays; the slab's residual creep is the one view object perallocate()(193/step) — cheap, not free.Also in this PR:
-PjmhIncludes=<regex>to run a subset of the JMH suite;runFlatRssSpikeJavaExec task.Open: the Android half (Cortex-A55-class device, direct
ByteBuffer) is pending the reference-device choice (PRD §9); recorded on #1016 when available — the JVM result already sets the rule.Test plan
Full local gate (
scripts/pr-gate.sh, JDK 25); results in the first comment. Benchmarks themselves are not run by the gate (JMH source set); they were run for the report as described.Closes #1016
🤖 Generated with Claude Code