docs(skeep): SKEEP-002 draft — off-heap tensor storage on Android (#921) - #926
Merged
michalharakal merged 2 commits intoAug 10, 2026
Merged
Conversation
Design proposal for file-backed (mmap) weight storage on Android, removing the ART managed-heap cap as the model-size ceiling. Four phases: AndroidMappedMemoryChunk (FileChannel.map, windowed past the 2 GiB per-mapping limit), a readable BufferHandle.FileBacked (today it is inert metadata — copyMaterialize throws), an opt-in mapped placement in the streaming GGUF loader (pass-through layouts only), and a buffer-aware kernel SPI aligned with the #920 JNI bridge for zero-copy consumption. Registered in nav.adoc and the index proposal table as Draft. Refs #921 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
📖 Documentation Preview The documentation has been built successfully for this PR. Generated Files:
Artifacts:
This comment will be updated automatically when the PR is updated. |
Phase 3 was written GGUF-only; the capability is format-agnostic and the SafeTensors streaming reader has the identical shape (per-tensor data_offsets ranges, loadTensorData copying to a heap ByteArray) — and is structurally the easier case: raw dense payloads, no block structure, and the KEEP_NATIVE narrow-float path already wraps on-disk bytes verbatim. Restructure phase 3 as an io-core-level capability with per-format adoption: GGUF first (the measured mobile path), SafeTensors as immediate follow-up on the same core API, ONNX explicitly out of scope (tensors are embedded in protobuf, contiguous region mapping does not apply). Current State documents both readers' copy-out behavior; acceptance criteria gain a both-formats mapped-vs-heap parity requirement. Refs #921 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
📖 Documentation Preview The documentation has been built successfully for this PR. Generated Files:
Artifacts:
This comment will be updated automatically when the PR is updated. |
aharakal
approved these changes
Aug 10, 2026
michalharakal
added a commit
that referenced
this pull request
Aug 10, 2026
Resolve nav.adoc and index.adoc conflicts with the merged SKEEP-002 (PR #926): keep both proposal entries in numeric order.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Design proposal (Status: Draft) for #921 — file-backed (
mmap) weight storage on Android, removing the ART managed-heap cap as the model-size ceiling. Docs-only PR: the SKEEP + its registration innav.adocand the index table, per the CONTRIBUTING SKEEP procedure.Shape of the proposal
Four phases, each independently shippable, all additive (default behavior unchanged on every platform):
AndroidMappedMemoryChunk—FileChannel.map(READ_ONLY)implementation ofMappedMemoryChunkinskainet-io-coreandroidMain (following theAndroidRandomAccessSourceprecedent from fix(io): implement createRandomAccessSource on Android — streaming loads instead of full-file OOM (#922) #924), windowed past the 2 GiB per-mapping limit.BufferHandle.FileBacked— today it is inert metadata (copyMaterialize()throws); it gains a real accessor path, makingMemoryDomain.MMAP_FILEreachable.loadTensorDataMappedhands tensor regions through as mapped views; applies only to pass-through layouts (GGUF's on-disk block layout is the packed layout for Q8_0/Q4_K), repacking/dequant paths keep materializing.ByteBufferoverload (copy-and-delegate default, existing providers untouched). Stands alone; compounds with the native-kernel work proposed in Ship the aarch64-verified NEON kernels to mobile: Apple targets + Android JNI for skainet-backend-native-cpu (measured 21 → 1.0 → 0.11 tok/s cliff) #920, on which nothing here depends.Motivation is measured, not estimated: SmolLM2-135M Q8_0 (138 MiB) costs ~145 MB resident managed heap today and a 1 GB model is impossible by construction on a 512 MB heap, while llama.cpp/ORT/TFLite load the same files on the same phones. Acceptance criteria in the doc are concrete (<40 MB heap delta for the SmolLM2 load; a ~600 MB Q4_K model on a default 256 MB heap; mapped-vs-heap decode parity within 10% warm).
Feedback wanted on the three open questions:
ByteBuffervs anexpect PlatformBuffer(would generalize to posixmmapfor Apple/Linux native), one windowed mapping of the data section vs per-tensor mappings, and whether mapped placement should eventually become the Android default.We can measure any phase on physical devices as it lands.
Refs #921
🤖 Generated with Claude Code