Skip to content

docs(skeep): SKEEP-002 draft — off-heap tensor storage on Android (#921) - #926

Merged
michalharakal merged 2 commits into
developfrom
feature/skeep-002-android-offheap-storage
Aug 10, 2026
Merged

michalharakal merged 2 commits into
developfrom
feature/skeep-002-android-offheap-storage

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Design proposal (Status: Draft) for #921 — file-backed (mmap) weight storage on Android, removing the ART managed-heap cap as the model-size ceiling. Docs-only PR: the SKEEP + its registration in nav.adoc and the index table, per the CONTRIBUTING SKEEP procedure.

Shape of the proposal

Four phases, each independently shippable, all additive (default behavior unchanged on every platform):

  1. AndroidMappedMemoryChunk — FileChannel.map(READ_ONLY) implementation of MappedMemoryChunk in skainet-io-core androidMain (following the AndroidRandomAccessSource precedent from fix(io): implement createRandomAccessSource on Android — streaming loads instead of full-file OOM (#922) #924), windowed past the 2 GiB per-mapping limit.
  2. Readable BufferHandle.FileBacked — today it is inert metadata (copyMaterialize() throws); it gains a real accessor path, making MemoryDomain.MMAP_FILE reachable.
  3. Opt-in mapped placement in the streaming GGUF loader — loadTensorDataMapped hands tensor regions through as mapped views; applies only to pass-through layouts (GGUF's on-disk block layout is the packed layout for Q8_0/Q4_K), repacking/dequant paths keep materializing.
  4. Buffer-aware kernel SPI — defaulted ByteBuffer overload (copy-and-delegate default, existing providers untouched). Stands alone; compounds with the native-kernel work proposed in Ship the aarch64-verified NEON kernels to mobile: Apple targets + Android JNI for skainet-backend-native-cpu (measured 21 → 1.0 → 0.11 tok/s cliff) #920, on which nothing here depends.

Motivation is measured, not estimated: SmolLM2-135M Q8_0 (138 MiB) costs ~145 MB resident managed heap today and a 1 GB model is impossible by construction on a 512 MB heap, while llama.cpp/ORT/TFLite load the same files on the same phones. Acceptance criteria in the doc are concrete (<40 MB heap delta for the SmolLM2 load; a ~600 MB Q4_K model on a default 256 MB heap; mapped-vs-heap decode parity within 10% warm).

Feedback wanted on the three open questions: ByteBuffer vs an expect PlatformBuffer (would generalize to posix mmap for Apple/Linux native), one windowed mapping of the data section vs per-tensor mappings, and whether mapped placement should eventually become the Android default.

We can measure any phase on physical devices as it lands.

Refs #921

🤖 Generated with Claude Code

Design proposal for file-backed (mmap) weight storage on Android,
removing the ART managed-heap cap as the model-size ceiling. Four
phases: AndroidMappedMemoryChunk (FileChannel.map, windowed past the
2 GiB per-mapping limit), a readable BufferHandle.FileBacked (today it
is inert metadata — copyMaterialize throws), an opt-in mapped placement
in the streaming GGUF loader (pass-through layouts only), and a
buffer-aware kernel SPI aligned with the #920 JNI bridge for zero-copy
consumption.

Registered in nav.adoc and the index proposal table as Draft.

Refs #921

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

📖 Documentation Preview

The documentation has been built successfully for this PR.

Generated Files:

  • Operator documentation: docs/modules/operators/_generated_/
  • JSON schema output: operators.json

Artifacts:

  • Download the documentation-preview-926 artifact to view the complete documentation locally.

This comment will be updated automatically when the PR is updated.

Phase 3 was written GGUF-only; the capability is format-agnostic and the
SafeTensors streaming reader has the identical shape (per-tensor
data_offsets ranges, loadTensorData copying to a heap ByteArray) — and is
structurally the easier case: raw dense payloads, no block structure, and
the KEEP_NATIVE narrow-float path already wraps on-disk bytes verbatim.

Restructure phase 3 as an io-core-level capability with per-format
adoption: GGUF first (the measured mobile path), SafeTensors as immediate
follow-up on the same core API, ONNX explicitly out of scope (tensors are
embedded in protobuf, contiguous region mapping does not apply). Current
State documents both readers' copy-out behavior; acceptance criteria gain
a both-formats mapped-vs-heap parity requirement.

Refs #921

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

📖 Documentation Preview

The documentation has been built successfully for this PR.

Generated Files:

  • Operator documentation: docs/modules/operators/_generated_/
  • JSON schema output: operators.json

Artifacts:

  • Download the documentation-preview-926 artifact to view the complete documentation locally.

This comment will be updated automatically when the PR is updated.

@michalharakal
michalharakal merged commit 12305f0 into develop Aug 10, 2026
15 checks passed
@michalharakal
michalharakal deleted the feature/skeep-002-android-offheap-storage branch August 10, 2026 08:57
@michalharakal
michalharakal requested review from aharakal and removed request for aharakal August 10, 2026 09:01
michalharakal added a commit that referenced this pull request Aug 10, 2026
Resolve nav.adoc and index.adoc conflicts with the merged SKEEP-002
(PR #926): keep both proposal entries in numeric order.
michalharakal added a commit that referenced this pull request Aug 10, 2026
Resolve the [Unreleased] CHANGELOG conflict with the merged #930 entry
(the earlier remote merge predated PRs #934/#926): keep all bullets.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants