Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions .github/workflows/build.yml
Original file line number Diff line number Diff line change
Expand Up @@ -41,6 +41,12 @@ jobs:
tasks: verifyNpmPins jsTest wasmJsTest wasmWasiTest
- name: native
tasks: linuxX64Test
# Host-side (JVM) unit tests of the *Android* compilations: the mmap weight path (#921)
# and the Android loading facade + device fit check (#1038, SKEEP-002). They compile
# against androidMain, so no other leg proves that code builds — assemble compiles it but
# runs nothing.
- name: android
tasks: testAndroidHostTest
# golden-parity: the SKEEP-003 packed-encoding gate (#1005). Bit-identical decode /
# scalar-kernel / TurboQuant digests on JVM and Kotlin/Native plus the dispatch parity
# tests, and the binary-compatibility check (apiCheck) that the contributing docs require
Expand Down
38 changes: 38 additions & 0 deletions docs/modules/skeep/pages/002-android-offheap-tensor-storage.adoc
Original file line number Diff line number Diff line change
Expand Up @@ -232,6 +232,44 @@ keeps working unchanged and the overload is purely opt-in. Expected consumers:
(packed quant blocks) and a synthesized SafeTensors file (dense
F32/F16/BF16) produce bit-identical tensors under either placement.

== Implementation status (2026-08-24)

Tracked by SKEEP-003's M2 slice
https://github.com/SKaiNET-developers/SKaiNET/issues/1038[#1038]. This SKEEP stays *Draft*: the
mechanism is in `develop`, the acceptance criteria that need a physical device are not met yet, and
one of them cannot be met until an unrelated contract is fixed.

Landed:

* *Phase 1 and 2* — `JvmMappedMemoryChunk` / `MappedRandomAccessSource` and a readable
`BufferHandle.FileBacked` with `JvmFileBackedResolver`, shared between the JVM and Android
compilations, covered by `androidHostTest`
(https://github.com/SKaiNET-developers/SKaiNET/issues/921[#921],
https://github.com/SKaiNET-developers/SKaiNET/issues/922[#922]).
* *Phase 3* — mapped loading is a *configuration of the ordinary loader* rather than a separate
helper: `StreamingGgufParametersLoader(staging = StagingPolicy.MAPPED)`
(https://github.com/SKaiNET-developers/SKaiNET/issues/1037[#1037]), with `AndroidGguf.loader()`
making it the Android default.
* *The fit check this SKEEP did not have* — `MemoryPlan.fitOn(DeviceMemory, weightsMapped)` treats
a phone as two pools (the ART cap and physical RAM) instead of one total, and
`AndroidGguf.fits(context, path, ctx)` answers from the GGUF header before a byte of payload is
read, naming the pool that runs out and what to do about it.

Not landed, and why:

* *Packed weights still reach the managed heap.* Mapped staging serves dense F32 tensors as
zero-heap views, but every quantized tensor is still handed to its kernel as a `ByteArray`,
because the packed matmul SPI takes arrays. That is Phase 4, and it is blocked on
https://github.com/SKaiNET-developers/SKaiNET/issues/973[#973] — the packed-quant byte-order
contract — since a buffer-aware kernel first needs an unambiguous answer to *which* byte order it
is reading. Until then the heap ceiling is lifted for dense checkpoints, not for a Q4_K_M one,
which is exactly the acceptance criterion below that remains open.
* *The device numbers.* "≤ 40 MB managed heap for SmolLM2-135M Q8_0" and "a ~600 MB Q4_K model
loads on a 256 MB heap" are measurements on a physical device; the repository's CI has none. They
belong to the `skainet-decode` sample in SKaiNET-transformers, which owns a model and a device
lane. `androidHostTest` covers what can be proven without hardware: the Android compilation
builds, maps, loads, and produces bit-identical tensors under either staging.

== Risks

* *Page-fault latency.* First-touch of cold pages during decode adds jitter.
Expand Down
6 changes: 6 additions & 0 deletions scripts/pr-gate.sh
Original file line number Diff line number Diff line change
Expand Up @@ -56,6 +56,12 @@ step "assemble"
step "Java consumer API tests"
"${GRADLE[@]}" :skainet-test:skainet-test-java:test

# The Android compilations have host-side (JVM) unit tests — the mmap weight path (#921) and the
# Android loading facade (#1038). They compile against androidMain, so they are the only thing that
# proves that code builds and runs on the Android variant; nothing else in the gate touches it.
step "Android host tests"
"${GRADLE[@]}" testAndroidHostTest

if [[ "$mode" == "--bench" ]]; then
step "benchmarks (compare against the committed baseline before/after)"
"${GRADLE[@]}" :skainet-lang:skainet-lang-core:jvmBenchmark
Expand Down
9 changes: 9 additions & 0 deletions skainet-io/skainet-io-gguf/build.gradle.kts
Original file line number Diff line number Diff line change
Expand Up @@ -86,5 +86,14 @@ kotlin {
implementation(libs.kotlinx.coroutines.test)
}
}

// Host-side tests of the *Android* compilation (#1038): they drive the suspending loader,
// so they need coroutines like jvmTest does. The GGUF fixtures they write are their own —
// jvmTest's SyntheticGguf is not visible from this compilation.
getByName("androidHostTest") {
dependencies {
implementation(libs.kotlinx.coroutines)
}
}
}
}
Original file line number Diff line number Diff line change
@@ -0,0 +1,145 @@
package sk.ainet.io.gguf

import kotlinx.coroutines.runBlocking
import sk.ainet.context.DefaultDataExecutionContext
import sk.ainet.io.model.QuantPolicy
import sk.ainet.io.model.StagingPolicy
import sk.ainet.lang.memory.ExperimentalMemoryApi
import sk.ainet.lang.memory.plan.DeviceMemory
import sk.ainet.lang.tensor.Tensor
import sk.ainet.lang.tensor.data.FloatArrayTensorData
import sk.ainet.lang.tensor.data.MmapFloatTensorData
import sk.ainet.lang.types.FP32
import java.io.File
import java.io.RandomAccessFile
import java.nio.ByteBuffer
import java.nio.ByteOrder
import kotlin.test.Test
import kotlin.test.assertContentEquals
import kotlin.test.assertEquals
import kotlin.test.assertFalse
import kotlin.test.assertTrue

/**
* Host-side test of the *Android compilation* (#1038, SKEEP-002): this source set compiles against
* androidMain, so it proves the Android loading facade builds and behaves on the Android variant —
* mapped staging by default, and a fit check that answers before the load rather than during it.
*
* `AndroidGguf.deviceMemory(context)` needs a real `Context` and belongs to the instrumented smoke
* test; everything downstream of it takes a [DeviceMemory] so it can be checked here.
*/
@OptIn(ExperimentalMemoryApi::class)
class AndroidGgufLoadingHostTest {

private val mb = 1024L * 1024L

/**
* A one-tensor GGUF v3 file with a known F32 payload. Written here rather than reused from
* jvmTest's `SyntheticGguf`, which this compilation cannot see.
*/
private fun model(elements: Int = 4096): File {
val file = File.createTempFile("android-gguf-", ".gguf")
file.deleteOnExit()
val head = ByteBuffer.allocate(4096).order(ByteOrder.LITTLE_ENDIAN)
head.putInt(0x46554747) // "GGUF"
head.putInt(3) // version
head.putLong(1) // tensor count
head.putLong(1) // kv count
val key = "general.architecture".encodeToByteArray()
head.putLong(key.size.toLong()); head.put(key)
head.putInt(GGUFValueType.STRING.value)
val value = "test".encodeToByteArray()
head.putLong(value.size.toLong()); head.put(value)
val name = "w_f32".encodeToByteArray()
head.putLong(name.size.toLong()); head.put(name)
head.putInt(1) // rank
head.putLong(elements.toLong())
head.putInt(GGMLQuantizationType.F32.value)
head.putLong(0L) // data offset
val padding = (32 - (head.position() % 32)) % 32
repeat(padding) { head.put(0) }

RandomAccessFile(file, "rw").use { raf ->
raf.write(head.array(), 0, head.position())
val payload = ByteBuffer.allocate(elements * 4).order(ByteOrder.LITTLE_ENDIAN)
repeat(elements) { payload.putFloat(it * 0.5f) }
raf.write(payload.array())
}
return file
}

private fun load(loader: StreamingGgufParametersLoader): Map<String, Tensor<FP32, Float>> {
val ctx = DefaultDataExecutionContext()
val out = LinkedHashMap<String, Tensor<FP32, Float>>()
runBlocking { loader.load<FP32, Float>(ctx, FP32::class) { name, t -> out[name] = t } }
return out
}

@Test
fun `the android loader maps weights by default`() {
val f = model()
try {
val mapped = load(AndroidGguf.loader(f.absolutePath))
assertTrue(
mapped.getValue("w_f32").data is MmapFloatTensorData<*>,
"dense F32 must come from file-backed pages on Android, got ${mapped.getValue("w_f32").data::class.simpleName}",
)
// and the heap path is still reachable, producing the same numbers
val onHeap = load(AndroidGguf.loader(f.absolutePath, staging = StagingPolicy.HEAP))
assertTrue(onHeap.getValue("w_f32").data is FloatArrayTensorData<*>)
assertContentEquals(
onHeap.getValue("w_f32").data.copyToFloatArray(),
mapped.getValue("w_f32").data.copyToFloatArray(),
"staging must not change the numbers",
)
val values = mapped.getValue("w_f32").data.copyToFloatArray()
assertEquals(0f, values[0]); assertEquals(0.5f, values[1]); assertEquals(2047.5f, values[4095])
} finally {
f.delete()
}
}

@Test
fun `the fit check reads the plan from the header and answers before loading`() {
val f = model()
try {
val plentiful = DeviceMemory(
totalRamBytes = 4096 * mb, availableRamBytes = 2048 * mb,
heapMaxBytes = 512 * mb, heapUsedBytes = 32 * mb, lowMemoryThresholdBytes = 180 * mb,
)
val fit = AndroidGguf.fits(plentiful, f.absolutePath, ctx = 128)
assertTrue(fit.fits, fit.render())
assertTrue(fit.weightsMapped, "the Android default is mapped weights")
assertTrue(fit.plan.weightsBytes > 0, "the plan comes from the header")

// a phone with almost nothing left says so, and says which pool ran out
val squeezed = plentiful.copy(availableRamBytes = 190 * mb, heapMaxBytes = 8 * mb, heapUsedBytes = 7 * mb)
val tight = AndroidGguf.fits(squeezed, f.absolutePath, ctx = 4096)
assertFalse(tight.fits, tight.render())
assertEquals("managed heap", tight.blockingPool)
assertTrue(tight.suggestions.isNotEmpty(), "a failing fit must say what to do")
} finally {
f.delete()
}
}

@Test
fun `an unmapped load is charged for its weights, a mapped one is not`() {
val f = model()
try {
val device = DeviceMemory(
totalRamBytes = 2048 * mb, availableRamBytes = 900 * mb,
heapMaxBytes = 512 * mb, heapUsedBytes = 40 * mb, lowMemoryThresholdBytes = 180 * mb,
)
val mapped = AndroidGguf.fits(device, f.absolutePath, ctx = 512, weightsMapped = true)
val heap = AndroidGguf.fits(device, f.absolutePath, ctx = 512, weightsMapped = false)
assertEquals(
mapped.plan.weightsBytes,
heap.heap.neededBytes - mapped.heap.neededBytes,
"the difference between the two is exactly the weights",
)
} finally {
f.delete()
}
}
}
Original file line number Diff line number Diff line change
@@ -0,0 +1,97 @@
package sk.ainet.io.gguf

import android.app.ActivityManager
import android.content.Context
import sk.ainet.io.RandomAccessSource
import sk.ainet.io.openRandomAccessSource
import sk.ainet.io.model.QuantPolicy
import sk.ainet.io.model.StagingPolicy
import sk.ainet.lang.memory.ExperimentalMemoryApi
import sk.ainet.lang.memory.plan.Budget
import sk.ainet.lang.memory.plan.DeviceFit
import sk.ainet.lang.memory.plan.DeviceMemory
import sk.ainet.lang.memory.plan.MemoryPlan
import sk.ainet.lang.memory.plan.MemoryPlans
import sk.ainet.lang.memory.plan.fitOn

/**
* Loading a GGUF on Android: mapped weights by default, and a fit check *before* the load
* (SKEEP-002, #921, #922, #1038).
*
* The managed heap is the binding constraint on a phone — hard-capped at 256 MB (512 MB with
* `largeHeap`) no matter how much RAM the device has — so the Android configuration of the loader
* is `staging = MAPPED`: weights come from file-backed pages the OS pages in on demand and evicts
* under pressure, and never count against the cap.
*
* What is *not* solved yet: packed (quantized) tensors still arrive as heap arrays, because the
* packed kernels take `ByteArray`s until the view contract of #973 lands. Mapping therefore lifts
* the ceiling for dense-F32 weights today, and for a Q4_K_M checkpoint only once #973 does. The
* fit check tells you which of the two pools you are about to run out of, rather than letting the
* app find out by being killed.
*/
@OptIn(ExperimentalMemoryApi::class)
public object AndroidGguf {

/**
* The loader Android should use: positional reads for the metadata, mapped pages for tensor
* payloads. [quantPolicy] is the caller's choice as usual; [staging] defaults to
* [StagingPolicy.MAPPED] and is a parameter only so a test or a benchmark can ask for the
* heap path explicitly.
*/
public fun loader(
filePath: String,
quantPolicy: QuantPolicy = QuantPolicy.NATIVE_OPTIMIZED,
staging: StagingPolicy = StagingPolicy.MAPPED,
onProgress: (current: Long, total: Long, message: String?) -> Unit = { _, _, _ -> },
): StreamingGgufParametersLoader = StreamingGgufParametersLoader(
sourceProvider = { openSource(filePath) },
onProgress = onProgress,
quantPolicy = quantPolicy,
staging = staging,
)

/**
* What this device has to offer: `ActivityManager.MemoryInfo` for physical RAM plus the ART
* heap cap, which is what actually stops a model from loading.
*/
public fun deviceMemory(context: Context): DeviceMemory {
val am = context.getSystemService(Context.ACTIVITY_SERVICE) as ActivityManager
val info = ActivityManager.MemoryInfo()
am.getMemoryInfo(info)
val runtime = Runtime.getRuntime()
return DeviceMemory(
totalRamBytes = info.totalMem,
availableRamBytes = info.availMem,
heapMaxBytes = runtime.maxMemory(),
heapUsedBytes = runtime.totalMemory() - runtime.freeMemory(),
lowMemory = info.lowMemory,
lowMemoryThresholdBytes = info.threshold,
)
}

/**
* The plan a GGUF's *header* predicts at [ctx] — tensor table and metadata only, no payload
* (M0-F1), so this costs a few kilobytes and a couple of reads.
*/
public fun plan(filePath: String, ctx: Int, budget: Budget? = null): MemoryPlan =
openSource(filePath).use { source ->
MemoryPlans.plan(StreamingGGUFReader.open(source).planInput(ctx), budget)
}

private fun openSource(filePath: String): RandomAccessSource =
openRandomAccessSource(filePath)
?: throw IllegalArgumentException("Cannot open for random access: $filePath")

/**
* Will this model load on this device? Checks the header-derived plan against both pools —
* managed heap and physical RAM — before a byte of payload is read.
*
* @param weightsMapped whether the load will use [StagingPolicy.MAPPED] (what [loader] does)
*/
public fun fits(context: Context, filePath: String, ctx: Int, weightsMapped: Boolean = true): DeviceFit =
fits(deviceMemory(context), filePath, ctx, weightsMapped)

/** [fits] against an explicit [DeviceMemory] — the form a test or a simulation uses. */
public fun fits(device: DeviceMemory, filePath: String, ctx: Int, weightsMapped: Boolean = true): DeviceFit =
plan(filePath, ctx).fitOn(device, weightsMapped)
}
Loading
Loading