feat(memory): M2 acceptance — process probe, ternary decode harness, BitNet-2B plan, release record (SKEEP-003 M2, S2.10) - #1093
Merged
Conversation
…ness, BitNet-2B plan, release record Closes #1042 (SKEEP-003 M2, S2.10; M2-A1, M2-A3, M2-A6). The milestone's criteria are about what the *operating system* sees, and nothing in the repository could see it. This adds the probe, the harness that exercises the M2 machinery, the arithmetic for the model M2 is named after, and the record of what is and is not closed. - `MemoryProbe` + `ProcessMemorySample`: RSS and major/minor page faults from `/proc/self` on JVM, Android and Kotlin/Native-on-Linux; `null` where the platform cannot answer (a browser has no /proc, and callers print "—" rather than guessing). Emits the `Counters.RSS` and `Counters.PAGE_FAULTS` that #1035 declared and nothing produced. - `TernaryDecodeHarness`: the M1 harness in M2 shape — TQ2_0 weights in a ModelScope, `I8Absmax` requantization into the Forward scope every step, dispatch onto `bitnet_gemv`, a KV ring that wraps. - `M2AcceptanceTest`: memory flat across steps with ternary weights, the weights costing exactly 66 bytes per 256 elements, one activation adapter per matmul and nothing else, zero major faults in steady state, and a resident set that does not grow with the step count. - `BitNet2BPlanTest`: M2-A1's arithmetic from BitNet-b1.58-2B-4T's geometry — 1.19 GB resident with a quantized cache, 1.29 GB with bf16, both under 1.3 GB; the margin is 148 MB against 42 MB, and at ctx 8192 only the quantized cache still fits, which is why the mobile profile quantizes it automatically. - `docs/design/memory/m2-acceptance.md`: the M1 and M2 tables side by side, each row saying met / partial / open and where the number came from. Measured on an ARMv8.2 Cortex-A55 reference board (2 cores, 1.9 GB RAM), running the M1 and M2 suites from the Kotlin/Native binary — 15 tests, all green: [m2] steps=12 before: rss=14 MB majflt=0 after: rss=14 MB majflt=0 [m2] rss growth: 4 steps → 0 bytes, 48 steps → 0 bytes Two findings the record states plainly rather than burying. BitNet-2B's bf16 embedding table (657 MB) outweighs its entire ternary stack (512 MB), so past a point a "2-bit model" is an embedding-table problem. And a 2 GB device does *not* hold this checkpoint: after the mobile profile's 700 MB reserve, 1.5 GB free leaves 800 MB, and 1.19 GB does not fit in 800 MB however it is staged — the planner refuses before loading, and the test asserts both the refusal and that a 4 GB device does hold it. M2-A5 and the measured half of M2-A1 stay open, for reasons named in the record: packed weights still reach the managed heap until #973 settles the byte-order contract, and decoding a real checkpoint needs the model stack in SKaiNET-transformers. Gate: scripts/pr-gate.sh — all legs passed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
📖 Documentation Preview The documentation has been built successfully for this PR. Generated Files:
Artifacts:
This comment will be updated automatically when the PR is updated. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #1042 · Milestone M2 · PRD M2-A1, M2-A3, M2-A6 · PRD §6.3
The gap this closes
M2's criteria are about what the operating system sees — resident set, page faults — and nothing in the repository could see it. Allocation events say what SKaiNET asked for; they cannot tell you the process is quietly growing, or that mapped weights are being paged back in.
MemoryProbereads RSS and major/minor faults from/proc/selfon JVM, Android and Kotlin/Native-on-Linux, and returnsnullwhere the platform cannot answer (a browser has no/proc; callers print "—" rather than guess). It emits theCounters.RSSandCounters.PAGE_FAULTSthat [S2.2] P4: phase markers + effective-bandwidth metric in the generation loop (core API; emitted by the decode sample) #1035 declared and nothing had produced.TernaryDecodeHarnessis the M1 harness in M2 shape: TQ2_0 weights in aModelScope,I8Absmaxrequantization into the Forward scope every step, dispatch ontobitnet_gemv, a KV ring that wraps.M2AcceptanceTest: memory flat across steps with ternary weights; the weights costing exactly 66 bytes per 256 elements; one activation adapter per matmul and nothing else; zero major faults in steady state; a resident set that does not grow with the step count.BitNet2BPlanTest: M2-A1's arithmetic, computed from BitNet-b1.58-2B-4T's geometry the wayskainet-plancomputes it from a header.docs/design/memory/m2-acceptance.md: M1 and M2 side by side, each row saying met / partial / open and where the number came from.Measured on the reference board
An ARMv8.2 Cortex-A55, two cores, 1.9 GB RAM — a 2 GB-class device, which is the class M2 targets. The M1 and M2 suites ran there from the Kotlin/Native binary: 15 tests, all green.
Zero major faults — nothing went to disk during steady-state decode — and the resident set is identical after forty-eight steps and after four. That is M1-A1's flatness and M2-A3's fault rate, observed in the kernel's own numbers instead of inferred from allocation events.
Two findings the record states plainly
BitNet-2B's embedding table outweighs its ternary stack. 657 MB of bf16 embeddings against 512 MB of 1.58-bit weights. Past a point, a "2-bit model" is an embedding-table problem — and the planner says so before anything is loaded, which is what M0 was for.
A 2 GB device does not hold this checkpoint. With the mobile profile's 700 MB reserve, 1.5 GB free leaves 800 MB, and 1.19 GB does not fit in 800 MB however it is staged. The test asserts the refusal and that a 4 GB device holds it with mapped weights. M2's title is aspirational for this particular model; the machinery it names is in place and measured.
Resident totals: 1.19 GB with a quantized KV cache at ctx 2048, 1.29 GB with bf16 — both under the 1.3 GB budget, but with 148 MB of headroom against 42 MB. At ctx 8192 only the quantized cache still fits, which is exactly why #1039's mobile profile quantizes it automatically rather than leaving it to the caller.
What stays open, and why
ByteArrays and a buffer-aware kernel needs the byte-order contract of Packed-quant byte-order (block layout) is an unwritten, contradictory contract across the engine and downstream converters #973 settled first. Decoding a real checkpoint also needs the model stack in SKaiNET-transformers, which this repository deliberately does not have.Each of these is a row in the record with the blocker named, not a silent omission.
Gate
scripts/pr-gate.sh— all legs passed. The new tests run on JVM, JS, Wasm and Kotlin/Native, plus the reference board.Keeps develop green by
Records and tests, plus one new
expect objectwith an actual for every target. Nothing existing changes behaviour.🤖 Generated with Claude Code