Skip to content

feat(memory): M2 acceptance — process probe, ternary decode harness, BitNet-2B plan, release record (SKEEP-003 M2, S2.10) - #1093

Merged
michalharakal merged 1 commit into
developfrom
feature/1042-m2-acceptance
Aug 24, 2026
Merged

michalharakal merged 1 commit into
developfrom
feature/1042-m2-acceptance

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Closes #1042 · Milestone M2 · PRD M2-A1, M2-A3, M2-A6 · PRD §6.3

The gap this closes

M2's criteria are about what the operating system sees — resident set, page faults — and nothing in the repository could see it. Allocation events say what SKaiNET asked for; they cannot tell you the process is quietly growing, or that mapped weights are being paged back in.

  • MemoryProbe reads RSS and major/minor faults from /proc/self on JVM, Android and Kotlin/Native-on-Linux, and returns null where the platform cannot answer (a browser has no /proc; callers print "—" rather than guess). It emits the Counters.RSS and Counters.PAGE_FAULTS that [S2.2] P4: phase markers + effective-bandwidth metric in the generation loop (core API; emitted by the decode sample) #1035 declared and nothing had produced.
  • TernaryDecodeHarness is the M1 harness in M2 shape: TQ2_0 weights in a ModelScope, I8Absmax requantization into the Forward scope every step, dispatch onto bitnet_gemv, a KV ring that wraps.
  • M2AcceptanceTest: memory flat across steps with ternary weights; the weights costing exactly 66 bytes per 256 elements; one activation adapter per matmul and nothing else; zero major faults in steady state; a resident set that does not grow with the step count.
  • BitNet2BPlanTest: M2-A1's arithmetic, computed from BitNet-b1.58-2B-4T's geometry the way skainet-plan computes it from a header.
  • docs/design/memory/m2-acceptance.md: M1 and M2 side by side, each row saying met / partial / open and where the number came from.

Measured on the reference board

An ARMv8.2 Cortex-A55, two cores, 1.9 GB RAM — a 2 GB-class device, which is the class M2 targets. The M1 and M2 suites ran there from the Kotlin/Native binary: 15 tests, all green.

[m2] steps=12  before: rss=14 MB majflt=0 minflt=4813   after: rss=14 MB majflt=0 minflt=5037
[m2] rss growth: 4 steps → 0 bytes, 48 steps → 0 bytes

Zero major faults — nothing went to disk during steady-state decode — and the resident set is identical after forty-eight steps and after four. That is M1-A1's flatness and M2-A3's fault rate, observed in the kernel's own numbers instead of inferred from allocation events.

Two findings the record states plainly

BitNet-2B's embedding table outweighs its ternary stack. 657 MB of bf16 embeddings against 512 MB of 1.58-bit weights. Past a point, a "2-bit model" is an embedding-table problem — and the planner says so before anything is loaded, which is what M0 was for.

A 2 GB device does not hold this checkpoint. With the mobile profile's 700 MB reserve, 1.5 GB free leaves 800 MB, and 1.19 GB does not fit in 800 MB however it is staged. The test asserts the refusal and that a 4 GB device holds it with mapped weights. M2's title is aspirational for this particular model; the machinery it names is in place and measured.

Resident totals: 1.19 GB with a quantized KV cache at ctx 2048, 1.29 GB with bf16 — both under the 1.3 GB budget, but with 148 MB of headroom against 42 MB. At ctx 8192 only the quantized cache still fits, which is exactly why #1039's mobile profile quantizes it automatically rather than leaving it to the caller.

What stays open, and why

Each of these is a row in the record with the blocker named, not a silent omission.

Gate

scripts/pr-gate.sh — all legs passed. The new tests run on JVM, JS, Wasm and Kotlin/Native, plus the reference board.

Keeps develop green by

Records and tests, plus one new expect object with an actual for every target. Nothing existing changes behaviour.

🤖 Generated with Claude Code

…ness, BitNet-2B plan, release record

Closes #1042 (SKEEP-003 M2, S2.10; M2-A1, M2-A3, M2-A6).

The milestone's criteria are about what the *operating system* sees, and
nothing in the repository could see it. This adds the probe, the harness
that exercises the M2 machinery, the arithmetic for the model M2 is named
after, and the record of what is and is not closed.

- `MemoryProbe` + `ProcessMemorySample`: RSS and major/minor page faults
  from `/proc/self` on JVM, Android and Kotlin/Native-on-Linux; `null`
  where the platform cannot answer (a browser has no /proc, and callers
  print "—" rather than guessing). Emits the `Counters.RSS` and
  `Counters.PAGE_FAULTS` that #1035 declared and nothing produced.
- `TernaryDecodeHarness`: the M1 harness in M2 shape — TQ2_0 weights in a
  ModelScope, `I8Absmax` requantization into the Forward scope every step,
  dispatch onto `bitnet_gemv`, a KV ring that wraps.
- `M2AcceptanceTest`: memory flat across steps with ternary weights, the
  weights costing exactly 66 bytes per 256 elements, one activation adapter
  per matmul and nothing else, zero major faults in steady state, and a
  resident set that does not grow with the step count.
- `BitNet2BPlanTest`: M2-A1's arithmetic from BitNet-b1.58-2B-4T's geometry
  — 1.19 GB resident with a quantized cache, 1.29 GB with bf16, both under
  1.3 GB; the margin is 148 MB against 42 MB, and at ctx 8192 only the
  quantized cache still fits, which is why the mobile profile quantizes it
  automatically.
- `docs/design/memory/m2-acceptance.md`: the M1 and M2 tables side by side,
  each row saying met / partial / open and where the number came from.

Measured on an ARMv8.2 Cortex-A55 reference board (2 cores, 1.9 GB RAM),
running the M1 and M2 suites from the Kotlin/Native binary — 15 tests, all
green:

  [m2] steps=12  before: rss=14 MB majflt=0  after: rss=14 MB majflt=0
  [m2] rss growth: 4 steps → 0 bytes, 48 steps → 0 bytes

Two findings the record states plainly rather than burying. BitNet-2B's
bf16 embedding table (657 MB) outweighs its entire ternary stack (512 MB),
so past a point a "2-bit model" is an embedding-table problem. And a 2 GB
device does *not* hold this checkpoint: after the mobile profile's 700 MB
reserve, 1.5 GB free leaves 800 MB, and 1.19 GB does not fit in 800 MB
however it is staged — the planner refuses before loading, and the test
asserts both the refusal and that a 4 GB device does hold it.

M2-A5 and the measured half of M2-A1 stay open, for reasons named in the
record: packed weights still reach the managed heap until #973 settles the
byte-order contract, and decoding a real checkpoint needs the model stack
in SKaiNET-transformers.

Gate: scripts/pr-gate.sh — all legs passed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

📖 Documentation Preview

The documentation has been built successfully for this PR.

Generated Files:

  • Operator documentation: docs/modules/operators/_generated_/
  • JSON schema output: operators.json

Artifacts:

  • Download the documentation-preview-1093 artifact to view the complete documentation locally.

This comment will be updated automatically when the PR is updated.

@michalharakal
michalharakal merged commit 5b4367e into develop Aug 24, 2026
18 checks passed
@michalharakal
michalharakal deleted the feature/1042-m2-acceptance branch August 24, 2026 15:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[S2.10] M2 acceptance: BitNet-b1.58-2B on the reference device — resident ≤ 1.3 GB, flat RSS, page faults ≈ 0; release tables

1 participant