Skip to content

feat(memory): planner device profiles — 2 GB mobile and desktop, automatic KV quantization, dequant limit (SKEEP-003 P5, S2.6) - #1085

Merged
michalharakal merged 1 commit into
developfrom
feature/1039-planner-2gb-profile
Aug 24, 2026
Merged

michalharakal merged 1 commit into
developfrom
feature/1039-planner-2gb-profile

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Closes #1039 · Phase P5 · Milestone M2 · PRD M2-F6 · Proposal §8 item 1, §10 decision #11

Decision #11 was prose

Its numbers — 700 MB reserve on Android/JVM, 300 MB on Native, prefill chunked at 256, off-heap above 256 KB, TurboQuant KV once the plan passes 80 % of the budget, dequantization over 5 % a warning — lived in a design document. A profile makes them explicit, testable, and impossible to "improve" by feel without a test turning red.

A 2 GB phone and a workstation genuinely want different rules: on the phone the reserve is large relative to the budget, weights must be mapped, the cache has to shrink before the model does, and a dispatcher-inserted dequantization is a defect rather than a slow path.

What a profile decides

PlannerProfile carries the reserve, prefill chunk, KV mode, the budget fraction above which KV auto-quantizes, the heap/off-heap threshold, the dequant warn limit and whether it is strict, and whether weights are mapped:

  • MOBILE_2GB — 700 MB reserved, prefill 256, off-heap ≥ 256 KB, KV → TurboQuant-4 over 80 % of budget, dequant warns over 5 %, weights mapped. The Android default.
  • DESKTOP — today's behaviour: no automatic re-quantization, heap staging.
  • NATIVE — the smaller 300 MB reserve.
  • forDevice(device) — mobile at or below 3 GB of RAM.

profile.plan(input, availableBytes) returns a ProfiledPlan that carries the decisions it made: the KV switch appears as a note rather than a silent rewrite, so a plan can be read back and understood. requireFits(device) refuses before anything is allocated, against #1038's two pools rather than one total.

checkDequant(share) turns #1035's adapterShareOfBytesRead into a verdict — fine, a warning that names the missing kernel, or a failure under strict(). That is the same number the decode loop already measures, so the rule is enforced against reality rather than a guess.

Where it shows up

skainet-plan --profile mobile|desktop|native plans under the rules and prints them above the table:

profile mobile-2gb · reserve 700.0 MB · prefill 256 · off-heap ≥ 256.0 KB · KV auto-quantize over 80.0%
  note: KV cache switched to TurboQuant 4-bit: the plan needed 118.4% of the budget (over 80.0%), saving 1.1 GB
  note: weights are counted resident and mapped — off the managed heap

On Android, AndroidGguf.profiledPlan(...) defaults to MOBILE_2GB, and its host test asserts both that default and the refusal path.

A CI gap this closed on the way

The skainet-plan CLI and the engine benchmark publisher are plain-JVM modules: they have a test task, not jvmTest, so the repo-wide leg never reached them and their tests had never run in CI — including the CLI tests added when the tool landed. The jvm leg now names them explicitly, in build.yml and in scripts/pr-gate.sh. (Same class of gap as the Android host tests in #1038.)

Acceptance — one test per rule

PlannerProfileTest (12 cases): budget arithmetic for each profile including the never-negative floor; every default of MOBILE_2GB and DESKTOP; the off-heap threshold at and just under the boundary; the profile overriding the caller's prefill chunk; KV quantizing only when tight (the same model at the same context length, planned against two budgets — so the comparison is like for like); the desktop profile leaving the cache alone; refusal against a budget and against a device's two pools; dequant OK / warn / strict-fails; and profile selection by RAM. Plus SkainetPlanTest for the CLI flag and the rendered header, and the Android host test for the platform default.

Gate

scripts/pr-gate.sh — all legs passed, including the new JVM-tool leg and the Android leg.

Keeps develop green by

Profile selection is explicit — nothing plans under a profile unless asked — and DESKTOP is today's behaviour by construction. The CLI's --profile defaults to none, which takes exactly the path it took before.

🤖 Generated with Claude Code

… automatic KV quantization and a dequant limit

Closes #1039 (SKEEP-003 P5, S2.6, proposal §8 item 1, decision #11; M2-F6).

Decision #11's numbers existed as prose. A 2 GB phone and a workstation do
not want the same defaults, and the difference is not taste: on the phone
the reserve is large relative to the budget, weights must be mapped, the KV
cache has to shrink before the model does, and a dispatcher-inserted
dequantization is a defect rather than a slow path.

- `PlannerProfile`: reserve, prefill chunk, KV mode, the fraction of the
  budget above which KV auto-quantizes, the heap/off-heap threshold, the
  dequant warn limit and whether it is strict, and whether weights are
  mapped. `MOBILE_2GB` (700 MB reserved, prefill 256, off-heap ≥ 256 KB, KV
  → TurboQuant-4 over 80 %, dequant warns over 5 %, weights mapped),
  `DESKTOP` (today's behaviour: no automatic re-quantization, heap
  staging), `NATIVE` (the smaller 300 MB reserve), and `forDevice()` which
  picks mobile at or below 3 GB of RAM.
- `profile.plan(input, availableBytes)` returns a `ProfiledPlan` carrying
  the decisions it made — the KV switch is a note, not a silent rewrite —
  and `requireFits(device)` refuses before anything is allocated, against
  the two pools of #1038 rather than one total.
- `checkDequant(share)` turns #1035's `adapterShareOfBytesRead` into a
  verdict: fine, a warning naming the missing kernel, or a failure under
  `strict()`.
- `skainet-plan --profile mobile|desktop|native` plans under the rules and
  prints them above the table, so a plan read months later says which rules
  produced it. On Android `AndroidGguf.profiledPlan(...)` defaults to
  `MOBILE_2GB`.

CI and the gate gain a leg for plain-JVM modules: the skainet-plan CLI and
the benchmark publisher have `test`, not `jvmTest`, so the repo-wide leg
never reached them and their tests had never run in CI.

PlannerProfileTest covers one rule per test — budget arithmetic, each
default, the off-heap threshold at and under the boundary, the profile
overriding the caller's prefill chunk, KV quantizing only when tight (same
model, same ctx, two budgets), the desktop profile leaving the cache alone,
refusal against a budget and against a device's two pools, dequant OK/warn/
strict, and profile selection by RAM.

Gate: scripts/pr-gate.sh — all legs passed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

📖 Documentation Preview

The documentation has been built successfully for this PR.

Generated Files:

  • Operator documentation: docs/modules/operators/_generated_/
  • JSON schema output: operators.json

Artifacts:

  • Download the documentation-preview-1085 artifact to view the complete documentation locally.

This comment will be updated automatically when the PR is updated.

@michalharakal
michalharakal merged commit 06a466e into develop Aug 24, 2026
18 checks passed
@michalharakal
michalharakal deleted the feature/1039-planner-2gb-profile branch August 24, 2026 11:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[S2.6] P5: planner 2 GB reference profile — defaults, fit check, TurboQuant KV auto ≥ 80 %, dequant warn/strict, desktop profile

1 participant