Skip to content

M2-A5: a Q4_K_M model loading under a real heap cap on an Android device #1130

Description

@michalharakal

Split out of #932's closing summary, where this was recorded as deliberately open but never given an issue.

What exists and what does not

The mechanism is built: mapped staging (#921/#922), the two-pool DeviceFit check, PlannerProfile.MOBILE_2GB — which is now strict (#1128), so a missing kernel fails rather than silently costing several times the weight. MemoryProbe reads RSS and fault counters on every target.

What is missing is the measurement, on hardware with ART. A JVM host is not a substitute: the heap cap, the GC and the page-cache behaviour that M2-A5 is about are all ART's.

Acceptance

  • A Q4_K_M model loads and decodes on an Android device under its real heap cap
  • RSS and major-fault counts recorded across a decode run
  • The MemoryPlan compared against what was actually allocated, within the plan-vs-actual tolerance
  • Numbers recorded as measurements, not turned into CI assertions — a shared runner is not the instrument for this claim (test: judge page faults by whether they scale with steps, not by a flat zero #1107)

Note

A suitable device is available; see the maintainer's notes rather than this issue for addressing.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    assessmentAssessment task (DARC: A)platformPlatform support (WASM, JVM, native)

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions