Skip to content

Extend mapped-servable staging beyond Q4_K/Q6_K (Q8_0, Q4_0/Q5_x, Q5_K, BITNET_B1_58) #1192

Description

@michalharakal

Follow-up to #1189/#1190. StorageCapabilities.MAPPED_SERVABLE_DEFAULT is honest but small: dense F32, Q4_K, Q6_K. Every other encoding under WeightResidency.MAPPED heap-stages, and the plan now (correctly) charges it against the heap budget.

Per format the recipe is mechanical (proven by #1190): row-major C variant reading (o·blocksPerRow + b)·bytesPerBlock + JNI direct-buffer shim + BufferPackedTensorData scratch-decoder arm + entry in the servable set + parity test vs the feed-order kernel on permuted bytes.

  • Q8_0 / Q4_0 / Q5_0 / Q5_1 / Q5_K — direct application of the recipe.
  • BITNET_B1_58 (I2_S) — different: the load repacks stock GROUP layouts to sequential ([ternary-f32] Phase 4: GGUF I2_S (type 36) import with group→sequential repack, keep-packed BITNET_B1_58 #1140), and a repacked copy cannot page from the source file. Needs the repack cached once to an app-files sidecar and mapped back on subsequent loads (the NeoGPU-converter SEQUENTIAL case could map directly, minus the scale trailer question). The ternary LUT kernel is already row-major by construction, so no new C is needed — only the direct-buffer entry and the cache file.
  • Keep AllocationResolver.explain / the servable set in lockstep with what loaders actually do — MappedBudgetPlanTest.encodings_nobody_serves_from_a_mapping_stay_heap_charged_even_under_MAPPED is the guard.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions