Skip to content

docs(#1205): why the GROUP_128/GROUP_64 native decode kernel was closed - #1214

Merged
michalharakal merged 1 commit into
developfrom
docs/1205-close-note
Aug 29, 2026
Merged

michalharakal merged 1 commit into
developfrom
docs/1205-close-note

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Closes the doc/code side of #1205 (already closed on GitHub — this PR just makes the reasoning
discoverable in-repo instead of only in a closed issue's comment).

Decision

A native kernel that decodes BitNet.cpp's grouped I2_S layout (`GROUP_128`/`GROUP_64`) directly,
skipping the repack, was considered and closed as not planned — AOT conversion
(`I2sAotConverter`, #1207) already gets a converted file to the same zero-copy mmap path (#1203)
for a fraction of the ongoing cost of maintaining a second decode variant per SIMD backend. Full
reasoning, and what such a kernel would actually need if a real use case someday can't tolerate an
AOT step, is in #1205's closing comment.

What this PR adds

  • `I2sGgufLayout.GROUP_128`/`GROUP_64` doc comments in `I2sRepack.kt` now point to this decision
    instead of leaving it undocumented in the code someone would actually be reading when they hit
    this limitation.
  • `ternary-getting-started.adoc` gains a "Loading a BitNet.cpp-quantized GGUF" section explaining
    the grouped-vs-sequential layout distinction and pointing at `I2sAotConverter` as the supported
    fix, plus a "Status and roadmap" update noting 0.51.0's off-heap/mmap/AOT work and [Deferred] Native decode kernel for BitNet.cpp's GROUP_128 layout (no-repack mmap) #1205's
    closure.

Test plan

  • `skainet-io-gguf` compiles clean (doc-comment-only change to `I2sRepack.kt`)
  • N/A — docs + KDoc only, no behavior change

…as closed, not built

#1205 (native decode kernel for BitNet.cpp's grouped I2_S layout) is closed
as not planned: AOT conversion (I2sAotConverter, #1207) already reaches the
same zero-copy mmap path (#1203) for a fraction of the ongoing cost of
maintaining a second decode variant per SIMD backend, and the layout is a
property of which pipeline quantized the file, not which device runs it --
so a native kernel would still need runtime dispatch across every layout
regardless of build target. Full reasoning and what a kernel would need if
a real use case someday can't tolerate an AOT step: #1205's closing
comment.

- I2sGgufLayout.GROUP_128/GROUP_64 doc comments point here instead of
  leaving the decision as tribal knowledge in a closed GitHub issue only.
- ternary-getting-started.adoc gains a "Loading a BitNet.cpp-quantized
  GGUF" section covering the same ground for someone hitting the grouped
  layout for the first time, plus a Status and roadmap update.
@github-actions

Copy link
Copy Markdown

📖 Documentation Preview

The documentation has been built successfully for this PR.

Generated Files:

  • Operator documentation: docs/modules/operators/_generated_/
  • JSON schema output: operators.json

Artifacts:

  • Download the documentation-preview-1214 artifact to view the complete documentation locally.

This comment will be updated automatically when the PR is updated.

@michalharakal
michalharakal merged commit f0e6115 into develop Aug 29, 2026
18 checks passed
@michalharakal
michalharakal deleted the docs/1205-close-note branch August 29, 2026 22:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant