Add checkpoint-native Vulkan NVFP4 projections - #207
Merged
Merged
Conversation
SnowCheetos
force-pushed
the
ft/vulkan-nvfp4-tensor
branch
from
September 22, 2026 05:33
fecd8de to
f1a2926
Compare
SnowCheetos
force-pushed
the
ft/vulkan-nvfp4-tensor
branch
from
September 22, 2026 05:57
f1a2926 to
785f05c
Compare
| pub struct Nvfp4MoeArtifact { | ||
| bytes: [u8; MOE_HEADER_BYTES], | ||
| } | ||
|
|
Contributor
There was a problem hiding this comment.
This is fine for now. Prefer compute shader sources from here on, which is more readable.
micro-perceptron
approved these changes
Sep 23, 2026
micro-perceptron
left a comment
Contributor
There was a problem hiding this comment.
Go for v0.4.1; see comment for the future.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Nemotron NVFP4 checkpoints store weights as packed E2M1 nibbles with one E4M3 scale per 16 values. Expanding those matrices before Vulkan execution removes the storage and bandwidth advantage, and TOSA has no FP4 tensor type.
This adds provider-native Vulkan artifacts for direct F32 x NVFP4 projections and fused routed experts. The provider supports shared, contiguous-batch, and per-expert weight bindings; imports cache-slot mappings without copies; and keeps the up-projection squared-ReLU intermediate in a private device arena across the down projection. Fused artifacts carry each expert's own byte offsets, lower them to SPIR-V word addressing, and read back only final outputs.
The shader set includes a subgroup32 four-output-row decode kernel and Intel Xe2/Xe3 cooperative-matrix FP16 8x16x16 support. The fused routed path normalizes its intermediate algebraically by folding two powers of the up scale into the down scale, avoiding denorm flushing without a host reduction.
Measured on Intel Graphics (Lunar Lake):
1x2688x2688: about 0.33 ms after warmupThe fused path is numerically admitted but currently loses the measured placement gate on this device. The result is retained as reusable zero-copy infrastructure rather than enabled as an unconditional optimization.
Validation:
cargo test -p virtio-accel-vulkan --lib(36 passed)VIRTIO_ACCEL_VULKAN_REQUIRE_SPIRV_VAL=1 cargo test -p virtio-accel-vulkan --test targets every_kernel_variant_passes_spirv_val -- --nocapture(179 variants)Release preparation for v0.4.1:
virtio-accel-device, required by Kore's accelerator service. This preserves the existing single-submit API and limit, and changes no virtio wire bytes.axmodelwith Vulkan passed 48 unit tests; Kore's Xe driver and accelerator services compiled, andkore-svc-accelpassed 89 unit tests after its old dependency pins were replaced locally.8beedb0. Bare-metal Xe2 execution remains a separate Kore validation gate.