Skip to content

Add checkpoint-native Vulkan NVFP4 projections - #207

Merged
micro-perceptron merged 18 commits into
mainfrom
ft/vulkan-nvfp4-tensor
Sep 23, 2026
Merged

micro-perceptron merged 18 commits into
mainfrom
ft/vulkan-nvfp4-tensor

Conversation

@SnowCheetos

@SnowCheetos SnowCheetos commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Nemotron NVFP4 checkpoints store weights as packed E2M1 nibbles with one E4M3 scale per 16 values. Expanding those matrices before Vulkan execution removes the storage and bandwidth advantage, and TOSA has no FP4 tensor type.

This adds provider-native Vulkan artifacts for direct F32 x NVFP4 projections and fused routed experts. The provider supports shared, contiguous-batch, and per-expert weight bindings; imports cache-slot mappings without copies; and keeps the up-projection squared-ReLU intermediate in a private device arena across the down projection. Fused artifacts carry each expert's own byte offsets, lower them to SPIR-V word addressing, and read back only final outputs.

The shader set includes a subgroup32 four-output-row decode kernel and Intel Xe2/Xe3 cooperative-matrix FP16 8x16x16 support. The fused routed path normalizes its intermediate algebraically by folding two powers of the up scale into the down scale, avoiding denorm flushing without a host reduction.

Measured on Intel Graphics (Lunar Lake):

  • standalone decode 1x2688x2688: about 0.33 ms after warmup
  • fused six-expert scalar Vulkan unit: 13.8 ms, relative L2 4.11e-3
  • matching CPU routed unit: 2.0 ms
  • per-expert Xe2 XMX experiment: 19.9 ms; reverted because one useful row wastes seven rows of each native tile

The fused path is numerically admitted but currently loses the measured placement gate on this device. The result is retained as reusable zero-copy infrastructure rather than enabled as an unconditional optimization.

Validation:

  • cargo test -p virtio-accel-vulkan --lib (36 passed)
  • VIRTIO_ACCEL_VULKAN_REQUIRE_SPIRV_VAL=1 cargo test -p virtio-accel-vulkan --test targets every_kernel_variant_passes_spirv_val -- --nocapture (179 variants)
  • hardware Axiom full-model calibration at the real Nemotron routed geometry

Release preparation for v0.4.1:

  • Bumps all 18 workspace crates to 0.4.1. The Rust 1.85, Clippy, rustdoc, and Lavapipe CI gates pass, as does the ordered publication dry run.
  • Adds provider-side ordered event retention to virtio-accel-device, required by Kore's accelerator service. This preserves the existing single-submit API and limit, and changes no virtio wire bytes.
  • In disposable consumer checkouts using this PR's 0.4.1 crates, Axiom axmodel with Vulkan passed 48 unit tests; Kore's Xe driver and accelerator services compiled, and kore-svc-accel passed 89 unit tests after its old dependency pins were replaced locally.
  • Full CI run passed on commit 8beedb0. Bare-metal Xe2 execution remains a separate Kore validation gate.

@SnowCheetos
SnowCheetos force-pushed the ft/vulkan-nvfp4-tensor branch from fecd8de to f1a2926 Compare September 22, 2026 05:33
@SnowCheetos
SnowCheetos force-pushed the ft/vulkan-nvfp4-tensor branch from f1a2926 to 785f05c Compare September 22, 2026 05:57
@SnowCheetos SnowCheetos self-assigned this Sep 22, 2026
@SnowCheetos
SnowCheetos marked this pull request as ready for review September 22, 2026 19:27
Copilot AI lite review requested due to automatic review settings September 22, 2026 19:27
@SnowCheetos SnowCheetos added the area: engine Command dispatch, state, ownership, and reset label Sep 22, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

pub struct Nvfp4MoeArtifact {
bytes: [u8; MOE_HEADER_BYTES],
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is fine for now. Prefer compute shader sources from here on, which is more readable.

@micro-perceptron micro-perceptron left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Go for v0.4.1; see comment for the future.

@micro-perceptron
micro-perceptron merged commit 8b66e79 into main Sep 23, 2026
18 checks passed
@micro-perceptron
micro-perceptron deleted the ft/vulkan-nvfp4-tensor branch September 23, 2026 22:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area: engine Command dispatch, state, ownership, and reset

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants