Skip to content

[feat]: FastH3 on RTX GPUs: NVFP4/FP8 loading, sparse FP4 and INT8 attention, single-GPU memory path - #1919

Open
aryan5v wants to merge 61 commits into
hao-ai-lab:mainfrom
aryan5v:fasth3-rtx
Open

aryan5v wants to merge 61 commits into
hao-ai-lab:mainfrom
aryan5v:fasth3-rtx

Conversation

@aryan5v

@aryan5v aryan5v commented Oct 5, 2026

Copy link
Copy Markdown
Collaborator

Summary

FastH3 on single RTX and Blackwell GPUs: the code behind the FastH3 consumer-GPU release. It runs pruned FastH3 (42 blocks, 8 steps) and the full 8-step V2 on one RTX 5090, RTX 4090 or RTX PRO 6000, and on lower-memory cards through FP8 and offload. Every new path is opt-in, and default behavior is unchanged.

This combines three PRs from my fork. All of their commits are kept as they were:

Fork PR Scope
aryan5v#45 Release core: packed NVFP4/FP8 checkpoint loading, calibrated activation scales, single-GPU memory path, headline benchmark
aryan5v#44 RTX PRO 6000 / RTX 5090: block-sparse SageAttention3 FP4 attention for VSA, ModelOpt NVFP4 converter, docs
aryan5v#46 RTX 4090 and lower-memory GPUs: INT8-QK sparse attention on sm89, streamed NVFP4 encoder, INT8 VAE epilogues, memory caps

The #44 kernel and loader commits were already folded into #45's history. This branch adds the two #44 files #45 had dropped: the RTX PRO 6000 doc and the sparse-kernel tests. The DGX Spark and Apple Silicon work (aryan5v#47) will be a separate PR stacked on this one.

What's in it

Checkpoints

  • Packed NVFP4 DiT export loader (layer_profile h3_dit / h3_dit_ffn / h3_dit_vsa). Each linear can carry a calibrated static activation scale (_nvfp4_input_global_sf), from max calibration over 1,000 prompts and all steps. Without one, H3's ff.fc_out saturates in 40 of 42 blocks.
  • scripts/checkpoint_conversion/convert_minimax_h3_modelopt_nvfp4_dit.py: ModelOpt NVFP4 → FastVideo packed format. Also quantizes the attention, gate and FFN linears from BF16 (--quantize-*), and embeds calibrated scales (--act-amax).
  • Pre-quantized FP8 W8A8 checkpoints load directly as FP8 buffers. The NVFP4 text encoder dequantizes per layer on sm80–sm90.

Single-GPU memory

  • AdaLN modulation tables: the AdaLN projection weights are no longer needed (24 GiB on V2).
  • CPU-first DiT load, and pinned encoder/VAE swaps through exact-size cudaHostRegister arenas. Before, PyTorch rounded each block up to a power of two: 2.87 GiB of FP4 weights took 5.06 GiB of host RAM.
  • Layerwise offload streams packed buffers and keeps the first N blocks resident (FASTVIDEO_LAYERWISE_RESIDENT_BLOCKS).
  • FFN token chunking, FASTVIDEO_CUDA_MEMORY_CAP_GIB and per-stage memory logs.

Kernels

  • sm_120 FP4: block-sparse SageAttention3 forward (fwd_sparse) that runs VSA's 64-token tiles exactly. The VSA-H3 path uses it under FASTVIDEO_H3_VSA_FP4=1.
  • sm89: INT8-QK / BF16-PV sparse attention over the existing layout, per-tensor FP8 GEMM with a fused per-token × per-channel epilogue, a streamed NVFP4 encoder with fused dequant, and INT8 VAE epilogues.

Docs and benchmarks

  • Cookbook recipes, docs/inference/fasth3_rtx_pro_6000.md, the 4090 benchmark README.
  • bench_headline.py with a Modal entrypoint: 2 fixed prompts, 1 warmup + 2 timed runs, end-to-end wall time.

Results

End to end, from prompt to MP4 with audio, with a warm server. Each number is the median of the timed runs.

Hardware Model 832×480, 124 frames 1344×768, 243 frames
1× RTX 5090 V1 4-step, NVFP4 12.9 s 48.0 s
1× RTX 5090 Pruned 8-step, NVFP4 17.4 s 80.5 s
1× RTX 5090 V2 8-step, NVFP4 — 90.1 s
1× RTX PRO 6000 Pruned 8-step, NVFP4 MLP 15.1 s 83.9 s
1× RTX 4090 Pruned 8-step, FP8 41.8 s —

On the 4090, a 243-frame 832×480 clip takes 79.7 s. It also completes with the allocator capped at 16 GiB (104.0 s) and 12 GiB (107.3 s).

Review follow-ups in this branch

The Greptile findings on #44, #45 and #46:

  • Fixed before this branch: _build_block_mask arguments, shared Q/K/V quantization only at unit scale (calibrated layers quantize separately), reuse of the tile buffer.
  • Fixed here:
    • sageattn_blackwell_sparse{,_bshd}(validate=True) rejects a zero q2k_num or out-of-range q2k_idx before launch.
    • The converter keeps projections ModelOpt already quantized off the dense re-quantize path.
    • The dropped input_scale key is declared with its reason.
    • Headline runs record HEADLINE_* and the resolved config.
    • A / in a Modal tag no longer breaks the result writer.
    • Failed cgroup sampling reports null instead of 0 GiB.
  • Converter parity test: test_minimax_h3_modelopt_nvfp4_converter.py converts a synthetic ModelOpt checkpoint and checks two things. The ModelOpt bytes carry over bit for bit, and one forward per loaded linear matches the reference with its calibrated scale.
  • Left as is: ff.fc_in / ff.fc_out in fp8_config._FP8_SUFFIXES. The tuple already lists Kandinsky5's FFN names, and H3's MLP names need the same treatment. Happy to move both if there's a preferred place.

Testing

I ran all 16 test files this branch adds or changes on one GB200 (sm_100, FlashInfer 0.6.18). Result: 125 passed, 25 failed, 9 skipped. None of the failures come from this branch:

  • 24 test_minimax_h3_tile_first.py cases fail in setup: that environment has no fastvideo_kernel build for the VIDEO_SPARSE_ATTN_H3 backend they select.
  • test_nvfp4_purge.py::test_auto_purges_always_fp4_and_retains_refine_only fails on upstream main as well, in the same environment (the log receipt isn't captured).
  • Skipped:
    • test_attn_qat_infer_sparse.py needs an sm_120a kernel build.
    • 8 sm89 INT8 attention tests need an RTX 4090. They passed on the 4090 for lora related merge to fsdp  #46, but I couldn't rerun them because the pod is unreachable.

Fixed here after the GB200 run: the 2 streamed-NVFP4-encoder cases failed on GB200 because their FP4 scalars stayed on the host. They pass after the last commit.

Also run:

  • the new converter parity test on GB200: passes, along with the existing H3 export tests;
  • validate_cookbook after the merge with main: passes.

Pre-commit's excludes skip the files changed in the review-fix commits.

Not included

  • DGX Spark and Apple Silicon: coming in a separate PR.
  • Multi-GPU sm_120 Ulysses path: not executed yet.
  • 8 GiB tier on the 4090: a 7.25 GiB allocator cap still peaked at ~8.28 GiB on the device.

aryan5v and others added 30 commits October 3, 2026 10:32
Keep encoder/DiT/VAE off disk between clips, overlay the Comfy int8-convrot decoder for quality, and let the playground queue prompts while switching clips. TAEH3 stays an optional preview config only.

(cherry picked from commit 15e5517)
Drop playground queue experiments from this PR, reject ConvRot overlays that cannot rotate activations, and skip FSDP2 modules that mix dense params with DTensors.

(cherry picked from commit f16b87e)
Add H3 cookbook recipes for the 42-block checkpoint, reject ConvRot overlays whose group size cannot rotate activations, and skip FSDP2 modules that mix dense parameters with DTensors.

(cherry picked from commit 7852505)
Dense CompactH3 constructed VIDEO_SPARSE_ATTN_H3 tiles with all-zero gates. Fail at load and first denoise, and drop restating CompactH3 banners.

(cherry picked from commit cc59773)
Snapshot of every RTX PRO 6000 experiment, including the Ulysses FP8
q/k/v exchange path that has not executed yet (8-GPU capacity was
unavailable) and the Modal drivers used for every measurement. The
ship-ready subset is on h3-sm120-sparse-fp4.

(cherry picked from commit da835b4)
- Pre-quantized FP8 W8A8 checkpoint loader (float8 weight + per-channel weight_scale)
- NVFP4 export: optional calibrated activation scale (_nvfp4_input_global_sf);
  converter gains --quantize-ffn (from bf16) and --act-amax
- NVFP4 static/dynamic activation scales via env; FP8 attention projections next to NVFP4 FFN
- AdaLN modulation host cache and precomputed tables (skips 24 GiB of AdaLN weights)
- Layerwise offload streams large buffers (fix: onload by name, placeholders fail the size test)
- CPU-first DiT load for layerwise/AdaLN-cache paths; pinned encoder/VAE swaps; parked modules
- NVFP4 text encoder bf16 de-quant fallback for pre-Blackwell GPUs
- VSA guard ignores offloaded placeholders; per-stage memory logging; memory cap / report knobs
- SP stage profiling and FP8 all-to-all simulation; Modal PRO 6000 bench steps

(cherry picked from commit cf04434)
- FP8 on sm89: per-tensor GEMM + Triton per-token x per-channel scale epilogue
  (torch rowwise _scaled_mm runs ~70 TFLOPS there, below bf16) and a fused
  one-launch per-token quantize (5x faster than the torch chain)
- FASTVIDEO_H3_FFN_CHUNK_TOKENS: inference-only FFN token chunking
- FASTVIDEO_LAYERWISE_RESIDENT_BLOCKS: keep the first N blocks resident
- Serialized NVFP4 text encoder: allow sm80-sm90 through the bf16 de-quant path
- FASTVIDEO_H3_SPLICE_TRANSFORMER / _FROM_STEP: second checkpoint runs late DMD steps

(cherry picked from commit e2ed39c)
…and Modal

bench_headline.py times the release protocol (two fixed prompts, one warmup,
two timed runs each, generate_video wall time) from a checkpoint's
fastvideo_inference.json contract, with optional W&B logging.
headline_app.py runs it on 1/4/8 RTX PRO 6000 Blackwell GPUs on Modal from a
FastVideo HF repo.
…K/V quantization only at unit scale

Upstream's _build_block_mask now takes per-region video tile spans and
sparsities; the FP4 VSA paths still passed the old five arguments and raised
TypeError on the first block. The shared Q/K/V quantization assumed every
NVFP4 layer used the unit activation scale; with calibrated (export or env)
or dynamic scales each projection now quantizes its own input.

Found by Greptile review on #45.
…endent NVFP4/FP8 conversion, splice scope

- fp8_kernels: int64 row offsets (a 78k-token fc_in output has 2.2e9 elements)
- _maybe_quantize_model: handle NVFP4 (+ mixed FP8) before the per-module walk
- step splice: primary transformer component only; no env mutation
- layerwise offload: tolerate malformed / all-resident FASTVIDEO_LAYERWISE_RESIDENT_BLOCKS
- NVFP4 encoder fallback: same input dtype contract as the FP4 path
- benchmarks: new _build_block_mask contract in bench_code.py, extra_env overrides,
  --timed validation, generator shutdown on failure
- drop the stale playground 'Older clip' assertion (that UI change did not survive the rebase)
torch's pinned allocator rounds every block up to a power of two, so parking
the pruned NVFP4 DiT (~20 GB) next to the NVFP4 encoder (15 GB) overran a
60 GB container on the RTX 5090. One registered arena per module pins exactly
the bytes needed; buffers now keep a persistent host copy as well.
Usage (conversion, profiles, env switches), the end-to-end and per-block
measurements, the block densities each kernel granularity computes, what
was tried and not shipped (dense FP4 for VSA students, the 8-GPU FP8
Ulysses exchange kept on h3-sm120-experimental), and known limitations.
 sparse kernel tests

A zero q2k_num or an out-of-range q2k_idx entry made fwd_sparse read outside
its index row or the KV tensors. sageattn_blackwell_sparse{,_bshd} take
validate=True to run check_sparse_block_lists (host sync) first.
test_attn_qat_infer_sparse.py from #44 was dropped when the sm_120 work was
folded into the release core; it is restored with a block-list check test.
…ze path; converter parity test

With --quantize-attention, an attention projection ModelOpt already
quantized matched both lists, and its packed uint8 weight was quantized
again as if it were BF16. It now stays on the ModelOpt path. The dropped
input_scale key is declared in DROPPED_MODELOPT_SUFFIXES with its reason.
The new GB200 test converts a synthetic ModelOpt checkpoint, checks the
ModelOpt bytes carry over bit for bit, loads the export and compares one
forward per linear (calibrated static scales) against the reference.
…ver reports unmeasured memory as 0 GiB

bench_headline.py records HEADLINE_* switches and the resolved engine and
experimental config in results.json and W&B. A slash in a Modal --tag no
longer creates a nested run directory the local writer cannot write. When
cgroup sampling fails, bench_pod.py reports host peaks as null with the
error instead of 0.0 GiB.
Only the cookbook cache-bust conflicted: both sides edited
cookbook-recipes.json, so the version moves to 14 on the data and every
cookbook page. validate_cookbook passes.
… GEMM's device

A layerwise-streamed encoder layer is finalized on the host, so its alpha
and activation global scale stayed on the CPU while the packed weights
visited the GPU. FlashInfer 0.6.18 rejects a CPU globalScale; both
test_streamed_encoder_matches_resident_and_releases_layers NVFP4 cases
failed on GB200 and now pass.
@mergify mergify Bot added type: feat New feature or capability scope: inference Inference pipeline, serving, CLI scope: attention Attention backends (VSA, STA, Flash, etc.) scope: kernel CUDA kernels, fastvideo-kernel scope: infra CI, tests, Docker, build scope: docs Documentation scope: model Model architecture (DiTs, encoders, VAEs) labels Oct 5, 2026
@mergify

mergify Bot commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Merge Protections

🔴 1 of 1 protections blocking · waiting on 👀 reviews and 🤖 CI

Protection Waiting on
🔴 PR merge requirements 👀 reviews and 🤖 CI

🔴 PR merge requirements

Waiting for

  • #approved-reviews-by>=1
  • check-success=fastcheck-passed
  • check-success=full-suite-passed
This rule is failing.
  • #approved-reviews-by>=1
  • check-success=fastcheck-passed
  • check-success=full-suite-passed
  • check-success~=pre-commit
  • title~=(?i)^\[(feat|feature|bugfix|fix|refactor|perf|ci|doc|docs|misc|chore|kernel|new.?model|skill|skills|infra)\]

@mergify

mergify Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

Pre-commit checks failed

Hi @aryan5v, the pre-commit checks have failed. To fix them locally:

# Install pre-commit if you haven't already
uv pip install pre-commit
pre-commit install

# Run all checks and auto-fix what's possible
pre-commit run --all-files

Common fixes:

  • yapf: yapf -i <file> (formatting)
  • ruff: ruff check --fix <file> (linting)
  • codespell: codespell --write-changes <file> (spelling)

After fixing, commit and push the changes. The checks will re-run automatically.

For future commits, pre-commit will run automatically on changed files before each commit.

@greptile-apps

greptile-apps Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

[High risk] Adds quantization kernels and model loading paths for GPU inference.

The PR appears safe to merge based on the reviewed changes and resolved prior findings.

Summary

The PR adds opt-in FastH3 paths for RTX GPUs, including packed NVFP4/FP8 loading, sparse attention, lower-memory component placement, benchmark tooling, and usage examples. Since the previous review, it also preserves ModelOpt activation scales, checks packed VSA gates, clones retained VAE tile batches, and validates sparse lists by default. Greptile automatically discovered a related ticket that helped explain the purpose of this PR: reducing MiniMax H3 text-encoder loading memory for single-device inference.

  • The four previous Greptile threads are resolved.
  • No new actionable finding was established from the changes since that review.

Reviews (2) · Last reviewed commit: "[bugfix]: sparse FP4 entry points valida..."

Comment thread fastvideo/models/vaes/minimax_h3_video.py Outdated
Comment thread fastvideo/pipelines/basic/minimax_h3/vsa_guard.py
Comment thread fastvideo-kernel/attn_qat_infer/blackwell/api.cu
Aryan Kumar and others added 5 commits October 4, 2026 22:19
…play

Under the opt-in reduce-overhead decoder compile, each batch's output lives
in the graph pool; retaining split views let the next batch overwrite
earlier tiles before stitching.
With layer_profile h3_dit_vsa the gate weight is purged and lives in
_nvfp4_weight, so an all-zero packed gate skipped the guard. Codes that
are +-0 in every nibble now count as a zero gate.
sageattn_blackwell_sparse{,_bshd} now check q2k_num / q2k_idx unless the
caller passes validate=False. The H3 VSA path builds its lists with
vsa_tile_mask_to_fp4_blocks and opts out to avoid the host sync. The
converter comment now says input_scale is carried, not dropped.
Upstream hao-ai-lab#1898 requires every FASTVIDEO_* variable to be declared in
fastvideo/envs.py and read with envs.NAME.get(). The 29 H3, NVFP4, offload
and memory-debug switches this branch added read os.environ directly and
tests set them with monkeypatch.setenv, which fails
test_env_access_follows_policy in the unit lane. Declare them with types
and defaults matching the old parsing, read them through the registry,
switch tests to envs.NAME.override(), and regenerate the env-var table.
…fallback

validate_runtime now accepts sm80-sm99 with a per-call bf16 de-quantization
and checks the tensor-parallel size first, and _apply_finalized takes the
fallback on any device without FP4 GEMM. Stub the TP size and assert the
sm75 refusal / sm89 acceptance, pin the FP4 path in the apply test, and
cover the fallback against the de-quantized reference linear. These two
tests failed in the encoder lane on non-Blackwell CI GPUs.
@mergify

mergify Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

Pre-commit checks failed

Hi @aryan5v, the pre-commit checks have failed. To fix them locally:

# Install pre-commit if you haven't already
uv pip install pre-commit
pre-commit install

# Run all checks and auto-fix what's possible
pre-commit run --all-files

Common fixes:

  • yapf: yapf -i <file> (formatting)
  • ruff: ruff check --fix <file> (linting)
  • codespell: codespell --write-changes <file> (spelling)

After fixing, commit and push the changes. The checks will re-run automatically.

For future commits, pre-commit will run automatically on changed files before each commit.

…s that predate validate=

Passing validate=False to an older sageattn_blackwell_sparse_bshd raised TypeError on the first forward.
Fall back to the positional call there. Also: headline benchmark adds a 768p 5 s setting and a showcase
mode that renders a prompt set per seed at 480p 5 s.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

scope: attention Attention backends (VSA, STA, Flash, etc.) scope: docs Documentation scope: inference Inference pipeline, serving, CLI scope: infra CI, tests, Docker, build scope: kernel CUDA kernels, fastvideo-kernel scope: model Model architecture (DiTs, encoders, VAEs) type: feat New feature or capability

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant