[feat]: FastH3 on RTX GPUs: NVFP4/FP8 loading, sparse FP4 and INT8 attention, single-GPU memory path - #1919
[feat]: FastH3 on RTX GPUs: NVFP4/FP8 loading, sparse FP4 and INT8 attention, single-GPU memory path#1919aryan5v wants to merge 61 commits into
Conversation
Keep encoder/DiT/VAE off disk between clips, overlay the Comfy int8-convrot decoder for quality, and let the playground queue prompts while switching clips. TAEH3 stays an optional preview config only. (cherry picked from commit 15e5517)
Drop playground queue experiments from this PR, reject ConvRot overlays that cannot rotate activations, and skip FSDP2 modules that mix dense params with DTensors. (cherry picked from commit f16b87e)
Add H3 cookbook recipes for the 42-block checkpoint, reject ConvRot overlays whose group size cannot rotate activations, and skip FSDP2 modules that mix dense parameters with DTensors. (cherry picked from commit 7852505)
(cherry picked from commit 19a408f)
(cherry picked from commit 7d2a059)
Dense CompactH3 constructed VIDEO_SPARSE_ATTN_H3 tiles with all-zero gates. Fail at load and first denoise, and drop restating CompactH3 banners. (cherry picked from commit cc59773)
Snapshot of every RTX PRO 6000 experiment, including the Ulysses FP8 q/k/v exchange path that has not executed yet (8-GPU capacity was unavailable) and the Modal drivers used for every measurement. The ship-ready subset is on h3-sm120-sparse-fp4. (cherry picked from commit da835b4)
- Pre-quantized FP8 W8A8 checkpoint loader (float8 weight + per-channel weight_scale) - NVFP4 export: optional calibrated activation scale (_nvfp4_input_global_sf); converter gains --quantize-ffn (from bf16) and --act-amax - NVFP4 static/dynamic activation scales via env; FP8 attention projections next to NVFP4 FFN - AdaLN modulation host cache and precomputed tables (skips 24 GiB of AdaLN weights) - Layerwise offload streams large buffers (fix: onload by name, placeholders fail the size test) - CPU-first DiT load for layerwise/AdaLN-cache paths; pinned encoder/VAE swaps; parked modules - NVFP4 text encoder bf16 de-quant fallback for pre-Blackwell GPUs - VSA guard ignores offloaded placeholders; per-stage memory logging; memory cap / report knobs - SP stage profiling and FP8 all-to-all simulation; Modal PRO 6000 bench steps (cherry picked from commit cf04434)
- FP8 on sm89: per-tensor GEMM + Triton per-token x per-channel scale epilogue (torch rowwise _scaled_mm runs ~70 TFLOPS there, below bf16) and a fused one-launch per-token quantize (5x faster than the torch chain) - FASTVIDEO_H3_FFN_CHUNK_TOKENS: inference-only FFN token chunking - FASTVIDEO_LAYERWISE_RESIDENT_BLOCKS: keep the first N blocks resident - Serialized NVFP4 text encoder: allow sm80-sm90 through the bf16 de-quant path - FASTVIDEO_H3_SPLICE_TRANSFORMER / _FROM_STEP: second checkpoint runs late DMD steps (cherry picked from commit e2ed39c)
…and Modal bench_headline.py times the release protocol (two fixed prompts, one warmup, two timed runs each, generate_video wall time) from a checkpoint's fastvideo_inference.json contract, with optional W&B logging. headline_app.py runs it on 1/4/8 RTX PRO 6000 Blackwell GPUs on Modal from a FastVideo HF repo.
…ross-module import)
…e local headline results
…K/V quantization only at unit scale Upstream's _build_block_mask now takes per-region video tile spans and sparsities; the FP4 VSA paths still passed the old five arguments and raised TypeError on the first block. The shared Q/K/V quantization assumed every NVFP4 layer used the unit activation scale; with calibrated (export or env) or dynamic scales each projection now quantizes its own input. Found by Greptile review on #45.
…endent NVFP4/FP8 conversion, splice scope - fp8_kernels: int64 row offsets (a 78k-token fc_in output has 2.2e9 elements) - _maybe_quantize_model: handle NVFP4 (+ mixed FP8) before the per-module walk - step splice: primary transformer component only; no env mutation - layerwise offload: tolerate malformed / all-resident FASTVIDEO_LAYERWISE_RESIDENT_BLOCKS - NVFP4 encoder fallback: same input dtype contract as the FP4 path - benchmarks: new _build_block_mask contract in bench_code.py, extra_env overrides, --timed validation, generator shutdown on failure - drop the stale playground 'Older clip' assertion (that UI change did not survive the rebase)
…emory-limited GPUs
torch's pinned allocator rounds every block up to a power of two, so parking the pruned NVFP4 DiT (~20 GB) next to the NVFP4 encoder (15 GB) overran a 60 GB container on the RTX 5090. One registered arena per module pins exactly the bytes needed; buffers now keep a persistent host copy as well.
Usage (conversion, profiles, env switches), the end-to-end and per-block measurements, the block densities each kernel granularity computes, what was tried and not shipped (dense FP4 for VSA students, the 8-GPU FP8 Ulysses exchange kept on h3-sm120-experimental), and known limitations.
sparse kernel tests A zero q2k_num or an out-of-range q2k_idx entry made fwd_sparse read outside its index row or the KV tensors. sageattn_blackwell_sparse{,_bshd} take validate=True to run check_sparse_block_lists (host sync) first. test_attn_qat_infer_sparse.py from #44 was dropped when the sm_120 work was folded into the release core; it is restored with a block-list check test.
…ze path; converter parity test With --quantize-attention, an attention projection ModelOpt already quantized matched both lists, and its packed uint8 weight was quantized again as if it were BF16. It now stays on the ModelOpt path. The dropped input_scale key is declared in DROPPED_MODELOPT_SUFFIXES with its reason. The new GB200 test converts a synthetic ModelOpt checkpoint, checks the ModelOpt bytes carry over bit for bit, loads the export and compares one forward per linear (calibrated static scales) against the reference.
…ver reports unmeasured memory as 0 GiB bench_headline.py records HEADLINE_* switches and the resolved engine and experimental config in results.json and W&B. A slash in a Modal --tag no longer creates a nested run directory the local writer cannot write. When cgroup sampling fails, bench_pod.py reports host peaks as null with the error instead of 0.0 GiB.
Only the cookbook cache-bust conflicted: both sides edited cookbook-recipes.json, so the version moves to 14 on the data and every cookbook page. validate_cookbook passes.
… GEMM's device A layerwise-streamed encoder layer is finalized on the host, so its alpha and activation global scale stayed on the CPU while the packed weights visited the GPU. FlashInfer 0.6.18 rejects a CPU globalScale; both test_streamed_encoder_matches_resident_and_releases_layers NVFP4 cases failed on GB200 and now pass.
Merge Protections🔴 1 of 1 protections blocking · waiting on 👀 reviews and 🤖 CI
🔴 PR merge requirementsWaiting for
This rule is failing.
|
Pre-commit checks failedHi @aryan5v, the pre-commit checks have failed. To fix them locally: # Install pre-commit if you haven't already
uv pip install pre-commit
pre-commit install
# Run all checks and auto-fix what's possible
pre-commit run --all-filesCommon fixes:
After fixing, commit and push the changes. The checks will re-run automatically. For future commits, |
|
…play Under the opt-in reduce-overhead decoder compile, each batch's output lives in the graph pool; retaining split views let the next batch overwrite earlier tiles before stitching.
With layer_profile h3_dit_vsa the gate weight is purged and lives in _nvfp4_weight, so an all-zero packed gate skipped the guard. Codes that are +-0 in every nibble now count as a zero gate.
sageattn_blackwell_sparse{,_bshd} now check q2k_num / q2k_idx unless the
caller passes validate=False. The H3 VSA path builds its lists with
vsa_tile_mask_to_fp4_blocks and opts out to avoid the host sync. The
converter comment now says input_scale is carried, not dropped.
Upstream hao-ai-lab#1898 requires every FASTVIDEO_* variable to be declared in fastvideo/envs.py and read with envs.NAME.get(). The 29 H3, NVFP4, offload and memory-debug switches this branch added read os.environ directly and tests set them with monkeypatch.setenv, which fails test_env_access_follows_policy in the unit lane. Declare them with types and defaults matching the old parsing, read them through the registry, switch tests to envs.NAME.override(), and regenerate the env-var table.
…fallback validate_runtime now accepts sm80-sm99 with a per-call bf16 de-quantization and checks the tensor-parallel size first, and _apply_finalized takes the fallback on any device without FP4 GEMM. Stub the TP size and assert the sm75 refusal / sm89 acceptance, pin the FP4 path in the apply test, and cover the fallback against the de-quantized reference linear. These two tests failed in the encoder lane on non-Blackwell CI GPUs.
Pre-commit checks failedHi @aryan5v, the pre-commit checks have failed. To fix them locally: # Install pre-commit if you haven't already
uv pip install pre-commit
pre-commit install
# Run all checks and auto-fix what's possible
pre-commit run --all-filesCommon fixes:
After fixing, commit and push the changes. The checks will re-run automatically. For future commits, |
…s that predate validate= Passing validate=False to an older sageattn_blackwell_sparse_bshd raised TypeError on the first forward. Fall back to the positional call there. Also: headline benchmark adds a 768p 5 s setting and a showcase mode that renders a prompt set per seed at 480p 5 s.
Summary
FastH3 on single RTX and Blackwell GPUs: the code behind the FastH3 consumer-GPU release. It runs pruned FastH3 (42 blocks, 8 steps) and the full 8-step V2 on one RTX 5090, RTX 4090 or RTX PRO 6000, and on lower-memory cards through FP8 and offload. Every new path is opt-in, and default behavior is unchanged.
This combines three PRs from my fork. All of their commits are kept as they were:
The #44 kernel and loader commits were already folded into #45's history. This branch adds the two #44 files #45 had dropped: the RTX PRO 6000 doc and the sparse-kernel tests. The DGX Spark and Apple Silicon work (aryan5v#47) will be a separate PR stacked on this one.
What's in it
Checkpoints
layer_profileh3_dit/h3_dit_ffn/h3_dit_vsa). Each linear can carry a calibrated static activation scale (_nvfp4_input_global_sf), from max calibration over 1,000 prompts and all steps. Without one, H3'sff.fc_outsaturates in 40 of 42 blocks.scripts/checkpoint_conversion/convert_minimax_h3_modelopt_nvfp4_dit.py: ModelOpt NVFP4 → FastVideo packed format. Also quantizes the attention, gate and FFN linears from BF16 (--quantize-*), and embeds calibrated scales (--act-amax).Single-GPU memory
cudaHostRegisterarenas. Before, PyTorch rounded each block up to a power of two: 2.87 GiB of FP4 weights took 5.06 GiB of host RAM.FASTVIDEO_LAYERWISE_RESIDENT_BLOCKS).FASTVIDEO_CUDA_MEMORY_CAP_GIBand per-stage memory logs.Kernels
fwd_sparse) that runs VSA's 64-token tiles exactly. The VSA-H3 path uses it underFASTVIDEO_H3_VSA_FP4=1.Docs and benchmarks
docs/inference/fasth3_rtx_pro_6000.md, the 4090 benchmark README.bench_headline.pywith a Modal entrypoint: 2 fixed prompts, 1 warmup + 2 timed runs, end-to-end wall time.Results
End to end, from prompt to MP4 with audio, with a warm server. Each number is the median of the timed runs.
On the 4090, a 243-frame 832×480 clip takes 79.7 s. It also completes with the allocator capped at 16 GiB (104.0 s) and 12 GiB (107.3 s).
Review follow-ups in this branch
The Greptile findings on #44, #45 and #46:
_build_block_maskarguments, shared Q/K/V quantization only at unit scale (calibrated layers quantize separately), reuse of the tile buffer.sageattn_blackwell_sparse{,_bshd}(validate=True)rejects a zeroq2k_numor out-of-rangeq2k_idxbefore launch.input_scalekey is declared with its reason.HEADLINE_*and the resolved config./in a Modal tag no longer breaks the result writer.nullinstead of 0 GiB.test_minimax_h3_modelopt_nvfp4_converter.pyconverts a synthetic ModelOpt checkpoint and checks two things. The ModelOpt bytes carry over bit for bit, and one forward per loaded linear matches the reference with its calibrated scale.ff.fc_in/ff.fc_outinfp8_config._FP8_SUFFIXES. The tuple already lists Kandinsky5's FFN names, and H3's MLP names need the same treatment. Happy to move both if there's a preferred place.Testing
I ran all 16 test files this branch adds or changes on one GB200 (sm_100, FlashInfer 0.6.18). Result: 125 passed, 25 failed, 9 skipped. None of the failures come from this branch:
test_minimax_h3_tile_first.pycases fail in setup: that environment has nofastvideo_kernelbuild for the VIDEO_SPARSE_ATTN_H3 backend they select.test_nvfp4_purge.py::test_auto_purges_always_fp4_and_retains_refine_onlyfails on upstreammainas well, in the same environment (the log receipt isn't captured).test_attn_qat_infer_sparse.pyneeds an sm_120a kernel build.Fixed here after the GB200 run: the 2 streamed-NVFP4-encoder cases failed on GB200 because their FP4 scalars stayed on the host. They pass after the last commit.
Also run:
validate_cookbookafter the merge withmain: passes.Pre-commit's excludes skip the files changed in the review-fix commits.
Not included