Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
100 commits
Select commit Hold shift + click to select a range
2d0ebf1
[feat]: load packed FastH3 NVFP4 and Comfy int8 VAE on Blackwell
aryan5v Sep 19, 2026
c38b4e2
[feat]: add CompactH3 NVFP4 cookbook paths for Blackwell GPUs
aryan5v Sep 19, 2026
a068234
[feat]: document CompactH3 NVFP4 on RTX 5090 and RTX PRO 6000
aryan5v Sep 19, 2026
d224e52
[bugfix]: refuse FSDP NVFP4 retain and drop restating comments
aryan5v Sep 19, 2026
d078283
[feat]: enable CompactH3 VAE compile on RTX PRO 6000
aryan5v Sep 19, 2026
a766348
[bugfix]: refuse zero-gate CompactH3 VSA
aryan5v Sep 19, 2026
1fb455d
[wip]: sm_120 H3 FP4 experiments: SP FP8 exchange path and Modal drivers
aryan5v Oct 1, 2026
0a4737c
[wip]: H3 single-GPU 5090/4090 shipping path
aryan5v Oct 3, 2026
4012c37
[wip]: H3 4090 groundwork, step splice, encoder fallback on sm89
aryan5v Oct 3, 2026
4a72920
[feat]: H3 headline benchmark (480p 5 s, 768p 10 s) for GPU clusters …
aryan5v Oct 3, 2026
f0688b0
[bugfix]: ship app.py into the headline Modal image
aryan5v Oct 3, 2026
0a23563
[bugfix]: ship app.py into the headline Modal image as local python s…
aryan5v Oct 3, 2026
a87344c
[bugfix]: run the headline benchmark from app.py (one Modal app, no c…
aryan5v Oct 3, 2026
edbe18b
[bugfix]: import os in gpu_worker; apply pre-commit formatting; ignor…
aryan5v Oct 3, 2026
dd10fc7
[bugfix]: H3 FP4 attention: new _build_block_mask signature; share Q/…
aryan5v Oct 3, 2026
a6faddd
[bugfix]: address #45 review: int64 FP8 epilogue offsets, order-indep…
aryan5v Oct 3, 2026
ccebd84
[misc]: headline Modal runs keep the full log on the volume and repor…
aryan5v Oct 3, 2026
54f335c
[bugfix]: restore import os in fsdp_load (dropped in the rebase onto …
aryan5v Oct 3, 2026
4d33717
[feat]: headline benchmark: HEADLINE_VAE_PARALLEL=1 decodes VAE tiles…
aryan5v Oct 3, 2026
5386585
[feat]: headline benchmark: engine/pipeline placement overrides for m…
aryan5v Oct 3, 2026
cb79431
[feat]: headline Modal entrypoint takes extra env and a run tag
aryan5v Oct 3, 2026
a97d23f
[bugfix]: H3 pinned swaps use exact-size cudaHostRegister arenas
aryan5v Oct 3, 2026
a216768
[wip]: exact-size pinned arenas for H3 offload and 4090 benchmarks
aryan5v Oct 3, 2026
a491f15
[wip]: document 4090 setup validation and reproducible baselines
aryan5v Oct 3, 2026
4ae791c
[wip]: record checkpoint revision and GPU driver in 4090 benchmark re…
aryan5v Oct 3, 2026
21f9899
[wip]: record completed 480p baseline and stage summary tooling
aryan5v Oct 3, 2026
c716408
[perf]: reduce H3 attention activation copies and share FP8 input qua…
aryan5v Oct 3, 2026
bd713dd
[wip]: prototype tile-64 INT8 QK and FP8 PV attention on sm89
aryan5v Oct 3, 2026
5bf9804
[perf]: retain offloaded H3 VAEs on host until their consuming stages
aryan5v Oct 3, 2026
59e6946
[perf]: add opt-in sm89 tile-64 BF16 and INT8 QK attention
aryan5v Oct 3, 2026
b3ab6e1
[docs]: record cached 4090 timings and sm89 precision validation
aryan5v Oct 3, 2026
994c220
[wip]: validate tilewise FP8 values and dynamic probability scales
aryan5v Oct 3, 2026
6b18535
[docs]: record resident-block timings and encoder constraints
aryan5v Oct 3, 2026
92fc26c
[docs]: record five-second 4090 timing and FP8 encoder footprint
aryan5v Oct 3, 2026
d986589
[feat]: stream the H3 encoder for consumer VRAM limits
aryan5v Oct 3, 2026
887deaa
[perf]: fuse serialized NVFP4 encoder weight dequantization
aryan5v Oct 3, 2026
2aa19c4
[perf]: release H3 fine-attention copies before the gated merge
aryan5v Oct 3, 2026
9c9f1ed
[perf]: share VAE INT8 input preparation and avoid weight copies
aryan5v Oct 3, 2026
e457b68
[docs]: record 12 GiB and streamed 4090 results
aryan5v Oct 3, 2026
7a0d7d3
[perf]: read H3 INT8 sparse attention from existing layouts
aryan5v Oct 3, 2026
3c0668f
[perf]: fuse H3 INT8 VAE scaling and bias without intermediates
aryan5v Oct 3, 2026
a7f7ede
[bugfix]: retain consumer attention environment parsing after core re…
aryan5v Oct 3, 2026
87b22a5
[test]: adapt consumer VSA validation to packed segment metadata
aryan5v Oct 3, 2026
fb92af1
[bench]: sample total GPU memory for consumer budget checks
aryan5v Oct 3, 2026
5ac85ad
[docs]: record 42-second 4090 clips and consumer memory validation
aryan5v Oct 3, 2026
98cdc7d
[docs]: record Track B PR and strict 8 GiB budget limit
aryan5v Oct 3, 2026
5a377ca
[docs]: FastH3 NVFP4 on RTX PRO 6000
aryan5v Oct 1, 2026
3d727ff
[bugfix]: validate sparse FP4 block lists before launch; restore the …
aryan5v Oct 5, 2026
5efdd8b
[bugfix]: keep ModelOpt-quantized projections off the dense re-quanti…
aryan5v Oct 5, 2026
74a190f
[bugfix]: headline benchmarks record their decode mode; 4090 bench ne…
aryan5v Oct 5, 2026
5db0cfa
[misc]: merge upstream main into the FastH3 RTX release branch
aryan5v Oct 5, 2026
4784019
[bugfix]: streamed NVFP4 encoder layers pass their FP4 scalars on the…
aryan5v Oct 5, 2026
7863e08
[bugfix]: preserve ModelOpt H3 activation calibration in NVFP4 export
Oct 4, 2026
9ca4eba
[bugfix]: declare NVFP4 static activation-scale cache attributes for …
aryan5v Oct 5, 2026
3c7e89c
[bugfix]: clone batched VAE tile output before the next CUDA-graph re…
aryan5v Oct 5, 2026
f4df26c
[bugfix]: VSA zero-gate guard checks packed NVFP4 gates
aryan5v Oct 5, 2026
2964f15
[bugfix]: sparse FP4 entry points validate block lists by default
aryan5v Oct 5, 2026
826b5d0
[misc]: register FastH3 single-GPU switches in fastvideo.envs
aryan5v Oct 5, 2026
a59e842
[test]: NVFP4 H3 encoder tests follow the pre-Blackwell de-quantized …
aryan5v Oct 5, 2026
e9a35b2
[misc]: drop the os imports layerwise offload no longer uses
aryan5v Oct 5, 2026
84eb65d
[bugfix]: FP4 VSA path stays compatible with fastvideo-kernel release…
aryan5v Oct 5, 2026
d60e209
[feat]: add resident FastH3 V2 Spark recipe
Oct 3, 2026
4874682
[feat]: benchmark resident FastH3 Spark recipe
Oct 3, 2026
8c01b5a
[docs]: pin Spark benchmark environment
Oct 3, 2026
6ab83c3
[bugfix]: retain per-stage CUDA peaks in Spark benchmarks
Oct 3, 2026
b1c3495
[feat]: add resident pruned and paired Spark recipes
Oct 3, 2026
4ce033a
[bugfix]: preserve H3 release helpers after main integration
Oct 3, 2026
40f6c3c
[feat]: support pruned FastH3 AdaLN in MLX
Oct 3, 2026
159168d
[feat]: trim MLX H3 conditioner to required layers
Oct 3, 2026
f4f903d
[feat]: allow FP8 source for MLX H3 conversion
Oct 3, 2026
c5d233e
[docs]: add pruned FastH3 MLX conversion recipe
Oct 3, 2026
89d3986
[feat]: add resident NVFP4 conditioning for MLX H3
Oct 3, 2026
cfe9d5f
[feat]: enable experimental MXFP8 H3 conversion
Oct 3, 2026
aef7cab
[test]: check MXFP8 H3 checkpoint round trip
Oct 3, 2026
d14f666
[bugfix]: omit unused layers from single-shard MLX encoder
Oct 3, 2026
2f400ed
[test]: verify rank-16 H3 modulation and cache reload
Oct 4, 2026
7735e12
[bugfix]: fence NVFP4 activation quantization on DGX Spark
Oct 4, 2026
0301008
[docs]: explain calibrated and reproducible Spark NVFP4 inference
Oct 4, 2026
5830879
[bugfix]: keep Spark NVFP4 fence opaque to Dynamo
Oct 4, 2026
ad06a38
[bugfix]: order fresh NVFP4 global scales on DGX Spark
Oct 4, 2026
c7969ea
[bugfix]: carry H3 performance switches to Ray workers
Oct 4, 2026
fd8b213
[docs]: record verified Track C Spark release timings
Oct 4, 2026
532fd6b
[perf]: bound MLX H3 sparse attention gather memory
Oct 4, 2026
e8bcd31
[perf]: batch 32 keys in opt-in MLX H3 sparse attention
Oct 4, 2026
ef1f277
[bugfix]: apply H3 wired residency through the MLX wired API
Oct 4, 2026
6457d71
[docs]: explain H3 Metal wired-memory limits
Oct 4, 2026
ecc91b2
[feat]: cache packed H3 encoder in MLX layout
Oct 4, 2026
10697e8
[feat]: support scaled native H3 floating quantization
Oct 5, 2026
32c8428
[perf]: parallelize H3 SIMD softmax across four lanes
Oct 5, 2026
ff8bbcc
[bugfix]: publish the packed MLX encoder cache with one rename
aryan5v Oct 5, 2026
e2f89da
[bugfix]: MLX H3 resident preload skips dropped AdaLN weights; restor…
aryan5v Oct 5, 2026
a1c024b
[test]: update MLX H3 regression stubs for model_root and NVFP4 condi…
aryan5v Oct 5, 2026
2539357
[test]: update MLX fast-spatial stubs for model_root and decode tile …
aryan5v Oct 5, 2026
6ccdbc7
[misc]: one registry entry per FastH3 switch after stacking on the RT…
aryan5v Oct 5, 2026
e331f01
[docs]: Align Spark recipes with released FastH3 stacks
aryan5v Oct 5, 2026
3168746
[docs]: Keep Spark launch table on released models
aryan5v Oct 5, 2026
ecbdbe5
[misc]: ship validated Track C MLX release paths
aryan5v Oct 5, 2026
8f33ee9
[bugfix]: reject packed H3 transformer sources
aryan5v Oct 5, 2026
ec7c971
[docs]: use released Track C INT6 Mac recipe
aryan5v Oct 5, 2026
17987f6
[test]: MLX dequant-GEMM test allows CPU BF16 accumulation error
aryan5v Oct 5, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -142,3 +142,4 @@ fastvideo/tests/ssim/reference_videos/**
*.nvimlog
.nvimlog
.python-version
scripts/benchmarks/minimax_h3_pro6000/headline_results/
73 changes: 72 additions & 1 deletion docs/assets/cookbook-recipes.json
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
{
"version": 12,
"version": 14,
"recipes": [
{
"id": "fastwan21-t2v",
Expand Down Expand Up @@ -601,6 +601,77 @@
"Height, width, frames, and steps in the YAML are examples. Edit them or pass CLI flags. See docs/getting_started/installation/spark_pair.md."
]
},
{
"id": "compacth3-rtx5090",
"group": "compacth3-rtx5090",
"group_label": "CompactH3 on RTX 5090",
"group_task": "4-step 42-block text to video + audio",
"family": "minimax_h3",
"stage": "inference",
"task": "Few-step text to video (with audio)",
"label": "CompactH3 NVFP4 on RTX 5090",
"summary": "Run the 42-block CompactH3 NVFP4 DiT with the NVFP4 Qwen3-VL encoder, Comfy int8-convrot VAE, and SageAttention3 FP4 on one 32 GB RTX 5090. Sequential load parks the encoder in pinned host RAM.",
"model": "./CompactH3",
"source": "examples/inference/basic/basic_compacth3_rtx5090.yaml",
"serving": {
"source": "examples/serving/openai_compacth3_rtx5090.yaml",
"install": "UV_TORCH_BACKEND=cu128 uv pip install -e \".[fasth3]\"",
"prepare": "hf download FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree --local-dir ./CompactH3 --include model_index.json --include \"tokenizer/**\" --include \"processor/**\" --include \"scheduler/**\" --include \"audio_scheduler/**\" --include \"audio_vae/**\" --include \"vae/**\"\nhf download aryan5v/FastH3-20B-42block-dmd2-ckpt1400-bf16 --local-dir ./CompactH3/transformer\nhf download aryan5v/FastH3-20B-42block-dmd2-ckpt1400-nvfp4 --local-dir ./CompactH3/transformer --include nvfp4_weights.safetensors\nhf download KyleNeverGivesUp/FastH3-text-encoder-nvfp4 --local-dir ./CompactH3/text_encoder\nhf download Comfy-Org/MiniMax-H3 --local-dir ./Comfy-MiniMax-H3 --include vae/minimax_h3_video_vae_int8_convrot.safetensors\ncp ./Comfy-MiniMax-H3/vae/minimax_h3_video_vae_int8_convrot.safetensors ./CompactH3/vae/",
"env": "FASTVIDEO_ATTENTION_BACKEND=ATTN_QAT_INFER FASTVIDEO_VSA_SM100A=0 FASTVIDEO_FA4=0 FLASHINFER_CUDA_ARCH_LIST=12.0a"
},
"command": "FASTVIDEO_ATTENTION_BACKEND=ATTN_QAT_INFER FASTVIDEO_VSA_SM100A=0 FASTVIDEO_FA4=0 FLASHINFER_CUDA_ARCH_LIST=12.0a FASTVIDEO_STAGE_LOGGING=1 fastvideo generate --config examples/inference/basic/basic_compacth3_rtx5090.yaml",
"gpu_types": ["NVIDIA"],
"hardware": {
"platform": "cuda",
"gpu_count": 1,
"evidence": "source-configured"
},
"evidence": "Source-backed",
"expected_artifact": "MP4 under outputs/compacth3_rtx5090/",
"modes": ["T2VA", "CompactH3 NVFP4", "RTX 5090"],
"limitations": [
"Assemble ./CompactH3 before running. The DiT NVFP4 export and encoder snapshot are gated; run huggingface-cli login and accept each repo license.",
"32 GB cannot keep the NVFP4 encoder and DiT on the GPU together. Keep h3_sequential_load on and lazy_module_load off.",
"Blackwell sm_120 needs ATTN_QAT_INFER, FLASHINFER_CUDA_ARCH_LIST=12.0a, and a CUDA 12.8 PyTorch wheel.",
"Legal num_frames values are 17n+5, capped at 362 (15.08 s). Native 16:9 sizes include 832x480 and 1344x768; 1344x768 on 32 GB is unmeasured."
]
},
{
"id": "compacth3-rtx-pro6000",
"group": "compacth3-rtx-pro6000",
"group_label": "CompactH3 on RTX PRO 6000",
"group_task": "4-step 42-block text to video + audio",
"family": "minimax_h3",
"stage": "inference",
"task": "Few-step text to video (with audio)",
"label": "CompactH3 NVFP4 on RTX PRO 6000 Blackwell",
"summary": "Run CompactH3 NVFP4 with the encoder, DiT, and int8-convrot VAE resident on one 96 GB RTX PRO 6000 Blackwell. The checked-in example is 1344x768 and 124 frames (5.17 s) with VAE torch.compile. Use 832x480 for clip-queue playground traffic.",
"model": "./CompactH3",
"source": "examples/inference/basic/basic_compacth3_rtx_pro6000.yaml",
"serving": {
"source": "examples/serving/openai_compacth3_rtx_pro6000.yaml",
"install": "UV_TORCH_BACKEND=cu128 uv pip install -e \".[fasth3]\"",
"prepare": "hf download FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree --local-dir ./CompactH3 --include model_index.json --include \"tokenizer/**\" --include \"processor/**\" --include \"scheduler/**\" --include \"audio_scheduler/**\" --include \"audio_vae/**\" --include \"vae/**\"\nhf download aryan5v/FastH3-20B-42block-dmd2-ckpt1400-bf16 --local-dir ./CompactH3/transformer\nhf download aryan5v/FastH3-20B-42block-dmd2-ckpt1400-nvfp4 --local-dir ./CompactH3/transformer --include nvfp4_weights.safetensors\nhf download KyleNeverGivesUp/FastH3-text-encoder-nvfp4 --local-dir ./CompactH3/text_encoder\nhf download Comfy-Org/MiniMax-H3 --local-dir ./Comfy-MiniMax-H3 --include vae/minimax_h3_video_vae_int8_convrot.safetensors\ncp ./Comfy-MiniMax-H3/vae/minimax_h3_video_vae_int8_convrot.safetensors ./CompactH3/vae/",
"env": "FASTVIDEO_ATTENTION_BACKEND=ATTN_QAT_INFER FASTVIDEO_VSA_SM100A=0 FASTVIDEO_FA4=0 FLASHINFER_CUDA_ARCH_LIST=12.0a"
},
"command": "FASTVIDEO_ATTENTION_BACKEND=ATTN_QAT_INFER FASTVIDEO_VSA_SM100A=0 FASTVIDEO_FA4=0 FLASHINFER_CUDA_ARCH_LIST=12.0a FASTVIDEO_STAGE_LOGGING=1 fastvideo generate --config examples/inference/basic/basic_compacth3_rtx_pro6000.yaml",
"gpu_types": ["NVIDIA"],
"hardware": {
"platform": "cuda",
"gpu_count": 1,
"evidence": "source-configured"
},
"evidence": "Source-backed",
"expected_artifact": "MP4 under outputs/compacth3_rtx_pro6000/",
"modes": ["T2VA", "CompactH3 NVFP4", "RTX PRO 6000"],
"limitations": [
"Assemble ./CompactH3 before running. The DiT NVFP4 export and encoder snapshot are gated; run huggingface-cli login and accept each repo license.",
"96 GB keeps the encoder, DiT, and VAE on GPU. Do not enable h3_sequential_load or lazy_module_load on this box.",
"Enable compile.vae_enabled. Leave inference_torch_compile off: FlashInfer and Sage3 custom ops cannot be compiled.",
"Blackwell sm_120 needs ATTN_QAT_INFER, FLASHINFER_CUDA_ARCH_LIST=12.0a, and a CUDA 12.8 PyTorch wheel.",
"Legal num_frames values are 17n+5, capped at 362 (15.08 s). Native 16:9 sizes include 832x480 and 1344x768. Dense CompactH3 has zero VSA gates; the pipeline raises if VIDEO_SPARSE_ATTN_H3 is loaded with all-zero to_gate_compress weights."
]
},
{
"id": "fasth3-8step-v2-cuda",
"group": "fasth3-8step-v2",
Expand Down
Loading
Loading