Conversation
Keep encoder/DiT/VAE off disk between clips, overlay the Comfy int8-convrot decoder for quality, and let the playground queue prompts while switching clips. TAEH3 stays an optional preview config only. (cherry picked from commit 15e5517)
Drop playground queue experiments from this PR, reject ConvRot overlays that cannot rotate activations, and skip FSDP2 modules that mix dense params with DTensors. (cherry picked from commit f16b87e)
Add H3 cookbook recipes for the 42-block checkpoint, reject ConvRot overlays whose group size cannot rotate activations, and skip FSDP2 modules that mix dense parameters with DTensors. (cherry picked from commit 7852505)
(cherry picked from commit 19a408f)
(cherry picked from commit 7d2a059)
Dense CompactH3 constructed VIDEO_SPARSE_ATTN_H3 tiles with all-zero gates. Fail at load and first denoise, and drop restating CompactH3 banners. (cherry picked from commit cc59773)
Snapshot of every RTX PRO 6000 experiment, including the Ulysses FP8 q/k/v exchange path that has not executed yet (8-GPU capacity was unavailable) and the Modal drivers used for every measurement. The ship-ready subset is on h3-sm120-sparse-fp4. (cherry picked from commit da835b4)
- Pre-quantized FP8 W8A8 checkpoint loader (float8 weight + per-channel weight_scale) - NVFP4 export: optional calibrated activation scale (_nvfp4_input_global_sf); converter gains --quantize-ffn (from bf16) and --act-amax - NVFP4 static/dynamic activation scales via env; FP8 attention projections next to NVFP4 FFN - AdaLN modulation host cache and precomputed tables (skips 24 GiB of AdaLN weights) - Layerwise offload streams large buffers (fix: onload by name, placeholders fail the size test) - CPU-first DiT load for layerwise/AdaLN-cache paths; pinned encoder/VAE swaps; parked modules - NVFP4 text encoder bf16 de-quant fallback for pre-Blackwell GPUs - VSA guard ignores offloaded placeholders; per-stage memory logging; memory cap / report knobs - SP stage profiling and FP8 all-to-all simulation; Modal PRO 6000 bench steps (cherry picked from commit cf04434)
- FP8 on sm89: per-tensor GEMM + Triton per-token x per-channel scale epilogue (torch rowwise _scaled_mm runs ~70 TFLOPS there, below bf16) and a fused one-launch per-token quantize (5x faster than the torch chain) - FASTVIDEO_H3_FFN_CHUNK_TOKENS: inference-only FFN token chunking - FASTVIDEO_LAYERWISE_RESIDENT_BLOCKS: keep the first N blocks resident - Serialized NVFP4 text encoder: allow sm80-sm90 through the bf16 de-quant path - FASTVIDEO_H3_SPLICE_TRANSFORMER / _FROM_STEP: second checkpoint runs late DMD steps (cherry picked from commit e2ed39c)
…and Modal bench_headline.py times the release protocol (two fixed prompts, one warmup, two timed runs each, generate_video wall time) from a checkpoint's fastvideo_inference.json contract, with optional W&B logging. headline_app.py runs it on 1/4/8 RTX PRO 6000 Blackwell GPUs on Modal from a FastVideo HF repo.
…ross-module import)
…e local headline results
…K/V quantization only at unit scale Upstream's _build_block_mask now takes per-region video tile spans and sparsities; the FP4 VSA paths still passed the old five arguments and raised TypeError on the first block. The shared Q/K/V quantization assumed every NVFP4 layer used the unit activation scale; with calibrated (export or env) or dynamic scales each projection now quantizes its own input. Found by Greptile review on #45.
…endent NVFP4/FP8 conversion, splice scope - fp8_kernels: int64 row offsets (a 78k-token fc_in output has 2.2e9 elements) - _maybe_quantize_model: handle NVFP4 (+ mixed FP8) before the per-module walk - step splice: primary transformer component only; no env mutation - layerwise offload: tolerate malformed / all-resident FASTVIDEO_LAYERWISE_RESIDENT_BLOCKS - NVFP4 encoder fallback: same input dtype contract as the FP4 path - benchmarks: new _build_block_mask contract in bench_code.py, extra_env overrides, --timed validation, generator shutdown on failure - drop the stale playground 'Older clip' assertion (that UI change did not survive the rebase)
…emory-limited GPUs
torch's pinned allocator rounds every block up to a power of two, so parking the pruned NVFP4 DiT (~20 GB) next to the NVFP4 encoder (15 GB) overran a 60 GB container on the RTX 5090. One registered arena per module pins exactly the bytes needed; buffers now keep a persistent host copy as well.
The export wrote straight into the output directory, so an out-of-memory or out-of-disk failure left partial files and the retry hit FileExistsError. It now builds in a sibling staging directory, renames it into place on success and removes it on failure.
…e the wired limit on close With the converter's AdaLN cache, dropped projection weights are None and mx.eval rejected them during resident preparation. metal_wired_limit_gib is process-wide, so close() now restores the previous limit instead of leaking it into later pipelines.
…X env migration Track C already declared FASTVIDEO_NVFP4_MM_BACKEND and FASTVIDEO_H3_VAE_TILE_BATCH (and FASTVIDEO_VSA_TRITON for Ray workers); drop the duplicate declarations from the RTX migration, override the registered FASTVIDEO_VSA_TRITON in the tile-first test, and regenerate the env-var table.
aryan5v
force-pushed
the
fasth3-spark-mlx
branch
from
October 5, 2026 19:10
82c1aa1 to
6ccdbc7
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Runs FastH3 V2 and FastH3 Trim on DGX Spark (one or two) and Apple Silicon Macs. Stacked on #1919; review the commits after
aryan5v:fasth3-rtx.What's in it
Results
5 s clips with audio at 832×480, end to end (median of two runs on each of two prompts).
Blog post: hao-ai-lab/hao-ai-lab.github.io#108.