Skip to content

[feat]: FastH3 on DGX Spark and Apple Silicon - #1920

Open
aryan5v wants to merge 100 commits into
hao-ai-lab:mainfrom
aryan5v:fasth3-spark-mlx
Open

aryan5v wants to merge 100 commits into
hao-ai-lab:mainfrom
aryan5v:fasth3-spark-mlx

Conversation

@aryan5v

@aryan5v aryan5v commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Runs FastH3 V2 and FastH3 Trim on DGX Spark (one or two) and Apple Silicon Macs. Stacked on #1919; review the commits after aryan5v:fasth3-rtx.

What's in it

  • DGX Spark: resident NVFP4 encoder, transformer and light VAE recipes for one or two Sparks, and the GB10 activation-quantization fence needed for correct repeated runs.
  • Apple Silicon: BF16 → MLX INT6 conversion with VSA gates and Trim's rank-16 AdaLN, the 8-step schedule, native NVFP4 text encoder, and phased placement on 36 GB Macs.
  • Tests: MLX CPU checks for the conversion, schedule and placement; the dequant-GEMM test now allows MLX CPU accumulation error.

Results

5 s clips with audio at 832×480, end to end (median of two runs on each of two prompts).

Device V2 Trim
DGX Spark 141.4 s 125.8 s
2× DGX Spark 87.2 s 78.3 s
Mac, M4 Max (INT6) — 925.2 s

Blog post: hao-ai-lab/hao-ai-lab.github.io#108.

aryan5v and others added 30 commits October 3, 2026 10:32
Keep encoder/DiT/VAE off disk between clips, overlay the Comfy int8-convrot decoder for quality, and let the playground queue prompts while switching clips. TAEH3 stays an optional preview config only.

(cherry picked from commit 15e5517)
Drop playground queue experiments from this PR, reject ConvRot overlays that cannot rotate activations, and skip FSDP2 modules that mix dense params with DTensors.

(cherry picked from commit f16b87e)
Add H3 cookbook recipes for the 42-block checkpoint, reject ConvRot overlays whose group size cannot rotate activations, and skip FSDP2 modules that mix dense parameters with DTensors.

(cherry picked from commit 7852505)
Dense CompactH3 constructed VIDEO_SPARSE_ATTN_H3 tiles with all-zero gates. Fail at load and first denoise, and drop restating CompactH3 banners.

(cherry picked from commit cc59773)
Snapshot of every RTX PRO 6000 experiment, including the Ulysses FP8
q/k/v exchange path that has not executed yet (8-GPU capacity was
unavailable) and the Modal drivers used for every measurement. The
ship-ready subset is on h3-sm120-sparse-fp4.

(cherry picked from commit da835b4)
- Pre-quantized FP8 W8A8 checkpoint loader (float8 weight + per-channel weight_scale)
- NVFP4 export: optional calibrated activation scale (_nvfp4_input_global_sf);
  converter gains --quantize-ffn (from bf16) and --act-amax
- NVFP4 static/dynamic activation scales via env; FP8 attention projections next to NVFP4 FFN
- AdaLN modulation host cache and precomputed tables (skips 24 GiB of AdaLN weights)
- Layerwise offload streams large buffers (fix: onload by name, placeholders fail the size test)
- CPU-first DiT load for layerwise/AdaLN-cache paths; pinned encoder/VAE swaps; parked modules
- NVFP4 text encoder bf16 de-quant fallback for pre-Blackwell GPUs
- VSA guard ignores offloaded placeholders; per-stage memory logging; memory cap / report knobs
- SP stage profiling and FP8 all-to-all simulation; Modal PRO 6000 bench steps

(cherry picked from commit cf04434)
- FP8 on sm89: per-tensor GEMM + Triton per-token x per-channel scale epilogue
  (torch rowwise _scaled_mm runs ~70 TFLOPS there, below bf16) and a fused
  one-launch per-token quantize (5x faster than the torch chain)
- FASTVIDEO_H3_FFN_CHUNK_TOKENS: inference-only FFN token chunking
- FASTVIDEO_LAYERWISE_RESIDENT_BLOCKS: keep the first N blocks resident
- Serialized NVFP4 text encoder: allow sm80-sm90 through the bf16 de-quant path
- FASTVIDEO_H3_SPLICE_TRANSFORMER / _FROM_STEP: second checkpoint runs late DMD steps

(cherry picked from commit e2ed39c)
…and Modal

bench_headline.py times the release protocol (two fixed prompts, one warmup,
two timed runs each, generate_video wall time) from a checkpoint's
fastvideo_inference.json contract, with optional W&B logging.
headline_app.py runs it on 1/4/8 RTX PRO 6000 Blackwell GPUs on Modal from a
FastVideo HF repo.
…K/V quantization only at unit scale

Upstream's _build_block_mask now takes per-region video tile spans and
sparsities; the FP4 VSA paths still passed the old five arguments and raised
TypeError on the first block. The shared Q/K/V quantization assumed every
NVFP4 layer used the unit activation scale; with calibrated (export or env)
or dynamic scales each projection now quantizes its own input.

Found by Greptile review on #45.
…endent NVFP4/FP8 conversion, splice scope

- fp8_kernels: int64 row offsets (a 78k-token fc_in output has 2.2e9 elements)
- _maybe_quantize_model: handle NVFP4 (+ mixed FP8) before the per-module walk
- step splice: primary transformer component only; no env mutation
- layerwise offload: tolerate malformed / all-resident FASTVIDEO_LAYERWISE_RESIDENT_BLOCKS
- NVFP4 encoder fallback: same input dtype contract as the FP4 path
- benchmarks: new _build_block_mask contract in bench_code.py, extra_env overrides,
  --timed validation, generator shutdown on failure
- drop the stale playground 'Older clip' assertion (that UI change did not survive the rebase)
torch's pinned allocator rounds every block up to a power of two, so parking
the pruned NVFP4 DiT (~20 GB) next to the NVFP4 encoder (15 GB) overran a
60 GB container on the RTX 5090. One registered arena per module pins exactly
the bytes needed; buffers now keep a persistent host copy as well.
Aryan Kumar and others added 20 commits October 5, 2026 12:10
The export wrote straight into the output directory, so an out-of-memory
or out-of-disk failure left partial files and the retry hit
FileExistsError. It now builds in a sibling staging directory, renames it
into place on success and removes it on failure.
…e the wired limit on close

With the converter's AdaLN cache, dropped projection weights are None and
mx.eval rejected them during resident preparation. metal_wired_limit_gib
is process-wide, so close() now restores the previous limit instead of
leaking it into later pipelines.
…X env migration

Track C already declared FASTVIDEO_NVFP4_MM_BACKEND and
FASTVIDEO_H3_VAE_TILE_BATCH (and FASTVIDEO_VSA_TRITON for Ray workers);
drop the duplicate declarations from the RTX migration, override the
registered FASTVIDEO_VSA_TRITON in the tile-first test, and regenerate the
env-var table.
@aryan5v aryan5v changed the title [feat]: FastH3 on DGX Spark and Apple Silicon: resident Spark recipes, pruned MLX H3, NVFP4 conditioner [feat]: Track C FastH3 on DGX Spark and Apple Silicon Oct 5, 2026
@aryan5v aryan5v changed the title [feat]: Track C FastH3 on DGX Spark and Apple Silicon [feat]: FastH3 on DGX Spark and Apple Silicon Oct 5, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

scope: attention Attention backends (VSA, STA, Flash, etc.) scope: docs Documentation scope: inference Inference pipeline, serving, CLI scope: infra CI, tests, Docker, build scope: kernel CUDA kernels, fastvideo-kernel scope: model Model architecture (DiTs, encoders, VAEs) type: feat New feature or capability

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant