Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions docs/cookbook/minimax-h3.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,11 @@ a full model, not a demo. **V2** is the eight-step checkpoint. More forwards
is why V2 is the higher-quality FastH3. The V2 schedule contract is in
[FastH3 distilled checkpoint schedules](../inference/fasth3-distilled.md).

The 42-block pruned checkpoint has an [MLX INT8/INT6 conversion and
eight-forward T2VA command](../getting_started/installation/mlx.md#pruned-eight-forward-checkpoint).
It reads `fastvideo_inference.json` for the trained schedule. The command
uses native 832x480 resolution and all requested frames.

<div class="cookbook-shell cookbook-family-page" data-cookbook data-family="minimax_h3" data-default-recipe="fasth3-preview-cuda" data-recipes="../../assets/cookbook-recipes.json?v=11">
<header class="cookbook-family-header">
<a class="cookbook-back-link" href="../"><span aria-hidden="true">←</span> All model families</a>
Expand Down
82 changes: 82 additions & 0 deletions docs/getting_started/installation/mlx.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,6 +46,88 @@ is the higher-quality FastH3.
Recorded shapes and evidence live in the
[support matrix](../../inference/support_matrix.md#apple-silicon-native-runtime).

## Pruned eight-forward checkpoint

The pruned FastH3 checkpoint has 42 transformer blocks and rank-16 AdaLN.
Its `fastvideo_inference.json` fixes eight denoising forwards, video/audio
shifts of 10/3, and VSA sparsity 0.8. Keep that file beside the transformer
when converting. The converter reads its schedule to build the AdaLN cache.

```bash
hf download FastVideo/FastH3-Pruned-8Step-BF16-ckpt300 \
--local-dir ./FastH3-Pruned-8Step-BF16-ckpt300 \
--exclude 'text_encoder/*'

# Optional BF16 encoder fallback: stream the first 50 language layers.
# The last three shards are unused. The packed NVFP4 option is described below.
hf download MiniMaxAI/MiniMax-H3 \
--local-dir ./FastH3-Pruned-8Step-BF16-ckpt300 \
--include 'text_encoder/model-0000[1-9]-of-00014.safetensors' \
--include 'text_encoder/model-0001[0-1]-of-00014.safetensors' \
--include 'text_encoder/model.safetensors.index.json' \
--include 'text_encoder/config.json'

python scripts/checkpoint_conversion/convert_minimax_h3_mlx.py \
--model-root ./FastH3-Pruned-8Step-BF16-ckpt300/transformer \
--out ./FastH3-Pruned-MLX-vsa \
--formats "int8 int6" --include-vsa

python examples/inference/basic/mlx_fasth3.py \
--model-root ./FastH3-Pruned-8Step-BF16-ckpt300 \
--mlx-checkpoint ./FastH3-Pruned-MLX-vsa/int8 \
--prompt "(S1) A potter asks <d>[English] Is the rim ready?</d>" \
--height 480 --width 832 --num-frames 243 --steps 8 \
--vsa --vsa-sparsity 0.8 --vsa-tile-size 64 \
--output-path ./outputs/fasth3_pruned_int8_480p.mp4
```

At 24 fps, 124 frames is the legal H3 count for a roughly five-second clip.
Use `--num-frames 124` and a separate output path for that run. The `--fast`
and `--fast-spatial` options change the workload and are not part of the
native-resolution benchmark. A 36 GB Mac may need INT6 and phased loading;
measure memory before claiming all-resident operation.

### Packed encoder and resident loading

The experimental MLX conditioner can read the released FastVideo NVFP4
text encoder directly, using native `nvfp4` matrix multiplication. It keeps
the packed weights and BF16 embedding table in memory, with FP32
activations. CUDA uses quantized activations, so the two encoders are not
bit-exact. Validate generated video and audio before publishing a timing.
MLX 0.32.2 supports the required operator on Apple Silicon.

Pass the packed encoder directory as `conditioner_dir`; `conditioner_mode="auto"`
selects it from `config.json`. The BF16 fallback continues to stream layers.
To request all-resident generation through the Python API:

```python
from fastvideo.mlx_runtime.minimax_h3_pipeline import MiniMaxH3MLXPipeline

pipeline = MiniMaxH3MLXPipeline(
model_root="./FastH3-Pruned-8Step-BF16-ckpt300",
mlx_dit_checkpoint="./FastH3-Pruned-MLX-vsa/int6",
conditioner_dir="./FastH3-NVFP4-encoder",
conditioner_mode="nvfp4",
resident=True,
vae_dtype="fp16",
)
try:
pipeline.prepare_resident() # Load encoder, DiT, video VAE and audio VAE.
result = pipeline.generate(
"(S1) A potter asks <d>[English] Is the rim ready?</d>",
output_path="./outputs/fasth3_pruned_resident.mp4",
height=480, width=832, num_frames=243, num_steps=8,
vsa=True, vsa_sparsity=0.8, vsa_tile_size=64,
)
finally:
pipeline.close()
```

Resident placement requires space for activations as well as all four
components. On a 36 GiB Mac, try INT6 first and measure peak allocation.
If loading or inference runs out of memory, use phased loading by leaving
`resident=False`. Changing placement does not change frames or resolution.

## Hardware

- FastMetal 1.3B and 5B: 16 GB unified memory and up
Expand Down
3 changes: 3 additions & 0 deletions docs/getting_started/installation/spark.md
Original file line number Diff line number Diff line change
Expand Up @@ -144,6 +144,9 @@ for which models are practical on the GB10, what makes them faster, and what
won't help on this hardware (and why) — so you don't spend a night tuning knobs
that can't move here.

For the eight-forward FastH3 V2 NVFP4 stack with a trimmed encoder and light
VAE, use the [one-Spark resident recipe](spark_performance.md#fasth3-v2-nvfp4-on-one-spark).

Two Sparks with QSFP cables: [Pair two NVIDIA DGX Sparks](spark_pair.md) for
one FastH3 clip across both GPUs (`sp_size=2` over Ray). Copy-paste commands
for one or two Sparks also live on the
Expand Down
111 changes: 104 additions & 7 deletions docs/getting_started/installation/spark_performance.md
Original file line number Diff line number Diff line change
Expand Up @@ -161,14 +161,16 @@ is power-cycled. To avoid it:
on: "CPU" offload uses the same unified RAM. Multi-GPU FSDP sharding remains
available because it partitions weights without parking them in a separate
host pool.
- **MiniMax H3 / FastH3** still needs deferred loading on one GB10. The Qwen3-VL
conditioner is tens of gigabytes of BF16. If the DiT and VAEs load while that
encoder is still resident, the process is a typical `earlyoom` kill (Python is
preferred). On unified memory, `lazy_module_load` auto-enables and owns that
- **Older MiniMax H3 / FastH3 bf16 weights** need deferred loading on one GB10.
The full Qwen3-VL conditioner is tens of gigabytes of BF16. If the DiT and
VAEs load while that encoder is still resident, the process can be killed by
`earlyoom`. On unified memory, `lazy_module_load` auto-enables and owns that
split (encoder, then DiT, then VAE; DiT can drop before decode). Sequential
load is the H3-only fallback when lazy is off; do not pass
`--no-lazy-module-load` here. Geometry scalars come from checkpoint
`config.json`, not live weights. See [Offloading](../../inference/offloading.md).
load is the H3-only fallback when lazy is off. Keep deferred loading for
those older checkpoints. The trimmed NVFP4 encoder and light VAE in the
[V2 resident recipe](#fasth3-v2-nvfp4-on-one-spark) are a different memory
profile. Geometry scalars come from checkpoint `config.json`, not live
weights. See [Offloading](../../inference/offloading.md).
- **FastH3 TAEH3** (`--video-decode-backend taeh3`) is an opt-in preview decoder.
T2VA never materializes the 9.7 GiB video VAE (DiT still loads after Qwen via
sequential start). On this box, alpine 768×1344×124 decoded in **2.4 s** versus
Expand Down Expand Up @@ -207,6 +209,101 @@ A few things that surprise people on this box (beyond the memory notes above):
`Released MiniMax-H3 text encoder after conditioning` before
`Loading MiniMax-H3 denoise modules`).

## FastH3 V2 NVFP4 on one Spark

This recipe uses the full V2 eight-forward transformer, the 50-layer NVFP4
Qwen3-VL encoder, and the light H3 video VAE. Its configuration keeps all
three resident on one GB10. Runtime, memory fit, and quality still need a run
on that device. The earlier bf16 H3 memory guidance above concerns a larger
checkpoint.

Install FastVideo from a checkout that includes the ModelOpt converter and
FlashInfer FP4 support, following [the Spark install guide](spark.md). Sign in
to Hugging Face with access to the FastVideo model repositories. Download the
V2 scheduler and audio components, the compact encoder and VAE from the pruned
repo, and the ModelOpt V2 transformer. The pruned model's encoder and VAE are
the same components used by V2.

```bash
SPARK_STACK=./FastH3-V2-Spark-NVFP4
V2_FP4_SRC=./FastH3-V2-ModelOpt-NVFP4

hf download FastVideo/FastVideo-FastH3-8-Step-V2 \
--local-dir "$SPARK_STACK" \
--exclude 'transformer/*' --exclude 'text_encoder/*' --exclude 'vae/*'
hf download FastVideo/FastH3-Pruned-8Step-BF16-ckpt300 \
--local-dir "$SPARK_STACK" \
--include 'text_encoder/*' --include 'vae/*'
hf download FastVideo/FastVideo-FastH3-8-Step-V2-NVFP4 \
--local-dir "$V2_FP4_SRC" --include 'transformer/*'

nice -n 19 python scripts/checkpoint_conversion/convert_minimax_h3_modelopt_nvfp4_dit.py \
--src "$V2_FP4_SRC/transformer" --dst "$SPARK_STACK/transformer" \
--quantize-attention --quantize-gate

test -f "$SPARK_STACK/transformer/nvfp4_weights.safetensors"
test -f "$SPARK_STACK/text_encoder/config.json"
test -f "$SPARK_STACK/vae/config.json"
test -f "$SPARK_STACK/fastvideo_inference.json"
python -m json.tool "$SPARK_STACK/fastvideo_inference.json" >/dev/null
```

The converter probes each packed linear through FlashInfer `mm_fp4`. If that
probe fails on `sm_121`, convert the transformer on another Blackwell GPU and
copy the resulting `transformer/` directory to the Spark. Do not omit
`fastvideo_inference.json`: it supplies V2's trained denoising ladder. The
recipe's `num_inference_steps: 9` means nine sigma points and eight DiT
forwards.

Run `examples/inference/basic/basic_fasth3_spark_v2_nvfp4.yaml` from the
repository root. It uses 832x480, 243 frames, VSA sparsity 0.8 with
64-token tiles, and the full H3 VAE. It does not use frame dropping or spatial
upscaling.

```bash
FASTVIDEO_MINIMAX_H3_FUSIONS=all \
FASTVIDEO_NVFP4_MM_BACKEND=cutlass \
FASTVIDEO_H3_VAE_TILE_BATCH=1 \
FASTVIDEO_VSA_TRITON=1 FASTVIDEO_VSA_SM100A=0 FASTVIDEO_FA4=0 \
FASTVIDEO_ATTENTION_BACKEND=VIDEO_SPARSE_ATTN_H3 \
FASTVIDEO_STAGE_LOGGING=1 \
nice -n 19 fastvideo generate \
--config examples/inference/basic/basic_fasth3_spark_v2_nvfp4.yaml
```

For a roughly five-second clip, set `--request.sampling.num_frames 124` and
write to a separate output path. H3 permits frame counts of `17n+5`; 124 is
the closest legal count above five seconds at 24 fps. For the secondary
10-second setting, set `--request.sampling.width 1344` and
`--request.sampling.height 768`, keeping 243 frames. Use the two prompts in
`handoff_spark_mac/benchmark_prompts.json` from the local release handoff.
The benchmark script runs one warmup and at least two timed generations for
each prompt in one process. It saves the MP4s and prints the wall time, stage
times, peak memory, and median. Set the Spark environment before running it:

```bash
export FASTVIDEO_MINIMAX_H3_FUSIONS=all
export FASTVIDEO_NVFP4_MM_BACKEND=cutlass FASTVIDEO_H3_VAE_TILE_BATCH=1
export FASTVIDEO_VSA_TRITON=1 FASTVIDEO_VSA_SM100A=0 FASTVIDEO_FA4=0
export FASTVIDEO_ATTENTION_BACKEND=VIDEO_SPARSE_ATTN_H3 FASTVIDEO_STAGE_LOGGING=1
nice -n 19 python examples/inference/basic/benchmark_fasth3_spark_nvfp4.py \
--config examples/inference/basic/basic_fasth3_spark_v2_nvfp4.yaml \
--prompts /path/to/fasth3-local-release/handoff_spark_mac/benchmark_prompts.json \
--output-dir outputs/fasth3_spark_v2_nvfp4/benchmark-243 --frames 243

# Repeat with --frames 124 and a different output directory for the five-second check.
```

Record the exact command and commit with the measurements. Review every clip's
video and audio before publishing a quality or speed claim.

After the V2 baseline works, sweep `FASTVIDEO_H3_VAE_TILE_BATCH` and
`FASTVIDEO_NVFP4_MM_BACKEND` on the same prompts. Compare the optional AdaLN
table and VAE compile only with the same frame count, schedule, and VSA
sparsity. The V2 converter packs VSA gates, so its `h3_dit_vsa` profile must
match the recipe. A later pruned NVFP4 transformer uses the separate
`h3_dit_ffn` profile, with attention and VSA gates left dense.

## Reproduce these numbers

Two scripts under `examples/inference/optimizations/` reproduce the claims on
Expand Down
63 changes: 63 additions & 0 deletions examples/inference/basic/basic_fasth3_spark_pair_pruned_nvfp4.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,63 @@
# Start Ray on both Sparks and source spark_pair_env.sh first; see spark_pair.md.
# FastH3 pruned ckpt300 eight-forward video+audio on two DGX Sparks over QSFP RoCE.
# Download the complete checkpoint, including its trained schedule, as described in
# docs/getting_started/installation/spark_performance.md.
#
# GB10 uses Triton VSA. The packed transformer uses NVFP4 FFN weights with bf16 attention
# and VSA gates (layer_profile: h3_dit_ffn). No temporal or spatial fast mode.
#
# FASTVIDEO_MINIMAX_H3_FUSIONS=all FASTVIDEO_NVFP4_MM_BACKEND=cutlass \
# FASTVIDEO_VSA_TRITON=1 FASTVIDEO_VSA_SM100A=0 FASTVIDEO_FA4=0 \
# FASTVIDEO_ATTENTION_BACKEND=VIDEO_SPARSE_ATTN_H3 FASTVIDEO_STAGE_LOGGING=1 \
# fastvideo generate --config examples/inference/basic/basic_fasth3_spark_pair_pruned_nvfp4.yaml
generator:
model_path: FastVideo/FastH3-Pruned-8Step-NVFP4-ckpt300
engine:
num_gpus: 2
execution_backend: ray
use_fsdp_inference: false
quantization:
transformer_quant: NVFP4
layer_profile: h3_dit_ffn
parallelism:
tp_size: 1
sp_size: 2
offload:
dit: false
dit_layerwise: false
text_encoder: false
image_encoder: false
vae: false
pin_cpu_memory: false
lazy_module_load: false
compile:
enabled: false
vae_enabled: false
pipeline:
workload_type: t2v
vae_tiling: true
experimental:
attention_backend: VIDEO_SPARSE_ATTN_H3
VSA_sparsity: 0.8
VSA_tile_size: 64
h3_sequential_load: false
inference_torch_compile: false
vae_parallel_decode: true
vae_parallel_decode_strategy: gather
video_decode_backend: h3-vae
request:
prompt: A quiet pottery studio with a potter finishing a bowl at the wheel.
negative_prompt: ""
sampling:
seed: 2026
height: 480
width: 832
num_frames: 243
fps: 24
num_inference_steps: 9 # nine sigma points, eight DiT forwards
guidance_scale: 1.0
batch_cfg: false
output:
output_path: outputs/fasth3_spark_pair_pruned_nvfp4/
save_video: true
return_frames: false
63 changes: 63 additions & 0 deletions examples/inference/basic/basic_fasth3_spark_pair_v2_nvfp4.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,63 @@
# Start Ray on both Sparks and source spark_pair_env.sh first; see spark_pair.md.
# FastH3 V2 eight-forward video+audio on two DGX Sparks over QSFP RoCE.
# Assemble ./FastH3-V2-Spark-NVFP4 as described in
# docs/getting_started/installation/spark_performance.md.
#
# GB10 uses Triton VSA. The packed transformer must include NVFP4 attention
# and VSA gates (layer_profile: h3_dit_vsa). No temporal or spatial fast mode.
#
# FASTVIDEO_MINIMAX_H3_FUSIONS=all FASTVIDEO_NVFP4_MM_BACKEND=cutlass \
# FASTVIDEO_VSA_TRITON=1 FASTVIDEO_VSA_SM100A=0 FASTVIDEO_FA4=0 \
# FASTVIDEO_ATTENTION_BACKEND=VIDEO_SPARSE_ATTN_H3 FASTVIDEO_STAGE_LOGGING=1 \
# fastvideo generate --config examples/inference/basic/basic_fasth3_spark_pair_v2_nvfp4.yaml
generator:
model_path: ./FastH3-V2-Spark-NVFP4
engine:
num_gpus: 2
execution_backend: ray
use_fsdp_inference: false
quantization:
transformer_quant: NVFP4
layer_profile: h3_dit_vsa
parallelism:
tp_size: 1
sp_size: 2
offload:
dit: false
dit_layerwise: false
text_encoder: false
image_encoder: false
vae: false
pin_cpu_memory: false
lazy_module_load: false
compile:
enabled: false
vae_enabled: false
pipeline:
workload_type: t2v
vae_tiling: true
experimental:
attention_backend: VIDEO_SPARSE_ATTN_H3
VSA_sparsity: 0.8
VSA_tile_size: 64
h3_sequential_load: false
inference_torch_compile: false
vae_parallel_decode: true
vae_parallel_decode_strategy: gather
video_decode_backend: h3-vae
request:
prompt: A quiet pottery studio with a potter finishing a bowl at the wheel.
negative_prompt: ""
sampling:
seed: 2026
height: 480
width: 832
num_frames: 243
fps: 24
num_inference_steps: 9 # nine sigma points, eight DiT forwards
guidance_scale: 1.0
batch_cfg: false
output:
output_path: outputs/fasth3_spark_pair_v2_nvfp4/
save_video: true
return_frames: false
Loading
Loading