Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
16 changes: 16 additions & 0 deletions docs/quantization/h3_int8_affine.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
# INT8 affine (group-64) for MiniMax-H3

Weight-only load-time INT8 for the H3 DiT on CUDA.
Implementation: `fastvideo/layers/quantization/int8_affine_config.py`.

```python
from fastvideo.layers.quantization.int8_affine_config import INT8AffineConfig

fastvideo_args.transformer_quant = INT8AffineConfig.for_minimax_h3()
```

Quantizes attention and FFN linears. Excludes `attn.to_gate_compress`,
`adaln_basis`, fp32-pinned I/O projections, and norms.

Inference only; not the MLX QAT callback. Tests:
`fastvideo/tests/ops/quantization/test_int8_affine_config.py`.
16 changes: 16 additions & 0 deletions docs/quantization/h3_nvfp4.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
# NVFP4 for MiniMax-H3

Load-time NVFP4 for the H3 DiT on Blackwell (sm100+).
Implementation: `fastvideo/layers/quantization/nvfp4_config.py` (FlashInfer).

```python
from fastvideo.layers.quantization.nvfp4_config import NVFP4Config

fastvideo_args.transformer_quant = NVFP4Config.for_minimax_h3()
```

Requires FlashInfer with NVFP4 support. Excludes `attn.to_gate_compress`.
Compact checkpoints use the NVFP4 sidecar helpers in the same module.

Tests: `fastvideo/tests/ops/quantization/test_nvfp4_h3_prefixes.py`,
`test_nvfp4_sidecar.py`.
15 changes: 15 additions & 0 deletions docs/quantization/h3_w4a16.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,15 @@
# W4A16 for MiniMax-H3

Weight-only 4-bit storage with bf16/fp16 activations.
Implementation: `fastvideo/layers/quantization/w4a16_config.py`.

```python
from fastvideo.layers.quantization.w4a16_config import W4A16Config

fastvideo_args.transformer_quant = W4A16Config.for_minimax_h3()
```

This is a memory lane: weights are stored in 4-bit form and dequantized before
each dense GEMM. There is no fused W4A16 kernel in-tree yet.

Tests: `fastvideo/tests/ops/quantization/test_w4a16_config.py`.
17 changes: 17 additions & 0 deletions docs/quantization/loader_quant_params.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
# Quantized models and loader allowlists

Some quantization methods register scale tensors that are absent from dense
checkpoints. Without an allowlist entry, loading fails with:

```
Unsupported new parameter: ...scale_weight...
```

`ALLOWED_NEW_PARAM_PATTERNS` in `fastvideo/models/loader/fsdp_load.py` admits
expected new parameter leaf names via substring match. Prefer
`register_buffer(..., persistent=False)` for values recomputed at load time so
they never enter `state_dict()`.

When adding a quant config that calls `register_parameter` for scales, add the
leaf name to `ALLOWED_NEW_PARAM_PATTERNS` and mirror it in
`fastvideo/models/loader/shard_cache.py`.
2 changes: 1 addition & 1 deletion docs/training/attn_qat.md
Original file line number Diff line number Diff line change
Expand Up @@ -105,7 +105,7 @@ The migrated recipe preserves these behaviors:
|---|---|
| Student fake-quantized attention | `models.student.attention_backend: ATTN_QAT_TRAIN` |
| Teacher and critic full-precision attention | Role-local `FLASH_ATTN` |
| Generator update every five critic steps | `method.generator_update_interval: 5` |
| Four critic-only steps, then one student-only step | `method.generator_update_interval: 5` |
| Three-step rollout | `method.dmd_denoising_steps: [1000, 757, 522]` |
| Score timestep range | `method.min_timestep_ratio: 0.02`, `max_timestep_ratio: 0.98` |
| Legacy guidance `cond + 2(cond - uncond)` | Standard CFG scale `3.0` |
Expand Down
10 changes: 10 additions & 0 deletions examples/compacth3/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,10 @@
# CompactH3

| Area | Path |
|---|---|
| Block scoring, folding, recovery | `scripts/fasth3_sprint/` |
| Recovery configs | `examples/train/configs/fasth3_*.yaml` |
| DMD2 / QAD | `examples/train/configs/distribution_matching/minimax_h3/` |
| Eval, QAD, export | `scripts/compacth3/` |

Checkpoints and generated media are excluded. Override cluster paths in launchers when deploying elsewhere.
26 changes: 26 additions & 0 deletions examples/distill/MiniMax-H3/distill_dmd.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,26 @@
#!/usr/bin/env bash

set -euo pipefail

SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
REPO_ROOT="$(cd "${SCRIPT_DIR}/../../.." && pwd)"
cd "${REPO_ROOT}"

export MASTER_PORT="${MASTER_PORT:-29513}"
export FASTVIDEO_FA4="${FASTVIDEO_FA4:-1}"

export NUM_GPUS="${NUM_GPUS:-4}"
WORLD_SIZE="${NUM_GPUS}"
SP_SIZE="${SP_SIZE:-1}"
HSDP_REPLICATE="${HSDP_REPLICATE:-1}"
HSDP_SHARD="${HSDP_SHARD:-${WORLD_SIZE}}"
CONFIG="${CONFIG:-examples/train/configs/distribution_matching/minimax_h3/dmd2_sp1_fsdp40_vidprom_v6.yaml}"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Default config is missing

The default points to dmd2_sp1_fsdp40_vidprom_v6.yaml, but that file is not present in the added MiniMax-H3 configuration directory. Running this launcher without manually setting CONFIG therefore passes a nonexistent file to examples/train/run.sh and aborts before training begins.

Prompt To Fix With AI
This is a comment left during a code review.
Path: examples/distill/MiniMax-H3/distill_dmd.sh
Line: 17

Comment:
**Default config is missing**

The default points to `dmd2_sp1_fsdp40_vidprom_v6.yaml`, but that file is not present in the added MiniMax-H3 configuration directory. Running this launcher without manually setting `CONFIG` therefore passes a nonexistent file to `examples/train/run.sh` and aborts before training begins.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Codex Fix in Claude Code Fix in Cursor Fix in Conductor

OUTPUT_DIR="${OUTPUT_DIR:-outputs/minimax_h3_dmd2_local}"

exec bash examples/train/run.sh "${CONFIG}" \
--training.distributed.num_gpus "${WORLD_SIZE}" \
--training.distributed.sp_size "${SP_SIZE}" \
--training.distributed.hsdp_replicate_dim "${HSDP_REPLICATE}" \
--training.distributed.hsdp_shard_dim "${HSDP_SHARD}" \
--training.checkpoint.output_dir "${OUTPUT_DIR}" \
"$@"
63 changes: 63 additions & 0 deletions examples/inference/minimax_h3/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,63 @@
# MiniMax-H3 inference examples

Basic single-request H3 examples live in `examples/inference/basic/`
(`basic_minimax_h3_t2v.py`, `basic_minimax_h3_fl2va.py`,
`basic_minimax_h3_ref2va.py`). This directory holds H3-specific benchmark
tooling.

## `h3_vsa_dmd.py` — VSA-H3 vs dense attention, few-step DMD inference

Benchmarks 3-step (DMD-style) H3 T2VA inference under two attention
backends and prints a latency/speedup table:

- `dense` — `FASTVIDEO_ATTENTION_BACKEND=FLASH_ATTN` with FA4
(`FASTVIDEO_FA4=1`). If the flash-attn package is not installed the
FLASH_ATTN request falls back to Torch SDPA (the worker log prints
"Using Torch SDPA backend"); the baseline is then SDPA, not FA4.
- `vsa` — `FASTVIDEO_ATTENTION_BACKEND=VIDEO_SPARSE_ATTN_H3` at
`--sparsity` (default 0.9), applied at generator boot through
`FastVideoArgs.VSA_sparsity` (`pipeline.experimental`).
`--vsa-tile-size {64,256}` (default 256) flows the same way
(`FastVideoArgs.VSA_tile_size`). At tile 256, `--vsa-kernel triton`
(default, no optional dependencies) uses the 256-to-64 expansion path
and `cutedsl` opts into the FA4 CuTe 256-tile forward (requires the
optional FA4 CuTe build, `flash_attn.cute`); at tile 64 the forward is
always the native 64-token Triton kernel and `--vsa-kernel` is ignored.
- `microbench` — model-free per-attention-layer proxy on the exact packed
H3 sequence geometry (dense FA4/SDPA vs the full `MiniMaxH3VSAImpl`
tile/pool/top-k/kernel/untile path). Useful standalone, and as the
speedup proxy when the full VSA pipeline leg is unavailable.

Each mode boots its own generator in a fresh subprocess (the backend env
var is resolved at boot), runs `--warmup` untimed request(s), then times
`--num-prompts` requests with fixed seeds shared across modes so the
per-mode videos can be eyeballed against each other. Model-load time is
reported separately from per-request latency. A crash in one mode is
contained: its signature is saved to `<output>/<mode>/crash_signature.txt`
and the remaining modes still report.

```bash
FASTVIDEO_FA4=1 python examples/inference/minimax_h3/h3_vsa_dmd.py \
--model-path /path/to/MiniMax-H3 \
--prompts-json /path/to/validation.json \
--num-prompts 4 \
--output-dir outputs/h3_vsa_dmd \
--modes dense,vsa,microbench \
--dmd-steps 1000,667,333 \
--num-gpus 4
```

`--prompts-json` expects `{"data": [{"caption": ...}]}`; without it a
built-in prompt set is used. Results land in
`<output>/<mode>/results.json`, per-mode videos in `<output>/<mode>/`, and
an aggregate `summary.json` plus a final table on stdout.

### Caveat: dense-trained checkpoints under VSA

With the base (dense-trained) H3 checkpoint this benchmark measures SPEED
only. The base model was never trained under VSA top-k masks, so at 90%
sparsity output-quality parity is not expected — judge quality with a
VSA-trained (sparse-student) DMD checkpoint. The 3-step DMD ladder applied
to the base checkpoint is likewise a latency proxy for a distilled
student, not a quality reference: real few-step quality requires a DMD
student checkpoint.
Loading
Loading