Skip to content

[feat]: Track C FastH3 resident Spark and MLX support - #47

Open
aryan5v wants to merge 15 commits into
h3-spark-main-basefrom
h3-spark-mlx
Open

aryan5v wants to merge 15 commits into
h3-spark-main-basefrom
h3-spark-mlx

Conversation

@aryan5v

@aryan5v aryan5v commented Oct 3, 2026 •

Copy link
Copy Markdown
Owner

Track C: Spark and Apple Silicon

This is the Track C release branch. It is a draft while full Mac benchmarks and the Spark repeat-quality failure are being resolved.

Base

Stacked on h3-spark-main-base, which integrates h3-sm120-experimental with the latest upstream main checked on October 3 (0cc41a22). The base carries the experimental CUDA NVFP4 prerequisites. This keeps the Track C diff focused; upstream delivery will stack on the reviewed release core.

Changes

  • Add rank-16 AdaLN support for the 42-block pruned FastH3 MLX transformer, retaining upstream's eight-forward DMD schedule validation.
  • Read FP8 transformer sources for local MLX conversion and expose experimental native MXFP8 output; affine INT8/INT6/INT4 remain the defaults.
  • Load the released 50-layer NVFP4 encoder through native MLX packed matrix multiplication, with FP32 activations and BF16 embedding storage.
  • Add resident encoder/DiT/video-VAE/audio-VAE loading, reuse, cleanup and a phased fallback.
  • Add resident one-Spark and two-Spark V2/pruned recipes and a benchmark harness. Preserve per-stage CUDA peaks before reset.
  • Resolve small loader/worker/H3 helper incompatibilities from integrating current main.

Weights and workload

  • Mac target: FastVideo/FastH3-Pruned-8Step-BF16-ckpt300, locally converted to INT8 and INT6. Normal V2 BF16-source conversion is also required. Complete local conversions and end-to-end clips are pending weight transfer.
  • Spark V2: FastVideo/FastVideo-FastH3-8-Step-V2-NVFP4, converted to the experimental packed DiT layout; released NVFP4 encoder and light H3 VAE. A separate diagnostic artifact restores 100 source FFN activation scales without further quantization.
  • Spark pruned: FastVideo/FastH3-Pruned-8Step-NVFP4-ckpt300; testing in progress.
  • Native 832×480, 243 frames is the headline; 124 frames is the five-second comparison. Both release prompts, VSA 0.8, eight forwards, video/audio decode and MP4 writing are included. No temporal dropping or spatial fast modes.

Measurements and quality status

Full records and commands: release plan, section 6.

Device / format Published old 124-frame e2e New result
M4 Max INT8 481 s pending
M4 Max INT6 456 s pending
One Spark 243 s V2 harbor 141.077 s median; both timed clips visually coherent; overall V2 reliability blocked
Two Sparks 209 s pending

Old values are from FastH3 Goes Local: four-step Preview and full H3 VAE. New runs use a different checkpoint, eight forwards, VSA 0.8 and the light VAE. These are not controlled model-only speedup comparisons. The blog has no matching 243-frame baseline or release-prompt results.

The first V2 243-frame probe produced colored noise. Restoring the omitted ModelOpt FFN input scales fixes a real activation-clipping defect, but one 124-frame ceramics repeat remains gray noise. Its 142.943 s median is invalid and must not be advertised. Stage hashes, activation ranges and encoder backend controls are being used to locate the remaining failure. Saved clips and review sheets are listed in the plan.

Validation

  • Focused MLX parity, scheduler, conditioner, packed-storage and resident-lifecycle checks: 31 passed on the M4 Max.
  • Integrated Spark H3 schedule/VSA/sequential tests: 62 passed.
  • Pre-commit passes for the owned runtime and docs paths. Scripts/test paths excluded by repository configuration were not linted separately.
  • Native MXFP8 matrix checkpoint round trip passes; no full MXFP8 checkpoint quality claim.
  • Light VAE FP16 versus FP32 storage gives identical finite output for a same-latent 256×256, 22-frame decoder control; this does not establish full-generation quality.

Remaining release gates

  • Fix and verify the Spark repeat-quality issue; complete one-Spark and two-Spark V2/pruned runs.
  • Complete Mac INT8/INT6 V2/pruned conversion, resident memory attempts and both frame-count benchmarks.
  • Review every saved video and its audio, update all measured medians/peaks and the old/new comparison, and remove draft status only for validated code/results.

Mac component measurement update

The released NVFP4 encoder now runs natively on the M4 Max with packed W4 weights and FP32 activations. The first 50 language-model layers and BF16 embedding use 14.224 GiB active and 15.209 GiB peak. Both required prompts return finite 1000×5120 hidden states. After warmup, ceramics encoding is 8.890 / 8.871 s and harbor is 8.819 / 8.846 s. These are encoder component timings; Mac end-to-end INT6 and INT8 timings remain pending source transfer and conversion. No prompt cache is used for clip timing.

A focused rank-16 test checks all 42 blocks and final modulation against independent NumPy projection calculations, then verifies cache reload. It passed; pre-commit respects the deliberate test-file excludes. The prior focused suite had 31 passing tests.

The official pruned Spark artifact also fails repeated visual quality: ceramics 142.436 s warmup, 131.959 / 132.905 s timed calls, with corrupted timed outputs. The 132.432 s median is invalid for release. Harbor warmup was 132.509 s before stopping. All raw records are in the release plan; no speedup is claimed from corrupted clips.

RetriggerConfidence Score: 4/5

The PR does not yet appear safe to merge because resident preparation still fails on checkpoints with dropped AdaLN weights.

Fix All in CodexFindings

  1. P1 Resident preload evaluates dropped weights ▶
  2. P2 FP8 source conversion lacks coverage ▶
Fix with agent prompt
### Issue 1
fastvideo/mlx_runtime/minimax_h3_pipeline.py:445-447
When an MLX checkpoint includes the AdaLN cache produced by the documented converter, its block dictionaries contain `None` where AdaLN projection weights were dropped. This loop passes those values to `mx.eval`, so resident preparation fails before generation. Skip the dropped values and test preload with a converted checkpoint.

```suggestion
            for group in [dit.weights, *dit.blocks, *dit.refiner]:
                for value in group.values():
                    if value is not None:
                        _eval_value(value)
```

### Issue 2
fastvideo/mlx_runtime/tests/test_minimax_h3_fp8_checkpoint.py:11-12
This test starts with an already-quantized MLX matrix, so it checks checkpoint save/load but not the new conversion of uint8 FP8 source weights. A scale-decoding or AdaLN-cache regression in the documented FP8-source workflow could go unnoticed. Add a small source safetensors fixture that checks the converted values and cache against a reference.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.
Summary

This PR adds Spark NVFP4 recipes and benchmarking, MLX support for the pruned FastH3 checkpoint and packed encoder, and resident component loading. Since the previous review, it adds a synthetic rank-16 AdaLN arithmetic and checkpoint round-trip test.

Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[Source checkpoint] --> B[MLX conversion and AdaLN cache]
  B --> C[Resident component loading]
  C --> D[Prompt conditioning]
  D --> E[DiT denoising]
  E --> F[Video and audio decoding]
Loading

Reviews (2) · Last reviewed commit: "[test]: verify rank-16 H3 modulation and..."

@coderabbitai

coderabbitai Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration
  • Configuration used: defaults
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: a245965b-9de6-4c13-a46d-d44686a72377

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@aryan5v
aryan5v marked this pull request as ready for review October 3, 2026 23:57
Comment on lines +445 to +447
for group in [dit.weights, *dit.blocks, *dit.refiner]:
for value in group.values():
_eval_value(value)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Resident preload evaluates dropped weights When an MLX checkpoint includes the AdaLN cache produced by the documented converter, its block dictionaries contain None where AdaLN projection weights were dropped. This loop passes those values to mx.eval, so resident preparation fails before generation. Skip the dropped values and test preload with a converted checkpoint.

Suggested change
for group in [dit.weights, *dit.blocks, *dit.refiner]:
for value in group.values():
_eval_value(value)
for group in [dit.weights, *dit.blocks, *dit.refiner]:
for value in group.values():
if value is not None:
_eval_value(value)
Prompt To Fix With AI
This is a comment left during a code review.
Path: fastvideo/mlx_runtime/minimax_h3_pipeline.py
Line: 445-447

Comment:
**Resident preload evaluates dropped weights** When an MLX checkpoint includes the AdaLN cache produced by the documented converter, its block dictionaries contain `None` where AdaLN projection weights were dropped. This loop passes those values to `mx.eval`, so resident preparation fails before generation. Skip the dropped values and test preload with a converted checkpoint.

```suggestion
            for group in [dit.weights, *dit.blocks, *dit.refiner]:
                for value in group.values():
                    if value is not None:
                        _eval_value(value)
```

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Codex Fix in Claude Code Fix in Cursor Fix in Conductor

Comment on lines +11 to +12
def test_mxfp8_checkpoint_preserves_quantized_matrix(tmp_path):
spec = MLXQuantizationSpec.from_name('mxfp8')

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 FP8 source conversion lacks coverage This test starts with an already-quantized MLX matrix, so it checks checkpoint save/load but not the new conversion of uint8 FP8 source weights. A scale-decoding or AdaLN-cache regression in the documented FP8-source workflow could go unnoticed. Add a small source safetensors fixture that checks the converted values and cache against a reference.

Prompt To Fix With AI
This is a comment left during a code review.
Path: fastvideo/mlx_runtime/tests/test_minimax_h3_fp8_checkpoint.py
Line: 11-12

Comment:
**FP8 source conversion lacks coverage** This test starts with an already-quantized MLX matrix, so it checks checkpoint save/load but not the new conversion of uint8 FP8 source weights. A scale-decoding or AdaLN-cache regression in the documented FP8-source workflow could go unnoticed. Add a small source safetensors fixture that checks the converted values and cache against a reference.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Fix in Codex Fix in Claude Code Fix in Cursor Fix in Conductor

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant