Skip to content

[perf]: Track B FastH3 on RTX 4090 and lower-VRAM GPUs - #46

Open
aryan5v wants to merge 24 commits into
h3-release-corefrom
h3-consumer-fp8
Open

aryan5v wants to merge 24 commits into
h3-release-corefrom
h3-consumer-fp8

Conversation

@aryan5v

@aryan5v aryan5v commented Oct 3, 2026 •

Copy link
Copy Markdown
Owner

FastH3's FP8 consumer path previously spent substantial time copying attention layouts, decoding INT8 VAE projections, and moving the text encoder. This adds opt-in RTX 4090 kernels and memory controls that reduce warmed end-to-end generation to 41.75 s for 124 frames at 832×480 and 79.67 s for 243 frames at 832×480, including audio and MP4 export.

Stacked on the shared release core in #45 (h3-release-core, a97d23f09). This is a draft staging PR in the fork; upstream submission follows the shared-core split.

Changes:

  • Exact-size registered pinned arenas with safe ownership and fallback, reducing host pinning overhead.
  • Tile-first DiT projections, shared FP8 input preparation, and direct-strided INT8 QK / BF16 PV sparse attention on sm89. The default attention route remains unchanged.
  • Layerwise streaming of the trimmed Qwen3-VL NVFP4 encoder, with fused dequantization into BF16. Streaming currently supports text-only conditioning.
  • Shared INT8 VAE QKV preparation, transposed weight views, and an exact-rounding Triton output epilogue that avoids large FP32 intermediates.
  • Reproducible consumer benchmark recipes, stage/host memory measurements, and sampled NVML total GPU memory.

Measured on one RTX 4090, checkpoint FastH3-Pruned-8Step-FP8-ckpt300, source fb92af176, eight DMD forwards, sparsity 0.8, tile 64, light H3 VAE. Each configuration uses one warmup and two timed prompts; construction and cold compilation are outside the headline median.

Geometry Timed runs Median Denoise / video decode medians
832×480, 124 frames (~5.17 s) 42.16 / 41.33 s 41.75 s 32.69 / 6.63 s
832×480, 243 frames (~10.13 s) 79.64 / 79.70 s 79.67 s 63.51 / 13.23 s

Earlier 16 and 12 GiB PyTorch allocator-cap trials completed at 104.03 and 107.27 s for the 243-frame 480p clip. These emulate capacity, not the throughput of smaller GPUs. The latest 7.25 GiB allocator-cap trial exceeded the strict 8 GiB total-device target (sampled NVML ~8.28 GiB); further tightening and the updated 768p benchmark are pending. The pod stopped accepting SSH connections before those results could be collected; the last verified 768p median remains 279.94 s. Current uncapped recipes use essentially all 24 GB of the board. A 32 GB system-RAM configuration has not been established.

Validation:

  • 120 distinct targeted CUDA/CPU checks passed after rebasing: pinned arenas and offload, encoder dequantization/streaming, tile-first and strided attention, VAE INT8 projections and fused epilogue, compile contracts, and stage behavior. The full pipeline CUDA-graph test is excluded for the eager recipe.
  • Scoped pre-commit checks passed, respecting the repository's intentional exclusions.
  • Both new 243-frame clips have identical decoded RGB and PCM audio hashes to the prior INT8 candidate, despite the memory changes and faster VAE decode.
  • INT8 attention is a numerical approximation; visual/audio quality against the original BF16 attention reference still needs broader review. Original attention remains the default. Actual 30-series hardware and V2 FP8 weights are not yet tested.

Exact commands, environment flags, checkpoint identity, historical measurements, limitations, and reproduction details are in scripts/benchmarks/minimax_h3_4090/README.md.

RetriggerConfidence Score: 4/5

The PR appears safe to merge, with a non-blocking benchmark reporting issue to address.

Fix All in CodexFindings

  1. P2 Failed samples report zero memory ▶
Fix with agent prompt
### Issue 1
scripts/benchmarks/minimax_h3_4090/bench_pod.py:48-49
If the cgroup memory files are unavailable or a read fails, the sampling thread stops, but the benchmark still records host-memory peaks as 0.0 GiB. That can make an unmeasured run look like a valid low-memory result. Mark the measurement unavailable or fail the run instead of publishing zero peaks.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.
Summary

This PR adds opt-in consumer-GPU attention, encoder-streaming, pinned-memory, and VAE INT8 optimizations for FastH3, alongside tests and RTX 4090 benchmark recipes. The default attention route remains unchanged.

Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[Text conditioning] --> B[Stream encoder layers]
  B --> C[DiT denoising]
  C --> D[Optional tile-first VSA]
  D --> E[Original or sm89 fine attention]
  E --> F[On-demand VAE decode]
  F --> G[Clip and benchmark measurements]
Loading

Reviews (1) · Last reviewed commit: "[docs]: record Track B PR and strict 8 G..."

aryan5v added 23 commits October 3, 2026 13:04
@coderabbitai

coderabbitai Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration
  • Configuration used: defaults
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: e19a92b2-c701-4548-bdba-ffcc76ff63ce

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@aryan5v
aryan5v marked this pull request as ready for review October 3, 2026 23:58
Comment on lines +48 to +49
return
self.stop.wait(0.1)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Failed samples report zero memory If the cgroup memory files are unavailable or a read fails, the sampling thread stops, but the benchmark still records host-memory peaks as 0.0 GiB. That can make an unmeasured run look like a valid low-memory result. Mark the measurement unavailable or fail the run instead of publishing zero peaks.

Prompt To Fix With AI
This is a comment left during a code review.
Path: scripts/benchmarks/minimax_h3_4090/bench_pod.py
Line: 48-49

Comment:
**Failed samples report zero memory** If the cgroup memory files are unavailable or a read fails, the sampling thread stops, but the benchmark still records host-memory peaks as 0.0 GiB. That can make an unmeasured run look like a valid low-memory result. Mark the measurement unavailable or fail the run instead of publishing zero peaks.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Codex Fix in Claude Code Fix in Cursor Fix in Conductor

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant