Skip to content

[perf] Enable cached H3 generation on RAM-limited RTX hosts - #48

Open
aryan5v wants to merge 9 commits into
fasth3-rtxfrom
fasth3-rtx-4090-launch
Open

aryan5v wants to merge 9 commits into
fasth3-rtxfrom
fasth3-rtx-4090-launch

Conversation

@aryan5v

@aryan5v aryan5v commented Oct 5, 2026 •

Copy link
Copy Markdown
Owner

Purpose

Layerwise offload ignored offload.pin_cpu_memory=False and always allocated a private pinned copy. On RAM-limited RTX 4090 hosts, this prevented testing cached components without repeatedly loading the encoder, DiT, and VAEs. Honor the existing option for DiT and streamed H3 encoder layers while keeping pinned storage as the default. CPU-targeted checkpoint loading also staged tensors on the GPU before layer hooks and skipped V2 AdaLN weights could be applied; load those tensors on the requested CPU device instead.

Stacked on aryan5v:fasth3-rtx, the consumer launch branch behind hao-ai-lab#1919.

Changes

  • Retain existing pageable CPU storage when pinning is disabled. File-backed checkpoint tensors preserve their storage pointers and remain reclaimable by the OS.
  • Pass the engine option through both component loaders and H3 encoder preparation. In the pageable streamed path, retain compatible encoder checkpoint mappings and finalize NVFP4 layers on CPU; moving them to CUDA and back would recreate an anonymous host copy. Unknown custom loaders and dtype/shape conversions retain their existing behavior.
  • Add benchmark --pageable-host, --once, warm-average reporting, and --seed options; preserve the prior seed default and document the launch protocol.
  • Keep CPU-targeted transformer checkpoint reads on CPU, preventing V2 loading OOM before offload can take effect.
  • Add a reproducible fixed-ladder V2 modulation-table builder, checking every block/rung against the original module and recording checkpoint/config/table hashes.
  • Preserve cgroup v1/v2 host peaks and NVML GPU measurements, including failed generation receipts.
  • Cover repeated GPU forwards, mapped-storage retention, and both pinning modes for unquantized and serialized quantized encoder execution.

Test Plan

Scoped pre-commit run --files passed for all changed paths, respecting configured exclusions.

On the RTX 4090 pod, with single-rank distributed environment:

python -P -m pytest \
  fastvideo/tests/hooks/test_layerwise_offload.py \
  fastvideo/tests/hooks/test_pinned_memory.py \
  fastvideo/tests/encoders/test_minimax_h3_encoder_layerwise.py \
  fastvideo/tests/contract/test_env_policy.py -q

Test Results

26 passed in 10.46 seconds for offload/encoder/env tests. An additional loader suite passed all 12 tests, including CPU checkpoint placement, exact forward parity, source release, and FSDP quantization policy. Seven isolated CPU benchmark tests verify cgroup accounting, failure receipts, and exactly one gallery request without extra warmup clips. The latest mapped-encoder suite passed all 40 tests on a 4090: streamed-vs-resident exact forward parity with real safetensors mappings, NVFP4 loading/validation, custom-loader fallback, CPU target placement, unified-memory policies, and rejection of malformed padded-vocabulary checkpoints. The new mapped-storage and encoder placement checks require exact equality.

Trim FP8, 832x480, 124 frames, seed 1234, 8 DMD forwards, sparsity 0.8, tile 64; one warmup and two timed samples per prompt. The cached pageable recipe also disables pinned VAE swaps and uses 34 resident DiT blocks. On this 31 GB host pod:

Prompt Lazy median Cached pageable median
Ceramics 155.60 s 43.905 s
Harbor 157.505 s 43.945 s

Cached runs peaked at 23.975 GiB total sampled GPU memory and 28.871 GiB pod-wide host usage, including file cache. Decoded RGB frames and PCM audio for both prompts exactly match the lazy recipe by SHA-256. This validates placement equivalence within the existing INT8 attention recipe; it is not a new BF16-reference quality claim. V2 completed three validation generations, with all 400 live worker timestep-input checks exactly matching the independently precomputed modulation inputs. The table builder passed every block/rung check on both 4090 pods (driver 570 and 580), producing identical SHA-256 table files. Both corgi clips are published at 24, 16, 12, and 8 GiB; independently rendered exports have identical SHA-256 values across tiers for each model. With the mapped encoder, 16 GiB Trim 480p medians are 70.79 s / 73.42 s (ceramics / harbor), down from 154.32 s / 140.385 s. The corrected allocator cap is 14.25 GiB and measured total GPU peak is 15.225 GiB. Both decoded RGB and PCM hashes exactly match the previous 16 GiB recipe for ceramics and harbor. 12 GiB Trim 480p medians are 74.705 s / 73.435 s, with an 11.707 GiB GPU peak. The full launch protocol is complete for both models at 24 and 12 GiB, at both 832x480 and 1344x768. V2 per-prompt medians are 51.665 s / 57.455 s at 24 GiB 480p, 153.74 s / 155.47 s at 24 GiB 768p, 92.075 s / 90.27 s at 12 GiB 480p, and 170.855 s / 170.535 s at 12 GiB 768p. The 16 GiB V2 480p medians are 80.665 s / 79.2 s. All values include text encoding, denoising, video/audio decode, and MP4 write. Capacity tiers are emulated on the 4090 and do not measure smaller-GPU throughput. The 16 GiB 768p protocol also completed: Trim medians are 137.64 s / 142.09 s and V2 medians are 153.35 s / 154.00 s, using VAE tile batch 14 for decode headroom. Required 24/16/12 GiB rows and all eight independently generated corgi exports are available on the launch data branch. The 8 GiB Trim 480p protocol completed with medians 82.25 s / 81.65 s and a sampled 7.752 GiB GPU peak. V2 480p completed with warm medians 92.97 s / 91.895 s, but its first warmup peaked at 8.088 GiB, so the recipe is marked does_not_fit for an 8 GiB target; its timed requests stayed at or below 7.758 GiB. Both 8 GiB 768p recipes failed on the attention workspace, including the 7.0 GiB allocator retry. Their measured peaks and errors are published without fabricated medians. All launch results and clips are pushed; both benchmark pods are idle. These timing-only recipe adjustments do not change the implementation in this PR.

Checklist

  • I ran pre-commit run --all-files and fixed all issues (scoped checks passed)
  • I added or updated tests for my changes
  • I updated documentation if needed
  • I considered GPU memory impact of my changes
  • I verified SSIM regression tests pass (exact placement/output parity checked; full reference SSIM not rerun)

@coderabbitai

coderabbitai Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration
  • Configuration used: defaults
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 42226fe9-3952-4225-a057-4cf12dd826d5

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@aryan5v aryan5v changed the title [perf] Respect pageable CPU storage in layerwise offload [perf] Enable cached H3 generation on RAM-limited RTX hosts Oct 5, 2026
@aryan5v
aryan5v marked this pull request as ready for review October 5, 2026 21:53
ids = a.prompts.split(",") if a.prompts else list(texts)
order = [ids[i % len(ids)] for i in range(a.warmup + a.timed)]
order = ids if a.once else [ids[i % len(ids)] for i in range(a.warmup + a.timed)]
warmup = 1 if a.once else a.warmup

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 First showcase render excluded With --once, every selected prompt is rendered once, but the first is marked as a warmup. Its duration is left out of the median and mean; with one prompt, neither aggregate is reported. Mark showcase renders as non-warmup so their timing is included.

Suggested change
warmup = 1 if a.once else a.warmup
warmup = 0 if a.once else a.warmup
Prompt To Fix With AI
This is a comment left during a code review.
Path: scripts/benchmarks/minimax_h3_4090/bench_pod.py
Line: 180

Comment:
**First showcase render excluded** With `--once`, every selected prompt is rendered once, but the first is marked as a warmup. Its duration is left out of the median and mean; with one prompt, neither aggregate is reported. Mark showcase renders as non-warmup so their timing is included.

```suggestion
    warmup = 0 if a.once else a.warmup
```

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Fix in Codex Fix in Claude Code Fix in Cursor Fix in Conductor

Comment on lines +88 to +95
metadata = {'config_sha256': hashlib.sha256((args.model / 'transformer/config.json').read_bytes()).hexdigest(),
'contract_sha256': hashlib.sha256((args.model / 'fastvideo_inference.json').read_bytes()).hexdigest(),
'model_revision': revision_file.read_text().splitlines()[0] if revision_file.is_file() else None,
'source_commit': args.source_commit, 'torch': str(torch.__version__), 'cuda': torch.version.cuda,
'gpu': torch.cuda.get_device_name(0),
'helper_sha256': hashlib.sha256(Path(__file__).read_bytes()).hexdigest(),
'contract': contract, 'blocks': len(tables), 'timestep_keys': list(inputs),
'table_sha256': hashlib.sha256(args.output.read_bytes()).hexdigest(),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Source weights lack hashes The table uses weights from checkpoint shards, but its provenance records only configuration, contract, and output hashes. If a local checkpoint has no revision metadata, different weights with the same configuration and contract cannot be distinguished from that record. Hash the source weight shards so a saved table can be matched to its checkpoint before reuse.

Prompt To Fix With AI
This is a comment left during a code review.
Path: scripts/benchmarks/minimax_h3_4090/precompute_v2_adaln.py
Line: 88-95

Comment:
**Source weights lack hashes** The table uses weights from checkpoint shards, but its provenance records only configuration, contract, and output hashes. If a local checkpoint has no revision metadata, different weights with the same configuration and contract cannot be distinguished from that record. Hash the source weight shards so a saved table can be matched to its checkpoint before reuse.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Codex Fix in Claude Code Fix in Cursor Fix in Conductor

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant