Conversation
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. Please check the settings in the CodeRabbit UI or the ⚙️ Run configuration
You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
| ids = a.prompts.split(",") if a.prompts else list(texts) | ||
| order = [ids[i % len(ids)] for i in range(a.warmup + a.timed)] | ||
| order = ids if a.once else [ids[i % len(ids)] for i in range(a.warmup + a.timed)] | ||
| warmup = 1 if a.once else a.warmup |
There was a problem hiding this comment.
First showcase render excluded With
--once, every selected prompt is rendered once, but the first is marked as a warmup. Its duration is left out of the median and mean; with one prompt, neither aggregate is reported. Mark showcase renders as non-warmup so their timing is included.
| warmup = 1 if a.once else a.warmup | |
| warmup = 0 if a.once else a.warmup |
Prompt To Fix With AI
This is a comment left during a code review.
Path: scripts/benchmarks/minimax_h3_4090/bench_pod.py
Line: 180
Comment:
**First showcase render excluded** With `--once`, every selected prompt is rendered once, but the first is marked as a warmup. Its duration is left out of the median and mean; with one prompt, neither aggregate is reported. Mark showcase renders as non-warmup so their timing is included.
```suggestion
warmup = 0 if a.once else a.warmup
```
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
| metadata = {'config_sha256': hashlib.sha256((args.model / 'transformer/config.json').read_bytes()).hexdigest(), | ||
| 'contract_sha256': hashlib.sha256((args.model / 'fastvideo_inference.json').read_bytes()).hexdigest(), | ||
| 'model_revision': revision_file.read_text().splitlines()[0] if revision_file.is_file() else None, | ||
| 'source_commit': args.source_commit, 'torch': str(torch.__version__), 'cuda': torch.version.cuda, | ||
| 'gpu': torch.cuda.get_device_name(0), | ||
| 'helper_sha256': hashlib.sha256(Path(__file__).read_bytes()).hexdigest(), | ||
| 'contract': contract, 'blocks': len(tables), 'timestep_keys': list(inputs), | ||
| 'table_sha256': hashlib.sha256(args.output.read_bytes()).hexdigest(), |
There was a problem hiding this comment.
Source weights lack hashes The table uses weights from checkpoint shards, but its provenance records only configuration, contract, and output hashes. If a local checkpoint has no revision metadata, different weights with the same configuration and contract cannot be distinguished from that record. Hash the source weight shards so a saved table can be matched to its checkpoint before reuse.
Prompt To Fix With AI
This is a comment left during a code review.
Path: scripts/benchmarks/minimax_h3_4090/precompute_v2_adaln.py
Line: 88-95
Comment:
**Source weights lack hashes** The table uses weights from checkpoint shards, but its provenance records only configuration, contract, and output hashes. If a local checkpoint has no revision metadata, different weights with the same configuration and contract cannot be distinguished from that record. Hash the source weight shards so a saved table can be matched to its checkpoint before reuse.
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.
Purpose
Layerwise offload ignored
offload.pin_cpu_memory=Falseand always allocated a private pinned copy. On RAM-limited RTX 4090 hosts, this prevented testing cached components without repeatedly loading the encoder, DiT, and VAEs. Honor the existing option for DiT and streamed H3 encoder layers while keeping pinned storage as the default. CPU-targeted checkpoint loading also staged tensors on the GPU before layer hooks and skipped V2 AdaLN weights could be applied; load those tensors on the requested CPU device instead.Stacked on
aryan5v:fasth3-rtx, the consumer launch branch behind hao-ai-lab#1919.Changes
--pageable-host,--once, warm-average reporting, and--seedoptions; preserve the prior seed default and document the launch protocol.Test Plan
Scoped
pre-commit run --filespassed for all changed paths, respecting configured exclusions.On the RTX 4090 pod, with single-rank distributed environment:
Test Results
26 passedin 10.46 seconds for offload/encoder/env tests. An additional loader suite passed all 12 tests, including CPU checkpoint placement, exact forward parity, source release, and FSDP quantization policy. Seven isolated CPU benchmark tests verify cgroup accounting, failure receipts, and exactly one gallery request without extra warmup clips. The latest mapped-encoder suite passed all 40 tests on a 4090: streamed-vs-resident exact forward parity with real safetensors mappings, NVFP4 loading/validation, custom-loader fallback, CPU target placement, unified-memory policies, and rejection of malformed padded-vocabulary checkpoints. The new mapped-storage and encoder placement checks require exact equality.Trim FP8, 832x480, 124 frames, seed 1234, 8 DMD forwards, sparsity 0.8, tile 64; one warmup and two timed samples per prompt. The cached pageable recipe also disables pinned VAE swaps and uses 34 resident DiT blocks. On this 31 GB host pod:
Cached runs peaked at 23.975 GiB total sampled GPU memory and 28.871 GiB pod-wide host usage, including file cache. Decoded RGB frames and PCM audio for both prompts exactly match the lazy recipe by SHA-256. This validates placement equivalence within the existing INT8 attention recipe; it is not a new BF16-reference quality claim. V2 completed three validation generations, with all 400 live worker timestep-input checks exactly matching the independently precomputed modulation inputs. The table builder passed every block/rung check on both 4090 pods (driver 570 and 580), producing identical SHA-256 table files. Both corgi clips are published at 24, 16, 12, and 8 GiB; independently rendered exports have identical SHA-256 values across tiers for each model. With the mapped encoder, 16 GiB Trim 480p medians are 70.79 s / 73.42 s (ceramics / harbor), down from 154.32 s / 140.385 s. The corrected allocator cap is 14.25 GiB and measured total GPU peak is 15.225 GiB. Both decoded RGB and PCM hashes exactly match the previous 16 GiB recipe for ceramics and harbor. 12 GiB Trim 480p medians are 74.705 s / 73.435 s, with an 11.707 GiB GPU peak. The full launch protocol is complete for both models at 24 and 12 GiB, at both 832x480 and 1344x768. V2 per-prompt medians are 51.665 s / 57.455 s at 24 GiB 480p, 153.74 s / 155.47 s at 24 GiB 768p, 92.075 s / 90.27 s at 12 GiB 480p, and 170.855 s / 170.535 s at 12 GiB 768p. The 16 GiB V2 480p medians are 80.665 s / 79.2 s. All values include text encoding, denoising, video/audio decode, and MP4 write. Capacity tiers are emulated on the 4090 and do not measure smaller-GPU throughput. The 16 GiB 768p protocol also completed: Trim medians are 137.64 s / 142.09 s and V2 medians are 153.35 s / 154.00 s, using VAE tile batch 14 for decode headroom. Required 24/16/12 GiB rows and all eight independently generated corgi exports are available on the launch data branch. The 8 GiB Trim 480p protocol completed with medians 82.25 s / 81.65 s and a sampled 7.752 GiB GPU peak. V2 480p completed with warm medians 92.97 s / 91.895 s, but its first warmup peaked at 8.088 GiB, so the recipe is marked
does_not_fitfor an 8 GiB target; its timed requests stayed at or below 7.758 GiB. Both 8 GiB 768p recipes failed on the attention workspace, including the 7.0 GiB allocator retry. Their measured peaks and errors are published without fabricated medians. All launch results and clips are pushed; both benchmark pods are idle. These timing-only recipe adjustments do not change the implementation in this PR.Checklist
pre-commit run --all-filesand fixed all issues (scoped checks passed)