Skip to content

[feat]: CompactH3 recovery, DMD2, and quantization - #41

Open
aryan5v wants to merge 5 commits into
mainfrom
aryan/compacth3-consolidated
Open

aryan5v wants to merge 5 commits into
mainfrom
aryan/compacth3-consolidated

Conversation

@aryan5v

@aryan5v aryan5v commented Sep 17, 2026 •

Copy link
Copy Markdown
Owner

Summary

  • consolidate the activation-guided 42-block and 34-block pruning/recovery workflow on the latest upstream FastVideo main
  • integrate the corrected joint audio-video four-call DMD2 implementation, export path, resume tooling, checkpoint grading, and validation media logging
  • add INT8, NVFP4, W4A16, and MXFP8 loader/export support plus the NVFP4 QAD recipe and checkpoint gates
  • retain current upstream MiniMax-H3 inference features, including FastH3 V2 schedules, LoRA, TAEH3, and multi-device loading
  • exclude checkpoints, generated media, W&B outputs, run logs, research drafts, and internal handoff documents

Lineage represented

The promoted 42-block path uses activation-guided block selection, recovery through the selected checkpoint-750 parent, and the corrected DMD2 run whose selected checkpoint is 1400. The 34-block model is derived from the recovered 42-block model and currently includes recovery infrastructure only; it has not yet been promoted through DMD2.

Validation performed

  • git diff --cached --check
  • Python syntax compilation for every changed Python file
  • bash -n for every changed shell and Slurm launcher
  • secret-pattern and merge-marker scans

Remaining validation

This is intentionally a draft. Targeted CPU/GPU tests, cluster path parameterization, one clean recovery/DMD2 launch smoke, and final removal of any redundant historical sprint scripts remain before promotion.

RetriggerConfidence Score: 0/5

This PR is not safe to merge until the broken launch paths, distributed inference-checkpoint deadlock, and pruning-score normalization are corrected.

Fix All in CodexFindings

  1. P1 Sweep scripts fail at startup ▶
  2. P1 Default config is missing ▶
  3. P1 Stage launcher path is missing ▶
  4. P1 Checkpoint failures can deadlock ▶
  5. P1 Shard count biases pruning ▶
  6. P2 Documented cadence default is stale ▶
  7. P2 Logger bypasses shared initializer ▶
Fix with agent prompt
### Issue 1
scripts/compacth3/sweep/measure_rank.py:13
This module evaluates `os.environ` without importing `os` and also references the undefined Python global `SPRINT_ROOT`. `sweep_driver.py` and `emit_folds.py` contain the same undefined global. Because the launchers execute these files directly without injecting that name, the sweep and measurement commands raise `NameError` before argument parsing and never start.

### Issue 2
examples/distill/MiniMax-H3/distill_dmd.sh:17
The default points to `dmd2_sp1_fsdp40_vidprom_v6.yaml`, but that file is not present in the added MiniMax-H3 configuration directory. Running this launcher without manually setting `CONFIG` therefore passes a nonexistent file to `examples/train/run.sh` and aborts before training begins.

### Issue 3
scripts/fasth3_sprint/prepare_h3_stage34.py:33
This unconditionally reads `scripts/fasth3_sprint/slurm_h3_base42_recovery.sbatch`, which is absent from this revision; only the differently named `slurm_h3_base42_recovery500.sbatch` is checked in. Every stage-34 preparation therefore raises `FileNotFoundError` after already creating a partial output directory.

### Issue 4
fastvideo/train/utils/checkpoint.py:501-508
If rank-zero staging, `dcp.save`, or publishing `export-status.json` fails before status propagation, surviving ranks either wait at a barrier or poll for this file forever. This loop has no timeout, so a one-rank filesystem or DCP failure can indefinitely block the entire distributed training job.

### Issue 5
scripts/fasth3_sprint/select_h3_block_map.py:32
The producer stores per-category sums and emits zero accumulators when a category is absent from a shard, but this code divides those sums by the total number of shard files. Sparse categories such as the eight-example `multiple_shots` group are therefore diluted relative to other categories. The resulting importance scores can rank blocks incorrectly and affect the model's permanent pruning map; the canonical aggregator instead divides by each category's actual sample count.

### Issue 6
fastvideo/train/methods/distribution_matching/dmd2.py:763-764
An omitted `generator_update_interval` now resolves to five, while the public training documentation and example configuration still state that the default is one. Users relying on that documented default will unknowingly run four critic-only iterations between student updates, so the documentation must be updated with this behavior change.

### Issue 7
fastvideo/layers/quantization/int8_affine_config.py:22
This new core-package module calls `logging.getLogger(__name__)` directly. The repository directive requires FastVideo package code to use `from fastvideo.logger import init_logger` and `logger = init_logger(__name__)` so formatting and distributed logging behavior remain consistent. This repository requirement must be satisfied before merging.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.
Summary

This PR introduces the CompactH3 recovery and pruning workflow, joint audio-video DMD2 training and export infrastructure, native-shape preprocessing, validation media logging, and several low-bit quantization formats. The review found startup failures in newly added launch tooling, a distributed checkpoint deadlock path, and biased pruning-score aggregation.

  • Adds INT8 affine, NVFP4/QAT, and W4A16 quantization implementations and sidecar formats.
  • Adds four-call MiniMax-H3 DMD2 training, checkpoint export, resume, and validation workflows.
  • Adds activation-guided 42/34-block recovery, evaluation, and AdaLN rank-sweep tooling.
  • Adds native-shape T2VA preprocessing, bucketing, immutable dataset receipts, and validation support.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[Native T2VA / text data] --> B[Exact-shape dataloader]
  B --> C[CompactH3 recovery]
  C --> D[42-block recovered parent]
  D --> E[Four-call DMD2]
  E --> F[Training DCP checkpoints]
  F --> G[Inference checkpoint export]
  G --> H[Validation and checkpoint grading]
  D --> I[34-block recovery preparation]
  E --> J[INT8 / NVFP4 / W4A16 export]
Loading

Reviews (1) · Last reviewed commit: "[feat]: add AdaLN rank sweep and rank-sp..."

@coderabbitai

coderabbitai Bot commented Sep 17, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Too many files!

This PR contains 211 files, which is 111 over the limit of 100.

To get a review, reduce the PR to 100 files or fewer by splitting it into smaller PRs or changing its base branch.

Upgrade to a paid plan to raise the limit.

This review couldn't start because sufficient usage credits or metered capacity aren't available. Add credits or update usage-based reviews in the billing tab, then retry.

Check out review usage here.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: 86379f75-2291-4c73-9750-fed3c84d3538

📥 Commits

Reviewing files that changed from the base of the PR and between c4824c7 and 8879dfa.

📒 Files selected for processing (211)
  • docs/quantization/h3_int8_affine.md
  • docs/quantization/h3_nvfp4.md
  • docs/quantization/h3_w4a16.md
  • docs/quantization/loader_quant_params.md
  • docs/training/attn_qat.md
  • examples/compacth3/README.md
  • examples/distill/MiniMax-H3/distill_dmd.sh
  • examples/inference/minimax_h3/README.md
  • examples/inference/minimax_h3/h3_vsa_dmd.py
  • examples/train/configs/compacth3/release14b_recovery_wandb.yaml
  • examples/train/configs/compacth3/release14b_validation_five.json
  • examples/train/configs/distribution_matching/minimax_h3/README.md
  • examples/train/configs/distribution_matching/minimax_h3/qad_nvfp4_4call.yaml
  • examples/train/configs/distribution_matching/minimax_h3/qad_nvfp4_4call_r16.yaml
  • examples/train/configs/distribution_matching/minimax_h3/qad_nvfp4_4call_r768.yaml
  • examples/train/configs/distribution_matching/minimax_h3/release20b_dmd2_v12_dense.yaml
  • examples/train/configs/distribution_matching/minimax_h3/release20b_validation_five.json
  • examples/train/configs/fasth3_14b_recovery.yaml
  • examples/train/configs/fasth3_base42_recovery.yaml
  • examples/train/configs/fasth3_detail_band_recovery.yaml
  • examples/train/configs/fasth3_detail_band_recovery34.yaml
  • examples/train/configs/fasth3_release_long_recovery.yaml
  • examples/train/configs/overfit_minimax_h3_t2va.yaml
  • examples/training/fasth3_14b_2step_qad/comparison_five_prompts.json
  • examples/training/fasth3_14b_2step_qad/quick_gate_prompts.json
  • examples/training/fasth3_14b_2step_qad/release_sentinel_24.json
  • examples/training/fasth3_14b_2step_qad/rescue_gate_prompts.json
  • examples/training/fasth3_14b_2step_qad/showcase_prompt.json
  • fastvideo/attention/utils/flash_attn_cute.py
  • fastvideo/configs/models/dits/minimax_h3.py
  • fastvideo/dataset/parquet_dataset_map_style.py
  • fastvideo/dataset/shape_bucket.py
  • fastvideo/dataset/validation_dataset.py
  • fastvideo/fastvideo_args.py
  • fastvideo/layers/quantization/__init__.py
  • fastvideo/layers/quantization/int8_affine_config.py
  • fastvideo/layers/quantization/nvfp4_config.py
  • fastvideo/layers/quantization/nvfp4_qat_config.py
  • fastvideo/layers/quantization/w4a16_config.py
  • fastvideo/models/dits/minimax_h3.py
  • fastvideo/models/loader/component_loader.py
  • fastvideo/models/loader/fsdp_load.py
  • fastvideo/models/loader/shard_cache.py
  • fastvideo/models/schedulers/scheduling_minimax_h3.py
  • fastvideo/pipelines/basic/minimax_h3/stages/minimax_h3_decoding.py
  • fastvideo/pipelines/basic/minimax_h3/stages/minimax_h3_denoising.py
  • fastvideo/pipelines/pipeline_batch_info.py
  • fastvideo/pipelines/preprocess/preprocess_minimax_h3_overfit.py
  • fastvideo/pipelines/preprocess/preprocess_minimax_h3_text_only.py
  • fastvideo/tests/attention/test_compile_policy.py
  • fastvideo/tests/attention/test_flash_attn_cute_custom_op.py
  • fastvideo/tests/attention/test_vsa_h3_inference_metadata_parity.py
  • fastvideo/tests/attention/test_vsa_h3_metadata.py
  • fastvideo/tests/dataset/test_exact_shape_bucket_sampler.py
  • fastvideo/tests/dataset/test_parquet_dataset_map_style.py
  • fastvideo/tests/dataset/test_validation_dataset.py
  • fastvideo/tests/loader/test_shard_cache.py
  • fastvideo/tests/ops/quantization/test_allowlist_mirror.py
  • fastvideo/tests/ops/quantization/test_int8_affine_config.py
  • fastvideo/tests/ops/quantization/test_int8_dispatch.py
  • fastvideo/tests/ops/quantization/test_nvfp4_h3_prefixes.py
  • fastvideo/tests/ops/quantization/test_nvfp4_sidecar.py
  • fastvideo/tests/ops/quantization/test_quant_param_allowlist.py
  • fastvideo/tests/ops/quantization/test_quant_sidecar_roundtrip.py
  • fastvideo/tests/ops/quantization/test_w4a16_config.py
  • fastvideo/tests/train/callbacks/test_callback.py
  • fastvideo/tests/train/callbacks/test_ema.py
  • fastvideo/tests/train/callbacks/test_latent_vis_shape_context.py
  • fastvideo/tests/train/callbacks/test_validation.py
  • fastvideo/tests/train/callbacks/test_validation_sampling_contract.py
  • fastvideo/tests/train/fixtures/minimax_h3_dmd2_min.yaml
  • fastvideo/tests/train/methods/test_dmd2_data_forcing.py
  • fastvideo/tests/train/methods/test_dmd2_fake_score_loss_space.py
  • fastvideo/tests/train/methods/test_dmd2_fastgen_parity.py
  • fastvideo/tests/train/methods/test_dmd2_rollout_carry.py
  • fastvideo/tests/train/methods/test_dmd2_timestep_bounds.py
  • fastvideo/tests/train/methods/test_dmd2_vsd_normalizer.py
  • fastvideo/tests/train/methods/test_minimax_h3_dmd2.py
  • fastvideo/tests/train/methods/test_minimax_h3_finetune.py
  • fastvideo/tests/train/trainer/test_validation.py
  • fastvideo/tests/train/utils/test_checkpoint.py
  • fastvideo/tests/train/utils/test_config.py
  • fastvideo/tests/train/utils/test_inference_checkpoint.py
  • fastvideo/tests/train/utils/test_inference_checkpoint_distributed.py
  • fastvideo/tests/train/utils/test_moduleloader_attention_backend.py
  • fastvideo/tests/train/utils/test_torch_compile.py
  • fastvideo/tests/train/utils/test_tracking.py
  • fastvideo/train/attn_qat/README.md
  • fastvideo/train/callbacks/callback.py
  • fastvideo/train/callbacks/ema.py
  • fastvideo/train/callbacks/latent_vis.py
  • fastvideo/train/callbacks/validation.py
  • fastvideo/train/entrypoint/dcp_to_diffusers.py
  • fastvideo/train/entrypoint/train.py
  • fastvideo/train/methods/base.py
  • fastvideo/train/methods/distribution_matching/dmd2.py
  • fastvideo/train/models/base.py
  • fastvideo/train/models/minimax_h3/__init__.py
  • fastvideo/train/models/minimax_h3/minimax_h3.py
  • fastvideo/train/models/minimax_h3/minimax_h3_dmd.py
  • fastvideo/train/trainer.py
  • fastvideo/train/utils/checkpoint.py
  • fastvideo/train/utils/config.py
  • fastvideo/train/utils/dataloader.py
  • fastvideo/train/utils/inference_checkpoint.py
  • fastvideo/train/utils/moduleloader.py
  • fastvideo/train/utils/optimizer.py
  • fastvideo/train/utils/tracking.py
  • fastvideo/train/utils/training_config.py
  • mkdocs.yml
  • scripts/checkpoint_conversion/convert_minimax_h3_adaln_rank.py
  • scripts/checkpoint_conversion/export_h3_dmd2_student.py
  • scripts/compacth3/analysis/adaln/analyze_adaln_rank.py
  • scripts/compacth3/analysis/adaln/compare_parent_dmd2.py
  • scripts/compacth3/analysis/adaln/run_adaln_rank.sh
  • scripts/compacth3/analysis/adaln/summarize_adaln_rank.py
  • scripts/compacth3/analysis/adaln_lowrank.py
  • scripts/compacth3/analysis/checkpoint_sweep_metrics.py
  • scripts/compacth3/eval_corrected_dmd_inference_exports.sbatch
  • scripts/compacth3/eval_dmd_export_1400.sh
  • scripts/compacth3/grade_dmd2_all_checkpoints.sbatch
  • scripts/compacth3/hardmotion_set.json
  • scripts/compacth3/qad/gen_qad_configs.py
  • scripts/compacth3/qad/qad_checkpoint_gate.py
  • scripts/compacth3/qad/setup_qad.py
  • scripts/compacth3/qad/tune_qad_cadence.py
  • scripts/compacth3/quantization/export_lane_int8.sh
  • scripts/compacth3/quantization/export_lane_nvfp4.sh
  • scripts/compacth3/quantization/export_lane_w4a16.sh
  • scripts/compacth3/quantization/export_quant_dit.py
  • scripts/compacth3/quantization/run_export_nvfp4.sh
  • scripts/compacth3/resume_release20b_dmd2_paired_generic.sh
  • scripts/compacth3/run_eval_lane.sh
  • scripts/compacth3/submit_release14b_hardened.sbatch
  • scripts/compacth3/sweep.sh
  • scripts/compacth3/sweep/adaln_rank_patch.py
  • scripts/compacth3/sweep/analysis.sbatch
  • scripts/compacth3/sweep/check_identity_gate.py
  • scripts/compacth3/sweep/compare_ranks.py
  • scripts/compacth3/sweep/emit_folds.py
  • scripts/compacth3/sweep/measure_rank.py
  • scripts/compacth3/sweep/run_measure.sh
  • scripts/compacth3/sweep/run_precheck.sh
  • scripts/compacth3/sweep/run_smoke.sh
  • scripts/compacth3/sweep/run_sweep.sh
  • scripts/compacth3/sweep/sweep.sbatch
  • scripts/compacth3/sweep/sweep_driver.py
  • scripts/compacth3/sweep/sweep_env.sh
  • scripts/compacth3/sweep_prompts.json
  • scripts/fasth3_sprint/admit_h3_fold_for_recovery.py
  • scripts/fasth3_sprint/aggregate_minimax_h3_block_scores.py
  • scripts/fasth3_sprint/assemble_h3_stage.py
  • scripts/fasth3_sprint/audio_fidelity_gate.py
  • scripts/fasth3_sprint/audit_recovery_weights.py
  • scripts/fasth3_sprint/build_h3_audio_stratified_prompt_index.py
  • scripts/fasth3_sprint/build_h3_recovery_config.py
  • scripts/fasth3_sprint/check_h3_base_gate.py
  • scripts/fasth3_sprint/extract_review_assets.py
  • scripts/fasth3_sprint/h3_serve_eval_prompts.json
  • scripts/fasth3_sprint/h3_serve_eval_prompts_cool_480p.json
  • scripts/fasth3_sprint/h3_serve_eval_prompts_gateway_fashion_480p.json
  • scripts/fasth3_sprint/h3_stage_promotion.py
  • scripts/fasth3_sprint/h3_taeh3_kitchen_prompt.json
  • scripts/fasth3_sprint/prepare_h3_prompt_pool.py
  • scripts/fasth3_sprint/prepare_h3_stage34.py
  • scripts/fasth3_sprint/run_baseline_matrix.py
  • scripts/fasth3_sprint/run_recovery_comparison.sh
  • scripts/fasth3_sprint/run_recovery_diagnostics.sh
  • scripts/fasth3_sprint/score_minimax_h3_blocks.py
  • scripts/fasth3_sprint/seams_for_block_map.py
  • scripts/fasth3_sprint/select_h3_block_map.py
  • scripts/fasth3_sprint/slurm_block_score.sbatch
  • scripts/fasth3_sprint/slurm_block_score_aggregate.sbatch
  • scripts/fasth3_sprint/slurm_h3_34_latest_five.sbatch
  • scripts/fasth3_sprint/slurm_h3_activation34_from42_recovery500.sbatch
  • scripts/fasth3_sprint/slurm_h3_activation42_five.sbatch
  • scripts/fasth3_sprint/slurm_h3_base42_recovery500.sbatch
  • scripts/fasth3_sprint/slurm_h3_release_eval.sbatch
  • scripts/fasth3_sprint/slurm_h3_release_long_phase.sbatch
  • scripts/fasth3_sprint/slurm_prune_candidate.sbatch
  • scripts/fasth3_sprint/slurm_tests.sbatch
  • scripts/fasth3_sprint/split_h3_prompt_pool.py
  • scripts/fasth3_sprint/submit_block_score.sh
  • scripts/fasth3_sprint/submit_pruned_candidates.sh
  • scripts/fasth3_sprint/validate_h3_candidate.py
  • scripts/fasth3_sprint/verify_media.py
  • scripts/fasth3_sprint/verify_recovery_export_prediction_parity.py
  • scripts/fasth3_sprint/verify_speech_asr.py
  • scripts/preprocess/minimax_h3_native_t2va/README.md
  • scripts/preprocess/minimax_h3_native_t2va/derive_filtered_dataset.py
  • scripts/preprocess/minimax_h3_native_t2va/derive_filtered_dataset_1tray.sbatch
  • scripts/preprocess/minimax_h3_native_t2va/encode_native_t2va_1tray.sbatch
  • scripts/preprocess/minimax_h3_native_t2va/encode_worker.py
  • scripts/preprocess/minimax_h3_native_t2va/finalize_dataset.py
  • scripts/preprocess/minimax_h3_native_t2va/finalize_extension_1tray.sbatch
  • scripts/preprocess/minimax_h3_native_t2va/freeze_sources.py
  • scripts/preprocess/minimax_h3_native_t2va/prepare_extension_1tray.sbatch
  • scripts/preprocess/minimax_h3_native_t2va/probe_native_t2va.sbatch
  • scripts/preprocess/minimax_h3_native_t2va/schema_dry_run.py
  • scripts/preprocess/minimax_h3_native_t2va/seed_encoded_chunks.py
  • scripts/preprocess/minimax_h3_native_t2va/v10_sources.json
  • scripts/preprocess/minimax_h3_native_t2va/v10_sources_v2.json
  • scripts/preprocess/minimax_h3_native_t2va/validate_heldout_media.py
  • scripts/run_release20b_dmd2_v12_16gpu.sh
  • scripts/submit_release20b_dmd2_v12_corrected_16gpu.sbatch
  • scripts/submit_release20b_dmd2_v12_corrected_32gpu.sbatch
  • scripts/train/mfu_calc_minimax_h3.py
  • tests/local_tests/minimax_h3/README.md
  • tests/local_tests/models/test_fsdp_load_mixed_dtype.py
  • tests/local_tests/preprocess/test_minimax_h3_native_t2va.py
  • tests/local_tests/preprocess/test_minimax_h3_native_t2va_filtered.py

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Remove pstack-style intros, broken handoff links, and stale README
references. No functional code changes.

Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
@cursor cursor Bot changed the title [feat] Consolidate CompactH3 recovery, DMD2, and quantization [feat]: CompactH3 recovery, DMD2, and quantization Sep 17, 2026
cursoragent and others added 3 commits September 17, 2026 21:22
Strip narrative comments, section banners, and YAML header prose across
the PR scope. Shorten analysis script module docstrings. No logic or CI
changes.

Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
Collapse quantization docs to operational minimums, shorten quant module
docstrings, and trim CompactH3 index. No logic or CI changes.

Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
Adds the centered-affine rank compression and behavioral sweep that were
missing from the consolidation: fold emission, the runtime AdaLN patch, the
hard-motion metric set, rank comparison, and the identity gate.

Adds hard-motion eval set (14 cases) and the r768/r16 QAD configs.

Cluster roots are parameterized (COMPACTH3_ROOT, COMPACTH3_EVAL_ROOT,
COMPACTH3_CODE_ROOT) so the scripts run outside the original tree.
@aryan5v
aryan5v marked this pull request as ready for review September 19, 2026 13:05
@aryan5v

aryan5v commented Sep 19, 2026

Copy link
Copy Markdown
Owner Author

@greptileai

import numpy as np
import cv2

EVAL_DIR = Path(os.environ.get("COMPACTH3_EVAL_ROOT", str(Path(SPRINT_ROOT).parent / "fasth3-eval")))

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Sweep scripts fail at startup

This module evaluates os.environ without importing os and also references the undefined Python global SPRINT_ROOT. sweep_driver.py and emit_folds.py contain the same undefined global. Because the launchers execute these files directly without injecting that name, the sweep and measurement commands raise NameError before argument parsing and never start.

Prompt To Fix With AI
This is a comment left during a code review.
Path: scripts/compacth3/sweep/measure_rank.py
Line: 13

Comment:
**Sweep scripts fail at startup**

This module evaluates `os.environ` without importing `os` and also references the undefined Python global `SPRINT_ROOT`. `sweep_driver.py` and `emit_folds.py` contain the same undefined global. Because the launchers execute these files directly without injecting that name, the sweep and measurement commands raise `NameError` before argument parsing and never start.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Codex Fix in Claude Code Fix in Cursor Fix in Conductor

SP_SIZE="${SP_SIZE:-1}"
HSDP_REPLICATE="${HSDP_REPLICATE:-1}"
HSDP_SHARD="${HSDP_SHARD:-${WORLD_SIZE}}"
CONFIG="${CONFIG:-examples/train/configs/distribution_matching/minimax_h3/dmd2_sp1_fsdp40_vidprom_v6.yaml}"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Default config is missing

The default points to dmd2_sp1_fsdp40_vidprom_v6.yaml, but that file is not present in the added MiniMax-H3 configuration directory. Running this launcher without manually setting CONFIG therefore passes a nonexistent file to examples/train/run.sh and aborts before training begins.

Prompt To Fix With AI
This is a comment left during a code review.
Path: examples/distill/MiniMax-H3/distill_dmd.sh
Line: 17

Comment:
**Default config is missing**

The default points to `dmd2_sp1_fsdp40_vidprom_v6.yaml`, but that file is not present in the added MiniMax-H3 configuration directory. Running this launcher without manually setting `CONFIG` therefore passes a nonexistent file to `examples/train/run.sh` and aborts before training begins.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Codex Fix in Claude Code Fix in Cursor Fix in Conductor

cfg['method']['feature_local_block_indices'] = [i for i in range(1, len(original)) if original[i] > original[i-1]+1]
config_path = a.output.resolve() / 'stage34.yaml'
config_path.write_text(yaml.safe_dump(cfg, sort_keys=False))
script = (root / 'scripts/fasth3_sprint/slurm_h3_base42_recovery.sbatch').read_text()

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Stage launcher path is missing

This unconditionally reads scripts/fasth3_sprint/slurm_h3_base42_recovery.sbatch, which is absent from this revision; only the differently named slurm_h3_base42_recovery500.sbatch is checked in. Every stage-34 preparation therefore raises FileNotFoundError after already creating a partial output directory.

Prompt To Fix With AI
This is a comment left during a code review.
Path: scripts/fasth3_sprint/prepare_h3_stage34.py
Line: 33

Comment:
**Stage launcher path is missing**

This unconditionally reads `scripts/fasth3_sprint/slurm_h3_base42_recovery.sbatch`, which is absent from this revision; only the differently named `slurm_h3_base42_recovery500.sbatch` is checked in. Every stage-34 preparation therefore raises `FileNotFoundError` after already creating a partial output directory.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Codex Fix in Claude Code Fix in Cursor Fix in Conductor

Comment on lines +501 to +508
else:
last_log = time.monotonic()
while not export_status_path.is_file():
time.sleep(2.0)
now = time.monotonic()
if now - last_log >= 60.0:
logger.info("Waiting for rank-zero inference export at step %s", step)
last_log = now

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Checkpoint failures can deadlock

If rank-zero staging, dcp.save, or publishing export-status.json fails before status propagation, surviving ranks either wait at a barrier or poll for this file forever. This loop has no timeout, so a one-rank filesystem or DCP failure can indefinitely block the entire distributed training job.

Prompt To Fix With AI
This is a comment left during a code review.
Path: fastvideo/train/utils/checkpoint.py
Line: 501-508

Comment:
**Checkpoint failures can deadlock**

If rank-zero staging, `dcp.save`, or publishing `export-status.json` fails before status propagation, surviving ranks either wait at a barrier or poll for this file forever. This loop has no timeout, so a one-rank filesystem or DCP failure can indefinitely block the entire distributed training job.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Codex Fix in Claude Code Fix in Cursor Fix in Conductor

for dst, src in zip(acc, rows, strict=True):
for metric in METRICS:
dst[metric] += float(src[metric])
return {cat: [{m: v / count for m, v in row.items()} for row in rows] for cat, rows in totals.items()}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Shard count biases pruning

The producer stores per-category sums and emits zero accumulators when a category is absent from a shard, but this code divides those sums by the total number of shard files. Sparse categories such as the eight-example multiple_shots group are therefore diluted relative to other categories. The resulting importance scores can rank blocks incorrectly and affect the model's permanent pruning map; the canonical aggregator instead divides by each category's actual sample count.

Prompt To Fix With AI
This is a comment left during a code review.
Path: scripts/fasth3_sprint/select_h3_block_map.py
Line: 32

Comment:
**Shard count biases pruning**

The producer stores per-category sums and emits zero accumulators when a category is absent from a shard, but this code divides those sums by the total number of shard files. Sparse categories such as the eight-example `multiple_shots` group are therefore diluted relative to other categories. The resulting importance scores can rank blocks incorrectly and affect the model's permanent pruning map; the canonical aggregator instead divides by each category's actual sample count.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Codex Fix in Claude Code Fix in Cursor Fix in Conductor

Comment on lines 763 to +764
if interval is None:
interval = 1
interval = 5

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Documented cadence default is stale

An omitted generator_update_interval now resolves to five, while the public training documentation and example configuration still state that the default is one. Users relying on that documented default will unknowingly run four critic-only iterations between student updates, so the documentation must be updated with this behavior change.

Prompt To Fix With AI
This is a comment left during a code review.
Path: fastvideo/train/methods/distribution_matching/dmd2.py
Line: 763-764

Comment:
**Documented cadence default is stale**

An omitted `generator_update_interval` now resolves to five, while the public training documentation and example configuration still state that the default is one. Users relying on that documented default will unknowingly run four critic-only iterations between student updates, so the documentation must be updated with this behavior change.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Codex Fix in Claude Code Fix in Cursor Fix in Conductor

)
from fastvideo.models.utils import set_weight_attrs

logger = logging.getLogger(__name__)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Logger bypasses shared initializer

This new core-package module calls logging.getLogger(__name__) directly. The repository directive requires FastVideo package code to use from fastvideo.logger import init_logger and logger = init_logger(__name__) so formatting and distributed logging behavior remain consistent. This repository requirement must be satisfied before merging.

Context Used: fastvideo/AGENTS.md (source)

Prompt To Fix With AI
This is a comment left during a code review.
Path: fastvideo/layers/quantization/int8_affine_config.py
Line: 22

Comment:
**Logger bypasses shared initializer**

This new core-package module calls `logging.getLogger(__name__)` directly. The repository directive requires FastVideo package code to use `from fastvideo.logger import init_logger` and `logger = init_logger(__name__)` so formatting and distributed logging behavior remain consistent. This repository requirement must be satisfied before merging.

**Context Used:** fastvideo/AGENTS.md ([source](https://github.com/aryan5v/fastvideo/blob/main/fastvideo/AGENTS.md))

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Fix in Codex Fix in Claude Code Fix in Cursor Fix in Conductor

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants