Skip to content

Community Pipeline: Add HRDiT (training-free high-resolution FLUX.1-dev) - #25

Closed
smellslikeml wants to merge 12 commits into
mainfrom
hrdit-training-free-high-resolution-image-generation-with-of
Closed

smellslikeml wants to merge 12 commits into
mainfrom
hrdit-training-free-high-resolution-image-generation-with-of

Conversation

@smellslikeml

Copy link
Copy Markdown
Owner

Draft for internal review/coordination on this fork before an upstream PR to huggingface/diffusers. Let's track open items via issues on this repo.

What

Adds a community pipeline examples/community/pipeline_flux_hrdit.py implementing HRDiT (arXiv:2608.07003, official repo) — training-free high-resolution (up to 4096×4096) text-to-image on off-the-shelf FLUX.1-dev. No fine-tuning, no new weights, no src/diffusers changes.

Validation (FLUX.1-dev · A100-80GB · torch 2.11)

HRDiT high-res

Both 2048² and 4096² are coherent — no tiling, no washout. A 4096² generation runs in ~2 min at ~26 GB peak.

Runnable notebook: https://colab.research.google.com/drive/19IS19j2BvQm4Qnf07v-FDYiiqYZfD7Ja?usp=sharing

Method (on top of the stock FluxPipeline denoise loop)

  • NTK-aware RoPE scaling — the primary high-res mechanism; per-stage scaling of the RoPE base brings out-of-range positions back into the trained band, applied on every upscale step.
  • Spatial Position Alignment (SPA) — leading-steps-only nudge; a monotonic bundle coarsening of position ids (no wrapping → no periodic tiling), averaged over sliding variants inside attention.
  • Structure-guided progressive ladder (1024 → 2048 → 4096) — each stage decodes / upscales / re-encodes the previous latent as a structural prior, then injects its low-frequency band every step (Butterworth FFT split, weight alpha) with a velocity-momentum term (beta). This is what keeps the highest stage from drifting to a washed-out mean:

Structure guidance ablation

Deferred, documented in-file (not required for the results above): HAP head-scope attention pruning, swin_pachify, and DWT (as opposed to FFT) guidance.

Files

  • examples/community/pipeline_flux_hrdit.py — the pipeline
  • examples/community/README.md — usage entry
  • tests/others/test_community_pipeline_flux_hrdit.py — CPU-only unit tests (SPA variants, NTK RoPE, structure-guidance helpers; 13 passing)
  • benchmarks/benchmarking_flux_hrdit.py — naive FluxPipeline vs HRDiT timing / peak memory

Open items to coordinate (issues)

  • Multi-prompt / multi-seed sweep at 4096² — validated on one prompt/seed so far
  • Decide whether to port HAP (speed) before the upstream PR
  • Host the figures as PR attachments (currently on the hrdit-validation-assets branch of this fork) before upstream submission

Disclosure

Training-free / adds no weights. Initially drafted by Outrider (remyx) and then corrected and completed with AI assistance (Claude) against the reference implementation, with each iteration validated on GPU before merge.

github-actions Bot and others added 11 commits August 14, 2026 16:30
…ion (OOM: autograd graph + eager full-score materialization)
…ttention variant averaging + proportional scale

The first draft mis-ported SPA two ways, producing periodic-tiled mush worse than
naive high-res generation (caught on GPU validation @2048):
  - position map used modulo wrapping ((y-shift)%H)%bundle -> periodic tiling;
    the paper uses a monotonic bundle coarsening phi(x)=ceil((x+1-n1)/size).
  - averaging happened over 4 full model forwards at the output; the paper
    averages per-RoPE-variant attention *outputs* inside each layer (identity
    mean(softmax(A_n))@v == mean(softmax(A_n)@v)), one forward per step.
Also adds the reference's proportional attention scale and corrects group_num
default to 80. HAP pruning, NTK RoPE scaling and per-step SPA scheduling are
documented as follow-ups (not yet ported). Ref: https://github.com/zylwithxy/HRDiT

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…-res mechanism)

GPU validation of the SPA-only port: 2048 coherent but cross-hatched, 4096 washed
out. Root cause vs reference inference.py: SPA is a *leading-few-steps* nudge
(spa_steps=[3,0] -> 3 steps at 2048, none at 4096), while NTK-aware RoPE scaling
(theta*=ntk_factor, [4,10] per stage) applied on *every* step is the primary
high-res mechanism. Running SPA every step over-averaged -> washout.

  - flux_rope(): NTK RoPE (== diffusers FluxPosEmbed at factor 1).
  - _SPAState: base NTK rope on every step; SPA variants only while spa_active.
  - per-stage ntk_factor / spa_steps / guidance_scale_highres, gated per step.

Remaining (documented): frequency-domain structure guidance + HAP pruning.
Ref: https://github.com/zylwithxy/HRDiT (inference.py, hrdit/transformer.py)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
… for the top stage

Adds the DemoFusion-style structure guidance the reference relies on at high res,
so 4096 stops drifting to a washed-out mean:
  - shared flow-match schedule across stages (pred_x0 references align per timestep)
  - per-stage prior: decode -> bicubic upscale -> sharpen -> re-encode
  - custom flowmatch step: per-step low-frequency injection from the upsampled
    previous-stage pred_x0 (alpha, Butterworth FFT split) + velocity momentum (beta)
Keeps validated NTK RoPE (every step) + gated SPA (leading steps). Reference
hyperparameters wired: ntk=[4,10], spa_steps=[3,0], steps=[17,10], guidance=[4.5,6],
alphas=[1.0,0.25], betas=[0.5,0.5], filter_ratio=0.2.

Component-verified offline vs diffusers: flux_rope==FluxPosEmbed, pack/unpack dims,
scheduler tail-index alignment, bundle variants in-range. Deferred: HAP, swin, DWT.
Ref: https://github.com/zylwithxy/HRDiT (pipeline.py flowmatch_step, inference.py)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…e NTK, structure-guidance helpers; drop HAP)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: remyx-ai[bot] <289541483+remyx-ai[bot]@users.noreply.github.com>
@smellslikeml smellslikeml changed the title [Community pipeline] HRDiT: training-free high-resolution (4K) generation for FLUX.1-dev Community Pipeline: Add HRDiT (training-free high-resolution FLUX.1-dev) Aug 14, 2026
@smellslikeml
smellslikeml deleted the hrdit-training-free-high-resolution-image-generation-with-of branch August 14, 2026 21:57
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant