Stream targeted block re-quantization's calibration pass - #2064
aquilarubra wants to merge 8 commits into
Conversation
11ba30e to
3be3fae
Compare
c6e097e to
3841bb3
Compare
Rebase of PR intel#2064 onto current main. Most of the original stack (disk-streaming-core intel#2061, resumability intel#2062, AutoScheme streaming intel#2063) is already merged verbatim or superseded by intel#2220, so this commit carries only the calibration-pass streaming logic and its tests that weren't already upstream.
e193420 to
3f26e3c
Compare
|
Rebased onto current |
| export AR_RESUME_DIR=/path/to/resume/state | ||
| ``` | ||
|
|
||
| ### AR_DISK_STREAM_MODEL |
There was a problem hiding this comment.
We already have this variable; please clean up the relevant documentation.
There was a problem hiding this comment.
Good catch — the merge with main had duplicated the AR_DISK_STREAM_MODEL and AR_RESUME_DIR sections verbatim. Removed the duplicate copy, kept the richer one (mentions the parallel-scoring interaction). Pushed.
xin3he
left a comment
There was a problem hiding this comment.
Nice catch, please clean up the document
| export AR_DISK_STREAM_MODEL=1 | ||
| ``` | ||
|
|
||
| ### AR_RESUME_DIR |
There was a problem hiding this comment.
We already have this variable; please clean up the relevant documentation.
The merge with main duplicated these two sections verbatim; keep the richer copy (mentions parallel-scoring interaction) and drop the dupe.
38b10ff to
8f89538
Compare
|
/azp run Unit-Test-CUDA-AutoRound. |
|
/azp run Unit-Test-CUDA-AutoRound |
|
No pipelines are associated with this pull request. |
|
Azure Pipelines successfully started running 1 pipeline(s). |
|
/azp run Unit-Test-CUDA-AutoRound |
|
Azure Pipelines successfully started running 1 pipeline(s). |
|
/azp run Unit-Test-CUDA-AutoRound |
|
Azure Pipelines successfully started running 1 pipeline(s). |
|
@Copilot move the cpu test to cuda test |
|
@copilot move the cpu test to cuda test |
|
@aquilarubra Hi, the cpu ut is time out, we may need to move the new ut to cuda |
|
/azp run Unit-Test-CUDA-AutoRound |
|
Azure Pipelines successfully started running 1 pipeline(s). |
| # held-out-loss-eval case. | ||
| stream_ctx = None | ||
| _moved_tensors = [] | ||
| if envs.AR_DISK_STREAM_MODEL: |
There was a problem hiding this comment.
For non-MoE models, Transformers already works this way, so no extra ops or code are needed. For MoE models, however, the current PR is not sufficient. The main branch already includes an attempt to address the MoE issue.
So I suppose there is no need for this PR. Please feel free to correct me if I am wrong.
There was a problem hiding this comment.
Thanks for looking — agreed that non-MoE models need nothing here, and this PR isn't about MoE. The remaining diff addresses a narrower case that I can still reproduce on current main (6afaecd): with AR_DISK_STREAM_MODEL=1 the model is a meta-device skeleton, and when to_quant_block_names targets a block other than the first, the blocks before it are never materialized. The calibration forward then propagates meta tensors until it fails with RuntimeError: Tensor on device meta is not on the expected device cpu!.
test_targeted_block_with_disk_streaming_does_not_crash fails on main with exactly that error and passes with this PR. It's gated on the env var, so default behavior is unchanged.
I've also moved the new tests to test/unit/test_cuda/ as requested (the CPU run was timing out) and removed an unused variable. If you'd prefer this not live in calibration/llm.py, I'm glad to restructure it or close the PR.
There was a problem hiding this comment.
@wenhuach21 pushed 852b112 with the tests moved to the CUDA tier — could you please re-run /azp run Unit-Test-CUDA-AutoRound on it when you get a chance? Thanks!
73a716f to
6075eee
Compare
|
/azp run Unit-Test-CUDA-AutoRound |
|
Azure Pipelines successfully started running 1 pipeline(s). |
|
/azp run Unit-Test-CUDA-AutoRound |
|
Azure Pipelines successfully started running 1 pipeline(s). |
Summary
Depends on #2061 (disk-streaming core) — stacked on top of it, so the diff includes those commits until it merges; only the last commit is new here.
to_quant_block_nameslets a caller restrict tuning to a subset of decoder blocks — useful for cheaply re-quantizing just a couple of blocks in an already-produced checkpoint at higher precision instead of redoing a full multi-hour run. Combined withAR_DISK_STREAM_MODEL, this crashes: the "cache block inputs" calibration forward pass needs real weights in every block leading up to (and sometimes through) the target block(s), but nothing materializes blocks outsidequant_block_listfor this specific pass — they stay meta forever, and the forward silently propagates meta-ness until it collides with a genuinely-materialized module. Full (unrestricted) runs never hit this, sincequant_block_listalready covers every block in that case.What's in this PR
auto_round/calibration/llm.py: when disk streaming is active and any decoder block still has meta parameters at the point this calibration forward runs, wraps it with the existingstream_block_forwardprimitive (from #2061), scoped to just the still-meta blocks. Only activates when there's something left meta to fix, so it's a no-op for the normal full-quantization path.Also adds an
AR_CALIB_STREAM_DEVICEenv-gated fast path: the default keeps this forward entirely on CPU (mixing a GPU-streamed block with CPU-resident hidden states crashes with a device mismatch), but a full CPU forward through every pre-target block of a 100B+ model is unusably slow for this specific targeted-requant use case. When set, every already-real tensor is moved to that device for the pass's duration and moved back afterward.Validation
Reproduced and fixed against a tiny hybrid-MoE fixture with an MTP head, with both a single restricted block and two adjacent ones, plus a control run of the full (unrestricted) path confirming no regression.