Use one PyTorch nightly pin across main - #23359
JacobSzwejbka wants to merge 1 commit into
Conversation
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23359
Note: Links to docs will display an error until the docs builds have been completed. ❌ 1 New Failure, 1 Cancelled JobAs of commit 62f47d7 with merge base 7dc8dd9 ( NEW FAILURE - The following job has failed:
CANCELLED JOB - The following job was cancelled. Please retry:
This comment was automatically generated by Dr. CI and updates every 15 minutes. |
This PR needs a
|
| # exceptional train requires data here, not another branch in its control flow. | ||
| # TorchAO's local suffix is selected separately because its aarch64 wheel is | ||
| # published on the CPU channel. | ||
| CUDA_DEPENDENCY_VERSIONS = { |
There was a problem hiding this comment.
@shoumikhin what was the issue with cuda pins that motivated us having these btw?
ab64ec0 to
e7603f5
Compare
e7603f5 to
649a219
Compare
649a219 to
ce39b9f
Compare
ce39b9f to
a73b5e6
Compare
shoumikhin
left a comment
There was a problem hiding this comment.
A few things from a read of the pin move. The first two look like they need a fix before this lands, the rest are smaller.
| # The one PyTorch wheel train used by main. Its embedded nightly date also | ||
| # selects the matching source commit and c10 headers; those are derived outputs, | ||
| # not an independently chosen pin. | ||
| PYTORCH_VERSION = "2.15.0.dev20260922" |
There was a problem hiding this comment.
On Linux aarch64, the CUDA builds of torchvision 0.30.0.dev20260922 ask for torch 2.15.0.dev20260921, not 0922. The x86_64 builds of cu130 and cu132 ask for 0922. So this set cannot be installed together on an aarch64 CUDA machine, and because the domain libraries go in a separate pip call, pip will most likely move torch off the pin there.
The picker matches wheels by the date in the filename and never reads what each wheel requires, so the same skew will come back on the next bump. Could the picker read each torchvision wheel's torch requirement and only accept a date that matches on every train and architecture?
| _DEPENDENCY_CONFIG = runpy.run_path("torch_pin.py") | ||
| PYTORCH_INDEX_URL = _DEPENDENCY_CONFIG["PYTORCH_INDEX_URL"] | ||
| TORCHAO_INDEX_URL = _DEPENDENCY_CONFIG["TORCHAO_INDEX_URL"] | ||
| PYTORCH_PACKAGES = ("torch", "torchvision", "torchaudio") |
There was a problem hiding this comment.
PyTorch stopped publishing cp310 nightlies after 2026-09-23. cp311 through cp314 are still published daily, the newest being 2026-10-02. Because the picker only accepts cp310 wheels, it can never choose a date later than 2026-09-23, on any channel.
It will not fail. It will keep returning the same date, the weekly bump will report that the pin is already current, and nobody gets told. The nightly index keeps about 60 days, so around late November the picker will start saying it found nothing.
Could this require a tag that is still published, or fail loudly when the chosen date falls more than a few days behind the newest nightly on the index?
| # not an independently chosen pin. | ||
| PYTORCH_VERSION = "2.15.0.dev20260922" | ||
| PYTORCH_INDEX_URL = "https://download.pytorch.org/whl/nightly" | ||
| TORCHVISION_VERSION = "0.30.0.dev20260922" |
There was a problem hiding this comment.
The source-built CI jobs and the binary install now disagree on torchvision. The setup scripts still build torchvision from the release/0.29 branch while this pin moves users to the 0.30 nightly. torchaudio lines up (release/2.11 against the 2.11 nightly), so it is only torchvision.
Could those scripts take the version from this file, or could the description say that source builds keep their own domain pins?
| TORCHAUDIO_VERSION = "2.11.0.dev20260922" | ||
|
|
||
| TORCHAO_INDEX_URL = PYTORCH_INDEX_URL | ||
| TORCHAO_NIGHTLY_VERSION = "0.19.0.dev20260907" |
There was a problem hiding this comment.
The commit hook runs the full pin updater whenever this file is staged, and this file now also holds the TorchAO pin, the ROCm pins and the CUDA train list.
So a commit that only rolls back TorchAO or drops a CUDA train still runs the updater. The updater writes the newest TorchAO version back into this file and moves the ao submodule, after your staged copy was taken, and the hook stages neither of them. It also needs network access, so offline the commit just fails.
Could the hook run only when the PyTorch version itself changes?
| ROCM_TORCHAO_NIGHTLY_VERSION = "0.19.0.dev20260805" | ||
| # PyTorch no longer publishes current ROCm nightlies. These jobs remain on the | ||
| # newest compatible test-index wheel instead of silently weakening the main pin. | ||
| ROCM_PYTORCH_VERSION = "2.14.0" |
There was a problem hiding this comment.
The comment says these jobs sit on the newest compatible test-index wheel, but the rocm7.2 test index also carries 2.14.1, with cp310 through cp315 Linux x86_64 wheels. Either the pin can move up, or the comment could say why 2.14.1 does not fit.
| self.installer.install_optional_example_requirements(nightly) | ||
| return [call.args[0] for call in run.call_args_list] | ||
|
|
||
| def test_every_binary_install_uses_one_pytorch_pin(self): |
There was a problem hiding this comment.
The deleted test file also covered things that were not about cu134, and two of them have no replacement here: that a normal install pins torchao at all, and that the torchao bound in setup.py accepts the pinned version.
The only torchao test left is the negative one, that an explicit source build omits the wheel pin. So removing the torchao pin from the installer would keep every test green. A couple of assertions in the every-train test would bring this back.
| contains(needs.changed-files.outputs.changed-files, 'extension/flat_tensor') || | ||
| contains(needs.changed-files.outputs.changed-files, 'extension/pytree') || | ||
| contains(needs.changed-files.outputs.changed-files, 'pyproject.toml') || | ||
| contains(needs.changed-files.outputs.changed-files, 'torch_pin.py') || |
There was a problem hiding this comment.
You added torch_pin.py here, but two jobs that gate on install_requirements.py did not get it: test-coreml-bc-macos further down this file, and test-huggingface-transformers-macos in the trunk workflow. The pins now live only in torch_pin.py, so a pin-only change, including the weekly bot bump, skips both. Could you add it to those two as well?
Make one full PyTorch nightly version authoritative for binary installs and source synchronization. Remove the cu134-only dependency path and drop cu126 from main's supported wheel trains, while keeping the unavailable ROCm nightly as an explicit compatibility exception. Authored with assistance from Codex.
a73b5e6 to
62f47d7
Compare
|
On this nightly, two TurboQuant SDPA tests in |
/Users/jakeszwe/.zshenv:.:1: no such file or directory: /tmp/codex-rust/env
Summary
Use one full PyTorch nightly version as the authoritative pin for ExecuTorch main.
2.15.0.dev20260922now drives CPU, cu130, cu132, and cu134 binary installs; its embedded date also selects the matching PyTorch source commit and vendored c10 headers.This removes the independent stable
TORCH_VERSION, date-onlyNIGHTLY_VERSION, and the cu134-only dependency mapping/control flow. It also removes cu126 from main's supported install and CUDA compatibility matrices.ROCm is deliberately explicit rather than pretending to share the PyTorch pin: the rocm7.2 channel no longer publishes a current matching PyTorch nightly, so the two ROCm jobs retain
2.14.0as a named compatibility exception. TorchAO does not need a ROCm-compiled wheel for these export paths, so ROCm now uses the same0.19.0.dev20260907TorchAO pin as every other platform via its platform-neutral wheel.CUDA_WHEEL_VERSIONSis the only CUDA wheel list on main. If release preparation needs to filter that set temporarily, the preparation layer will own that release state rather than adding a second global list. This PR does not add release preparation or branch-cut behavior; those remain in the stacked follow-ups.Why each file changes
torch_pin.py: defines the one full PyTorch nightly pin, matching torchvision/torchaudio versions, one TorchAO nightly pin, nightly index, explicit ROCm PyTorch exception, and one shared CUDA wheel list..ci/docker/ci_commit_pins/pytorch.txt: records the exact PyTorch source commit underlying that nightly.install_requirements.py: consumes that pin for CPU/cu130/cu132/cu134 and removes all cu134-specific package selection.install_utils.py: removes CUDA 12.6 from the versions supported by main.setup.py: removes the cu134-specific TorchAO metadata branch..github/scripts/update_pytorch_pin.py: selects the newest complete package date with Python 3.10 wheels for every configured platform and nightly channel, updates the three exact package versions, and derives source synchronization from the full PyTorch version..github/workflows/weekly-pytorch-pin-bump.yml: gives the updater a maximum date and names the bot branch/commit/PR after the version actually selected..github/scripts/filter_cuda_matrix.py: reads the release wheel trains fromtorch_pin.pyand removes the obsolete cu126 rationale..ci/scripts/tests/test_pytorch_dependencies.py: replaces cu134-only coverage with tests proving one pin is used across CPU and all three CUDA trains, including ARM64, Windows, source builds, and TorchAO source builds..ci/scripts/tests/test_cu134_dependencies.py: deleted because cu134 no longer has special dependency behavior..ci/scripts/tests/test_update_pytorch_pin.py: tests full-version date derivation, complete-wheel-set selection, and centralized pin updates..ci/scripts/tests/test_filter_cuda_matrix.py: verifies the matrix uses the one centralized CUDA wheel list..ci/scripts/tests/test_cuda_workflow.py: updates the expected compatibility matrix after dropping CUDA 12.6..github/workflows/cuda.yml: stops testing CUDA 12.6 on main..github/workflows/lint.yml: installs the exact centralized torch, torchvision, and torchaudio nightlies..github/workflows/windows-msvc.yml: installs the centralized CPU nightly on Windows and runs when its config changes..ci/scripts/test_wheel_package_qnn.sh: installs the same centralized CPU nightly for QNN wheel tests..ci/scripts/test-rocm-aoti.sh: uses the explicit ROCm PyTorch compatibility pin with the shared, platform-neutral TorchAO nightly instead of requiring compiled ROCm TorchAO kernels..ci/scripts/test-rocm-voxtral.sh: uses the same PyTorch exception and standard TorchAO pin, removing the custom wheel URL and dependency-filtering workaround..github/workflows/build-wheels-aarch64-linux.yml: reruns when install logic or the pin changes..github/workflows/build-wheels-cuda-aarch64-linux.yml: reruns when the pin changes..github/workflows/build-wheels-cuda-linux.yml: reruns when the pin changes..github/workflows/build-wheels-linux.yml: reruns when install logic or the pin changes..github/workflows/build-wheels-macos.yml: reruns when install logic or the pin changes..github/workflows/build-wheels-windows.yml: reruns when install logic or the pin changes..github/workflows/pull.yml: includes pin-only changes in the relevant pull-request test decision..githooks/README.md: describes the full version pin users now troubleshoot.Test plan
python .ci/scripts/tests/test_pytorch_dependencies.py— 6 testspython .ci/scripts/tests/test_update_pytorch_pin.py— 3 testspython .ci/scripts/tests/test_filter_cuda_matrix.py— 26 testspython -m unittest discover -s .ci/scripts/tests -p test_cuda_workflow.py— 6 testsgit diff --checkAuthored with OpenAI Codex.