Skip to content

Use one PyTorch nightly pin across main - #23359

Open
JacobSzwejbka wants to merge 1 commit into
mainfrom
refactor/release-config
Open

JacobSzwejbka wants to merge 1 commit into
mainfrom
refactor/release-config

Conversation

@JacobSzwejbka

@JacobSzwejbka JacobSzwejbka commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

/Users/jakeszwe/.zshenv:.:1: no such file or directory: /tmp/codex-rust/env

Summary

Use one full PyTorch nightly version as the authoritative pin for ExecuTorch main. 2.15.0.dev20260922 now drives CPU, cu130, cu132, and cu134 binary installs; its embedded date also selects the matching PyTorch source commit and vendored c10 headers.

This removes the independent stable TORCH_VERSION, date-only NIGHTLY_VERSION, and the cu134-only dependency mapping/control flow. It also removes cu126 from main's supported install and CUDA compatibility matrices.

ROCm is deliberately explicit rather than pretending to share the PyTorch pin: the rocm7.2 channel no longer publishes a current matching PyTorch nightly, so the two ROCm jobs retain 2.14.0 as a named compatibility exception. TorchAO does not need a ROCm-compiled wheel for these export paths, so ROCm now uses the same 0.19.0.dev20260907 TorchAO pin as every other platform via its platform-neutral wheel.

CUDA_WHEEL_VERSIONS is the only CUDA wheel list on main. If release preparation needs to filter that set temporarily, the preparation layer will own that release state rather than adding a second global list. This PR does not add release preparation or branch-cut behavior; those remain in the stacked follow-ups.

Why each file changes

  • torch_pin.py: defines the one full PyTorch nightly pin, matching torchvision/torchaudio versions, one TorchAO nightly pin, nightly index, explicit ROCm PyTorch exception, and one shared CUDA wheel list.
  • .ci/docker/ci_commit_pins/pytorch.txt: records the exact PyTorch source commit underlying that nightly.
  • install_requirements.py: consumes that pin for CPU/cu130/cu132/cu134 and removes all cu134-specific package selection.
  • install_utils.py: removes CUDA 12.6 from the versions supported by main.
  • setup.py: removes the cu134-specific TorchAO metadata branch.
  • .github/scripts/update_pytorch_pin.py: selects the newest complete package date with Python 3.10 wheels for every configured platform and nightly channel, updates the three exact package versions, and derives source synchronization from the full PyTorch version.
  • .github/workflows/weekly-pytorch-pin-bump.yml: gives the updater a maximum date and names the bot branch/commit/PR after the version actually selected.
  • .github/scripts/filter_cuda_matrix.py: reads the release wheel trains from torch_pin.py and removes the obsolete cu126 rationale.
  • .ci/scripts/tests/test_pytorch_dependencies.py: replaces cu134-only coverage with tests proving one pin is used across CPU and all three CUDA trains, including ARM64, Windows, source builds, and TorchAO source builds.
  • .ci/scripts/tests/test_cu134_dependencies.py: deleted because cu134 no longer has special dependency behavior.
  • .ci/scripts/tests/test_update_pytorch_pin.py: tests full-version date derivation, complete-wheel-set selection, and centralized pin updates.
  • .ci/scripts/tests/test_filter_cuda_matrix.py: verifies the matrix uses the one centralized CUDA wheel list.
  • .ci/scripts/tests/test_cuda_workflow.py: updates the expected compatibility matrix after dropping CUDA 12.6.
  • .github/workflows/cuda.yml: stops testing CUDA 12.6 on main.
  • .github/workflows/lint.yml: installs the exact centralized torch, torchvision, and torchaudio nightlies.
  • .github/workflows/windows-msvc.yml: installs the centralized CPU nightly on Windows and runs when its config changes.
  • .ci/scripts/test_wheel_package_qnn.sh: installs the same centralized CPU nightly for QNN wheel tests.
  • .ci/scripts/test-rocm-aoti.sh: uses the explicit ROCm PyTorch compatibility pin with the shared, platform-neutral TorchAO nightly instead of requiring compiled ROCm TorchAO kernels.
  • .ci/scripts/test-rocm-voxtral.sh: uses the same PyTorch exception and standard TorchAO pin, removing the custom wheel URL and dependency-filtering workaround.
  • .github/workflows/build-wheels-aarch64-linux.yml: reruns when install logic or the pin changes.
  • .github/workflows/build-wheels-cuda-aarch64-linux.yml: reruns when the pin changes.
  • .github/workflows/build-wheels-cuda-linux.yml: reruns when the pin changes.
  • .github/workflows/build-wheels-linux.yml: reruns when install logic or the pin changes.
  • .github/workflows/build-wheels-macos.yml: reruns when install logic or the pin changes.
  • .github/workflows/build-wheels-windows.yml: reruns when install logic or the pin changes.
  • .github/workflows/pull.yml: includes pin-only changes in the relevant pull-request test decision.
  • .githooks/README.md: describes the full version pin users now troubleshoot.

Test plan

  • python .ci/scripts/tests/test_pytorch_dependencies.py — 6 tests
  • python .ci/scripts/tests/test_update_pytorch_pin.py — 3 tests
  • python .ci/scripts/tests/test_filter_cuda_matrix.py — 26 tests
  • python -m unittest discover -s .ci/scripts/tests -p test_cuda_workflow.py — 6 tests
  • shell syntax checks for the three changed shell scripts
  • parsed all 11 changed workflow files as YAML
  • ufmt, flake8, mypy on standalone scripts, Python compilation, and git diff --check
  • live nightly-index selection returns the pinned Sep 22 torch/vision/audio set across CPU, cu130, cu132, and cu134
  • exact CPU package set resolves successfully in a clean pip dry run
  • the shared TorchAO pin resolves to a platform-neutral wheel with no Torch dependency; ROCm execution is covered by PR CI
  • vendored c10 files match the exact pinned PyTorch source commit

Authored with OpenAI Codex.

@pytorch-bot

pytorch-bot Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23359

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Cancelled Job

As of commit 62f47d7 with merge base 7dc8dd9 (image):

NEW FAILURE - The following job has failed:

CANCELLED JOB - The following job was cancelled. Please retry:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Oct 2, 2026
@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

Comment thread torch_pin.py Outdated
# exceptional train requires data here, not another branch in its control flow.
# TorchAO's local suffix is selected separately because its aarch64 wheel is
# published on the CPU channel.
CUDA_DEPENDENCY_VERSIONS = {

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@shoumikhin what was the issue with cuda pins that motivated us having these btw?

@JacobSzwejbka
JacobSzwejbka force-pushed the refactor/release-config branch from ab64ec0 to e7603f5 Compare October 2, 2026 15:57
@JacobSzwejbka JacobSzwejbka changed the title Centralize release dependency configuration Use one PyTorch nightly pin across main Oct 2, 2026
@JacobSzwejbka
JacobSzwejbka force-pushed the refactor/release-config branch from e7603f5 to 649a219 Compare October 2, 2026 16:31
@JacobSzwejbka
JacobSzwejbka force-pushed the refactor/release-config branch from 649a219 to ce39b9f Compare October 2, 2026 17:02
@github-actions github-actions Bot added the ciflow/docker Trigger docker-builds to rebuild and push the CI images for this PR label Oct 2, 2026
@JacobSzwejbka
JacobSzwejbka force-pushed the refactor/release-config branch from ce39b9f to a73b5e6 Compare October 2, 2026 17:27

@shoumikhin shoumikhin left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A few things from a read of the pin move. The first two look like they need a fix before this lands, the rest are smaller.

Comment thread torch_pin.py
# The one PyTorch wheel train used by main. Its embedded nightly date also
# selects the matching source commit and c10 headers; those are derived outputs,
# not an independently chosen pin.
PYTORCH_VERSION = "2.15.0.dev20260922"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

On Linux aarch64, the CUDA builds of torchvision 0.30.0.dev20260922 ask for torch 2.15.0.dev20260921, not 0922. The x86_64 builds of cu130 and cu132 ask for 0922. So this set cannot be installed together on an aarch64 CUDA machine, and because the domain libraries go in a separate pip call, pip will most likely move torch off the pin there.

The picker matches wheels by the date in the filename and never reads what each wheel requires, so the same skew will come back on the next bump. Could the picker read each torchvision wheel's torch requirement and only accept a date that matches on every train and architecture?

_DEPENDENCY_CONFIG = runpy.run_path("torch_pin.py")
PYTORCH_INDEX_URL = _DEPENDENCY_CONFIG["PYTORCH_INDEX_URL"]
TORCHAO_INDEX_URL = _DEPENDENCY_CONFIG["TORCHAO_INDEX_URL"]
PYTORCH_PACKAGES = ("torch", "torchvision", "torchaudio")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

PyTorch stopped publishing cp310 nightlies after 2026-09-23. cp311 through cp314 are still published daily, the newest being 2026-10-02. Because the picker only accepts cp310 wheels, it can never choose a date later than 2026-09-23, on any channel.

It will not fail. It will keep returning the same date, the weekly bump will report that the pin is already current, and nobody gets told. The nightly index keeps about 60 days, so around late November the picker will start saying it found nothing.

Could this require a tag that is still published, or fail loudly when the chosen date falls more than a few days behind the newest nightly on the index?

Comment thread torch_pin.py
# not an independently chosen pin.
PYTORCH_VERSION = "2.15.0.dev20260922"
PYTORCH_INDEX_URL = "https://download.pytorch.org/whl/nightly"
TORCHVISION_VERSION = "0.30.0.dev20260922"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The source-built CI jobs and the binary install now disagree on torchvision. The setup scripts still build torchvision from the release/0.29 branch while this pin moves users to the 0.30 nightly. torchaudio lines up (release/2.11 against the 2.11 nightly), so it is only torchvision.

Could those scripts take the version from this file, or could the description say that source builds keep their own domain pins?

Comment thread torch_pin.py
TORCHAUDIO_VERSION = "2.11.0.dev20260922"

TORCHAO_INDEX_URL = PYTORCH_INDEX_URL
TORCHAO_NIGHTLY_VERSION = "0.19.0.dev20260907"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The commit hook runs the full pin updater whenever this file is staged, and this file now also holds the TorchAO pin, the ROCm pins and the CUDA train list.

So a commit that only rolls back TorchAO or drops a CUDA train still runs the updater. The updater writes the newest TorchAO version back into this file and moves the ao submodule, after your staged copy was taken, and the hook stages neither of them. It also needs network access, so offline the commit just fails.

Could the hook run only when the PyTorch version itself changes?

Comment thread torch_pin.py
ROCM_TORCHAO_NIGHTLY_VERSION = "0.19.0.dev20260805"
# PyTorch no longer publishes current ROCm nightlies. These jobs remain on the
# newest compatible test-index wheel instead of silently weakening the main pin.
ROCM_PYTORCH_VERSION = "2.14.0"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The comment says these jobs sit on the newest compatible test-index wheel, but the rocm7.2 test index also carries 2.14.1, with cp310 through cp315 Linux x86_64 wheels. Either the pin can move up, or the comment could say why 2.14.1 does not fit.

self.installer.install_optional_example_requirements(nightly)
return [call.args[0] for call in run.call_args_list]

def test_every_binary_install_uses_one_pytorch_pin(self):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The deleted test file also covered things that were not about cu134, and two of them have no replacement here: that a normal install pins torchao at all, and that the torchao bound in setup.py accepts the pinned version.

The only torchao test left is the negative one, that an explicit source build omits the wheel pin. So removing the torchao pin from the installer would keep every test green. A couple of assertions in the every-train test would bring this back.

contains(needs.changed-files.outputs.changed-files, 'extension/flat_tensor') ||
contains(needs.changed-files.outputs.changed-files, 'extension/pytree') ||
contains(needs.changed-files.outputs.changed-files, 'pyproject.toml') ||
contains(needs.changed-files.outputs.changed-files, 'torch_pin.py') ||

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You added torch_pin.py here, but two jobs that gate on install_requirements.py did not get it: test-coreml-bc-macos further down this file, and test-huggingface-transformers-macos in the trunk workflow. The pins now live only in torch_pin.py, so a pin-only change, including the weekly bot bump, skips both. Could you add it to those two as well?

Make one full PyTorch nightly version authoritative for binary installs and source synchronization. Remove the cu134-only dependency path and drop cu126 from main's supported wheel trains, while keeping the unavailable ROCm nightly as an explicit compatibility exception.

Authored with assistance from Codex.
@JacobSzwejbka
JacobSzwejbka force-pushed the refactor/release-config branch from a73b5e6 to 62f47d7 Compare October 2, 2026 18:48
@shoumikhin

Copy link
Copy Markdown
Contributor

On this nightly, two TurboQuant SDPA tests in unittest-cuda fail inside inductor: test_export_cuda and test_e2e_cpp_runner, both with 'NotImplementedType' object has no attribute 'expr'. They pass on main with 2.14, so the pin move likely needs a fix or a skip.

This branch was successfully deployed

1 active deployment
cadence — 62f47d70 Deployed Oct 2, 2026 by JacobSzwejbka via hifi-op-test / hifi4 #31140
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/docker Trigger docker-builds to rebuild and push the CI images for this PR CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants