Skip to content

Stop publishing CUDA 13.0 wheels - #23380

Merged
shoumikhin merged 1 commit into
pytorch:mainfrom
shoumikhin:stop-publishing-cuda-13-0-wheels
Oct 3, 2026
Merged

shoumikhin merged 1 commit into
pytorch:mainfrom
shoumikhin:stop-publishing-cuda-13-0-wheels

Conversation

@shoumikhin

Copy link
Copy Markdown
Contributor

Summary

ExecuTorch publishes a CUDA wheel for each CUDA version that PyTorch's build matrix offers. PyTorch is moving off CUDA 13.0. It now keeps 13.0 on its nightly builds only as a temporary hold for other projects, and its release-candidate builds do not carry it (see pytorch/test-infra#8989 and pytorch/pytorch#199520).

If ExecuTorch keeps publishing cu130, nightlies carry a wheel that no release gets, and that wheel goes away again when the hold ends. It is easy to miss when it goes. The wheel filter skips a CUDA version that the matrix no longer offers, and the run stays green. That is what happened from Sep 29 to Oct 2, when PyTorch briefly dropped 13.0 and no cu130 wheel was built.

This change publishes CUDA 13.2 and 13.4 only:

  • filter_cuda_matrix.py: remove cu130 from the published list. The single row built for a pull request moves from cu130 to cu132.
  • test_filter_cuda_matrix.py: update the pinned published list.
  • Docs: remove the cu130 row from the install table and use cu132 in the examples. A machine on CUDA 13.0 can use the cu132 wheel, because CUDA minor versions are compatible.

Not changed: install_utils.py (it still accepts a CUDA 13.0 toolkit for source builds), the GPU architecture table, and the CI jobs that run on a CUDA 13.0 image. None of those decide which wheels get published.

Test plan

  • python .ci/scripts/tests/test_filter_cuda_matrix.py: 26 tests pass. With the original filter, the two tests that pin the published list fail.
  • Ran the shared matrix generator for Linux x86_64 and aarch64 on the nightly and test channels and for a pull request, then fed its output through the filter. Nightly goes from 15 rows to 10 (cu132 and cu134, five Python versions each). The test channel stays at 10. The pull request row is cu132.
  • black and flake8 give the same results as on main.

ExecuTorch publishes a wheel for each CUDA version the PyTorch build
matrix offers. PyTorch is moving off CUDA 13.0. It now keeps 13.0 on its
nightly builds only as a temporary hold for other projects, and its
release-candidate builds do not carry it.

If ExecuTorch keeps publishing cu130, nightlies carry a train that no
release gets, and the train goes away again when PyTorch ends the hold.
Nobody notices when it goes, because the filter skips a CUDA version the
generator no longer offers and the run stays green. That is how the cu130
wheel stopped without a failure for four days (Sep 29 to Oct 2).

This change publishes CUDA 13.2 and 13.4 only. The single row built for a
pull request moves from cu130 to cu132, and the install docs drop the
cu130 row. A machine on CUDA 13.0 can use the cu132 wheel, because CUDA
minor versions are compatible. A machine on an older CUDA can still build
from source.

What was tested:
- The filter unit tests pass (26 of 26). With the original filter, the two
  tests that pin the published list fail.
- Replayed the shared matrix generator for Linux x86_64 and aarch64, on the
  nightly and test channels and for a pull request. Nightly goes from 15
  rows to 10 (cu132 and cu134, five Pythons each). The test channel stays
  at 10. The pull request row is cu132.
- black and flake8 give the same results as on main.
Copilot AI balanced review requested due to automatic review settings October 2, 2026 23:38
@shoumikhin shoumikhin added the release notes: build Changes related to build, including dependency upgrades, build flags, optimizations, etc. label Oct 2, 2026
@shoumikhin shoumikhin added the release notes: build Changes related to build, including dependency upgrades, build flags, optimizations, etc. label Oct 2, 2026
@pytorch-bot

pytorch-bot Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23380

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit 51d716e with merge base fbf91c2 (image):

NEW FAILURE - The following job has failed:

  • Cadence Build & Test / Resolve CI docker image / resolve (gh)
    ##[error]Refusing to check out fork pull request code from a 'pull_request_target' workflow. This workflow runs with the base repository's GITHUB_TOKEN, secrets, default-branch cache scope, and runner access. Fetching and executing a fork's code in that trusted context commonly leads to "pwn request" vulnerabilities. To opt in, review the risks at https://gh.io/securely-using-pull_request_target and set 'allow-unsafe-pr-checkout: true' on the actions/checkout step.

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Oct 2, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@huydhn huydhn left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The change LGTM! so stamped, but plz check if there is any communication that needs to be done before merging this one. I'll leave this to the team to decide when to merge. Here is the RFC from PyTorch to remove 13.0 if it helps pytorch/pytorch#190385

@shoumikhin
shoumikhin merged commit 3794e44 into pytorch:main Oct 3, 2026
399 of 408 checks passed
Gasoonjia added a commit that referenced this pull request Oct 3, 2026
Stacked on #23121, which ships linkable libraries in the Windows wheel.

A Windows user can run a CUDA program but cannot lower one: lowering compiles the model
with a toolchain only the Linux side has, and the CUDA Windows CI job already exports on
Linux and runs on Windows for exactly this reason. What was missing is a wheel that gives
that Windows user the runtime. This adds a Windows CUDA wheel carrying the CUDA delegate
as C++ libraries, the same components the Linux CUDA wheel ships (#21645, #21668), and
nothing CUDA in the Python module, which on Windows has no CUDA program to lower. The
export (AOT) side is unchanged and identical to the CPU wheel's; only the runtime links CUDA.

    lib/executorch_backend_cuda.dll    + .lib   the delegate          executorch::backend_cuda
    lib/executorch_extension_cuda.dll  + .lib   the stream helper     executorch::extension_cuda
    backends/cuda/aoti_cuda_shims.dll            the AOTI shim layer   (runtime DLL of the delegate)
    data/lib/aoti_cuda_shims.lib                 already shipped as the lowering stub; also the
                                                 shim layer's import library here
    extension/pybindings/_C.pyd                  unchanged: no CUDA dependency, no CudaBackend

A C++ application links it exactly as on Linux:

    find_package(executorch REQUIRED COMPONENTS kernels_optimized backend_cuda)
    target_link_libraries(app PRIVATE executorch::runtime executorch::kernels_optimized
                                      executorch::backend_cuda)
    # plus the $<TARGET_RUNTIME_DLLS:app> copy from the CPU PR, which now also brings the
    # stream helper and the shim layer

The wheel's DLLs import no CUDA runtime DLL: CMake links cudart.lib into them on Windows,
and that reaches the driver (nvcuda.dll). The compiled model is different: the library
AOTInductor builds into model.pte imports cudart64_13.dll, which comes from the CUDA
Toolkit's bin directory on PATH, as the guide says. PyPI's nvidia-cuda-runtime has win_amd64
builds that carry that DLL, but a C++ program does not search site-packages, so declaring it
would not make it loadable; the wheel declares no NVIDIA package for Windows.

## How

- Generator. The CUDA toolkit's Visual Studio integration fails compiler identification
  under the ClangCL toolset on some toolkit and Visual Studio pairs (MSB4023 in its targets
  file), and the MSVC toolset cannot compile the Python extension's sources. A Windows CUDA
  build therefore uses Ninja Multi-Config with clang-cl, the compiler the CPU wheel uses,
  and nvcc with cl.exe as its host compiler, when ninja is on PATH. Without it, as in a
  --no-build-isolation source install, the build keeps the Visual Studio generator it always
  used. CMAKE_CUDA_COMPILER is passed only when CUDACXX is unset, so options a user put in
  CUDACXX (such as -allow-unsupported-compiler) still apply. The toolkit is the one
  install_utils already reports, so what compiles is the train packaging declares.
- CUDA 13.2 and newer. Their CCCL stops with an #error under cl.exe's traditional
  preprocessor, so the CUDA sources pass -Xcompiler=/Zc:preprocessor on Windows (13.0
  accepts it too).
- flatc / flatcc byproducts and imported locations share one host executable suffix chosen
  by CMAKE_HOST_WIN32. CMAKE_EXECUTABLE_SUFFIX is the target's: a Zephyr toolchain sets
  .elf, and naming the byproduct after it left Ninja with no rule for the host tool.
- The shims resolve the runtime from executorch.dll rather than a static core. (The PAL's
  strong symbols are in #23121; C++20 for the delegate and shims comes from #23135.)
- The shims keep WINDOWS_EXPORT_ALL_SYMBOLS OFF in a static build, as upstream sets it,
  and turn it ON in the shared layout: the delegate DLL imports the stream guard functions
  (get/set/peek/clearCurrentCUDAStream) from the shim DLL, and those carry no
  AOTI_SHIM_EXPORT, so the annotations alone would not export them.
- #23121 keeps a Windows CUDA build static; this drops that guard, so the pybind preset
  now builds the shared layout with CUDA too, and adds the CUDA import libraries to the
  Windows packaging list.
- extension/cuda links CUDA::cudart before backends/cuda finds the toolkit, so the root
  finds it first; a CPU torch does not define that target.
- The package config adds the stream helper and the shim layer to the delegate's runtime
  DLL set, and reports EXECUTORCH_RUNTIME_DLLS_EXTRA for the pre-3.28 route, since the
  shim layer ships outside lib/.

## CI

- build-wheels-cuda-windows.yml mirrors the Linux CUDA workflow on the Windows wheel
  builder and reuses filter_cuda_matrix.py. pre_build_script.sh drops the blanket Windows
  CUDA=OFF (a CPU row is still forced off by the generic row classification), resolves
  TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh, and writes it to BUILD_ENV_FILE as well as
  GITHUB_ENV: the reusable workflow sources that file after the hook, and its own list had
  replaced the row's (sm_75 in, sm_89 out). It also installs ninja for the generator.
- End-to-end, per published CUDA train, the route a Windows user has, modeled on
  cuda-windows.yml. The trains are read from filter_cuda_matrix.py's SUPPORTED_CUDA_VERSIONS
  (today cu132 and cu134; main stopped publishing cu130 in #23380), the same list the release
  rows are built from, so adding or
  dropping a train there changes what is tested. (install_utils.py's list also has 12.6 for
  source builds, but 12.6 is no longer a published wheel train.) A Linux GPU job exports a
  Windows-target program with that train's torch (`test_cuda_windows.py --export`), and a
  Windows GPU job builds the wheel with that train's nvcc, installs it and runs
  test_cuda_windows.py with the program (test_cuda_windows_e2e.ps1). Trains the images do not
  carry are assembled from NVIDIA's checksummed redistributable archives by
  install_cuda_redist.py, which needs no installer or admin rights and fails on a train it
  has no release for. Each train's Windows job waits only for its own program, so one train
  failing to export does not skip the others. The wheel build's own smoke test has no GPU
  program and says so.

## Tests

test_cuda_windows.py, the Windows counterpart of test_cuda_linux.py:

- the CUDA DLLs and import libraries ship;
- _C depends on no CUDA DLL and registers no CUDA backend;
- no shipped CUDA DLL imports cudart64_*.dll (only the driver), and no Windows NVIDIA
  requirement is declared;
- every DLL with device code carries exactly the row's GPUs, both directions, with the
  newest also as PTX; the expected list comes from cuda_arch_list.sh keyed by the wheel's
  +cuXYZ tag, not from the TORCH_CUDA_ARCH_LIST the build consumed;
- a C++ application linking executorch::backend_cuda gets the three DLLs copied beside it
  and sees CudaBackend registered;
- a program lowered on Linux for Windows runs through the C++ SDK and matches eager, after
  checking it was exported with the same CUDA train as the wheel (every CUDA 13 train loads
  cudart64_13.dll, so a mismatch would not fail on its own);
- then what a CPU Windows row checks: test_shared_libraries.py (with the three CUDA
  ownership rows), test_cpp_sdk.py and the MobileNetV3/XNNPACK model run. The Arm Cortex-M
  Python module is not built with Ninja, so its install check is left out.

## Test plan

On Windows 11 x64 with an RTX 5080 (sm_120, driver 595.95), VS 2022 BuildTools, CPython 3.12,
and WSL Ubuntu for the export, once per CUDA train. 13.2 and 13.4 are the published trains
and the ones CI covers; 13.0 is no longer published (#23380) but still builds from source, so
it was run too:

| train | Windows nvcc | WSL export torch | wheel | test_cuda_windows.py | Linux-exported program |
|---|---|---|---|---|---|
| 13.0 | 13.0.88 (installer) | 2.14.1+cu130 | executorch-1.6.0+cu130 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 |
| 13.2 | 13.2.86 (redist) | 2.14.1+cu132 | executorch-1.6.0+cu132 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 |
| 13.4 | 13.4.59 (redist) | 2.14.0.dev20260810+cu134 | executorch-1.6.0+cu134 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 |

- Each wheel built with TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh: device code for exactly
  sm_80/86/89/90/100/120, plus sm_120 PTX. 13.4 runs on a driver that reports 13.2, through
  CUDA's minor-version compatibility.
- Before the /Zc:preprocessor change, 13.2 and 13.4 failed to compile the .cu shims with
  CCCL's "traditional preprocessor" #error; 13.0 was unaffected.
- Negative: the cu134 wheel pointed at the 13.0-exported program fails the train check.
- The export used `test_cuda_windows.py --export` from this branch in WSL, with mingw-w64 and
  the Windows CUDA runtime of the same train (13.0 from install_cuda_windows_cross_compile.sh,
  13.2/13.4 from the redist archives).
- The new CI end-to-end jobs could not be run before upload; they follow cuda-windows.yml's
  linux_job_v3 / windows_job pairing and the steps above.
- Linux: the CMake changes are behind WIN32 / MSVC, or pick the same suffix there (empty on a
  Linux host); setup.py's new path runs only on Windows.
Gasoonjia added a commit that referenced this pull request Oct 3, 2026
Stacked on #23121, which ships linkable libraries in the Windows wheel.

A Windows user can run a CUDA program but cannot lower one: lowering compiles the model
with a toolchain only the Linux side has, and the CUDA Windows CI job already exports on
Linux and runs on Windows for exactly this reason. What was missing is a wheel that gives
that Windows user the runtime. This adds a Windows CUDA wheel carrying the CUDA delegate
as C++ libraries, the same components the Linux CUDA wheel ships (#21645, #21668), and
nothing CUDA in the Python module, which on Windows has no CUDA program to lower. The
export (AOT) side is unchanged and identical to the CPU wheel's; only the runtime links CUDA.

    lib/executorch_backend_cuda.dll    + .lib   the delegate          executorch::backend_cuda
    lib/executorch_extension_cuda.dll  + .lib   the stream helper     executorch::extension_cuda
    backends/cuda/aoti_cuda_shims.dll            the AOTI shim layer   (runtime DLL of the delegate)
    data/lib/aoti_cuda_shims.lib                 already shipped as the lowering stub; also the
                                                 shim layer's import library here
    extension/pybindings/_C.pyd                  unchanged: no CUDA dependency, no CudaBackend

A C++ application links it exactly as on Linux:

    find_package(executorch REQUIRED COMPONENTS kernels_optimized backend_cuda)
    target_link_libraries(app PRIVATE executorch::runtime executorch::kernels_optimized
                                      executorch::backend_cuda)
    # plus the $<TARGET_RUNTIME_DLLS:app> copy from the CPU PR, which now also brings the
    # stream helper and the shim layer

The wheel's DLLs import no CUDA runtime DLL: CMake links cudart.lib into them on Windows,
and that reaches the driver (nvcuda.dll). The compiled model is different: the library
AOTInductor builds into model.pte imports cudart64_13.dll, which comes from the CUDA
Toolkit's bin directory on PATH, as the guide says. PyPI's nvidia-cuda-runtime has win_amd64
builds that carry that DLL, but a C++ program does not search site-packages, so declaring it
would not make it loadable; the wheel declares no NVIDIA package for Windows.

## How

- Generator. The CUDA toolkit's Visual Studio integration fails compiler identification
  under the ClangCL toolset on some toolkit and Visual Studio pairs (MSB4023 in its targets
  file), and the MSVC toolset cannot compile the Python extension's sources. A Windows CUDA
  build therefore uses Ninja Multi-Config with clang-cl, the compiler the CPU wheel uses,
  and nvcc with cl.exe as its host compiler, when ninja is on PATH. Without it, as in a
  --no-build-isolation source install, the build keeps the Visual Studio generator it always
  used. CMAKE_CUDA_COMPILER is passed only when CUDACXX is unset, so options a user put in
  CUDACXX (such as -allow-unsupported-compiler) still apply. The toolkit is the one
  install_utils already reports, so what compiles is the train packaging declares.
- CUDA 13.2 and newer. Their CCCL stops with an #error under cl.exe's traditional
  preprocessor, so the CUDA sources pass -Xcompiler=/Zc:preprocessor on Windows (13.0
  accepts it too).
- flatc / flatcc byproducts and imported locations share one host executable suffix chosen
  by CMAKE_HOST_WIN32. CMAKE_EXECUTABLE_SUFFIX is the target's: a Zephyr toolchain sets
  .elf, and naming the byproduct after it left Ninja with no rule for the host tool.
- The shims resolve the runtime from executorch.dll rather than a static core. (The PAL's
  strong symbols are in #23121; C++20 for the delegate and shims comes from #23135.)
- The shims keep WINDOWS_EXPORT_ALL_SYMBOLS OFF in a static build, as upstream sets it,
  and turn it ON in the shared layout: the delegate DLL imports the stream guard functions
  (get/set/peek/clearCurrentCUDAStream) from the shim DLL, and those carry no
  AOTI_SHIM_EXPORT, so the annotations alone would not export them.
- #23121 keeps a Windows CUDA build static; this drops that guard, so the pybind preset
  now builds the shared layout with CUDA too, and adds the CUDA import libraries to the
  Windows packaging list.
- extension/cuda links CUDA::cudart before backends/cuda finds the toolkit, so the root
  finds it first; a CPU torch does not define that target.
- The package config adds the stream helper and the shim layer to the delegate's runtime
  DLL set, and reports EXECUTORCH_RUNTIME_DLLS_EXTRA for the pre-3.28 route, since the
  shim layer ships outside lib/.

## CI

- build-wheels-cuda-windows.yml mirrors the Linux CUDA workflow on the Windows wheel
  builder and reuses filter_cuda_matrix.py. pre_build_script.sh drops the blanket Windows
  CUDA=OFF (a CPU row is still forced off by the generic row classification), resolves
  TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh, and writes it to BUILD_ENV_FILE as well as
  GITHUB_ENV: the reusable workflow sources that file after the hook, and its own list had
  replaced the row's (sm_75 in, sm_89 out). It also installs ninja for the generator.
- End-to-end, per published CUDA train, the route a Windows user has, modeled on
  cuda-windows.yml. The trains are read from filter_cuda_matrix.py's SUPPORTED_CUDA_VERSIONS
  (today cu132 and cu134; main stopped publishing cu130 in #23380), the same list the release
  rows are built from, so adding or
  dropping a train there changes what is tested. (install_utils.py's list also has 12.6 for
  source builds, but 12.6 is no longer a published wheel train.) A Linux GPU job exports a
  Windows-target program with that train's torch (`test_cuda_windows.py --export`), and a
  Windows GPU job builds the wheel with that train's nvcc, installs it and runs
  test_cuda_windows.py with the program (test_cuda_windows_e2e.ps1). Trains the images do not
  carry are assembled from NVIDIA's checksummed redistributable archives by
  install_cuda_redist.py, which needs no installer or admin rights and fails on a train it
  has no release for. Each train's Windows job waits only for its own program, so one train
  failing to export does not skip the others. The wheel build's own smoke test has no GPU
  program and says so.

## Tests

test_cuda_windows.py, the Windows counterpart of test_cuda_linux.py:

- the CUDA DLLs and import libraries ship;
- _C depends on no CUDA DLL and registers no CUDA backend;
- no shipped CUDA DLL imports cudart64_*.dll (only the driver), and no Windows NVIDIA
  requirement is declared;
- every DLL with device code carries exactly the row's GPUs, both directions, with the
  newest also as PTX; the expected list comes from cuda_arch_list.sh keyed by the wheel's
  +cuXYZ tag, not from the TORCH_CUDA_ARCH_LIST the build consumed;
- a C++ application linking executorch::backend_cuda gets the three DLLs copied beside it
  and sees CudaBackend registered;
- a program lowered on Linux for Windows runs through the C++ SDK and matches eager, after
  checking it was exported with the same CUDA train as the wheel (every CUDA 13 train loads
  cudart64_13.dll, so a mismatch would not fail on its own);
- then what a CPU Windows row checks: test_shared_libraries.py (with the three CUDA
  ownership rows), test_cpp_sdk.py and the MobileNetV3/XNNPACK model run. The Arm Cortex-M
  Python module is not built with Ninja, so its install check is left out.

## Test plan

On Windows 11 x64 with an RTX 5080 (sm_120, driver 595.95), VS 2022 BuildTools, CPython 3.12,
and WSL Ubuntu for the export, once per CUDA train. 13.2 and 13.4 are the published trains
and the ones CI covers; 13.0 is no longer published (#23380) but still builds from source, so
it was run too:

| train | Windows nvcc | WSL export torch | wheel | test_cuda_windows.py | Linux-exported program |
|---|---|---|---|---|---|
| 13.0 | 13.0.88 (installer) | 2.14.1+cu130 | executorch-1.6.0+cu130 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 |
| 13.2 | 13.2.86 (redist) | 2.14.1+cu132 | executorch-1.6.0+cu132 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 |
| 13.4 | 13.4.59 (redist) | 2.14.0.dev20260810+cu134 | executorch-1.6.0+cu134 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 |

- Each wheel built with TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh: device code for exactly
  sm_80/86/89/90/100/120, plus sm_120 PTX. 13.4 runs on a driver that reports 13.2, through
  CUDA's minor-version compatibility.
- Before the /Zc:preprocessor change, 13.2 and 13.4 failed to compile the .cu shims with
  CCCL's "traditional preprocessor" #error; 13.0 was unaffected.
- Negative: the cu134 wheel pointed at the 13.0-exported program fails the train check.
- The export used `test_cuda_windows.py --export` from this branch in WSL, with mingw-w64 and
  the Windows CUDA runtime of the same train (13.0 from install_cuda_windows_cross_compile.sh,
  13.2/13.4 from the redist archives).
- The new CI end-to-end jobs could not be run before upload; they follow cuda-windows.yml's
  linux_job_v3 / windows_job pairing and the steps above.
- Linux: the CMake changes are behind WIN32 / MSVC, or pick the same suffix there (empty on a
  Linux host); setup.py's new path runs only on Windows.
Gasoonjia added a commit that referenced this pull request Oct 3, 2026
Stacked on #23121, which ships linkable libraries in the Windows wheel.

A Windows user can run a CUDA program but cannot lower one: lowering compiles the model
with a toolchain only the Linux side has, and the CUDA Windows CI job already exports on
Linux and runs on Windows for exactly this reason. What was missing is a wheel that gives
that Windows user the runtime. This adds a Windows CUDA wheel carrying the CUDA delegate
as C++ libraries, the same components the Linux CUDA wheel ships (#21645, #21668), and
nothing CUDA in the Python module, which on Windows has no CUDA program to lower. The
export (AOT) side is unchanged and identical to the CPU wheel's; only the runtime links CUDA.

    lib/executorch_backend_cuda.dll    + .lib   the delegate          executorch::backend_cuda
    lib/executorch_extension_cuda.dll  + .lib   the stream helper     executorch::extension_cuda
    backends/cuda/aoti_cuda_shims.dll            the AOTI shim layer   (runtime DLL of the delegate)
    data/lib/aoti_cuda_shims.lib                 already shipped as the lowering stub; also the
                                                 shim layer's import library here
    extension/pybindings/_C.pyd                  unchanged: no CUDA dependency, no CudaBackend

A C++ application links it exactly as on Linux:

    find_package(executorch REQUIRED COMPONENTS kernels_optimized backend_cuda)
    target_link_libraries(app PRIVATE executorch::runtime executorch::kernels_optimized
                                      executorch::backend_cuda)
    # plus the $<TARGET_RUNTIME_DLLS:app> copy from the CPU PR, which now also brings the
    # stream helper and the shim layer

The wheel's DLLs import no CUDA runtime DLL: CMake links cudart.lib into them on Windows,
and that reaches the driver (nvcuda.dll). The compiled model is different: the library
AOTInductor builds into model.pte imports cudart64_13.dll, which comes from the CUDA
Toolkit's bin directory on PATH, as the guide says. PyPI's nvidia-cuda-runtime has win_amd64
builds that carry that DLL, but a C++ program does not search site-packages, so declaring it
would not make it loadable; the wheel declares no NVIDIA package for Windows.

## How

- Generator. The CUDA toolkit's Visual Studio integration fails compiler identification
  under the ClangCL toolset on some toolkit and Visual Studio pairs (MSB4023 in its targets
  file), and the MSVC toolset cannot compile the Python extension's sources. A Windows CUDA
  build therefore uses Ninja Multi-Config with clang-cl, the compiler the CPU wheel uses,
  and nvcc with cl.exe as its host compiler, when ninja is on PATH. Without it, as in a
  --no-build-isolation source install, the build keeps the Visual Studio generator it always
  used. CMAKE_CUDA_COMPILER is passed only when CUDACXX is unset, so options a user put in
  CUDACXX (such as -allow-unsupported-compiler) still apply. The toolkit is the one
  install_utils already reports, so what compiles is the train packaging declares.
- CUDA 13.2 and newer. Their CCCL stops with an #error under cl.exe's traditional
  preprocessor, so the CUDA sources pass -Xcompiler=/Zc:preprocessor on Windows (13.0
  accepts it too).
- flatc / flatcc byproducts and imported locations share one host executable suffix chosen
  by CMAKE_HOST_WIN32. CMAKE_EXECUTABLE_SUFFIX is the target's: a Zephyr toolchain sets
  .elf, and naming the byproduct after it left Ninja with no rule for the host tool.
- The shims resolve the runtime from executorch.dll rather than a static core. (The PAL's
  strong symbols are in #23121; C++20 for the delegate and shims comes from #23135.)
- The shims keep WINDOWS_EXPORT_ALL_SYMBOLS OFF in a static build, as upstream sets it,
  and turn it ON in the shared layout: the delegate DLL imports the stream guard functions
  (get/set/peek/clearCurrentCUDAStream) from the shim DLL, and those carry no
  AOTI_SHIM_EXPORT, so the annotations alone would not export them.
- #23121 keeps a Windows CUDA build static; this drops that guard, so the pybind preset
  now builds the shared layout with CUDA too, and adds the CUDA import libraries to the
  Windows packaging list.
- extension/cuda links CUDA::cudart before backends/cuda finds the toolkit, so the root
  finds it first; a CPU torch does not define that target.
- The package config adds the stream helper and the shim layer to the delegate's runtime
  DLL set, and reports EXECUTORCH_RUNTIME_DLLS_EXTRA for the pre-3.28 route, since the
  shim layer ships outside lib/.

## CI

- build-wheels-cuda-windows.yml mirrors the Linux CUDA workflow on the Windows wheel
  builder and reuses filter_cuda_matrix.py. pre_build_script.sh drops the blanket Windows
  CUDA=OFF (a CPU row is still forced off by the generic row classification), resolves
  TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh, and writes it to BUILD_ENV_FILE as well as
  GITHUB_ENV: the reusable workflow sources that file after the hook, and its own list had
  replaced the row's (sm_75 in, sm_89 out). It also installs ninja for the generator.
- End-to-end, per published CUDA train, the route a Windows user has, modeled on
  cuda-windows.yml. The trains are read from filter_cuda_matrix.py's SUPPORTED_CUDA_VERSIONS
  (today cu132 and cu134; main stopped publishing cu130 in #23380), the same list the release
  rows are built from, so adding or
  dropping a train there changes what is tested. (install_utils.py's list also has 12.6 for
  source builds, but 12.6 is no longer a published wheel train.) A Linux GPU job exports a
  Windows-target program with that train's torch (`test_cuda_windows.py --export`), and a
  Windows GPU job builds the wheel with that train's nvcc, installs it and runs
  test_cuda_windows.py with the program (test_cuda_windows_e2e.ps1). Trains the images do not
  carry are assembled from NVIDIA's checksummed redistributable archives by
  install_cuda_redist.py, which needs no installer or admin rights and fails on a train it
  has no release for. Each train's Windows job waits only for its own program, so one train
  failing to export does not skip the others. The wheel build's own smoke test has no GPU
  program and says so.

## Tests

test_cuda_windows.py, the Windows counterpart of test_cuda_linux.py:

- the CUDA DLLs and import libraries ship;
- _C depends on no CUDA DLL and registers no CUDA backend;
- no shipped CUDA DLL imports cudart64_*.dll (only the driver), and no Windows NVIDIA
  requirement is declared;
- every DLL with device code carries exactly the row's GPUs, both directions, with the
  newest also as PTX; the expected list comes from cuda_arch_list.sh keyed by the wheel's
  +cuXYZ tag, not from the TORCH_CUDA_ARCH_LIST the build consumed;
- a C++ application linking executorch::backend_cuda gets the three DLLs copied beside it
  and sees CudaBackend registered;
- a program lowered on Linux for Windows runs through the C++ SDK and matches eager, after
  checking it was exported with the same CUDA train as the wheel (every CUDA 13 train loads
  cudart64_13.dll, so a mismatch would not fail on its own);
- then what a CPU Windows row checks: test_shared_libraries.py (with the three CUDA
  ownership rows), test_cpp_sdk.py and the MobileNetV3/XNNPACK model run. The Arm Cortex-M
  Python module is not built with Ninja, so its install check is left out.

## Test plan

On Windows 11 x64 with an RTX 5080 (sm_120, driver 595.95), VS 2022 BuildTools, CPython 3.12,
and WSL Ubuntu for the export, once per CUDA train. 13.2 and 13.4 are the published trains
and the ones CI covers; 13.0 is no longer published (#23380) but still builds from source, so
it was run too:

| train | Windows nvcc | WSL export torch | wheel | test_cuda_windows.py | Linux-exported program |
|---|---|---|---|---|---|
| 13.0 | 13.0.88 (installer) | 2.14.1+cu130 | executorch-1.6.0+cu130 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 |
| 13.2 | 13.2.86 (redist) | 2.14.1+cu132 | executorch-1.6.0+cu132 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 |
| 13.4 | 13.4.59 (redist) | 2.14.0.dev20260810+cu134 | executorch-1.6.0+cu134 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 |

- Each wheel built with TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh: device code for exactly
  sm_80/86/89/90/100/120, plus sm_120 PTX. 13.4 runs on a driver that reports 13.2, through
  CUDA's minor-version compatibility.
- Before the /Zc:preprocessor change, 13.2 and 13.4 failed to compile the .cu shims with
  CCCL's "traditional preprocessor" #error; 13.0 was unaffected.
- Negative: the cu134 wheel pointed at the 13.0-exported program fails the train check.
- The export used `test_cuda_windows.py --export` from this branch in WSL, with mingw-w64 and
  the Windows CUDA runtime of the same train (13.0 from install_cuda_windows_cross_compile.sh,
  13.2/13.4 from the redist archives).
- The new CI end-to-end jobs could not be run before upload; they follow cuda-windows.yml's
  linux_job_v3 / windows_job pairing and the steps above.
- Linux: the CMake changes are behind WIN32 / MSVC, or pick the same suffix there (empty on a
  Linux host); setup.py's new path runs only on Windows.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. release notes: build Changes related to build, including dependency upgrades, build flags, optimizations, etc.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants