Stop publishing CUDA 13.0 wheels - #23380
Merged
shoumikhin merged 1 commit intoOct 3, 2026
Merged
Conversation
ExecuTorch publishes a wheel for each CUDA version the PyTorch build matrix offers. PyTorch is moving off CUDA 13.0. It now keeps 13.0 on its nightly builds only as a temporary hold for other projects, and its release-candidate builds do not carry it. If ExecuTorch keeps publishing cu130, nightlies carry a train that no release gets, and the train goes away again when PyTorch ends the hold. Nobody notices when it goes, because the filter skips a CUDA version the generator no longer offers and the run stays green. That is how the cu130 wheel stopped without a failure for four days (Sep 29 to Oct 2). This change publishes CUDA 13.2 and 13.4 only. The single row built for a pull request moves from cu130 to cu132, and the install docs drop the cu130 row. A machine on CUDA 13.0 can use the cu132 wheel, because CUDA minor versions are compatible. A machine on an older CUDA can still build from source. What was tested: - The filter unit tests pass (26 of 26). With the original filter, the two tests that pin the published list fail. - Replayed the shared matrix generator for Linux x86_64 and aarch64, on the nightly and test channels and for a pull request. Nightly goes from 15 rows to 10 (cu132 and cu134, five Pythons each). The test channel stays at 10. The pull request row is cu132. - black and flake8 give the same results as on main.
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23380
Note: Links to docs will display an error until the docs builds have been completed. ❌ 1 New FailureAs of commit 51d716e with merge base fbf91c2 ( NEW FAILURE - The following job has failed:
This comment was automatically generated by Dr. CI and updates every 15 minutes. |
huydhn
approved these changes
Oct 2, 2026
huydhn
left a comment
Contributor
There was a problem hiding this comment.
The change LGTM! so stamped, but plz check if there is any communication that needs to be done before merging this one. I'll leave this to the team to decide when to merge. Here is the RFC from PyTorch to remove 13.0 if it helps pytorch/pytorch#190385
Gasoonjia
added a commit
that referenced
this pull request
Oct 3, 2026
Stacked on #23121, which ships linkable libraries in the Windows wheel. A Windows user can run a CUDA program but cannot lower one: lowering compiles the model with a toolchain only the Linux side has, and the CUDA Windows CI job already exports on Linux and runs on Windows for exactly this reason. What was missing is a wheel that gives that Windows user the runtime. This adds a Windows CUDA wheel carrying the CUDA delegate as C++ libraries, the same components the Linux CUDA wheel ships (#21645, #21668), and nothing CUDA in the Python module, which on Windows has no CUDA program to lower. The export (AOT) side is unchanged and identical to the CPU wheel's; only the runtime links CUDA. lib/executorch_backend_cuda.dll + .lib the delegate executorch::backend_cuda lib/executorch_extension_cuda.dll + .lib the stream helper executorch::extension_cuda backends/cuda/aoti_cuda_shims.dll the AOTI shim layer (runtime DLL of the delegate) data/lib/aoti_cuda_shims.lib already shipped as the lowering stub; also the shim layer's import library here extension/pybindings/_C.pyd unchanged: no CUDA dependency, no CudaBackend A C++ application links it exactly as on Linux: find_package(executorch REQUIRED COMPONENTS kernels_optimized backend_cuda) target_link_libraries(app PRIVATE executorch::runtime executorch::kernels_optimized executorch::backend_cuda) # plus the $<TARGET_RUNTIME_DLLS:app> copy from the CPU PR, which now also brings the # stream helper and the shim layer The wheel's DLLs import no CUDA runtime DLL: CMake links cudart.lib into them on Windows, and that reaches the driver (nvcuda.dll). The compiled model is different: the library AOTInductor builds into model.pte imports cudart64_13.dll, which comes from the CUDA Toolkit's bin directory on PATH, as the guide says. PyPI's nvidia-cuda-runtime has win_amd64 builds that carry that DLL, but a C++ program does not search site-packages, so declaring it would not make it loadable; the wheel declares no NVIDIA package for Windows. ## How - Generator. The CUDA toolkit's Visual Studio integration fails compiler identification under the ClangCL toolset on some toolkit and Visual Studio pairs (MSB4023 in its targets file), and the MSVC toolset cannot compile the Python extension's sources. A Windows CUDA build therefore uses Ninja Multi-Config with clang-cl, the compiler the CPU wheel uses, and nvcc with cl.exe as its host compiler, when ninja is on PATH. Without it, as in a --no-build-isolation source install, the build keeps the Visual Studio generator it always used. CMAKE_CUDA_COMPILER is passed only when CUDACXX is unset, so options a user put in CUDACXX (such as -allow-unsupported-compiler) still apply. The toolkit is the one install_utils already reports, so what compiles is the train packaging declares. - CUDA 13.2 and newer. Their CCCL stops with an #error under cl.exe's traditional preprocessor, so the CUDA sources pass -Xcompiler=/Zc:preprocessor on Windows (13.0 accepts it too). - flatc / flatcc byproducts and imported locations share one host executable suffix chosen by CMAKE_HOST_WIN32. CMAKE_EXECUTABLE_SUFFIX is the target's: a Zephyr toolchain sets .elf, and naming the byproduct after it left Ninja with no rule for the host tool. - The shims resolve the runtime from executorch.dll rather than a static core. (The PAL's strong symbols are in #23121; C++20 for the delegate and shims comes from #23135.) - The shims keep WINDOWS_EXPORT_ALL_SYMBOLS OFF in a static build, as upstream sets it, and turn it ON in the shared layout: the delegate DLL imports the stream guard functions (get/set/peek/clearCurrentCUDAStream) from the shim DLL, and those carry no AOTI_SHIM_EXPORT, so the annotations alone would not export them. - #23121 keeps a Windows CUDA build static; this drops that guard, so the pybind preset now builds the shared layout with CUDA too, and adds the CUDA import libraries to the Windows packaging list. - extension/cuda links CUDA::cudart before backends/cuda finds the toolkit, so the root finds it first; a CPU torch does not define that target. - The package config adds the stream helper and the shim layer to the delegate's runtime DLL set, and reports EXECUTORCH_RUNTIME_DLLS_EXTRA for the pre-3.28 route, since the shim layer ships outside lib/. ## CI - build-wheels-cuda-windows.yml mirrors the Linux CUDA workflow on the Windows wheel builder and reuses filter_cuda_matrix.py. pre_build_script.sh drops the blanket Windows CUDA=OFF (a CPU row is still forced off by the generic row classification), resolves TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh, and writes it to BUILD_ENV_FILE as well as GITHUB_ENV: the reusable workflow sources that file after the hook, and its own list had replaced the row's (sm_75 in, sm_89 out). It also installs ninja for the generator. - End-to-end, per published CUDA train, the route a Windows user has, modeled on cuda-windows.yml. The trains are read from filter_cuda_matrix.py's SUPPORTED_CUDA_VERSIONS (today cu132 and cu134; main stopped publishing cu130 in #23380), the same list the release rows are built from, so adding or dropping a train there changes what is tested. (install_utils.py's list also has 12.6 for source builds, but 12.6 is no longer a published wheel train.) A Linux GPU job exports a Windows-target program with that train's torch (`test_cuda_windows.py --export`), and a Windows GPU job builds the wheel with that train's nvcc, installs it and runs test_cuda_windows.py with the program (test_cuda_windows_e2e.ps1). Trains the images do not carry are assembled from NVIDIA's checksummed redistributable archives by install_cuda_redist.py, which needs no installer or admin rights and fails on a train it has no release for. Each train's Windows job waits only for its own program, so one train failing to export does not skip the others. The wheel build's own smoke test has no GPU program and says so. ## Tests test_cuda_windows.py, the Windows counterpart of test_cuda_linux.py: - the CUDA DLLs and import libraries ship; - _C depends on no CUDA DLL and registers no CUDA backend; - no shipped CUDA DLL imports cudart64_*.dll (only the driver), and no Windows NVIDIA requirement is declared; - every DLL with device code carries exactly the row's GPUs, both directions, with the newest also as PTX; the expected list comes from cuda_arch_list.sh keyed by the wheel's +cuXYZ tag, not from the TORCH_CUDA_ARCH_LIST the build consumed; - a C++ application linking executorch::backend_cuda gets the three DLLs copied beside it and sees CudaBackend registered; - a program lowered on Linux for Windows runs through the C++ SDK and matches eager, after checking it was exported with the same CUDA train as the wheel (every CUDA 13 train loads cudart64_13.dll, so a mismatch would not fail on its own); - then what a CPU Windows row checks: test_shared_libraries.py (with the three CUDA ownership rows), test_cpp_sdk.py and the MobileNetV3/XNNPACK model run. The Arm Cortex-M Python module is not built with Ninja, so its install check is left out. ## Test plan On Windows 11 x64 with an RTX 5080 (sm_120, driver 595.95), VS 2022 BuildTools, CPython 3.12, and WSL Ubuntu for the export, once per CUDA train. 13.2 and 13.4 are the published trains and the ones CI covers; 13.0 is no longer published (#23380) but still builds from source, so it was run too: | train | Windows nvcc | WSL export torch | wheel | test_cuda_windows.py | Linux-exported program | |---|---|---|---|---|---| | 13.0 | 13.0.88 (installer) | 2.14.1+cu130 | executorch-1.6.0+cu130 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 | | 13.2 | 13.2.86 (redist) | 2.14.1+cu132 | executorch-1.6.0+cu132 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 | | 13.4 | 13.4.59 (redist) | 2.14.0.dev20260810+cu134 | executorch-1.6.0+cu134 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 | - Each wheel built with TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh: device code for exactly sm_80/86/89/90/100/120, plus sm_120 PTX. 13.4 runs on a driver that reports 13.2, through CUDA's minor-version compatibility. - Before the /Zc:preprocessor change, 13.2 and 13.4 failed to compile the .cu shims with CCCL's "traditional preprocessor" #error; 13.0 was unaffected. - Negative: the cu134 wheel pointed at the 13.0-exported program fails the train check. - The export used `test_cuda_windows.py --export` from this branch in WSL, with mingw-w64 and the Windows CUDA runtime of the same train (13.0 from install_cuda_windows_cross_compile.sh, 13.2/13.4 from the redist archives). - The new CI end-to-end jobs could not be run before upload; they follow cuda-windows.yml's linux_job_v3 / windows_job pairing and the steps above. - Linux: the CMake changes are behind WIN32 / MSVC, or pick the same suffix there (empty on a Linux host); setup.py's new path runs only on Windows.
Gasoonjia
added a commit
that referenced
this pull request
Oct 3, 2026
Stacked on #23121, which ships linkable libraries in the Windows wheel. A Windows user can run a CUDA program but cannot lower one: lowering compiles the model with a toolchain only the Linux side has, and the CUDA Windows CI job already exports on Linux and runs on Windows for exactly this reason. What was missing is a wheel that gives that Windows user the runtime. This adds a Windows CUDA wheel carrying the CUDA delegate as C++ libraries, the same components the Linux CUDA wheel ships (#21645, #21668), and nothing CUDA in the Python module, which on Windows has no CUDA program to lower. The export (AOT) side is unchanged and identical to the CPU wheel's; only the runtime links CUDA. lib/executorch_backend_cuda.dll + .lib the delegate executorch::backend_cuda lib/executorch_extension_cuda.dll + .lib the stream helper executorch::extension_cuda backends/cuda/aoti_cuda_shims.dll the AOTI shim layer (runtime DLL of the delegate) data/lib/aoti_cuda_shims.lib already shipped as the lowering stub; also the shim layer's import library here extension/pybindings/_C.pyd unchanged: no CUDA dependency, no CudaBackend A C++ application links it exactly as on Linux: find_package(executorch REQUIRED COMPONENTS kernels_optimized backend_cuda) target_link_libraries(app PRIVATE executorch::runtime executorch::kernels_optimized executorch::backend_cuda) # plus the $<TARGET_RUNTIME_DLLS:app> copy from the CPU PR, which now also brings the # stream helper and the shim layer The wheel's DLLs import no CUDA runtime DLL: CMake links cudart.lib into them on Windows, and that reaches the driver (nvcuda.dll). The compiled model is different: the library AOTInductor builds into model.pte imports cudart64_13.dll, which comes from the CUDA Toolkit's bin directory on PATH, as the guide says. PyPI's nvidia-cuda-runtime has win_amd64 builds that carry that DLL, but a C++ program does not search site-packages, so declaring it would not make it loadable; the wheel declares no NVIDIA package for Windows. ## How - Generator. The CUDA toolkit's Visual Studio integration fails compiler identification under the ClangCL toolset on some toolkit and Visual Studio pairs (MSB4023 in its targets file), and the MSVC toolset cannot compile the Python extension's sources. A Windows CUDA build therefore uses Ninja Multi-Config with clang-cl, the compiler the CPU wheel uses, and nvcc with cl.exe as its host compiler, when ninja is on PATH. Without it, as in a --no-build-isolation source install, the build keeps the Visual Studio generator it always used. CMAKE_CUDA_COMPILER is passed only when CUDACXX is unset, so options a user put in CUDACXX (such as -allow-unsupported-compiler) still apply. The toolkit is the one install_utils already reports, so what compiles is the train packaging declares. - CUDA 13.2 and newer. Their CCCL stops with an #error under cl.exe's traditional preprocessor, so the CUDA sources pass -Xcompiler=/Zc:preprocessor on Windows (13.0 accepts it too). - flatc / flatcc byproducts and imported locations share one host executable suffix chosen by CMAKE_HOST_WIN32. CMAKE_EXECUTABLE_SUFFIX is the target's: a Zephyr toolchain sets .elf, and naming the byproduct after it left Ninja with no rule for the host tool. - The shims resolve the runtime from executorch.dll rather than a static core. (The PAL's strong symbols are in #23121; C++20 for the delegate and shims comes from #23135.) - The shims keep WINDOWS_EXPORT_ALL_SYMBOLS OFF in a static build, as upstream sets it, and turn it ON in the shared layout: the delegate DLL imports the stream guard functions (get/set/peek/clearCurrentCUDAStream) from the shim DLL, and those carry no AOTI_SHIM_EXPORT, so the annotations alone would not export them. - #23121 keeps a Windows CUDA build static; this drops that guard, so the pybind preset now builds the shared layout with CUDA too, and adds the CUDA import libraries to the Windows packaging list. - extension/cuda links CUDA::cudart before backends/cuda finds the toolkit, so the root finds it first; a CPU torch does not define that target. - The package config adds the stream helper and the shim layer to the delegate's runtime DLL set, and reports EXECUTORCH_RUNTIME_DLLS_EXTRA for the pre-3.28 route, since the shim layer ships outside lib/. ## CI - build-wheels-cuda-windows.yml mirrors the Linux CUDA workflow on the Windows wheel builder and reuses filter_cuda_matrix.py. pre_build_script.sh drops the blanket Windows CUDA=OFF (a CPU row is still forced off by the generic row classification), resolves TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh, and writes it to BUILD_ENV_FILE as well as GITHUB_ENV: the reusable workflow sources that file after the hook, and its own list had replaced the row's (sm_75 in, sm_89 out). It also installs ninja for the generator. - End-to-end, per published CUDA train, the route a Windows user has, modeled on cuda-windows.yml. The trains are read from filter_cuda_matrix.py's SUPPORTED_CUDA_VERSIONS (today cu132 and cu134; main stopped publishing cu130 in #23380), the same list the release rows are built from, so adding or dropping a train there changes what is tested. (install_utils.py's list also has 12.6 for source builds, but 12.6 is no longer a published wheel train.) A Linux GPU job exports a Windows-target program with that train's torch (`test_cuda_windows.py --export`), and a Windows GPU job builds the wheel with that train's nvcc, installs it and runs test_cuda_windows.py with the program (test_cuda_windows_e2e.ps1). Trains the images do not carry are assembled from NVIDIA's checksummed redistributable archives by install_cuda_redist.py, which needs no installer or admin rights and fails on a train it has no release for. Each train's Windows job waits only for its own program, so one train failing to export does not skip the others. The wheel build's own smoke test has no GPU program and says so. ## Tests test_cuda_windows.py, the Windows counterpart of test_cuda_linux.py: - the CUDA DLLs and import libraries ship; - _C depends on no CUDA DLL and registers no CUDA backend; - no shipped CUDA DLL imports cudart64_*.dll (only the driver), and no Windows NVIDIA requirement is declared; - every DLL with device code carries exactly the row's GPUs, both directions, with the newest also as PTX; the expected list comes from cuda_arch_list.sh keyed by the wheel's +cuXYZ tag, not from the TORCH_CUDA_ARCH_LIST the build consumed; - a C++ application linking executorch::backend_cuda gets the three DLLs copied beside it and sees CudaBackend registered; - a program lowered on Linux for Windows runs through the C++ SDK and matches eager, after checking it was exported with the same CUDA train as the wheel (every CUDA 13 train loads cudart64_13.dll, so a mismatch would not fail on its own); - then what a CPU Windows row checks: test_shared_libraries.py (with the three CUDA ownership rows), test_cpp_sdk.py and the MobileNetV3/XNNPACK model run. The Arm Cortex-M Python module is not built with Ninja, so its install check is left out. ## Test plan On Windows 11 x64 with an RTX 5080 (sm_120, driver 595.95), VS 2022 BuildTools, CPython 3.12, and WSL Ubuntu for the export, once per CUDA train. 13.2 and 13.4 are the published trains and the ones CI covers; 13.0 is no longer published (#23380) but still builds from source, so it was run too: | train | Windows nvcc | WSL export torch | wheel | test_cuda_windows.py | Linux-exported program | |---|---|---|---|---|---| | 13.0 | 13.0.88 (installer) | 2.14.1+cu130 | executorch-1.6.0+cu130 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 | | 13.2 | 13.2.86 (redist) | 2.14.1+cu132 | executorch-1.6.0+cu132 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 | | 13.4 | 13.4.59 (redist) | 2.14.0.dev20260810+cu134 | executorch-1.6.0+cu134 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 | - Each wheel built with TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh: device code for exactly sm_80/86/89/90/100/120, plus sm_120 PTX. 13.4 runs on a driver that reports 13.2, through CUDA's minor-version compatibility. - Before the /Zc:preprocessor change, 13.2 and 13.4 failed to compile the .cu shims with CCCL's "traditional preprocessor" #error; 13.0 was unaffected. - Negative: the cu134 wheel pointed at the 13.0-exported program fails the train check. - The export used `test_cuda_windows.py --export` from this branch in WSL, with mingw-w64 and the Windows CUDA runtime of the same train (13.0 from install_cuda_windows_cross_compile.sh, 13.2/13.4 from the redist archives). - The new CI end-to-end jobs could not be run before upload; they follow cuda-windows.yml's linux_job_v3 / windows_job pairing and the steps above. - Linux: the CMake changes are behind WIN32 / MSVC, or pick the same suffix there (empty on a Linux host); setup.py's new path runs only on Windows.
Gasoonjia
added a commit
that referenced
this pull request
Oct 3, 2026
Stacked on #23121, which ships linkable libraries in the Windows wheel. A Windows user can run a CUDA program but cannot lower one: lowering compiles the model with a toolchain only the Linux side has, and the CUDA Windows CI job already exports on Linux and runs on Windows for exactly this reason. What was missing is a wheel that gives that Windows user the runtime. This adds a Windows CUDA wheel carrying the CUDA delegate as C++ libraries, the same components the Linux CUDA wheel ships (#21645, #21668), and nothing CUDA in the Python module, which on Windows has no CUDA program to lower. The export (AOT) side is unchanged and identical to the CPU wheel's; only the runtime links CUDA. lib/executorch_backend_cuda.dll + .lib the delegate executorch::backend_cuda lib/executorch_extension_cuda.dll + .lib the stream helper executorch::extension_cuda backends/cuda/aoti_cuda_shims.dll the AOTI shim layer (runtime DLL of the delegate) data/lib/aoti_cuda_shims.lib already shipped as the lowering stub; also the shim layer's import library here extension/pybindings/_C.pyd unchanged: no CUDA dependency, no CudaBackend A C++ application links it exactly as on Linux: find_package(executorch REQUIRED COMPONENTS kernels_optimized backend_cuda) target_link_libraries(app PRIVATE executorch::runtime executorch::kernels_optimized executorch::backend_cuda) # plus the $<TARGET_RUNTIME_DLLS:app> copy from the CPU PR, which now also brings the # stream helper and the shim layer The wheel's DLLs import no CUDA runtime DLL: CMake links cudart.lib into them on Windows, and that reaches the driver (nvcuda.dll). The compiled model is different: the library AOTInductor builds into model.pte imports cudart64_13.dll, which comes from the CUDA Toolkit's bin directory on PATH, as the guide says. PyPI's nvidia-cuda-runtime has win_amd64 builds that carry that DLL, but a C++ program does not search site-packages, so declaring it would not make it loadable; the wheel declares no NVIDIA package for Windows. ## How - Generator. The CUDA toolkit's Visual Studio integration fails compiler identification under the ClangCL toolset on some toolkit and Visual Studio pairs (MSB4023 in its targets file), and the MSVC toolset cannot compile the Python extension's sources. A Windows CUDA build therefore uses Ninja Multi-Config with clang-cl, the compiler the CPU wheel uses, and nvcc with cl.exe as its host compiler, when ninja is on PATH. Without it, as in a --no-build-isolation source install, the build keeps the Visual Studio generator it always used. CMAKE_CUDA_COMPILER is passed only when CUDACXX is unset, so options a user put in CUDACXX (such as -allow-unsupported-compiler) still apply. The toolkit is the one install_utils already reports, so what compiles is the train packaging declares. - CUDA 13.2 and newer. Their CCCL stops with an #error under cl.exe's traditional preprocessor, so the CUDA sources pass -Xcompiler=/Zc:preprocessor on Windows (13.0 accepts it too). - flatc / flatcc byproducts and imported locations share one host executable suffix chosen by CMAKE_HOST_WIN32. CMAKE_EXECUTABLE_SUFFIX is the target's: a Zephyr toolchain sets .elf, and naming the byproduct after it left Ninja with no rule for the host tool. - The shims resolve the runtime from executorch.dll rather than a static core. (The PAL's strong symbols are in #23121; C++20 for the delegate and shims comes from #23135.) - The shims keep WINDOWS_EXPORT_ALL_SYMBOLS OFF in a static build, as upstream sets it, and turn it ON in the shared layout: the delegate DLL imports the stream guard functions (get/set/peek/clearCurrentCUDAStream) from the shim DLL, and those carry no AOTI_SHIM_EXPORT, so the annotations alone would not export them. - #23121 keeps a Windows CUDA build static; this drops that guard, so the pybind preset now builds the shared layout with CUDA too, and adds the CUDA import libraries to the Windows packaging list. - extension/cuda links CUDA::cudart before backends/cuda finds the toolkit, so the root finds it first; a CPU torch does not define that target. - The package config adds the stream helper and the shim layer to the delegate's runtime DLL set, and reports EXECUTORCH_RUNTIME_DLLS_EXTRA for the pre-3.28 route, since the shim layer ships outside lib/. ## CI - build-wheels-cuda-windows.yml mirrors the Linux CUDA workflow on the Windows wheel builder and reuses filter_cuda_matrix.py. pre_build_script.sh drops the blanket Windows CUDA=OFF (a CPU row is still forced off by the generic row classification), resolves TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh, and writes it to BUILD_ENV_FILE as well as GITHUB_ENV: the reusable workflow sources that file after the hook, and its own list had replaced the row's (sm_75 in, sm_89 out). It also installs ninja for the generator. - End-to-end, per published CUDA train, the route a Windows user has, modeled on cuda-windows.yml. The trains are read from filter_cuda_matrix.py's SUPPORTED_CUDA_VERSIONS (today cu132 and cu134; main stopped publishing cu130 in #23380), the same list the release rows are built from, so adding or dropping a train there changes what is tested. (install_utils.py's list also has 12.6 for source builds, but 12.6 is no longer a published wheel train.) A Linux GPU job exports a Windows-target program with that train's torch (`test_cuda_windows.py --export`), and a Windows GPU job builds the wheel with that train's nvcc, installs it and runs test_cuda_windows.py with the program (test_cuda_windows_e2e.ps1). Trains the images do not carry are assembled from NVIDIA's checksummed redistributable archives by install_cuda_redist.py, which needs no installer or admin rights and fails on a train it has no release for. Each train's Windows job waits only for its own program, so one train failing to export does not skip the others. The wheel build's own smoke test has no GPU program and says so. ## Tests test_cuda_windows.py, the Windows counterpart of test_cuda_linux.py: - the CUDA DLLs and import libraries ship; - _C depends on no CUDA DLL and registers no CUDA backend; - no shipped CUDA DLL imports cudart64_*.dll (only the driver), and no Windows NVIDIA requirement is declared; - every DLL with device code carries exactly the row's GPUs, both directions, with the newest also as PTX; the expected list comes from cuda_arch_list.sh keyed by the wheel's +cuXYZ tag, not from the TORCH_CUDA_ARCH_LIST the build consumed; - a C++ application linking executorch::backend_cuda gets the three DLLs copied beside it and sees CudaBackend registered; - a program lowered on Linux for Windows runs through the C++ SDK and matches eager, after checking it was exported with the same CUDA train as the wheel (every CUDA 13 train loads cudart64_13.dll, so a mismatch would not fail on its own); - then what a CPU Windows row checks: test_shared_libraries.py (with the three CUDA ownership rows), test_cpp_sdk.py and the MobileNetV3/XNNPACK model run. The Arm Cortex-M Python module is not built with Ninja, so its install check is left out. ## Test plan On Windows 11 x64 with an RTX 5080 (sm_120, driver 595.95), VS 2022 BuildTools, CPython 3.12, and WSL Ubuntu for the export, once per CUDA train. 13.2 and 13.4 are the published trains and the ones CI covers; 13.0 is no longer published (#23380) but still builds from source, so it was run too: | train | Windows nvcc | WSL export torch | wheel | test_cuda_windows.py | Linux-exported program | |---|---|---|---|---|---| | 13.0 | 13.0.88 (installer) | 2.14.1+cu130 | executorch-1.6.0+cu130 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 | | 13.2 | 13.2.86 (redist) | 2.14.1+cu132 | executorch-1.6.0+cu132 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 | | 13.4 | 13.4.59 (redist) | 2.14.0.dev20260810+cu134 | executorch-1.6.0+cu134 | 52 checks pass | runs on the GPU, maxdiff 2.4e-7 | - Each wheel built with TORCH_CUDA_ARCH_LIST from cuda_arch_list.sh: device code for exactly sm_80/86/89/90/100/120, plus sm_120 PTX. 13.4 runs on a driver that reports 13.2, through CUDA's minor-version compatibility. - Before the /Zc:preprocessor change, 13.2 and 13.4 failed to compile the .cu shims with CCCL's "traditional preprocessor" #error; 13.0 was unaffected. - Negative: the cu134 wheel pointed at the 13.0-exported program fails the train check. - The export used `test_cuda_windows.py --export` from this branch in WSL, with mingw-w64 and the Windows CUDA runtime of the same train (13.0 from install_cuda_windows_cross_compile.sh, 13.2/13.4 from the redist archives). - The new CI end-to-end jobs could not be run before upload; they follow cuda-windows.yml's linux_job_v3 / windows_job pairing and the steps above. - Linux: the CMake changes are behind WIN32 / MSVC, or pick the same suffix there (empty on a Linux host); setup.py's new path runs only on Windows.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
ExecuTorch publishes a CUDA wheel for each CUDA version that PyTorch's build matrix offers. PyTorch is moving off CUDA 13.0. It now keeps 13.0 on its nightly builds only as a temporary hold for other projects, and its release-candidate builds do not carry it (see pytorch/test-infra#8989 and pytorch/pytorch#199520).
If ExecuTorch keeps publishing
cu130, nightlies carry a wheel that no release gets, and that wheel goes away again when the hold ends. It is easy to miss when it goes. The wheel filter skips a CUDA version that the matrix no longer offers, and the run stays green. That is what happened from Sep 29 to Oct 2, when PyTorch briefly dropped 13.0 and nocu130wheel was built.This change publishes CUDA 13.2 and 13.4 only:
filter_cuda_matrix.py: removecu130from the published list. The single row built for a pull request moves fromcu130tocu132.test_filter_cuda_matrix.py: update the pinned published list.cu130row from the install table and usecu132in the examples. A machine on CUDA 13.0 can use thecu132wheel, because CUDA minor versions are compatible.Not changed:
install_utils.py(it still accepts a CUDA 13.0 toolkit for source builds), the GPU architecture table, and the CI jobs that run on a CUDA 13.0 image. None of those decide which wheels get published.Test plan
python .ci/scripts/tests/test_filter_cuda_matrix.py: 26 tests pass. With the original filter, the two tests that pin the published list fail.cu132andcu134, five Python versions each). The test channel stays at 10. The pull request row iscu132.blackandflake8give the same results as on main.